EDBT 2026 Demo / reviewers in the wild / expert
Vishnu Monn Baskaran
dblp:123/3772
· DBLP profile ↗
33ranked-venue papers
4as first author
29since 2021 · last 2026
0000-0001-6809-5817ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 first-author · 11 since 2021Systems, architecture and hardware · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Deep Probabilistic Flow-Based Framework for Unsupervised Cross-Domain Soft SensingabstractIndustrial soft sensing is crucial for accurate process monitoring through reliable inference of dominant sensor variables. However, developing effective data-driven soft sensor models presents challenges, such as achieving domain adaptability, addressing incomplete sensor labels, and learning stochastic data variability. To overcome these challenges, we propose a deep variational potential flow (DVPF) framework for cross-domain soft sensor modeling, taking into account the lack of sensor labels in the target domain. Our framework introduces sequential variational Bayes with recurrent neural network (RNN) parameterization to address the maximum likelihood estimation problem that characterizes cross-domain soft sensing. Central to the framework is a potential flow that performs unsupervised Bayesian inference on the RNN-extracted features to obtain an exact representation of the intractable posterior distribution. Together, these DVPF components learn domain-adaptable features that effectively capture complex cross-domain process dynamics and data variability. We validate the proposed DVPF on a real industrial multiphase flow process across varying operating modes. The results show that the DVPF demonstrates superior performance in cross-domain soft sensing compared to existing deep feature-based domain adaptation methods. Junn Yong Loo, Hwa Hui Tew, Fang Yu Leong, Ze Yang Ding, Vishnu Monn Baskaran, Chee-Ming Ting, Chee Pin Tan |
IEEE Trans. Ind. Informatics | 5 |
| 2025 | Conceptualizing Multi-scale Wavelet Attention and Ray-based Encoding for Human-Object Interaction DetectionabstractHuman-object interaction (HOI) detection is essential for accurately localizing and characterizing interactions between humans and objects, providing a comprehensive understanding of complex visual scenes across various domains. However, existing HOI detectors often struggle to deliver reliable predictions efficiently, relying on resource-intensive training methods and inefficient architectures. To address these challenges, we conceptualize a wavelet attention-like backbone and a novel ray-based encoder architecture tailored for HOI detection. Our wavelet backbone addresses the limitations of expressing middle-order interactions by aggregating discriminative features from the low- and high-order interactions extracted from diverse convolutional filters. Concurrently, the ray-based encoder facilitates multi-scale attention by optimizing the focus of the decoder on relevant regions of interest and mitigating computational overhead. As a result of harnessing the attenuated intensity of learnable ray origins, our decoder aligns query embeddings with emphasized regions of interest for accurate predictions. Experimental results on benchmark datasets, including ImageNet and HICO-DET, showcase the potential of our proposed architecture. The code is publicly available at [https://github.com/henrypay/RayEncoder]. Quan Bi Pay, Vishnu Monn Baskaran, Junn Yong Loo, Koksheik Wong, Simon See |
IJCNN | 2 |
| 2025 | SpaRTAN: Spatial Reinforcement Token-based Aggregation Network for Visual RecognitionabstractThe resurgence of convolutional neural networks (CNNs) in visual recognition tasks, exemplified by ConvNeXt, has demonstrated their capability to rival transformer-based architectures through advanced training methodologies and ViTinspired design principles. However, both CNNs and transformers exhibit a simplicity bias, favoring straightforward features over complex structural representations. Furthermore, modern CNNs often integrate MLP-like blocks akin to those in transformers, but these blocks suffer from significant information redundancies, necessitating high expansion ratios to sustain competitive performance. To address these limitations, we propose SpaRTAN, a lightweight architectural design that enhances spatial and channel-wise information processing. SpaRTAN employs kernels with varying receptive fields, controlled by kernel size and dilation factor, to capture discriminative multi-order spatial features effectively. A wave-based channel aggregation module further modulates and reinforces pixel interactions, mitigating channel-wise redundancies. Combining the two modules, the proposed network can efficiently gather and dynamically contextualize discriminative features. Experimental results in ImageNet and COCO demonstrate that SpaRTAN achieves remarkable parameter efficiency while maintaining competitive performance. In particular, on the ImageNet-1k benchmark, SpaRTAN achieves 77. 7% accuracy with only 3.8M parameters and approximately 1.0 GFLOPs, demonstrating its ability to deliver strong performance through an efficient design. On the COCO benchmark, it achieves 50.0% AP, surpassing the previous benchmark by 1.2% with only 21.5M parameters. The code is publicly available at [https://github.com/henry-pay/SpaRTAN]. Quan Bi Pay, Vishnu Monn Baskaran, Junn Yong Loo, Koksheik Wong, Simon See |
IJCNN | 2 |
| 2025 | Contrastive Denoising Variational Recurrent Neural Network for Noise-Agnostic State Modeling of Soft Robotic SystemsabstractSoft robotic systems are highly susceptible to sensor noise arising from hardware imperfections, environmental disturbances, and the intrinsic compliance of soft materials. These noisy measurements can obscure essential state information and degrade performance in both perception and control tasks. In this paper, we introduce a Contrastive Denoising Variational Recurrent Neural Network (CD-VRNN) designed to address this challenge. The proposed model learns to decompose time-series data into clean signal and noise pathways, improving state estimation accuracy for soft robotic platforms. By incorporating a contrastive objective that enforces a separation between signal-related and noise-related latent representations, the proposed CD-VRNN more effectively isolates noise while retaining critical temporal features. In addition, we incorporate conditional flow-based priors for expressive, state-dependent distributions and skip-connected decoders that preserve subtle signal variations. Experimental results on a pneumatic soft robot and multiple public time-series datasets show that CD-VRNN consistently outperforms existing approaches for denoising and downstream task performance, demonstrating the model’s robustness and strong generalization capacity. Shageenderan Sapai, Vishnu Monn Baskaran, Junn Yong Loo, Surya Girinatha Nurzaman, Chee Pin Tan |
IJCNN | 2 |
| 2025 | Joint Optimisation of Electric Vehicle Routing and Scheduling: A Deep Learning-Driven Approach for Dynamic Fleet SizesabstractElectric Vehicles (EVs) are becoming increasingly prevalent nowadays, with studies highlighting their potential as mobile energy storage systems to provide grid support. Realising this potential requires effective charging coordination, which are often formulated as mixed-integer programming (MIP) problems. However, MIP problems are NP-hard and often intractable when applied to time-sensitive tasks. To address this limitation, we propose a deep learning assisted approach for optimising a day-ahead EV joint routing and scheduling problem with varying number of EVs. This problem simultaneously optimises EV routing, charging, discharging and generator scheduling within a distribution network with renewable energy sources. A convolutional neural network is trained to predict the binary variables, thereby reducing the solution search space and enabling solvers to determine the remaining variables more efficiently. Additionally, a padding mechanism is included to handle the changes in input and output sizes caused by varying number of EVs, thus eliminating the need for re-training. In a case study on the IEEE 33-bus system and Nguyen-Dupius transportation network, our approach reduced runtime by 97.8% when compared to an unassisted MIP solver, while retaining 99.5% feasibility and deviating less than 0.01% from the optimal solution. Jun Kang Yap, Vishnu Monn Baskaran, Wen-Shan Tan, Ze Yang Ding, Hao Wang 0016, David L. Dowe |
IJCNN | 2 |
| 2025 | Contrastive Autoencoder for Robust State Modelling of Soft Robots in Incomplete and Noisy EnvironmentsabstractSoft robotic systems heavily depend on accurate sensor data for perception and control; however, this data is often corrupted by missing observations, due to partial sensor coverage, communication failures, or occlusions and noisy measurements stemming from hardware imperfections, environmental disturbances, and the intrinsic compliance of soft materials. Such corruption can obscure critical state information, causing unreliable modeling of soft robotics and degrading control accuracy. To address these challenges, we propose a Contrastive Dual-Latent Autoencoder (CDLAE) that jointly handles missing and noisy data in a single end-to-end framework. Our approach leverages an attention based autoencoder architecture with dual latent pathways, where one focuses on capturing the underlying clean signals while the other isolates noise-related components. A contrastive loss encourages strong separation between these pathways, enhancing the model’s ability to filter noise while reconstructing missing values. Additionally, the autoencoder is trained jointly with a downstream predictive network, ensuring that signal imputation is optimized with respect to the ultimate control task. Experimental evaluations on a pneumatic soft robot platform and multiple public time-series datasets demonstrate that CDLAE consistently outperforms existing methods in handling corrupted data, offering robust, high-fidelity reconstructions that significantly improve soft robot perception and control in real-world conditions. Shageenderan Sapai, Vishnu Monn Baskaran, Junn Yong Loo, Surya Girinatha Nurzaman, Chee Pin Tan |
IROS | 2 |
| 2025 | @LM DeceptionNet: A multimodal approach for efficient transfer learning-based deception detectionabstractIn terms of deception detection, traditional contact-based techniques often require collecting physiological signals, which can negatively impact device accuracy and participant comfort. While multimodal features extracted from audio and video modalities have been shown to outperform human observers on public datasets, the generalizability of existing audio and visual-based deception detection methods in different scenarios remains insufficiently explored. To narrow this gap, this work proposes a novel domain knowledge transfer learning method for deception detection in cross-scenario applications, which enhances its generalization and adaptability. Additionally, we designed a multimodal framework that filters out irrelevant information from other modalities when a particular modality yields reliable results, further improving overall system accuracy and robustness. We evaluate the proposed method on different public datasets, achieving promising generalizability results with consistent enhancements using four variations and networks. Apart from this, the proposed @LM DeceptionNet demonstrates better generalization capacity in computational efficiency, feature extraction, and adaptability compared to a larger model when employing fewer parameters. Yuanya Zhuo, Vishnu Monn Baskaran, Lillian Yee Kiaw Wang, Raphael C.-W. Phan |
Knowl. Based Syst. | 2 |
| 2025 | Trade-off independent image watermarking using enhanced structured matrix decompositionabstractAbstract Image watermarking plays a vital role in providing protection from copyright violation. However, conventional watermarking techniques typically exhibit trade-offs in terms of image quality, robustness and capacity constrains. More often than not, these techniques optimize on one constrain while settling with the two other constraints. Therefore, in this paper, an enhanced saliency detection based watermarking method is proposed to simultaneously improve quality, capacity, and robustness. First, the enhanced structured matrix decomposition (E-SMD) is proposed to extract salient regions in the host image for producing a saliency mask. This mask is then applied to partition the foreground and background of the host and watermark images. Subsequently, the watermark (with the same dimension of host image) is shuffled using multiple Arnold and Logistic chaotic maps, and the resulting shuffled-watermark is embedded into the wavelet domain of the host image. Furthermore, a filtering operation is put forward to estimate the original host image so that the proposed watermarking method can also operate in blind mode. In the best case scenario, we could embed a 24-bit image as the watermark into another 24-bit image while maintaining an average SSIM of 0.9999 and achieving high robustness against commonly applied watermark attacks. Furthermore, as per our best knowledge, with high payload embedding, the significant improvement in these features (in terms of saliency, PSNR, SSIM, and NC) has not been achieved by the state-of-the-art methods. Thus, the outcomes of this research realizes a trade-off independent image watermarking method, which is a first of its kind in this domain. Koksheik Wong, Vishnu Monn Baskaran |
Multim. Tools Appl. | 3 |
| 2025 | DAP-CBR: enhancing Bitcoin block propagation efficiency using dynamic compact block relay's prefilling of transactionsabstractAbstract This study examines the potential of BIP-152’s Compact Block Relay (CBR) to enhance the Bitcoin network. This work explores the block propagation efficiency through dynamic prefilling of transactions. In addition, an enhanced CBR model is proposed to reduce superfluous transaction requests, thus improving the block distribution process. The analysis considers the impact of the dynamically prefilled transactions on Bitcoin network scalability, comparing the advantages and disadvantages of this approach. We also conduct a comparative study of fixed-size and dynamically sized prefilled transactions to highlight the importance of adapting to network demands. Prefilling a fixed number of transactions without considering demand can cause inefficiencies and strain the network with unnecessary bandwidth use. Indiscriminate prefilling exacerbates these issues by inflating data packets unnecessarily, increasing latency and reducing network responsiveness. Our research indicates that the proposed solution can significantly reduce the number of round-trips between network nodes by an average of 29.77% and block reconstruction latency by 39.10% when compared with the CBR. Zi Hau Chin, Vishnu Monn Baskaran, Chee Keong Tan, Ian K. T. Tan, Timothy Tzen Vun Yap |
J. Supercomput. | 2 |
| 2024 | Reimagining Violent Action Detection with Human-Object InteractionabstractThe rising urban crime rates globally underscore the need for advanced video surveillance systems capable of autonomously detecting violent actions. Current deep learning models face limitations, struggling with subtle motions and lacking real-time capabilities. In response, we advocate for a paradigm shift in surveillance oriented violent action detection, emphasizing the pivotal role of human-object interaction (HOI) detection as opposed to conventional action recognition methodologies. Our contributions include unveiling Violence-HOI (V-HOI), a dataset capturing HOI interactions in static surveillance images. Additionally, we introduce Violence-Net (V Net), a novel convolutional-transformer network architecture, which outperforms existing HOI approaches by 5.25 percentage points in mean average precision. Moreover, when trained on V-HOI, V-Net achieves near real-time processing at 10.43 frames per second, demonstrating its practicality in dynamic surveillance scenarios. The code and dataset is available at https://github.com/MarcusLimJunYi/vhoi. Vishnu Monn Baskaran, Ricky Sutopo, JunYi Lim, Joanne Mun-Yee Lim, Koksheik Wong |
AVSS | 1 |
| 2024 | Music Form Analysis: A Case Study of The Theme and Variations FormabstractThe theme and variation music form is a hierarchical structure in music. It has a theme segment at the beginning, followed by a series of variation segments imitating the theme segment. Hence, the primary features of the theme and variation form are repetition and variation at different levels. However, due to the lack of available datasets, the theme and variation form analysis method has not been explored much. Therefore, in this work, we curate and contribute a dataset named Performance of Theme and Variation Form (PTV) and propose a theme and variation form segmentation framework to analyze the theme and variation form. Experiment results show that our method achieves an F1 score of 92.8% on our dataset with some constraints. In addition, we conduct analysis to support and encourage future studies of the theme and variation form. Jing Zhao 0033, Koksheik Wong, Vishnu Monn Baskaran, Kiki Maulana, David Taniar |
ICME | 3 |
| 2024 | Video Deception Detection through the Fusion of Multimodal Feature Extraction and Neural NetworksabstractDetecting deceptive behavior in videos is a complex task within several domains, including academic fraud assessment, commercial anti-fraud activities, judicial system evidence analysis, suspicious activity detection in security monitoring systems, and behavioral intent analysis in psychological research. In this study, we present a novel approach to video deception detection by integrating visual and audio models for deep feature fusion, primarily targeting advanced deception detection datasets. Our visual model leverages hierarchical image feature learning to enhance deceptive cue detection, complemented by an audio model that processes acoustic signals for precise speech pattern analysis. This multimodal method significantly boosts detection accuracy and lessens reliance on extensive training data. Notably, our visual model incorporates knowledge distillation technology, improving efficiency and reducing computational resource needs without compromising performance. We implement a transformer architecture using distillation tokens for effective learning and incorporate convolutional neural network insights to enrich our model’s interpretative capabilities. Experimental results demonstrate that our approach surpasses existing technologies in various standards and scenarios, offering enhanced deception recognition capabilities and addressing the challenge of limited training data. Yuanya Zhuo, Vishnu Monn Baskaran, Lillian Yee Kiaw Wang, Raphael C.-W. Phan |
IJCNN | 2 |
| 2024 | MDHA: Multi-Scale Deformable Transformer with Hybrid Anchors for Multi-View 3D Object DetectionabstractMulti-view 3D object detection is a crucial component of autonomous driving systems. Contemporary query-based methods primarily depend either on dataset-specific initialization of 3D anchors, introducing bias, or utilize dense attention mechanisms, which are computationally inefficient and unscalable. To overcome these issues, we present MDHA, a novel sparse query-based framework, which constructs adaptive 3D output proposals using hybrid anchors from multi-view, multi-scale image input. Fixed 2D anchors are combined with depth predictions to form 2.5D anchors, which are projected to obtain 3D proposals. To ensure high efficiency, our proposed Anchor Encoder performs sparse refinement and selects the top-k anchors and features. Moreover, while existing multi-view attention mechanisms rely on projecting reference points to multiple images, our novel Circular Deformable Attention mechanism only projects to a single image but allows reference points to seamlessly attend to adjacent images, improving efficiency without compromising on performance. On the nuScenes val set, it achieves 46.4% mAP and 55.0% NDS with a ResNet101 backbone. MDHA significantly outperforms the baseline where anchor proposals are modelled as learnable embeddings. Code is available at https://github.com/NaomiEX/MDHA. Michelle Adeline, Junn Yong Loo, Vishnu Monn Baskaran |
IROS | 3 |
| 2024 | Emotion-specific AUs for micro-expression recognitionabstractAbstract The Facial Action Coding System (FACS) comprehensively describes facial expressions with facial action units (AUs). It is a well-used technique by researchers in emotions research to understand human emotions better. Most micro-expression datasets provide FACS-coded AU ground truths corresponding to micro-expressions classes. It is commonly accepted in computer vision-based emotions research that certain emotions are reliably revealed when specific combinations of AUs occur. However, the reliability of the ground truth AUs in the micro-expression datasets is lower than that of normal expressions, as they have lower AU intensities. Moreover, these micro-expression datasets only report the overall reliability of all AUs. It could not be identified which AUs had been accurately coded. This work aims to revisit the ground truth AUs of popular micro-expression datasets, namely CASME II, SAMM and CAS(ME) $$^2$$ 2 , and inspect whether any AUs crucial for micro-expression recognition may need to be reconsidered. This paper also provides a detailed AU analysis which yields new AU-based RoIs for each dataset. These new RoIs improve the micro-expression recognition performances compared to the baselines considered in this work. The proposed RoIs for CASME II, SAMM and CAS(ME) $$^2$$ 2 improve the recognition rates by $$2\%$$ 2 % , $$1\%$$ 1 % and $$4\%$$ 4 % , respectively, when compared with the existing RoIs. Shu-Min Leong, Raphael C.-W. Phan, Vishnu Monn Baskaran |
Multim. Tools Appl. | 3 |
| 2024 | Sigma-Point Kalman Filter With Nonlinear Unknown Input Estimation via Optimization and Data-Driven Approach for Dynamic SystemsabstractMost works on joint state and unknown input (UI) estimation require the assumption that the UIs are linear; this is potentially restrictive as it does not hold in many intelligent autonomous systems. To overcome this restriction and circumvent the need to linearize the system, we propose a derivative-free UI sigma-point Kalman filter (SPKF-nUI), where the SPKF is interconnected with a general nonlinear UI estimator that can be implemented via nonlinear optimization and data-driven approaches. The nonlinear UI estimator uses the posterior state estimate, which is less susceptible to state prediction error. In addition, we introduce a joint sigma-point transformation scheme to incorporate both the state and UI uncertainties in the estimation of SPKF-nUI. An in-depth stochastic stability analysis proves that the proposed SPKF-nUI yields exponentially converging estimation error bounds under reasonable assumptions. Finally, two case studies are carried out on a simulation-based rigid robot and a physical soft robot, i.e., the robots made of soft materials with complex dynamics, to validate the effectiveness of the proposed filter on nonlinear dynamic systems. Our results demonstrate that the proposed SPKF-nUI achieves the lowest state and UI estimation errors when compared to the existing nonlinear state-UI filters. Junn Yong Loo, Ze Yang Ding, Vishnu Monn Baskaran, Surya Girinatha Nurzaman, Chee Pin Tan |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2023 | Computational Music: Analysis of Music Forms
Jing Zhao 0033, Koksheik Wong, Vishnu Monn Baskaran, Kiki Maulana, David Taniar |
ICCSA (1) | 3 |
| 2023 | ScratchHOI: Training Human-Object Interaction Detectors from ScratchabstractTransformer-based approaches have exhibited outstanding performances in the field of human-object interaction (HOI) detection. However, these approaches rely on underlying object detectors that have undergone large-scale pre-trainings on the ImageNet and MS-COCO dataset. This limits the potential of unique architectural designs and induces a learning bias, causing ineffective HOI representation learning. In this paper, we propose ScratchHOI, a transformer-based method for human-object interaction detection that can be trained from scratch, eliminating the need for pre-trained object detectors. ScratchHOI employs dynamic and static affinity-based feature aggregation for processing local and long-range visual information. Additional techniques are also employed to improve detection performance, such as dynamic and interactive anchor refinement for objects and interactions. Experiments on the HICO-Det dataset show that ScratchHOI achieves competitive performance against other state-of-the-art approaches over a variety of different evaluation measures. JunYi Lim, Vishnu Monn Baskaran, Joanne Mun-Yee Lim, Ricky Sutopo, Koksheik Wong, Massimo Tistarelli |
ICIP | 2 |
| 2023 | Cross-domain Transfer Learning and State Inference for Soft Robots via a Semi-supervised Sequential Variational Bayes FrameworkabstractRecently, data-driven models such as deep neural networks have shown to be promising tools for modelling and state inference in soft robots. However, voluminous amounts of data are necessary for deep models to perform effectively, which requires exhaustive and quality data collection, particularly of state labels. Consequently, obtaining labelled state data for soft robotic systems is challenged for various reasons, including difficulty in the sensorization of soft robots and the inconvenience of collecting data in unstructured environments. To address this challenge, in this paper, we propose a semi-supervised sequential variational Bayes (DSVB) framework for transfer learning and state inference in soft robots with missing state labels on certain robot configurations. Considering that soft robots may exhibit distinct dynamics under different robot configurations, a feature space transfer strategy is also incorporated to promote the adaptation of latent features across multiple configurations. Unlike existing transfer learning approaches, our proposed DSVB employs a recurrent neural network to model the nonlinear dynamics and temporal coherence in soft robot data. The proposed framework is validated on multiple setup configurations of a pneumatic-based soft robot finger. Experimental results on four transfer scenarios demonstrate that DSVB performs effective transfer learning and accurate state inference amidst missing state labels. Shageenderan Sapai, Junn Yong Loo, Ze Yang Ding, Chee Pin Tan, Raphael C.-W. Phan, Vishnu Monn Baskaran, Surya Girinatha Nurzaman |
ICRA | 6 |
| 2023 | Multi-mmlg: a novel framework of extracting multiple main melodies from MIDI filesabstractAbstract As an essential part of music, main melody is the cornerstone of music information retrieval. In the MIR’s sub-field of main melody extraction, the mainstream methods assume that the main melody is unique. However, the assumption cannot be established, especially for music with multiple main melodies such as symphony or music with many harmonies. Hence, the conventional methods ignore some main melodies in the music. To solve this problem, we propose a deep learning-based Multiple Main Melodies Generator (Multi-MMLG) framework that can automatically predict potential main melodies from a MIDI file. This framework consists of two stages: (1) main melody classification using a proposed MIDIXLNet model and (2) conditional prediction using a modified MuseBERT model. Experiment results suggest that the proposed MIDIXLNet model increases the accuracy of main melody classification from 89.62 to 97.37%. In addition, this model requires fewer parameters (71.8 million) than the previous state-of-art approaches. We also conduct ablation experiments on the Multi-MMLG framework. In the best-case scenario, predicting meaningful multiple main melodies for the music are achieved. Jing Zhao 0033, David Taniar, Kiki Maulana, Vishnu Monn Baskaran, Koksheik Wong |
Neural Comput. Appl. | 4 |
| 2023 | A Zero-Shot Soft Sensor Modeling Approach Using Adversarial Learning for Robustness Against Sensor FaultabstractSoft sensors are widely used in many industrial systems to monitor key variables that are difficult to measure, using measurements from other available physical sensors. Because physical sensors are susceptible to faults, it is crucial for soft sensor models to be robust against them. Recently, deep learning has shown promising results in developing data-driven soft sensors for various applications. However, existing learning-based soft sensors are still vulnerable to sensor faults, which could deteriorate the performance of the models. In this article, we propose a deep learning-based modeling framework for developing soft sensor models that are robust to sensor faults. Due to the difficulty in obtaining datasets that cover all possible sensor fault characteristics, the proposed framework is developed to be zero-shot such that the model can be trained with only fault-free dataset without requiring any sensor fault patterns, thus greatly saving the time and resources needed to collect such data. Instead, adversarial examples are used as a proxy for faulty sensor inputs so that the model can learn to be adaptive through the proposed two-stage, uncertainty-aware recurrent neural network architecture. We demonstrate our approach to the TE benchmark process and a real industrial multiphase flow process and show that robustness is achieved as the accuracy does not degrade significantly when sensor faults are present during the model evaluation. Ze Yang Ding, Junn Yong Loo, Surya Girinatha Nurzaman, Chee Pin Tan, Vishnu Monn Baskaran |
IEEE Trans. Ind. Informatics | 5 |
| 2023 | ERNet: An Efficient and Reliable Human-Object Interaction Detection NetworkabstractHuman-Object Interaction (HOI) detection recognizes how persons interact with objects, which is advantageous in autonomous systems such as self-driving vehicles and collaborative robots. However, current HOI detectors are often plagued by model inefficiency and unreliability when making a prediction, which consequently limits its potential for real-world scenarios. In this paper, we address these challenges by proposing ERNet, an end-to-end trainable convolutional-transformer network for HOI detection. The proposed model employs an efficient multi-scale deformable attention to effectively capture vital HOI features. We also put forward a novel detection attention module to adaptively generate semantically rich instance and interaction tokens. These tokens undergo pre-emptive detections to produce initial region and vector proposals that also serve as queries which enhances the feature refinement process in the transformer decoders. Several impactful enhancements are also applied to improve the HOI representation learning. Additionally, we utilize a predictive uncertainty estimation framework in the instance and interaction classification heads to quantify the uncertainty behind each prediction. By doing so, we can accurately and reliably predict HOIs even under challenging scenarios. Experiment results on the HICO-Det, V-COCO, and HOI-A datasets demonstrate that the proposed model achieves state-of-the-art performance in detection accuracy and training efficiency. Codes are publicly available at https://github.com/Monash-CyPhi-AI-Research-Lab/ernet. JunYi Lim, Vishnu Monn Baskaran, Joanne Mun-Yee Lim, Koksheik Wong, John See, Massimo Tistarelli |
IEEE Trans. Image Process. | 2 |
| 2023 | Is it Violin or Viola? Classifying the Instruments' Music Pieces using Descriptive StatisticsabstractClassifying music pieces based on their instrument sounds is pivotal for analysis and application purposes. Given its importance, techniques using machine learning have been proposed to classify violin and viola music pieces. The violin and viola are two different instruments with three overlapping strings of the same notes, and it is challenging for ordinary people or even musicians to distinguish the sound produced by these instruments. However, the classification of musical instrument pieces was barely performed by prior research. To solve this problem, we propose a technique using descriptive statistics to reliably distinguish between violin and viola music pieces. Likewise, a similar technique on the basis of histogram is introduced alongside the main descriptive statistics approach. These approaches are derived based on the nature of the instruments’ strings and the range of their pieces. We also solve the problem in the current literature which divide the audio into segments for processing instead of managing the whole song. Thereby, we compile a dataset of recordings that comprises of violin and viola solo pieces from the Baroque, Classical, Romantic, and Modern eras. Experiment results suggest that our approach achieves high accuracy on solo pieces as compared to other methods with 0.97 accuracy on Baroque pieces. Chong Hong Tan, Koksheik Wong, Vishnu Monn Baskaran, Kiki Maulana, David Taniar |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2022 | GraphEx: Facial Action Unit Graph for Micro-Expression ClassificationabstractFacial micro-expressions are crucial cues for expressing human emotions. Existing works have shown substantial progress in detecting micro-expressions for various applications in the computer vision field. However, it is still onerous for existing methods to handle and interpret micro-expressions efficiently. This paper proposes a deep learning-based approach leveraging spatio-temporal and graph representation learning for micro-expression classification. We design a novel Spatial-Temporal Info Extraction Network (STIENet) for learning facial appearance and muscle motion from high dimensional video clip frames and summarizes them into more meaningful feature maps. We construct an action unit (AU) relation graph to further represent the AU co-occurrence in the same micro-expression video clip. A graph neural network (GNN) is used to learn AU-related graph embedding for the downstream classification task. Performance evaluation on two mainstream micro-expression datasets, i.e., CASME II and SAMM, show that the proposed framework outperforms other state-of-the-art methods for micro-expression classification. Shu-Min Leong, Fuad Noman, Raphael C.-W. Phan, Vishnu Monn Baskaran, Chee-Ming Ting |
ICIP | 4 |
| 2022 | Invisible emotion magnification algorithm (IEMA) for real-time micro-expression recognition with graph-based features
Adamu Muhammad Buhari, Chee-Pun Ooi, Vishnu Monn Baskaran, Raphael C.-W. Phan, Koksheik Wong, Wooi-Haw Tan |
Multim. Tools Appl. | 3 |
| 2022 | Efficient Long-Term Dependencies Learning for Passenger Flow Prediction With Selective Feedback MechanismabstractWith the rapid growth of worldwide urbanization, the increasing demand for public transportation is indispensable. To improve the service quality, predicting the flow of passengers is important for the transport operators. Information on density of passengers can be used as early warnings of overcrowding and to determine if additional fleet is required. However, passenger flow forecasting is a challenging task, as it is affected by many complex factors such as spatial dependencies, temporal dependencies, and external influences. Furthermore, the ability to learn the long-term dependency of the data is also crucial, as the distant past flow information contributes to the flow over time. Most of the existing studies struggle to solve this issue, especially to learn the long-term dependency of the data, as they rely heavily on the raw handcrafted features and require high memory bandwidth to compute. To address these issues, we propose a Selective Feedback Transformer (SFT) capable of learning long-term dependency efficiently, where the selective feedback mechanism only computes the important feedback from the dominant query-key pairs in the memory. Experimental results demonstrate that the proposed model outperforms all the benchmarked methods by 27% - 37% in terms of RMSE and 36% - 50% in terms of MAE. Additionally, when the proposed model is tested with shallowed model (less number of decoding layer), it exhibits a substantial improvement of 14% - 57% in the training time and 15% - 46% in the inference time, with minimal impact on the accuracies. Ricky Sutopo, Joanne Mun-Yee Lim, Vishnu Monn Baskaran |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | Paying Attention to Varying Receptive Fields: Object Detection with Atrous Filters and Vision Transformers
Arthur Jian Shun Lam, JunYi Lim, Ricky Sutopo, Vishnu Monn Baskaran |
BMVC | 4 |
| 2021 | Deep multi-level feature pyramids: Application for non-canonical firearm detection in video surveillance
JunYi Lim, Md Istiaque Al Jobayer, Vishnu Monn Baskaran, Joanne Mun-Yee Lim, John See, Koksheik Wong |
Eng. Appl. Artif. Intell. | 3 |
| 2021 | Faceless identification based on temporal strips
Shu-Min Leong, Raphael C.-W. Phan, Vishnu Monn Baskaran, Chee-Pun Ooi |
Multim. Tools Appl. | 3 |
| 2021 | Appearance-based passenger counting in cluttered scenes with lateral movement compensation
Ricky Sutopo, Joanne Mun-Yee Lim, Vishnu Monn Baskaran, Koksheik Wong, Massimo Tistarelli, Heng Fui Liau |
Neural Comput. Appl. | 3 |
| 2018 | Dominant speaker detection in multipoint video communication using Markov chain with non-linear weights and dynamic transition window
Vishnu Monn Baskaran, Yoong Choon Chang, Jonathan Loo, Koksheik Wong, Ming-Tao Gan |
Inf. Sci. | 1 |
| 2016 | Fast watermarking scheme for real-time spatial scalable video coding
Adamu Muhammad Buhari, Huo-Chong Ling, Vishnu Monn Baskaran, Koksheik Wong |
Signal Process. Image Commun. | 3 |
| 2015 | Design and implementation of parallel video combiner architecture for multi-user video conferencing at ultra-high definition resolution
Vishnu Monn Baskaran, Yoong Choon Chang, Jonathan Loo, Koksheik Wong |
Multim. Tools Appl. | 1 |
| 2013 | Software-based serverless endpoint video combiner architecture for high-definition multiparty video conferencing
Vishnu Monn Baskaran, Yoong Choon Chang, Jonathan Loo, Koksheik Wong |
J. Netw. Comput. Appl. | 1 |