EDBT 2026 Demo / reviewers in the wild / expert
Chee-Ming Ting
dblp:86/10527
· DBLP profile ↗
38ranked-venue papers
7as first author
30since 2021 · last 2026
0000-0002-6037-3728ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 3 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 3 first-author · 12 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FLAG-4D: Flow-Guided Local-Global Dual-Deformation Model for 4D ReconstructionabstractWe introduce FLAG-4D, a novel framework for generating novel views of dynamic scenes by reconstructing how 3D Gaussian primitives evolve through space and time. Existing methods typically rely on a single Multilayer Perceptron(MLP) to model temporal deformations, and they often struggle to capture complex point motions and fine-grained dynamic details consistently over time, especially from sparse input views. Our approach, FLAG-4D overcomes this by employing a dual-deformation network that dynamically warps a canonical set of 3D Gaussians over time into new positions and anisotropic shapes. This dual-deformation network consists of an Instantaneous Deformation Network (IDN) for modeling fine-grained, local deformations, and Global Motion Network (GMN) for capturing long-range dynamics, refined via mutual learning. To ensure these deformations are both accurate and temporally smooth, FLAG-4D incorporates dense motion features from a pretrained optical flow backbone. We fuse these motion cues from adjacent timeframes and use a deformation-guided attention mechanism to align this flow information with the current state of each evolving 3D Gaussian. Extensive experiments demonstrate that FLAG-4D achieves higher-fidelity and more temporally coherent reconstructions with finer detail preservation than state-of-the-art methods. Guan Yuan Tan, Ngoc Tuan Vu, Arghya Pal, Sailaja Rajanala, Raphael C.-W. Phan, Mettu Srinivas, Chee-Ming Ting |
AAAI | 7 |
| 2026 | Multi-Modal masked autoencoder and parallel Mamba for 3D brain tumor segmentationabstractAccurate segmentation of brain tumors from multimodal MRI is essential for diagnosis and treatment planning. However, most existing approaches can only process single type of data modality, without exploiting the complementary information across different modalities. To overcome this limitation, a novel framework called MFMamba which integrates modality-aware masked autoencoder pretraining, a gated fusion strategy, and a Mamba-based backbone for efficient long-range modeling is proposed. In this design, one modality is fully masked while others are partially masked, forcing the network to reconstruct missing data through cross-modal learning. The gated fusion module then selectively incorporates generative priors into task-specific features, enhancing multimodal representations. Experimental results on the BraTS 2023 dataset show that MFMamba achieves Dice score of 93.77% for Whole Tumor and 92.69% for Tumor Core, corresponding to 1.6–2.1% improvements over state-of-the-art baselines. The gains are statistically significant ( p < 0 . 05 ), indicating the framework’s ability to deliver more precise tumor boundary delineation. Overall, the results suggest that modality-aware fusion can enhance segmentation quality while maintaining computational efficiency, underscoring its potential application for clinical image analysis. The implementation is publicly available at https://github.com/ministerhuang/MFMamba . Yaya Huang, Litong Liu, Tianzhen Zhang, Chee-Ming Ting |
Pattern Recognit. Lett. | 5 |
| 2026 | A Deep Probabilistic Flow-Based Framework for Unsupervised Cross-Domain Soft SensingabstractIndustrial soft sensing is crucial for accurate process monitoring through reliable inference of dominant sensor variables. However, developing effective data-driven soft sensor models presents challenges, such as achieving domain adaptability, addressing incomplete sensor labels, and learning stochastic data variability. To overcome these challenges, we propose a deep variational potential flow (DVPF) framework for cross-domain soft sensor modeling, taking into account the lack of sensor labels in the target domain. Our framework introduces sequential variational Bayes with recurrent neural network (RNN) parameterization to address the maximum likelihood estimation problem that characterizes cross-domain soft sensing. Central to the framework is a potential flow that performs unsupervised Bayesian inference on the RNN-extracted features to obtain an exact representation of the intractable posterior distribution. Together, these DVPF components learn domain-adaptable features that effectively capture complex cross-domain process dynamics and data variability. We validate the proposed DVPF on a real industrial multiphase flow process across varying operating modes. The results show that the DVPF demonstrates superior performance in cross-domain soft sensing compared to existing deep feature-based domain adaptation methods. Junn Yong Loo, Hwa Hui Tew, Fang Yu Leong, Ze Yang Ding, Vishnu Monn Baskaran, Chee-Ming Ting, Chee Pin Tan |
IEEE Trans. Ind. Informatics | 6 |
| 2026 | A Unified Framework for Sparse Reconstruction via Preconditioning and Nonconvex RegularizationabstractCompressed Sensing (CS) is an effective technique to recover sparse signals with fewer samples than what is required by the classical Shannon Nyquist sampling theorem. The sensing matrix, sparsifying transform, and sparse recovery algorithm are three key factors for accurate reconstruction in CS. Traditional CS uses a convex $l_{1}$-norm sparse regularizer which may lead to biased estimates and is suboptimal in promoting sparsity. Another challenge is the design of incoherent sensing matrices which is crucial for accurate sparse recovery. In this paper, we propose a novel CS framework combining a preconditioned sensing matrix and nonconvex regularization for improved sparse signal recovery. First, we formulate an optimization problem to find an incoherent sensing matrix via a preconditioner. It allows for a direct computation of the optimal preconditioner and preconditioned sensing matrix, simultaneously. Secondly, we consider a generalized CS model for signal recovery based on the incoherent sensing matrix and a nonconvex $\ell _{1/2}$-norm regularizer. We then derive an Alternating Direction Method of Multipliers (ADMM) algorithm to solve this nonconvex optimization problem. The proposed model is applied to sparse-view Computed Tomography (CT) reconstruction with highly-undersampled and noisy data. Qualitative and quantitative results show significantly better image reconstruction using the preconditioned sensing matrix and $\ell _{1/2}$ regularizer, compared to methods without preconditioning and using the $\ell _{1}$ regularizer. Prasad Theeda, Fuad Noman, Arghya Pal, Raphael C.-W. Phan, Hernando C. Ombao, Chee-Ming Ting |
IEEE J. Biomed. Health Informatics | 6 |
| 2025 | ST-HCSS: Deep Spatio-Temporal Hypergraph Convolutional Neural Network for Soft SensingabstractHigher-order sensor networks are more accurate in characterizing the nonlinear dynamics of sensory time-series data in modern industrial settings by allowing multi-node connections beyond simple pairwise graph edges. In light of this, we propose a deep spatio-temporal hypergraph convolutional neural network for soft sensing (ST-HCSS). In particular, our proposed framework is able to construct and leverage a higher-order graph (hypergraph) to model the complex multi-interactions between sensor nodes in the absence of prior structural knowledge. To capture rich spatio-temporal relationships underlying sensor data, our proposed ST-HCSS incorporates stacked gated temporal and hypergraph convolution layers to effectively aggregate and update hypergraph information across time and nodes. Our results validate the superiority of ST-HCSS compared to existing state-of-the-art soft sensors, and demonstrates that the learned hypergraph feature representations aligns well with the sensor data correlations. The code is available at https://github.com/htew0001/ST-HCSS.git Hwa Hui Tew, Gaoxuan Li, Junn Yong Loo, Chee-Ming Ting, Ze Yang Ding, Chee Pin Tan |
ICASSP | 5 |
| 2025 | HG-YOLO: Improving Tumor Detection with PP-HGNet and Global Attention Mechanism
Kien Trang, Bao Quoc Vuong, An Hoang Nguyen, Fung Fung Ting, Chee-Ming Ting |
ICCSA (1) | 5 |
| 2025 | Guided Diffusion For Class-Conditioned Synthesis & Classification Of Microscopic Blood Cell ImagesabstractMicroscopic visualization of diseased cells plays a vital role in the diagnosis and understanding of various medical conditions. Recent advances in deep learning generative models have shown remarkable potential as formidable tools for generating high-quality medical images. However, training these models generally requires large, annotated datasets, which are often costly and time-consuming to obtain. To overcome this challenge, we propose a fast-sampling guided score-based diffusion model with a classifier-free guidance strategy for class-conditioned generation of microscopic peripheral blood cell images. Our model achieves a Fréchet Inception Distance (FID) score of 10.24, demonstrating its ability to generate realistic synthetic blood cell images. Furthermore, our experimental results show that augmenting training data with these synthetic images significantly improves classification accuracy compared to relying solely on real data, highlighting the potential of synthetic data augmentation in hematology. Kar-Ee Hoh, Junn Yong Loo, Yee-Fan Tan, Raphael C.-W. Phan, Chee-Ming Ting |
ICIP | 5 |
| 2025 | Enhancing Autonomous Driving Perception Under Complex Weather Conditions Through Cyclegan-Based Driving Scene GenerationabstractComplex and dynamic weather conditions pose significant challenges to the perception capabilities of autonomous driving systems. Existing datasets often lack sufficient cover-age of adverse weather scenarios, resulting in degraded performance of vision-based modules. To address this limitation, we propose a weather-aware driving scene generation framework based on CycleGAN. Our enhanced architecture integrates a multi-scale discriminator, feature matching loss, and channel-level structural similarity (SSIM) loss to generate high-fidelity driving images under various weather conditions. To evaluate the effectiveness of our approach, we perform cross-validation using three widely adopted detection models: YOLOv8, Faster R-CNN, and EfficientDet. The results demonstrate that the generated multi-weather images significantly improve detection accuracy and enhance the robustness and reliability of the perception system in complex environments. In addition to improving dataset diversity, our method preserves fine-grained structural details and enhances the visual quality of the generated scenes. Litong Liu, Gaoxuan Li, Yaya Huang, Yixuan Dong, Chee-Ming Ting, Junn Yong Loo |
ICIP | 5 |
| 2025 | Polyfit generative model: can a group of lower-order polynomials generate high resolution diverse images?abstractImplicit neural representations (INRs) have recently gained popularity as a means to model images as continuous functions of spatial coordinates, synthesizing each pixel independently and yielding impressive results in tasks such as scene reconstruction and image generation. A notable advancement, PolyINR, utilizes element-wise multiplications between features and affine-transformed coordinates to achieve higher-order polynomial functions, eliminating the need for positional encodings. However, the finite encoding capacity of INRs, coupled with PolyINR’s recursive polynomial estimation, necessitates substantial training parameters, resulting in high computational costs and limiting applicability across diverse computer vision domains. In this work, we address these challenges by representing images as grids of smaller patches, within which we fit low-degree polynomials to capture local intensity variations. Our approach substantially reduces parameter requirements and computational demands. We evaluate our model qualitatively and quantitatively on large-scale datasets, ImageNet, CelebA, LSUN Bedroom, and Flower102; thus demonstrating competitive performance with state-of-the-art generative models, despite the absence of convolutional, normalization, or self-attention layers. Arghya Pal, Ai-Fang Chai, Sailaja Rajanala, Raphael C.-W. Phan, Koksheik Wong, Chee-Ming Ting |
IJCNN | 6 |
| 2025 | Adaptive Graph Learning with Multi-graph Convolutions for Brain Disorder Classification
Fuad Noman, Raphael C.-W. Phan, Hernando C. Ombao, Chee-Ming Ting |
MICCAI (12) | 4 |
| 2025 | T2I-Diff: fMRI Signal Generation via Time-Frequency Image Transform and Classifier-Free Denoising Diffusion Models
Hwa Hui Tew, Junn Yong Loo, Yee-Fan Tan, Hernando C. Ombao, Fuad Noman, Raphael C.-W. Phan, Chee-Ming Ting |
MICCAI (3) | 8 |
| 2025 | Res-SH: Unbiased Residual Learning for Self-Healing Interface Toughness Prediction with Limited DataabstractThe development of self-healing materials is often hindered by the high costs and material waste associated with traditional characterization methods. Current approaches to toughness prediction, primarily based on convolutional neural networks (CNNs), are limited by their tendency to capture only surface-level features, which can lead to biased predictions. Moreover, working with small datasets, which is common in materials science, further increases the risk of biased training due to overfitting, posing a critical challenge to the reliability and generalizability of predictive models. This study introduces an unbiased residual learning framework designed explicitly for predicting self-healing interface toughness under limiteddata conditions. Our approach, ResNet-inspired approach for predicting self-healing material toughness, named Res-SH, used the power of residual networks to capture deeper, more complex patterns in the data, thereby addressing critical challenges in materials research. Res-SH minimises resource consumption and experimental overhead by focusing on unbiased learning, achieving accurate predictions with fewer training epochs and lower R2score and root mean square prediction errors compared to conventional CNN and lightweight model MobileNetv2. This novel framework provides a cost-effective and resource-efficient alternative to traditional material characterization methods, reducing material waste and accelerating the discovery and optimization of self-healing material systems. Pei-Sze Tan, Karen Jia-Jun Koh, Sailaja Rajanala, Arghya Pal, Raphael C.-W. Phan, Nan Ze, Fuad Noman, Chee-Ming Ting, Norfadilah Dolmat, Nik Nur Wahidah Nik Hashim, Afidalina Tumian |
TENCON | 8 |
| 2025 | PK-YOLO: Pretrained Knowledge Guided YOLO for Brain Tumor Detection in Multiplanar MRI SlicesabstractBrain tumor detection in multiplane Magnetic Resonance Imaging (MRI) slices is a challenging task due to the various appearances and relationships in the structure of the multiplane images. In this paper, we propose a new You Only Look Once (YOLO)-based detection model that incorporates Pretrained Knowledge (PK), called PK-YOLO, to improve the performance for brain tumor detection in multiplane MRI slices. To our best knowledge, PK-YOLO is the first pretrained knowledge guided YOLO-based object detector. The main components of the new method are a pretrained pure lightweight convolutional neural network-based backbone via sparse masked modeling, a YOLO architecture with the pretrained backbone, and a regression loss function for improving small object detection. The pre-trained backbone allows for feature transferability of object queries on individual plane MRI slices into the model encoders, and the learned domain knowledge base can improve in-domain detection. The improved loss function can further boost detection performance on small-size brain tumors in multiplanar two-dimensional MRI slices. Experimental results show that the proposed PK-YOLO achieves competitive performance on the multiplanar MRI brain tumor detection datasets compared to state-of-the-art YOLO-like and DETR-like object detectors. The code is available at https://github.com/mkang315/PK-YOLO. Ming Kang 0002, Fung Fung Ting, Raphael C.-W. Phan, Chee-Ming Ting |
WACV | 4 |
| 2024 | BrainFC-CGAN: A Conditional Generative Adversarial Network for Brain Functional Connectivity Augmentation and Aging SynthesisabstractBrain functional connectivity (FC) changes are associated with neuropsychiatric disorders and other underlying factors, such as age and gender. Due to small training sample, data augmentation has been increasingly used for deep learning-based classification of brain FC. Although deep generative models could generate brain FCs to enhance downstream classification, most existing methods neglect the underlying factors involved in the generation process and fail to preserve the subject identity. We propose a novel brain FC conditional Generative Adversarial Network (GAN) called BrainFC-CGAN with specialized layers and filters to preserve the symmetry property and topological structure of brain FCs. We design a FC generator that captures the complex variations between brain FCs, ages, and health statuses to generate synthetic FCs that preserve the subject identity. We categorized true brain FCs into different age groups; an augmented age-specific dataset generated from BrainFC-CGAN is combined with the training set for classification. Experimental results on major depressive disorder (MDD) resting-state functional magnetic resonance imaging data show that the proposed method synthesizes realistic brain FCs of different target age groups, significantly improving downstream classification performance over baseline without augmentation, and also outperforming several state-of-the-art GANs. Yee-Fan Tan, Junn Yong Loo, Chee-Ming Ting, Fuad Noman, Raphael C.-W. Phan, Hernando C. Ombao |
ICASSP | 3 |
| 2024 | Cafct-Net: A Cnn-Transformer Hybrid Network With Contextual And Attentional Feature Fusion For Liver Tumor SegmentationabstractMedical image semantic segmentation techniques can help identify tumors automatically from computed tomography (CT) scans. In this paper, we propose a Contextual and Attentional feature Fusions enhanced Convolutional Neural Network (CNN) and Transformer hybrid network (CAFCT-Net) for liver tumor segmentation. We incorporate three novel modules in the CAFCT-Net architecture: Attentional Feature Fusion (AFF), Atrous Spatial Pyramid Pooling (ASPP) of DeepLabv3, and Attention Gates (AGs) to improve contextual information related to tumor boundaries for accurate segmentation. Experimental results show that the proposed model achieves a mean Intersection over Union (IoU) of 76.54% and Dice coefficient of 84.29%, respectively, on the Liver Tumor Segmentation Benchmark (LiTS) dataset, outperforming pure CNN or Transformer methods, e.g., Attention U-Net and PVTFormer. Ming Kang 0002, Chee-Ming Ting, Fung Fung Ting, Raphael C.-W. Phan |
ICIP | 2 |
| 2024 | CST-Yolo: A Novel Method For Blood Cell Detection Based On Improved Yolov7 And CNN-Swin TransformerabstractBlood cell detection is a typical small-scale object detection problem in computer vision. In this paper, we propose a CST-YOLO model for blood cell detection based on YOLOv7 architecture and enhance it with the CNN-Swin Transformer (CST), which is a new attempt at CNN-Transformer fusion. We also introduce three other useful modules: Weighted Efficient Layer Aggregation Networks (W-ELAN), Multiscale Channel Split (MCS), and Concatenate Convolutional Layers (CatConv) in our CST-YOLO to improve small-scale object detection precision. Experimental results show that the proposed CST-YOLO achieves 92.7%, 95.6%, and 91.1% mAP @ 0.5, respectively, on three blood cell datasets, outperforming state-of-the-art object detectors, e.g., RT-DETR, YOLOv5, and YOLOv7. Our code is available at https://github.com/mkang315/CST-YOLO. Ming Kang 0002, Chee-Ming Ting, Fung Fung Ting, Raphael C.-W. Phan |
ICIP | 2 |
| 2024 | Deep Multi-Graph Embedded Clustering for Community Detection in FMRI Functional Brain Networks Across IndividualsabstractAnalyzing the community structure of brain networks provides new insights into human brain function. Existing studies broadly use conventional network clustering approaches. While graph neural networks have recently shown promise in modeling brain functional connectivity (FC) networks, their applications to brain community detection still need improvement and further refinement. Moreover, identifying common community structure while resolving the single-subject partitions across multiple individual networks remains underexplored. We propose a Deep Multi-Graph Embedded Clustering (DMGEC) framework to identify shared community partition in brain FC networks over a cohort of individuals. By incorporating the consensus information aggregated across network structures, DMGEC leverages a graph autoencoder to produce consensus-aware latent representations of individual networks, and applies deep embedded clustering on the multi-subject network representation to produce common community assignment of brain nodes. Simulations show superior community recovery by our method compared to conventional approaches, especially for networks with large number of communities. When applied to functional magnetic resonance imaging (fMRI) data, the DMGEC achieves outstanding alikeness over individual partitions, and uncovers group-level differences in brain community motifs between major depressive disorder patients and normal controls. Kai-Jun See, Chee-Ming Ting, Fuad Noman, Junn Yong Loo, Yee-Fan Tan, Hernando C. Ombao, Raphael C.-W. Phan |
ICIP | 2 |
| 2024 | A Preconditioning Approach To Optimizing Sensing Matrix For Improved Compressed Sensing CT ReconstructionabstractCompressed sensing (CS) exploiting inherent sparsity prior of signals has been proven effective for sparse-view computed tomography (CT) image reconstruction from undersampled projection data. However, most CS-based CT studies focused on formulating different sparsity regularizers, e.g., total variation (TV) minimization, and neglect design of an incoherent sensing matrix - a key factor of CS performance. The sensing matrix formed by an incomplete set of Radon projections in CT typically exhibits large coherence. In this paper, we propose a novel method for optimizing the sensing matrix via preconditioning to improve CS-CT reconstruction. A well-conditioned preconditioner is designed to optimally reduce the coherence of the sensing matrix and thus improving the CS systems. The desired preconditioner is obtained by solving a nonconvex optimization problem via gradient descent method. The preconditioned systems solved by TV-based sparse recovery algorithms can provide better reconstruction accuracy with fewer measurements even in noisy settings. Evaluated on brain and COVID-19 chest CT datasets, the proposed method when used for preconditioning of Radon sensing matrix reconstructed images with substantially higher quality with faster speed than baselines without preconditioning. Prasad Theeda, Chee-Ming Ting, Arghya Pal, Hernando C. Ombao |
ICIP | 2 |
| 2024 | Dynamic MRI Reconstruction Using Low-Rank Plus Sparse Decomposition With Smoothness RegularizationabstractThe low-rank plus sparse (L+S) decomposition model has enabled better reconstruction of dynamic magnetic resonance imaging (dMRI) with separation into background (L) and dynamic (S) component. However, use of low-rank prior alone may not fully explain the slow variations or smoothness of the background part at the local scale. In this paper, we propose a smoothness-regularized L+S (SR-L+S) model for dMRI reconstruction from highly undersampled k-t-space data. We exploit joint low-rank and smooth priors on the background component of dMRI to better capture both its global and local temporal correlated structures. Extending the L+S formulation, the low-rank property is encoded by the nuclear norm, while the smoothness by a general $\ell_{p}$-norm penalty on the local differences of the columns of L. The additional smoothness regularizer can promote piecewise local consistency between neighboring frames. By smoothing out the noise and dynamic activities, it allows accurate recovery of the background part, and subsequently more robust dMRI reconstruction. Extensive experiments on multi-coil cardiac and synthetic data shows that the SR-L+S model outperforms several state-of-the-art methods in terms of recovery accuracy. Chee-Ming Ting, Fuad Noman, Raphael C.-W. Phan, Hernando C. Ombao |
ICIP | 1 |
| 2024 | A Deep Probabilistic Spatiotemporal Framework for Dynamic Graph Representation Learning with Application to Brain Disorder Identification
Sin-Yee Yap, Junn Yong Loo, Chee-Ming Ting, Fuad Noman, Raphael C.-W. Phan, Adeel Razi, David L. Dowe |
IJCAI | 3 |
| 2024 | BGF-YOLO: Enhanced YOLOv8 with Multiscale Attentional Feature Fusion for Brain Tumor Detection
Ming Kang 0002, Chee-Ming Ting, Fung Fung Ting, Raphael C.-W. Phan |
MICCAI (8) | 2 |
| 2024 | KANS: Knowledge Discovery Graph Attention Network for Soft Sensing in Multivariate Industrial ProcessesabstractSoft sensing of hard-to-measure variables is often crucial in industrial processes. Current practices rely heavily on conventional modeling techniques that show success in improving accuracy. However, they overlook the non-linear nature, dynamics characteristics, and non-Euclidean dependencies between complex process variables. To tackle these challenges, we present a framework known as a Knowledge discovery graph Attention Network for effective Soft sensing (KANS). Unlike the existing deep learning soft sensor models, KANS can discover the intrinsic correlations and irregular relationships between the multivariate industrial processes without a predefined topology. First, an unsupervised graph structure learning method is introduced, incorporating the cosine similarity between different sensor embedding to capture the correlations between sensors. Next, we present a graph attention-based representation learning that can compute the multivariate data parallelly to enhance the model in learning complex sensor nodes and edges. To fully explore KANS, knowledge discovery analysis has also been conducted to demonstrate the interpretability of the model. Experimental results demonstrate that KANS significantly outperforms all the baselines and state-of-the-art methods in soft sensing performance. Furthermore, the analysis shows that KANS can find sensors closely related to different process variables without domain knowledge, significantly improving soft sensing accuracy. Hwa Hui Tew, Gaoxuan Li, Xuewen Luo, Junn Yong Loo, Chee-Ming Ting, Ze Yang Ding, Chee Pin Tan |
SMC | 6 |
| 2024 | ASF-YOLO: A novel YOLO model with attentional scale sequence fusion for cell instance segmentationabstractWe propose a novel Attentional Scale Sequence Fusion based You Only Look Once (YOLO) framework (ASF-YOLO) which combines spatial and scale features for accurate and fast cell instance segmentation. Built on the YOLO segmentation framework, we employ the Scale Sequence Feature Fusion (SSFF) module to enhance the multiscale information extraction capability of the network, and the Triple Feature Encoder (TFE) module to fuse feature maps of different scales to increase detailed information. We further introduce a Channel and Position Attention Mechanism (CPAM) to integrate both the SSFF and TFE modules, which focus on informative channels and spatial position-related small objects for improved detection and segmentation performance. Experimental validations on two cell datasets show remarkable segmentation accuracy and speed of the proposed ASF-YOLO model. It achieves a box mAP of 0.91, mask mAP of 0.887, and an inference speed of 47.3 FPS on the 2018 Data Science Bowl dataset, outperforming the state-of-the-art methods. The source code is available at https://github.com/mkang315/ASF-YOLO. Ming Kang 0002, Chee-Ming Ting, Fung Fung Ting, Raphael C.-W. Phan |
Image Vis. Comput. | 2 |
| 2024 | Graph Autoencoders for Embedding Learning in Brain Networks and Major Depressive Disorder IdentificationabstractBrain functional connectivity (FC) networks inferred from functional magnetic resonance imaging (fMRI) have shown altered or aberrant brain functional connectome in various neuropsychiatric disorders. Recent application of deep neural networks to connectome-based classification mostly relies on traditional convolutional neural networks (CNNs) using input FCs on a regular Euclidean grid to learn spatial maps of brain networks neglecting the topological information of the brain networks, leading to potentially sub-optimal performance in brain disorder identification. We propose a novel graph deep learning framework that leverages non-Euclidean information inherent in the graph structure for classifying brain networks in major depressive disorder (MDD). We introduce a novel graph autoencoder (GAE) architecture, built upon graph convolutional networks (GCNs), to embed the topological structure and node content of large fMRI networks into low-dimensional representations. For constructing the brain networks, we employ the Ledoit-Wolf (LDW) shrinkage method to efficiently estimate high-dimensional FC metrics from fMRI data. We explore both supervised and unsupervised techniques for graph embedding learning. The resulting embeddings serve as feature inputs for a deep fully-connected neural network (FCNN) to distinguish MDD from healthy controls (HCs). Evaluating our model on resting-state fMRI MDD dataset, we observe that the GAE-FCNN outperforms several state-of-the-art methods for brain connectome classification, achieving the highest accuracy when using LDW-FC edges as node features. The graph embeddings of fMRI FC networks also reveal significant group differences between MDD and HCs. Our framework demonstrates the feasibility of learning graph embeddings from brain networks, providing valuable discriminative information for diagnosing brain disorders. Fuad Noman, Chee-Ming Ting, Hakmook Kang, Raphael C.-W. Phan, Hernando C. Ombao |
IEEE J. Biomed. Health Informatics | 2 |
| 2023 | A Unified Framework for Static and Dynamic Functional Connectivity Augmentation for Multi-Domain Brain Disorder ClassificationabstractDeep learning (DL) methods recently show promise on accurate brain disorder classification using functional connectivity (FC) estimated from functional magnetic resonance imaging (fMRI). However, DL model building can be hindered by small sample-size settings of fMRI. Moreover, most studies utilize either static (sFC) or dynamic FC (dFC) for classification. We propose a unified framework for data augmentation of both sFC and dFC for multi-domain joint classification of brain disorders. We exploit generative adversarial networks (GAN) to synthesize realistic FCs for data augmentation. Notably, we adopted the TimeGAN for dFC generation that can capture temporal dependencies in real dFC, and the GR-SPD-GAN for sFC generation that preserves the spatial connectivity structure. We further develop BrainFusionNet - a specialized DL model for multi-domain FC that simultaneously learns embedded features from both sFC and dFC to provide complementary spatio-temporal information for downstream classification. The synthetic FC data are augmented in training data to improve the BrainFusionNet performance and generalizability. Experimental results on major depressive disorder (MDD) identification using resting-state fMRI show substantial improvement in classification accuracy by our framework, outperforming competing models without FC augmentation and using sFC or dFC features alone. Yee-Fan Tan, Chee-Ming Ting, Fuad Noman, Raphael C.-W. Phan, Hernando C. Ombao |
ICIP | 2 |
| 2023 | RCS-YOLO: A Fast and High-Accuracy Object Detector for Brain Tumor Detection
Ming Kang 0002, Chee-Ming Ting, Fung Fung Ting, Raphael C.-W. Phan |
MICCAI (4) | 2 |
| 2022 | GraphEx: Facial Action Unit Graph for Micro-Expression ClassificationabstractFacial micro-expressions are crucial cues for expressing human emotions. Existing works have shown substantial progress in detecting micro-expressions for various applications in the computer vision field. However, it is still onerous for existing methods to handle and interpret micro-expressions efficiently. This paper proposes a deep learning-based approach leveraging spatio-temporal and graph representation learning for micro-expression classification. We design a novel Spatial-Temporal Info Extraction Network (STIENet) for learning facial appearance and muscle motion from high dimensional video clip frames and summarizes them into more meaningful feature maps. We construct an action unit (AU) relation graph to further represent the AU co-occurrence in the same micro-expression video clip. A graph neural network (GNN) is used to learn AU-related graph embedding for the downstream classification task. Performance evaluation on two mainstream micro-expression datasets, i.e., CASME II and SAMM, show that the proposed framework outperforms other state-of-the-art methods for micro-expression classification. Shu-Min Leong, Fuad Noman, Raphael C.-W. Phan, Vishnu Monn Baskaran, Chee-Ming Ting |
ICIP | 5 |
| 2022 | Graph Autoencoder-Based Embedded Learning in Dynamic Brain Networks for Autism Spectrum Disorder IdentificationabstractRecent applications of pattern recognition techniques to brain connectome-based classification focus on static functional connectivity (FC) neglecting the dynamics of FC over time, and use input connectivity matrices on a regular Euclidean grid. We exploit the graph convolutional networks (GCNs) to learn irregular structural patterns in brain FC networks and propose extensions to capture dynamic changes in network topology. We develop a dynamic graph autoencoder (DyGAE)-based framework to leverage the time-varying topological structures of dynamic brain networks for identification of autism spectrum disorder (ASD). The framework combines a GCN-based DyGAE to encode individual-level dynamic networks into time-varying low-dimensional network embeddings, and classifiers based on weighted fully-connected neural network (FCNN) and long short-term memory (LSTM) to facilitate dynamic graph classification via the learned spatial-temporal information. Evaluation on a large ABIDE resting-state functional magnetic resonance imaging (rs-fMRI) dataset shows that our method outperformed state-of-the-art methods in detecting altered FC in ASD. Dynamic FC analyses with DyGAE learned embeddings also reveal apparent group difference between ASD and healthy controls in network profiles and switching dynamics of brain states. Fuad Noman, Sin-Yee Yap, Raphael C.-W. Phan, Hernando C. Ombao, Chee-Ming Ting |
ICIP | 5 |
| 2022 | Separating Stimulus-Induced and Background Components of Dynamic Functional Connectivity in Naturalistic fMRIabstractWe consider the challenges in extracting stimulus-related neural dynamics from other intrinsic processes and noise in naturalistic functional magnetic resonance imaging (fMRI). Most studies rely on inter-subject correlations (ISC) of low-level regional activity and neglect varying responses in individuals. We propose a novel, data-driven approach based on low-rank plus sparse ( [Formula: see text]) decomposition to isolate stimulus-driven dynamic changes in brain functional connectivity (FC) from the background noise, by exploiting shared network structure among subjects receiving the same naturalistic stimuli. The time-resolved multi-subject FC matrices are modeled as a sum of a low-rank component of correlated FC patterns across subjects, and a sparse component of subject-specific, idiosyncratic background activities. To recover the shared low-rank subspace, we introduce a fused version of principal component pursuit (PCP) by adding a fusion-type penalty on the differences between the columns of the low-rank matrix. The method improves the detection of stimulus-induced group-level homogeneity in the FC profile while capturing inter-subject variability. We develop an efficient algorithm via a linearized alternating direction method of multipliers to solve the fused-PCP. Simulations show accurate recovery by the fused-PCP even when a large fraction of FC edges are severely corrupted. When applied to natural fMRI data, our method reveals FC changes that were time-locked to auditory processing during movie watching, with dynamic engagement of sensorimotor systems for speech-in-noise. It also provides a better mapping to auditory content in the movie than ISC. Chee-Ming Ting, Jeremy I. Skipper, Fuad Noman, Steven L. Small, Hernando C. Ombao |
IEEE Trans. Medical Imaging | 1 |
| 2021 | Detecting Dynamic Community Structure in Functional Brain Networks Across Individuals: A Multilayer ApproachabstractOBJECTIVE: We present a unified statistical framework for characterizing community structure of brain functional networks that captures variation across individuals and evolution over time. Existing methods for community detection focus only on single-subject analysis of dynamic networks; while recent extensions to multiple-subjects analysis are limited to static networks. METHOD: To overcome these limitations, we propose a multi-subject, Markov-switching stochastic block model (MSS-SBM) to identify state-related changes in brain community organization over a group of individuals. We first formulate a multilayer extension of SBM to describe the time-dependent, multi-subject brain networks. We develop a novel procedure for fitting the multilayer SBM that builds on multislice modularity maximization which can uncover a common community partition of all layers (subjects) simultaneously. By augmenting with a dynamic Markov switching process, our proposed method is able to capture a set of distinct, recurring temporal states with respect to inter-community interactions over subjects and the change points between them. RESULTS: Simulation shows accurate community recovery and tracking of dynamic community regimes over multilayer networks by the MSS-SBM. Application to task fMRI reveals meaningful non-assortative brain community motifs, e.g., core-periphery structure at the group level, that are associated with language comprehension and motor functions suggesting their putative role in complex information integration. Our approach detected dynamic reconfiguration of modular connectivity elicited by varying task demands and identified unique profiles of intra and inter-community connectivity across different task conditions. CONCLUSION: The proposed multilayer network representation provides a principled way of detecting synchronous, dynamic modularity in brain networks across subjects. Chee-Ming Ting, S. Balqis Samdin, Meini Tang, Hernando C. Ombao |
IEEE Trans. Medical Imaging | 1 |
| 2020 | A Markov-Switching Model Approach to Heart Sound Segmentation and ClassificationabstractOBJECTIVE: We consider challenges in accurate segmentation of heart sound signals recorded under noisy clinical environments for subsequent classification of pathological events. Existing state-of-the-art solutions to heart sound segmentation use probabilistic models such as hidden Markov models (HMMs), which, however, are limited by its observation independence assumption and rely on pre-extraction of noise-robust features. METHODS: We propose a Markov-switching autoregressive (MSAR) process to model the raw heart sound signals directly, which allows efficient segmentation of the cyclical heart sound states according to the distinct dependence structure in each state. To enhance robustness, we extend the MSAR model to a switching linear dynamic system (SLDS) that jointly model both the switching AR dynamics of underlying heart sound signals and the noise effects. We introduce a novel algorithm via fusion of switching Kalman filter and the duration-dependent Viterbi algorithm, which incorporates the duration of heart sound states to improve state decoding. RESULTS: Evaluated on Physionet/CinC Challenge 2016 dataset, the proposed MSAR-SLDS approach significantly outperforms the hidden semi-Markov model (HSMM) in heart sound segmentation based on raw signals and comparable to a feature-based HSMM. The segmented labels were then used to train Gaussian-mixture HMM classifier for identification of abnormal beats, achieving high average precision of 86.1% on the same dataset including very noisy recordings. CONCLUSION: The proposed approach shows noticeable performance in heart sound segmentation and classification on a large noisy dataset. SIGNIFICANCE: It is potentially useful in developing automated heart monitoring systems for pre-screening of heart pathologies. Fuad Noman, Sheikh Hussain Shaikh Salleh, Chee-Ming Ting, S. Balqis Samdin, Hernando C. Ombao, Hadri Hussain |
IEEE J. Biomed. Health Informatics | 3 |
| 2020 | A Multi-Domain Connectome Convolutional Neural Network for Identifying Schizophrenia From EEG Connectivity PatternsabstractOBJECTIVE: We exploit altered patterns in brain functional connectivity as features for automatic discriminative analysis of neuropsychiatric patients. Deep learning methods have been introduced to functional network classification only very recently for fMRI, and the proposed architectures essentially focused on a single type of connectivity measure. METHODS: We propose a deep convolutional neural network (CNN) framework for classification of electroencephalogram (EEG)-derived brain connectome in schizophrenia (SZ). To capture complementary aspects of disrupted connectivity in SZ, we explore combination of various connectivity features consisting of time and frequency-domain metrics of effective connectivity based on vector autoregressive model and partial directed coherence, and complex network measures of network topology. We design a novel multi-domain connectome CNN (MDC-CNN) based on a parallel ensemble of 1D and 2D CNNs to integrate the features from various domains and dimensions using different fusion strategies. We also consider an extension to dynamic brain connectivity using the recurrent neural networks. RESULTS: Hierarchical latent representations learned by the multiple convolutional layers from EEG connectivity reveals apparent group differences between SZ and healthy controls (HC). Results on a large resting-state EEG dataset show that the proposed CNNs significantly outperform traditional support vector machine classifier. The MDC-CNN with combined connectivity features further improves performance over single-domain CNNs using individual features, achieving remarkable accuracy of 91.69% with a decision-level fusion. CONCLUSION: The proposed MDC-CNN by integrating information from diverse brain connectivity descriptors is able to accurately discriminate SZ from HC. SIGNIFICANCE: The new framework is potentially useful for developing diagnostic tools for SZ and other disorders. Chun-Ren Phang, Fuad Noman, Hadri Hussain, Chee-Ming Ting, Hernando C. Ombao |
IEEE J. Biomed. Health Informatics | 4 |
| 2019 | Short-segment Heart Sound Classification Using an Ensemble of Deep Convolutional Neural NetworksabstractThis paper proposes a framework based on deep convolutional neural networks (CNNs) for automatic heart sound classification using short-segments of individual heart beats. We design a 1D-CNN that directly learns features from raw heart-sound signals, and a 2D-CNN that takes inputs of two-dimensional time-frequency feature maps based on Mel-frequency cepstral coefficients. We further develop a time-frequency CNN ensemble (TF-ECNN) combining the 1D-CNN and 2D-CNN based on score-level fusion of the class probabilities. On the large PhysioNet CinC challenge 2016 database, the proposed CNN models outperformed traditional classifiers based on support vector machine and hidden Markov models with various hand-crafted time- and frequency-domain features. Best classification scores with 89.22% accuracy and 89.94% sensitivity were achieved by the ECNN, and 91.55% specificity and 88.82% modified accuracy by the 2D-CNN alone on the test set. Fuad Noman, Chee-Ming Ting, Sheikh Hussain Shaikh Salleh, Hernando C. Ombao |
ICASSP | 2 |
| 2018 | Robust Facial Expression Recognition for MuCI: A Comprehensive Neuromuscular Signal AnalysisabstractThis paper presents a comprehensive study on the analysis of neuromuscular signal activities to recognize 11 facial expressions for muscle computer interfacing applications. A robust denoising protocol comprised of Wavelet transform and Kalman filtering is proposed to enhance the electromyogram (EMG) signal-to-noise ratio and improve classification performance. The effectiveness of eight different time-domain facial EMG features on system performance is examined and compared in order to identify the most discriminative one. Fourteen pattern recognition-based algorithms are employed to classify the extracted features. These classifiers are evaluated in terms of classification accuracy and processing time. Finally, the best methods that obtain almost identical system performance are compared through the Normalized Mutual Information (NMI) criterion and a repeated measure analysis of variance (ANOVA) for a statistical significant test.To clarify the impact of signal denoising, all considered EMG features and classifiers are assessed with and without this stage. Results show that: (1) the proposed denosing step significantly improves the system performance; (2) root mean square is the most discriminative facial EMG feature; (3) discriminant analysis when the parameters are estimated by the Maximum Likelihood algorithm achieves the highest classification accuracy and NMI; however, ANOVA reveals no significant difference among the best methods with almost similar performance. Mahyar Hamedi, Sheikh Hussain Shaikh Salleh, Chee-Ming Ting, Mehdi Astaraki, Alias Mohd Noor |
IEEE Trans. Affect. Comput. | 3 |
| 2018 | Estimating Dynamic Connectivity States in fMRI Using Regime-Switching Factor ModelsabstractWe consider the challenges in estimating the state-related changes in brain connectivity networks with a large number of nodes. Existing studies use the sliding-window analysis or time-varying coefficient models, which are unable to capture both smooth and abrupt changes simultaneously, and rely on ad-hoc approaches to the high-dimensional estimation. To overcome these limitations, we propose a Markov-switching dynamic factor model, which allows the dynamic connectivity states in functional magnetic resonance imaging (fMRI) data to be driven by lower-dimensional latent factors. We specify a regime-switching vector autoregressive (SVAR) factor process to quantity the time-varying directed connectivity. The model enables a reliable, data-adaptive estimation of change-points of connectivity regimes and the massive dependencies associated with each regime. We develop a three-step estimation procedure: 1) extracting the factors using principal component analysis, 2) identifying connectivity regimes in a low-dimensional subspace based on the factor-based SVAR model, and 3) constructing high-dimensional state connectivity metrics based on the subspace estimates. Simulation results show that our estimator outperforms -means clustering of time-windowed coefficients, providing more accurate estimate of time-evolving connectivity. It achieves percentage of reduction in mean squared error by 60% when the network dimension is comparable to the sample size. When applied to the resting-state fMRI data, our method successfully identifies modular organization in the resting-statenetworksin consistencywith other studies. It further reveals changes in brain states with variations across subjects and distinct large-scale directed connectivity patterns across states. Chee-Ming Ting, Hernando C. Ombao, S. Balqis Samdin, Sheikh Hussain Shaikh Salleh |
IEEE Trans. Medical Imaging | 1 |
| 2015 | Is First-Order Vector Autoregressive Model Optimal for fMRI Data?abstractWe consider the problem of selecting the optimal orders of vector autoregressive (VAR) models for fMRI data. Many previous studies used model order of one and ignored that it may vary considerably across data sets depending on different data dimensions, subjects, tasks, and experimental designs. In addition, the classical information criteria (IC) used (e.g., the Akaike IC (AIC)) are biased and inappropriate for the high-dimensional fMRI data typically with a small sample size. We examine the mixed results on the optimal VAR orders for fMRI, especially the validity of the order-one hypothesis, by a comprehensive evaluation using different model selection criteria over three typical data types--a resting state, an event-related design, and a block design data set--with varying time series dimensions obtained from distinct functional brain networks. We use a more balanced criterion, Kullback's IC (KIC) based on Kullback's symmetric divergence combining two directed divergences. We also consider the bias-corrected versions (AICc and KICc) to improve VAR model selection in small samples. Simulation results show better small-sample selection performance of the proposed criteria over the classical ones. Both bias-corrected ICs provide more accurate and consistent model order choices than their biased counterparts, which suffer from overfitting, with KICc performing the best. Results on real data show that orders greater than one were selected by all criteria across all data sets for the small to moderate dimensions, particularly from small, specific networks such as the resting-state default mode network and the task-related motor networks, whereas low orders close to one but not necessarily one were chosen for the large dimensions of full-brain networks. Chee-Ming Ting, Abd-Krim Seghouane, Muhammad Usman Khalid, Sheikh Hussain Shaikh Salleh |
Neural Comput. | 1 |
| 2015 | Estimating Effective Connectivity from fMRI Data Using Factor-based Subspace Autoregressive ModelsabstractWe consider the problem of identifying large-scale effective connectivity of brain networks from fMRI data. Standard vector autoregressive (VAR) models fail to estimate reliably networks with large number of nodes. We propose a new method based on factor modeling for reliable and efficient high-dimensional VAR analysis of large networks. We develop a subspace VAR (SVAR) model from a factor model (FM), where observations are driven by a lower-dimensional subspace of common latent factors with an AR dynamics. We consider two variants of principal components (PC) methods that provide consistent estimates for the FM hence the implied SVAR model, even of large dimensions. Information criterion is used to select the optimal subspace dimension. We established asymptotic normality and convergence rates for the estimated SVAR coefficients matrix. Evaluation on simulated resting-state fMRI shows that the SVAR models are more robust and produce better connectivity estimates than the classical model for a moderately-large network analysis. Results on real data by varying the subspace dimensions identify strong connections in the default mode network and reveal hierarchical connectivity of resting-state networks with distinct functional relevance. Chee-Ming Ting, Abd-Krim Seghouane, Sheikh Hussain Shaikh Salleh, A. B. Mohd Noor |
IEEE Signal Process. Lett. | 1 |
| 2014 | Artifact Removal from Single-Trial ERPs using Non-Gaussian Stochastic Volatility Models and Particle FilterabstractThis paper considers improved modeling of artifactual noise for denoising of single-trial event-related potentials (ERPs) by state-space approach. Instead of the inadequate constant variance models used in existing studies, we propose to use stochastic volatility (SV) models to better describe the time-varying volatility in real ERP noise sources. We further propose a class of non-Gaussian SV models to capture the abrupt volatility changes typically present in impulsive noise, to improve artifact removal from ERPs. Two specifications are considered: (1) volatility driven by a heavy-tailed component and (2) transformation of volatility. Both result in volatility processes with heavy-tailed transition densities which can predict the impulsive noise volatility dynamics, more accurately than the Gaussian models. These SV noise models are incorporated in an autoregressive (AR) state-space ERP dynamic model. Parameter estimation is done using a Rao-Blackwellized particle filter (RBPF). Evaluation on simulated auditory brainstem responses (ABRs), corrupted by real eye-blink artifacts, shows that the non-Gaussian models can accurately detect the artifact-induced abrupt volatility spikes, and able to uncover the underlying inter-trial dynamics. Among them, the log-SV model performs the best. The results on real data demonstrate significant artifact suppression. Chee-Ming Ting, Sheikh Hussain Shaikh Salleh, Zaitul Marlizawati Zainuddin, Arifah Bahar |
IEEE Signal Process. Lett. | 1 |