Mengling Feng

dblp:31/7025 · DBLP profile ↗
← Back
59ranked-venue papers
8as first author
38since 2021 · last 2026
0000-0002-5338-6248ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 28 · 4 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 21 · 2 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 2 first-author · 9 since 2021Databases, data management, data science and information retrieval · 11 · 3 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Recovering Coherent Affective Patterns: Addressing Modality Missing in Multimodal Sentiment Analysis
abstract
Multimodal sentiment analysis (MSA) seeks to decode human emotions by integrating heterogeneous modalities. However, real-world scenarios often involve missing or misaligned data due to sensor failures or transmission errors, leading to disrupted temporal dynamics and degraded cross-modal correlations. To address these challenges, we propose RECAP (REcovery of Coherent Affective Patterns), a robust two-stage framework to restore temporal and structural emotional integrity under modality incompleteness. The first stage employs a causality-aware adversarial generator for multi-granularity temporal reconstruction, complemented by a contrastive mutual information factorization module that disentangles shared and modality-specific semantics. The second stage introduces a mutual information-guided attention fusion mechanism with a ranking-based objective, enabling adaptive integration of complementary signals for refined prediction. Extensive experiments on MOSI, MOSEI, and SIMS under various missing-modality conditions demonstrate that RECAP consistently outperforms state-of-the-art methods. Notably, it improves ACC-7 on MOSI by 2.71 percentage points and F1 on SIMS by 6.38 percentage points. These results verify the performance of RECAP in terms of capturing fine-grained emotional cues and robustness.
Huiting Huang, Tieliang Gong, Kai He 0001, Wen Wen 0013, Weizhan Zhang, Mengling Feng
AAAI6
2026 MedForge: Interpretable Medical Deepfake Detection via Forgery-aware Reasoning
abstract
Text-guided image editors can now manipulate authentic medical scans with high fidelity, enabling lesion implantation/removal that threatens clinical trust and safety. Existing defenses are inadequate for healthcare. Medical detectors are largely black-box, while MLLM-based explainers are typically post-hoc, lack medical expertise, and may hallucinate evidence on ambiguous cases. We present MedForge, a data-and-method solution for pre-hoc, evidence-grounded medical forgery detection. We introduce MedForge-90K, a large-scale benchmark of realistic lesion edits across 19 pathologies with expert-guided reasoning supervision via doctor inspection guidelines and gold edit locations. Building on it, MedForge-Reasoner performs localize-then-analyze reasoning, predicting suspicious regions before producing a verdict, and is further aligned with Forgery-aware GSPO to strengthen grounding and reduce hallucinations. Experiments demonstrate state-of-the-art detection accuracy and trustworthy, expert-aligned explanations.
Kai He 0001, Qingyuan Lei, Bin Pu, Jian Zhang 0087, Yuling Xu, Mengling Feng
ACL (1)7
2026 medDreamer: Model-Based Reinforcement Learning with Latent Imagination on Complex EHRs for Clinical Decision Support
abstract
Timely and personalized treatment decisions are essential across a wide range of healthcare settings where patient responses can vary significantly and evolve over time. Clinical data used to support these treatment decisions are often irregularly sampled, where missing data frequencies may implicitly convey information about the patient's condition. Existing Reinforcement Learning (RL) based clinical decision support systems often ignore the missing patterns and distort them with coarse discretization and simple imputation. They are also predominantly model-free and largely depend on retrospective data, which could lead to insufficient exploration and bias by historical behaviors. To address these limitations, we propose medDreamer, a novel model-based reinforcement learning framework for personalized treatment recommendation. medDreamer contains a world model with an Adaptive Feature Integration module that simulates latent patient states from irregular data and a two-phase policy trained on a hybrid of real and imagined trajectories. This enables learning optimal policies that go beyond the sub-optimality of historical clinical decisions, while remaining close to real clinical data. We evaluate medDreamer on both sepsis and mechanical ventilation treatment tasks using two large-scale Electronic Health Records (EHRs) datasets. Comprehensive evaluations show that medDreamer significantly outperforms model-free and model-based baselines in both clinical outcomes and off-policy metrics.
Qianyi Xu, Gousia Habib, Feng Wu 0001, Dilruk Perera, Mengling Feng
KDD (1)5
2026 Multi-granularity semantic extraction and multi-task fusion for Chinese medical entity normalization
Kai He 0001, Rui Mao 0010, Mengling Feng
Expert Syst. Appl.6
2026 Smart Imitator: Learning from Imperfect Clinical Decisions
abstract
OBJECTIVES: This study introduces Smart Imitator (SI), a 2-phase reinforcement learning (RL) solution enhancing personalized treatment policies in healthcare, addressing challenges from imperfect clinician data and complex environments. MATERIALS AND METHODS: Smart Imitator's first phase uses adversarial cooperative imitation learning with a novel sample selection schema to categorize clinician policies from optimal to nonoptimal. The second phase creates a parameterized reward function to guide the learning of superior treatment policies through RL. Smart Imitator's effectiveness was validated on 2 datasets: a sepsis dataset with 19 711 patient trajectories and a diabetes dataset with 7234 trajectories. RESULTS: Extensive quantitative and qualitative experiments showed that SI significantly outperformed state-of-the-art baselines in both datasets. For sepsis, SI reduced estimated mortality rates by 19.6% compared to the best baseline. For diabetes, SI reduced HbA1c-High rates by 12.2%. The learned policies aligned closely with successful clinical decisions and deviated strategically when necessary. These deviations aligned with recent clinical findings, suggesting improved outcomes. DISCUSSION: Smart Imitator advances RL applications by addressing challenges such as imperfect data and environmental complexities, demonstrating effectiveness within the tested conditions of sepsis and diabetes. Further validation across diverse conditions and exploration of additional RL algorithms are needed to enhance precision and generalizability. CONCLUSION: This study shows potential in advancing personalized healthcare learning from clinician behaviors to improve treatment outcomes. Its methodology offers a robust approach for adaptive, personalized strategies in various complex and uncertain environments.
Dilruk Perera, Kay Choong See, Mengling Feng
J. Am. Medical Informatics Assoc.4
2026 DeepEN: A deep reinforcement learning framework for personalized enteral nutrition in critical care
Daniel J. Tan, Jiayang Chen, Dilruk Perera, Kay Choong See, Mengling Feng
J. Biomed. Informatics5
2026 PAL: Prompting analytic learning with missing modality for multi-modal class-incremental learning
Xianghu Yue, Yiming Chen 0010, Xueyi Zhang 0001, Xiaoxue Gao, Mengling Feng, Mingrui Lao, Huiping Zhuang, Haizhou Li 0001
Pattern Recognit.5
2026 Incorporating Large Vision Model Distillation and Fuzzy Perception for Improving Disease Diagnosis
abstract
Early disease diagnosis is critical for timely clinical intervention and treatment. Intelligent models have shown significant potential in addressing the challenges of misdiagnosis, especially given the shortage of experienced experts. However, there exists complexity of ultrasound image information, and subtle differences between positive and negative samples, combined with limited disease data, pose challenges in feature extraction, class imbalance, and the lack of representative positive class prototypes. To this end, we propose a large vision model distillation framework with fuzzy perception to improve rare disease diagnosis. Specifically, we first fine-tune the pre-trained Medical Segment Anything Model (MedSAM) to adapt it to the target domain. Through image augmentation and latent feature relationship distillation, we enhance feature extraction robustness and reduce inter-class ambiguity, which helps mitigate the impact of class imbalance on the lightweight student model. Second, as imaging style differences in ultrasound images are often more pronounced than subtle variations between positive and negative samples, we construct image style subsets and introduce a fuzzy style matching strategy to perceive these differences. Finally, we combine features from both the differential perception and semantic enhancement branches to strengthen disease classification. Extensive experiments on an internal fetal spina bifida dataset and three widely used imbalanced benign-malignant medical datasets demonstrate the effectiveness of the proposed method.
Qika Lin, Huaxuan Wen, Bin Pu, Mengling Feng, Kenli Li 0001
IEEE Trans. Fuzzy Syst.5
2026 External Retrievals or Internal Priors? From RAG to Epitome-Augmented Generation by Fuzzy Selection
abstract
Retrieval-Augmented Generation (RAG) offers a promising solution to the limitations of static knowledge and hallucinations in Large Language Models (LLMs). While prior research has introduced numerous enhancements to RAG systems, a significant challenge remains under-explored: the potential conflict between external retrievals and LLMs' internal priors, which can undermine the quality of generated outputs. To tackle this issue, we present theEpitome-AugmentedGeneration (EAG) framework, which strategically aligns queries, external retrievals, and internal priors to produce high-quality LLM generations by selecting fuzzy inputs. EAG employs two novel lightweight modules, Criticism and Distillation, allowing traditional RAGs to be upgraded to EAGs without the need for specialized training data. Extensive experiments on five datasets across general and medical domains, including both open-ended and closed-ended tasks, validate the effectiveness of EAG. Our framework achieves substantial F1 score improvements: 7.03%, 23.35%, and 21.58% over baseline RAGs in medical QA tasks, 11.80% in law domain, 7.95% in finance domain, and 4.13% and 5.16% in general domain. Beyond performance gains, our study delves into the interplay between LLMs' internal priors and external retrievals, uncovering key principles that govern generation quality and providing valuable insights for future retrieval-augmented frameworks.
Kai He 0001, Jiaxing Xu, Qika Lin, Zeyu Gao 0001, Jialun Wu, Mengling Feng
IEEE Trans. Fuzzy Syst.8
2026 EsurvFusion: An Evidential Multimodal Survival Fusion Model Based on Epistemic Random Fuzzy Sets
abstract
Multimodal survival analysis aims to combine heterogeneous data sources to improve the prediction quality of survival outcomes. However, this task is particularly challenging due to high heterogeneity and noise across data sources. Additionally, the exact survival time is often censored (partially known) due to incomplete event observation. To address the above challenges, we propose a novel interpretable evidential multimodal survival fusion model, EsurvFusion. This model is designed to combine multimodal data at the decision level using Epistemic Random Fuzzy Sets that jointly handle both data and model uncertainty while incorporating modality-level reliability. Specifically, EsurvFusion first models unimodal data with newly introduced Gaussian random fuzzy numbers, producing possible unimodal survival predictions along with corresponding aleatory and epistemic uncertainty. It then estimates modality-level reliability through a reliability discounting layer to correct the misleading impact of noisy data modalities. Finally, a multimodal evidence fusion layer is introduced to combine the discounted predictions, revealing modality-level influence based on the learned reliability coefficients. Extensive experiments on four multimodal cancer survival datasets demonstrate the effectiveness of our model in handling highly heterogeneous data, establishing a new state-of-the-art performance on several benchmarks.
Ling Huang 0003, Yucheng Xing, Qika Lin, Jinming Duan 0001, Su Ruan, Mengling Feng
IEEE Trans. Fuzzy Syst.6
2026 BrainPrompt+: Multi-Level Brain Prompt Learning for Knowledge-Guided Neurological Disorder Identification
abstract
Accurate identification of neurological disorders such as Alzheimer's disease (AD), Parkinson's disease (PD), and Autism Spectrum Disorder (ASD) is challenging due to subtle early-stage symptoms and heterogeneous brain dynamics. Resting-state functional MRI (rs-fMRI) enables the construction of functional brain networks, where Graph Neural Networks (GNNs) have shown promise for disease classification. However, existing GNN-based methods face three key limitations: correlation-based graph construction introduces noise and negative edges; domain knowledge about brain regions is ignored; and demographic or clinical metadata are fused through simplistic encodings. To overcome these limitations, we propose BrainPrompt+, a knowledge-guided framework that integrates Large Language Models (LLMs) with multi-level natural language prompts. Five types of prompts are introduced: spectral (frequency-domain BOLD features), spatial (inter-ROI connectivity), ROI (anatomical and functional knowledge), disease (progression stages), and subject (demographic context). These prompts are encoded by a frozen LLM and incorporated into a GNN pipeline, unifying imaging, clinical, and external knowledge in a semantically enriched and interpretable manner. Experiments on three rs-fMRI datasets show that BrainPrompt+ consistently outperforms state-of-the-art baselines, achieving accuracy gains of up to 8.93%. Biomarker analysis further demonstrates that the highlighted ROIs align with established neuroscience findings, confirming the interpretability of the model. BrainPrompt+ thus establishes a flexible and generalizable paradigm for knowledge-guided brain network analysis. The source code is available at https://github.com/AngusMonroe/BrainPromptPlus.
Jiaxing Xu, Kai He 0001, Wei Li 0231, Mengcheng Lan, Yue Xun, Qika Lin, Peifan Ran, Yiping Ke, Mengling Feng
IEEE Trans. Medical Imaging10
2025 Crab: A Novel Configurable Role-Playing LLM with Assessing Benchmark
abstract
Kai He, Yucheng Huang, Wenqing Wang, Delong Ran, Dongming Sheng, Junxuan Huang, Qika Lin, Jiaxing Xu, Wenqiang Liu, Mengling Feng. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Kai He 0001, Delong Ran, Dongming Sheng, Junxuan Huang, Qika Lin, Jiaxing Xu, Mengling Feng
ACL (1)10
2025 Self-supervised Quantized Representation for Seamlessly Integrating Knowledge Graphs with Large Language Models
abstract
Qika Lin, Tianzhe Zhao, Kai He, Zhen Peng, Fangzhi Xu, Ling Huang, Jingying Ma, Mengling Feng. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Qika Lin, Tianzhe Zhao, Kai He 0001, Zhen Peng 0005, Fangzhi Xu, Ling Huang 0003, Jingying Ma, Mengling Feng
ACL (1)8
2025 PTransformer: A Prompt-Based Multimodal Transformer Architecture For Medical Tabular Data
Yucheng Ruan, Xiang Lan 0004, Daniel J. Tan, Hairil Rizal Abdullah, Mengling Feng
AIME (1)5
2025 DivScore: Zero-Shot Detection of LLM-Generated Text in Specialized Domains
abstract
Detecting LLM-generated text in specialized and high-stakes domains like medicine and law is crucial for combating misinformation and ensuring authenticity.However, current zeroshot detectors, while effective on general text, often fail when applied to specialized content due to domain shift.We provide a theoretical analysis showing this failure is fundamentally linked to the KL divergence between human, detector, and source text distributions.To address this, we propose DivScore, a zero-shot detection framework using normalized entropybased scoring and domain knowledge distillation to robustly identify LLM-generated text in specialized domains.We also release a domainspecific benchmark for LLM-generated text detection in the medical and legal domains.Experiments on our benchmark show that Di-vScore consistently outperforms state-of-theart detectors, with 14.4% higher AUROC and 64.0% higher recall (0.1% false positive rate threshold).In adversarial settings, DivScore demonstrates superior robustness to other baselines, achieving on average 22.8% advantage in AUROC and 29.5% in recall.Code and data are publicly available 1 .
Kai He 0001, Yunxiao Zhu, Mengling Feng
EMNLP5
2025 Teaching AI the Anatomy Behind the Scan: Addressing Anatomical Flaws in Medical Image Segmentation with Learnable Prior
Young Seok Jeon, Hongfei Yang, Huazhu Fu, Mengling Feng
ICCV4
2025 ST-USleepNet: A Spatial-Temporal Coupling Prominence Network for Multi-Channel Sleep Staging
abstract
Sleep staging is critical to assess sleep quality and diagnose disorders. Despite advancements in artificial intelligence enabling automated sleep staging, significant challenges remain: (1) Simultaneously extracting prominent temporal and spatial sleep features from multi-channel raw signals, including characteristic sleep waveforms and salient spatial brain networks. (2) Capturing the spatial-temporal coupling patterns essential for accurate sleep staging. To address these challenges, we propose a novel framework named ST-USleepNet, comprising a spatial-temporal graph construction module (ST) and a U-shaped sleep network (USleepNet). The ST module converts raw signals into a spatial-temporal graph based on signal similarity, temporal, and spatial relationships to model spatial-temporal coupling patterns. The USleepNet employs a U-shaped structure for both the temporal and spatial streams, mirroring its original use in image segmentation to isolate significant targets. Applied to raw sleep signals and graph data from the ST module, USleepNet effectively segments these inputs, simultaneously extracting prominent temporal and spatial sleep features. Testing on three datasets demonstrates that ST-USleepNet outperforms existing baselines, and model visualizations confirm its efficacy in extracting prominent sleep features and temporal-spatial coupling patterns across various sleep stages. The code is available at https://github.com/Majy-Yuji/ST-USleepNet.
Jingying Ma, Qika Lin, Ziyu Jia, Mengling Feng
IJCAI4
2025 No More Sliding Window: Efficient 3D Medical Image Segmentation with Differentiable Top-K Patch Sampling
Young Seok Jeon, Hongfei Yang, Huazhu Fu, Yeshe M. Kway, Mengling Feng
MICCAI (16)5
2025 BrainPrompt: Multi-level Brain Prompt Enhancement for Neurological Condition Identification
Jiaxing Xu, Kai He 0001, Wei Li 0231, Mengcheng Lan, Xia Dong, Yiping Ke, Mengling Feng
MICCAI (12)8
2025 GEM: Empowering MLLM for Grounded ECG Understanding with Time Series and Images
abstract
While recent multimodal large language models (MLLMs) have advanced automated ECG interpretation, they still face two key limitations: (1) insufficient multimodal synergy between ECG time series and ECG images, and (2) limited explainability in linking diagnoses to granular waveform evidence. We introduce GEM, the first MLLM unifying ECG time series, 12-lead ECG images and text for grounded and clinician-aligned ECG interpretation. GEM enables feature-grounded analysis, evidence-driven reasoning, and a clinician-like diagnostic process through three core innovations: a dual-encoder framework extracting complementary time series and image features, cross-modal alignment for effective multimodal understanding, and knowledge-guided instruction data generation for generating high-granularity grounding data (ECG-Grounding) linking diagnoses to measurable parameters ($e.g.$, QRS/PR Intervals). Additionally, we propose the Grounded ECG Understanding task, a clinically motivated benchmark designed to comprehensively assess the MLLM's capability in grounded ECG understanding. Experimental results on both existing and our proposed benchmarks show GEM significantly improves predictive performance (CSN $7.4\%$ $\uparrow$), explainability ($22.7\%$ $\uparrow$), and grounding ($25.3\%$ $\uparrow$), making it a promising approach for real-world clinical applications. Codes, model, and data are available at https://github.com/lanxiang1017/GEM.
Xiang Lan 0004, Feng Wu 0001, Kai He 0001, Qinghao Zhao, Shenda Hong, Mengling Feng
NeurIPS6
2025 Evidential time-to-event prediction with calibrated uncertainty quantification
abstract
Time-to-event analysis provides insights into clinical prognosis and treatment recommendations. However, this task is more challenging than standard regression problems due to the presence of censored observations. Additionally, the lack of confidence assessment, model robustness, and prediction calibration raises concerns about the reliability of predictions. To address these challenges, we propose an evidential regression model specifically designed for time-to-event prediction. Our approach computes a degree of belief for the event time occurring within a time interval, without any strict distribution assumption. Meanwhile, the proposed model quantifies both epistemic and aleatory uncertainties using Gaussian Random Fuzzy Numbers and belief functions, providing clinicians with uncertainty-aware survival time predictions. Experimental evaluations using simulated and real-world survival datasets highlight the potential of our approach for enhancing clinical decision-making in survival analysis. • Calculate survival time without strict distribution assumptions. • Quantifies epistemic and aleatory uncertainties via GRFN and belief functions. • Provides uncertainty-aware survival predictions for clinicians. • Validated on simulated & real-world survival datasets.
Ling Huang 0003, Yucheng Xing, Swapnil Mishra, Thierry Denoeux, Mengling Feng
Int. J. Approx. Reason.5
2025 Guest Editorial: Multimodal Representation and Reasoning for Social Computing
Mengling Feng, Erik Cambria, Qika Lin, Kaize Shi, Weiping Li 0005
IEEE Trans. Comput. Soc. Syst.1
2025 Cross-Modal Knowledge Diffusion-Based Generation for Difference-Aware Medical VQA
abstract
Multimodal medical applications have garnered considerable attention due to their potential to offer comprehensive and robust support for medical assistance. Specifically, within this domain, difference-aware medical Visual Question Answering (VQA) has emerged as a topic of increasing interest that enables the recognition of changes in physical conditions over time when compared to previous states and provides customized suggestions accordingly. However, it is challenging because samples usually exhibit characteristics of complexity, diversity, and inherent noise. Besides, there is a need for multimodal knowledge understanding of the medical domain. The difference-aware setting requiring image comparison further intensifies these situations. To this end, we propose a cross-Modal knowlEdge diffusioN-baseD gEneration netwoRk (MENDER), where the diffusion mechanism with multi-step denoising and knowledge injection from global to local level are employed to tackle the aforementioned challenges, respectively. The diffusion process is to gradually generate answers with the sequence input of questions, random noises for the answer masks and virtual vision prompts of images. The strategy of answer nosing and knowledge cascading is specifically tailored for this task and is implemented during forward and reverse diffusion processes. Moreover, the visual and structure knowledge injection are proposed to learn virtual vision prompts to guide the diffusion process, where the former is realized using a pre-trained medical image-text network and the latter is modeled with spatial and semantic graph structures processed by the heterogeneous graph Transformer models. Experiment results demonstrate the effectiveness of MENDER for difference-aware medical VQA. Furthermore, it also exhibits notable performance in the low-resource setting and conventional medical VQA tasks.
Qika Lin, Kai He 0001, Yifan Zhu 0001, Fangzhi Xu, Erik Cambria, Mengling Feng
IEEE Trans. Image Process.6
2025 Frailty Modeling Using Machine Learning Methodologies: A Systematic Review With Discussions on Outstanding Questions
abstract
Studying frailty is crucial for enhancing the health and quality of life among older adults, refining healthcare delivery methods, and tackling the obstacles linked to an aging demographic. Approaches to frailty modeling often utilise simple analytic techniques rather than available advanced machine learning methods, which may be sub-optimal. There is no large-scale systematic review on applications of machine learning methods on frailty modeling. In this study we explore the use of machine learning methods to predict or classify frailty in older persons in routinely collected data. We reviewed 181 research articles, and categorised analytic methods into three categories: generalised linear models, survival models, and non-linear models. These methods have a moderate agreement with existing frailty scores and predictive validity for adverse outcomes. Limited evidence suggests that non-linear methods outperform generalised linear methods. The top-three predictor/input variables are specific diagnosis or groups of diagnoses, functional performance (e.g., ADLs), and impaired cognition. Mortality, hospital admissions and prolonged hospital stay are the mainly predicted outcomes. Most studies utilise classical machine learning methods with cross-sectional data. Longitudinal data collected by wearable sensors have been used for frailty modeling. We also discuss the opportunities to use more advanced machine learning methods with high dimensional longitudinal data for more personalised and accessible frailty tools.
Hongfei Yang, Jiangeng Chang, Wenbo He 0004, Caitlin Fern Wee, John Soong Tshon Yit, Mengling Feng
IEEE J. Biomed. Health Informatics6
2024 Learning the Unlearned: Mitigating Feature Suppression in Contrastive Learning
Jihai Zhang 0002, Xiang Lan 0004, Xiaoye Qu, Yu Cheng 0001, Mengling Feng, Bryan Hooi
ECCV (83)5
2024 Towards Enhancing Time Series Contrastive Learning: A Dynamic Bad Pair Mining Approach
abstract
*Not all positive pairs are beneficial to time series contrastive learning*. In this paper, we study two types of bad positive pairs that can impair the quality of time series representation learned through contrastive learning: the noisy positive pair and the faulty positive pair. We observe that, with the presence of noisy positive pairs, the model tends to simply learn the pattern of noise (Noisy Alignment). Meanwhile, when faulty positive pairs arise, the model wastes considerable amount of effort aligning non-representative patterns (Faulty Alignment). To address this problem, we propose a Dynamic Bad Pair Mining (DBPM) algorithm, which reliably identifies and suppresses bad positive pairs in time series contrastive learning. Specifically, DBPM utilizes a memory module to dynamically track the training behavior of each positive pair along training process. This allows us to identify potential bad positive pairs at each epoch based on their historical training behaviors. The identified bad pairs are subsequently down-weighted through a transformation module, thereby mitigating their negative impact on the representation learning process. DBPM is a simple algorithm designed as a lightweight **plug-in** without learnable parameters to enhance the performance of existing state-of-the-art methods. Through extensive experiments conducted on four large-scale, real-world time series datasets, we demonstrate DBPM's efficacy in mitigating the adverse effects of bad positive pairs.
Xiang Lan 0004, Hanshu Yan, Shenda Hong, Mengling Feng
ICLR4
2024 Artificial Intelligence and Data Science for Healthcare: Bridging Data-Centric AI and People-Centric Healthcare
abstract
KDD AIDSH 2024 aims to foster discussions and developments that push the boundaries of Artificial Intelligence (AI) and Data Science (DS) in healthcare, enhance diagnostic accuracy and promote human-centric approaches to healthcare, thus stimulating future interdisciplinary collaborations. This year's symposium will focus on expanding the application of AI/DS in healthcare/medicine and bridging existing gaps. The workshop invites submissions of full papers as well as work-in-progress on the application of AI/DS in healthcare. The workshop will feature three invited talks from eminent speakers, spanning academia, industry, and clinical researchers. In addition, selected papers will be invited to publish in Health Data Science, a Science Partner Journal. This summary provides a brief description of the half-day workshop to be held on August 26th, 2024. The webpage for the workshop can be found at https://aimel.ai/kdd2024aidsh.
Shenda Hong, Daoxin Yin, Gongzheng Tang, Tianfan Fu, Liantao Ma, Mengling Feng, Mai Wang, Fei Wang 0001, Luxia Zhang
KDD7
2024 A review of uncertainty quantification in medical image analysis: Probabilistic and non-probabilistic methods
abstract
The comprehensive integration of machine learning healthcare models within clinical practice remains suboptimal, notwithstanding the proliferation of high-performing solutions reported in the literature. A predominant factor hindering widespread adoption pertains to an insufficiency of evidence affirming the reliability of the aforementioned models. Recently, uncertainty quantification methods have been proposed as a potential solution to quantify the reliability of machine learning models and thus increase the interpretability and acceptability of the results. In this review, we offer a comprehensive overview of the prevailing methods proposed to quantify the uncertainty inherent in machine learning models developed for various medical image tasks. Contrary to earlier reviews that exclusively focused on probabilistic methods, this review also explores non-probabilistic approaches, thereby furnishing a more holistic survey of research pertaining to uncertainty quantification for machine learning models. Analysis of medical images with the summary and discussion on medical applications and the corresponding uncertainty evaluation protocols are presented, which focus on the specific challenges of uncertainty in medical image analysis. We also highlight some potential future research work at the end. Generally, this review aims to allow researchers from both clinical and technical backgrounds to gain a quick and yet in-depth understanding of the research in uncertainty quantification for medical image analysis machine learning models.
Ling Huang 0003, Su Ruan, Yucheng Xing, Mengling Feng
Medical Image Anal.4
2024 Selective HuBERT: Self-Supervised Pre-Training for Target Speaker in Clean and Mixture Speech
abstract
Self-supervised pre-trained speech models were shown effective for various downstream speech processing tasks. Since they are mainly pre-trained to map input speech to pseudo-labels, the resulting representations are only effective for the type of pre-train data used, either clean or mixture speech. With the idea of selective auditory attention, we propose a novel pre-training solution called Selective-HuBERT, or SHuBERT, which learns the selective extraction of target speech representations from either clean or mixture speech. Specifically, SHuBERT is trained to predict pseudo labels of a target speaker, conditioned on an enrolment speech from the target speaker. By doing so, SHuBERT is expected to selectively attend to the target speaker in a complex acoustic environment, thus benefiting various downstream tasks. We further introduce a dual-path training strategy and use the cross-correlation constraint between the two branches to encourage the model to generate noise-invariant representation. Experiments on SUPERB benchmark and LibriMix dataset demonstrate the universality and noise-robustness of SHuBERT. Furthermore, we find that our high-quality representation can be easily integrated with conventional supervised learning methods to achieve significant performance, even under extremely low-resource labeled data.
Jingru Lin, Meng Ge, Wupeng Wang, Haizhou Li 0001, Mengling Feng
IEEE Signal Process. Lett.5
2024 FCSN: Global Context Aware Segmentation by Learning the Fourier Coefficients of Objects in Medical Images
abstract
The encoder-decoder model is a commonly used Deep Neural Network (DNN) model for medical image segmentation. Conventional encoder-decoder models make pixel-wise predictions focusing heavily on local patterns around the pixel. This makes it challenging to give segmentation that preserves the object's shape and topology, which often requires an understanding of the global context. In this work, we propose a Fourier Coefficient Segmentation Network (FCSN)—a novel global context-aware DNN model that segments an object by learning the complex Fourier coefficients of the object's masks. The Fourier coefficients are calculated by integrating over the whole contour. Therefore, for our model to make a precise estimation of the coefficients, the model is motivated to incorporate the global context of the object, leading to a more accurate segmentation of the object's shape. This global context awareness also makes our model robust to unseen local perturbations during inference, such as additive noise or motion blur that are prevalent in medical images. We compare FCSN with other state-of-the-art global context-aware models (UNet++, DeepLabV3+, UNETR) on 5 medical image segmentation tasks, of which 3 are camera imaging datasets (ISIC_2018, RIM_CUP, RIM_DISC) and 2 are medical imaging datasets (PROSTATE, FETAL). When FCSN is compared with UNETR, FCSN attains significantly lower Hausdorff scores with 19.14 (6%), 17.42 (6%), 9.16 (14%), 11.18 (22%), and 5.98 (6%) for ISIC_2018, RIM_CUP, RIM_DISC, PROSTATE, and FETAL tasks respectively. Moreover, FCSN is lightweight by discarding the decoder module, which incurs significant computational overhead. FCSN only requires 29.7 M parameters which are 75.6 M and 9.9 M fewer parameters than UNETR and DeepLabV3+, respectively. FCSN attains inference and training speeds of 1.6 ms/img and 6.3 ms/img, which is 8× and 3× faster than UNet and UNETR. The code for FCSN is made publicly available athttps://github.com/nus-mornin-lab/FCSN.
Young Seok Jeon, Hongfei Yang, Mengling Feng
IEEE J. Biomed. Health Informatics3
2023 Demystifying Complex Treatment Recommendations: A Hierarchical Cooperative Multi-Agent RL Approach
abstract
Reinforcement Learning (RL) techniques are widely adopted in healthcare research to learn effective policies for complex sequential decision-making tasks, such as treatment recommendation to achieve better clinical outcomes. However, existing RL based solutions suffer from two major limitations. (1) Burdens a single agent to make complex treatment recommendations: existing RL algorithms do not explicitly support the common hierarchical decision-making processes used in practise. That is, they consider treatment recommendation as a singular decision and the treatment space as a flat search space, and expect a single RL agent to select the best treatment outright. (2) Sub optimal patient state representation techniques: accurate representation of patient states and their dynamics is essential to make effective clinical decisions. However, there are three main limitations in current state representation techniques in the clinical environment: (i) lack of clear structure and high variability of input features; (ii) lack of a temporal context, which makes it difficult to fully capture patient conditions and their dynamics at a given time; (iii) and the use of imputation techniques, which undermines personalization, introduce biases and errors that can cause harm to patients. To address the complexities of treatment recommendation and generate superior treatment policies, we propose a consolidated RL solution that incorporates a hierarchical multi-agent architecture and advanced patient state representation technique. In our proposed architecture, the complex task is decomposed into a hierarchy of subtasks, each handled by a specialized agent that coordinates skills learned by lower level agents to determine a high-level treatment policy. The patient state representation technique is designed to effectively encode the changing states of patients throughout their hospital stay, addressing the three primary limitations of current approaches. Our experiments using the sepsis treatment task shows that the proposed solution consistently outperforms state-of-the-art baselines and improves patient survival.
Dilruk Perera, Mengling Feng
IJCNN3
2023 Automatic Calcification Morphology and Distribution Classification for Breast Mammograms With Multi-Task Graph Convolutional Neural Network
abstract
The morphology and distribution of microcalcifications are the most important descriptors for radiologists to diagnose breast cancer based on mammograms. However, it is very challenging and time-consuming for radiologists to characterize these descriptors manually, and there also lacks of effective and automatic solutions for this problem. We observed that the distribution and morphology descriptors are determined by the radiologists based on the spatial and visual relationships among calcifications. Thus, we hypothesize that this information can be effectively modelled by learning a relationship-aware representation using graph convolutional networks (GCNs). In this study, we propose a multi-task deep GCN method for automatic characterization of both the morphology and distribution of microcalcifications in mammograms. Our proposed method transforms morphology and distribution characterization into node and graph classification problem and learns the representations concurrently. We trained and validated the proposed method in an in-house dataset and public DDSM dataset with 195 and 583 cases,respectively. The proposed method reaches good and stable results with distribution AUC at 0.812 ± 0.043 and 0.873 ± 0.019, morphology AUC at 0.663 ± 0.016 and 0.700 ± 0.044 for both in-house and public datasets. In both datasets, our proposed method demonstrates statistically significant improvements compared to the baseline models. The performance improvements brought by our proposed multi-task mechanism can be attributed to the association between the distribution and morphology of calcifications in mammograms, which is interpretable using graphical visualizations and consistent with the definitions of descriptors in the standard BI-RADS guideline. In short, we explore, for the first time, the application of GCNs in microcalcification characterization that suggests the potential of using graph learning for more robust understanding of medical images.
Hao Du 0005, Min-Szu Yao, Liangyu Chen 0005, Wing P. Chan, Mengling Feng
IEEE J. Biomed. Health Informatics6
2022 Intra-Inter Subject Self-Supervised Learning for Multivariate Cardiac Signals
abstract
Learning information-rich and generalizable representations effectively from unlabeled multivariate cardiac signals to identify abnormal heart rhythms (cardiac arrhythmias) is valuable in real-world clinical settings but often challenging due to its complex temporal dynamics. Cardiac arrhythmias can vary significantly in temporal patterns even for the same patient (i.e., intra subject difference). Meanwhile, the same type of cardiac arrhythmia can show different temporal patterns among different patients due to different cardiac structures (i.e., inter subject difference). In this paper, we address the challenges by proposing an Intra-Inter Subject Self-Supervised Learning (ISL) model that is customized for multivariate cardiac signals. Our proposed ISL model integrates medical knowledge into self-supervision to effectively learn from intra-inter subject differences. In intra subject self-supervision, ISL model first extracts heartbeat-level features from each subject using a channel-wise attentional CNN-RNN encoder. Then a stationarity test module is employed to capture the temporal dependencies between heartbeats. In inter subject self-supervision, we design a set of data augmentations according to the clinical characteristics of cardiac signals and perform contrastive learning among subjects to learn distinctive representations for various types of patients. Extensive experiments on three real-world datasets were conducted. In a semi-supervised transfer learning scenario, our pre-trained ISL model leads about 10% improvement over supervised training when only 1% labeled data is available, suggesting strong generalizability and robustness of the model.
Xiang Lan 0004, Dianwen Ng, Shenda Hong, Mengling Feng
AAAI4
2022 Deep learning for temporal data representation in electronic health records: A systematic review of challenges and methodologies
Feng Xie 0004, Yilin Ning, Marcus Eng Hock Ong, Mengling Feng, Wynne Hsu, Bibhas Chakraborty, Nan Liu 0003
J. Biomed. Informatics5
2022 Federated Learning for Electronic Health Records
abstract
In data-driven medical research, multi-center studies have long been preferred over single-center ones due to a single institute sometimes not having enough data to obtain sufficient statistical power for certain hypothesis testings as well as predictive and subgroup studies. The wide adoption of electronic health records (EHRs) has made multi-institutional collaboration much more feasible. However, concerns over infrastructures, regulations, privacy, and data standardization present a challenge to data sharing across healthcare institutions. Federated Learning (FL), which allows multiple sites to collaboratively train a global model without directly sharing data, has become a promising paradigm to break the data isolation. In this study, we surveyed existing works on FL applications in EHRs and evaluated the performance of current state-of-the-art FL algorithms on two EHR machine learning tasks of significant clinical importance on a real world multi-center EHR dataset.
Trung Kien Dang, Xiang Lan 0004, Jianshu Weng, Mengling Feng
ACM Trans. Intell. Syst. Technol.4
2022 A Clustering-Based Optimization Method for the Driving Cycle Construction: A Case Study in Fuzhou and Putian, China
abstract
Driving cycle is a crucial topic for the auto industry. It is developed to provide a quantitative measure on the fuel consumption and emission of a vehicle. In recent years, massive amount of driving data has been collected but has not yet been commonly used for the evaluation of driving cycle. We believe the collection of such data and the advancement in analytics models may provide a fresh perspective for the construction of driving cycle. Therefore, we propose a novel clustering-based optimization method for the construction of driving cycles. We employ the principal component analysis and spectral clustering algorithms to eliminate redundant features and analyze data structure. We further develop an adaptive optimization algorithm to select the appropriate kinematic segments to form a representative driving cycle. To demonstrate the effectiveness of our method, we compare our performance against the baselines including the New European Driving Cycle (NEDC), Federal Test Procedure (FTP), and Markov chain-based methods. The model performance is evaluated with real driving data from two cities in Fujian, China. Our proposed method is shown to be superior to all baselines. In addition, based on our optimized driving cycle, we can also estimate the fuel consumption to evaluate its energy economy. To sum up, this study offers a novel methodology to establish the driving cycle based on real and localized traffic data, where the constructed driving cycle can further be used for the development of energy economy and emission control.
Huaxin Qiu 0002, Shaoze Cui, Sutong Wang, Yanzhang Wang, Mengling Feng
IEEE Trans. Intell. Transp. Syst.5
2021 Adversarial Domain Adaptation with Correlation-Based Association Networks for Longitudinal Disk Fault Prediction
abstract
Disk fault is known to be the key cause of data loss in the modern large-scale data center, which affects the reliability and stability of the server and even the whole IT infrastructure, resulting in high financial cost. Recent works on disk fault prediction demonstrate the ability of machine learning techniques in the early prediction of disk failure. However, two limitations hinder the real-world application of current methods. First, they ignore the data heterogeneity in the data center, where distribution shifts commonly exist across different disk model types that decrease the model's performance. Second, the number of disks of different disk model types varies greatly in the data center, and current methods failed to deliver an acceptable performance over disk model types with few training samples. To address these limitations, we propose adversarial domain adaptation with correlation-based association networks (ADA-CBAN) to both mitigate the distribution shift problem and also to boost the performance on small-scaled data. Extensive experiments prove that our proposed model is effective and achieves new state-of-the-art. In addition, our post model analysis can also reveal important feature interactions and highlight the crucial period before disk faults, where both are useful information in real-world applications.
Xiang Lan 0004, Dianwen Ng, Jiongzhou Liu, Mengling Feng
IJCNN7
2021 Interpretable and Lightweight 3-D Deep Learning Model for Automated ACL Diagnosis
abstract
We propose an interpretable and lightweight 3D deep neural network model that diagnoses anterior cruciate ligament (ACL) tears from a knee MRI exam. Previous works focused primarily on achieving better diagnostic accuracy but paid less attention to practical aspects such as explainability and model size. They mainly relied on ImageNet pre-trained 2D deep neural network backbones, such as AlexNet or ResNet, which are computationally expensive. Some of them tried to interpret the models using post-inference visualization tools, such as CAM or Grad-CAM, which lack in generating accurate heatmaps. Our work addresses the two limitations by understanding the characteristics of ACL tear diagnosis. We argue that the semantic features required for classifying ACL tears are locally confined and highly homogeneous. We harness the unique characteristics of the task by incorporating: 1) attention modules and Gaussian positional encoding to reinforce the seeking of local features; 2) squeeze modules and fewer convolutional filters to reflect the homogeneity of the features. As a result, our model is interpretable: our attention modules can precisely highlight the ACL region without any location information given to them. Our model is extremely lightweight: consisting of only 43 K trainable parameters and 7.1 G of Floating-point operations per second (FLOPs), that is 225 times smaller and 91 times lesser than the previous state-of-the-art, respectively. Our model is accurate: our model outperforms the previous state-of-the-art with the average ROC-AUC of 0.983 and 0.980 on the Chiba and Stanford knee datasets, respectively.
Young Seok Jeon, Kensuke Yoshino, Shigeo Hagiwara, Atsuya Watanabe, Swee Tian Quek, Hiroshi Yoshioka, Mengling Feng
IEEE J. Biomed. Health Informatics7
2019 Serial Heart Rate Variability Measures for Risk Prediction of Septic Patients in the Emergency Department
Calvin Chiew, Han Wang 0001, Marcus Eng Hock Ong, Ting Hway Wong, Zhixiong Koh, Nan Liu 0003, Mengling Feng
AMIA7
2017 Understanding vasopressor intervention and weaning: risk prediction in a public heterogeneous clinical time series database
abstract
BACKGROUND: The widespread adoption of electronic health records allows us to ask evidence-based questions about the need for and benefits of specific clinical interventions in critical-care settings across large populations. OBJECTIVE: We investigated the prediction of vasopressor administration and weaning in the intensive care unit. Vasopressors are commonly used to control hypotension, and changes in timing and dosage can have a large impact on patient outcomes. MATERIALS AND METHODS: We considered a cohort of 15 695 intensive care unit patients without orders for reduced care who were alive 30 days post-discharge. A switching-state autoregressive model (SSAM) was trained to predict the multidimensional physiological time series of patients before, during, and after vasopressor administration. The latent states from the SSAM were used as predictors of vasopressor administration and weaning. RESULTS: The unsupervised SSAM features were able to predict patient vasopressor administration and successful patient weaning. Features derived from the SSAM achieved areas under the receiver operating curve of 0.92, 0.88, and 0.71 for predicting ungapped vasopressor administration, gapped vasopressor administration, and vasopressor weaning, respectively. We also demonstrated many cases where our model predicted weaning well in advance of a successful wean. CONCLUSION: Models that used SSAM features increased performance on both predictive tasks. These improvements may reflect an underlying, and ultimately predictive, latent state detectable from the physiological time series.
Mike Wu, Marzyeh Ghassemi, Mengling Feng, Leo A. Celi, Peter Szolovits, Finale Doshi-Velez
J. Am. Medical Informatics Assoc.3
2016 Hypotension Risk Prediction via Sequential Contrast Patterns of ICU Blood Pressure
abstract
Acute hypotension is a significant risk factor for in-hospital mortality at intensive care units. Prolonged hypotension can cause tissue hypoperfusion, leading to cellular dysfunction and severe injuries to multiple organs. Prompt medical interventions are thus extremely important for dealing with acute hypotensive episodes (AHE). Population level prognostic scoring systems for risk stratification of patients are suboptimal in such scenarios. However, the design of an efficient risk prediction system can significantly help in the identification of critical care patients, who are at risk of developing an AHE within a future time span. Toward this objective, a pattern mining algorithm is employed to extract informative sequential contrast patterns from hemodynamic data, for the prediction of hypotensive episodes. The hypotensive and normotensive patient groups are extracted from the MIMIC-II critical care research database, following an appropriate clinical inclusion criteria. The proposed method consists of a data preprocessing step to convert the blood pressure time series into symbolic sequences, using a symbolic aggregate approximation algorithm. Then, distinguishing subsequences are identified using the sequential contrast mining algorithm. These subsequences are used to predict the occurrence of an AHE in a future time window separated by a user-defined gap interval. Results indicate that the method performs well in terms of the prediction performance as well as in the generation of sequential patterns of clinical significance. Hence, the novelty of sequential patterns is in their usefulness as potential physiological biomarkers for building optimal patient risk stratification systems and for further clinical investigation of interesting patterns in critical care patients.
Shameek Ghosh, Mengling Feng, Hung T. Nguyen 0001, Jinyan Li 0001
IEEE J. Biomed. Health Informatics2
2015 A Multivariate Timeseries Modeling Approach to Severity of Illness Assessment and Forecasting in ICU with Sparse, Heterogeneous Clinical Data
abstract
The ability to determine patient acuity (or severity of illness) has immediate practical use for clinicians. We evaluate the use of multivariate timeseries modeling with the multi-task Gaussian process (GP) models using noisy, incomplete, sparse, heterogeneous and unevenly-sampled clinical data, including both physiological signals and clinical notes. The learned multi-task GP (MTGP) hyperparameters are then used to assess and forecast patient acuity. Experiments were conducted with two real clinical data sets acquired from ICU patients: firstly, estimating cerebrovascular pressure reactivity, an important indicator of secondary damage for traumatic brain injury patients, by learning the interactions between intracranial pressure and mean arterial blood pressure signals, and secondly, mortality prediction using clinical progress notes. In both cases, MTGPs provided improved results: an MTGP model provided better results than single-task GP models for signal interpolation and forecasting (0.91 vs 0.69 RMSE), and the use of MTGP hyperparameters obtained improved results when used as additional classification features (0.812 vs 0.788 AUC).
Marzyeh Ghassemi, Marco A. F. Pimentel, Tristan Naumann, Thomas Brennan, David A. Clifton, Peter Szolovits, Mengling Feng
AAAI7
2015 Supporting Exploratory Hypothesis Testing and Analysis
abstract
Conventional hypothesis testing is carried out in a hypothesis-driven manner. A scientist must first formulate a hypothesis based on what he or she sees and then devise a variety of experiments to test it. Given the rapid growth of data, it has become virtually impossible for a person to manually inspect all data to find all of the interesting hypotheses for testing. In this article, we propose and develop a data-driven framework for automatic hypothesis testing and analysis. We define a hypothesis as a comparison between two or more subpopulations. We find subpopulations for comparison using frequent pattern mining techniques and then pair them up for statistical hypothesis testing. We also generate additional information for further analysis of the hypotheses that are deemed significant. The number of hypotheses generated can be very large, and many of them are very similar. We develop algorithms to remove redundant hypotheses and present a succinct set of significant hypotheses to users. We conducted a set of experiments to show the efficiency and effectiveness of the proposed algorithms. The results show that our system can help users (1) identify significant hypotheses efficiently, (2) isolate the reasons behind significant hypotheses efficiently, and (3) find confounding factors that form Simpson’s paradoxes with discovered significant hypotheses.
Guimei Liu, Haojun Zhang, Mengling Feng, Limsoon Wong, See-Kiong Ng
ACM Trans. Knowl. Discov. Data3
2014 Big Data for Critical Care with Cloud-based In-Memory Database
Mengling Feng, Mohammad M. Ghassemi, Thomas Brennan, John Ellenberger, Ishrar Hussain, Roger G. Mark
AMIA1
2014 Risk Prediction for Acute Hypotensive Patients by Using Gap Constrained Sequential Contrast Patterns
Shameek Ghosh, Mengling Feng, Hung T. Nguyen 0001, Jinyan Li 0001
AMIA2
2014 Management and analytic of biomedical big data with cloud-based in-memory database and dynamic querying: a hands-on experience with real-world data
abstract
Analyzing Biomedical Big Data (BBD) is computationally expensive due to high dimensionality and large data volume. Performance and scalability issues of traditional database management systems (DBMS) often limit the usage of more sophisticated and complex data queries and analytic models. Moreover, in the conventional setting, data management and analysis use separate software platforms. Exporting and importing large amounts of data across platforms require a significant amount of computational and I/O resources, as well as potentially putting sensitive data at a security risk. In this tutorial, the participants will learn the difference between in-memory DBMS and traditional DBMS through hands-on exercises using SAP's cloud-based HANA in-memory DBMS in conjunction with the Multi-parameter Intelligent Monitoring in Intensive Care (MIMIC) dataset. MIMIC is an open-access critical care EHR archive (over 4TB in size) and consists of structured, unstructured and waveform data. Furthermore, this tutorial will seek to educate the participants on how a combination of dynamic querying, and in-memory DBMS may enhance the management and analysis of complex clinical data.
Mengling Feng, Mohammad M. Ghassemi, Thomas Brennan, John Ellenberger, Ishrar Hussain, Roger G. Mark
KDD1
2013 An online approach for intracranial pressure forecasting based on signal decomposition and robust statistics
abstract
Intracranial pressure (ICP) is an important physiological signal for patients with traumatic brain injuries. Accurate ICP forecasting enables active and early interventions for more effective control of ICP levels. To achieve high accuracy, most existing methods require a high sampling rate (100 Hz), which is infeasible for online medical applications. Therefore, we propose an online ICP forecasting method requiring only low rate signal sampling (0.1 Hz). Our ARIMA based forecasting method applies empirical mode decomposition (EMD) to remove non-stationarities from the ICP signal, and robust estimation to mitigate the influence of motion induced artifacts. Experimental performance assessment with simulated and clinically collected data demonstrate that the proposed method is more accurate compared to previously proposed and standard methods.
Michael Muma, Mengling Feng, Abdelhak M. Zoubir
ICASSP3
2012 Artifact correction with robust statistics for non-stationary intracranial pressure signal monitoring
Mengling Feng, Liang Yu Loy, Kelvin Sim, Clifton Phua, Cuntai Guan
ICPR1
2012 Online ICP forecast for patients with traumatic brain injury
Mengling Feng, Liang Yu Loy, Zhuo Zhang 0001, Cuntai Guan
ICPR2
2012 AssocExplorer: an association rule visualization system for exploratory data analysis
abstract
We present a system called AssocExplorer to support exploratory data analysis via association rule visualization and exploration. AssocExplorer is designed by following the visual information-seeking mantra: overview first, zoom and filter, then details on demand. It effectively uses coloring to deliver information so that users can easily detect things that are interesting to them. If users find a rule interesting, they can explore related rules for further analysis, which allows users to find interesting phenomenon that are difficult to detect when rules are examined separately. Our system also allows users to compare rules and inspect rules with similar item composition but different statistics so that the key factors that contribute to the difference can be isolated.
Guimei Liu, Andre Suchitra, Haojun Zhang, Mengling Feng, See-Kiong Ng, Limsoon Wong
KDD4
2012 Response: an empirical comparison of several recent epistatic interaction detection methods
abstract
Abstract Contact: [email protected]
Yue Wang 0006, Guimei Liu, Mengling Feng, Limsoon Wong
Bioinform.3
2011 Towards exploratory hypothesis testing and analysis
abstract
Hypothesis testing is a well-established tool for scientific discovery. Conventional hypothesis testing is carried out in a hypothesis-driven manner. A scientist must first formulate a hypothesis based on his/her knowledge and experience, and then devise a variety of experiments to test it. Given the rapid growth of data, it has become virtually impossible for a person to manually inspect all the data to find all the interesting hypotheses for testing. In this paper, we propose and develop a data-driven system for automatic hypothesis testing and analysis. We define a hypothesis as a comparison between two or more sub-populations. We find sub-populations for comparison using frequent pattern mining techniques and then pair them up for statistical testing. We also generate additional information for further analysis of the hypotheses that are deemed significant. We conducted a set of experiments to show the efficiency of the proposed algorithms, and the usefulness of the generated hypotheses. The results show that our system can help users (1) identify significant hypotheses; (2) isolate the reasons behind significant hypotheses; and (3) find confounding factors that form Simpson's Paradoxes with discovered significant hypotheses.
Guimei Liu, Mengling Feng, Yue Wang 0006, Limsoon Wong, See-Kiong Ng, Tzia Liang Mah, Edmund Jon Deoon Lee
ICDE2
2011 An empirical comparison of several recent epistatic interaction detection methods
abstract
MOTIVATION: Many new methods have recently been proposed for detecting epistatic interactions in GWAS data. There is, however, no in-depth independent comparison of these methods yet. RESULTS: Five recent methods-TEAM, BOOST, SNPHarvester, SNPRuler and Screen and Clean (SC)-are evaluated here in terms of power, type-1 error rate, scalability and completeness. In terms of power, TEAM performs best on data with main effect and BOOST performs best on data without main effect. In terms of type-1 error rate, TEAM and BOOST have higher type-1 error rates than SNPRuler and SNPHarvester. SC does not control type-1 error rate well. In terms of scalability, we tested the five methods using a dataset with 100 000 SNPs on a 64 bit Ubuntu system, with Intel (R) Xeon(R) CPU 2.66 GHz, 16 GB memory. TEAM takes ~36 days to finish and SNPRuler reports heap allocation problems. BOOST scales up to 100 000 SNPs and the cost is much lower than that of TEAM. SC and SNPHarvester are the most scalable. In terms of completeness, we study how frequently the pruning techniques employed by these methods incorrectly prune away the most significant epistatic interactions. We find that, on average, 20% of datasets without main effect and 60% of datasets with main effect are pruned incorrectly by BOOST, SNPRuler and SNPHarvester. AVAILABILITY: The software for the five methods tested are available from the URLs below. TEAM: http://csbio.unc.edu/epistasis/download.php BOOST: http://ihome.ust.hk/~eeyang/papers.html. SNPHarvester: http://bioinformatics.ust.hk/SNPHarvester.html. SNPRuler: http://bioinformatics.ust.hk/SNPRuler.zip. Screen and Clean: http://wpicr.wpic.pitt.edu/WPICCompGen/. CONTACT: [email protected].
Yue Wang 0006, Guimei Liu, Mengling Feng, Limsoon Wong
Bioinform.3
2010 Efficiently Finding the Best Parameter for the Emerging Pattern-Based Classifier PCL
Thanh-Son Ngo, Mengling Feng, Guimei Liu, Limsoon Wong
PAKDD (1)2
2010 Pattern Space Maintenance for Data Updates and Interactive Mining
abstract
This article addresses the incremental and decremental maintenance of the frequent pattern space. We conduct an in‐depth investigation on how the frequent pattern space evolves under both incremental and decremental updates. Based on the evolution analysis, a new data structure, Generator‐Enumeration Tree (GE‐tree), is developed to facilitate the maintenance of the frequent pattern space. With the concept of GE‐tree, we propose two novel algorithms, Pattern Space Maintainer+ (PSM+) and Pattern Space Maintainer− (PSM−), for the incremental and decremental maintenance of frequent patterns. Experimental results demonstrate that the proposed algorithms, on average, outperform the representative state‐of‐the‐art methods by an order of magnitude.
Mengling Feng, Guozhu Dong, Jinyan Li 0001, Yap-Peng Tan, Limsoon Wong
Comput. Intell.1
2008 Negative Generator Border for Effective Pattern Maintenance
Mengling Feng, Jinyan Li 0001, Limsoon Wong, Yap-Peng Tan
ADMA1
2007 Evolution and Maintenance of Frequent Pattern Space When Transactions Are Removed
Mengling Feng, Guozhu Dong, Jinyan Li 0001, Yap-Peng Tan, Limsoon Wong
PAKDD1
2005 Relative risk and odds ratio: a data mining perspective
abstract
We are often interested to test whether a given cause has a given effect. If we cannot specify the nature of the factors involved, such tests are called model-free studies. There are two major strategies to demonstrate associations between risk factors (ie. patterns) and outcome phenotypes (ie. class labels). The first is that of prospective study designs, and the analysis is based on the concept of "relative risk": What fraction of the exposed (ie. has the pattern) or unexposed (ie. lacks the pattern) individuals have the phenotype (ie. the class label)? The second is that of retrospective designs, and the analysis is based on the concept of "odds ratio": The odds that a case has been exposed to a risk factor is compared to the odds for a case that has not been exposed. The efficient extraction of patterns that have good relative risk and/or odds ratio has not been previously studied in the data mining context. In this paper, we investigate such patterns. We show that this pattern space can be systematically stratified into plateaus of convex spaces based on their support levels. Exploiting convexity, we formulate a number of sound and complete algorithms to extract the most general and the most specific of such patterns at each support level. We compare these algorithms. We further demonstrate that the most efficient among these algorithms is able to mine these sophisticated patterns at a speed comparable to that of mining frequent closed patterns, which are patterns that satisfy considerably simpler conditions.
Haiquan Li, Jinyan Li 0001, Limsoon Wong, Mengling Feng, Yap-Peng Tan
PODS4
2004 Adaptive binarization method for document image analysis
abstract
This paper proposes an adaptive binarization method, based on the criterion of maximizing local contrast, for document image analysis. The proposed method has overcome, to a large extent, the general problems of poor quality document images, such as non-uniform illumination, undesirable shadows and random noise. It was tested against a variety of challenging images, and the experimental results are presented to show the effectiveness and superiority of the proposed method.
Mengling Feng, Yap-Peng Tan
ICME1