VLDB 2026 Research / reviewers in the wild / expert
Jinman Kim
dblp:43/2801
· DBLP profile ↗
134ranked-venue papers
8as first author
78since 2021 · last 2027
0000-0001-5960-1060ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 66 · 5 first-author · 36 since 2021Applied, interdisciplinary, general and emerging computing · 58 · 3 first-author · 30 since 2021Artificial intelligence and machine learning · 30 · 22 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Language-guided medical image segmentation with target-informed multi-level contrastive alignmentsabstractMedical image segmentation is a fundamental task in numerous medical applications. Recently, language-guided segmentation has shown promise in medical scenarios where textual clinical reports are readily available as semantic guidance. Clinical reports contain diagnostic information provided by clinicians, which can provide auxiliary textual semantics to guide segmentation. However, existing language-guided segmentation methods neglect the inherent pattern gaps between image and text modalities, resulting in sub-optimal visual-language integration. Contrastive learning is a well-recognized approach to align image-text patterns, but it has not been optimized for medical image segmentation, where clinically meaningful semantics are often concentrated in localized target regions rather than the entire image. In this study, we propose TMCA, a Target-informed Multi-level Contrastive Alignment framework to bridge image-text pattern gaps for medical language-guided segmentation. The core innovation is to reformulate image-text contrastive alignment from conventional instance-level matching to segmentation-oriented semantic matching, where image-text samples are aligned according to their segmentation targets rather than merely whether they come from the same patient. Specifically, TMCA enables target-informed image-text alignments and fine-grained textual guidance by introducing: (i) a target-sensitive semantic distance module that utilizes target information for more granular image-text alignment modeling, (ii) a multi-level contrastive alignment strategy that directs fine-grained textual guidance to multi-scale image details, and (iii) a language-guided target enhancement module that reinforces attention to critical regions based on the aligned image-text patterns. Extensive experiments on four public benchmarks, involving three medical imaging modalities with clinical reports, show that TMCA enabled superior performance over state-of-the-art language-guided medical segmentation methods. Mingjian Li, Mingyuan Meng, Shuchang Ye, Mingye Zou, Michael J. Fulham, Lei Bi 0001, Jinman Kim |
Expert Syst. Appl. | 7 |
| 2026 | SynTaskNet: A synergistic multi-task network for joint segmentation and classification of small anatomical structures in ultrasound imaging
Abdulrhman H. Al-Jebrni, Saba Ghazanfar Ali, Bin Sheng 0001, Huating Li, Xiao Lin 0012, Ping Li 0016, Younhyun Jung, Jinman Kim, Lixin Jiang |
Comput. Vis. Image Underst. | 8 |
| 2026 | A novel language model for predicting serious adverse event results in clinical trials from their prospective registrationsabstractOBJECTIVES: With accurate estimates of expected safety results, clinical trials could be better designed and monitored. We evaluated methods for predicting serious adverse event (SAE) results in clinical trials using information only from their registrations prior to the trial. MATERIAL AND METHODS: We analyzed 22,107 clinical trials from ClinicalTrials.gov alongside their summary results. We developed a classifier predicting significant differences in the proportion of participants with SAEs (area under the receiver operating characteristic curve; AUC) across experimental and control arms, and a regression model predicting the proportion of participants with SAEs in the control arms (root mean squared error; RMSE). A transfer learning approach using pretrained language models (e.g., ClinicalT5, BioBERT) was used to build a prediction model. To maintain semantic representation in long trial texts, a sliding window method was applied. RESULTS: The best performing model (BioBERT + Transformer + MLP) had 85.1% AUC when predicting which trial arm had a higher proportion of SAEs. When predicting SAE proportion in the control arm, the same model achieved RMSE of 18.8%. The sliding window approach consistently outperformed direct comparisons; the average absolute AUC increase was 3.5%, and absolute RMSE reduction was 1.58%. Classifiers built using language model-based representations consistently outperformed baseline models trained solely on structured data, with an average AUC difference of 11.60%. DISCUSSION: Summary results data from ClinicalTrials.gov remains underutilized. Predicted results of publicly reported trials provides an opportunity to identify discrepancies between expected and reported safety results. Qixuan Hu, Xumou Zhang, Jinman Kim, Florence T. Bourgeois, Adam G. Dunn |
J. Biomed. Informatics | 3 |
| 2026 | Improving lesion segmentation in medical images by global and regional feature compensationabstract• Dual-feature compensation framework to improve medical image segmentation of challenging lesions. • Addresses the loss of detailed global features caused by downsampling. • Uses SSL residual maps to provide additional pixel-level feature representations. • Applies patch-based cross-attention to integrate SSL residual maps, targeting likely lesion regions. Automated lesion segmentation of medical images has made tremendous improvements in recent years due to deep learning advancements. However, accurately capturing fine-grained global and regional feature representations remains a challenge. Many existing methods achieve suboptimal performance in complex lesion segmentation due to information loss during typical downsampling operations and insufficient capture of either regional or global features. To address these issues, we propose the Global and Regional Compensation Segmentation Framework (GRCSF), which introduces two key innovations: the Global Compensation Unit (GCU) and the Region Compensation Unit (RCU). The proposed GCU addresses resolution loss in the U-shaped backbone by preserving global contextual features and fine-grained details during multiscale downsampling. Meanwhile, the RCU introduces a self-supervised learning (SSL) residual map generated by Masked Autoencoders (MAE), obtained as pixel-wise differences between reconstructed and original images, to highlight regions with potential lesions. These SSL residual maps guide precise lesion localization and segmentation through a patch-based cross-attention mechanism that integrates regional spatial and pixel-level features. Additionally, the RCU incorporates patch-level importance scoring to enhance feature fusion by leveraging global spatial information from the backbone. Experiments on three publicly available medical image segmentation datasets, including brain stroke lesion, lung tumor and coronary artery calcification datasets, demonstrate that our GRCSF outperforms state-of-the-art methods, confirming its effectiveness across diverse lesion types and its potential as a generalizable lesion segmentation solution. Jean Y. H. Yang, Jinman Kim |
Pattern Recognit. | 4 |
| 2026 | Hierarchical Deep Decision Tree-Based Network for Odontogenic Cystic Lesion Classification in CBCT ImagesabstractOdontogenic cystic lesions (OCLs) are complex jaw abnormalities that require a precise diagnosis of the disease for treatment. Visual OCL diagnosis is commonly based on reviewing cone-beam computed tomography (CBCT) to identify morpho-pathological features associated with specific lesion types in a hierarchical manner. Current state-of-the-art methods focus on extracting features from the image without any guidance beyond the lesion diagnosis, and do not fully leverage the hierarchical relationship between the lesion diagnosis and morphological features. In this study, we propose a hierarchical deep decision tree network (H2DT-Net) with three modules: a deep decision tree-based hierarchical learning module (DHLM) to leverage inter-categorical relationships; a feature category embedding module (FCEM) to capture representations from both diagnostic and morpho-pathological domains and support the DHLM; and a lesion localised attention module (LLAM) to facilitate the feature extraction process by generating lesion-focused attention maps. Evaluated on 289 CBCT images, H2DT-Net achieved state-of-the-art performance in OCL classification. We further demonstrate that our method is effective in clinical settings, where it outperformed six maxillofacial clinicians in diagnostic assessment. Zimo Huang, Hao Wang 0143, Eduardo Delamare, Shengfu Huang, Lei Bi 0001, Jinman Kim |
IEEE J. Biomed. Health Informatics | 7 |
| 2026 | Dynamic Traceback Learning for Medical Report GenerationabstractAutomated medical report generation has demonstrated the potential to significantly reduce the workload associated with time-consuming medical reporting. Recent generative representation learning methods have shown promise in integrating vision and language modalities for medical report generation. However, when trained end-to-end and applied directly to medical image-to-text generation, they face two significant challenges: i) difficulty in accurately capturing subtle yet crucial pathological details, and ii) reliance on both visual and textual inputs during inference, leading to performance degradation in zero-shot inference when only images are available. To address these challenges, this study proposes a novel multimodal dynamic traceback learning framework (DTrace)1. Specifically, we introduce a traceback mechanism to supervise the semantic validity of generated content and a dynamic learning strategy to adapt to various proportions of image and text input, enabling text generation without strong reliance on the input from both modalities during inference. The learning of cross-modal knowledge is enhanced by supervising the model to recover masked semantic information from a complementary counterpart. Extensive experiments conducted on two benchmark datasets, IU-Xray and MIMIC-CXR, demonstrate that the proposedDTraceframework outperforms state-of-the-art methods for medical report generation. Shuchang Ye, Mingyuan Meng, Mingjian Li, David Dagan Feng, Usman Naseem, Jinman Kim |
IEEE Trans. Multim. | 6 |
| 2025 | Enhancing Zero-Shot Learning of Pathology Vision-Language Foundation Models in Tumour Malignancy RecognitionabstractVision-Language Foundation Models (VLFMs) demonstrate promise in zero-shot learning through joint visual-textual representations. However, in histopathology image analysis, their effectiveness is limited by weak image-text alignment due to coarse-grained textual descriptions that fail to capture critical fine-grained visual details. This misalignment introduces semantic noise and imprecision in zero-shot retrieval, hindering the identification of relevant cases and degrading downstream classification. To address this, we introduce Retrieval-based De-noising Causal Language Modelling (RDCLM), a novel framework that refines noisy retrieval outputs from pathology VLFMs. RDCLM constructs a pathology-specific knowledge base of fine-grained, discriminative tumour malignancy descriptions using a large language model (LLM). Given a query histopathology image, a pathology VLFM retrieves candidate descriptions from this knowledge base. Our de-noising module, leveraging a frozen language model, integrates visual features with these retrieved texts, filtering irrelevant content and enhancing semantic alignment. This significantly improves retrieval precision (by an average of 10% across datasets) and enables more accurate zero-shot image classification. To further bolster performance and generalization, we propose two retrieval augmentation strategies: Retrieval Negatives Replacement (RNR) and Description-wise Shuffling (DS). Extensive evaluations across four histopathology cancer datasets demonstrate that RDCLM significantly outperforms state-of-the-art methods in both zero-shot image-text retrieval and malignancy classification, achieving an average improvement of 12.7% in F1-score and 9.6% in accuracy over the second-best competitor. These results highlight the importance of retrieval de-noising for advancing VLFM-based zero-shot learning in histopathology. Our code available at: https://github.com/xw18958/RDCLM Usman Naseem, Jinman Kim |
ECAI | 3 |
| 2025 | Alleviating Textual Reliance in Medical Language-Guided Segmentation via Prototype-Driven Semantic ApproximationabstractMedical language-guided segmentation, integrating textual clinical reports as auxiliary guidance to enhance image segmentation, has demonstrated significant improvements over unimodal approaches. However, its inherent reliance on paired image-text input, which we refer to as ``textual reliance", presents two fundamental limitations: 1) many medical segmentation datasets lack paired reports, leaving a substantial portion of image-only data underutilized for training; and 2) inference is limited to retrospective analysis of cases with paired reports, limiting its applicability in most clinical scenarios where segmentation typically precedes reporting. To address these limitations, we propose ProLearn, the first Prototype-driven Learning framework for language-guided segmentation that fundamentally alleviates textual reliance. At its core, we introduce a novel Prototype-driven Semantic Approximation (PSA) module to enable approximation of semantic guidance from textual input. PSA initializes a discrete and compact prototype space by distilling segmentation-relevant semantics from textual reports. Once initialized, it supports a query-and-respond mechanism which approximates semantic guidance for images without textual input, thereby alleviating textual reliance. Extensive experiments on QaTa-COV19, MosMedData+ and Kvasir-SEG demonstrate that ProLearn outperforms state-of-the-art language-guided methods when limited text is available. Shuchang Ye, Usman Naseem, Mingyuan Meng, Jinman Kim |
ICCV | 4 |
| 2025 | A Generative Adversarial Network for Upsampling of Direct Volume Rendering ImagesabstractAbstract Direct volume rendering (DVR) is an important tool for scientific and medical imaging visualization. Modern GPU acceleration has made DVR more accessible; however, the production of high‐quality rendered images with high frame rates is computationally expensive. We propose a deep learning method with a reduced computational demand. We leveraged a conditional generative adversarial network (cGAN) to upsample DVR images (a rendered scene), with a reduced sampling rate to obtain similar visual quality to that of a fully sampled method. Our dvrGAN is combined with a colour‐based loss function that is optimized for DVR images where different structures such as skin, bone, etc. are distinguished by assigning them distinct colours. The loss function highlights the structural differences between images, by examining pixel‐level colour, and thus helps identify, for instance, small bones in the limbs that may not be evident with reduced sampling rates. We evaluated our method in DVR of human computed tomography (CT) and CT angiography (CTA) volumes. Our method retained image quality and reduced computation time when compared to fully sampled methods and outperformed existing state‐of‐the‐art upsampling methods. Ge Jin 0001, Younhyun Jung, Michael J. Fulham, David Dagan Feng, Jinman Kim |
Comput. Graph. Forum | 5 |
| 2025 | AutoFuse: Automatic fusion networks for deformable medical image registration
Mingyuan Meng, Michael J. Fulham, David Dagan Feng, Lei Bi 0001, Jinman Kim |
Pattern Recognit. | 5 |
| 2025 | Head-Mounted Displays in Context-Aware Systems for Open Surgery: A State-of-the-Art ReviewabstractSurgical context-aware systems (SCAS), which leverage real-time data and analysis from the operating room to inform surgical activities, can be enhanced through the integration of head-mounted displays (HMDs). Rather than user-agnostic data derived from conventional, and often static, external sensors, HMD-based SCAS relies on dynamic user-centric sensing of the surgical context. The analyzed context-aware information is then augmented directly into a user's field of view via augmented reality (AR) to directly improve their task and decision-making capability. This state-of-the-art review complements previous reviews by exploring the advancement of HMD- based SCAS, including their development and impact on enhancing situational awareness and surgical outcomes in the operating room. The survey demonstrates that this technology can mitigate risks associated with gaps in surgical expertise, increase procedural efficiency, and improve patient outcomes. We also highlight key limitations still to be addressed by the research community, including improving prediction accuracy, robustly handling data heterogeneity, and reducing system latency. Mingxiao Tu, Hoijoon Jung, Jinman Kim, Andre Kyme |
IEEE J. Biomed. Health Informatics | 3 |
| 2025 | Explicit Abnormality Extraction for Unsupervised Motion Artifact Reduction in Magnetic Resonance ImagingabstractMotion artifacts compromise the quality of magnetic resonance imaging (MRI) and pose challenges to achieving diagnostic outcomes and image-guided therapies. In recent years, supervised deep learning approaches have emerged as successful solutions for motion artifact reduction (MAR). One disadvantage of these methods is their dependency on acquiring paired sets of motion artifact-corrupted (MA-corrupted) and motion artifact-free (MA-free) MR images for training purposes. Obtaining such image pairs is difficult and therefore limits the application of supervised training. In this paper, we propose a novel UNsupervised Abnormality Extraction Network (UNAEN) to alleviate this problem. Our network is capable of working with unpaired MA-corrupted and MA-free images. It converts the MA-corrupted images to MA-reduced images by extracting abnormalities from the MA-corrupted images using a proposed artifact extractor, which intercepts the residual artifact maps from the MA-corrupted MR images explicitly, and a reconstructor to restore the original input from the MA-reduced images. The performance of UNAEN was assessed by experimenting with various publicly available MRI datasets and comparing them with state-of-the-art methods. The quantitative evaluation demonstrates the superiority of UNAEN over alternative MAR methods and visually exhibits fewer residual artifacts. Our results substantiate the potential of UNAEN as a promising solution applicable in real-world clinical environments, with the capability to enhance diagnostic accuracy and facilitate image-guided therapies. Hao Li 0034, Zhengmin Kong, Tao Huang 0008, Euijoon Ahn, Zhihan Lyu, Jinman Kim, David Dagan Feng |
IEEE J. Biomed. Health Informatics | 8 |
| 2025 | Enhancing Medical Vision-Language Contrastive Learning via Inter-Matching Relation ModelingabstractMedical image representations can be learned through medical vision-language contrastive learning (mVLCL) where medical imaging reports are used as weak supervision through image-text alignment. These learned image representations can be transferred to and benefit various downstream medical vision tasks such as disease classification and segmentation. Recent mVLCL methods attempt to align image sub-regions and the report keywords as local-matchings. However, these methods aggregate all local-matchings via simple pooling operations while ignoring the inherent relations between them. These methods therefore fail to reason between local-matchings that are semantically related, e.g., local-matchings that correspond to the disease word and the location word (semantic-relations), and also fail to differentiate such clinically important local-matchings from others that correspond to less meaningful words, e.g., conjunction words (importance-relations). Hence, we propose a mVLCL method that models the inter-matching relations between local-matchings via a relation-enhanced contrastive learning framework (RECLF). In RECLF, we introduce a semantic-relation reasoning module (SRM) and an importance-relation reasoning module (IRM) to enable more fine-grained report supervision for image representation learning. We evaluated our method using six public benchmark datasets on four downstream tasks, including segmentation, zero-shot classification, linear classification, and cross-modal retrieval. Our results demonstrated the superiority of our RECLF over the state-of-the-art mVLCL methods with consistent improvements across single-modal and cross-modal tasks. These results suggest that our RECLF, by modeling the inter-matching relations, can learn improved medical image representations with better generalization capabilities. Mingjian Li, Mingyuan Meng, Michael J. Fulham, David Dagan Feng, Lei Bi 0001, Jinman Kim |
IEEE Trans. Medical Imaging | 6 |
| 2025 | EGDNet: an efficient glomerular detection network for multiple anomalous pathological feature in glomerulonephritis
Saba Ghazanfar Ali, Ping Li 0016, Huating Li, Po Yang 0001, Younhyun Jung, Harry Qin, Jinman Kim, Bin Sheng 0001 |
Vis. Comput. | 8 |
| 2025 | Attention-driven visual emphasis for medical volumetric image visualizationabstractAbstract Direct volume rendering (DVR) is a commonly utilized technique for three-dimensional visualization of volumetric medical images. A key goal of DVR is to enable users to visually emphasize regions of interest (ROIs) which may be occluded by other structures. Conventional methods for ROIs visual emphasis require extensive user involvement for the adjustment of rendering parameters to reduce the occlusion, dependent on the user’s viewing direction. Several works have been proposed to automatically preserve the view of the ROIs by eliminating the occluding structures of lower importance in a view-dependent manner. However, they require pre-segmentation labeling and manual importance assignment on the images. An alternative to ROIs segmentation is to use ‘saliency’ to identify important regions. This however lacks semantic information and thus leads to the inclusion of false positive regions. In this study, we propose an attention-driven visual emphasis method for volumetric medical image visualization. We developed a deep learning attention model, termed as focused-class attention map (F-CAM), trained with only image-wise labels for automated ROIs localization and importance estimation. Our F-CAM transfers the semantic information from the classification task for use in the localization of ROIs, with a focus on small ROIs that characterize medical images. Additionally, we propose an attention compositing module that integrates the generated attention map with transfer function within the DVR pipeline to automate the view-dependent visual emphasis of the ROIs. We demonstrate the superiority of our method compared to existing methods on a multi-modality PET-CT dataset and an MRI dataset. Mingjian Li, Younhyun Jung, Shaoli Song, Jinman Kim |
Vis. Comput. | 4 |
| 2025 | Deep contour attention learning for scleral deformation from OCT images
Hao Chen 0011, Yupeng Xu, Huating Li, Yuan Xie 0006, David Dagan Feng, Jinman Kim, Lei Bi 0001, Xiangui He, Bin Sheng 0001 |
Vis. Comput. | 8 |
| 2025 | HRDC challenge: a public benchmark for hypertension and hypertensive retinopathy classification from fundus images
Xiangning Wang, Zhouyu Guan, An-ran Ran, Tingyao Li, Zheyuan Wang, Xinming Shu, Jinyang Xie, Shichang Liu, Guanyu Xing, Julio Silva-Rodríguez, Riadh Kobbi, Ping Li 0016, Tingli Chen, Lei Bi 0001, Jinman Kim, Weiping Jia, Huating Li, Harry Qin, Ping Zhang 0016, Ching Yu Cheng, Pheng-Ann Heng, Tien Yin Wong, Carol Y. Cheung, Nadia Magnenat-Thalmann, Bin Sheng 0001 |
Vis. Comput. | 18 |
| 2024 | Semi-Mamba: Improving Medical Image Segmentation via Semi-Automatic Mamba Network
Lei Bi 0001, Yige Peng, David Dagan Feng, Jinman Kim |
CGI (3) | 4 |
| 2024 | Mixed Reality Hologram Slicer (mxdR-HS): A Markerless Tangible User Interface for Interactive Holographic Medical Volume Visualization
Hoijoon Jung, Younhyun Jung, Michael J. Fulham, Jinman Kim |
CGI (3) | 4 |
| 2024 | Correlation-aware Coarse-to-fine MLPs for Deformable Medical Image RegistrationabstractDeformable image registration is a fundamental step for medical image analysis. Recently, transformers have been used for registration and outperformed Convolutional Neural Networks (CNNs). Transformers can capture long-range dependence among image features, which have been shown beneficial for registration. However, due to the high computation/memory loads of self-attention, transformers are typically used at downsampled feature resolutions and cannot capture fine-grained long-range dependence at the full image resolution. This limits deformable registration as it necessitates precise dense correspondence between each image pixel. Multi-layer Perceptrons (MLPs) without self-attention are efficient in computation/memory usage, enabling the feasibility of capturing fine-grained long-range dependence at full resolution. Nevertheless, MLPs have not been extensively explored for image registration and are lacking the consideration of inductive bias crucial for medical registration tasks. In this study, we propose the first correlation-aware MLP-based registration network (CorrMLP) for deformable medical image registration. Our CorrMLP introduces a correlation-aware multi-window MLP block in a novel coarse-to-fine registration architecture, which captures fine-grained multi-range dependence to perform correlation-aware coarse-to-fine registration. Extensive experiments with seven public medical datasets show that our CorrMLP outperforms state-of-the-art deformable registration methods. Mingyuan Meng, David Dagan Feng, Lei Bi 0001, Jinman Kim |
CVPR | 4 |
| 2024 | 3DPX: Progressive 2D-to-3D Oral Image Reconstruction with Hybrid MLP-CNN Networks
Xiaoshuang Li, Mingyuan Meng, Zimo Huang, Lei Bi 0001, Eduardo Delamare, David Dagan Feng, Bin Sheng 0001, Jinman Kim |
MICCAI (7) | 8 |
| 2024 | Enabling Text-Free Inference in Language-Guided Segmentation of Chest X-Rays via Self-guidance
Shuchang Ye, Mingyuan Meng, Mingjian Li, David Dagan Feng, Jinman Kim |
MICCAI (8) | 5 |
| 2024 | Vaccine Misinformation Detection in X using Cooperative Multimodal FrameworkabstractIdentifying social media posts that spread vaccine misinformation can inform emerging public health risks and aid in designing effective communication interventions. Existing studies, while promising, often rely on single user posts, potentially leading to flawed conclusions. This highlights the necessity to model users' historical posts for a comprehensive understanding of their stance towards vaccines. However, users' historical posts may contain a diverse range of content that adds noise and leads to low performance. To address this gap, in this study, we present VaxMine, a cooperative multi-agent reinforcement learning method that automatically selects relevant textual and visual content from a user's posts, reducing noise. To evaluate the performance of the proposed method, we create and release a new dataset of 2,072 users with historical posts due to the unavailability of publicly available datasets. The experimental results show that our approach outperforms state-of-the-art methods with an F1-Score of 0.94 (an absolute increase of 13%), demonstrating that extracting relevant content from users' historical posts and understanding both modalities are essential to detecting anti-vaccine users on social media. We further analyze the robustness and generalizability of VaxMine, showing that extracting relevant textual and visual content from a user's posts improves performance. We conclude with a discussion of the practical implications of our study by explaining how computational methods used in surveillance can benefit from our work, with flow-on effects on the design of health communication interventions to counter vaccine misinformation on social media. Usman Naseem, Adam G. Dunn, Matloob Khushi, Jinman Kim |
ACM Multimedia | 4 |
| 2024 | A Linguistic Grounding-Infused Contrastive Learning Approach for Health Mention Classification on Social MediaabstractSocial media users use disease and symptoms words in different ways, including describing their personal health experiences figuratively or in other general discussions. The health mention classification (HMC) task aims to separate how people use terms, which is important in public health applications. Existing HMC studies address this problem using pretrained language models (PLMs). However, the remaining gaps in the area include the need for linguistic grounding, the requirement for large volumes of labelled data, and that solutions are often only tested on Twitter or Reddit, which provides limited evidence of the transportability of models. To address these gaps, we propose a novel method that uses a transformer-based PLM to obtain a contextual representation of target (disease or symptom) terms coupled with a contrastive loss to establish a larger gap between target terms' literal and figurative uses using linguistic theories. We introduce the use of a simple and effective approach for harvesting candidate instances from the broad corpus and generalising the proposed method using self-training to address the label scarcity challenge. Our experiments on publicly available health-mention datasets from Twitter (HMC2019) and Reddit (RHMD) demonstrate that our method outperforms the state-of-the-art HMC methods on both datasets for the HMC task. We further analyse the transferability and generalisability of our method and conclude with a discussion on the empirical and ethical considerations of our study. Usman Naseem, Jinman Kim, Matloob Khushi, Adam G. Dunn |
WSDM | 2 |
| 2024 | A Transfer Function Design for Medical Volume Data Using a Knowledge Database Based on Deep Image and Primitive Intensity Profile Features Retrieval
Younhyun Jung, Jim Kong, Bin Sheng 0001, Jinman Kim |
J. Comput. Sci. Technol. | 4 |
| 2024 | A multi-resolution self-supervised learning framework for semantic segmentation in histopathologyabstractModern whole slide imaging technique together with supervised deep learning approaches have been advancing the field of histopathology, enabling accurate analysis of tissues. These approaches use whole slide images (WSIs) at various resolutions, utilising low-resolution WSIs to identify regions of interest in the tissue and high-resolution for detailed analysis of cellular structures. Due to the labour-intensive process of annotating gigapixels WSIs, accurate analysis of WSIs remains challenging for supervised approaches. Self-supervised learning (SSL) has emerged as an approach to build efficient and robust models using unlabelled data. It has been successfully used to pre-train models to learn meaningful image features which are then fine-tuned with downstream tasks for improved performance compared to training models from scratch. Yet, existing SSL methods optimised for WSI are unable to leverage the multi-resolutions and instead, work only in an individual resolution neglecting the hierarchical structure of multi-resolution inputs. This limitation prevents from the effective utilisation of complementary information between different resolutions, hampering discriminative WSI representation learning. In this paper we propose a Multi-resolution SSL Framework for WSI semantic segmentation (MSF-WSI) that effectively learns histopathological features. Our MSF-WSI learns complementary information from multiple WSI resolutions during the pre-training stage; this contrasts with existing works that only learn between the resolutions at the fine-tuning stage. Our pre-training initialises the model with a comprehensive understanding of multi-resolution features which can lead to improved performance in the subsequent tasks. To achieve this, we introduced a novel Context-Target Fusion Module (CTFM) and a masked jigsaw pretext task to facilitate the learning of multi-resolution features. Additionally, we designed Dense SimSiam Learning (DSL) strategy to maximise the similarities of image features from early model layers to enable discriminative learned representations. We evaluated our method using three public datasets on breast and liver cancer segmentation tasks. Our experiment results demonstrated that our MSF-WSI surpassed the accuracy of other state-of-the-art methods in downstream fine-tuning and semi-supervised settings. Hao Wang 0143, Euijoon Ahn, Jinman Kim |
Pattern Recognit. | 3 |
| 2024 | ClarityDiffuseNet: Enhancing fundus image quality under black shadows with diffusion model-based research
Jiadi Dong, Tianwei Qian, Yuxian Jiang, Lei Bi 0001, Jinman Kim, Lisheng Wang |
Pattern Recognit. Lett. | 5 |
| 2024 | Hybrid Text Representation for Explainable Suicide Risk Identification on Social MediaabstractSocial media data that characterize users can provide mental health signals, including suicide risks. Existing methods for suicide risk identification on social media have demonstrated promising results; however, the limitation of existing methods is that they are unable to capture low-and high-level features with complex structured data on social media and are incapable of explaining the predicted labels. Explainable models are more useful when translated, so we aimed to evaluate a novel method that would produce explainable models. This article presents a hybrid text representation method that integrates word and document-level text representations to explain suicide risk identification on social media. The proposed method is then fed to a transformer-based encoder with ordinal classification to determine suicide risk. Our results show that our method outperforms state-of-the-art baselines with an FScore of 0.79 (an absolute increase of 15%) on a public suicide dataset. Our method shows that an explainable model can perform at a comparable level to the best nonexplainable models but has advantages if translated for use in clinical and public health practice. Usman Naseem, Matloob Khushi, Jinman Kim, Adam G. Dunn |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2024 | K-PathVQA: Knowledge-Aware Multimodal Representation for Pathology Visual Question AnsweringabstractPathology imaging is routinely used to detect the underlying effects and causes of diseases or injuries. Pathology visual question answering (PathVQA) aims to enable computers to answer questions about clinical visual findings from pathology images. Prior work on PathVQA has focused on directly analyzing the image content using conventional pretrained encoders without utilizing relevant external information when the image content is inadequate. In this paper, we present a knowledge-driven PathVQA (K-PathVQA), which uses a medical knowledge graph (KG) from a complementary external structured knowledge base to infer answers for the PathVQA task. K-PathVQA improves the question representation with external medical knowledge and then aggregates vision, language, and knowledge embeddings to learn a joint knowledge-image-question representation. Our experiments using a publicly available PathVQA dataset showed that our K-PathVQA outperformed the best baseline method with an increase of 4.15% in accuracy for the overall task, an increase of 4.40% in open-ended question type and an absolute increase of 1.03% in closed-ended question types. Ablation testing shows the impact of each of the contributions. Generalizability of the method is demonstrated with a separate medical VQA dataset. Usman Naseem, Matloob Khushi, Adam G. Dunn, Jinman Kim |
IEEE J. Biomed. Health Informatics | 4 |
| 2024 | SparseVoxNet: 3-D Object Recognition With Sparsely Aggregation of 3-D Dense BlocksabstractAutomatic recognition of 3-D objects in a 3-D model by convolutional neural network (CNN) methods has been successfully applied to various tasks, e.g., robotics and augmented reality. Three-dimensional object recognition is mainly performed by analyzing the object using multi-view images, depth images, graphs, or volumetric data. In some cases, using volumetric data provides the most promising results. However, existing recognition techniques on volumetric data have many drawbacks, such as losing object details on converting points to voxels and the large size of the input volume data that leads to substantial 3-D CNNs. Using point clouds could also provide very promising results; however, point-cloud-based methods typically need sparse data entry and time-consuming training stages. Thus, using volumetric could be a more efficient and flexible recognizer for our special case in the School of Medicine, Shanghai Jiao Tong University. In this article, we propose a novel solution to 3-D object recognition from volumetric data using a combination of three compact CNN models, low-cost SparseNet, and feature representation technique. We achieve an optimized network by estimating extra geometrical information comprising the surface normal and curvature into two separated neural networks. These two models provide supplementary information to each voxel data that consequently improve the results. The primary network model takes advantage of all the predicted features and uses these features in Random Forest (RF) for recognition purposes. Our method outperforms other methods in training speed in our experiments and provides an accurate result as good as the state-of-the-art. Ahmad Karambakhsh, Bin Sheng 0001, Ping Li 0016, Huating Li, Jinman Kim, Younhyun Jung, C. L. Philip Chen |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Deep choroid layer segmentation using hybrid features extraction from OCT images
Saleha Masood, Saba Ghazanfar Ali, Xiangning Wang, Afifa Masood, Ping Li 0016, Huating Li, Younhyun Jung, Bin Sheng 0001, Jinman Kim |
Vis. Comput. | 9 |
| 2024 | Importance-aware 3D volume visualization for medical content-based image retrieval-a preliminary studyabstractA medical content-based image retrieval (CBIR) system is designed to retrieve images from large imaging repositories that are visually similar to a user′s query image. CBIR is widely used in evidence- based diagnosis, teaching, and research. Although the retrieval accuracy has largely improved, there has been limited development toward visualizing important image features that indicate the similarity of retrieved images. Despite the prevalence of3D volumetric data in medical imaging such as computed tomography (CT), current CBIR systems still rely on 2D cross-sectional views for the visualization of retrieved images. Such 2D visualization requires users to browse through the image stacks to confirm the similarity of the retrieved images and often involves mental reconstruction of 3D information, including the size, shape, and spatial relations of multiple structures. This process is time-consuming and reliant on users’ experience. In this study, we proposed an importance-aware 3D volume visualization method. The rendering parameters were automatically optimized to maximize the visibility of important structures that were detected and prioritized in the retrieval process. We then integrated the proposed visualization into a CBIR system, thereby complementing the 2D cross-sectional views for relevance feedback and further analyses. Our preliminary results demonstrate that 3D visualization can provide additional information using multimodal positron emission tomography and computed tomography (PET- CT) images of a non-small cell lung cancer dataset. Mingjian Li, Younhyun Jung, Michael J. Fulham, Jinman Kim |
Virtual Real. Intell. Hardw. | 4 |
| 2023 | Challenges and Constraints in Deformation-Based Medical Mesh Representation
Ge Jin 0001, Younhyun Jung, Jinman Kim |
CGI (4) | 3 |
| 2023 | Merging-Diverging Hybrid Transformer Networks for Survival Prediction in Head and Neck Cancer
Mingyuan Meng, Lei Bi 0001, Michael J. Fulham, David Dagan Feng, Jinman Kim |
MICCAI (6) | 5 |
| 2023 | Non-iterative Coarse-to-Fine Transformer Networks for Joint Affine and Deformable Image Registration
Mingyuan Meng, Lei Bi 0001, Michael J. Fulham, David Dagan Feng, Jinman Kim |
MICCAI (10) | 5 |
| 2023 | Improving Automatic Fetal Biometry Measurement with Swoosh Activation Function
Shijia Zhou, Euijoon Ahn, Hao Wang 0143, Ann Quinton, Narelle Kennedy, Pradeeba Sridar, Ralph Nanan, Jinman Kim |
MICCAI (7) | 8 |
| 2023 | A Multimodal Framework for the Identification of Vaccine Critical Memes on TwitterabstractMemes can be a useful way to spread information because they are funny, easy to share, and can spread quickly and reach further than other forms. With increased interest in COVID-19 vaccines, vaccination-related memes have grown in number and reach. Memes analysis can be difficult because they use sarcasm and often require contextual understanding. Previous research has shown promising results but could be improved by capturing global and local representations within memes to model contextual information. Further, the limited public availability of annotated vaccine critical memes datasets limit our ability to design computational methods to help design targeted interventions and boost vaccine uptake. To address these gaps, we present VaxMeme, which consists of 10,244 manually labelled memes. With VaxMeme, we propose a new multimodal framework designed to improve the memes' representation by learning the global and local representations of memes. The improved memes' representations are then fed to an attentive representation learning module to capture contextual information for classification using an optimised loss function. Experimental results show that our framework outperformed state-of-the-art methods with an F1-Score of 84.2%. We further analyse the transferability and generalisability of our framework and show that understanding both modalities is important to identify vaccine critical memes on Twitter. Finally, we discuss how understanding memes can be useful in designing shareable vaccination promotion, myth debunking memes and monitoring their uptake on social media platforms. Usman Naseem, Jinman Kim, Matloob Khushi, Adam G. Dunn |
WSDM | 2 |
| 2023 | EditorialabstractThis special issue is the third issue dedicated to the best papers of the second Call of the CGI 2022 conference. In 2022, CGI (Computer Graphics International) was organized by MIRALab at the University of Geneva. It was supported by the Computer Graphics Society. The conference was very successful and attracted more than 150 online participants through Zoom. Twenty-six papers were selected from more than 100 papers submitted to the second call of the conference. This issue contains the last eight full papers. Nine papers were already published in the 33.5 issue and nine papers in the 33.6 issue. All papers have been reviewed by two or three reviewers of the CGI 2022 Program Committee, revised according to the reviewers' comments, and checked and reviewed again by the CAVW Editorial Board. Data-driven based double-layer bicycle simulation model by Tianlu Mao, Zhong Fang, Qinyuan Yan, and Zhaoqi Wang, from Academy of Sciences, Beijing, and Ruoyu Meng and Shaohua Liu, all in China. This special issue is edited by the program-co-chairs of CGI2022 conference. Jinman Kim, George Papagiannakis, Bin Sheng 0001, Daniel Thalmann |
Comput. Animat. Virtual Worlds | 1 |
| 2023 | RHMD: A Real-World Dataset for Health Mention Classification on RedditabstractPeople on social media share their thoughts and experiences using diseases and symptoms words other than to mention their health, which can introduce biases in data-driven public health applications. For the advancement of HMC research, in this study, we present a Reddit health mention dataset (RHMD), a new dataset of multi-domain Reddit data for the HMC. RHMD is composed of 10015 manually annotated Reddit posts that include 15 common disease or symptom terms and are labeled with four labels: personal health mentions (HMs), nonpersonal HMs, figurative HMs, and hyperbolic HMs. Empirical evaluation using recently proposed methods demonstrates the challenge of labeling user-generated text across these four types. Contributions to this work include the public release of a robustly annotated Reddit dataset (RHMD) for HM tasks and a comprehensive performance analysis of baseline methods. We expect the release of the dataset, and the evaluations will help facilitate the development of new methods for detecting HMs in the user-generated text. The dataset is available athttps://github.com/usmaann/RHMD-Health-Mention-Dataset. Usman Naseem, Matloob Khushi, Jinman Kim, Adam G. Dunn |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2023 | Unsupervised Landmark Detection-Based Spatiotemporal Motion Estimation for 4-D Dynamic Medical ImagesabstractMotion estimation is a fundamental step in dynamic medical image processing for the assessment of target organ anatomy and function. However, existing image-based motion estimation methods, which optimize the motion field by evaluating the local image similarity, are prone to produce implausible estimation, especially in the presence of large motion. In addition, the correct anatomical topology is difficult to be preserved as the image global context is not well incorporated into motion estimation. In this study, we provide a novel motion estimation framework of dense-sparse-dense (DSD), which comprises two stages. In the first stage, we process the raw dense image to extract sparse landmarks to represent the target organ's anatomical topology, and discard the redundant information that is unnecessary for motion estimation. For this purpose, we introduce an unsupervised 3-D landmark detection network to extract spatially sparse but representative landmarks for the target organ's motion estimation. In the second stage, we derive the sparse motion displacement from the extracted sparse landmarks of two images of different time points. Then, we present a motion reconstruction network to construct the motion field by projecting the sparse landmarks' displacement back into the dense image domain. Furthermore, we employ the estimated motion field from our two-stage DSD framework as initialization and boost the motion estimation quality in light-weight yet effective iterative optimization. We evaluate our method on two dynamic medical imaging tasks to model cardiac motion and lung respiratory motion, respectively. Our method has produced superior motion estimation accuracy compared to the existing comparative methods. Besides, the extensive experimental results demonstrate that our solution can extract well-representative anatomical landmarks without any requirement of manual annotation. Our code is publicly available online: https://github.com/yyguo-sjtu/DSD-3D-Unsupervised-Landmark-Detection-Based-Motion-Estimation. Yuyu Guo 0002, Lei Bi 0001, Dongming Wei, Liyun Chen, Zhengbin Zhu, David Dagan Feng, Ruiyan Zhang, Qian Wang 0001, Jinman Kim |
IEEE Trans. Cybern. | 9 |
| 2023 | Vision-Language Transformer for Interpretable Pathology Visual Question AnsweringabstractPathology visual question answering (PathVQA) attempts to answer a medical question posed by pathology images. Despite its great potential in healthcare, it is not widely adopted because it requires interactions on both the image (vision) and question (language) to generate an answer. Existing methods focused on treating vision and language features independently, which were unable to capture the high and low-level interactions that are required for VQA. Further, these methods failed to offer capabilities to interpret the retrieved answers, which are obscure to humans where the models' interpretability to justify the retrieved answers has remained largely unexplored. Motivated by these limitations, we introduce a vision-language transformer that embeds vision (images) and language (questions) features for an interpretable PathVQA. We present an interpretable transformer-based Path-VQA (TraP-VQA), where we embed transformers' encoder layers with vision and language features extracted using pre-trained CNN and domain-specific language model (LM), respectively. A decoder layer is then embedded to upsample the encoded features for the final prediction for PathVQA. Our experiments showed that our TraP-VQA outperformed the state-of-the-art comparative methods with public PathVQA dataset. Our experiments validated the robustness of our model on another medical VQA dataset, and the ablation study demonstrated the capability of our integrated transformer-based vision-language model for PathVQA. Finally, we present the visualization results of both text and images, which explain the reason for a retrieved answer in PathVQA. Usman Naseem, Matloob Khushi, Jinman Kim |
IEEE J. Biomed. Health Informatics | 3 |
| 2023 | A Shortened Model for Logan Reference Plot Implemented via the Self-Supervised Neural Network for Parametric PET ImagingabstractDynamic PET imaging provides superior physiological information than conventional static PET imaging. However, the dynamic information is gained at the cost of a long scanning protocol; this limits the clinical application of dynamic PET imaging. We developed a modified Logan reference plot model to shorten the acquisition procedure in dynamic PET imaging by omitting the early-time information necessary for the conventional reference Logan model. The proposed model is accurate theoretically, but the straightforward approach raises the sampling problem in implementation and results in noisy parametric images. We then designed a self-supervised convolutional neural network to increase the noise performance of parametric imaging, with dynamic images of only a single subject for training. The proposed method was validated via simulated and real dynamic [Formula: see text]-fallypride PET data. Results showed that it accurately estimated the distribution volume ratio (DVR) in dynamic PET with a shortened scanning protocol, e.g., 20 minutes, where the estimations were comparable with those obtained from a standard dynamic PET study of 120 minutes of acquisition. Further comparisons illustrated that our method outperformed the shortened Logan model implemented with Gaussian filtering, regularization, BM4D and the 4D deep image prior methods in terms of the trade-off between bias and variance. Since the proposed method uses data acquired in a short period of time upon the equilibrium, it has the potential to add clinical values by providing both DVR and Standard Uptake Value (SUV) simultaneously. It thus promotes clinical applications of dynamic PET studies when neuronal receptor functions are studied. Wenxiang Ding, Qiaoqiao Ding, Kewei Chen 0001, Miao Zhang 0038, David Dagan Feng, Lei Bi 0001, Jinman Kim, Qiu Huang |
IEEE Trans. Medical Imaging | 8 |
| 2023 | SThy-Net: a feature fusion-enhanced dense-branched modules network for small thyroid nodule classification from ultrasound images
Abdulrhman H. Al-Jebrni, Saba Ghazanfar Ali, Huating Li, Xiao Lin 0012, Ping Li 0016, Younhyun Jung, Jinman Kim, David Dagan Feng, Bin Sheng 0001, Lixin Jiang |
Vis. Comput. | 7 |
| 2022 | Self-Supervised Representation Learning Framework for Remote Physiological Measurement Using Spatiotemporal Augmentation LossabstractRecent advances in supervised deep learning methods are enabling remote measurements of photoplethysmography-based physiological signals using facial videos. The performance of these supervised methods, however, are dependent on the availability of large labelled data. Contrastive learning as a self-supervised method has recently achieved state-of-the-art performances in learning representative data features by maximising mutual information between different augmented views. However, existing data augmentation techniques for contrastive learning are not designed to learn physiological signals from videos and often fail when there are complicated noise and subtle and periodic colour/shape variations between video frames. To address these problems, we present a novel self-supervised spatiotemporal learning framework for remote physiological signal representation learning, where there is a lack of labelled training data. Firstly, we propose a landmark-based spatial augmentation that splits the face into several informative parts based on the Shafer’s dichromatic reflection model to characterise subtle skin colour fluctuations. We also formulate a sparsity-based temporal augmentation exploiting Nyquist–Shannon sampling theorem to effectively capture periodic temporal changes by modelling physiological signal features. Furthermore, we introduce a constrained spatiotemporal loss which generates pseudo-labels for augmented video clips. It is used to regulate the training process and handle complicated noise. We evaluated our framework on 3 public datasets and demonstrated superior performances than other self-supervised methods and achieved competitive accuracy compared to the state-of-the-art supervised methods. Code is available at https://github.com/Dylan-H-Wang/SLF-RPM. Hao Wang 0143, Euijoon Ahn, Jinman Kim |
AAAI | 3 |
| 2022 | Non-iterative Coarse-to-Fine Registration Based on Single-Pass Deep Cumulative Learning
Mingyuan Meng, Lei Bi 0001, David Dagan Feng, Jinman Kim |
MICCAI (6) | 4 |
| 2022 | Early Identification of Depression Severity Levels on Reddit Using Ordinal ClassificationabstractUser-generated text on social media is a promising avenue for public health surveillance and has been actively explored for its feasibility in the early identification of depression. Existing methods in the identification of depression have shown promising results; however, these methods were all focused on treating the identification as a binary classification problem. To date, there has been little effort towards identifying users’ depression severity level and disregard the inherent ordinal nature across these fine-grain levels. This paper aims to make early identification of depression severity levels on social media data. To accomplish this, we built a new dataset based on the inherent ordinal nature over depression severity levels using clinical depression standards on Reddit posts. The posts were classified into 4 depression severity levels covering the clinical depression standards on social media. Accordingly, we reformulate the early identification of depression as an ordinal classification task over clinical depression standards such as Beck’s Depression Inventory and the Depressive Disorder Annotation scheme to identify depression severity levels. With these, we propose a hierarchical attention method optimized to factor in the increasing depression severity levels through a soft probability distribution. We experimented using two datasets (a public dataset having more than one post from each user and our built dataset with a single user post) using real-world Reddit posts that have been classified according to questionnaires built by clinical experts and demonstrated that our method outperforms state-of-the-art models. Finally, we conclude by analyzing the minimum number of posts required to identify depression severity level followed by a discussion of empirical and practical considerations of our study. Usman Naseem, Adam G. Dunn, Jinman Kim, Matloob Khushi |
WWW | 3 |
| 2022 | Identification of Disease or Symptom terms in Reddit to Improve Health Mention ClassificationabstractIn a user-generated text such as on social media platforms and online forums, people often use disease or symptom terms in ways other than to describe their health. In data-driven public health surveillance, the health mention classification (HMC) task aims to identify posts where users are discussing health conditions rather than using disease and symptom terms for other reasons. Existing computational research typically only studies health mentions in Twitter, with limited coverage of disease or symptom terms, ignore user behavior information, and other ways people use disease or symptom terms. To advance the HMC research, we present a Reddit health mention dataset (RHMD), a new dataset of multi-domain Reddit data for the HMC. RHMD consists of 10,015 manually labeled Reddit posts that mention 15 common disease or symptom terms and are annotated with four labels: namely personal health mentions, non-personal health mentions, figurative health mentions, and hyperbolic health mentions. With RHMD, we propose HMCNET that combines a target keyword (disease or symptom term) identification and user behavior hierarchically to improve HMC. Experimental results demonstrate that the proposed approach outperforms state-of-the-art methods with an F1-Score of 0.75 (an increase of 11% over the state-of-the-art) and shows that our new dataset poses a strong challenge to the existing HMC methods. Usman Naseem, Jinman Kim, Matloob Khushi, Adam G. Dunn |
WWW | 2 |
| 2022 | Deep multi-scale resemblance network for the sub-class differentiation of adrenal masses on computed tomography images
Lei Bi 0001, Jinman Kim, Tingwei Su, Michael J. Fulham, David Dagan Feng, Guang Ning |
Artif. Intell. Medicine | 2 |
| 2022 | Benchmarking for biomedical natural language processing tasks with a domain specific ALBERTabstractBACKGROUND: The abundance of biomedical text data coupled with advances in natural language processing (NLP) is resulting in novel biomedical NLP (BioNLP) applications. These NLP applications, or tasks, are reliant on the availability of domain-specific language models (LMs) that are trained on a massive amount of data. Most of the existing domain-specific LMs adopted bidirectional encoder representations from transformers (BERT) architecture which has limitations, and their generalizability is unproven as there is an absence of baseline results among common BioNLP tasks. RESULTS: We present 8 variants of BioALBERT, a domain-specific adaptation of a lite bidirectional encoder representations from transformers (ALBERT), trained on biomedical (PubMed and PubMed Central) and clinical (MIMIC-III) corpora and fine-tuned for 6 different tasks across 20 benchmark datasets. Experiments show that a large variant of BioALBERT trained on PubMed outperforms the state-of-the-art on named-entity recognition (+ 11.09% BLURB score improvement), relation extraction (+ 0.80% BLURB score), sentence similarity (+ 1.05% BLURB score), document classification (+ 0.62% F1-score), and question answering (+ 2.83% BLURB score). It represents a new state-of-the-art in 5 out of 6 benchmark BioNLP tasks. CONCLUSIONS: The large variant of BioALBERT trained on PubMed achieved a higher BLURB score than previous state-of-the-art models on 5 of the 6 benchmark BioNLP tasks. Depending on the task, 5 different variants of BioALBERT outperformed previous state-of-the-art models on 17 of the 20 benchmark datasets, showing that our model is robust and generalizable in the common BioNLP tasks. We have made BioALBERT freely available which will help the BioNLP community avoid computational cost of training and establish a new set of baselines for future efforts across a broad range of BioNLP tasks. Usman Naseem, Adam G. Dunn, Matloob Khushi, Jinman Kim |
BMC Bioinform. | 4 |
| 2022 | Clinical applications of machine learning in predicting 3D shapes of the human body: a systematic reviewabstractBACKGROUND: Predicting morphological changes to anatomical structures from 3D shapes such as blood vessels or appearance of the face is a growing interest to clinicians. Machine learning (ML) has had great success driving predictions in 2D, however, methods suitable for 3D shapes are unclear and the use cases unknown. OBJECTIVE AND METHODS: This systematic review aims to identify the clinical implementation of 3D shape prediction and ML workflows. Ovid-MEDLINE, Embase, Scopus and Web of Science were searched until 28th March 2022. RESULTS: 13,754 articles were identified, with 12 studies meeting final inclusion criteria. These studies involved prediction of the face, head, aorta, forearm, and breast, with most aiming to visualize shape changes after surgical interventions. ML algorithms identified were regressions (67%), artificial neural networks (25%), and principal component analysis (8%). Meta-analysis was not feasible due to the heterogeneity of the outcomes. CONCLUSION: 3D shape prediction is a nascent but growing area of research in medicine. This review revealed the feasibility of predicting 3D shapes using ML clinically, which could play an important role for clinician-patient visualization and communication. However, all studies were early phase and there were inconsistent language and reporting. Future work could develop guidelines for publication and promote open sharing of source code. Joyce Zhanzi Wang, Jonathon Lillia, Ashnil Kumar, Paula Bray, Jinman Kim, Joshua Burns, Tegan L. Cheng |
BMC Bioinform. | 5 |
| 2022 | Experimental protocol designed to employ Nd: YAG laser surgery for anterior chamber glaucoma detection via UBMabstractAbstract Angle closure glaucoma leads to fluid deposition in eye, and intraocular pressure occurs that damage the optic nerve, causes blindness and vision loss. Anterior chamber (AC) evaluation is imperative for determining the risk of angle‐closure. Previously, techniques were dependent on either Pentacam–Scheimpflug that interprets poor visual information, anterior segment optical coherence tomography is injurious to intercede opaque optical structures. Therefore, in this paper, an experimental protocol is designed for detailed disease analysis based on IBM SPSS statistics via ultrasound biomicroscopy which is superior in evaluating deep structures; first, the affected parameter for AC is analysed, and afterwards the direction that needs laser surgery is explored. Experiments are conducted on large‐scale clinical studies from an affiliated hospital in Shanghai, China. The dataset comprised 600 AC images in five directions of 60 subjects. The mean with standard deviation for anterior open distance is 0.158790.096779 mm, 0.158630.081435 mm, and anterior chamber angle is 18.74908.0315, 18.74108.3889 for left and right eye respectively. It is found that anterior chamber angle in the downside of the AC is wider than the upside. However, this decision is partly based on the narrowest part of the angle to widen the depth of the direction and eliminate pupil block. Saba Ghazanfar Ali, Riaz Ali, Bin Sheng 0001, Huating Li, Po Yang 0001, Ping Li 0016, Younhyun Jung, Ping Lu 0008, Jinman Kim |
IET Image Process. | 11 |
| 2022 | Special issue on computer graphics international 2022 part 1abstractThis special issue is dedicated to the best papers of the second Call of the CGI 2022 conference. This year CGI (Computer Graphics International) was organized by MIRALab at the University of Geneva. It was supported by the Computer Graphics Society. The conference was very successful and attracted more than 150 online participants through Zoom. Twenty-six papers were selected from more than 100 papers submitted to the second call of the conference. This issue contains nine full papers. The other selected papers will be published in the two next issues respectively 33.6 and 34.1. All papers have been reviewed by two or three reviewers of the CGI 2022 Program Committee, revised according to the reviewers' comments, and checked and reviewed again by the CAVW Editorial Board. The first paper in this issue is the recipient of the CAVW Best Paper Award given at CGI 2022 by the Award Committee chaired by Professor Nadia Magnenat Thalmann, MIRALab-University of Geneva, Switzerland, and Professor Constantine Stephanidis, ICS—FORTH, Greece. This awarded paper is: Simulation of collective pursuit-evasion behavior with runtime situational awareness by Zhenjing Yu, Tan Junyin, and Sheng Li, from Peking University, China. Jinman Kim, George Papagiannakis, Bin Sheng 0001, Daniel Thalmann |
Comput. Animat. Virtual Worlds | 1 |
| 2022 | EditorialabstractThis special issue is the second issue dedicated to the best papers of the second Call of the CGI 2022 conference. This year CGI (Computer Graphics International) was organized by MIRALab at the University of Geneva. It was supported by the Computer Graphics Society. The conference was very successful and attracted more than 150 online participants through Zoom. Twenty-six papers were selected from more than 100 papers submitted to the second call of the conference. This issue contains nine full papers. Nine papers were already published in the 33.5 issue and the last papers will appear in the issue 34.1. All papers have been reviewed by two or three reviewers of the CGI 2022 Program Committee, revised according to the reviewers' comments, and checked and reviewed again by the CAVW Editorial Board. This special issue is edited by the program co-chairs of CGI2022 conference: Jinman Kim, Sydney University, Australia, George Papagiannakis, University of Crete, Greece, Bin Sheng, Shanghai Jiao Tong University, China, Daniel Thalmann, EPFL, Switzerland. Jinman Kim, George Papagiannakis, Bin Sheng 0001, Daniel Thalmann |
Comput. Animat. Virtual Worlds | 1 |
| 2022 | Hyper-fusion network for semi-automatic segmentation of skin lesions
Lei Bi 0001, Michael J. Fulham, Jinman Kim |
Medical Image Anal. | 3 |
| 2022 | Deep Cognitive Gate: Resembling Human Cognition for Saliency DetectionabstractSaliency detection by human refers to the ability to identify pertinent information using our perceptive and cognitive capabilities. While human perception is attracted by visual stimuli, our cognitive capability is derived from the inspiration of constructing concepts of reasoning. Saliency detection has gained intensive interest with the aim of resembling human 'perceptual' system. However, saliency related to human 'cognition', particularly the analysis of complex salient regions ('cogitating' process), is yet to be fully exploited. We propose to resemble human cognition, coupled with human perception, to improve saliency detection. We recognize saliency in three phases ('Seeing' - 'Perceiving' - 'Cogitating), mimicking human's perceptive and cognitive thinking of an image. In our method, 'Seeing' phase is related to human perception, and we formulate the 'Perceiving' and 'Cogitating' phases related to the human cognition systems via deep neural networks (DNNs) to construct a new module (Cognitive Gate) that enhances the DNN features for saliency detection. To the best of our knowledge, this is the first work that established DNNs to resemble human cognition for saliency detection. In our experiments, our approach outperformed 17 benchmarking DNN methods on six well-recognized datasets, demonstrating that resembling human cognition improves saliency detection. Ke Yan 0005, Xiuying Wang 0001, Jinman Kim, Wangmeng Zuo, David Dagan Feng |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | An attention-enhanced cross-task network to analyse lung nodule attributes in CT images
Xiaohang Fu, Lei Bi 0001, Ashnil Kumar, Michael J. Fulham, Jinman Kim |
Pattern Recognit. | 5 |
| 2022 | ECSU-Net: An Embedded Clustering Sliced U-Net Coupled With Fusing Strategy for Efficient Intervertebral Disc Segmentation and ClassificationabstractAutomatic vertebra segmentation from computed tomography (CT) image is the very first and a decisive stage in vertebra analysis for computer-based spinal diagnosis and therapy support system. However, automatic segmentation of vertebra remains challenging due to several reasons, including anatomic complexity of spine, unclear boundaries of the vertebrae associated with spongy and soft bones. Based on 2D U-Net, we have proposed an Embedded Clustering Sliced U-Net (ECSU-Net). ECSU-Net comprises of three modules named segmentation, intervertebral disc extraction (IDE) and fusion. The segmentation module follows an instance embedding clustering approach, where our three sliced sub-nets use axis of CT images to generate a coarse 2D segmentation along with embedding space with the same size of the input slices. Our IDE module is designed to classify vertebra and find the inter-space between two slices of segmented spine. Our fusion module takes the coarse segmentation (2D) and outputs the refined 3D results of vertebra. A novel adaptive discriminative loss (ADL) function is introduced to train the embedding space for clustering. In the fusion strategy, three modules are integrated via a learnable weight control component, which adaptively sets their contribution. We have evaluated classical and deep learning methods on Spineweb dataset-2. ECSU-Net has provided comparable performance to previous neural network based algorithms achieving the best segmentation dice score of 95.60% and classification accuracy of 96.20%, while taking less time and computation resources. Anam Nazir, Muhammad Nadeem Cheema, Bin Sheng 0001, Ping Li 0016, Huating Li, Guangtao Xue, Harry Qin, Jinman Kim, David Dagan Feng |
IEEE Trans. Image Process. | 8 |
| 2022 | DeepMTS: Deep Multi-Task Learning for Survival Prediction in Patients With Advanced Nasopharyngeal Carcinoma Using Pretreatment PET/CTabstractNasopharyngeal Carcinoma (NPC) is a malignant epithelial cancer arising from the nasopharynx. Survival prediction is a major concern for NPC patients, as it provides early prognostic information to plan treatments. Recently, deep survival models based on deep learning have demonstrated the potential to outperform traditional radiomics-based survival prediction models. Deep survival models usually use image patches covering the whole target regions (e.g., nasopharynx for NPC) or containing only segmented tumor regions as the input. However, the models using the whole target regions will also include non-relevant background information, while the models using segmented tumor regions will disregard potentially prognostic information existing out of primary tumors (e.g., local lymph node metastasis and adjacent tissue invasion). In this study, we propose a 3D end-to-end Deep Multi-Task Survival model (DeepMTS) for joint survival prediction and tumor segmentation in advanced NPC from pretreatment PET/CT. Our novelty is the introduction of a hard-sharing segmentation backbone to guide the extraction of local features related to the primary tumors, which reduces the interference from non-relevant background information. In addition, we also introduce a cascaded survival network to capture the prognostic information existing out of primary tumors and further leverage the global tumor information (e.g., tumor size, shape, and locations) derived from the segmentation backbone. Our experiments with two clinical datasets demonstrate that our DeepMTS can consistently outperform traditional radiomics-based survival prediction models and existing deep survival models. Mingyuan Meng, Bingxin Gu, Lei Bi 0001, Shaoli Song, David Dagan Feng, Jinman Kim |
IEEE J. Biomed. Health Informatics | 6 |
| 2022 | Improving Breast Tumor Segmentation in PET via Attentive Transformation Based NormalizationabstractPositron Emission Tomography (PET) has become a preferred imaging modality for cancer diagnosis, radiotherapy planning, and treatment responses monitoring. Accurate and automatic tumor segmentation is the fundamental requirement for these clinical applications. Deep convolutional neural networks have become the state-of-the-art in PET tumor segmentation. The normalization process is one of the key components for accelerating network training and improving the performance of the network. However, existing normalization methods either introduce batch noise into the instance PET image by calculating statistics on batch level or introduce background noise into every single pixel by sharing the same learnable parameters spatially. In this paper, we proposed an attentive transformation (AT)-based normalization method for PET tumor segmentation. We exploit the distinguishability of breast tumor in PET images and dynamically generate dedicated and pixel-dependent learnable parameters in normalization via the transformation on a combination of channel-wise and spatial-wise attentive responses. The attentive learnable parameters allow to re-calibrate features pixel-by-pixel to focus on the high-uptake area while attenuating the background noise of PET images. Our experimental results on two real clinical datasets show that the AT-based normalization method improves breast tumor segmentation performance when compared with the existing normalization methods. Xiaoya Qiao, Chunjuan Jiang, Panli Li, Yuan Yuan 0022, Qinglong Zeng, Lei Bi 0001, Shaoli Song, Jinman Kim, David Dagan Feng, Qiu Huang |
IEEE J. Biomed. Health Informatics | 8 |
| 2022 | Graph-Based Intercategory and Intermodality Network for Multilabel Classification and Melanoma Diagnosis of Skin Lesions in Dermoscopy and Clinical ImagesabstractThe identification of melanoma involves an integrated analysis of skin lesion images acquired using clinical and dermoscopy modalities. Dermoscopic images provide a detailed view of the subsurface visual structures that supplement the macroscopic details from clinical images. Visual melanoma diagnosis is commonly based on the 7-point visual category checklist (7PC), which involves identifying specific characteristics of skin lesions. The 7PC contains intrinsic relationships between categories that can aid classification, such as shared features, correlations, and the contributions of categories towards diagnosis. Manual classification is subjective and prone to intra- and interobserver variability. This presents an opportunity for automated methods to aid in diagnostic decision support. Current state-of-the-art methods focus on a single image modality (either clinical or dermoscopy) and ignore information from the other, or do not fully leverage the complementary information from both modalities. Furthermore, there is not a method to exploit the 'intercategory' relationships in the 7PC. In this study, we address these issues by proposing a graph-based intercategory and intermodality network (GIIN) with two modules. A graph-based relational module (GRM) leverages intercategorical relations, intermodal relations, and prioritises the visual structure details from dermoscopy by encoding category representations in a graph network. The category embedding learning module (CELM) captures representations that are specialised for each category and support the GRM. We show that our modules are effective at enhancing classification performance using three public datasets (7PC, ISIC 2017, and ISIC 2018), and that our method outperforms state-of-the-art methods at classifying the 7PC categories and diagnosis. Xiaohang Fu, Lei Bi 0001, Ashnil Kumar, Michael J. Fulham, Jinman Kim |
IEEE Trans. Medical Imaging | 5 |
| 2022 | Real-time spatial normalization for dynamic gesture classification
Sofiane Zeghoud, Saba Ghazanfar Ali, Egemen Ertugrul, Aouaidjia Kamel, Bin Sheng 0001, Ping Li 0016, Xiaoyu Chi, Jinman Kim, Lijuan Mao |
Vis. Comput. | 8 |
| 2022 | Computer graphics for metaverseabstractCGI is one of the oldest international conferences in Computer Graphics in the world.It is the official conference of the Computer Graphics Society (CGS), a long-standing international computer graphics organization.CGI conference has been held annually in many different countries across the world and has gained a reputation as one of the key conferences for researchers and practitioners to share their achievements and discover the latest advances in Computer Graphics.With the change in the form of networking and intelligence in industry, manufacturing and all aspects of society, and the development of technology, we are aware of the increasingly obvious trend of evolution of intelligence in human society.Among them, metaverse is increasingly becoming a hot spot for research in various industries and has broad application prospects.It absorbs the results of the information revolution, the Internet revolution, the artificial intelligence revolution, and the virtual reality technology revolution including VR, AR, MR, and especially game engines, showing mankind the possibility of building a holographic digital world parallel to the traditional physical world.The core of the metaverse lies in the hosting of virtual assets and virtual identities.Unlike traditional games, users can experience different content, make different friends, create their own creations, and perform a series of virtual activities in the metaverse.With the popularity of smart terminals and the rise of applications such as e-commerce/short videos/games, "metaverse" has become an inevitable trend in the development of digital society.In a broad sense, the "metaverse" is a virtual space-time consisting of a series of augmented reality (AR), virtual reality (VR) and the Internet; in a narrow sense, the "metaverse" is a virtual world parallel to the real world.By wearing a helmet and headset device, one can enter a three-dimensional world constructed by computer simulation through a terminal connection".The new mode of "virtual reality" presentation and scene interaction for scene visualization will be more conducive to better visual effects and interactive operations in the digital world.The metaverse becomes the best track and new growth point for AI applications because of its huge imagination, close social attention and rich landing scenes, while AI and related arithmetic, big data and other technical fields are the technical base for the metaverse to become a kind of concrete expression in the future.Overall, with the further development of human technology and the improvement of hardware level, it becomes possible for humans to build a "meta" world.This year, CGI 2022 is still online as the pandemic prevents many researchers to come to Geneva.The conference CGI is organized from September 12 to September 16, 2022, by MIRALab at the Computer Research Centre (CUI) of the University of Geneva, in Switzerland.All presentations are online.In addition to the Visual Computer journal published by Springer, and the CAVW journal (Computer Animation and Virtual Worlds) published by Wiley, we have also included the twenty-three accepted papers in the VRIH journal (Virtual Reality and Intelligent Hardware journal published by Science Press).This special issue is composed of the six papers related to the topic of metaverse from these twenty-three accepted papers. Nadia Magnenat-Thalmann, Jinman Kim, George Papagiannakis, Daniel Thalmann, Bin Sheng 0001 |
Virtual Real. Intell. Hardw. | 2 |
| 2021 | Dynamic Shadow Synthesis Using Silhouette Edge Optimization
Saba Ghazanfar Ali, Bin Sheng 0001, Ping Li 0016, Xiaoyu Chi, Jinman Kim, Lijuan Mao |
CGI | 7 |
| 2021 | A Classification Network for Ocular Diseases Based on Structure Feature and Visual Attention
Yupeng Xu, Bin Sheng 0001, Lei Bi 0001, Jinman Kim, Xiangui He |
CGI | 6 |
| 2021 | Multi-Stream Fusion Network for Multi-Distortion Image Super-Resolution
Yupeng Xu, Bin Sheng 0001, Ping Li 0016, Lei Bi 0001, Jinman Kim, Xiangui He |
CGI | 6 |
| 2021 | Classifying vaccine sentiment tweets by modelling domain-specific representation and commonsense knowledge into context-aware attentive GRUabstractVaccines are an important public health measure, but vaccine hesitancy and refusal can create clusters of low vaccine coverage and reduce the effectiveness of vaccination programs. Social media provides an opportunity to estimate emerging risks to vaccine acceptance by including geographical location and detailing vaccine-related concerns. Methods for classifying social media posts, such as vaccine-related tweets, use language models (LMs) trained on general domain text. However, challenges to measuring vaccine sentiment at scale arise from the absence of tonal stress and gestural cues and may not always have additional information about the user, e.g., past tweets or social connections. Another challenge in LMs is the lack of ‘commonsense’ knowledge that are apparent in users' metadata, i.e., emoticons, positive and negative words etc. In this study, to classify vaccine sentiment tweets with limited information, we present a novel end-to-end framework consisting of interconnected components that use domain-specific LM trained on vaccine-related tweets and models commonsense knowledge into a bidirectional gated recurrent network (CK-BiGRU) with context-aware attention. We further leverage syntactical, user metadata and sentiment information to capture the sentiment of a tweet. We experimented using two popular vaccine-related Twitter datasets and demonstrate that our proposed approach outperforms state-of-the-art models in identifying pro-vaccine, anti-vaccine and neutral tweets. Usman Naseem, Matloob Khushi, Jinman Kim, Adam G. Dunn |
IJCNN | 3 |
| 2021 | BioALBERT: A Simple and Effective Pre-trained Language Model for Biomedical Named Entity RecognitionabstractIn recent years, with the growing amount of biomedical documents, coupled with advancement in natural language processing algorithms, the research on biomedical named entity recognition (BioNER) has increased exponentially. However, BioNER research is challenging as NER in the biomedical domain are: (i) often restricted due to limited amount of training data, (ii) an entity can refer to multiple types and concepts depending on its context and, (iii) heavy reliance on acronyms that are sub-domain specific. Existing BioNER approaches often neglect these issues and directly adopt the state-of-the-art (SOTA) models trained in general corpora, which often yields unsatisfactory results. We propose biomedical ALBERT (A Lite Bidirectional Encoder Representations from Transformers for Biomedical Text Mining) - bioALBERT - an effective domain-specific pre-trained language model trained on a huge biomedical corpus designed to capture biomedical context-dependent NER. We adopted a self-supervised loss function used in ALBERT that targets modelling inter-sentence coherence to better learn context-dependent representations and incorporated parameter reduction strategies to minimise memory usage and enhance the training time in BioNER. In our experiments, BioALBERT outperformed comparative SOTA BioNER models on 8 biomedical NER benchmark datasets with 4 different entity types. The performance is increased for; (i) disease type corpora by 7.47% (NCBI- disease) and 10.63% (BC5CDR-disease); (ii) drug-chem type corpora by 4.61 % (BC5CDR-Chem) and 3.89% (BC4CHEMD); (iii) gene-protein type corpora by 12.25% (BC2GM) and 6.42% (JNLPBA); and (iv) species type corpora by 6.19% (LINNAEUS) and 23.71 % (Species-800) is observed which leads to a state-of-the-art results. The performance of a proposed model on four different biomedical entity types shows that our model is robust and generalisable in recognising biomedical entities in text. Usman Naseem, Matloob Khushi, Vinay Reddy, Sakthivel Rajendran, Muhammad Imran Razzak, Jinman Kim |
IJCNN | 6 |
| 2021 | A Spatial Guided Self-supervised Clustering Network for Medical Image Segmentation
Euijoon Ahn, David Dagan Feng, Jinman Kim |
MICCAI (1) | 3 |
| 2021 | High-parallelism Inception-like Spiking Neural Networks for Unsupervised Feature Learning
Mingyuan Meng, Lei Bi 0001, Jinman Kim, Shanlin Xiao, Zhiyi Yu |
Neurocomputing | 4 |
| 2021 | Unsupervised brain tumor segmentation using a symmetric-driven adversarial network
Xinheng Wu, Lei Bi 0001, Michael J. Fulham, David Dagan Feng, Luping Zhou, Jinman Kim |
Neurocomputing | 6 |
| 2021 | AFF-Dehazing: Attention-based feature fusion network for low-light image DehazingabstractAbstract Images captured in haze conditions, especially at nighttime with low light, often suffer from degraded visibility, contrasts, and vividness, which makes it difficult to carry out the following vision tasks. In this article, we propose an attention‐based feature fusion network (AFF‐Dehazing) for low‐light image dehazing. Our method decomposes the low‐light image dehazing into two task‐independent streams containing four modules: image dehazing module, low‐light feature extractor module, feature fusion module, and image restoration module. The basic block of these modules is the proposed attention‐based residual dense block. Since the dual‐branch are used, AFF‐Dehazing can avoid learning the mixed degradation all‐in‐one and enhance the details of low‐light haze images. Extensive experiments show that our method surpasses previous state‐of‐the‐art image dehazing methods and low‐light enhancement methods by a very large margin both quantitatively and qualitatively. Yu Zhou 0066, Bin Sheng 0001, Ping Li 0016, Jinman Kim, Enhua Wu |
Comput. Animat. Virtual Worlds | 5 |
| 2021 | COVIDSenti: A Large-Scale Benchmark Twitter Data Set for COVID-19 Sentiment AnalysisabstractSocial media (and the world at large) have been awash with news of the COVID-19 pandemic. With the passage of time, news and awareness about COVID-19 spread like the pandemic itself, with an explosion of messages, updates, videos, and posts. Mass hysteria manifest as another concern in addition to the health risk that COVID-19 presented. Predictably, public panic soon followed, mostly due to misconceptions, a lack of information, or sometimes outright misinformation about COVID-19 and its impacts. It is thus timely and important to conduct anex post factoassessment of the early information flows during the pandemic on social media, as well as a case study of evolving public opinion on social media which is of general interest. This study aims to inform policy that can be applied to social media platforms; for example, determining what degree of moderation is necessary to curtail misinformation on social media. This study also analyzes views concerning COVID-19 by focusing on people who interact and share social media on Twitter. As a platform for our experiments, we present a new large-scale sentiment data set COVIDSENTI, which consists of 90 000 COVID-19-related tweets collected in the early stages of the pandemic, from February to March 2020. The tweets have been labeled into positive, negative, and neutral sentiment classes. We analyzed the collected tweets for sentiment classification using different sets of features and classifiers. Negative opinion played an important role in conditioning public sentiment, for instance, we observed that people favored lockdown earlier in the pandemic; however, as expected, sentiment shifted by mid-March. Our study supports the view that there is a need to develop a proactive and agile public health presence to combat the spread of negative sentiment on social media following a pandemic. Usman Naseem, Muhammad Imran Razzak, Matloob Khushi, Peter W. Eklund, Jinman Kim |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2021 | A New Aggregation of DNN Sparse and Dense Labeling for Saliency DetectionabstractAs a fundamental requirement to many computer vision systems, saliency detection has experienced substantial progress in recent years based on deep neural networks (DNNs). Most DNN-based methods rely on either sparse or dense labeling, and thus they are subject to the inherent limitations of the chosen labeling schemes. DNN dense labeling captures salient objects mainly from global features, which are often hampered by other visually distinctive regions. On the other hand, DNN sparse labeling is usually impeded by inaccurate presegmentation of the images that it depends on. To address these limitations, we propose a new framework consisting of two pathways and an Aggregator to progressively integrate the DNN sparse and DNN dense labeling schemes to derive the final saliency map. In our "zipper" type aggregation, we propose a multiscale kernels approach to extract optimal criteria for saliency detection where we suppress nonsalient regions in the sparse labeling while guiding the dense labeling to recognize more complete extent of the saliency. We demonstrate that our method outperforms in saliency detection compared to other 11 state-of-the-art methods across six well-recognized benchmarking datasets. Ke Yan 0005, Xiuying Wang 0001, Jinman Kim, David Dagan Feng |
IEEE Trans. Cybern. | 3 |
| 2021 | Optic Disk and Cup Segmentation Through Fuzzy Broad Learning System for Glaucoma ScreeningabstractGlaucoma is an ocular disease that causes permanent blindness if not cured at an early stage. Cup-to-disk ratio (CDR), obtained by dividing the height of optic cup (OC) with the height of optic disk (OD), is a widely adopted metric used for glaucoma screening. Therefore, accurately segmenting OD and OC is crucial for calculating a CDR. Most methods have employed deep learning methods for the segmentation of OD and OC. However, these methods are very time consuming. In this article, we present a new fuzzy broad learning system-based technique for OD and OC segmentation with glaucoma screening. We comprehensively integrated extracting a region of interest from RGB images, data augmentation, extracting red and green channel images, and inputting them to the two separate fuzzy broad learning system-based neural networks for segmenting the OD and OC, respectively, and then calculated CDR. Experiments show that our fuzzy broad learning system-based technique outperforms many state-of-the-art methods. Riaz Ali, Bin Sheng 0001, Ping Li 0016, Huating Li, Po Yang 0001, Younhyun Jung, Jinman Kim, C. L. Philip Chen |
IEEE Trans. Ind. Informatics | 8 |
| 2021 | Modified GAN-CAED to Minimize Risk of Unintentional Liver Major Vessels Cutting by Controlled Segmentation Using CTA/SPET-CTabstractThis article substantially advances upon state-of-the-art to enhance liver vessels segmentation accuracy by leveraging advantages of synthetic PET-CT (SPET-CT) images in addition to computed tomography angiography (CTA) volumes. Our setup makes a hybrid solution of modified generative adversarial network-convolutional autoencoder (GAN-cAED) combining synthetic ability of GAN to deliver SPET-CT images with generative ability of cAED network in terms of latent learning to more refined segmentation of major liver vessels. We improve time complexity through a novel concept of controlled segmentation by introducing a threshold metric to stop segmentation up to a desired level. The innovative concept of controlled vessel segmentation with a stopping criterion via variant threshold levels will help surgeons to avoid unintentional major blood vessels cutting, reducing the risk of excessive blood loss. Clinically, such solutions offer computer-aided liver surgeries and drug treatment evaluation in a CTA-only environment, shorten the requirement of radioactive and expensive fused PET-CT images. Muhammad Nadeem Cheema, Anam Nazir, Po Yang 0001, Bin Sheng 0001, Ping Li 0016, Huating Li, Xiaoer Wei, Harry Qin, Jinman Kim, David Dagan Feng |
IEEE Trans. Ind. Informatics | 9 |
| 2021 | Multimodal Spatial Attention Module for Targeting Multimodal PET-CT Lung Tumor SegmentationabstractMultimodal positron emission tomography-computed tomography (PET-CT) is used routinely in the assessment of cancer. PET-CT combines the high sensitivity for tumor detection of PET and anatomical information from CT. Tumor segmentation is a critical element of PET-CT but at present, the performance of existing automated methods for this challenging task is low. Segmentation tends to be done manually by different imaging experts, which is labor-intensive and prone to errors and inconsistency. Previous automated segmentation methods largely focused on fusing information that is extracted separately from the PET and CT modalities, with the underlying assumption that each modality contains complementary information. However, these methods do not fully exploit the high PET tumor sensitivity that can guide the segmentation. We introduce a deep learning-based framework in multimodal PET-CT segmentation with a multimodal spatial attention module (MSAM). The MSAM automatically learns to emphasize regions (spatial areas) related to tumors and suppress normal regions with physiologic high-uptake from the PET input. The resulting spatial attention maps are subsequently employed to target a convolutional neural network (CNN) backbone for segmentation of areas with higher tumor likelihood from the CT image. Our experimental results on two clinical PET-CT datasets of non-small cell lung cancer (NSCLC) and soft tissue sarcoma (STS) validate the effectiveness of our framework in these different cancer types. We show that our MSAM, with a conventional U-Net backbone, surpasses the state-of-the-art lung tumor segmentation approach by a margin of 7.6% in Dice similarity coefficient (DSC). Xiaohang Fu, Lei Bi 0001, Ashnil Kumar, Michael J. Fulham, Jinman Kim |
IEEE J. Biomed. Health Informatics | 5 |
| 2021 | Hybrid Refinement-Correction Heatmaps for Human Pose EstimationabstractIn this paper, we present a method (Hybrid-Pose) to improve human pose estimation in images. We adopt Stacked Hourglass Networks to design two convolutional neural network models, RNet for pose refinement and CNet for pose correction. The CNet (Correction Network) guides the pose refinement RNet (Refinement Network) to correct the joint location before generating the final pose. Each of the two models is composed of four hourglasses, and each hourglass generates a group of detection heatmaps for the joints. The RNet model hourglasses have the same structure. However, the CNet model is designed with hourglasses of different structures for pose guidance. Since the pose estimation in RGB images is very sensitive to the image scene, our proposed approach generates multiple outputs of detection heatmaps to broaden the searching scope for the correct joints locations. We use the RNet model to refine the joints locations in each hourglass stage horizontally, then the heatmaps of each stage are fused with the heatmaps of all the CNet model hourglasses vertically in a hybrid manner. Our method shows competitive results with the existing state-of-the-art approaches on MPII and FLIC benchmark datasets. Although our proposed method focuses on improving single-person pose estimation, we also show the influence of this improvement on multi-person pose estimation by detecting multiple people using SSD detector, then estimating the pose of each person individually. Aouaidjia Kamel, Bin Sheng 0001, Ping Li 0016, Jinman Kim, David Dagan Feng |
IEEE Trans. Multim. | 4 |
| 2021 | Efficient Body Motion Quantification and Similarity Evaluation Using 3-D Joints Skeleton CoordinatesabstractEvaluating whole-body motion is challenging because of the articulated nature of the skeleton structure. Each joint moves in an unpredictable way with uncountable possibilities of movements direction under the influence of one or many of its parent joints. This paper presents a method for human motion quantification via three-dimensional (3-D) body joints coordinates. We calculate a set of metrics that influence the joints movement considering the motion of its parent joints without requiring prior knowledge of the motion parameters. Only the raw joints coordinates data of a motion sequence are needed to automatically estimate the transformation matrix of the joints between frames. We also consider the angles between limbs as a fundamental factor to follow the joints directions. We classify the joints motion as global motion and local motion. The global motion represents the joint movement according to a fixed joint, and the local motion represents the joint movement according to its first parent joint. In order to evaluate the performance of the proposed method, we also propose a comparison algorithm between two skeletons motions based on the quantified metrics. We measured the comparative similarity between the 3-D joints coordinates on Microsoft Kinect V2 and UTD-MHAD dataset. User studies were conducted to evaluate the performance under different factors. Various results and comparisons have shown that our method effectively quantifies and evaluates the motion similarity. Aouaidjia Kamel, Bin Sheng 0001, Ping Li 0016, Jinman Kim, David Dagan Feng |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2020 | Hierarchical Rendering System Based on Viewpoint Prediction in Virtual Reality
Ping Lu 0008, Ping Li 0016, Jinman Kim, Bin Sheng 0001, Lijuan Mao |
CGI | 4 |
| 2020 | GHand: A Graph Convolution Network for 3D Hand Pose Estimation
Pengsheng Wang, Guangtao Xue, Ping Li 0016, Jinman Kim, Bin Sheng 0001, Lijuan Mao |
CGI | 4 |
| 2020 | Preserving Temporal Consistency in Videos Through Adaptive SLIC
Han Zhang 0053, Riaz Ali, Bin Sheng 0001, Ping Li 0016, Jinman Kim |
CGI | 5 |
| 2020 | A Spatiotemporal Volumetric Interpolation Network for 4D Dynamic Medical ImageabstractDynamic medical images are often limited in its application due to the large radiation doses and longer image scanning and reconstruction times. Existing methods attempt to reduce the volume samples in the dynamic sequence by interpolating the volumes between the acquired samples. However, these methods are limited to either 2D images and/or are unable to support large but periodic variations in the functional motion between the image volume samples. In this paper, we present a spatiotemporal volumetric interpolation network (SVIN) designed for 4D dynamic medical images. SVIN introduces dual networks: the first is the spatiotemporal motion network that leverages the 3D convolutional neural network (CNN) for unsupervised parametric volumetric registration to derive spatiotemporal motion field from a pair of image volumes; the second is the sequential volumetric interpolation network, which uses the derived motion field to interpolate image volumes, together with a new regression-based module to characterize the periodic motion cycles in functional organ structures. We also introduce an adaptive multi-scale architecture to capture the volumetric large anatomy motions. Experimental results demonstrated that our SVIN outperformed state-of-the-art temporal medical interpolation methods and natural video interpolation method that has been extended to support volumetric images. Code is available at [1]. Yuyu Guo 0002, Lei Bi 0001, Euijoon Ahn, David Dagan Feng, Qian Wang 0001, Jinman Kim |
CVPR | 6 |
| 2020 | Unsupervised Positron Emission Tomography Tumor Segmentation via GAN based Adversarial Auto-EncoderabstractFluorodeoxyglucose Positron emission tomography (FDG PET) is the imaging modality of choice for the diagnosis of lung cancer. The automated segmentation of tumors in PET images is a fundamental requirement for image analysis in computer aided diagnosis systems. Current tumor segmentation in PET generally relies on local features to discriminate tumor from the background. These methods are limited due to poor resolution, and subtle inter-class differences when there is normal FDG uptake region (i.e., in the heart and mediastinum) in the same field of view. We propose a new image based discriminative method to separate tumor regions from normal regions. We introduce a convolutional adversarial auto-encoder to learn a latent space which models normal (disease-free) variations of PET images, and then to compute a residual map that identifies where the PET image differs from this manifold due to anomalies, i.e., tumors. Our method is tolerant to normal intra-class variations among the PET images but is discriminative of the tumors with high sensitivity. Our experiments with a clinical lung cancer dataset show that our method outperformed the state-of-the-art unsupervised segmentation methods. We also achieved higher dice score (62.0%) and sensitivity (77.9%) than the supervised U-Net method (59.5% and 59.7%). Xinheng Wu, Lei Bi 0001, Michael J. Fulham, Jinman Kim |
ICARCV | 4 |
| 2020 | Malocclusion Treatment Planning via PointNet Based Spatial Transformation Network
Xiaoshuang Li, Lei Bi 0001, Jinman Kim, Tingyao Li, Peng Li 0079, Bin Sheng 0001, David Dagan Feng |
MICCAI (3) | 3 |
| 2020 | Multi-modality Information Fusion for Radiomics-Based Neural Architecture Search
Yige Peng, Lei Bi 0001, Michael J. Fulham, David Dagan Feng, Jinman Kim |
MICCAI (7) | 5 |
| 2020 | Emotion sharing in remote patient monitoring of patients with chronic kidney diseaseabstractOBJECTIVE: To investigate the relationship between emotion sharing and technically troubled dialysis (TTD) in a remote patient monitoring (RPM) setting. MATERIALS AND METHODS: A custom software system was developed for home hemodialysis patients to use in an RPM setting, with focus on emoticon sharing and sentiment analysis of patients' text data. We analyzed the outcome of emoticon and sentiment against TTD. Logistic regression was used to assess the relationship between patients' emotions (emoticon and sentiment) and TTD. RESULTS: Usage data were collected from January 1, 2015 to June 1, 2018 from 156 patients that actively used the app system, with a total of 31 159 dialysis sessions recorded. Overall, 122 patients (78%) made use of the emoticon feature while 146 patients (94%) wrote at least 1 or more session notes for sentiment analysis. In total, 4087 (13%) sessions were classified as TTD. In the multivariate model, when compared to sessions with self-reported very happy emoticons, those with sad emoticons showed significantly higher associations to TTD (aOR 4.97; 95% CI 4.13-5.99; P = < .001). Similarly, negative sentiments also revealed significant associations to TTD (aOR 1.56; 95% CI 1.22-2; P = .003) when compared to positive sentiments. DISCUSSION: The distribution of emoticons varied greatly when compared to sentiment analysis outcomes due to the differences in the design features. The emoticon feature was generally easier to understand and quicker to input while the sentiment analysis required patients to manually input their personal thoughts. CONCLUSION: Patients on home hemodialysis actively expressed their emotions during RPM. Negative emotions were found to have significant associations with TTD. The use of emoticons and sentimental analysis may be used as a predictive indicator for prolonged TTD. Robin Huang, Na Liu 0005, Mary Ann Nicdao, Mary Mikaheal, Tanya Baldacchino, Annabelle Albeos, Kathy Petoumenos, Kamal Sud, Jinman Kim |
J. Am. Medical Informatics Assoc. | 9 |
| 2020 | Multi-Label classification of multi-modality skin lesion via hyper-connected convolutional neural network
Lei Bi 0001, David Dagan Feng, Michael J. Fulham, Jinman Kim |
Pattern Recognit. | 4 |
| 2020 | Automated Decision Support System for Lung Cancer Detection and Classification via Enhanced RFCN With Multilayer Fusion RPNabstractDetection of lung cancer at early stages is critical, in most of the cases radiologists read computed tomography (CT) images to prescribe follow-up treatment. The conventional method for detecting nodule presence in CT images is tedious. In this article, we propose an enhanced multidimensional region-based fully convolutional network (mRFCN) based automated decision support system for lung nodule detection and classification. The mRFCN is used as an image classifier backbone for feature extraction along with the novel multilayer fusion region proposal network (mLRPN) with position-sensitive score maps being explored. We applied a median intensity projection to leverage three-dimensional information from CT scans and introduced deconvolutional layer to adopt proposed mLRPN in our architecture to automatically select the potential region of interest. Our system has been trained and evaluated using LIDC dataset, and the experimental results showed promising detection performance in comparison to the state-of-the-art nodule detection/classification methods, achieving a sensitivity of 98.1% and classification accuracy of 97.91%. Anum Masood, Bin Sheng 0001, Po Yang 0001, Ping Li 0016, Huating Li, Jinman Kim, David Dagan Feng |
IEEE Trans. Ind. Informatics | 6 |
| 2020 | OFF-eNET: An Optimally Fused Fully End-to-End Network for Automatic Dense Volumetric 3D Intracranial Blood Vessels SegmentationabstractIntracranial blood vessels segmentation from computed tomography angiography (CTA) volumes is a promising biomarker for diagnosis and therapeutic treatment in cerebrovascular diseases. These segmentation outputs are a fundamental requirement in the development of automated decision support systems for preoperative assessment or intraoperative guidance in neuropathology. The state-of-the-art in medical image segmentation methods are reliant on deep learning architectures based on convolutional neural networks. However, despite their popularity, there is a research gap in the current deep learning architectures optimized to address the technical challenges in blood vessel segmentation. These challenges include: (i) the extraction of concrete brain vessels close to the skull; and (ii) the precise marking of the vessel locations. We propose an Optimally Fused Fully end-to-end Network (OFF-eNET) for automatic segmentation of the volumetric 3D intracranial vascular structures. OFF-eNET comprises of three modules. In the first module, we exploit the up-skip connections to enhance information flow, and dilated convolution for detailed preservation of spatial feature map that are designed for thin blood vessels. In the second module, we employ residual mapping along with inception module for speedy network convergence and richer visual representation. For the third module, we make use of the transferred knowledge in the form of cascaded training strategy to gradually optimize the three segmentation stages (basic, complete, and enhanced) to segment thin vessels located close to the skull. All these modules are designed to be computationally efficient. Our OFF-eNET, evaluated using 70 CTA image volumes, resulted in 90.75% performance in the segmentation of intracranial blood vessels and outperformed the state-of-the-art counterparts. Anam Nazir, Muhammad Nadeem Cheema, Bin Sheng 0001, Huating Li, Ping Li 0016, Po Yang 0001, Younhyun Jung, Harry Qin, Jinman Kim, David Dagan Feng |
IEEE Trans. Image Process. | 9 |
| 2020 | Unsupervised Domain Adaptation to Classify Medical Images Using Zero-Bias Convolutional Auto-Encoders and Context-Based Feature AugmentationabstractThe accuracy and robustness of image classification with supervised deep learning are dependent on the availability of large-scale labelled training data. In medical imaging, these large labelled datasets are sparse, mainly related to the complexity in manual annotation. Deep convolutional neural networks (CNNs), with transferable knowledge, have been employed as a solution to limited annotated data through: 1) fine-tuning generic knowledge with a relatively smaller amount of labelled medical imaging data, and 2) learning image representation that is invariant to different domains. These approaches, however, are still reliant on labelled medical image data. Our aim is to use a new hierarchical unsupervised feature extractor to reduce reliance on annotated training data. Our unsupervised approach uses a multi-layer zero-bias convolutional auto-encoder that constrains the transformation of generic features from a pre-trained CNN (for natural images) to non-redundant and locally relevant features for the medical image data. We also propose a context-based feature augmentation scheme to improve the discriminative power of the feature representation. We evaluated our approach on 3 public medical image datasets and compared it to other state-of-the-art supervised CNNs. Our unsupervised approach achieved better accuracy when compared to other conventional unsupervised methods and baseline fine-tuned CNNs. Euijoon Ahn, Ashnil Kumar, Michael J. Fulham, David Dagan Feng, Jinman Kim |
IEEE Trans. Medical Imaging | 5 |
| 2020 | Co-Learning Feature Fusion Maps From PET-CT Images of Lung CancerabstractThe analysis of multi-modality positron emission tomography and computed tomography (PET-CT) images for computer aided diagnosis applications (e.g., detection and segmentation) requires combining the sensitivity of PET to detect abnormal regions with anatomical localization from CT. Current methods for PET-CT image analysis either process the modalities separately or fuse information from each modality based on knowledge about the image analysis task. These methods generally do not consider the spatially varying visual characteristics that encode different information across the different modalities, which have different priorities at different locations. For example, a high abnormal PET uptake in the lungs is more meaningful for tumor detection than physiological PET uptake in the heart. Our aim is to improve fusion of the complementary information in multi-modality PET-CT with a new supervised convolutional neural network (CNN) that learns to fuse complementary information for multi-modality medical image analysis. Our CNN first encodes modality-specific features and then uses them to derive a spatially varying fusion map that quantifies the relative importance of each modality's features across different spatial locations. These fusion maps are then multiplied with the modality-specific feature maps to obtain a representation of the complementary multi-modality information at different locations, which can then be used for image analysis. We evaluated the ability of our CNN to detect and segment multiple regions (lungs, mediastinum, tumors) with different fusion requirements using a dataset of PET-CT images of lung cancer. We compared our method to baseline techniques for multi-modality image fusion (fused inputs (FS), multi-branch (MB) techniques, and multichannel (MC) techniques) and segmentation. Our findings show that our CNN had a significantly higher foreground detection accuracy (99.29%, p < 0:05) than the fusion baselines (FS: 99.00%, MB: 99.08%, TC: 98.92%) and a significantly higher Dice score (63.85%) than recent PET-CT tumor segmentation methods. Ashnil Kumar, Michael J. Fulham, David Dagan Feng, Jinman Kim |
IEEE Trans. Medical Imaging | 4 |
| 2020 | Simplified non-locally dense network for single-image dehazing
Zhuoliang Hu, Bin Sheng 0001, Ping Li 0016, Jinman Kim, Enhua Wu |
Vis. Comput. | 5 |
| 2019 | Challenges for Brain Data Analysis in VR EnvironmentsabstractAnalysing and understanding brain function and disorder is the main focus of neuroscience. Due to the high complexity of the brain, directionality of the signal and changing activity over time, visual exploration and data analysis are difficult. For this reason, a vast amount of research challenges are still unsolved. We explored different challenges of the visual analysis of brain data and the design of corresponding immersive environments in collaboration with experts from the biomedical domain. We built a prototype of an immersive virtual reality environment to explore the design space and to investigate how brain data analysis can be supported by a variety of design choices. Our environment can be used to study the effect of different visualisations and combinations of brain data representation, as for example network layouts, anatomical mapping or time series. As a long-term goal, we aim to aid neuro-scientists in a better understanding of brain function and disorder. Sabrina Jaeger, Karsten Klein 0001, Lucas Joos, Johannes Zagermann, Michael de Ridder, Jinman Kim, Jean Y. H. Yang, Ulrike Pfeil, Harald Reiterer, Falk Schreiber |
PacificVis | 6 |
| 2019 | Optic Disc and Cup Segmentation Based on Enhanced SegNetabstractDue to imbalanced distributed and restricted medical resources, reliable analysis for medical images is hard to come by, and it is impractical to only rely on human beings to do all the analysis, which is time-consuming and not economic. Application of computer vision techniques in such fields emerges as the situation requires. In this paper, we use deep learning segmentation algorithm to segment the optic disc and the cup from each other and from the rest of the ophthalmoscopy photographs. For a better performance, we change the loss function and crop as a way of data augmentation. The segmentation results can be used to calculate the cup-to-disc ratio (CDR), which is further used to diagnose glaucoma. Challenges such as over-fitting, biased dataset, and poor generalization of the model exist in front of us. We illustrate our model and associated methods dealing with these challenges. Lianyi Wu, Yelin Shi, Bin Sheng 0001, Ping Li 0016, Lei Bi 0001, Jinman Kim |
CASA | 7 |
| 2019 | Detect Glaucoma with Image Segmentation and Transfer LearningabstractIn this paper, we aim to automatically detect glaucoma via deep learning. To do that, we need to calculate the cup-to-disc ratio (CDR) on fine segmented retina images. To get precise segmentation, we implemented SegNet together with adversarial discriminative domain adaptation (ADDA), the former is a famous artificial neural network with encoder-decoder architecture used in image segmentation area and the latter is a transfer learning method for domain adaptation. We are the first to combine them together to detect glaucoma on test dataset which have different brightness from our training dataset. We thoroughly evaluated the proposed method with various loss functions, normal cross entropy loss, weighted cross entropy loss and dice coefficient loss included. And we show that dice loss is the best for this task. Last but not least, our experiments on transfer learning have shown that our ADDA method reduces the mean square error (MSE) between the CDR of our segmentation and annotations greatly. Lianyi Wu, Yelin Shi, Bin Sheng 0001, Ping Li 0016, Lei Bi 0001, Jinman Kim |
CASA | 7 |
| 2019 | Deep Intrinsic Image Decomposition Using Joint Parallel Learning
Yuan Yuan 0022, Bin Sheng 0001, Ping Li 0016, Lei Bi 0001, Jinman Kim, Enhua Wu |
CGI | 5 |
| 2019 | Automatic Computer Aided System for Lung Cancer in Chest CTs Using MD-RFCN Combined with Tri-Level Region Proposal NetworkabstractPulmonary cancer is one of the major causes of deaths caused by cancer around the globe. Early stage lung cancer detection can prove to be essential for the patients, for which the computed tomography (CT) images are analyzed by the radiologists to determine the presence of nodules and diagnose the disease. Conventional techniques used by the radiologists for nodule detection in CT images is time-consuming and inefficient; to assist in the diagnosis process and further enhance its efficiency and accuracy, decision support systems have been developed in the past few years. In our paper, we proposed a Multi-Dimension Region-based Fully Convolutional Network based decision support system for detection and classification of lung nodule. The Multi-Dimension RFCN serves as an image classifier backbone for our feature extraction step in addition to the proposed Tri-Level Region Proposal Network (3L-RPN) along with the position-sensitive score maps (PSSM) being explored. A novel median intensity projection method is used to leverage the multi-dimensional information from CT images and introduced an additional deconvolutional layer to adopt the proposed Tri-Level Region Proposal Network in our architecture to automatically identify the potential Region of Interest. We trained and evaluated our proposed decision support system using LIDC-IDRI dataset. The evaluation results demonstrated the high level performance of our proposed model in comparison to the state-of-the-art nodule detection and classification methods by attaining classification accuracy of 97.61% and sensitivity of 97.4%. Anum Masood, Bin Sheng 0001, Ping Li 0016, Po Yang 0001, Jinman Kim |
INDIN | 5 |
| 2019 | Deep Local-Global Refinement Network for Stent Analysis in IVOCT Images
Yuyu Guo 0002, Lei Bi 0001, Ashnil Kumar, Yue Gao 0002, Ruiyan Zhang, David Dagan Feng, Qian Wang 0001, Jinman Kim |
MICCAI (5) | 8 |
| 2019 | SRNPD: Spatial rendering network for pencil drawing stylizationabstractAbstract Pencil drawing is a simple yet effective way to depict what people see by clearly presenting details of the scene. Existing methods usually extract strokes of the input image and adjust the result image tone to make it look like a pencil drawing. However, they do not consider the quality of the stroke image and the geometry information of lines in the stroke image, which unavoidably results in the violation of original essential structures and in a flatten pencil drawing with unrealistic appearance. We put forward a spatial rendering network for pencil drawing stylization. Spatial stroke images are extracted from the image pyramid by a single‐shot bottom‐up neural network to improve the quality of these stroke images. Unlike the former tone adjustment–based methods, we analyze perceptual cues of strokes at different stroke image levels and use the obtained geometry information to constrain the stroke shading procedure. The final pencil drawing result is achieved by the stroke shading fusion of different levels' shading results. The effectiveness of our spatial rendering network for pencil drawing stylization is demonstrated by an ablation study, comparison to the state of the art, and a user study. Yuxi Jin, Ping Li 0016, Bin Sheng 0001, Yongwei Nie, Jinman Kim, Enhua Wu |
Comput. Animat. Virtual Worlds | 5 |
| 2019 | Convolutional sparse kernel network for unsupervised medical image analysis
Euijoon Ahn, Ashnil Kumar, Michael J. Fulham, David Dagan Feng, Jinman Kim |
Medical Image Anal. | 5 |
| 2019 | Step-wise integration of deep class-specific learning for dermoscopic image segmentation
Lei Bi 0001, Jinman Kim, Euijoon Ahn, Ashnil Kumar, David Dagan Feng, Michael J. Fulham |
Pattern Recognit. | 2 |
| 2019 | Unsupervised Two-Path Neural Network for Cell Event Detection and Classification Using Spatiotemporal PatternsabstractAutomatic event detection in cell videos is essential for monitoring cell populations in biomedicine. Deep learning methods have advantages over traditional approaches for cell event detection due to their ability to capture more discriminative features of cellular processes. Supervised deep learning methods, however, are inherently limited due to the scarcity of annotated data. Unsupervised deep learning methods have shown promise in general (non-cell) videos because they can learn the visual appearance and motion of regularly occurring events. Cell videos, however, can have rapid, irregular changes in cell appearance and motion, such as during cell division and death, which are often the events of most interest. We propose a novel unsupervised two-path input neural network architecture to capture these irregular events with three key elements: 1) a visual encoding path to capture regular spatiotemporal patterns of observed objects with convolutional long short-term memory units; 2) an event detection path to extract information related to irregular events with max-pooling layers; and 3) integration of the hidden states of the two paths to provide a comprehensive representation of the video that is used to simultaneously locate and classify cell events. We evaluated our network in detecting cell division in densely packed stem cells in phase-contrast microscopy videos. Our unsupervised method achieved higher or comparable accuracy to standard and state-of-the-art supervised methods. Ha Tran Hong Phan, Ashnil Kumar, David Dagan Feng, Michael J. Fulham, Jinman Kim |
IEEE Trans. Medical Imaging | 5 |
| 2018 | Voxelized Facial Reconstruction Using Deep Neural NetworkabstractThis paper presents an approach to predicting variation tendency of human faces with regard to cranium changes based on deep learning. Our work focuses on generating individual customized facial models with high plausibility. Inspired by the performance of encoder-decoder convolutional neural network, the core trainable predicting engine of our learning network is designed for three-dimension voxelized data representation as the encoder-decoder structure and the encoder part is similar to the 7 layers of VGG16 network. To take full consideration of the cranium changes and features of original human face, a novel formation of channeled volumetric data structure is presented, and also the corresponding sub and up-sampling strategies for volume data. Our encoder-decoder neural network consumes discrete 3-channel volume data and generates 1-channel volume data as predicted post-variation human face. This framework is quantified with clinical dataset and it shows that its' performance improves in comparison with the state-of-the-art technologies. Xiaoshuang Li, Bin Sheng 0001, Ping Li 0016, Jinman Kim, David Dagan Feng |
CGI | 4 |
| 2018 | Prior Knowledge Driven Energy for Saliency DetectionabstractSaliency detection on images has experienced substantial progress in recent years on the basis of deep neural network (DNN). However, there may exist secondary saliency in the background that distracts DNN learning and mistakes the secondary salient regions as saliency. To address this issue, we propose a dual-term energy to improve the inference of saliency on top of DNN estimation, where dense term smoothens salient regions in pixel scale and sparse term extracts prior knowledge to differentiate saliency and non-saliency superpixels. Our prior knowledge including extra- and intra-region priors, contributes to improving overall saliency detection. The extra-region prior knowledge estimates the saliency probabilities for different pre-partitioned regions to eliminate the secondary saliency. The intra-region prior knowledge helps to group the salient regions that otherwise could be ignored by DNN predictor, and thus to provide more complete saliency definition. We evaluated our model on 8,465 images from four well-recognized saliency detection benchmarking datasets, and compared our model to six state-of-the-art comparative methods. Experimental results demonstrated that our model outperformed the state-of-the-art counterpart with improvements of up to 2.51% in terms of F-measure. Ke Yan 0005, Chaojie Zheng, Qiu Huang, Jinman Kim, David Dagan Feng, Xiuying Wang 0001 |
ICARCV | 4 |
| 2018 | Feature of Interest-Based Direct Volume Rendering Using Contextual Saliency-Driven Ray Profile AnalysisabstractAbstract Direct volume rendering (DVR) visualization helps interpretation because it allows users to focus attention on the subset of volumetric data that is of most interest to them. The ideal visualization of the features of interest (FOIs) in a volume, however, is still a major challenge. The clear depiction of FOIs depends on accurate identification of the FOIs and appropriate specification of the optical parameters via transfer function (TF) design and it is typically a repetitive trial‐and‐error process. We address this challenge by introducing a new method that uses contextual saliency information to group the voxels along a viewing ray into distinct FOIs where ‘contextual saliency’ is a biologically inspired attribute that aids the identification of features that the human visual system considers important. The saliency information is also used to automatically define the optical parameters that emphasize the visual depiction of the FOIs in DVR. We demonstrate the capabilities of our method by its application to a variety of volumetric data sets and highlight its advantages by comparison to current state‐of‐the‐art ray profile analysis methods. Younhyun Jung, Jinman Kim, Ashnil Kumar, David Dagan Feng, Michael J. Fulham |
Comput. Graph. Forum | 2 |
| 2018 | Dense and Sparse Labeling With Multidimensional Features for Saliency DetectionabstractConventional low-level feature-based saliency detection methods tend to use nonrobust prior knowledge and do not perform well in complex or low-contrast images. In this paper, to address these issues in existing methods, we propose a novel deep neural network (DNN)-based dense and sparse labeling (DSL) framework for saliency detection. DSL consists of three major steps, namely, dense labeling (DL), sparse labeling (SL), and deep convolutional (DC) network. The DL and SL steps conduct initial saliency estimations with macro object contours and low-level image features, respectively, which effectively approximate the location of the salient object and generate accurate guidance channels for the DC step; the DC step, on the other hand, takes in the results of DL and SL, establishes a six-channeled input data structure (including local superpixel information), and conducts accurate final saliency classification. Our DSL framework exploits the saliency estimation guidance from both macro object contours and local low-level features, as well as utilizing the DNN for high-level saliency feature extraction. Extensive experiments are conducted on six well-recognized public data sets against 16 state-of-the-art saliency detection methods, including ten conventional feature-based methods and six learning-based methods. The results demonstrate the superior performance of DSL on various challenging cases in terms of both accuracy and robustness. Yuchen Yuan, ChangYang Li, Jinman Kim, Tom Weidong Cai, David Dagan Feng |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Reversion Correction and Regularized Random Walk Ranking for Saliency DetectionabstractIn recent saliency detection research, many graph-based algorithms have applied boundary priors as background queries, which may generate completely "reversed" saliency maps if the salient objects are on the image boundaries. Moreover, these algorithms usually depend heavily on pre-processed superpixel segmentation, which may lead to notable degradation in image detail features. In this paper, a novel saliency detection method is proposed to overcome the above issues. First, we propose a saliency reversion correction process, which locates and removes the boundary-adjacent foreground superpixels, and thereby increases the accuracy and robustness of the boundary prior-based saliency estimations. Second, we propose a regularized random walk ranking model, which introduces prior saliency estimation to every pixel in the image by taking both region and pixel image features into account, thus leading to pixel-detailed and superpixel-independent saliency maps. Experiments are conducted on four well-recognized data sets; the results indicate the superiority of our proposed method against 14 state-of-the-art methods, and demonstrate its general extensibility as a saliency optimization algorithm. We further evaluate our method on a new data set comprised of images that we define as boundary adjacent object saliency, on which our method performs better than the comparison methods. Yuchen Yuan, ChangYang Li, Jinman Kim, Tom Weidong Cai, David Dagan Feng |
IEEE Trans. Image Process. | 3 |
| 2018 | Dual-Path Adversarial Learning for Fully Convolutional Network (FCN)-Based Medical Image Segmentation
Lei Bi 0001, David Dagan Feng, Jinman Kim |
Vis. Comput. | 3 |
| 2017 | Temporaltracks: visual analytics for exploration of 4D fMRI time-series coactivationabstractFunctional magnetic resonance imaging (fMRI) is a 4D medical imaging modality that depicts a proxy of neuronal activity in a series of temporal scans. Statistical processing of the modality shows promise in uncovering insights about the functioning of the brain, such as the default mode network, and characteristics of mental disorders. Current statistical processing generally summarises the temporal signals between brain regions into a single data point to represent the 'coactivation' of the regions. That is, how similar are their temporal patterns over the scans. However, the potential of such processing is limited by issues of possible data misrepresentation due to uncertainties, e.g. noise in the data. Moreover, it has been shown that brain signals are characterised by brief traces of coactivation, which are lost in the single value representations. To alleviate the issues, alternate statistical processes have been used, however creating effective techniques has proven difficult due to problems, e.g. issues with noise, which often require user input to uncover. Visual analytics, therefore, through its ability to interactively exploit human expertise, presents itself as an interesting approach of benefit to the domain. In this work, we present the conceptual design behind TemporalTracks, our visual analytics system for exploration of 4D fMRI time-series coactivation data, utilising a visual metaphor to effectively present coactivation data for easier understanding. We describe our design with a case study visually analysing Human Connectome Project data, demonstrating that TemporalTracks can uncover temporal events that would otherwise be hidden in standard analysis. Michael de Ridder, Karsten Klein 0001, Jinman Kim |
CGI | 3 |
| 2017 | Saliency-Based Lesion Segmentation Via Background Detection in Dermoscopic ImagesabstractThe segmentation of skin lesions in dermoscopic images is a fundamental step in automated computer-aided diagnosis of melanoma. Conventional segmentation methods, however, have difficulties when the lesion borders are indistinct and when contrast between the lesion and the surrounding skin is low. They also perform poorly when there is a heterogeneous background or a lesion that touches the image boundaries; this then results in under- and oversegmentation of the skin lesion. We suggest that saliency detection using the reconstruction errors derived from a sparse representation model coupled with a novel background detection can more accurately discriminate the lesion from surrounding regions. We further propose a Bayesian framework that better delineates the shape and boundaries of the lesion. We also evaluated our approach on two public datasets comprising 1100 dermoscopic images and compared it to other conventional and state-of-the-art unsupervised (i.e., no training required) lesion segmentation methods, as well as the state-of-the-art unsupervised saliency detection methods. Our results show that our approach is more accurate and robust in segmenting lesions compared to other methods. We also discuss the general extension of our framework as a saliency optimization algorithm for lesion segmentation. Euijoon Ahn, Jinman Kim, Lei Bi 0001, Ashnil Kumar, ChangYang Li, Michael J. Fulham, David Dagan Feng |
IEEE J. Biomed. Health Informatics | 2 |
| 2017 | Occlusion and Slice-Based Volume Rendering Augmentation for PET-CTabstractDual-modality positron emission tomography and computed tomography (PET-CT) depicts pathophysiological function with PET in an anatomical context provided by CT. Three-dimensional volume rendering approaches enable visualization of a two-dimensional slice of interest (SOI) from PET combined with direct volume rendering (DVR) from CT. However, because DVR depicts the whole volume, it may occlude a region of interest, such as a tumor in the SOI. Volume clipping can eliminate this occlusion by cutting away parts of the volume, but it requires intensive user involvement in deciding on the appropriate depth to clip. Transfer functions that are currently available can make the regions of interest visible, but this often requires complex parameter tuning and coupled preprocessing of the data to define the regions. Hence, we propose a new visualization algorithm where an SOI from PET is augmented by volumetric contextual information from a DVR of the counterpart CT so that the obtrusiveness from the CT in the SOI is minimized. Our approach automatically calculates an augmentation depth parameter by considering the occlusion information derived from the voxels of the CT in front of the PET SOI. The depth parameter is then used to generate an opacity weight function that controls the amount of contextual information visible from the DVR. We outline the improvements with our visualization approach compared to other slice-based and our previous approaches. We present the preliminary clinical evaluation of our visualization in a series of PET-CT studies from patients with nonsmall cell lung cancer. Younhyun Jung, Jinman Kim, David Dagan Feng, Michael J. Fulham |
IEEE J. Biomed. Health Informatics | 2 |
| 2017 | An Ensemble of Fine-Tuned Convolutional Neural Networks for Medical Image ClassificationabstractThe availability of medical imaging data from clinical archives, research literature, and clinical manuals, coupled with recent advances in computer vision offer the opportunity for image-based diagnosis, teaching, and biomedical research. However, the content and semantics of an image can vary depending on its modality and as such the identification of image modality is an important preliminary step. The key challenge for automatically classifying the modality of a medical image is due to the visual characteristics of different modalities: some are visually distinct while others may have only subtle differences. This challenge is compounded by variations in the appearance of images based on the diseases depicted and a lack of sufficient training data for some modalities. In this paper, we introduce a new method for classifying medical images that uses an ensemble of different convolutional neural network (CNN) architectures. CNNs are a state-of-the-art image classification technique that learns the optimal image features for a given classification task. We hypothesise that different CNN architectures learn different levels of semantic image representation and thus an ensemble of CNNs will enable higher quality features to be extracted. Our method develops a new feature extractor by fine-tuning CNNs that have been initialized on a large dataset of natural images. The fine-tuning process leverages the generic image features from natural images that are fundamental for all images and optimizes them for the variety of medical imaging modalities. These features are used to train numerous multiclass classifiers whose posterior probabilities are fused to predict the modalities of unseen images. Our experiments on the ImageCLEF 2016 medical image public dataset (30 modalities; 6776 training images, and 4166 test images) show that our ensemble of fine-tuned CNNs achieves a higher accuracy than established CNNs. Our ensemble also achieves a higher accuracy than methods in the literature evaluated on the same benchmark dataset and is only overtaken by those methods that source additional training data. Ashnil Kumar, Jinman Kim, David Lyndon, Michael J. Fulham, David Dagan Feng |
IEEE J. Biomed. Health Informatics | 2 |
| 2017 | Automatic Measurement of Thalamic Diameter in 2-D Fetal Ultrasound Brain Images Using Shape Prior Constrained Regularized Level SetsabstractWe derived an automated algorithm for accurately measuring the thalamic diameter from 2-D fetal ultrasound (US) brain images. The algorithm overcomes the inherent limitations of the US image modality: nonuniform density; missing boundaries; and strong speckle noise. We introduced a "guitar" structure that represents the negative space surrounding the thalamic regions. The guitar acts as a landmark for deriving the widest points of the thalamus even when its boundaries are not identifiable. We augmented a generalized level-set framework with a shape prior and constraints derived from statistical shape models of the guitars; this framework was used to segment US images and measure the thalamic diameter. Our segmentation method achieved a higher mean Dice similarity coefficient, Hausdorff distance, specificity, and reduced contour leakage when compared to other well-established methods. The automatic thalamic diameter measurement had an interobserver variability of -0.56 ± 2.29 mm compared to manual measurement by an expert sonographer. Our method was capable of automatically estimating the thalamic diameter, with the measurement accuracy on par with clinical assessment. Our method can be used as part of computer-assisted screening tools that automatically measure the biometrics of the fetal thalamus; these biometrics are linked to neurodevelopmental outcomes. Pradeeba Sridar, Ashnil Kumar, ChangYang Li, Joyce Woo, Ann Quinton, Ron Benzie, Michael J. Peek, David Dagan Feng, R. Krishna Kumar, Ralph Nanan, Jinman Kim |
IEEE J. Biomed. Health Informatics | 11 |
| 2017 | Stacked fully convolutional networks with multi-channel learning: application to medical image segmentation
Lei Bi 0001, Jinman Kim, Ashnil Kumar, Michael J. Fulham, David Dagan Feng |
Vis. Comput. | 2 |
| 2016 | Adaptive background search and foreground estimation for saliency detection via comprehensive autoencoderabstractIn saliency object detection, inappropriate boundary-background priors is known to degrade performance in challenging image datasets, and even may lead to `inverse' results when saliency regions are attached to the image boundaries. This is an active field where many works have proposed various techniques to lessen such degradation by inappropriate boundary-background priors. Although the use of boundary-background priors has shown to be capable of improving the detection, inherently, these techniques confront serious challenges in background suppression. To overcome this limitation, we propose an adaptive background extractor to search background seeds without the need of boundary-background priors. With the adaptive background seeds, the saliency objects can be then extracted via our proposed hierarchical foreground estimation model. We evaluate our adaptive Background Search and Foreground Estimation (BSFE) algorithm in comparison with six state-of-the-art methods on four well-recognized public datasets. The experimental results demonstrate that our BSFE algorithm outperforms compared methods in majority of the datasets and in particular achieves double-winners in terms of F-measure and mean absolute error on two challenging datasets. Ke Yan 0005, ChangYang Li, Xiuying Wang 0001, Yuchen Yuan, Jinman Kim, David Dagan Feng |
ICIP | 6 |
| 2016 | DeepGene: an advanced cancer type classifier based on deep learning and somatic point mutationsabstractBACKGROUND: With the developments of DNA sequencing technology, large amounts of sequencing data have become available in recent years and provide unprecedented opportunities for advanced association studies between somatic point mutations and cancer types/subtypes, which may contribute to more accurate somatic point mutation based cancer classification (SMCC). However in existing SMCC methods, issues like high data sparsity, small volume of sample size, and the application of simple linear classifiers, are major obstacles in improving the classification performance. RESULTS: To address the obstacles in existing SMCC studies, we propose DeepGene, an advanced deep neural network (DNN) based classifier, that consists of three steps: firstly, the clustered gene filtering (CGF) concentrates the gene data by mutation occurrence frequency, filtering out the majority of irrelevant genes; secondly, the indexed sparsity reduction (ISR) converts the gene data into indexes of its non-zero elements, thereby significantly suppressing the impact of data sparsity; finally, the data after CGF and ISR is fed into a DNN classifier, which extracts high-level features for accurate classification. Experimental results on our curated TCGA-DeepGene dataset, which is a reformulated subset of the TCGA dataset containing 12 selected types of cancer, show that CGF, ISR and DNN all contribute in improving the overall classification performance. We further compare DeepGene with three widely adopted classifiers and demonstrate that DeepGene has at least 24% performance improvement in terms of testing accuracy. CONCLUSIONS: Based on deep learning and somatic point mutation data, we devise DeepGene, an advanced cancer type classifier, which addresses the obstacles in existing SMCC studies. Experiments indicate that DeepGene outperforms three widely adopted existing classifiers, which is mainly attributed to its deep learning module that is able to extract the high level features between combinatorial somatic point mutations and cancer types. Yuchen Yuan, Yi Shi 0007, ChangYang Li, Jinman Kim, Tom Weidong Cai, Zeguang Han, David Dagan Feng |
BMC Bioinform. | 4 |
| 2015 | Guest-Editorial - Telehealth Systems and ApplicationsabstractIn the 21st century, the convergence of healthcare and information and communications technologies (ICT) offers an opportunity to give patients greater liberty from their health problems. Telehealth systems and applications, supported by advances in ICT, are fostering a diversity of cost-effective and efficient healthcare solutions. These solutions are becoming embedded in all aspects of clinical care and are enhancing the quality, equality, and accessibility of care, while playing a pivotal role in decreasing the rising costs from the growth in aging population. The emergence of affordable health sensors and accessible mobile computing devices, such as smartphones, wearables and tablets, offers opportunities to revolutionize healthcare solutions. David Dagan Feng, Jinman Kim, Mohamed Khadra, Donna L. Hudson, Christian Roux |
IEEE J. Biomed. Health Informatics | 2 |
| 2015 | A Visual Analytics Approach Using the Exploration of Multidimensional Feature Spaces for Content-Based Medical Image RetrievalabstractContent-based image retrieval (CBIR) is a search technique based on the similarity of visual features and has demonstrated potential benefits for medical diagnosis, education, and research. However, clinical adoption of CBIR is partially hindered by the difference between the computed image similarity and the user's search intent, the semantic gap, with the end result that relevant images with outlier features may not be retrieved. Furthermore, most CBIR algorithms do not provide intuitive explanations as to why the retrieved images were considered similar to the query (e.g., which subset of features were similar), hence, it is difficult for users to verify if relevant images, with a small subset of outlier features, were missed. Users, therefore, resort to examining irrelevant images and there are limited opportunities to discover these "missed" images. In this paper, we propose a new approach to medical CBIR by enabling a guided visual exploration of the search space through a tool, called visual analytics for medical image retrieval (VAMIR). The visual analytics approach facilitates interactive exploration of the entire dataset using the query image as a point-of-reference. We conducted a user study and several case studies to demonstrate the capabilities of VAMIR in the retrieval of computed tomography images and multimodality positron emission tomography and computed tomography images. Ashnil Kumar, Falk Nette, Karsten Klein 0001, Michael J. Fulham, Jinman Kim |
IEEE J. Biomed. Health Informatics | 5 |
| 2014 | Multi-stage Thresholded Region Classification for Whole-Body PET-CT Lymphoma Studies
Lei Bi 0001, Jinman Kim, David Dagan Feng, Michael J. Fulham |
MICCAI (1) | 2 |
| 2014 | A graph-based approach for the retrieval of multi-modality medical images
Ashnil Kumar, Jinman Kim, Lingfeng Wen, Michael J. Fulham, David Dagan Feng |
Medical Image Anal. | 2 |
| 2014 | Editorial
Jinman Kim, Daniel Thalmann, Kun Zhou 0001, David Dagan Feng, Holly E. Rushmeier |
Vis. Comput. | 1 |
| 2013 | Graph-based retrieval of PET-CT images using vector space embeddingabstractGraph-based content-based image retrieval (CBIR) techniques, which use graphs to represent image features and calculate image similarity using the graph edit distance, achieve high retrieval accuracy. However, such techniques suffer from high computational complexity. In this paper, we present a graph-based CBIR algorithm that achieves improved retrieval efficiency. We compute a vector space embedding for every graph, using their distances from a set of prototype graphs, so that each vector component represents a distortion from a prototype. This process is performed offline. We compare images by computing the Euclidean distance of the vector embeddings, which is a faster process than calculating the graph edit distance. We evaluated our work using 50 combined positron emission tomography and computed tomography (PET-CT) volumes of patients with lung tumours. Our results show that our method is at least 21 times faster than the graph edit distance with a mean average precision difference of less than 4%. Ashnil Kumar, Jinman Kim, David Dagan Feng, Michael J. Fulham |
CBMS | 2 |
| 2013 | A web-based medical multimedia visualisation interface for personal health recordsabstractThe healthcare industry has begun to utilise web-based systems and cloud computing infrastructure to develop an increasing array of online personal health record (PHR) systems. Although these systems provide the technical capacity to store and retrieve medical data in various multimedia formats, including images, videos, voice, and text, individual patient use remains limited by the lack of intuitive data representation and visualisation techniques. As such, further research is necessary to better visualise and present these records, in ways that make the complex medical data more intuitive. In this study, we present a web-based PHR visualisation system, called the 3D medical graphical avatar (MGA), which was designed to explore web-based delivery of a wide array of medical data types including multi-dimensional medical images; medical videos; text-based data; and spatial annotations. Mapping information was extracted from each of the data types and was used to embed spatial and textual annotations, such as regions of interest (ROIs) and time-based video annotations. Our MGA itself is built from clinical patient imaging studies, when available. We have taken advantage of the emerging web technologies of HTML5 and WebGL to make our application available to a wider base of users and devices. We analysed the performance of our proof-of-concept prototype system on mobile and desktop consumer devices. Our initial experiments indicate that our system can render the medical data in a fashion that enables interactive navigation of the MGA. Michael de Ridder, Liviu Constantinescu, Lei Bi 0001, Younhyun Jung, Ashnil Kumar, Jinman Kim, David Dagan Feng, Michael J. Fulham |
CBMS | 6 |
| 2013 | Visibility-driven PET-CT visualisation with region of interest (ROI) segmentation
Younhyun Jung, Jinman Kim, Stefan Eberl, Michael J. Fulham, David Dagan Feng |
Vis. Comput. | 2 |
| 2012 | Graph-based retrieval of multi-modality medical images: A comparison of representations using simulated imagesabstractContent-based image retrieval (CBIR) is an image search technique that utilises visual features as search criteria; it has potential clinical applications in evidence-based diagnosis, physician training, and biomedical research. Graph-based CBIR techniques have high accuracy when retrieving images by the similarity of the spatial arrangement of their constituent objects but these techniques were initially designed for single-modality images and have limited retrieval capabilities when multi-modality images, such as combined positron emission tomography and computed tomography (PET-CT), are considered. In this paper, we present a graph-based CBIR approach for multimodality images that integrates modality-specific features on graph vertices and adapts a well-established graph similarity scheme to account for varying vertex feature sets. Furthermore, we propose a graph pruning method that removes redundant edges using the spatial proximity of image regions. We evaluated our work using two simulated data sets, consisting of 2D liver shapes and 3D whole-body lymphoma images. In our experiments we achieved a higher level of retrieval precision using our graph method when compared to conventional graph-based retrieval, demonstrating that our proposed method enabled new capabilities and improved multi-modality CBIR. Ashnil Kumar, Jinman Kim, David Dagan Feng, Michael J. Fulham |
CBMS | 2 |
| 2012 | SparkMed: A Framework for Dynamic Integration of Multimedia Medical Data Into Distributed m-Health SystemsabstractWith the advent of 4G and other long-term evolution (LTE) wireless networks, the traditional boundaries of patient record propagation are diminishing as networking technologies extend the reach of hospital infrastructure and provide on-demand mobile access to medical multimedia data. However, due to legacy and proprietary software, storage and decommissioning costs, and the price of centralization and redevelopment, it remains complex, expensive, and often unfeasible for hospitals to deploy their infrastructure for online and mobile use. This paper proposes the SparkMed data integration framework for mobile healthcare (m-Health), which significantly benefits from the enhanced network capabilities of LTE wireless technologies, by enabling a wide range of heterogeneous medical software and database systems (such as the picture archiving and communication systems, hospital information system, and reporting systems) to be dynamically integrated into a cloud-like peer-to-peer multimedia data store. Our framework allows medical data applications to share data with mobile hosts over a wireless network (such as WiFi and 3G), by binding to existing software systems and deploying them as m-Health applications. SparkMed integrates techniques from multimedia streaming, rich Internet applications (RIA), and remote procedure call (RPC) frameworks to construct a Self-managing, Pervasive Automated netwoRK for Medical Enterprise Data (SparkMed). Further, it is resilient to failure, and able to use mobile and handheld devices to maintain its network, even in the absence of dedicated server devices. We have developed a prototype of the SparkMed framework for evaluation on a radiological workflow simulation, which uses SparkMed to deploy a radiological image viewer as an m-Health application for telemedical use by radiologists and stakeholders. We have evaluated our prototype using ten devices over WiFi and 3G, verifying that our framework meets its two main objectives: 1) interactive delivery of medical multimedia data to mobile devices; and 2) attaching to non-networked medical software processes without significantly impacting their performance. Consistent response times of under 500 ms and graphical frame rates of over 5 frames per second were observed under intended usage conditions. Further, overhead measurements displayed linear scalability and low resource requirements. Liviu Constantinescu, Jinman Kim, David Dagan Feng |
IEEE Trans. Inf. Technol. Biomed. | 2 |
| 2012 | A Multistage Discriminative Model for Tumor and Lymph Node Detection in Thoracic ImagesabstractAnalysis of primary lung tumors and disease in regional lymph nodes is important for lung cancer staging, and an automated system that can detect both types of abnormalities will be helpful for clinical routine. In this paper, we present a new method to automatically detect both tumors and abnormal lymph nodes simultaneously from positron emission tomography-computed tomography thoracic images. We perform the detection in a multistage approach, by first detecting all potential abnormalities, then differentiate between tumors and lymph nodes, and finally refine the detected tumors for false positive reduction. Each stage is designed with a discriminative model based on support vector machines and conditional random fields, exploiting intensity, spatial and contextual features. The method is designed to handle a wide and complex variety of abnormal patterns found in clinical datasets, consisting of different spatial contexts of tumors and abnormal lymph nodes. We evaluated the proposed method thoroughly on clinical datasets, and encouraging results were obtained. Yang Song 0001, Tom Weidong Cai, Jinman Kim, David Dagan Feng |
IEEE Trans. Medical Imaging | 3 |
| 2011 | Robust statistical shape models for MRI bone segmentation in presence of small field of view
Jérôme Schmid, Jinman Kim, Nadia Magnenat-Thalmann |
Medical Image Anal. | 2 |
| 2010 | Coupled Registration-Segmentation: Application to Femur Analysis with Intra-subject Multiple Levels of Detail MRI Data
Jérôme Schmid, Jinman Kim, Nadia Magnenat-Thalmann |
MICCAI (2) | 2 |
| 2010 | Collaborative telemedicine for interactive multiuser segmentation of volumetric medical images
Seunghyun Han, Niels A. Nijdam, Jérôme Schmid, Jinman Kim, Nadia Magnenat-Thalmann |
Vis. Comput. | 4 |
| 2007 | Real-Time Volume Rendering Visualization of Dual-Modality PET/CT Images With Interactive Fuzzy Thresholding SegmentationabstractThree-dimensional (3-D) visualization has become an essential part for imaging applications, including image-guided surgery, radiotherapy planning, and computer-aided diagnosis. In the visualization of dual-modality positron emission tomography and computed tomography (PET/CT), 3-D volume rendering is often limited to rendering of a single image volume and by high computational demand. Furthermore, incorporation of segmentation in volume rendering is usually restricted to visualizing the presegmented volumes of interest. In this paper, we investigated the integration of interactive segmentation into real-time volume rendering of dual-modality PET/CT images. We present and validate a fuzzy thresholding segmentation technique based on fuzzy cluster analysis, which allows interactive and real-time optimization of the segmentation results. This technique is then incorporated into a real-time multi-volume rendering of PET/CT images. Our method allows a real-time fusion and interchangeability of segmentation volume with PET or CT volumes, as well as the usual fusion of PET/CT volumes. Volume manipulations such as window level adjustments and lookup table can be applied to individual volumes, which are then fused together in real time as adjustments are made. We demonstrate the benefit of our method in integrating segmentation with volume rendering in its application to PET/CT images. Responsive frame rates are achieved by utilizing a texture-based volume rendering algorithm and the rapid transfer capability of the high-memory bandwidth available in low-cost graphic hardware. Jinman Kim, Tom Weidong Cai, Stefan Eberl, David Dagan Feng |
IEEE Trans. Inf. Technol. Biomed. | 1 |
| 2006 | Segmentation of VOI From Multidimensional Dynamic PET Images by Integrating Spatial and Temporal FeaturesabstractSegmentation of multidimensional dynamic positron emission tomography (PET) images into volumes of interest (VOIs) exhibiting similar temporal behavior and spatial features is a challenging task due to inherently poor signal-to-noise ratio and spatial resolution. In this study, we propose VOI segmentation of dynamic PET images by utilizing both the three-dimensional (3-D) spatial and temporal domain information in a hybrid technique that integrates two independent segmentation techniques of cluster analysis and region growing. The proposed technique starts with a cluster analysis that partitions the image based on temporal similarities. The resulting temporal partitions, together with the 3-D spatial information are utilized in the region growing segmentation. The technique was evaluated with dynamic 2-[18F] fluoro-2-deoxy-D-glucose PET simulations and clinical studies of the human brain and compared with the k-means and fuzzy c-means cluster analysis segmentation methods. The quantitative evaluation with simulated images demonstrated that the proposed technique can segment the dynamic PET images into VOIs of different kinetic structures and outperforms the cluster analysis approaches with notable improvements in the smoothness of the segmented VOIs with fewer disconnected or spurious segmentation clusters. In clinical studies, the hybrid technique was only superior to the other techniques in segmenting the white matter. In the gray matter segmentation, the other technique tended to perform slightly better than the hybrid technique, but the differences did not reach significance. The hybrid technique generally formed smoother VOIs with better separation of the background. Overall, the proposed technique demonstrated potential usefulness in the diagnosis and evaluation of dynamic PET neurological imaging studies. Jinman Kim, Tom Weidong Cai, David Dagan Feng, Stefan Eberl |
IEEE Trans. Inf. Technol. Biomed. | 1 |
| 2006 | A New Way for Multidimensional Medical Data Management: Volume of Interest (VOI)-Based Retrieval of Medical Images With Visual and Functional FeaturesabstractThe advances in digital medical imaging and storage in integrated databases are resulting in growing demands for efficient image retrieval and management. Content-based image retrieval (CBIR) refers to the retrieval of images from a database, using the visual features derived from the information in the image, and has become an attractive approach to managing large medical image archives. In conventional CBIR systems for medical images, images are often segmented into regions which are used to derive two-dimensional visual features for region-based queries. Although such approach has the advantage of including only relevant regions in the formulation of a query, medical images that are inherently multidimensional can potentially benefit from the multidimensional feature extraction which could open up new opportunities in visual feature extraction and retrieval. In this study, we present a volume of interest (VOI) based content-based retrieval of four-dimensional (three spatial and one temporal) dynamic PET images. By segmenting the images into VOIs consisting of functionally similar voxels (e.g., a tumor structure), multidimensional visual and functional features were extracted and used as region-based query features. A prototype VOI-based functional image retrieval system (VOI-FIRS) has been designed to demonstrate the proposed multidimensional feature extraction and retrieval. Experimental results show that the proposed system allows for the retrieval of related images that constitute similar visual and functional VOI features, and can find potential applications in medical data management, such as to aid in education, diagnosis, and statistical analysis. Jinman Kim, Tom Weidong Cai, David Dagan Feng |
IEEE Trans. Inf. Technol. Biomed. | 1 |
| 2002 | Content access and distribution of multimedia medical data in E-healthabstractE-health is greatly impacting on information distribution and availability within the health services, hospitals and to the public. Previous research has addressed the development of system architectures with the aim of integrating the distributed and heterogeneous medical information systems. Easing the difficulties in the sharing and management of multimedia medical data and the timely accessibility to these data are critical needs for health care providers. We have proposed a client-server agent that integrates and allows a portal to every permitted information system of the hospital that consists of picture archiving and communication systems (PACS), radiology information system (RIS) and hospital information system (HIS) via the intranet and the Internet. Our proposed agent enables remote access into the usually closed information system of the hospital and a server that manages all the multimedia medical data and allows for in-depth and complex search queries for content access and automatic creation of patient reports for distribution. Jinman Kim, David Dagan Feng, Tom Weidong Cai, Stefan Eberl |
ICME (2) | 1 |