Lanfen Lin

dblp:13/3315 · DBLP profile ↗
← Back
98ranked-venue papers
1as first author
67since 2021 · last 2026
0000-0003-4098-588XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 53 · 40 since 2021Applied, interdisciplinary, general and emerging computing · 32 · 26 since 2021Artificial intelligence and machine learning · 24 · 19 since 2021Databases, data management, data science and information retrieval · 9 · 5 since 2021Human-computer interaction and ubiquitous computing · 8 · 1 first-authorComputer networks · 1 · 1 since 2021Security and privacy · 1
YearPublicationVenuePosition
2026 Taming the Phantom: Token-Asymmetric Filtering for Hallucination Mitigation in Large Vision-Language Models
abstract
Hallucination in Large Vision-Language Models (LVLMs) remains a critical challenge, undermining their reliability in real-world applications. Existing studies have investigated the causes of hallucination at the modality level and proposed effective strategies. However, interaction patterns beyond the modality level remain insufficiently explored. In this paper, we conduct a token-level analysis and identify two key phenomena: (1) a small subset of textual tokens in LVLMs exert disproportionate influence in the visual-active layers, surpassing that of the visual modality and potentially misleading visual understanding; (2) while LVLMs can correctly identify key visual information, insufficient focus on these cues can sometimes lead to hallucinations. Based on such observation, we attribute hallucinations in LVLMs to two token-level causes: the disproportionate influence of certain textual tokens (phantom tokens) and the underutilization of critical visual cues (anchor tokens). To mitigate these issues, we introduce Token-Asymmetric Filtering (TAF)—a training-free, plug-and-play method that modulates intermediate attention maps in LVLMs. TAF isolates the influence of phantom tokens and emphasizes the influence of anchor tokens in the visual-active layers. Experimental results across multiple benchmarks demonstrate that TAF significantly mitigates hallucinations across a range of state-of-the-art LVLMs.
Shuyi Ouyang, Hongyi Wang 0002, Gongfan Fang, Xinyin Ma, Lanfen Lin, Xinchao Wang
AAAI5
2026 Enhancing Depression Detection Using Pretrained Multi-modal Sentiment Analysis Models with Deep Prefix Tuning
abstract
Depression, a pervasive mental health condition, affects millions globally, challenging early and accurate diagnosis due to its subtle and varied manifestations. Recognizing the critical link between emotional dysregulation and depressive symptoms, our research introduces a pioneering training paradigm that integrates sentiment analysis with depression detection. This approach is motivated by the potential of sentiment data to enrich models with a deeper understanding of emotional states, crucial for identifying depressive patterns. To leverage the nuanced sentiment information without compromising the pretrained model’s integrity, we employ deep prefix tuning. This novel technique allows for targeted model refinement, ensuring that the valuable pretrained structures are not overshadowed by the sparse and specific nature of depression-related data. The empirical results demonstrate superior performance across standard benchmarks, setting a new precedent for multimodal depression detection.
Shiyu Teng, Jiaqing Liu, Shurong Chai, Hao Sun 0013, Tomoko Tateyama, Lanfen Lin, Yen-Wei Chen 0001
ACM Trans. Comput. Heal.6
2026 An improved multi-instance learning model with clinical-guided cross-attention for postoperative early recurrence prediction of hepatocellular carcinoma using histopathological images
Gan Zhan, Fang Wang 0030, Yinhao Li 0002, Rahul Kumar Jain 0001, Qingqing Chen 0001, Lanfen Lin, Hongjie Hu, C. Krishna Mohan, Yen-Wei Chen 0001
Neurocomputing8
2026 One framework to rule them all: Unifying multimodal tasks with LLM neural-tuning
Hao Sun 0013, Yu Song 0008, Jiaqing Liu, Jihong Hu, Yen-Wei Chen 0001, Lanfen Lin
Pattern Recognit.6
2026 SPA: Leveraging the SAM With Spatial Priors Adapter for Enhanced Medical Image Segmentation
abstract
The Segment Anything Model (SAM) has gained renown for its success in image segmentation, benefiting significantly from its pretraining on extensive datasets and its interactive prompt-based segmentation approach. Although highly effective in natural (real-world) image segmentation tasks, the SAM model encounters significant challenges in medical imaging due to the inherent differences between these two domains. To address these challenges, we propose the Spatial Prior Adapter (SPA) scheme, a parameter-efficient fine-tuning strategy that enhances SAM's adaptability to medical imaging tasks. SPA introduces two novel modules: the Spatial Prior Module (SPM), which captures localized spatial features through convolutional layers, and the Feature Communication Module (FCM), which integrates these features into SAM's image encoder via cross-attention mechanisms. Furthermore, we develop a Multiscale Feature Fusion Module (MSFFM) to enhance SAM's end-to-end segmentation capabilities by effectively aggregating multiscale contextual information. These lightweight modules require minimal computational resources while significantly boosting segmentation performance. Our approach demonstrates superior performance in both prompt-based and end-to-end segmentation scenarios through extensive experiments on publicly available medical imaging datasets. Performance highlights the potential of the proposed method to bridge the gap between foundation models and domain-specific medical imaging tasks. This advancement paves the way for more effective AI-assisted medical diagnostic systems.
Jihong Hu, Yinhao Li 0002, Rahul Kumar Jain 0001, Lanfen Lin, Yen-Wei Chen 0001
IEEE J. Biomed. Health Informatics4
2026 Multimodal Graph Learning With Multi-Hypergraph Reasoning Networks for Focal Liver Lesion Classification in Multimodal Magnetic Resonance Imaging
abstract
Multimodal magnetic resonance imaging (MRI) is instrumental in differentiating liver lesions. The major challenge involves modeling reliable connections and simultaneously learning complementary information across various MRI sequences. While previous studies have primarily focused on multimodal integration in a pair-wise manner using few modalities, our research seeks to advance a more comprehensive understanding of interaction modeling by establishing complex high-order correlations among the diverse modalities in multimodal MRI. In this paper, we introduce a multimodal graph learning with multi-hypergraph reasoning network to capture the full spectrum of both pair-wise and group-wise relationships among different modalities. Specifically, a weight-shared encoder extracts features from regions of interest (ROI) images across all modalities. Subsequently, a collection of uniform hypergraphs are constructed with varying vertex configurations, allowing for the modeling of not only pair-wise correlations but also the high-order collaborations for relational reasoning. Following information propagation through the hypergraph message passing, adaptive intra-modality fusion module is proposed to effectively fuse feature representations from different hypergraphs of the same modality. Finally, all refined features are concatenated to prepare for the classification task. Our experimental evaluations, including focal liver lesions classification using the LLD-MMRI2023 dataset and early recurrence prediction of hepatocellular carcinoma using our internal datasets, demonstrate that our method significantly surpasses the performance of existing approaches, indicating the effectiveness of our model in handling both pair-wise and group-wise interactions across multiple modalities.
Shaocong Mo, Lanfen Lin, Ruofeng Tong 0001, Fang Wang 0030, Qingqing Chen 0001, Wenbin Ji, Yinhao Li 0002, Hongjie Hu, Yen-Wei Chen 0001
IEEE J. Biomed. Health Informatics3
2026 S2Match: Revisiting Weak-to-Strong Consistency From a Semantic Similarity Perspective for Semi-Supervised Medical Image Segmentation
abstract
Semi-supervised learning (SSL) for medical image segmentation is a challenging yet highly practical task, which reduces reliance on large-scale labeled datasets by leveraging unlabeled samples. Among SSL techniques, the weak-to-strong consistency framework, popularized by FixMatch, has emerged as a state-of-the-art method in classification tasks. Notably, such a simple pipeline has also shown competitive performance in medical image segmentation. However, two key limitations still persist, impeding its efficient adaptation: (1) the neglect of contextual dependencies results in inconsistent predictions for similar semantic features, leading to incomplete object segmentation; (2) the lack of exploitation on semantic similarity between labeled and unlabeled data induces considerable class-distribution discrepancy. To address these limitations, we propose a novel SSL framework for medical image segmentation, named S2Match, powered by two appealing designs from a semantic similarity perspective: (1) rectifying pixel-wise prediction by reasoning about the intra-image pair-wise affinity map, thus integrating contextual dependencies explicitly into the final prediction; (2) bridging labeled and unlabeled data via a feature querying mechanism for compact class representation learning, which fully considers cross-image anatomical similarities. As the reliable semantic similarity extraction depends on robust features, we further introduce an effective Spatial-aware Fusion Module (SFM) to explore distinctive information from multiple scales. Experiments show that S2Match yields consistent improvements over the state-of-the-art methods across five public medical image segmentation benchmarks, exhibiting competitive performance on both 2D and 3D tasks.
Shiao Xie, Hongyi Wang 0002, Ziwei Niu, Hao Sun 0013, Shuyi Ouyang, Yen-Wei Chen 0001, Lanfen Lin
IEEE J. Biomed. Health Informatics7
2026 EICSeg: Universal Medical Image Segmentation via Explicit In-Context Learning
abstract
Deep learning models for medical image segmentation often struggle with task-specific characteristics, limiting their generalization to unseen tasks with new anatomies, labels, or modalities. Retraining or fine-tuning these models requires substantial human effort and computational resources. To address this, in-context learning (ICL) has emerged as a promising paradigm, enabling query image segmentation by conditioning on example image-mask pairs provided as prompts. Unlike previous approaches that rely on implicit modeling or non-end-to-end pipelines, we redefine the core interaction mechanism in ICL as an explicit retrieval process, termed E-ICL, benefiting from the emergence of vision foundation models (VFMs). E-ICL captures dense correspondences between queries and prompts at minimal learning cost and leverages them to dynamically weight multi-class prompt masks. Built upon E-ICL, we propose EICSeg, the first end-to-end ICL framework that integrates complementary VFMs for universal medical image segmentation. Specifically, we introduce a lightweight SD-Adapter to bridge the distinct functionalities of the VFMs, enabling more accurate segmentation predictions. To fully exploit the potential of EICSeg, we further design a scalable self-prompt training strategy and an adaptive token-to-image prompt selection mechanism, facilitating both efficient training and inference. EICSeg is trained on 47 datasets covering diverse modalities and segmentation targets. Experiments on nine unseen datasets demonstrate its strong few-shot generalization ability, achieving an average Dice score of 74.0%, outperforming existing in-context and few-shot methods by 4.5%, and reducing the gap to task-specific models to 10.8%. Even with a single prompt, EICSeg achieves a competitive average Dice score of 60.1%. Notably, it performs automatic segmentation without manual prompt engineering, delivering results comparable to interactive models while requiring minimal labeled data. Source code will be available at https://github.com/zerone-fg/EICSeg.
Shiao Xie, Liangjun Zhang, Ziwei Niu, Fanfan Ye, Qiaoyong Zhong, Di Xie, Yen-Wei Chen 0001, Lanfen Lin
IEEE Trans. Medical Imaging8
2026 Disentangled Multimodal Tuning and Interaction for Human Perception Understanding
abstract
Understanding human perceptions poses a significant multimodal challenge for computers, involving textual, acoustic, and visual signals. Recently, large language models (LLMs) have garnered great attention, leading to numerous methods aimed at efficiently fine-tuning pretrained models for multimodal downstream tasks. However, there remains a scarcity of techniques that prioritize modality-invariant and -specific information during parameter-efficient tuning, despite evidence from previous studies showcasing the effectiveness of modality disentangling. To address this gap, we propose a novel multimodal tuning approach for LLMs, termed Disentangled Multimodal Tuning and Interaction. Specifically, we evaluate the independence among different modalities and disentangle corresponding modality-invariant and specific components, which are subsequently leveraged for prompt tuning. Following tuning, a newly designed independence-guided cross-attention module is introduced for modality interaction, where the attention mechanism is decoupled and bolstered with independence from the modality-disentangling process. This approach not only enables LLMs to efficiently assimilate information from various modalities but also cultivates an awareness of both modality-invariant and specific information. Compared to previous methods, our approach facilitates modality interaction at a more granular level, resulting in enhanced performance. We validate our method through experiments on four public datasets, demonstrating significant performance improvements.
Hao Sun 0013, Ziwei Niu, Jiaqing Liu, Yen-Wei Chen 0001, Lanfen Lin
ACM Trans. Multim. Comput. Commun. Appl.6
2025 M2OST: Many-to-one Regression for Predicting Spatial Transcriptomics from Digital Pathology Images
abstract
The advancement of Spatial Transcriptomics (ST) has facilitated the spatially-aware profiling of gene expressions based on histopathology images. Although ST data offers valuable insights into the micro-environment of tumors, its acquisition cost remains expensive. Therefore, directly predicting the ST expressions from digital pathology images is desired. Current methods usually adopt existing regression backbones along with patch-sampling for this task, which ignores the inherent multi-scale information embedded in the pyramidal data structure of digital pathology images, and wastes the inter-spot visual information crucial for accurate gene expression prediction. To address these limitations, we propose M2OST, a many-to-one regression Transformer that can accommodate the hierarchical structure of the pathology images via a decoupled multi-scale feature extractor. Unlike traditional models that are trained with one-to-one image-label pairs, M2OST uses multiple images from different levels of the digital pathology image to jointly predict the gene expressions in their common corresponding spot. Built upon our many-to-one scheme, M2OST can be easily scaled to fit different numbers of inputs, and its network structure inherently incorporates nearby inter-spot features, enhancing regression performance. We have tested M2OST on three public ST datasets and the experimental results show that M2OST can achieve state-of-the-art performance with fewer parameters and floating-point operations (FLOPs).
Hongyi Wang 0002, Xiuju Du, Jing Liu 0041, Shuyi Ouyang, Yen-Wei Chen 0001, Lanfen Lin
AAAI6
2025 Triple-Prompt Controllable Diffusion for Universal Data Augmentation in Medical Image Segmentation
abstract
Medical image segmentation is a crucial yet challenging task in image analysis across diverse anatomical structures. Current segmentation models heavily depend on large-scale datasets, which are laborious to collect and annotate. While generative models offer a promising alternative for data augmentation, most existing approaches are limited to single-modality outputs, either synthetic images or segmentation masks. Moreover, these methods often lack flexible conditioning mechanisms and struggle to capture the rich contextual dependencies inherent in anatomical structures. To address these challenges, in this paper, we propose TPCDM, a novel framework that co-synthesizes high-fidelity paired medical images and segmentation masks through a unified Triple-Prompt Conditional Diffusion Model. At the heart of TPCDM lies a newly defined joint image-label generation paradigm, termed Coordinated Distribution Learning, governed by three synergistic prompts: (1) a text prompt encoding global anatomical semantics; (2) a spatial prompt enforcing pixel-wise spatial coherence; (3) a task prompt dynamically adapting to diverse distributions. Furthermore, TPCDM disentangles instance-wise annotations into semantic masks and distance maps, enabling seamless extension to instance segmentation tasks. Extensive experiments on four benchmarks demonstrate that TPCDM achieves superior synthesis quality. Besides, incorporating the synthesized samples leads to state-of-the-art performance in both downstream semantic and instance segmentation tasks, while also delivering significant improvements under limited labeled data.
Shiao Xie, Hongyi Wang 0002, Liangjun Zhang, Ziwei Niu, Yen-Wei Chen 0001, Lanfen Lin
ECAI7
2025 Enhanced Multimodal Depression Detection With Emotion Prompts
abstract
Depression is a pervasive mental health disorder that remains frequently undiagnosed and untreated due to societal barriers and the subjective nature of its symptoms. Leveraging recent advances in large language models (LLMs), we propose a novel depression detection pipeline that generates emotion prompts tailored to individual data, enhancing detection accuracy. Our approach integrates cross-modality fusion via cross attention mechanisms to combine depressive and emotional features, creating a comprehensive representation of depression indicators. Evaluated on the E-DAIC and EATD datasets, our method outperforms state-of-the-art techniques, demonstrating its potential for more precise emotion-based depression detection.
Shiyu Teng, Jiaqing Liu, Hao Sun 0013, Shurong Chai, Tomoko Tateyama, Lanfen Lin, Yen-Wei Chen 0001
ICASSP6
2025 Region-Aware Anchoring Mechanism for Efficient Referring Visual Grounding
Shuyi Ouyang, Ziwei Niu, Hongyi Wang 0002, Yen-Wei Chen 0001, Lanfen Lin
ICCV5
2025 EPIC: Efficient Prompt Interaction for Text-Image Classification
abstract
In recent years, large-scale pre-trained multimodal models (LMMs) generally emerge to integrate the vision and language modalities, achieving considerable success in multimodal tasks, such as text-image classification. The growing size of LMMs, however, results in a significant computational cost for fine-tuning these models for downstream tasks. Hence, prompt-based interaction strategy is studied to align modalities more efficiently. In this context, we propose a novel efficient prompt-based multimodal interaction strategy, namely Efficient Prompt Interaction for text-image Classification (EPIC). Specifically, we utilize temporal prompts on intermediate layers, and integrate different modalities with similarity-based prompt interaction, to leverage sufficient information exchange between modalities. Utilizing this approach, our method achieves reduced computational resource consumption and fewer trainable parameters (about 1% of the foundation model) compared to other fine-tuning strategies. Furthermore, it demonstrates superior performance on the UPMC-Food101 and SNLI-VE datasets, while achieving comparable performance on the MM-IMDB dataset.
Xinyao Yu 0003, Hao Sun 0013, Zeyu Ling, Ziwei Niu, Zhenjia Bai, Yen-Wei Chen 0001, Lanfen Lin
ICME8
2025 Clinical Data-Driven Retrieval-Augmented Model for Lung Nodule Malignancy Prediction
Ruibo Hou, Shurong Chai, Rahul Kumar Jain 0001, Yinhao Li 0002, Jiaqing Liu, Shiyu Teng, Lanfen Lin, Yen-Wei Chen 0001
MICCAI (10)8
2025 TextBraTS: Text-Guided Volumetric Brain Tumor Segmentation with Innovative Dataset Development and Fusion Module Exploration
Rahul Kumar Jain 0001, Yinhao Li 0002, Ruibo Hou, Jingliang Cheng, Guohua Zhao, Lanfen Lin, Rui Xu 0002, Yen-Wei Chen 0001
MICCAI (6)8
2025 PD-UniST: Prompt-Driven Universal Model for Unpaired H&E-to-IHC Stain Translation
Chujie Zhang, Yangyang Xie, Yinhao Li 0002, Lanfen Lin, Yen-Wei Chen 0001
MICCAI (2)5
2025 EIR-SDG: Explore Invariant Representation for Single-source Domain Generalization in Medical Image Segmentation
Ziwei Niu, Shiao Xie, Ziyue Wang 0005, Yen-Wei Chen 0001, Yueming Jin, Lanfen Lin
ACM Multimedia6
2025 Multi-modal Medical SAM: An Adaptation Method of Segment Anything Model (SAM) for Glioma Segmentation Using Multi-modal MR Images
abstract
The segmentation of glioma is crucial for early diagnosis, according to a World Health Organization (WHO) 2021 report. For glioma diagnosis, 3D multi-modal brain MRI/CT imaging has become an essential tool, offering detailed information. Nowadays, deep learning frameworks have been applied to various medical imaging problems, including brain glioma segmentation. Recently, foundation models like Segment Anything Model (SAM) have emerged as pivotal tools in computer vision tasks. These models are trained using large (real-world) datasets, offering a generalized understanding of visual data and semantic key features. Therefore, the effective utilization of foundation models in medical imaging is a significant area of current research. However, the differences in data distribution between multi-modal medical images and real-world images present challenges in directly applying foundation models to medical imaging. Additionally, utilizing multi-modal images to extract crucial information and its fusion poses further challenges. To address these issues, we propose a framework using foundation model and novel strategies for multi-modal fusion. Our fusion adapters effectively integrate the information from different modalities to enhance glioma segmentation in multi-modal MRI scans. Our method outperforms current state-of-the-art methods for accurate segmentation of the glioma using private and publicly available brain MRI datasets, proving the effectiveness of our approach across different datasets and imaging modalities.
Rahul Kumar Jain 0001, Yinhao Li 0002, Shurong Chai, Jingliang Cheng, Guohua Zhao, Lanfen Lin, Yen-Wei Chen 0001
ACM Trans. Comput. Heal.8
2025 Multimodal Sentiment Analysis With Mutual Information-Based Disentangled Representation Learning
abstract
Multimodal sentiment analysis seeks to utilize various types of signals to identify underlying emotions and sentiments. A key challenge in this field lies in multimodal representation learning, which aims to develop effective methods for integrating multimodal features into cohesive representations. Recent advancements include two notable approaches: one focuses on decomposing multimodal features into modality-invariant and -specific components, while the other emphasizes the use of mutual information to enhance the fusion of modalities. Both strategies have demonstrated effectiveness and yielded remarkable results. In this paper, we propose a novel learning framework that combines the strengths of these two approaches, termed mutual information-based disentangled multimodal representation learning. Our approach involves estimating different types of information during feature extraction and fusion stages. Specifically, we quantitatively assess and adjust the proportions of modality-invariant, -specific, and -complementary information during feature extraction. Subsequently, during fusion, we evaluate the amount of information retained by each modality in the fused representation. We employ mutual information or conditional mutual information to estimate each type of information content. By reconciling the proportions of these different types of information, our approach achieves state-of-the-art performance on popular sentiment analysis benchmarks, including CMU-MOSI and CMU-MOSEI.
Hao Sun 0013, Ziwei Niu, Hongyi Wang 0002, Xinyao Yu 0003, Jiaqing Liu, Yen-Wei Chen 0001, Lanfen Lin
IEEE Trans. Affect. Comput.7
2025 SAMA: A Self-and-Mutual Attention Network for Accurate Recurrence Prediction of Non-Small Cell Lung Cancer Using Genetic and CT Data
abstract
Accurate preoperative recurrence prediction for non-small cell lung cancer (NSCLC) is a challenging issue in the medical field. Existing studies primarily conduct image and molecular analyses independently or directly fuse multimodal information through radiomics and genomics, which fail to fully exploit and effectively utilize the highly heterogeneous cross-modal information at different levels and model the complex relationships between modalities, resulting in poor fusion performance and becoming the bottleneck of precise recurrence prediction. To address these limitations, we propose a novel unified framework, the Self-and-Mutual Attention (SAMA) Network, designed to efficiently fuse and utilize macroscopic CT images and microscopic gene data for precise NSCLC recurrence prediction, integrating handcrafted features, deep features, and gene features. Specifically, we design a Self-and-Mutual Attention Module that performs three-stage fusion: the self-enhancement stage enhances modality-specific features; the gene-guided and CT-guided cross-modality fusion stages perform bidirectional cross-guidance on the self-enhanced features, complementing and refining each modality, enhancing heterogeneous feature expression; and the optimized feature aggregation stage ensures the refined interactive features for precise prediction. Extensive experiments on both publicly available datasets from The Cancer Imaging Archive (TCIA) and The Cancer Genome Atlas (TCGA) demonstrate that our method achieves state-of-the-art performance and exhibits broad applicability to various cancers.
Yang Ai, Jing Liu 0041, Yinhao Li 0002, Fang Wang 0030, Xiuju Du, Rahul Kumar Jain 0001, Lanfen Lin, Yen-Wei Chen 0001
IEEE J. Biomed. Health Informatics7
2024 Combinatorial CNN-Transformer Learning with Manifold Constraints for Semi-supervised Medical Image Segmentation
abstract
Semi-supervised learning (SSL), as one of the dominant methods, aims at leveraging the unlabeled data to deal with the annotation dilemma of supervised learning, which has attracted much attentions in the medical image segmentation. Most of the existing approaches leverage a unitary network by convolutional neural networks (CNNs) with compulsory consistency of the predictions through small perturbations applied to inputs or models. The penalties of such a learning paradigm are that (1) CNN-based models place severe limitations on global learning; (2) rich and diverse class-level distributions are inhibited. In this paper, we present a novel CNN-Transformer learning framework in the manifold space for semi-supervised medical image segmentation. First, at intra-student level, we propose a novel class-wise consistency loss to facilitate the learning of both discriminative and compact target feature representations. Then, at inter-student level, we align the CNN and Transformer features using a prototype-based optimal transport method. Extensive experiments show that our method outperforms previous state-of-the-art methods on three public medical image segmentation benchmarks.
Huimin Huang 0002, Yawen Huang, Shiao Xie, Lanfen Lin, Ruofeng Tong 0001, Yen-Wei Chen 0001, Yuexiang Li, Yefeng Zheng 0001
AAAI4
2024 Going Beyond Multi-Task Dense Prediction with Synergy Embedding Models
abstract
Multi-task visual scene understanding aims to leverage the relationships among a set of correlated tasks, which are solved simultaneously by embedding them within a unified network. However, most existing methods give rise to two primary concerns from a task-level perspective: (1) the lack of task-independent correspondences for distinct tasks, and (2) the neglect of explicit task-consensual dependencies among various tasks. To address these issues, we propose a novel synergy embedding models (SEM), which goes beyond multi-task dense prediction by leveraging two innovative designs: the intra-task hierarchy-adaptive module and the inter-task EM-interactive module. Specifically, the constructed intra-task module incorporates hierarchy-adaptive keys from multiple stages, enabling the efficient learning of specialized visual patterns with an optimal trade-off. In addition, the developed inter-task module learns interactions from a compact set of mutual bases among various tasks, benefiting from the expectation maximization (EM) algorithm. Extensive empirical evidence from two public benchmarks, NYUD-v2 and PASCAL-Context, demonstrates that SEM consistently outperforms state-of-the-art approaches across a range of metrics.
Huimin Huang 0002, Yawen Huang, Lanfen Lin, Ruofeng Tong 0001, Yen-Wei Chen 0001, Hao Zheng 0008, Yuexiang Li, Yefeng Zheng 0001
CVPR3
2024 IRLSG: Invariant Representation Learning for Single-Domain Generalization in Medical Image Segmentation
abstract
Single-domain generalization (SDG) can efficiently enhance model generalization while avoiding high annotation costs and privacy concerns. However, existing SDG methods are mainly based on data manipulation and meta-learning, which are not efficient enough due to the limited generalization performance and complex inference. In response to these challenges, we present a novel single domaininvariant representation learning approach for medical image segmentation, called IRLSG, with two appealing designs: (1) A Classscale Photo-metric Augmentation is first proposed to simulate unseen target domain that is sufficient in diversity and informativeness. After that, a Dual-Consistency Framework is further designed to constrain the consistency of intermediate features and segmentation results between the original and the augmented images, which helps to explore the domain-invariant representation. (2) A simple and effective Style Feature Whitening is designed to decouple and remove the domain-specific style from higher-order covariance statistics, which can further improve the modeling and generalization capability of the network. Experimental results on different benchmarks demonstrate that our IRLSG outperforms the current state-of-the-art methods in tackling single-domain generalization.
Ziwei Niu, Hao Sun 0013, Shuyi Ouyang, Shiao Xie, Yen-Wei Chen 0001, Ruofeng Tong 0001, Lanfen Lin
ICASSP7
2024 A Novel Adaptive Hypergraph Neural Network for Enhancing Medical Image Segmentation
Shurong Chai, Rahul Kumar Jain 0001, Shaocong Mo, Jiaqing Liu, Yinhao Li 0002, Tomoko Tateyama, Lanfen Lin, Yen-Wei Chen 0001
MICCAI (9)8
2024 LGA: A Language Guide Adapter for Advancing the SAM Model's Capabilities in Medical Image Segmentation
Jihong Hu, Yinhao Li 0002, Hao Sun 0013, Yu Song 0008, Chujie Zhang, Lanfen Lin, Yen-Wei Chen 0001
MICCAI (12)6
2024 Segmentation Guided Crossing Dual Decoding Generative Adversarial Network for Synthesizing Contrast-Enhanced Computed Tomography Images
abstract
Although contrast-enhanced computed tomography (CE-CT) images significantly improve the accuracy of diagnosing focal liver lesions (FLLs), the administration of contrast agents imposes a considerable physical burden on patients. The utilization of generative models to synthesize CE-CT images from non-contrasted CT images offers a promising solution. However, existing image synthesis models tend to overlook the importance of critical regions, inevitably reducing their effectiveness in downstream tasks. To overcome this challenge, we propose an innovative CE-CT image synthesis model called Segmentation Guided Crossing Dual Decoding Generative Adversarial Network (SGCDD-GAN). Specifically, the SGCDD-GAN involves a crossing dual decoding generator including an attention decoder and an improved transformation decoder. The attention decoder is designed to highlight some critical regions within the abdominal cavity, while the improved transformation decoder is responsible for synthesizing CE-CT images. These two decoders are interconnected using a crossing technique to enhance each other's capabilities. Furthermore, we employ a multi-task learning strategy to guide the generator to focus more on the lesion area. To evaluate the performance of proposed SGCDD-GAN, we test it on an in-house CE-CT dataset. In both CE-CT image synthesis tasks-namely, synthesizing ART images and synthesizing PV images-the proposed SGCDD-GAN demonstrates superior performance metrics across the entire image and liver region, including SSIM, PSNR, MSE, and PCC scores. Furthermore, CE-CT images synthetized from our SGCDD-GAN achieve remarkable accuracy rates of 82.68%, 94.11%, and 94.11% in a deep learning-based FLLs classification task, along with a pilot assessment conducted by two radiologists.
Qingqing Chen 0001, Yinhao Li 0002, Fang Wang 0030, Xianhua Han, Yutaro Iwamoto, Jing Liu 0041, Lanfen Lin, Hongjie Hu, Yen-Wei Chen 0001
IEEE J. Biomed. Health Informatics8
2024 Rethinking Multiple Instance Learning for Whole Slide Image Classification: A Bag-Level Classifier is a Good Instance-Level Teacher
abstract
Multiple Instance Learning (MIL) has demonstrated promise in Whole Slide Image (WSI) classification. However, a major challenge persists due to the high computational cost associated with processing these gigapixel images. Existing methods generally adopt a two-stage approach, comprising a non-learnable feature embedding stage and a classifier training stage. Though it can greatly reduce memory consumption by using a fixed feature embedder pre-trained on other domains, such a scheme also results in a disparity between the two stages, leading to suboptimal classification accuracy. To address this issue, we propose that a bag-level classifier can be a good instance-level teacher. Based on this idea, we design Iteratively Coupled Multiple Instance Learning (ICMIL) to couple the embedder and the bag classifier at a low cost. ICMIL initially fixes the patch embedder to train the bag classifier, followed by fixing the bag classifier to fine-tune the patch embedder. The refined embedder can then generate better representations in return, leading to a more accurate classifier for the next iteration. To realize more flexible and more effective embedder fine-tuning, we also introduce a teacher-student framework to efficiently distill the category knowledge in the bag classifier to help the instance-level embedder fine-tuning. Intensive experiments were conducted on four distinct datasets to validate the effectiveness of ICMIL. The experimental results consistently demonstrated that our method significantly improves the performance of existing MIL backbones, achieving state-of-the-art results. The code and the organized datasets can be accessed by: https://github.com/Dootmaan/ICMIL/tree/confidence-based.
Hongyi Wang 0002, Luyang Luo, Fang Wang 0030, Ruofeng Tong 0001, Yen-Wei Chen 0001, Hongjie Hu, Lanfen Lin, Hao Chen 0011
IEEE Trans. Medical Imaging7
2024 Knowledge Distillation-Based Domain-Invariant Representation Learning for Domain Generalization
abstract
Domain generalization (DG) aims to generalize the knowledge learned from multiple source domains to unseen target domains. Existing DG techniques can be subsumed under two broad categories, i.e., domain-invariant representation learning and domain manipulation. Nevertheless, it is extremely difficult to explicitly augment or generate the unseen target data. And when source domain variety increases, developing a domain-invariant model by simply aligning more domain-specific information becomes more challenging. In this paper, we propose a simple yet effective method for domain generalization, named Knowledge Distillation based Domain-invariant Representation Learning (KDDRL), that learns domain-invariant representation while encouraging the model to maintain domain-specific features, which recently turned out to be effective for domain generalization. To this end, our method incorporates multiple auxiliary student models and one student leader model to perform a two-stage distillation. In the first-stage distillation, each domain-specific auxiliary student treats the ensemble of other auxiliary students' predictions as a target, which helps to excavate the domain-invariant representation. Also, we present an error removal module to prevent the transfer of faulty information by eliminating incorrect predictions compared to the true labels. In the second-stage distillation, the student leader model with domain-specific features combines the domain-invariant representation learned from the group of auxiliary students to make the final prediction. Extensive experiments and in-depth analysis on popular DG benchmark datasets demonstrate that our KDDRL significantly outperforms the current state-of-the-art methods.
Ziwei Niu, Junkun Yuan, Jing Liu 0041, Yen-Wei Chen 0001, Ruofeng Tong 0001, Lanfen Lin
IEEE Trans. Multim.8
2023 ClassFormer: Exploring Class-Aware Dependency with Transformer for Medical Image Segmentation
abstract
Vision Transformers have recently shown impressive performances on medical image segmentation. Despite their strong capability of modeling long-range dependencies, the current methods still give rise to two main concerns in a class-level perspective: (1) intra-class problem: the existing methods lacked in extracting class-specific correspondences of different pixels, which may lead to poor object coverage and/or boundary prediction; (2) inter-class problem: the existing methods failed to model explicit category-dependencies among various objects, which may result in inaccurate localization. In light of these two issues, we propose a novel transformer, called ClassFormer, powered by two appealing transformers, i.e., intra-class dynamic transformer and inter-class interactive transformer, to address the challenge of fully exploration on compactness and discrepancy. Technically, the intra-class dynamic transformer is first designed to decouple representations of different categories with an adaptive selection mechanism for compact learning, which optimally highlights the informative features to reflect the salient keys/values from multiple scales. We further introduce the inter-class interactive transformer to capture the category dependency among different objects, and model class tokens as the representative class centers to guide a global semantic reasoning. As a consequence, the feature consistency is ensured with the expense of intra-class penalization, while inter-class constraint strengthens the feature discriminability between different categories. Extensive empirical evidence shows that ClassFormer can be easily plugged into any architecture, and yields improvements over the state-of-the-art methods in three public benchmarks.
Huimin Huang 0002, Shiao Xie, Lanfen Lin, Ruofeng Tong 0001, Yen-Wei Chen 0001, Hong Wang 0021, Yuexiang Li, Yawen Huang, Yefeng Zheng 0001
AAAI3
2023 SemiCVT: Semi-Supervised Convolutional Vision Transformer for Semantic Segmentation
abstract
Semi-supervised learning improves data efficiency of deep models by leveraging unlabeled samples to alleviate the reliance on a large set of labeled samples. These successes concentrate on the pixel-wise consistency by using convolutional neural networks (CNNs) but fail to address both global learning capability and class-level features for unlabeled data. Recent works raise a new trend that Transformer achieves superior performance on the entire feature map in various tasks. In this paper, we unify the current dominant Mean-Teacher approaches by reconciling intra-model and inter-model properties for semi-supervised segmentation to produce a novel algorithm, SemiCVT, that absorbs the quintessence of CNNs and Transformer in a comprehensive way. Specifically, we first design a parallel CNN-Transformer architecture (CVT) with introducing an intra-model local-global interaction schema (LGI) in Fourier domain for full integration. The inter-model class-wise consistency is further presented to complement the class-level statistics of CNNs and Transformer in a cross-teaching manner. Extensive empirical evidence shows that SemiCVT yields consistent improvements over the state-of-the-art methods in two public benchmarks.
Huimin Huang 0002, Shiao Xie, Lanfen Lin, Ruofeng Tong 0001, Yen-Wei Chen 0001, Yuexiang Li, Hong Wang 0021, Yawen Huang, Yefeng Zheng 0001
CVPR3
2023 MCKD: Mutually Collaborative Knowledge Distillation For Federated Domain Adaptation And Generalization
abstract
Conventional unsupervised domain adaptation (UDA) and domain generalization (DG) methods rely on the assumption that all source domains can be directly accessed and combined for model training. However, this centralized training strategy may violate privacy policies in many real-world applications. A paradigm for tackling this problem is to train multiple local models and aggregate a generalized central model without data sharing. Recent methods have made remarkable advancements in this paradigm by exploiting parameter alignment and aggregation. But when sources domain variety increases, directly aligning and aggregating local parameters becomes more challenging. Adapting a different approach in this work, we devised a data-free semantic collaborative distillation strategy to learn domain-invariant representation for both federated UDA and DG. Each local model transmits its predictions to the central server and derives its target distribution from the average of other local models' distributions to facilitate the mutual transfer of domain-specific knowledge. When unlabeled target data is available, we introduce a novel UDA strategy termed knowledge filter to adapt the central model to the target data. Extensive experiments on four UDA and DG datasets demonstrate that our method has a competitive performance compared with the state-of-the-art methods.
Ziwei Niu, Hongyi Wang 0002, Hao Sun 0013, Shuyi Ouyang, Yen-Wei Chen 0001, Lanfen Lin
ICASSP6
2023 MedFCT: A Frequency Domain Joint CNN-Transformer Network for Semi-supervised Medical Image Segmentation
abstract
Semi-supervised learning(SSL) is a data-efficient way in leveraging large-scale data without annotations and alleviating the dependence on labeled data. Mean-Teacher (MT) scheme with teacher-student model architecture has shown its effectiveness in semi-supervised medical image segmentation, where the student network learns from the teacher by minimizing pixel-wise consistency loss. However, existing MT-based SSLs still give rise to two main concerns: (1) limited learning ability of student network that neglects the union of local feature and global cues extraction which may impact the representation learning of variable objects. (2) limited knowledge-transferring ability of teacher network with only pixel-level consistency regularization that may result in inadequate and unstable guidance. To address these limitations, we propose a novel semi-supervised learning scheme, namely MedFCT, with two appealing designs: (1) A dual student architecture with parallel CNN and Transformer branches is designed for local-global feature extraction, where the full-frequency interaction between CNN and Transformer can be explored by a frequency domain cross-fusion (FDCF) module to learn complementarity of the two-paradigm features. (2) A comprehensive multi-level consistency regularization considering pixel-wise, feature-wise and class-wise information is presented to realize more effective guidance and knowledge transfer from teacher network. Experiments show that MedFCT outperforms previous state-of-the-art methods on two public medical image segmentation benchmarks.
Shiao Xie, Huimin Huang 0002, Ziwei Niu, Lanfen Lin, Yen-Wei Chen 0001
ICME4
2023 SLViT: Scale-Wise Language-Guided Vision Transformer for Referring Image Segmentation
abstract
Referring image segmentation aims to segment an object out of an image via a specific language expression. The main concept is establishing global visual-linguistic relationships to locate the object and identify boundaries using details of the image. Recently, various Transformer-based techniques have been proposed to efficiently leverage long-range cross-modal dependencies, enhancing performance for referring segmentation. However, existing methods consider visual feature extraction and cross-modal fusion separately, resulting in insufficient visual-linguistic alignment in semantic space. In addition, they employ sequential structures and hence lack multi-scale information interaction. To address these limitations, we propose a Scale-Wise Language-Guided Vision Transformer (SLViT) with two appealing designs: (1) Language-Guided Multi-Scale Fusion Attention, a novel attention mechanism module for extracting rich local visual information and modeling global visual-linguistic relationships in an integrated manner. (2) An Uncertain Region Cross-Scale Enhancement module that can identify regions of high uncertainty using linguistic features and refine them via aggregated multi-scale features. We have evaluated our method on three benchmark datasets. The experimental results demonstrate that SLViT surpasses state-of-the-art methods with lower computational cost. The code is publicly available at: https://github.com/NaturalKnight/SLViT.
Shuyi Ouyang, Hongyi Wang 0002, Shiao Xie, Ziwei Niu, Ruofeng Tong 0001, Yen-Wei Chen 0001, Lanfen Lin
IJCAI7
2023 Iteratively Coupled Multiple Instance Learning from Instance to Bag Classifier for Whole Slide Image Classification
Hongyi Wang 0002, Luyang Luo, Fang Wang 0030, Ruofeng Tong 0001, Yen-Wei Chen 0001, Hongjie Hu, Lanfen Lin, Hao Chen 0011
MICCAI (6)7
2023 Semi-Supervised Convolutional Vision Transformer with Bi-Level Uncertainty Estimation for Medical Image Segmentation
abstract
Semi-supervised learning (SSL) has attracted much attention in the field of medical image segmentation, which enables to alleviate the heavy burden of labelling pixel-wise annotation by extracting knowledge from unlabeled data. The existing methods basically benefit from the success of convolutional neural networks (CNNs) by keeping consistency of the predictions under small perturbations imposed on the networks or inputs. Two main concerns arise when learning such a paradigm: (1) CNNs tend to retain discriminative local features, neglecting global dependency and thus leading to inaccurate localization; (2) CNNs omit reliable feature-level and pixel-level information, resulting in sketchy pseudo-labels, especially around the confusing boundary. In this paper, we revisit the model of semi-supervised learning and develop a novel CNN-Transformer learning framework that allows for effective segmentation of medical images by producing complementary and reliable features and pseudo-label with bi-level uncertainty. Motivated by the uncertainty estimation to gain insight on feature discrimination, we explore the statistical and geometrical properties of features on network optimization and thus launching an alignment method in a more accurate and stable way. We attach equal significance to pixel-level uncertainty estimation for alleviating the influence of unreliable pseudo-labels in the training progress and advocating the reliability of predictions. Experimental results show that our method significantly surpasses existing semi-supervised approaches on two public medical image segmentation datasets.
Huimin Huang 0002, Yawen Huang, Shiao Xie, Lanfen Lin, Ruofeng Tong 0001, Yen-Wei Chen 0001, Yuexiang Li, Yefeng Zheng 0001
ACM Multimedia4
2023 HSVLT: Hierarchical Scale-Aware Vision-Language Transformer for Multi-Label Image Classification
abstract
The task of multi-label image classification involves recognizing multiple objects within a single image. Considering both valuable semantic information contained in the labels and essential visual features presented in the image, tight visual-linguistic interactions play a vital role in improving classification performance. Moreover, given the potential variance in object size and appearance within a single image, attention to features of different scales can help to discover possible objects in the image. Recently, Transformer-based methods have achieved great success in multi-label image classification by leveraging the advantage of modeling long-range dependencies, but they have several limitations. Firstly, existing methods treat visual feature extraction and cross-modal fusion as separate steps, resulting in insufficient visual-linguistic alignment in the joint semantic space. Additionally, they only extract visual features and perform cross-modal fusion at a single scale, neglecting objects with different characteristics. To address these issues, we propose a Hierarchical Scale-Aware Vision-Language Transformer (HSVLT) with two appealing designs: (1)A hierarchical multi-scale architecture that involves a Cross-Scale Aggregation module, which leverages joint multi-modal features extracted from multiple scales to recognize objects of varying sizes and appearances in images. (2)Interactive Visual-Linguistic Attention, a novel attention mechanism module that tightly integrates cross-modal interaction, enabling the joint updating of visual, linguistic and multi-modal features. We have evaluated our method on three benchmark datasets. The experimental results demonstrate that HSVLT surpasses state-of-the-art methods with lower computational cost.
Shuyi Ouyang, Hongyi Wang 0002, Ziwei Niu, Zhenjia Bai, Shiao Xie, Ruofeng Tong 0001, Yen-Wei Chen 0001, Lanfen Lin
ACM Multimedia9
2023 IS2Net: Intra-domain Semantic and Inter-domain Style Enhancement for Semi-supervised Medical Domain Generalization
abstract
Domain generalization (DG) demonstrates superior generalization ability in cross-center medical image segmentation. Despite its great success, existing fully supervised DG methods require collecting a large quantity of pixel-level annotations which is quite expensive and time-consuming. To address this challenge, several semi-supervised domain generalized (SSDG) methods have been proposed by simply coupling semi-supervised learning (SSL) with DG tasks, which give rise to two main concerns: (1) Intra-domain dubious semantic information: the quality of pseudo labels in each source domain suffers from the limited amount of labeled data and cross-domain discrepancy. (2) Inter-domain intangible style relationship: current models fail in integrating domain-level information and overlook the relationships among different domains, which degrades the generalization ability of model. In light of these two issues, we propose a novel SSDG framework, namely IS2Net, by arranging an inter-domain generalization branch and several intra-domain SSL branches in a parallel manner, powered by two appealing designs that build a positive interaction between them: (1) A style and semantic memory mechanism is designed to provide both high-quality class-wise representations for intra-domain semantic enhancement and stable domain-specific knowledge for inter-domain style relationship construction. (2) Confident pseudo labeling strategy aims at generating more reliable supervision for intra and inter domain branches, and thus facilitating the learning process of the whole framework. Extensive experiments show that IS2Net yields consistent improvements over the state-of- the-art methods in three public benchmarks.
Shiao Xie, Ziwei Niu, Huimin Huang 0002, Hao Sun 0013, Yen-Wei Chen 0001, Lanfen Lin
ACM Multimedia7
2023 HAP: Structure-Aware Masked Image Modeling for Human-Centric Perception
abstract
Model pre-training is essential in human-centric perception. In this paper, we first introduce masked image modeling (MIM) as a pre-training approach for this task. Upon revisiting the MIM training strategy, we reveal that human structure priors offer significant potential. Motivated by this insight, we further incorporate an intuitive human structure prior - human parts - into pre-training. Specifically, we employ this prior to guide the mask sampling process. Image patches, corresponding to human part regions, have high priority to be masked out. This encourages the model to concentrate more on body structure information during pre-training, yielding substantial benefits across a range of human-centric perception tasks. To further capture human characteristics, we propose a structure-invariant alignment loss that enforces different masked views, guided by the human part prior, to be closely aligned for the same image. We term the entire method as HAP. HAP simply uses a plain ViT as the encoder yet establishes new state-of-the-art performance on 11 human-centric benchmarks, and on-par result on one dataset. For example, HAP achieves 78.1% mAP on MSMT17 for person re-identification, 86.54% mA on PA-100K for pedestrian attribute recognition, 78.2% AP on MS COCO for 2D pose estimation, and 56.0 PA-MPJPE on 3DPW for 3D pose and shape estimation.
Junkun Yuan, Xinyu Zhang 0015, Hao Zhou 0039, Jian Wang 0066, Zhongwei Qiu, Zhiyin Shao, Shaofeng Zhang, Sifan Long 0001, Kun Kuang 0001, Junyu Han, Errui Ding, Lanfen Lin, Fei Wu 0001, Jingdong Wang 0001
NeurIPS13
2023 Domain-Specific Bias Filtering for Single Labeled Domain Generalization
Junkun Yuan, Defang Chen 0001, Kun Kuang 0001, Fei Wu 0001, Lanfen Lin
Int. J. Comput. Vis.6
2023 TensorFormer: A Tensor-Based Multimodal Transformer for Multimodal Sentiment Analysis and Depression Detection
abstract
Sentiment analysis is an important research field aiming to extract and fuse sentimental information from human utterances. Due to the diversity of human sentiment, analyzing from multiple modalities is usually more accurate than from a single modality. To complement the information between related modalities, one effective approach is performing cross-modality interactions. Recently, Transformer-based frameworks have shown a strong ability to capture long-range dependencies, leading to the introduction of several Transformer-based approaches for multimodal processing. However, due to the built-in attention mechanism of the Transformers, only two modalities can be engaged at once. As a result, the complementary information flow in these Transformer-based techniques is partial and constrained. To mitigate this, we propose, TensorFormer, a tensor-based multimodal Transformer framework that takes into account all relevant modalities for interactions. More precisely, we first construct a tensor utilizing the features extracted from each modality, assuming one modality is the target while the remaining tensors serve as the sources. We can generate the corresponding interacted features by calculating source-target attention. This strategy interacts with all involved modalities and generates complementing global information. Experiments on multimodal sentiment analysis benchmark datasets demonstrated the effectiveness of TensorFormer. In addition, we also evaluate TensorFormer in another related area: depression detection and the results reveal significant improvements when compared to other state-of-the-art methods.
Hao Sun 0013, Yen-Wei Chen 0001, Lanfen Lin
IEEE Trans. Affect. Comput.3
2023 Adaptive Decomposition and Shared Weight Volumetric Transformer Blocks for Efficient Patch-Free 3D Medical Image Segmentation
abstract
High resolution (HR) 3D medical image segmentation is vital for an accurate diagnosis. However, in the field of medical imaging, it is still a challenging task to achieve a high segmentation performance with cost-effective and feasible computation resources. Previous methods commonly use patch-sampling to reduce the input size, but this inevitably harms the global context and decreases the model's performance. In recent years, a few patch-free strategies have been presented to deal with this issue, but either they have limited performance due to their over-simplified model structures or they follow a complicated training process. In this study, to effectively address these issues, we present Adaptive Decomposition (A-Decomp) and Shared Weight Volumetric Transformer Blocks (SW-VTB). A-Decomp can adaptively decompose features and reduce their spatial size, which greatly lowers GPU memory consumption. SW-VTB is able to capture long-range dependencies at a low cost with its lightweight design and cross-scale weight-sharing mechanism. Our proposed cross-scale weight-sharing approach enhances the network's ability to capture scale-invariant core semantic information in addition to reducing parameter numbers. By combining these two designs together, we present a novel patch-free segmentation framework named VolumeFormer. Experimental results on two datasets show that VolumeFormer outperforms existing patch-based and patch-free methods with a comparatively fast inference speed and relatively compact design.
Hongyi Wang 0002, Qingqing Chen 0001, Ruofeng Tong 0001, Yen-Wei Chen 0001, Hongjie Hu, Lanfen Lin
IEEE J. Biomed. Health Informatics7
2023 Instrumental Variable-Driven Domain Generalization with Unobserved Confounders
abstract
Domain generalization (DG) aims to learn from multiple source domains a model that can generalize well on unseen target domains. Existing DG methods mainly learn the representations with invariant marginal distribution of the input features, however, the invariance of the conditional distribution of the labels given the input features is more essential for unknown domain prediction. Meanwhile, the existing of unobserved confounders which affect the input features and labels simultaneously cause spurious correlation and hinder the learning of the invariant relationship contained in the conditional distribution. Interestingly, with a causal view on the data generating process, we find that the input features of one domain are valid instrumental variables for other domains. Inspired by this finding, we propose an instrumental variable-driven DG method (IV-DG) by removing the bias of the unobserved confounders with two-stage learning. In the first stage, it learns the conditional distribution of the input features of one domain given input features of another domain. In the second stage, it estimates the relationship by predicting labels with the learned conditional distribution. Theoretical analyses and simulation experiments show that it accurately captures the invariant relationship. Extensive experiments on real-world datasets demonstrate that IV-DG method yields state-of-the-art results.
Junkun Yuan, Ruoxuan Xiong, Mingming Gong, Fei Wu 0001, Lanfen Lin, Kun Kuang 0001
ACM Trans. Knowl. Discov. Data7
2023 Collaborative Semantic Aggregation and Calibration for Federated Domain Generalization
abstract
Domain generalization (DG) aims to learn from multiple known source domains a model that can generalize well to unknown target domains. The existing DG methods usually exploit the fusion of shared multi-source data to train a generalizable model. However, tremendous data is distributed across lots of places nowadays that can not be shared due to privacy policies. In this paper, we tackle the problem of federated domain generalization where the source datasets can only be accessed and learned locally for privacy protection. We propose a novel framework called Collaborative Semantic Aggregation and Calibration (CSAC) to enable this challenging problem. To fully absorb multi-source semantic information while avoiding unsafe data fusion, we conduct data-free semantic aggregation by fusing the models trained on the separated domains layer-by-layer. To address the semantic dislocation problem caused by domain shift, we further design cross-layer semantic calibration with an attention mechanism to align each semantic level and enhance domain invariance. We unify multi-source semantic learning and alignment in a collaborative way by repeating the semantic aggregation and calibration alternately, keeping each dataset localized, and the data privacy is carefully protected. Extensive experiments show the significant performance of our method in addressing this challenging problem.
Junkun Yuan, Defang Chen 0001, Fei Wu 0001, Lanfen Lin, Kun Kuang 0001
IEEE Trans. Knowl. Data Eng.5
2023 Multi-Modal Tumor Segmentation With Deformable Aggregation and Uncertain Region Inpainting
abstract
Multi-modal tumor segmentation exploits complementary information from different modalities to help recognize tumor regions. Known multi-modal segmentation methods mainly have deficiencies in two aspects: First, the adopted multi-modal fusion strategies are built upon well-aligned input images, which are vulnerable to spatial misalignment between modalities (caused by respiratory motions, different scanning parameters, registration errors, etc). Second, the performance of known methods remains subject to the uncertainty of segmentation, which is particularly acute in tumor boundary regions. To tackle these issues, in this paper, we propose a novel multi-modal tumor segmentation method with deformable feature fusion and uncertain region refinement. Concretely, we introduce a deformable aggregation module, which integrates feature alignment and feature aggregation in an ensemble, to reduce inter-modality misalignment and make full use of cross-modal information. Moreover, we devise an uncertain region inpainting module to refine uncertain pixels using neighboring discriminative features. Experiments on two clinical multi-modal tumor datasets demonstrate that our method achieves promising tumor segmentation results and outperforms state-of-the-art methods.
Yue Zhang 0042, Chengtao Peng, Ruofeng Tong 0001, Lanfen Lin, Yen-Wei Chen 0001, Qingqing Chen 0001, Hongjie Hu, Shaohua Kevin Zhou
IEEE Trans. Medical Imaging4
2022 Mixed Transformer U-Net for Medical Image Segmentation
abstract
Though U-Net has achieved tremendous success in medical image segmentation tasks, it lacks the ability to explicitly model long-range dependencies. Therefore, Vision Transformers have emerged as alternative segmentation structures recently, for their innate ability of capturing long-range correlations through Self-Attention (SA). However, Transformers usually rely on large-scale pre-training and have high computational complexity. Furthermore, SA can only model self-affinities within a single sample, ignoring the potential correlations of the overall dataset. To address these problems, we propose a novel Transformer module named Mixed Transformer Module (MTM) for simultaneous inter- and intra- affinities learning. MTM first calculates self-affinities efficiently through our well-designed Local-Global Gaussian-Weighted Self-Attention (LGG-SA). Then, it mines inter-connections between data samples through External Attention (EA). By using MTM, we construct a U-shaped model named Mixed Transformer U-Net (MT-UNet) for accurate medical image segmentation. We test our method on two different public datasets, and the experimental results show that the proposed method achieves better performance over other state-of-the-art methods. The code is available at: https://github.com/Dootmaan/MT-UNet.
Hongyi Wang 0002, Shiao Xie, Lanfen Lin, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Ruofeng Tong 0001
ICASSP3
2022 ScaleFormer: Revisiting the Transformer-based Backbones from a Scale-wise Perspective for Medical Image Segmentation
abstract
Recently, a variety of vision transformers have been developed as their capability of modeling long-range dependency. In current transformer-based backbones for medical image segmentation, convolutional layers were replaced with pure transformers, or transformers were added to the deepest encoder to learn global context. However, there are mainly two challenges in a scale-wise perspective: (1) intra-scale problem: the existing methods lacked in extracting local-global cues in each scale, which may impact the signal propagation of small objects; (2) inter-scale problem: the existing methods failed to explore distinctive information from multiple scales, which may hinder the representation learning from objects with widely variable size, shape and location. To address these limitations, we propose a novel backbone, namely ScaleFormer, with two appealing designs: (1) A scale-wise intra-scale transformer is designed to couple the CNN-based local features with the transformer-based global cues in each scale, where the row-wise and column-wise global dependencies can be extracted by a lightweight Dual-Axis MSA. (2) A simple and effective spatial-aware inter-scale transformer is designed to interact among consensual regions in multiple scales, which can highlight the cross-scale dependency and resolve the complex scale variations. Experimental results on different benchmarks demonstrate that our Scale-Former outperforms the current state-of-the-art methods. The code is publicly available at: https://github.com/ZJUGiveLab/ScaleFormer.
Huimin Huang 0002, Shiao Xie, Lanfen Lin, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Ruofeng Tong 0001
IJCAI3
2022 An Accurate Unsupervised Liver Lesion Detection Method Using Pseudo-lesions
He Li 0042, Yutaro Iwamoto, Xianhua Han, Lanfen Lin, Hongjie Hu, Yen-Wei Chen 0001
MICCAI (8)4
2022 CubeMLP: An MLP-based Model for Multimodal Sentiment Analysis and Depression Estimation
abstract
Multimodal sentiment analysis and depression estimation are two important research topics that aim to predict human mental states using multimodal data. Previous research has focused on developing effective fusion strategies for exchanging and integrating mind-related information from different modalities. Some MLP-based techniques have recently achieved considerable success in a variety of computer vision tasks. Inspired by this, we explore multimodal approaches with a feature-mixing perspective in this study. To this end, we introduce CubeMLP, a multimodal feature processing framework based entirely on MLP. CubeMLP consists of three independent MLP units, each of which has two affine transformations. CubeMLP accepts all relevant modality features as input and mixes them across three axes. After extracting the characteristics using CubeMLP, the mixed multimodal features are flattened for task predictions. Our experiments are conducted on sentiment analysis datasets: CMU-MOSI and CMU-MOSEI, and depression estimation dataset: AVEC2019. The results show that CubeMLP can achieve state-of-the-art performance with a much lower computing cost.
Hao Sun 0013, Hongyi Wang 0002, Jiaqing Liu, Yen-Wei Chen 0001, Lanfen Lin
ACM Multimedia5
2022 Label-Efficient Domain Generalization via Collaborative Exploration and Generalization
abstract
Considerable progress has been made in domain generalization (DG) which aims to learn a generalizable model from multiple well-annotated source domains to unknown target domains. However, it can be prohibitively expensive to obtain sufficient annotation for source datasets in many real scenarios. To escape from the dilemma between domain generalization and annotation costs, in this paper, we introduce a novel task named label-efficient domain generalization (LEDG) to enable model generalization with label-limited source domains. To address this challenging task, we propose a novel framework called Collaborative Exploration and Generalization (CEG) which jointly optimizes active exploration and semi-supervised generalization. Specifically, in active exploration, to explore class and domain discriminability while avoiding information divergence and redundancy, we query the labels of the samples with the highest overall ranking of class uncertainty, domain representativeness, and information diversity. In semi-supervised generalization, we design MixUp-based intra- and inter-domain knowledge augmentation to expand domain knowledge and generalize domain invariance. We unify active exploration and semi-supervised generalization in a collaborative way and promote mutual enhancement between them, boosting model generalization with limited annotation. Extensive experiments show that CEG yields superior generalization performance. In particular, CEG can even use only 5% data annotation budget to achieve competitive results compared to the previous DG methods with fully labeled data on PACS dataset.
Junkun Yuan, Defang Chen 0001, Kun Kuang 0001, Fei Wu 0001, Lanfen Lin
ACM Multimedia6
2022 A multi-head pseudo nodes based spatial-temporal graph convolutional network for emotion perception from GAIT
Shurong Chai, Jiaqing Liu, Rahul Kumar Jain 0001, Tomoko Tateyama, Yutaro Iwamoto, Lanfen Lin, Yen-Wei Chen 0001
Neurocomputing6
2022 Attention-based cross-layer domain alignment for unsupervised domain adaptation
Junkun Yuan, Yen-Wei Chen 0001, Ruofeng Tong 0001, Lanfen Lin
Neurocomputing5
2022 Mutual Information-Based Graph Co-Attention Networks for Multimodal Prior-Guided Magnetic Resonance Imaging Segmentation
abstract
Multimodal magnetic resonance imaging (MRI) provides complementary information about targets, and the segmentation of multimodal MRI is widely used as an essential preprocessing step for initial diagnosis, stage differentiation, and post-treatment efficacy evaluation in clinical situations. For the main modality or each of the modalities, it is important to enhance the visual information by modeling the connection and effectively fusing the features among them. However, the existing methods for multimodal segmentation have a drawback; they coincidentally drop information of individual modality during the fusion process. Recently, graph learning-based methods have been applied in segmentation, and these methods have achieved considerable improvements by modeling the relationships across feature regions and reasoning using global information. In this paper, we propose a graph learning-based approach to efficiently extract modality-specific features and establish regional correspondence effectively among all modalities. In detail, after projecting features into a graph domain and employing graph convolution to propagate information across all regions for learning global modality-specific features, we propose a mutual information-based graph co-attention module to learn the weight coefficients of one bipartite graph constructed by the fully connected graphs having different modalities in the graph domain and by selectively fusing the node features. Based on the deformation diagram between the spatial-graph space and our proposed graph co-attention module, we present a multimodal prior-guided segmentation framework, which uses two strategies for two clinical situations:Modality-Specific Learning StrategyandCo-Modality Learning Strategy. Besides, the improvedCo-Modality Learning Strategyis used with trainable weights in the multi-task loss for the optimization of the proposed framework. We validated our proposed modules and frameworks on two multimodal MRI datasets: our private liver lesion dataset and a public prostate zone dataset. Our experimental results on both datasets prove the superiority of our proposed approaches.
Shaocong Mo, Lanfen Lin, Ruofeng Tong 0001, Qingqing Chen 0001, Fang Wang 0030, Hongjie Hu, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001
IEEE Trans. Circuits Syst. Video Technol.3
2022 MTL-ABS3Net: Atlas-Based Semi-Supervised Organ Segmentation Network With Multi-Task Learning for Medical Images
abstract
Organ segmentation is one of the most important step for various medical image analysis tasks. Recently, semi-supervised learning (SSL) has attracted much attentions by reducing labeling cost. However, most of the existing SSLs neglected the prior shape and position information specialized in the medical images, leading to unsatisfactory localization and non-smooth of objects. In this paper, we propose a novel atlas-based semi-supervised segmentation network with multi-task learning for medical organs, named MTL-ABS3Net, which incorporates the anatomical priors and makes full use of unlabeled data in a self-training and multi-task learning manner. The MTL-ABS3Net consists of two components: an Atlas-Based Semi-Supervised Segmentation Network (ABS3Net) and Reconstruction-Assisted Module (RAM). Specifically, the ABS3Net improves the existing SSLs by utilizing atlas prior, which generates credible pseudo labels in a self-training manner; while the RAM further assists the segmentation network by capturing the anatomical structures from the original images in a multi-task learning manner. Better reconstruction quality is achieved by using MS-SSIM loss function, which further improves the segmentation accuracy. Experimental results from the liver and spleen datasets demonstrated that the performance of our method was significantly improved compared to existing state-of-the-art methods.
Huimin Huang 0002, Qingqing Chen 0001, Lanfen Lin, Qiaowei Zhang, Yutaro Iwamoto, Xianhua Han, Akira Furukawa, Shuzo Kanasaki, Yen-Wei Chen 0001, Ruofeng Tong 0001, Hongjie Hu
IEEE J. Biomed. Health Informatics3
2022 DeepRecS: From RECIST Diameters to Precise Liver Tumor Segmentation
abstract
Liver tumor segmentation (LiTS) is of primary importance in diagnosis and treatment of hepatocellular carcinoma. Known automated LiTS methods could not yield satisfactory results for clinical use since they were hard to model flexible tumor shapes and locations. In clinical practice, radiologists usually estimate tumor shape and size by a Response Evaluation Criteria in Solid Tumor (RECIST) mark. Inspired by this, in this paper, we explore a deep learning (DL) based interactive LiTS method, which incorporates guidance from user-provided RECIST marks. Our method takes a three-step framework to predict liver tumor boundaries. Under this architecture, we develop a RECIST mark propagation network (RMP-Net) to estimate RECIST-like marks in off-RECIST slices. We also devise a context-guided boundary-sensitive network (CGBS-Net) to distill tumors' contextual and boundary information from corresponding RECIST(-like) marks, and then predict tumor maps. To further refine the segmentation results, we process the tumor maps using a 3D conditional random field (CRF) algorithm and a morphology hole-filling operation. Verified on two clinical contrast-enhanced abdomen computed tomography (CT) image datasets, our proposed approach can produce promising segmentation results, and outperforms the state-of-the-art interactive segmentation methods.
Yue Zhang 0042, Chengtao Peng, Liying Peng, Lanfen Lin, Ruofeng Tong 0001, Zhiyi Peng, Xiongwei Mao, Hongjie Hu, Yen-Wei Chen 0001, Jingsong Li 0001
IEEE J. Biomed. Health Informatics5
2022 Auto IV: Counterfactual Prediction via Automatic Instrumental Variable Decomposition
abstract
Instrumental variables (IVs), sources of treatment randomization that are conditionally independent of the outcome, play an important role in causal inference with unobserved confounders. However, the existing IV-based counterfactual prediction methods need well-predefined IVs, while it’s an art rather than science to find valid IVs in many real-world scenes. Moreover, the predefined hand-made IVs could be weak or erroneous by violating the conditions of valid IVs. These thorny facts hinder the application of the IV-based counterfactual prediction methods. In this article, we propose a novel Automatic Instrumental Variable decomposition (AutoIV) algorithm to automatically generate representations serving the role of IVs from observed variables (IV candidates). Specifically, we let the learned IV representations satisfy the relevance condition with the treatment and exclusion condition with the outcome via mutual information maximization and minimization constraints, respectively. We also learn confounder representations by encouraging them to be relevant to both the treatment and the outcome. The IV and confounder representations compete for the information with their constraints in an adversarial game, which allows us to get valid IV representations for IV-based counterfactual prediction. Extensive experiments demonstrate that our method generates valid IV representations for accurate IV-based counterfactual prediction.
Junkun Yuan, Anpeng Wu, Kun Kuang 0001, Bo Li 0064, Runze Wu 0001, Fei Wu 0001, Lanfen Lin
ACM Trans. Knowl. Discov. Data7
2021 Graph-Based Pyramid Global Context Reasoning With a Saliency- Aware Projection for Covid-19 Lung Infections Segmentation
abstract
Coronavirus Disease 2019 (COVID-19) has rapidly spread in 2020, emerging a mass of studies for lung infection segmentation from CT images. Though many methods have been proposed for this issue, it is a challenging task because of infections of various size appearing in different lobe zones. To tackle these issues, we propose a Graph-based Pyramid Global Context Reasoning (Graph-PGCR) module, which is capable of modeling long-range dependencies among disjoint infections as well as adapt size variation. We first incorporate graph convolution to exploit long-term contextual information from multiple lobe zones. Different from previous average pooling or maximum object probability, we propose a saliency-aware projection mechanism to pick up infection-related pixels as a set of graph nodes. After graph reasoning, the relation-aware features are reversed back to the original coordinate space for the down-stream tasks. We further construct multiple graphs with different sampling rates to handle the size variation problem. To this end, distinct multi-scale long-range contextual patterns can be captured. Our Graph- PGCR module is plug-and-play, which can be integrated into any architecture to improve its performance. Experiments demonstrated that the proposed method consistently boost the performance of state-of-the-art backbone architectures on both of public and our private COVID-19 datasets.
Huimin Huang 0002, Lanfen Lin, Xiongwei Mao, Xiaohan Qian, Zhiyi Peng, Jianying Zhou 0006, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Ruofeng Tong 0001
ICASSP3
2021 Graph-BAS3Net: Boundary-Aware Semi-Supervised Segmentation Network with Bilateral Graph Convolution
abstract
Semi-supervised learning (SSL) algorithms have attracted much attentions in medical image segmentation by leveraging unlabeled data, which challenge in acquiring massive pixel-wise annotated samples. However, most of the existing SSLs neglected the geometric shape constraint in object, leading to unsatisfactory boundary and non-smooth of object. In this paper, we propose a novel boundary-aware semi-supervised medical image segmentation network, named Graph-BAS3Net, which incorporates the boundary information and learns duality constraints between semantics and geometrics in the graph domain. Specifically, the proposed method consists of two components: a multi-task learning framework BAS3Net and a graph-based cross-task module BGCM. The BAS3Net improves the existing GAN-based SSL by adding a boundary detection task, which encodes richer features of object shape and surface. Moreover, the BGCM further explores the co-occurrence relations between the semantics segmentation and boundary detection task, so that the network learns stronger semantic and geometric correspondences from both labeled and unlabeled data. Experimental results on the LiTS dataset and COVID-19 dataset confirm that our proposed Graph-BAS3Net outperforms the state-of-the-art methods in semi-supervised segmentation task.
Huimin Huang 0002, Lanfen Lin, Yue Zhang 0042, Xiongwei Mao, Xiaohan Qian, Zhiyi Peng, Jianying Zhou 0006, Yen-Wei Chen 0001, Ruofeng Tong 0001
ICCV2
2021 3D Graph-S2Net: Shape-Aware Self-ensembling Network for Semi-supervised Segmentation with Bilateral Graph Convolution
Huimin Huang 0002, Lanfen Lin, Hongjie Hu, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Ruofeng Tong 0001
MICCAI (2)3
2021 Patch-Free 3D Medical Image Segmentation Driven by Super-Resolution Technique and Self-Supervised Guidance
Hongyi Wang 0002, Lanfen Lin, Hongjie Hu, Qingqing Chen 0001, Yinhao Li 0002, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Ruofeng Tong 0001
MICCAI (1)2
2021 Multi-phase Liver Tumor Segmentation with Spatial Aggregation and Uncertain Region Inpainting
Yue Zhang 0042, Chengtao Peng, Liying Peng, Huimin Huang 0002, Ruofeng Tong 0001, Lanfen Lin, Jingsong Li 0001, Yen-Wei Chen 0001, Qingqing Chen 0001, Hongjie Hu, Zhiyi Peng
MICCAI (1)6
2021 A Tensor Sparse Representation-Based CBMIR System for Computer-Aided Diagnosis of Focal Liver Lesions and its Pilot Trial
abstract
Clinicians refer to diagnosed medical cases in order to make correct diagnosis and take appropriate treatments, due to the complexity of focal liver lesions. It's a heavy burden, however, for medical doctors to find out similar and meaningful cases from the accumulated extreme large medical datasets. Content based medical image retrieval (CBMIR) that searches for similar images in a large database has been attracting increasing research interest recently. A CBMIR system provides doctors the diagnosed cases to improve the diagnosis accuracy and confidence. This paper proposed a tensor sparse representation method to extract temporal and spatial features of multi-phase CT images, so as to provide doctors medical cases more relevant to the query one. The proposed tensor sparse representation method is applied to the retrieval of focal liver lesions (FLLs). Experiments show that the proposed method achieved better retrieval performance than conventional methods. Pilot trial was conducted and results show that diagnosis accuracy and confidence was improved significantly by the developed CBMIR system based on the proposed method.
Jian Wang 0004, Xianhua Han, Lanfen Lin, Hongjie Hu, Yen-Wei Chen 0001
ICMR3
2021 M-DFNet: Multi-phase Discriminative Feature Network for Retrieval of Focal Liver Lesions
abstract
Content based medical image retrieval (CBMIR) plays a great role in computer aided diagnosis for assisting radiologists to detect and characterize focal liver lesions (FLLs). Deep learning has gained exciting performance on CBMIR. While the features generated by deep learning models trained using softmax loss are always separable but not discriminative enough, which is insufficient for retrieval task. In this paper, we propose a multi-phase discriminative feature network (M-DFNet) with a DeepExtracter and a feature refine module (FRModule) to learn discriminative and separable features under a joint supervision of center loss and softmax loss. The hybrid loss enables to minimize intra-class variations and enlarge inter-class differences as much as possible. The FRModule is proposed to recalibrate the deep features based on the learned class centers to tackle the complex imaging manifestations of FLLs and further enhance both the feature discrimination and generalization. Multi-phase computed tomography (CT) images contain pivotal information for diagnosis of FLLs. Thus the M-DFNet is designed to cope with multi-phase information and we explore an appropriate and effective method for multi-phase feature integration on limited data. Experimental results clearly demonstrate strong performance superiority by our proposed method.
Jing Liu 0041, Lanfen Lin, Hongjie Hu, Ruofeng Tong 0001, Jingsong Li 0001, Yen-Wei Chen 0001
ICMR3
2021 Accurate and fast mitotic detection using an anchor-free method based on full-scale connection with recurrent deep layer aggregation in 4D microscopy images
abstract
BACKGROUND: To effectively detect and investigate various cell-related diseases, it is essential to understand cell behaviour. The ability to detection mitotic cells is a fundamental step in diagnosing cell-related diseases. Convolutional neural networks (CNNs) have been successfully applied to object detection tasks, however, when applied to mitotic cell detection, most existing methods generate high false-positive rates due to the complex characteristics that differentiate normal cells from mitotic cells. Cell size and orientation variations in each stage make detecting mitotic cells difficult in 2D approaches. Therefore, effective extraction of the spatial and temporal features from mitotic data is an important and challenging task. The computational time required for detection is another major concern for mitotic detection in 4D microscopic images. RESULTS: In this paper, we propose a backbone feature extraction network named full scale connected recurrent deep layer aggregation (RDLA++) for anchor-free mitotic detection. We utilize a 2.5D method that includes 3D spatial information extracted from several 2D images from neighbouring slices that form a multi-stream input. CONCLUSIONS: Our proposed technique addresses the scale variation problem and can efficiently extract spatial and temporal features from 4D microscopic images, resulting in improved detection accuracy and reduced computation time compared with those of other state-of-the-art methods.
Titinunt Kitrungrotsakul, Yutaro Iwamoto, Satoko Takemoto, Hideo Yokota, Sari Ipponjima, Tomomi Nemoto, Lanfen Lin, Ruofeng Tong 0001, Jingsong Li 0001, Yen-Wei Chen 0001
BMC Bioinform.7
2021 VolumeNet: A Lightweight Parallel Network for Super-Resolution of MR and CT Volumetric Data
abstract
Deep learning-based super-resolution (SR) techniques have generally achieved excellent performance in the computer vision field. Recently, it has been proven that three-dimensional (3D) SR for medical volumetric data delivers better visual results than conventional two-dimensional (2D) processing. However, deepening and widening 3D networks increases training difficulty significantly due to the large number of parameters and small number of training samples. Thus, we propose a 3D convolutional neural network (CNN) for SR of magnetic resonance (MR) and computer tomography (CT) volumetric data called ParallelNet using parallel connections. We construct a parallel connection structure based on the group convolution and feature aggregation to build a 3D CNN that is as wide as possible with a few parameters. As a result, the model thoroughly learns more feature maps with larger receptive fields. In addition, to further improve accuracy, we present an efficient version of ParallelNet (called VolumeNet), which reduces the number of parameters and deepens ParallelNet using a proposed lightweight building block module called the Queue module. Unlike most lightweight CNNs based on depthwise convolutions, the Queue module is primarily constructed using separable 2D cross-channel convolutions. As a result, the number of network parameters and computational complexity can be reduced significantly while maintaining accuracy due to full channel fusion. Experimental results demonstrate that the proposed VolumeNet significantly reduces the number of model parameters and achieves high precision results compared to state-of-the-art methods in tasks of brain MR image SR, abdomen CT image SR, and reconstruction of super-resolution 7T-like images from their 3T counterparts.
Yinhao Li 0002, Yutaro Iwamoto, Lanfen Lin, Rui Xu 0002, Ruofeng Tong 0001, Yen-Wei Chen 0001
IEEE Trans. Image Process.3
2021 Attention-RefNet: Interactive Attention Refinement Network for Infected Area Segmentation of COVID-19
abstract
COVID-19 pneumonia is a disease that causes an existential health crisis in many people by directly affecting and damaging lung cells. The segmentation of infected areas from computed tomography (CT) images can be used to assist and provide useful information for COVID-19 diagnosis. Although several deep learning-based segmentation methods have been proposed for COVID-19 segmentation and have achieved state-of-the-art results, the segmentation accuracy is still not high enough (approximately 85%) due to the variations of COVID-19 infected areas (such as shape and size variations) and the similarities between COVID-19 and non-COVID-infected areas. To improve the segmentation accuracy of COVID-19 infected areas, we propose an interactive attention refinement network (Attention RefNet). The interactive attention refinement network can be connected with any segmentation network and trained with the segmentation network in an end-to-end fashion. We propose a skip connection attention module to improve the important features in both segmentation and refinement networks and a seed point module to enhance the important seeds (positions) for interactive refinement. The effectiveness of the proposed method was demonstrated on public datasets (COVID-19CTSeg and MICCAI) and our private multicenter dataset. The segmentation accuracy was improved to more than 90%. We also confirmed the generalizability of the proposed network on our multicenter dataset. The proposed method can still achieve high segmentation accuracy.
Titinunt Kitrungrotsakul, Qingqing Chen 0001, Huitao Wu, Yutaro Iwamoto, Hongjie Hu, Wenchao Zhu, Fangyi Xu, Lanfen Lin, Ruofeng Tong 0001, Jingsong Li 0001, Yen-Wei Chen 0001
IEEE J. Biomed. Health Informatics10
2021 Medical Image Segmentation With Deep Atlas Prior
abstract
Organ segmentation from medical images is one of the most important pre-processing steps in computer-aided diagnosis, but it is a challenging task because of limited annotated data, low-contrast and non-homogenous textures. Compared with natural images, organs in the medical images have obvious anatomical prior knowledge (e.g., organ shape and position), which can be used to improve the segmentation accuracy. In this paper, we propose a novel segmentation framework which integrates the medical image anatomical prior through loss into the deep learning models. The proposed prior loss function is based on probabilistic atlas, which is called as deep atlas prior (DAP). It includes prior location and shape information of organs, which are important prior information for accurate organ segmentation. Further, we combine the proposed deep atlas prior loss with the conventional likelihood losses such as Dice loss and focal loss into an adaptive Bayesian loss in a Bayesian framework, which consists of a prior and a likelihood. The adaptive Bayesian loss dynamically adjusts the ratio of the DAP loss and the likelihood loss in the training epoch for better learning. The proposed loss function is universal and can be combined with a wide variety of existing deep segmentation models to further enhance their performance. We verify the significance of our proposed framework with some state-of-the-art models, including fully-supervised and semi-supervised segmentation models on a public dataset (ISBI LiTS 2017 Challenge) for liver segmentation and a private dataset for spleen segmentation.
Huimin Huang 0002, Lanfen Lin, Hongjie Hu, Qiaowei Zhang, Qingqing Chen 0001, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Ruofeng Tong 0001
IEEE Trans. Medical Imaging3
2020 Black-Box Adversarial Attacks Against Deep Learning Based Malware Binaries Detection with GAN
abstract
For efficient malware detection, there are more and more deep learning methods based on raw software binaries. Recent studies show that deep learning models can easily be fooled to make a wrong decision by introducing subtle perturbations to inputs, which attracts a large influx of work in adversarial attacks. However, most of the existing attack methods are based on manual features (e.g., API calls) or in the white-box setting, making the attacks impractical in current real-world scenarios. In this work, we propose a novel attack framework called GAPGAN, which generates adversarial payloads (padding bytes) with generative adversarial networks (GANs). To the best of our knowledge, it is the first work that performs end-to-end black-box attacks at the byte-level against deep learning based malware binaries detection. In our attack framework, we map input discrete malware binaries to continuous space, then feed it to the generator of GAPGAN to generate adversarial payloads. We append payloads to the original binaries to craft an adversarial sample while preserving its functionality. We propose to use a dynamic threshold for reducing the loss of the effectiveness of the payloads when mapping it from continuous format back to the original discrete format. For balancing the attention of the generator to the payloads and the adversarial samples, we use an automatic weight tuning strategy. We train GAPGAN with both malicious and benign software. Once the training is finished, the generator can generate an adversarial sample with only the input malware in less than twenty milliseconds. We apply GAPGAN to attack the state-of-the-art detector MalConv and achieve 100% attack success rate with only appending payloads of 2.5% of the total length of the data for detection. We also attack deep learning models with different structures under different defense methods. The experiments show that GAPGAN outperforms other state-of-the-art attack models in efficiency and effectiveness.
Junkun Yuan, Shaofang Zhou, Lanfen Lin, Jia Cui
ECAI3
2020 UNet 3+: A Full-Scale Connected UNet for Medical Image Segmentation
abstract
Recently, a growing interest has been seen in deep learning-based semantic segmentation. UNet, which is one of deep learning networks with an encoder-decoder architecture, is widely used in medical image segmentation. Combining multi-scale features is one of important factors for accurate segmentation. UNet++ was developed as a modified Unet by designing an architecture with nested and dense skip connections. However, it does not explore sufficient information from full scales and there is still a large room for improvement. In this paper, we propose a novel UNet 3+, which takes advantage of full-scale skip connections and deep supervisions. The full-scale skip connections incorporate low-level details with high-level semantics from feature maps in different scales; while the deep supervision learns hierarchical representations from the full-scale aggregated feature maps. The proposed method is especially benefiting for organs that appear at varying scales. In addition to accuracy improvements, the proposed UNet 3+ can reduce the network parameters to improve the computation efficiency. We further propose a hybrid loss function and devise a classification-guided module to enhance the organ boundary and reduce the over-segmentation in a non-organ image, yielding more accurate segmentation results. The effectiveness of the proposed method is demonstrated on two datasets. The code is available at: github.com/ZJUGiveLab/UNet-Version.
Huimin Huang 0002, Lanfen Lin, Ruofeng Tong 0001, Hongjie Hu, Qiaowei Zhang, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Jian Wu 0001
ICASSP2
2020 Visual Relationship Detection With A Deep Convolutional Relationship Network
abstract
Visual relationship is crucial to image understanding and can be applied to many tasks (e.g., image caption and visual question answering). Despite great progress on many vision tasks, relationship detection remains a challenging problem due to the complexity of modeling the widely spread and imbalanced distribution of {subject - predicate - object} triplets. In this paper, we propose a new framework to capture the relative positions and sizes of the subject and object in the feature map and add a new branch to filter out some object pairs that are unlikely to have relationships. In addition, an activation function is trained to increase the probability of some feature maps given an object pair. Experiments on two large datasets, the Visual Relationship Detection (VRD) and Visual Genome (VG) datasets, demonstrate the superiority of our new approach over state-of-the-art methods. Further, ablation study verifies the effectiveness of our techniques.
Yaopeng Peng, Danny Ziyi Chen, Lanfen Lin
ICIP3
2020 Multimodal Priors Guided Segmentation of Liver Lesions in MRI Using Mutual Information Based Graph Co-Attention Networks
Shaocong Mo, Lanfen Lin, Ruofeng Tong 0001, Qingqing Chen 0001, Fang Wang 0030, Hongjie Hu, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001
MICCAI (4)3
2020 Tensor-based sparse representations of multi-phase medical images for classification of focal liver lesions
Jian Wang 0004, Jing Li 0046, Xianhua Han, Lanfen Lin, Hongjie Hu, Qingqing Chen 0001, Yutaro Iwamoto, Yen-Wei Chen 0001
Pattern Recognit. Lett.4
2020 Semi-Supervised Learning for Semantic Segmentation of Emphysema With Partial Annotations
abstract
Segmentation and quantification of each subtype of emphysema is helpful to monitor chronic obstructive pulmonary disease. Due to the nature of emphysema (diffuse pulmonary disease), it is very difficult for experts to allocate semantic labels to every pixel in the CT images. In practice, partially annotating is a better choice for the radiologists to reduce their workloads. In this paper, we propose a new end-to-end trainable semi-supervised framework for semantic segmentation of emphysema with partial annotations, in which a segmentation network is trained from both annotated and unannotated areas. In addition, we present a new loss function, referred to as Fisher loss, to enhance the discriminative power of the model and successfully integrate it into our proposed framework. Our experimental results show that the proposed methods have superior performance over the baseline supervised approach (trained with only annotated areas) and outperform the state-of-the-art methods for emphysema segmentation.
Liying Peng, Lanfen Lin, Hongjie Hu, Yue Zhang 0042, Huali Li, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001
IEEE J. Biomed. Health Informatics2
2019 A Dual-Attention Dilated Residual Network for Liver Lesion Classification and Localization on CT Images
abstract
Automatic liver lesion classification on computed tomography images is of great importance to early cancer diagnosis and remains a challenging task. State-of-the-art liver lesion classification algorithms are currently based on manually selected regions of interest (ROIs) or automatically detected ROIs. However, liver lesions usually vary in size and shape, which makes the ROI selection process labor-intensive and also poses an obstacle to automatic lesion detection. In this paper, we propose a dual-attention dilated residual network (DADRN) as a potential solution to lesion classification task without manual ROI selection or automatic lesion detection. We incorporated a novel dual-attention module in order to capture the non-local feature dependencies and help the deep neural network focus on the lesion area by enlarging the difference between the lesion area and nonlesion area. To the best of our knowledge, we are the first to employ the self-attention mechanism to address liver lesion classification task. In addition, the well-trained DADRN can be used for weakly-supervised lesion localization without any architectural change or retraining. Experiment results show that DADRN could achieve a lesion classification accuracy comparable to that of the state-of-the-art ROI-based method and outperformed state-of-the-art attention-based approaches in both liver lesion classification and localization tasks.
Xiao Chen 0016, Jian Wu 0001, Lanfen Lin, Hongjie Hu, Qiaowei Zhang, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Ruofeng Tong 0001
ICIP3
2019 Multi-Stream Scale-Insensitive Convolutional and Recurrent Neural Networks for Liver Tumor Detection in Dynamic Ct Images
abstract
Convolutional neural networks (CNNs) have achieved great success in numerous challenging vision tasks, and have great potential for object detection in natural images. Compared with the natural images, medical images exhibit some unique characteristics. Therefore, substantial challenges still remain in this field. The first challenge is to develop a method for effectively distilling enhancement patterns from the dynamic CT images. Moreover, since tumor sizes vary greatly and small lesions are important for early liver tumor detection, lesion detection with a widely variable scale is another challenge. In this paper, we propose a multi-stream scale-insensitive convolutional and recurrent neural network (MSCR) for liver tumor detection. Specifically, we propose the use of grouped convolutional long short-term memory (GCLSTM) to extract enhancement patterns, which is developed as a plug-and-play module. Experiments show that the MSCR framework exhibits superior performance over state-of-the-art approaches, achieving an average precision of 77.06% for detection of focal liver lesions. We have released the code of MSCR in1.
Ruofeng Tong 0001, Jian Wu 0001, Lanfen Lin, Xiao Chen 0016, Hongjie Hu, Qiaowei Zhang, Qingqing Chen 0001, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001
ICIP4
2019 CNN-based DGA Detection with High Coverage
abstract
Attackers often use domain generation algorithms (DGAs) to create various kinds of pseudorandom domains dynamically and select a part of them to connect with command and control servers, therefore it is important to automatically detect the algorithmically generated domains (AGDs). AGDs can be broken down into two categories: character-based domains and wordlist-based domains. Recently, methods based on machine learning and deep learning have been widely explored. However, much of the previous work perform well in detecting one kind of DGA families but poorly in classifying another kind. A general detection system which is applicable to both kinds of domains still remains a challenge. To address this problem, we propose a novel real-time detection method with high accuracy as well as high coverage. We first convey a domain name into a sequence of word-level or character-level components, then design a deep neural network based on temporal convolutional network to extract the implicit pattern and classify the domain into two or more categories. Our experimental results demonstrate that our model outperforms state-of-the-art approaches in both binary classification and multi-class classification, and shows a good performance in detecting different kinds of DGAs. Besides, the high training efficiency of our model makes it adjust to new malicious domains quickly.
Shaofang Zhou, Lanfen Lin, Junkun Yuan, Zhaoting Ling, Jia Cui
ISI2
2019 Semi-supervised Segmentation of Liver Using Adversarial Learning with Deep Atlas Prior
Lanfen Lin, Hongjie Hu, Qiaowei Zhang, Qingqing Chen 0001, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Ruofeng Tong 0001, Jian Wu 0001
MICCAI (6)2
2019 Classification and Quantification of Emphysema Using a Multi-Scale Residual Network
abstract
Automated tissue classification is an essential step for quantitative analysis and treatment of emphysema. Although many studies have been conducted in this area, there still remain two major challenges. First, different emphysematous tissue appears in different scales, which we call "inter-class variations." Second, the intensities of CT images acquired from different patients, scanners or scanning protocols may vary, which we call "intra-class variations". In this paper, we present a novel multi-scale residual network with two channels of raw CT image and its differential excitation component. We incorporate multi-scale information into our networks to address the challenge of inter-class variations. In addition to the conventional raw CT image, we use its differential excitation component as a pair of inputs to handle intra-class variations. Experimental results show that our approach has superior performance over the state-of-the- art methods, achieving a classification accuracy of 93.74% on our original emphysema database. Based on the classification results, we also perform the quantitative analysis of emphysema in 50 subjects by correlating the quantitative results (the area percentage of each class) with pulmonary functions. We show that centrilobular emphysema (CLE) and panlobular emphysema (PLE) have strong correlation with the pulmonary functions and the sum of CLE and PLE can be used as a new and accurate measure of emphysema severity instead of the conventional measure (sum of all subtypes of emphysema). The correlations between the new measure and various pulmonary functions are up to |r| = 0.922 (r is correlation coefficient).
Liying Peng, Yen-Wei Chen 0001, Lanfen Lin, Hongjie Hu, Huali Li, Qingqing Chen 0001, Xiaoli Ling, Xianhua Han, Yutaro Iwamoto
IEEE J. Biomed. Health Informatics3
2018 Classification of Pulmonary Emphysema in CT Images Based on Multi-Scale Deep Convolutional Neural Networks
abstract
In this work, we aim at classifying emphysema in computed tomography (CT) images of lungs. Most previous works are limited to extracting low-level features or mid-level features without enough high-level information. Moreover, these approaches do not take the characteristics (scales) of different emphysema into account, which are crucial for feature extraction. In contrast to previous works, we propose a novel deep learning method based on multiscale deep convolutional neural networks. There are three contributions for this paper. First, we propose to use a base residual network with 20 layers to extract more high-level information. To the best of our knowledge, this is the first deep learning method for classification of emphysema. Second, we incorporate multi-scale information into our deep neural networks so as to take full consideration of the characteristics of different emphysema. Finally, we established a high-quality emphysema dataset which contains 91 high-resolution computed tomography (HRCT) volumes, annotated manually by two experienced radiologists and checked by one experienced chest radiologist. A 92.68% classification accuracy is achieved on this dataset. The results show that (1) the multi-scale method is highly effective in comparison to the single scale setting; (2) the proposed approach is superior to the state-of-the-art techniques.
Liying Peng, Lanfen Lin, Hongjie Hu, Huali Li, Xiaoli Ling, Xianhua Han, Yutaro Iwamoto, Yen-Wei Chen 0001
ICIP2
2018 Combining Convolutional and Recurrent Neural Networks for Classification of Focal Liver Lesions in Multi-phase CT Images
Lanfen Lin, Hongjie Hu, Qiaowei Zhang, Qingqing Chen 0001, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001
MICCAI (2)2
2018 Residual Convolutional Neural Networks with Global and Local Pathways for Classification of Focal Liver Lesions
Lanfen Lin, Hongjie Hu, Qiaowei Zhang, Qingqing Chen 0001, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001
PRICAI (1)2
2017 Joint weber-based rotation invariant uniform local ternary pattern for classification of pulmonary emphysema in CT images
abstract
In this paper, we present a novel image representation approach for classifying emphysema in computed tomography (CT) images of the lung. Our proposed method extends rotation invariant uniform local binary pattern (RIULBP) and local ternary pattern (LTP), which are extensively used in a variety of computer vision applications, into rotation invariant uniform local ternary pattern (RIULTP) with a human perception principle: Weber's law. In addition, by integrating the upper pattern and the lower pattern of the Weber-based RIULTP (WRIULTP), we further put forward the joint Weber-based rotation invariant uniform local ternary pattern (JWRIULTP), which allows for a much richer representation and also takes the comprehensive information of the image into account. The proposed methods are tested on the Outex database (texture database) and the Bruijne and Srensen database (emphysema database). The results show the superiority of the proposed approaches to the state-of-the-art techniques for emphysema classification including rotation invariant local binary pattern (RILBP) and texton-based approach.
Liying Peng, Lanfen Lin, Hongjie Hu, Xiaoli Ling, Xianhua Han, Yen-Wei Chen 0001
ICIP2
2017 Tensor Sparse Representation of Temporal Features for Content-Based Retrieval of Focal Liver Lesions Using Multi-phase Medical Images
abstract
Content Based Image Retrieval (CBIR) systems that search similar images in a large database are attracting more and more research interests recently, and have been applied to medical image characterization for expert's experience sharing. One challenging task in CBIR is how to extract features for effective image representation. Therein sparse coding technique has been proven to be an effective way to learn inherent structure features for image analysis. However, it is necessary to first vectorize the 2- or 3-dimensional spatial structure for analysis with sparse coding, and then destroy the spatial relation of nearby voxels. In this study, we propose a multilinear sparse coding method to learn features from multi-dimensional medical images. We regard high dimensional local structures as tensors and propose a K-CP (CANDECOMP/PARAFAC) algorithm to learn a tensor dictionary in an iterative way. With the learned tensor dictionary, sparse coefficients of tensor local structures are calculated by multilinear orthogonal matching pursuit (MOMP) algorithm, which is an extended multilinear version of the conventional linear OMP. The proposed multilinear sparse coding method is prospected to be more efficient and effective for inherent feature extraction compared with conventional linear methods. The proposed method is applied to a CBIR system for retrieval of focal liver lesions (FLLs) using a medical database consisting of contrast-enhanced multi-phase computer-tomography (CT) images. Experiments show that the constructed CBIR with multilinear sparse coding method can achieve promising retrieval performance.
Jian Wang 0004, Xianhua Han, Lanfen Lin, Hongjie Hu, Chongwu Jin, Yen-Wei Chen 0001
ISM4
2017 A novel confidence estimation method for heterogeneous implicit feedback
abstract
Implicit feedback, which indirectly reflects opinion through user behaviors, has gained increasing attention in recommender system communities due to its accessibility and richness in real-world applications. A major way of exploiting implicit feedback is to treat the data as an indication of positive and negative preferences associated with vastly varying confidence levels. Such algorithms assume that the numerical value of implicit feedback, such as time of watching, indicates confidence, rather than degree of preference, and a larger value indicates a higher confidence, although this works only when just one type of implicit feedback is available. However, in real-world applications, there are usually various types of implicit feedback, which can be referred to as heterogeneous implicit feedback. Existing methods cannot efficiently infer confidence levels from heterogeneous implicit feedback. In this paper, we propose a novel confidence estimation approach to infer the confidence level of user preference based on heterogeneous implicit feedback. Then we apply the inferred confidence to both point-wise and pair-wise matrix factorization models, and propose a more generic strategy to select effective training samples for pair-wise methods. Experiments on real-world e-commerce datasets from Tmall.com show that our methods outperform the state-of-the-art approaches, considering several commonly used ranking-oriented evaluation criteria.
Jing Wang 0025, Lanfen Lin, Jiaqi Tu, Penghua Yu
Frontiers Inf. Technol. Electron. Eng.2
2016 Confidence-Learning Based Collaborative Filtering with Heterogeneous Implicit Feedbacks
Jing Wang 0025, Lanfen Lin, Jiaqi Tu
APWeb (1)2
2016 Bag of temporal co-occurrence words for retrieval of focal liver lesions using 3D multiphase contrast-enhanced CT images
abstract
Computer-aided diagnosis (CAD) systems have been verified to have the potential to assist radiologists in clinical diagnosis to detect and characterize focal liver lesions (FLLs) based on single- or multiphase contrast-enhanced computed tomography (CT) images. Features extracted from multiphase contrast-enhanced CT images carry more important diagnostic information i.e. enhancement pattern and demonstrate much stronger discriminative ability compared to those of single-phase CT images. In this paper, we propose a new method for multiphase image feature generation called the bag of temporal co-occurrence words (BoTCoW). A temporal co-occurrence image connecting intensity from multiphase images is constructed. Then the bag of visual word (BoVW) model is employed on the temporal co-occurrence images to extract temporal features. The proposed method effectively captures temporal enhancement information and demonstrates the distribution of the evolution patterns. The effectiveness of this method is validated in a retrieval system using 132 FLLs with confirmed pathology type. The preliminary results show that the proposed BoTCoW method outperforms the previously proposed temporal features and multiphase features based on the BoVW model.
Lanfen Lin, Hongjie Hu, Yitao Liu, Jian Wang 0004, Xianhua Han, Yen-Wei Chen 0001
ICPR2
2016 A Novel Framework to Process the Quantity and Quality of User Behavior Data in Recommender Systems
Penghua Yu, Lanfen Lin, Yuangang Yao
WAIM (1)2
2014 New word identification in social network text based on time series information
abstract
Different from the languages widely used in western countries such as English or French, there are no spaces between words in Chinese language, and a segmentation of the texts is necessary before other superior processes. New word identification is an important problem in the segmentation process, especially when the segmentation targets are social network texts which have more abbreviated words or other non-standard representations. Several methods have been proposed to detect Chinese new words. Most of these methods take the corpus as a static set and they don't consider the time domain information. Different from these studies, we regard our social network corpus as a text series spreading along the time line and design a new kind of features named dynamic features which can reflect the temporal variety of the string's statistical features. The experimental results on the dataset crawled from the biggest microblogging application in China show that this method can significantly improve the effect of Chinese new word identification.
Meng Wang 0016, Lanfen Lin
CSCWD2
2014 Improving Recommendations with Collaborative Factors
Penghua Yu, Lanfen Lin, Jing Wang 0025, Meng Wang 0016
WAIM2
2014 A new sketch-based 3D model retrieval approach by using global and local features
Lanfen Lin, Min Tang 0001
Graph. Model.2
2013 Extracting Novel Features for E-Commerce Page Quality Classification
Jing Wang 0025, Lanfen Lin, Penghua Yu, Jiaolong Liu
ADMA (1)2
2012 Research on enterprise interoperability for cluster supply chain
abstract
Cluster supply chain is a kind of supply chain coupling with the industrial cluster, and enterprises within this region, cooperating with others frequently, are in urgent need to promote the efficiency of business cooperation by means of system interoperation, which is now seriously impeded by the high heterogeneity of the data model and system among the enterprises. Under such circumstances, we study the requirements of enterprise interoperability for the cluster supply chain by analyzing the characteristics of cluster supply chain and the complex business relationships in it. A framework of enterprise interoperability for cluster supply chain is presented to give global description of the needs. After that, based on the framework, a three-step solution of enterprise interoperability is proposed, through realizing interoperability of each layer to support the business collaboration among the enterprises. In our solution, semantic mapping based on mixed ontologies is utilized to resolve the syntax and semantic differences, and Web service is employed to ensure interoperability and manageability. Due to the high flexibility and sematic enhancement of the service, the dynamic inter-enterprise business process is satisfied by the combination of services. Finally, a SOA-based platform is developed in accord with the proposed solution and the experimental results prove its practicability and effectiveness.
Zhan Jiang, Lanfen Lin, Guorong Wang
CSCWD2
2012 Research on semantic-based knowledge service for cluster supply chain
abstract
Knowledge service is critical to enterprise collaboration in cluster supply chain. However, current cluster enterprises have limited capabilities for knowledge sharing. In order to address the knowledge service needs of cluster supply chain, this paper proposes a novel semantic-based knowledge service approach. In the approach, a multi-level knowledge model is presented to achieve uniform and flexible representation for cluster public knowledge and enterprise private knowledge, which further forms complex knowledge network for cluster supply chain. Then the knowledge collaboration mechanism is utilized to support knowledge collaborative creating and enriching, and cooperative sharing. In addition, the knowledge access control scheme is employed to ensure high security. To demonstrate the practicality of the proposed approach, a knowledge service platform is implemented, including several core services such as knowledge collaborative editing, retrieval and visual navigation, etc. Applications show that this platform can support knowledge sharing between enterprises effectively and further promote the development of cluster supply chain efficiently.
Meng Wang 0016, Lanfen Lin
CSCWD4
2007 Research on Ontology-based Integrated Product modeling in Multidisciplinary Collaborative Environment
abstract
To effectively integrate product information spanning various stages throughout the whole product life cycle, and facilitate seamless sharing and interoperability at knowledge level in heterogeneous collaborative environment, a framework of integrated product modeling based on multiple ontologies is presented. This framework is composed of master model and domain models with the development of a shared basic ontology and domain-specific ontologies. Both the ontologies enable product knowledge acquisition and management. The layered master model contains semantic-enriched core and shared information and supports product and process integration. By extracting from master model and knowledge-based reusing from others in different contexts, the domain models are established with well-defined meaning, which permits each discipline to work on its own perspective while ensuring co-ordination and consistency at knowledge level in multidisciplinary collaborative product development.
Yi-Chao Lou, Lanfen Lin, Jinxiang Dong
CSCWD2
2005 Study of ASP service lifecycle management technologies for networked manufacturing system
abstract
Nowadays, networked manufacturing system based on ASP (application service provider) has become one of the hotspots of research and application. How to manage large amount of ASP services, and how to provide better service quality, has become very important. The paper mainly studies the ASP service management technologies. We present the concept of service lifecycle management (SLM), defines ASP service, and set up a state model for service lifecycle. Then, we give several key technologies' solutions for implementation of service management. Finally the application of these technologies in a real networked manufacturing system is introduced.
Lanfen Lin, Jinxiang Dong
CSCWD (2)3
2005 Integrated product modeling based on Web services in distributed environment
abstract
To flexibly integrate various systems at different stages and improve the openness and reusability of information in remote collaborative product development, a framework of integrated product modeling based on Web services is presented. The integrated model consists of a master model and application models that combine static and dynamic information supporting whole product lifecycle. The master model contains the core and shared information. The application model relating to specific manufacture domain is built up by extracting the master model and complementing domain information. The flexible integration of heterogeneous systems is ensured by the use of Web services. The openness and expansibility are preserved by combining STEP and XML for information representation and exchange. The high reusability of product information is achieved by using a service-oriented multi-granularity schema. The information consistency among the master model and application models is maintained through a coordinated-protocol for rapidly acquiring alteration and reacting to it.
Lanfen Lin, Yi-Chao Lou, Jinxiang Dong
CSCWD (1)1
2005 Study of networked manufacturing oriented cooperative CAPP system
abstract
Networked manufacturing oriented cooperative CAPP system has become one of the hotspots of research and application. The paper mainly studies how to realize cooperative planning among heterogeneous CAPP systems in a networked manufacturing oriented environment. We put forward an idea for cooperative process planning. Then, we study some key technologies of it, including cooperative control command management technology and shared data processing. Finally an example of the networked manufacturing oriented cooperative CAPP system is introduced.
Zhaomin Xu, Lanfen Lin, Jinxiang Dong
CSCWD (2)3
2002 Research on Technology of Product Data Exchange for Internet-Based Distributed Integrated Manufacturing System
abstract
One of the key technologies for distributed integrated manufacturing systems is to realize automatic transmission and conversion of product data among distributed and heterogeneous systems like CAD, CAPP, CAM, CAE, etc. This paper presents an Internet-based framework of product data exchange and the architecture of the framework, which uses Internet/CORBA as the communication platform and XML as product data exchange language. This framework within which, conversion of other data formats to XML format is implemented by software agents, supports exchange of structured and non-structured product data.
Lanfen Lin, Jinxiang Dong
CSCWD2