VLDB 2026 Research / reviewers in the wild / expert
Yen-Wei Chen 0001
dblp:55/1008-1
· DBLP profile ↗
218ranked-venue papers
23as first author
81since 2021 · last 2026
0000-0002-5952-0188ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 120 · 6 first-author · 47 since 2021Artificial intelligence and machine learning · 108 · 19 first-author · 23 since 2021Applied, interdisciplinary, general and emerging computing · 43 · 1 first-author · 33 since 2021Computer networks · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 2 since 2021Security and privacy · 1Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MIRTH: Mutual-Information Reasoning with Temporal Hubs for Vision-Language-Action AgentsabstractVLA models have emerged as a powerful paradigm for transferring semantic knowledge from web-scale data to physical robotic control.However, current single-frame architectures suffer from intrinsic limitations: temporal myopia that discards historical dynamics, reasoning gaps between high-level instructions and low-level motor commands, and inference inefficiency due to autoregressive scalar decoding.In this work, we propose MIRTH, a unified framework designed to address these challenges.MIRTH augments a pretrained VLA backbone with three key innovations: (1) dualscale temporal memory hubs that compress long-term scene evolution and short-term motion trends into compact embeddings; (2) latent reasoning tokens optimized via a mutualinformation objective carving out a semantic plan space to align multimodal context with action trajectories; and (3) a parallel action decoding scheme that replaces autoregressive generation with vector-wise prediction to maximize control throughput.Extensive evaluations on the LIBERO simulation benchmark and a real-world LeRobot platform demonstrate that MIRTH achieves state-of-the-art performance and exhibiting emergent error recovery capabilities.The codes and collected datasets are released at http://github.com/kiva12138/mirth. Hao Sun 0013, Yu Song 0008, Shiyu Teng, Ziwei Niu, Yen-Wei Chen 0001 |
ACL (1) | 5 |
| 2026 | Enhancing Depression Detection Using Pretrained Multi-modal Sentiment Analysis Models with Deep Prefix TuningabstractDepression, a pervasive mental health condition, affects millions globally, challenging early and accurate diagnosis due to its subtle and varied manifestations. Recognizing the critical link between emotional dysregulation and depressive symptoms, our research introduces a pioneering training paradigm that integrates sentiment analysis with depression detection. This approach is motivated by the potential of sentiment data to enrich models with a deeper understanding of emotional states, crucial for identifying depressive patterns. To leverage the nuanced sentiment information without compromising the pretrained model’s integrity, we employ deep prefix tuning. This novel technique allows for targeted model refinement, ensuring that the valuable pretrained structures are not overshadowed by the sparse and specific nature of depression-related data. The empirical results demonstrate superior performance across standard benchmarks, setting a new precedent for multimodal depression detection. Shiyu Teng, Jiaqing Liu, Shurong Chai, Hao Sun 0013, Tomoko Tateyama, Lanfen Lin, Yen-Wei Chen 0001 |
ACM Trans. Comput. Heal. | 7 |
| 2026 | An improved multi-instance learning model with clinical-guided cross-attention for postoperative early recurrence prediction of hepatocellular carcinoma using histopathological images
Gan Zhan, Fang Wang 0030, Yinhao Li 0002, Rahul Kumar Jain 0001, Qingqing Chen 0001, Lanfen Lin, Hongjie Hu, C. Krishna Mohan, Yen-Wei Chen 0001 |
Neurocomputing | 11 |
| 2026 | One framework to rule them all: Unifying multimodal tasks with LLM neural-tuning
Hao Sun 0013, Yu Song 0008, Jiaqing Liu, Jihong Hu, Yen-Wei Chen 0001, Lanfen Lin |
Pattern Recognit. | 5 |
| 2026 | SPA: Leveraging the SAM With Spatial Priors Adapter for Enhanced Medical Image SegmentationabstractThe Segment Anything Model (SAM) has gained renown for its success in image segmentation, benefiting significantly from its pretraining on extensive datasets and its interactive prompt-based segmentation approach. Although highly effective in natural (real-world) image segmentation tasks, the SAM model encounters significant challenges in medical imaging due to the inherent differences between these two domains. To address these challenges, we propose the Spatial Prior Adapter (SPA) scheme, a parameter-efficient fine-tuning strategy that enhances SAM's adaptability to medical imaging tasks. SPA introduces two novel modules: the Spatial Prior Module (SPM), which captures localized spatial features through convolutional layers, and the Feature Communication Module (FCM), which integrates these features into SAM's image encoder via cross-attention mechanisms. Furthermore, we develop a Multiscale Feature Fusion Module (MSFFM) to enhance SAM's end-to-end segmentation capabilities by effectively aggregating multiscale contextual information. These lightweight modules require minimal computational resources while significantly boosting segmentation performance. Our approach demonstrates superior performance in both prompt-based and end-to-end segmentation scenarios through extensive experiments on publicly available medical imaging datasets. Performance highlights the potential of the proposed method to bridge the gap between foundation models and domain-specific medical imaging tasks. This advancement paves the way for more effective AI-assisted medical diagnostic systems. Jihong Hu, Yinhao Li 0002, Rahul Kumar Jain 0001, Lanfen Lin, Yen-Wei Chen 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2026 | Multimodal Graph Learning With Multi-Hypergraph Reasoning Networks for Focal Liver Lesion Classification in Multimodal Magnetic Resonance ImagingabstractMultimodal magnetic resonance imaging (MRI) is instrumental in differentiating liver lesions. The major challenge involves modeling reliable connections and simultaneously learning complementary information across various MRI sequences. While previous studies have primarily focused on multimodal integration in a pair-wise manner using few modalities, our research seeks to advance a more comprehensive understanding of interaction modeling by establishing complex high-order correlations among the diverse modalities in multimodal MRI. In this paper, we introduce a multimodal graph learning with multi-hypergraph reasoning network to capture the full spectrum of both pair-wise and group-wise relationships among different modalities. Specifically, a weight-shared encoder extracts features from regions of interest (ROI) images across all modalities. Subsequently, a collection of uniform hypergraphs are constructed with varying vertex configurations, allowing for the modeling of not only pair-wise correlations but also the high-order collaborations for relational reasoning. Following information propagation through the hypergraph message passing, adaptive intra-modality fusion module is proposed to effectively fuse feature representations from different hypergraphs of the same modality. Finally, all refined features are concatenated to prepare for the classification task. Our experimental evaluations, including focal liver lesions classification using the LLD-MMRI2023 dataset and early recurrence prediction of hepatocellular carcinoma using our internal datasets, demonstrate that our method significantly surpasses the performance of existing approaches, indicating the effectiveness of our model in handling both pair-wise and group-wise interactions across multiple modalities. Shaocong Mo, Lanfen Lin, Ruofeng Tong 0001, Fang Wang 0030, Qingqing Chen 0001, Wenbin Ji, Yinhao Li 0002, Hongjie Hu, Yen-Wei Chen 0001 |
IEEE J. Biomed. Health Informatics | 10 |
| 2026 | S2Match: Revisiting Weak-to-Strong Consistency From a Semantic Similarity Perspective for Semi-Supervised Medical Image SegmentationabstractSemi-supervised learning (SSL) for medical image segmentation is a challenging yet highly practical task, which reduces reliance on large-scale labeled datasets by leveraging unlabeled samples. Among SSL techniques, the weak-to-strong consistency framework, popularized by FixMatch, has emerged as a state-of-the-art method in classification tasks. Notably, such a simple pipeline has also shown competitive performance in medical image segmentation. However, two key limitations still persist, impeding its efficient adaptation: (1) the neglect of contextual dependencies results in inconsistent predictions for similar semantic features, leading to incomplete object segmentation; (2) the lack of exploitation on semantic similarity between labeled and unlabeled data induces considerable class-distribution discrepancy. To address these limitations, we propose a novel SSL framework for medical image segmentation, named S2Match, powered by two appealing designs from a semantic similarity perspective: (1) rectifying pixel-wise prediction by reasoning about the intra-image pair-wise affinity map, thus integrating contextual dependencies explicitly into the final prediction; (2) bridging labeled and unlabeled data via a feature querying mechanism for compact class representation learning, which fully considers cross-image anatomical similarities. As the reliable semantic similarity extraction depends on robust features, we further introduce an effective Spatial-aware Fusion Module (SFM) to explore distinctive information from multiple scales. Experiments show that S2Match yields consistent improvements over the state-of-the-art methods across five public medical image segmentation benchmarks, exhibiting competitive performance on both 2D and 3D tasks. Shiao Xie, Hongyi Wang 0002, Ziwei Niu, Hao Sun 0013, Shuyi Ouyang, Yen-Wei Chen 0001, Lanfen Lin |
IEEE J. Biomed. Health Informatics | 6 |
| 2026 | EICSeg: Universal Medical Image Segmentation via Explicit In-Context LearningabstractDeep learning models for medical image segmentation often struggle with task-specific characteristics, limiting their generalization to unseen tasks with new anatomies, labels, or modalities. Retraining or fine-tuning these models requires substantial human effort and computational resources. To address this, in-context learning (ICL) has emerged as a promising paradigm, enabling query image segmentation by conditioning on example image-mask pairs provided as prompts. Unlike previous approaches that rely on implicit modeling or non-end-to-end pipelines, we redefine the core interaction mechanism in ICL as an explicit retrieval process, termed E-ICL, benefiting from the emergence of vision foundation models (VFMs). E-ICL captures dense correspondences between queries and prompts at minimal learning cost and leverages them to dynamically weight multi-class prompt masks. Built upon E-ICL, we propose EICSeg, the first end-to-end ICL framework that integrates complementary VFMs for universal medical image segmentation. Specifically, we introduce a lightweight SD-Adapter to bridge the distinct functionalities of the VFMs, enabling more accurate segmentation predictions. To fully exploit the potential of EICSeg, we further design a scalable self-prompt training strategy and an adaptive token-to-image prompt selection mechanism, facilitating both efficient training and inference. EICSeg is trained on 47 datasets covering diverse modalities and segmentation targets. Experiments on nine unseen datasets demonstrate its strong few-shot generalization ability, achieving an average Dice score of 74.0%, outperforming existing in-context and few-shot methods by 4.5%, and reducing the gap to task-specific models to 10.8%. Even with a single prompt, EICSeg achieves a competitive average Dice score of 60.1%. Notably, it performs automatic segmentation without manual prompt engineering, delivering results comparable to interactive models while requiring minimal labeled data. Source code will be available at https://github.com/zerone-fg/EICSeg. Shiao Xie, Liangjun Zhang, Ziwei Niu, Fanfan Ye, Qiaoyong Zhong, Di Xie, Yen-Wei Chen 0001, Lanfen Lin |
IEEE Trans. Medical Imaging | 7 |
| 2026 | Disentangled Multimodal Tuning and Interaction for Human Perception UnderstandingabstractUnderstanding human perceptions poses a significant multimodal challenge for computers, involving textual, acoustic, and visual signals. Recently, large language models (LLMs) have garnered great attention, leading to numerous methods aimed at efficiently fine-tuning pretrained models for multimodal downstream tasks. However, there remains a scarcity of techniques that prioritize modality-invariant and -specific information during parameter-efficient tuning, despite evidence from previous studies showcasing the effectiveness of modality disentangling. To address this gap, we propose a novel multimodal tuning approach for LLMs, termed Disentangled Multimodal Tuning and Interaction. Specifically, we evaluate the independence among different modalities and disentangle corresponding modality-invariant and specific components, which are subsequently leveraged for prompt tuning. Following tuning, a newly designed independence-guided cross-attention module is introduced for modality interaction, where the attention mechanism is decoupled and bolstered with independence from the modality-disentangling process. This approach not only enables LLMs to efficiently assimilate information from various modalities but also cultivates an awareness of both modality-invariant and specific information. Compared to previous methods, our approach facilitates modality interaction at a more granular level, resulting in enhanced performance. We validate our method through experiments on four public datasets, demonstrating significant performance improvements. Hao Sun 0013, Ziwei Niu, Jiaqing Liu, Yen-Wei Chen 0001, Lanfen Lin |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2025 | M2OST: Many-to-one Regression for Predicting Spatial Transcriptomics from Digital Pathology ImagesabstractThe advancement of Spatial Transcriptomics (ST) has facilitated the spatially-aware profiling of gene expressions based on histopathology images. Although ST data offers valuable insights into the micro-environment of tumors, its acquisition cost remains expensive. Therefore, directly predicting the ST expressions from digital pathology images is desired. Current methods usually adopt existing regression backbones along with patch-sampling for this task, which ignores the inherent multi-scale information embedded in the pyramidal data structure of digital pathology images, and wastes the inter-spot visual information crucial for accurate gene expression prediction. To address these limitations, we propose M2OST, a many-to-one regression Transformer that can accommodate the hierarchical structure of the pathology images via a decoupled multi-scale feature extractor. Unlike traditional models that are trained with one-to-one image-label pairs, M2OST uses multiple images from different levels of the digital pathology image to jointly predict the gene expressions in their common corresponding spot. Built upon our many-to-one scheme, M2OST can be easily scaled to fit different numbers of inputs, and its network structure inherently incorporates nearby inter-spot features, enhancing regression performance. We have tested M2OST on three public ST datasets and the experimental results show that M2OST can achieve state-of-the-art performance with fewer parameters and floating-point operations (FLOPs). Hongyi Wang 0002, Xiuju Du, Jing Liu 0041, Shuyi Ouyang, Yen-Wei Chen 0001, Lanfen Lin |
AAAI | 5 |
| 2025 | Triple-Prompt Controllable Diffusion for Universal Data Augmentation in Medical Image SegmentationabstractMedical image segmentation is a crucial yet challenging task in image analysis across diverse anatomical structures. Current segmentation models heavily depend on large-scale datasets, which are laborious to collect and annotate. While generative models offer a promising alternative for data augmentation, most existing approaches are limited to single-modality outputs, either synthetic images or segmentation masks. Moreover, these methods often lack flexible conditioning mechanisms and struggle to capture the rich contextual dependencies inherent in anatomical structures. To address these challenges, in this paper, we propose TPCDM, a novel framework that co-synthesizes high-fidelity paired medical images and segmentation masks through a unified Triple-Prompt Conditional Diffusion Model. At the heart of TPCDM lies a newly defined joint image-label generation paradigm, termed Coordinated Distribution Learning, governed by three synergistic prompts: (1) a text prompt encoding global anatomical semantics; (2) a spatial prompt enforcing pixel-wise spatial coherence; (3) a task prompt dynamically adapting to diverse distributions. Furthermore, TPCDM disentangles instance-wise annotations into semantic masks and distance maps, enabling seamless extension to instance segmentation tasks. Extensive experiments on four benchmarks demonstrate that TPCDM achieves superior synthesis quality. Besides, incorporating the synthesized samples leads to state-of-the-art performance in both downstream semantic and instance segmentation tasks, while also delivering significant improvements under limited labeled data. Shiao Xie, Hongyi Wang 0002, Liangjun Zhang, Ziwei Niu, Yen-Wei Chen 0001, Lanfen Lin |
ECAI | 6 |
| 2025 | Enhanced Multimodal Depression Detection With Emotion PromptsabstractDepression is a pervasive mental health disorder that remains frequently undiagnosed and untreated due to societal barriers and the subjective nature of its symptoms. Leveraging recent advances in large language models (LLMs), we propose a novel depression detection pipeline that generates emotion prompts tailored to individual data, enhancing detection accuracy. Our approach integrates cross-modality fusion via cross attention mechanisms to combine depressive and emotional features, creating a comprehensive representation of depression indicators. Evaluated on the E-DAIC and EATD datasets, our method outperforms state-of-the-art techniques, demonstrating its potential for more precise emotion-based depression detection. Shiyu Teng, Jiaqing Liu, Hao Sun 0013, Shurong Chai, Tomoko Tateyama, Lanfen Lin, Yen-Wei Chen 0001 |
ICASSP | 7 |
| 2025 | Region-Aware Anchoring Mechanism for Efficient Referring Visual Grounding
Shuyi Ouyang, Ziwei Niu, Hongyi Wang 0002, Yen-Wei Chen 0001, Lanfen Lin |
ICCV | 4 |
| 2025 | EPIC: Efficient Prompt Interaction for Text-Image ClassificationabstractIn recent years, large-scale pre-trained multimodal models (LMMs) generally emerge to integrate the vision and language modalities, achieving considerable success in multimodal tasks, such as text-image classification. The growing size of LMMs, however, results in a significant computational cost for fine-tuning these models for downstream tasks. Hence, prompt-based interaction strategy is studied to align modalities more efficiently. In this context, we propose a novel efficient prompt-based multimodal interaction strategy, namely Efficient Prompt Interaction for text-image Classification (EPIC). Specifically, we utilize temporal prompts on intermediate layers, and integrate different modalities with similarity-based prompt interaction, to leverage sufficient information exchange between modalities. Utilizing this approach, our method achieves reduced computational resource consumption and fewer trainable parameters (about 1% of the foundation model) compared to other fine-tuning strategies. Furthermore, it demonstrates superior performance on the UPMC-Food101 and SNLI-VE datasets, while achieving comparable performance on the MM-IMDB dataset. Xinyao Yu 0003, Hao Sun 0013, Zeyu Ling, Ziwei Niu, Zhenjia Bai, Yen-Wei Chen 0001, Lanfen Lin |
ICME | 7 |
| 2025 | Clinical Data-Driven Retrieval-Augmented Model for Lung Nodule Malignancy Prediction
Ruibo Hou, Shurong Chai, Rahul Kumar Jain 0001, Yinhao Li 0002, Jiaqing Liu, Shiyu Teng, Lanfen Lin, Yen-Wei Chen 0001 |
MICCAI (10) | 9 |
| 2025 | TextBraTS: Text-Guided Volumetric Brain Tumor Segmentation with Innovative Dataset Development and Fusion Module Exploration
Rahul Kumar Jain 0001, Yinhao Li 0002, Ruibo Hou, Jingliang Cheng, Guohua Zhao, Lanfen Lin, Rui Xu 0002, Yen-Wei Chen 0001 |
MICCAI (6) | 10 |
| 2025 | EchoCardMAE: Video Masked Auto-Encoders Customized for Echocardiography
Rui Xu 0002, Xinchen Ye, Zhihui Wang 0001, Miao Zhang 0004, Yi Wang 0037, Xin Fan 0001, Hongkai Wang 0002, Qingxiong Yue, Xiangjian He, Yen-Wei Chen 0001 |
MICCAI (13) | 11 |
| 2025 | PD-UniST: Prompt-Driven Universal Model for Unpaired H&E-to-IHC Stain Translation
Chujie Zhang, Yangyang Xie, Yinhao Li 0002, Lanfen Lin, Yen-Wei Chen 0001 |
MICCAI (2) | 6 |
| 2025 | EIR-SDG: Explore Invariant Representation for Single-source Domain Generalization in Medical Image Segmentation
Ziwei Niu, Shiao Xie, Ziyue Wang 0005, Yen-Wei Chen 0001, Yueming Jin, Lanfen Lin |
ACM Multimedia | 4 |
| 2025 | Semantic-assisted report generation with memory enhanced transformer using context-aware visual extractor
Peketi Divya, Partha Sarathi Chakraborty 0003, C. Krishna Mohan, Yen-Wei Chen 0001 |
Appl. Intell. | 4 |
| 2025 | Multi-modal Medical SAM: An Adaptation Method of Segment Anything Model (SAM) for Glioma Segmentation Using Multi-modal MR ImagesabstractThe segmentation of glioma is crucial for early diagnosis, according to a World Health Organization (WHO) 2021 report. For glioma diagnosis, 3D multi-modal brain MRI/CT imaging has become an essential tool, offering detailed information. Nowadays, deep learning frameworks have been applied to various medical imaging problems, including brain glioma segmentation. Recently, foundation models like Segment Anything Model (SAM) have emerged as pivotal tools in computer vision tasks. These models are trained using large (real-world) datasets, offering a generalized understanding of visual data and semantic key features. Therefore, the effective utilization of foundation models in medical imaging is a significant area of current research. However, the differences in data distribution between multi-modal medical images and real-world images present challenges in directly applying foundation models to medical imaging. Additionally, utilizing multi-modal images to extract crucial information and its fusion poses further challenges. To address these issues, we propose a framework using foundation model and novel strategies for multi-modal fusion. Our fusion adapters effectively integrate the information from different modalities to enhance glioma segmentation in multi-modal MRI scans. Our method outperforms current state-of-the-art methods for accurate segmentation of the glioma using private and publicly available brain MRI datasets, proving the effectiveness of our approach across different datasets and imaging modalities. Rahul Kumar Jain 0001, Yinhao Li 0002, Shurong Chai, Jingliang Cheng, Guohua Zhao, Lanfen Lin, Yen-Wei Chen 0001 |
ACM Trans. Comput. Heal. | 9 |
| 2025 | Multimodal Sentiment Analysis With Mutual Information-Based Disentangled Representation LearningabstractMultimodal sentiment analysis seeks to utilize various types of signals to identify underlying emotions and sentiments. A key challenge in this field lies in multimodal representation learning, which aims to develop effective methods for integrating multimodal features into cohesive representations. Recent advancements include two notable approaches: one focuses on decomposing multimodal features into modality-invariant and -specific components, while the other emphasizes the use of mutual information to enhance the fusion of modalities. Both strategies have demonstrated effectiveness and yielded remarkable results. In this paper, we propose a novel learning framework that combines the strengths of these two approaches, termed mutual information-based disentangled multimodal representation learning. Our approach involves estimating different types of information during feature extraction and fusion stages. Specifically, we quantitatively assess and adjust the proportions of modality-invariant, -specific, and -complementary information during feature extraction. Subsequently, during fusion, we evaluate the amount of information retained by each modality in the fused representation. We employ mutual information or conditional mutual information to estimate each type of information content. By reconciling the proportions of these different types of information, our approach achieves state-of-the-art performance on popular sentiment analysis benchmarks, including CMU-MOSI and CMU-MOSEI. Hao Sun 0013, Ziwei Niu, Hongyi Wang 0002, Xinyao Yu 0003, Jiaqing Liu, Yen-Wei Chen 0001, Lanfen Lin |
IEEE Trans. Affect. Comput. | 6 |
| 2025 | SAMA: A Self-and-Mutual Attention Network for Accurate Recurrence Prediction of Non-Small Cell Lung Cancer Using Genetic and CT DataabstractAccurate preoperative recurrence prediction for non-small cell lung cancer (NSCLC) is a challenging issue in the medical field. Existing studies primarily conduct image and molecular analyses independently or directly fuse multimodal information through radiomics and genomics, which fail to fully exploit and effectively utilize the highly heterogeneous cross-modal information at different levels and model the complex relationships between modalities, resulting in poor fusion performance and becoming the bottleneck of precise recurrence prediction. To address these limitations, we propose a novel unified framework, the Self-and-Mutual Attention (SAMA) Network, designed to efficiently fuse and utilize macroscopic CT images and microscopic gene data for precise NSCLC recurrence prediction, integrating handcrafted features, deep features, and gene features. Specifically, we design a Self-and-Mutual Attention Module that performs three-stage fusion: the self-enhancement stage enhances modality-specific features; the gene-guided and CT-guided cross-modality fusion stages perform bidirectional cross-guidance on the self-enhanced features, complementing and refining each modality, enhancing heterogeneous feature expression; and the optimized feature aggregation stage ensures the refined interactive features for precise prediction. Extensive experiments on both publicly available datasets from The Cancer Imaging Archive (TCIA) and The Cancer Genome Atlas (TCGA) demonstrate that our method achieves state-of-the-art performance and exhibits broad applicability to various cancers. Yang Ai, Jing Liu 0041, Yinhao Li 0002, Fang Wang 0030, Xiuju Du, Rahul Kumar Jain 0001, Lanfen Lin, Yen-Wei Chen 0001 |
IEEE J. Biomed. Health Informatics | 8 |
| 2024 | Combinatorial CNN-Transformer Learning with Manifold Constraints for Semi-supervised Medical Image SegmentationabstractSemi-supervised learning (SSL), as one of the dominant methods, aims at leveraging the unlabeled data to deal with the annotation dilemma of supervised learning, which has attracted much attentions in the medical image segmentation. Most of the existing approaches leverage a unitary network by convolutional neural networks (CNNs) with compulsory consistency of the predictions through small perturbations applied to inputs or models. The penalties of such a learning paradigm are that (1) CNN-based models place severe limitations on global learning; (2) rich and diverse class-level distributions are inhibited. In this paper, we present a novel CNN-Transformer learning framework in the manifold space for semi-supervised medical image segmentation. First, at intra-student level, we propose a novel class-wise consistency loss to facilitate the learning of both discriminative and compact target feature representations. Then, at inter-student level, we align the CNN and Transformer features using a prototype-based optimal transport method. Extensive experiments show that our method outperforms previous state-of-the-art methods on three public medical image segmentation benchmarks. Huimin Huang 0002, Yawen Huang, Shiao Xie, Lanfen Lin, Ruofeng Tong 0001, Yen-Wei Chen 0001, Yuexiang Li, Yefeng Zheng 0001 |
AAAI | 6 |
| 2024 | Going Beyond Multi-Task Dense Prediction with Synergy Embedding ModelsabstractMulti-task visual scene understanding aims to leverage the relationships among a set of correlated tasks, which are solved simultaneously by embedding them within a unified network. However, most existing methods give rise to two primary concerns from a task-level perspective: (1) the lack of task-independent correspondences for distinct tasks, and (2) the neglect of explicit task-consensual dependencies among various tasks. To address these issues, we propose a novel synergy embedding models (SEM), which goes beyond multi-task dense prediction by leveraging two innovative designs: the intra-task hierarchy-adaptive module and the inter-task EM-interactive module. Specifically, the constructed intra-task module incorporates hierarchy-adaptive keys from multiple stages, enabling the efficient learning of specialized visual patterns with an optimal trade-off. In addition, the developed inter-task module learns interactions from a compact set of mutual bases among various tasks, benefiting from the expectation maximization (EM) algorithm. Extensive empirical evidence from two public benchmarks, NYUD-v2 and PASCAL-Context, demonstrates that SEM consistently outperforms state-of-the-art approaches across a range of metrics. Huimin Huang 0002, Yawen Huang, Lanfen Lin, Ruofeng Tong 0001, Yen-Wei Chen 0001, Hao Zheng 0008, Yuexiang Li, Yefeng Zheng 0001 |
CVPR | 5 |
| 2024 | Hyperspectral Image Reconstruction Using Hierarchical Neural Architecture Search from A Snapshot ImageabstractHyperspectral imaging is a promising imaging modality, and has attracted increasing research attention by compressive sensing such as coded aperture snapshot spectral imaging (CASSI), for simultaneously capturing abundant information in spatial, spectral and temporal domains. Hyperspectral image (HSI) reconstruction in the CASSI aims to retrieve the original 3D signal upon the 2D compressed snapshot. Recently, deep learning has extensively been employed for HSI reconstruction via manually designing network architectures, and usually causes complicated and massive-computational models, which are difficult for being embedding in the real imaging systems. This study aims to leverage network architecture search to automatically design effective and efficient network architectures for HSI reconstruction. Specifically, we exploit gradient-based search strategies and prepare optional operations (cells) with adaptive receptive field such as dilate and deformable convolutional layers to construct a flexible hierarchical search space. Through sharing cells within different levels of features and utilizing an early stopping technique, we achieve a computational and memory efficient NAS strategy to automatically design an effective lightweight model for HSI reconstruction. Extensive experimental results have demonstrated that the network architecture achieved by our proposed NAS has much smaller model size and a lower computational cost while produce better or comparable HSI reconstruction performance compared with the state-of-the-art methods. Xianhua Han, Huiyan Jiang, Yen-Wei Chen 0001 |
ICASSP | 3 |
| 2024 | IRLSG: Invariant Representation Learning for Single-Domain Generalization in Medical Image SegmentationabstractSingle-domain generalization (SDG) can efficiently enhance model generalization while avoiding high annotation costs and privacy concerns. However, existing SDG methods are mainly based on data manipulation and meta-learning, which are not efficient enough due to the limited generalization performance and complex inference. In response to these challenges, we present a novel single domaininvariant representation learning approach for medical image segmentation, called IRLSG, with two appealing designs: (1) A Classscale Photo-metric Augmentation is first proposed to simulate unseen target domain that is sufficient in diversity and informativeness. After that, a Dual-Consistency Framework is further designed to constrain the consistency of intermediate features and segmentation results between the original and the augmented images, which helps to explore the domain-invariant representation. (2) A simple and effective Style Feature Whitening is designed to decouple and remove the domain-specific style from higher-order covariance statistics, which can further improve the modeling and generalization capability of the network. Experimental results on different benchmarks demonstrate that our IRLSG outperforms the current state-of-the-art methods in tackling single-domain generalization. Ziwei Niu, Hao Sun 0013, Shuyi Ouyang, Shiao Xie, Yen-Wei Chen 0001, Ruofeng Tong 0001, Lanfen Lin |
ICASSP | 5 |
| 2024 | Deep Versatile Hyperspectral Reconstruction Model from A Snapshot Measurement with Arbitrary MasksabstractRecently, coded aperture snapshot spectral imaging (CASSI) has been actively researched to capture three-dimensional (3D) hyperspectral (HS) images for dynamic scenes, where the optical systems detect a 2D snapshot measurement while a computational algorithm performs the inverse problem for recovering the latent HS cubic data. Benefiting from the powerful modeling capability of the deep convolution neural networks (DCNN), the reconstruction performance of the HS images has been significantly improved. However, the existing deep methods usually assume a particular hardware mask to train the reconstruction models, and restrict widely applicability to the snapshots measured in different hardwares. This study exploits a novel deep versatile HS reconstruction framework for adaptively handling the snapshots with arbitrary masks. Specifically, we employ a meta-learning like training procedure using the training paired samples of different distributions to learn a highly generalized model, and further incorporate a mask structure modeling module to produce effective knowledge for modulating the spectral recovering procedure. Moreover, we configure the deep reconstruction model with the spectral transformer for modeling the long-dependence in spectral domain, which is especially critical for high fidelity spectral recovering. Experiments on two benchmark HS datasets have demonstrated the superiority of our framework over the state-of-the-art methods. Takumi Takabe, Xianhua Han, Yen-Wei Chen 0001 |
ICASSP | 3 |
| 2024 | SOFIM: Stochastic Optimization Using Regularized Fisher Information MatrixabstractThis paper introduces a new stochastic optimization method based on the regularized Fisher information matrix (FIM), named SOFIM, which can efficiently utilize the FIM to approximate the Hessian matrix for finding Newton’s gradient update in large-scale stochastic optimization of machine learning models. It can be viewed as a variant of natural gradient descent, where the challenge of storing and calculating the full FIM is addressed through making use of the regularized FIM and directly finding the gradient update direction via Sherman-Morrison matrix inversion. Additionally, like the popular Adam method, SOFIM uses the first moment of the gradient to address the issue of non-stationary objectives across mini-batches due to heterogeneous data. The utilization of the regularized FIM and Sherman-Morrison matrix inversion leads to the improved convergence rate with the same space and time complexities as stochastic gradient descent (SGD) with momentum. The extensive experiments on training deep learning models using several benchmark image classification datasets demonstrate that the proposed SOFIM outperforms SGD with momentum and several state-of-the-art Newton optimization methods in term of the convergence speed for achieving the pre-specified objectives of training and test losses as well as test accuracy. Mrinmay Sen, A. K. Qin 0001, Gayathri C, Raghu Kishore N, Yen-Wei Chen 0001, Balasubramanian Raman |
IJCNN | 5 |
| 2024 | Ladder Fine-tuning Approach for SAM Integrating Complementary NetworkabstractRecently, foundation models have been introduced demonstrating various tasks in the field of computer vision. These models such as Segment Anything Model (SAM) are generalized models trained using huge datasets. Currently, ongoing research focuses on exploring the effective utilization of these generalized models for Specific domains, such as medical imaging. However, in medical imaging, the lack of training samples due to privacy concerns and other factors presents a major challenge for applying these generalized models to medical image segmentation task. To address this issue, the effective fine tuning of these models is crucial to ensure their optimal utilization. In this study, we propose to combine a complementary Convolutional Neural Network (CNN) along with the standard SAM network for medical image segmentation. To reduce the burden of fine tuning large foundation model and implement cost-efficient training scheme, we focus only on fine-tuning the additional CNN network and SAM decoder part. This strategy significantly reduces training time and achieves competitive results on publicly available dataset. The code is available at ">https://github.com/11yxk/SAM-LST . Shurong Chai, Rahul Kumar Jain 0001, Shiyu Teng, Jiaqing Liu, Yinhao Li 0002, Tomoko Tateyama, Yen-Wei Chen 0001 |
KES | 7 |
| 2024 | A Novel Adaptive Hypergraph Neural Network for Enhancing Medical Image Segmentation
Shurong Chai, Rahul Kumar Jain 0001, Shaocong Mo, Jiaqing Liu, Yinhao Li 0002, Tomoko Tateyama, Lanfen Lin, Yen-Wei Chen 0001 |
MICCAI (9) | 9 |
| 2024 | LGA: A Language Guide Adapter for Advancing the SAM Model's Capabilities in Medical Image Segmentation
Jihong Hu, Yinhao Li 0002, Hao Sun 0013, Yu Song 0008, Chujie Zhang, Lanfen Lin, Yen-Wei Chen 0001 |
MICCAI (12) | 7 |
| 2024 | FedEL: Federated ensemble learning for non-iid data
Xing Wu 0001, Jie Pei, Xianhua Han, Yen-Wei Chen 0001, Junfeng Yao, Yang Liu 0005, Quan Qian, Yike Guo |
Expert Syst. Appl. | 4 |
| 2024 | A motion-aware and temporal-enhanced Spatial-Temporal Graph Convolutional Network for skeleton-based human action segmentation
Shurong Chai, Rahul Kumar Jain 0001, Jiaqing Liu, Shiyu Teng, Tomoko Tateyama, Yinhao Li 0002, Yen-Wei Chen 0001 |
Neurocomputing | 7 |
| 2024 | Memory Guided Transformer With Spatio-Semantic Visual Extractor for Medical Report GenerationabstractMedicalimaging-based report writing for effective diagnosis in radiology is time-consuming and can be error-prone by inexperienced radiologists. Automatic reporting helps radiologists avoid missed diagnoses and saves valuable time. Recently, transformer-based medical report generation has become prominent in capturing long-term dependencies of sequential data with its attention mechanism. Nevertheless, input features obtained from traditional visual extractor of conventional transformers do not capture spatial and semantic information of an image. So, the transformer is unable to capture fine-grained details and may not produce detailed descriptive reports of radiology images. Therefore, we propose a spatio-semantic visual extractor (SSVE) to capture multi-scale spatial and semantic information from radiology images. Here, we incorporate two types of networks in ResNet 101 backbone architecture, i.e. (i) deformable network at the intermediate layer of ResNet 101 that utilizes deformable convolutions in order to obtain spatially invariant features, and (ii) semantic network at the final layer of backbone architecture which uses dilated convolutions to extract rich multi-scale semantic information. Further, these network representations are fused to encode fine-grained details of radiology images. The performance of our proposed model outperforms existing works on two radiology report datasets, i.e., IU X-ray and MIMIC-CXR. Peketi Divya, Sravani Yenduri, Chalavadi Vishnu, C. Krishna Mohan, Yen-Wei Chen 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2024 | Segmentation Guided Crossing Dual Decoding Generative Adversarial Network for Synthesizing Contrast-Enhanced Computed Tomography ImagesabstractAlthough contrast-enhanced computed tomography (CE-CT) images significantly improve the accuracy of diagnosing focal liver lesions (FLLs), the administration of contrast agents imposes a considerable physical burden on patients. The utilization of generative models to synthesize CE-CT images from non-contrasted CT images offers a promising solution. However, existing image synthesis models tend to overlook the importance of critical regions, inevitably reducing their effectiveness in downstream tasks. To overcome this challenge, we propose an innovative CE-CT image synthesis model called Segmentation Guided Crossing Dual Decoding Generative Adversarial Network (SGCDD-GAN). Specifically, the SGCDD-GAN involves a crossing dual decoding generator including an attention decoder and an improved transformation decoder. The attention decoder is designed to highlight some critical regions within the abdominal cavity, while the improved transformation decoder is responsible for synthesizing CE-CT images. These two decoders are interconnected using a crossing technique to enhance each other's capabilities. Furthermore, we employ a multi-task learning strategy to guide the generator to focus more on the lesion area. To evaluate the performance of proposed SGCDD-GAN, we test it on an in-house CE-CT dataset. In both CE-CT image synthesis tasks-namely, synthesizing ART images and synthesizing PV images-the proposed SGCDD-GAN demonstrates superior performance metrics across the entire image and liver region, including SSIM, PSNR, MSE, and PCC scores. Furthermore, CE-CT images synthetized from our SGCDD-GAN achieve remarkable accuracy rates of 82.68%, 94.11%, and 94.11% in a deep learning-based FLLs classification task, along with a pilot assessment conducted by two radiologists. Qingqing Chen 0001, Yinhao Li 0002, Fang Wang 0030, Xianhua Han, Yutaro Iwamoto, Jing Liu 0041, Lanfen Lin, Hongjie Hu, Yen-Wei Chen 0001 |
IEEE J. Biomed. Health Informatics | 10 |
| 2024 | Rethinking Multiple Instance Learning for Whole Slide Image Classification: A Bag-Level Classifier is a Good Instance-Level TeacherabstractMultiple Instance Learning (MIL) has demonstrated promise in Whole Slide Image (WSI) classification. However, a major challenge persists due to the high computational cost associated with processing these gigapixel images. Existing methods generally adopt a two-stage approach, comprising a non-learnable feature embedding stage and a classifier training stage. Though it can greatly reduce memory consumption by using a fixed feature embedder pre-trained on other domains, such a scheme also results in a disparity between the two stages, leading to suboptimal classification accuracy. To address this issue, we propose that a bag-level classifier can be a good instance-level teacher. Based on this idea, we design Iteratively Coupled Multiple Instance Learning (ICMIL) to couple the embedder and the bag classifier at a low cost. ICMIL initially fixes the patch embedder to train the bag classifier, followed by fixing the bag classifier to fine-tune the patch embedder. The refined embedder can then generate better representations in return, leading to a more accurate classifier for the next iteration. To realize more flexible and more effective embedder fine-tuning, we also introduce a teacher-student framework to efficiently distill the category knowledge in the bag classifier to help the instance-level embedder fine-tuning. Intensive experiments were conducted on four distinct datasets to validate the effectiveness of ICMIL. The experimental results consistently demonstrated that our method significantly improves the performance of existing MIL backbones, achieving state-of-the-art results. The code and the organized datasets can be accessed by: https://github.com/Dootmaan/ICMIL/tree/confidence-based. Hongyi Wang 0002, Luyang Luo, Fang Wang 0030, Ruofeng Tong 0001, Yen-Wei Chen 0001, Hongjie Hu, Lanfen Lin, Hao Chen 0011 |
IEEE Trans. Medical Imaging | 5 |
| 2024 | Knowledge Distillation-Based Domain-Invariant Representation Learning for Domain GeneralizationabstractDomain generalization (DG) aims to generalize the knowledge learned from multiple source domains to unseen target domains. Existing DG techniques can be subsumed under two broad categories, i.e., domain-invariant representation learning and domain manipulation. Nevertheless, it is extremely difficult to explicitly augment or generate the unseen target data. And when source domain variety increases, developing a domain-invariant model by simply aligning more domain-specific information becomes more challenging. In this paper, we propose a simple yet effective method for domain generalization, named Knowledge Distillation based Domain-invariant Representation Learning (KDDRL), that learns domain-invariant representation while encouraging the model to maintain domain-specific features, which recently turned out to be effective for domain generalization. To this end, our method incorporates multiple auxiliary student models and one student leader model to perform a two-stage distillation. In the first-stage distillation, each domain-specific auxiliary student treats the ensemble of other auxiliary students' predictions as a target, which helps to excavate the domain-invariant representation. Also, we present an error removal module to prevent the transfer of faulty information by eliminating incorrect predictions compared to the true labels. In the second-stage distillation, the student leader model with domain-specific features combines the domain-invariant representation learned from the group of auxiliary students to make the final prediction. Extensive experiments and in-depth analysis on popular DG benchmark datasets demonstrate that our KDDRL significantly outperforms the current state-of-the-art methods. Ziwei Niu, Junkun Yuan, Jing Liu 0041, Yen-Wei Chen 0001, Ruofeng Tong 0001, Lanfen Lin |
IEEE Trans. Multim. | 6 |
| 2023 | ClassFormer: Exploring Class-Aware Dependency with Transformer for Medical Image SegmentationabstractVision Transformers have recently shown impressive performances on medical image segmentation. Despite their strong capability of modeling long-range dependencies, the current methods still give rise to two main concerns in a class-level perspective: (1) intra-class problem: the existing methods lacked in extracting class-specific correspondences of different pixels, which may lead to poor object coverage and/or boundary prediction; (2) inter-class problem: the existing methods failed to model explicit category-dependencies among various objects, which may result in inaccurate localization. In light of these two issues, we propose a novel transformer, called ClassFormer, powered by two appealing transformers, i.e., intra-class dynamic transformer and inter-class interactive transformer, to address the challenge of fully exploration on compactness and discrepancy. Technically, the intra-class dynamic transformer is first designed to decouple representations of different categories with an adaptive selection mechanism for compact learning, which optimally highlights the informative features to reflect the salient keys/values from multiple scales. We further introduce the inter-class interactive transformer to capture the category dependency among different objects, and model class tokens as the representative class centers to guide a global semantic reasoning. As a consequence, the feature consistency is ensured with the expense of intra-class penalization, while inter-class constraint strengthens the feature discriminability between different categories. Extensive empirical evidence shows that ClassFormer can be easily plugged into any architecture, and yields improvements over the state-of-the-art methods in three public benchmarks. Huimin Huang 0002, Shiao Xie, Lanfen Lin, Ruofeng Tong 0001, Yen-Wei Chen 0001, Hong Wang 0021, Yuexiang Li, Yawen Huang, Yefeng Zheng 0001 |
AAAI | 5 |
| 2023 | SemiCVT: Semi-Supervised Convolutional Vision Transformer for Semantic SegmentationabstractSemi-supervised learning improves data efficiency of deep models by leveraging unlabeled samples to alleviate the reliance on a large set of labeled samples. These successes concentrate on the pixel-wise consistency by using convolutional neural networks (CNNs) but fail to address both global learning capability and class-level features for unlabeled data. Recent works raise a new trend that Transformer achieves superior performance on the entire feature map in various tasks. In this paper, we unify the current dominant Mean-Teacher approaches by reconciling intra-model and inter-model properties for semi-supervised segmentation to produce a novel algorithm, SemiCVT, that absorbs the quintessence of CNNs and Transformer in a comprehensive way. Specifically, we first design a parallel CNN-Transformer architecture (CVT) with introducing an intra-model local-global interaction schema (LGI) in Fourier domain for full integration. The inter-model class-wise consistency is further presented to complement the class-level statistics of CNNs and Transformer in a cross-teaching manner. Extensive empirical evidence shows that SemiCVT yields consistent improvements over the state-of-the-art methods in two public benchmarks. Huimin Huang 0002, Shiao Xie, Lanfen Lin, Ruofeng Tong 0001, Yen-Wei Chen 0001, Yuexiang Li, Hong Wang 0021, Yawen Huang, Yefeng Zheng 0001 |
CVPR | 5 |
| 2023 | MCKD: Mutually Collaborative Knowledge Distillation For Federated Domain Adaptation And GeneralizationabstractConventional unsupervised domain adaptation (UDA) and domain generalization (DG) methods rely on the assumption that all source domains can be directly accessed and combined for model training. However, this centralized training strategy may violate privacy policies in many real-world applications. A paradigm for tackling this problem is to train multiple local models and aggregate a generalized central model without data sharing. Recent methods have made remarkable advancements in this paradigm by exploiting parameter alignment and aggregation. But when sources domain variety increases, directly aligning and aggregating local parameters becomes more challenging. Adapting a different approach in this work, we devised a data-free semantic collaborative distillation strategy to learn domain-invariant representation for both federated UDA and DG. Each local model transmits its predictions to the central server and derives its target distribution from the average of other local models' distributions to facilitate the mutual transfer of domain-specific knowledge. When unlabeled target data is available, we introduce a novel UDA strategy termed knowledge filter to adapt the central model to the target data. Extensive experiments on four UDA and DG datasets demonstrate that our method has a competitive performance compared with the state-of-the-art methods. Ziwei Niu, Hongyi Wang 0002, Hao Sun 0013, Shuyi Ouyang, Yen-Wei Chen 0001, Lanfen Lin |
ICASSP | 5 |
| 2023 | MedFCT: A Frequency Domain Joint CNN-Transformer Network for Semi-supervised Medical Image SegmentationabstractSemi-supervised learning(SSL) is a data-efficient way in leveraging large-scale data without annotations and alleviating the dependence on labeled data. Mean-Teacher (MT) scheme with teacher-student model architecture has shown its effectiveness in semi-supervised medical image segmentation, where the student network learns from the teacher by minimizing pixel-wise consistency loss. However, existing MT-based SSLs still give rise to two main concerns: (1) limited learning ability of student network that neglects the union of local feature and global cues extraction which may impact the representation learning of variable objects. (2) limited knowledge-transferring ability of teacher network with only pixel-level consistency regularization that may result in inadequate and unstable guidance. To address these limitations, we propose a novel semi-supervised learning scheme, namely MedFCT, with two appealing designs: (1) A dual student architecture with parallel CNN and Transformer branches is designed for local-global feature extraction, where the full-frequency interaction between CNN and Transformer can be explored by a frequency domain cross-fusion (FDCF) module to learn complementarity of the two-paradigm features. (2) A comprehensive multi-level consistency regularization considering pixel-wise, feature-wise and class-wise information is presented to realize more effective guidance and knowledge transfer from teacher network. Experiments show that MedFCT outperforms previous state-of-the-art methods on two public medical image segmentation benchmarks. Shiao Xie, Huimin Huang 0002, Ziwei Niu, Lanfen Lin, Yen-Wei Chen 0001 |
ICME | 5 |
| 2023 | SLViT: Scale-Wise Language-Guided Vision Transformer for Referring Image SegmentationabstractReferring image segmentation aims to segment an object out of an image via a specific language expression. The main concept is establishing global visual-linguistic relationships to locate the object and identify boundaries using details of the image. Recently, various Transformer-based techniques have been proposed to efficiently leverage long-range cross-modal dependencies, enhancing performance for referring segmentation. However, existing methods consider visual feature extraction and cross-modal fusion separately, resulting in insufficient visual-linguistic alignment in semantic space. In addition, they employ sequential structures and hence lack multi-scale information interaction. To address these limitations, we propose a Scale-Wise Language-Guided Vision Transformer (SLViT) with two appealing designs: (1) Language-Guided Multi-Scale Fusion Attention, a novel attention mechanism module for extracting rich local visual information and modeling global visual-linguistic relationships in an integrated manner. (2) An Uncertain Region Cross-Scale Enhancement module that can identify regions of high uncertainty using linguistic features and refine them via aggregated multi-scale features. We have evaluated our method on three benchmark datasets. The experimental results demonstrate that SLViT surpasses state-of-the-art methods with lower computational cost. The code is publicly available at: https://github.com/NaturalKnight/SLViT. Shuyi Ouyang, Hongyi Wang 0002, Shiao Xie, Ziwei Niu, Ruofeng Tong 0001, Yen-Wei Chen 0001, Lanfen Lin |
IJCAI | 6 |
| 2023 | FLWGAN: Federated Learning with Wasserstein Generative Adversarial Network for Brain Tumor SegmentationabstractRecently, the potential of deep learning in identifying complex patterns is gaining research interest in medical applications specifically for brain tumor diagnosis. To segment tumors accurately in brain MRIs, there is a need for a large amount of data for training deep learning models. Also, hospitals cannot share patient data for centralization on the server since health records are prone to privacy and ownership challenges. To deal with these challenges, we set up an efficient federated learning (FL) pipeline with Wasserstein generative adversarial networks (FLWGAN) to ensure data privacy and data sufficiency. FL preserves the data privacy of clients by sharing only the trained model parameters to a centralized server instead of raw data. A modified 3D Wasserstein generative adversarial network with gradient penalty (WGAN-GP) and is incorporated at the client side to generate image-segmentation pairs for efficient training segmentation models. Here, 3D-UNet with an attention module is used for the brain MRI segmentation. The attention module is integrated into a 3D-UNet encoder network for effective brain tumor segmentation. Our approach aims to allow each client to benefit from locally available real data and synthetic data. This process enhances the learning performance while respecting data privacy. The efficacy of our proposed pipeline is demonstrated on the brain tumor task of the medical segmentation decathlon (MSD) dataset. We designed FLWGAN frameworks for predicting four segmentation tasks, i.e., whole tumor (WT), enhanced tumor (ET), tumor core (TC), and multiclass. Our proposed approach achieves state of the art performance in terms of various segmentation metrics. Peketi Divya, Chalavadi Vishnu, C. Krishna Mohan, Yen-Wei Chen 0001 |
IJCNN | 4 |
| 2023 | Iteratively Coupled Multiple Instance Learning from Instance to Bag Classifier for Whole Slide Image Classification
Hongyi Wang 0002, Luyang Luo, Fang Wang 0030, Ruofeng Tong 0001, Yen-Wei Chen 0001, Hongjie Hu, Lanfen Lin, Hao Chen 0011 |
MICCAI (6) | 5 |
| 2023 | Semi-Supervised Convolutional Vision Transformer with Bi-Level Uncertainty Estimation for Medical Image SegmentationabstractSemi-supervised learning (SSL) has attracted much attention in the field of medical image segmentation, which enables to alleviate the heavy burden of labelling pixel-wise annotation by extracting knowledge from unlabeled data. The existing methods basically benefit from the success of convolutional neural networks (CNNs) by keeping consistency of the predictions under small perturbations imposed on the networks or inputs. Two main concerns arise when learning such a paradigm: (1) CNNs tend to retain discriminative local features, neglecting global dependency and thus leading to inaccurate localization; (2) CNNs omit reliable feature-level and pixel-level information, resulting in sketchy pseudo-labels, especially around the confusing boundary. In this paper, we revisit the model of semi-supervised learning and develop a novel CNN-Transformer learning framework that allows for effective segmentation of medical images by producing complementary and reliable features and pseudo-label with bi-level uncertainty. Motivated by the uncertainty estimation to gain insight on feature discrimination, we explore the statistical and geometrical properties of features on network optimization and thus launching an alignment method in a more accurate and stable way. We attach equal significance to pixel-level uncertainty estimation for alleviating the influence of unreliable pseudo-labels in the training progress and advocating the reliability of predictions. Experimental results show that our method significantly surpasses existing semi-supervised approaches on two public medical image segmentation datasets. Huimin Huang 0002, Yawen Huang, Shiao Xie, Lanfen Lin, Ruofeng Tong 0001, Yen-Wei Chen 0001, Yuexiang Li, Yefeng Zheng 0001 |
ACM Multimedia | 6 |
| 2023 | HSVLT: Hierarchical Scale-Aware Vision-Language Transformer for Multi-Label Image ClassificationabstractThe task of multi-label image classification involves recognizing multiple objects within a single image. Considering both valuable semantic information contained in the labels and essential visual features presented in the image, tight visual-linguistic interactions play a vital role in improving classification performance. Moreover, given the potential variance in object size and appearance within a single image, attention to features of different scales can help to discover possible objects in the image. Recently, Transformer-based methods have achieved great success in multi-label image classification by leveraging the advantage of modeling long-range dependencies, but they have several limitations. Firstly, existing methods treat visual feature extraction and cross-modal fusion as separate steps, resulting in insufficient visual-linguistic alignment in the joint semantic space. Additionally, they only extract visual features and perform cross-modal fusion at a single scale, neglecting objects with different characteristics. To address these issues, we propose a Hierarchical Scale-Aware Vision-Language Transformer (HSVLT) with two appealing designs: (1)A hierarchical multi-scale architecture that involves a Cross-Scale Aggregation module, which leverages joint multi-modal features extracted from multiple scales to recognize objects of varying sizes and appearances in images. (2)Interactive Visual-Linguistic Attention, a novel attention mechanism module that tightly integrates cross-modal interaction, enabling the joint updating of visual, linguistic and multi-modal features. We have evaluated our method on three benchmark datasets. The experimental results demonstrate that HSVLT surpasses state-of-the-art methods with lower computational cost. Shuyi Ouyang, Hongyi Wang 0002, Ziwei Niu, Zhenjia Bai, Shiao Xie, Ruofeng Tong 0001, Yen-Wei Chen 0001, Lanfen Lin |
ACM Multimedia | 8 |
| 2023 | IS2Net: Intra-domain Semantic and Inter-domain Style Enhancement for Semi-supervised Medical Domain GeneralizationabstractDomain generalization (DG) demonstrates superior generalization ability in cross-center medical image segmentation. Despite its great success, existing fully supervised DG methods require collecting a large quantity of pixel-level annotations which is quite expensive and time-consuming. To address this challenge, several semi-supervised domain generalized (SSDG) methods have been proposed by simply coupling semi-supervised learning (SSL) with DG tasks, which give rise to two main concerns: (1) Intra-domain dubious semantic information: the quality of pseudo labels in each source domain suffers from the limited amount of labeled data and cross-domain discrepancy. (2) Inter-domain intangible style relationship: current models fail in integrating domain-level information and overlook the relationships among different domains, which degrades the generalization ability of model. In light of these two issues, we propose a novel SSDG framework, namely IS2Net, by arranging an inter-domain generalization branch and several intra-domain SSL branches in a parallel manner, powered by two appealing designs that build a positive interaction between them: (1) A style and semantic memory mechanism is designed to provide both high-quality class-wise representations for intra-domain semantic enhancement and stable domain-specific knowledge for inter-domain style relationship construction. (2) Confident pseudo labeling strategy aims at generating more reliable supervision for intra and inter domain branches, and thus facilitating the learning process of the whole framework. Extensive experiments show that IS2Net yields consistent improvements over the state-of- the-art methods in three public benchmarks. Shiao Xie, Ziwei Niu, Huimin Huang 0002, Hao Sun 0013, Yen-Wei Chen 0001, Lanfen Lin |
ACM Multimedia | 6 |
| 2023 | TensorFormer: A Tensor-Based Multimodal Transformer for Multimodal Sentiment Analysis and Depression DetectionabstractSentiment analysis is an important research field aiming to extract and fuse sentimental information from human utterances. Due to the diversity of human sentiment, analyzing from multiple modalities is usually more accurate than from a single modality. To complement the information between related modalities, one effective approach is performing cross-modality interactions. Recently, Transformer-based frameworks have shown a strong ability to capture long-range dependencies, leading to the introduction of several Transformer-based approaches for multimodal processing. However, due to the built-in attention mechanism of the Transformers, only two modalities can be engaged at once. As a result, the complementary information flow in these Transformer-based techniques is partial and constrained. To mitigate this, we propose, TensorFormer, a tensor-based multimodal Transformer framework that takes into account all relevant modalities for interactions. More precisely, we first construct a tensor utilizing the features extracted from each modality, assuming one modality is the target while the remaining tensors serve as the sources. We can generate the corresponding interacted features by calculating source-target attention. This strategy interacts with all involved modalities and generates complementing global information. Experiments on multimodal sentiment analysis benchmark datasets demonstrated the effectiveness of TensorFormer. In addition, we also evaluate TensorFormer in another related area: depression detection and the results reveal significant improvements when compared to other state-of-the-art methods. Hao Sun 0013, Yen-Wei Chen 0001, Lanfen Lin |
IEEE Trans. Affect. Comput. | 2 |
| 2023 | Adaptive Decomposition and Shared Weight Volumetric Transformer Blocks for Efficient Patch-Free 3D Medical Image SegmentationabstractHigh resolution (HR) 3D medical image segmentation is vital for an accurate diagnosis. However, in the field of medical imaging, it is still a challenging task to achieve a high segmentation performance with cost-effective and feasible computation resources. Previous methods commonly use patch-sampling to reduce the input size, but this inevitably harms the global context and decreases the model's performance. In recent years, a few patch-free strategies have been presented to deal with this issue, but either they have limited performance due to their over-simplified model structures or they follow a complicated training process. In this study, to effectively address these issues, we present Adaptive Decomposition (A-Decomp) and Shared Weight Volumetric Transformer Blocks (SW-VTB). A-Decomp can adaptively decompose features and reduce their spatial size, which greatly lowers GPU memory consumption. SW-VTB is able to capture long-range dependencies at a low cost with its lightweight design and cross-scale weight-sharing mechanism. Our proposed cross-scale weight-sharing approach enhances the network's ability to capture scale-invariant core semantic information in addition to reducing parameter numbers. By combining these two designs together, we present a novel patch-free segmentation framework named VolumeFormer. Experimental results on two datasets show that VolumeFormer outperforms existing patch-based and patch-free methods with a comparatively fast inference speed and relatively compact design. Hongyi Wang 0002, Qingqing Chen 0001, Ruofeng Tong 0001, Yen-Wei Chen 0001, Hongjie Hu, Lanfen Lin |
IEEE J. Biomed. Health Informatics | 5 |
| 2023 | Multi-Modal Tumor Segmentation With Deformable Aggregation and Uncertain Region InpaintingabstractMulti-modal tumor segmentation exploits complementary information from different modalities to help recognize tumor regions. Known multi-modal segmentation methods mainly have deficiencies in two aspects: First, the adopted multi-modal fusion strategies are built upon well-aligned input images, which are vulnerable to spatial misalignment between modalities (caused by respiratory motions, different scanning parameters, registration errors, etc). Second, the performance of known methods remains subject to the uncertainty of segmentation, which is particularly acute in tumor boundary regions. To tackle these issues, in this paper, we propose a novel multi-modal tumor segmentation method with deformable feature fusion and uncertain region refinement. Concretely, we introduce a deformable aggregation module, which integrates feature alignment and feature aggregation in an ensemble, to reduce inter-modality misalignment and make full use of cross-modal information. Moreover, we devise an uncertain region inpainting module to refine uncertain pixels using neighboring discriminative features. Experiments on two clinical multi-modal tumor datasets demonstrate that our method achieves promising tumor segmentation results and outperforms state-of-the-art methods. Yue Zhang 0042, Chengtao Peng, Ruofeng Tong 0001, Lanfen Lin, Yen-Wei Chen 0001, Qingqing Chen 0001, Hongjie Hu, Shaohua Kevin Zhou |
IEEE Trans. Medical Imaging | 5 |
| 2023 | IDH mutation status prediction by a radiomics associated modality attention network
Yutaro Iwamoto, Jingliang Cheng, Guohua Zhao, Xianhua Han, Yen-Wei Chen 0001 |
Vis. Comput. | 8 |
| 2022 | Mixed Transformer U-Net for Medical Image SegmentationabstractThough U-Net has achieved tremendous success in medical image segmentation tasks, it lacks the ability to explicitly model long-range dependencies. Therefore, Vision Transformers have emerged as alternative segmentation structures recently, for their innate ability of capturing long-range correlations through Self-Attention (SA). However, Transformers usually rely on large-scale pre-training and have high computational complexity. Furthermore, SA can only model self-affinities within a single sample, ignoring the potential correlations of the overall dataset. To address these problems, we propose a novel Transformer module named Mixed Transformer Module (MTM) for simultaneous inter- and intra- affinities learning. MTM first calculates self-affinities efficiently through our well-designed Local-Global Gaussian-Weighted Self-Attention (LGG-SA). Then, it mines inter-connections between data samples through External Attention (EA). By using MTM, we construct a U-shaped model named Mixed Transformer U-Net (MT-UNet) for accurate medical image segmentation. We test our method on two different public datasets, and the experimental results show that the proposed method achieves better performance over other state-of-the-art methods. The code is available at: https://github.com/Dootmaan/MT-UNet. Hongyi Wang 0002, Shiao Xie, Lanfen Lin, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Ruofeng Tong 0001 |
ICASSP | 6 |
| 2022 | Pixel-Level and Affinity-Level Knowledge Distillation for Unsupervised Segmentation of Covid-19 LesionsabstractAutomatic segmentation of COVID-19 lesions is essential for computer-aided diagnosis. However, this task remains challenging because widely-used supervised based methods require large-scale annotated data that is difficult to obtain. Although an unsupervised method based on anomaly detection has shown promising results in [1], its performance is relatively poor. We address this problem by proposing a pixel-level and affinity-level knowledge distillation method. It obtains a pre-trained teacher network with rich semantic knowledge of CT images by constructing and training an auto-encoder at first, and then trains a student network with the same architecture as the teacher by distilling the teacher’s knowledge only from normal CT images, and finally localizes COVID-19 lesions using the feature discrepancy between the teacher and the student networks. Besides, except for the traditional pixel-level distillation, we design the affinity-level distillation that takes into account the pairwise relationship of features to fully distill effective knowledge. We evaluate this method by using three different COVID-19 datasets and the experimental results show that the segmentation performance is largely improved when it is compared with the other existing unsupervised anomaly detection methods. Rui Xu 0002, Xinchen Ye, Yen-Wei Chen 0001, Fangyi Xu, Wenchao Zhu, Hongjie Hu, Xiaofeng Qu, Shoji Kido, Noriyuki Tomiyama |
ICASSP | 5 |
| 2022 | Unsupervised Generative Network for Blind Hyperspectral Image Super-ResolutionabstractHyperspectral (HS) imaging sacrifices spatial resolution to ensure a high spectral resolution when capturing the detailed spectral signature at each spatial location of the scene. To compensate for this deficiency, fusing low-resolution HS (LR-HS) images with high-resolution RGB (HR-RGB) images to obtain high-resolution HS (HR-HS) images has attracted remarkable attention. Recently, deep learning-based fusion methods in a fully-supervised manner have been proven to make great progress in hyperspectral image super-resolution (HSI-SR) tasks. However, these methods require collecting a large number of training samples and constructing a non-blind prediction model to super-resolve the observations captured under controlled imaging conditions. This study proposes a novel unsupervised generative network (UGN) for learning network parameters using the observed LR-HS, HR-RGB only without the corresponding ground-truth, and designs the spatial and spectral degradation blocks to automatically learn the image degradation operations for constructing an end-to-end blind HSI SR framework. To verify the effectiveness of our proposed method, we con-duct experiments on two benchmark HS image datasets and demonstrate superior performance compared with the super-vised and unsupervised blind/non-blind SoTA methods. Zhe Liu 0039, Xianhua Han, Jiande Sun 0001, Yen-Wei Chen 0001 |
ICIP | 4 |
| 2022 | ScaleFormer: Revisiting the Transformer-based Backbones from a Scale-wise Perspective for Medical Image SegmentationabstractRecently, a variety of vision transformers have been developed as their capability of modeling long-range dependency. In current transformer-based backbones for medical image segmentation, convolutional layers were replaced with pure transformers, or transformers were added to the deepest encoder to learn global context. However, there are mainly two challenges in a scale-wise perspective: (1) intra-scale problem: the existing methods lacked in extracting local-global cues in each scale, which may impact the signal propagation of small objects; (2) inter-scale problem: the existing methods failed to explore distinctive information from multiple scales, which may hinder the representation learning from objects with widely variable size, shape and location. To address these limitations, we propose a novel backbone, namely ScaleFormer, with two appealing designs: (1) A scale-wise intra-scale transformer is designed to couple the CNN-based local features with the transformer-based global cues in each scale, where the row-wise and column-wise global dependencies can be extracted by a lightweight Dual-Axis MSA. (2) A simple and effective spatial-aware inter-scale transformer is designed to interact among consensual regions in multiple scales, which can highlight the cross-scale dependency and resolve the complex scale variations. Experimental results on different benchmarks demonstrate that our Scale-Former outperforms the current state-of-the-art methods. The code is publicly available at: https://github.com/ZJUGiveLab/ScaleFormer. Huimin Huang 0002, Shiao Xie, Lanfen Lin, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Ruofeng Tong 0001 |
IJCAI | 6 |
| 2022 | An Accurate Unsupervised Liver Lesion Detection Method Using Pseudo-lesions
He Li 0042, Yutaro Iwamoto, Xianhua Han, Lanfen Lin, Hongjie Hu, Yen-Wei Chen 0001 |
MICCAI (8) | 6 |
| 2022 | Local-Region and Cross-Dataset Contrastive Learning for Retinal Vessel Segmentation
Rui Xu 0002, Xinchen Ye, Zhihui Wang 0001, Yen-Wei Chen 0001 |
MICCAI (2) | 7 |
| 2022 | CubeMLP: An MLP-based Model for Multimodal Sentiment Analysis and Depression EstimationabstractMultimodal sentiment analysis and depression estimation are two important research topics that aim to predict human mental states using multimodal data. Previous research has focused on developing effective fusion strategies for exchanging and integrating mind-related information from different modalities. Some MLP-based techniques have recently achieved considerable success in a variety of computer vision tasks. Inspired by this, we explore multimodal approaches with a feature-mixing perspective in this study. To this end, we introduce CubeMLP, a multimodal feature processing framework based entirely on MLP. CubeMLP consists of three independent MLP units, each of which has two affine transformations. CubeMLP accepts all relevant modality features as input and mixes them across three axes. After extracting the characteristics using CubeMLP, the mixed multimodal features are flattened for task predictions. Our experiments are conducted on sentiment analysis datasets: CMU-MOSI and CMU-MOSEI, and depression estimation dataset: AVEC2019. The results show that CubeMLP can achieve state-of-the-art performance with a much lower computing cost. Hao Sun 0013, Hongyi Wang 0002, Jiaqing Liu, Yen-Wei Chen 0001, Lanfen Lin |
ACM Multimedia | 4 |
| 2022 | A multi-head pseudo nodes based spatial-temporal graph convolutional network for emotion perception from GAIT
Shurong Chai, Jiaqing Liu, Rahul Kumar Jain 0001, Tomoko Tateyama, Yutaro Iwamoto, Lanfen Lin, Yen-Wei Chen 0001 |
Neurocomputing | 7 |
| 2022 | Attention-based cross-layer domain alignment for unsupervised domain adaptation
Junkun Yuan, Yen-Wei Chen 0001, Ruofeng Tong 0001, Lanfen Lin |
Neurocomputing | 3 |
| 2022 | Mutual Information-Based Graph Co-Attention Networks for Multimodal Prior-Guided Magnetic Resonance Imaging SegmentationabstractMultimodal magnetic resonance imaging (MRI) provides complementary information about targets, and the segmentation of multimodal MRI is widely used as an essential preprocessing step for initial diagnosis, stage differentiation, and post-treatment efficacy evaluation in clinical situations. For the main modality or each of the modalities, it is important to enhance the visual information by modeling the connection and effectively fusing the features among them. However, the existing methods for multimodal segmentation have a drawback; they coincidentally drop information of individual modality during the fusion process. Recently, graph learning-based methods have been applied in segmentation, and these methods have achieved considerable improvements by modeling the relationships across feature regions and reasoning using global information. In this paper, we propose a graph learning-based approach to efficiently extract modality-specific features and establish regional correspondence effectively among all modalities. In detail, after projecting features into a graph domain and employing graph convolution to propagate information across all regions for learning global modality-specific features, we propose a mutual information-based graph co-attention module to learn the weight coefficients of one bipartite graph constructed by the fully connected graphs having different modalities in the graph domain and by selectively fusing the node features. Based on the deformation diagram between the spatial-graph space and our proposed graph co-attention module, we present a multimodal prior-guided segmentation framework, which uses two strategies for two clinical situations:Modality-Specific Learning StrategyandCo-Modality Learning Strategy. Besides, the improvedCo-Modality Learning Strategyis used with trainable weights in the multi-task loss for the optimization of the proposed framework. We validated our proposed modules and frameworks on two multimodal MRI datasets: our private liver lesion dataset and a public prostate zone dataset. Our experimental results on both datasets prove the superiority of our proposed approaches. Shaocong Mo, Lanfen Lin, Ruofeng Tong 0001, Qingqing Chen 0001, Fang Wang 0030, Hongjie Hu, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 10 |
| 2022 | A Graph Convolutional Multiple Instance Learning on a Hypersphere Manifold Approach for Diagnosing Chronic Obstructive Pulmonary Disease in CT ImagesabstractChronic obstructive pulmonary disease (COPD) is a prevalent chronic disease with high morbidity and mortality. The early diagnosis of COPD is vital for clinical treatment, which helps patients to have a better quality of life. Because COPD can be ascribed to chronic bronchitis and emphysema, lesions in a computed tomography (CT) image can present anywhere inside the lung with different types, shapes and sizes. Multiple instance learning (MIL) is an effective tool for solving COPD discrimination. In this study, a novel graph convolutional MIL with the adaptive additive margin loss (GCMIL-AAMS) approach is proposed to diagnose COPD by CT. Specifically, for those early stage patients, the selected instance-level features can be more discriminative if they were learned by our proposed graph convolution and pooling with self-attention mechanism. The AAMS loss can utilize the information of COPD severity on a hypersphere manifold by adaptively setting the angular margins to improve the performance, as the severity can be quantified as four grades by pulmonary function test. The results show that our proposed GCMIL-AAMS method provides superior discrimination and generalization abilities in COPD discrimination, with areas under a receiver operating characteristic curve (AUCs) of 0.960 ± 0.014 and 0.862 ± 0.010 in the test set and external testing set, respectively, in 5-fold stratified cross validation; moreover, it demonstrates that graph learning is applicable to MIL and suggests that MIL may be adaptable to graph learning. Qixing Feng, Xi Yin 0009, Xiangde Min, Defu Yang, Yen-Wei Chen 0001, Daoqiang Zhang, Wentao Zhu 0002 |
IEEE J. Biomed. Health Informatics | 7 |
| 2022 | MTL-ABS3Net: Atlas-Based Semi-Supervised Organ Segmentation Network With Multi-Task Learning for Medical ImagesabstractOrgan segmentation is one of the most important step for various medical image analysis tasks. Recently, semi-supervised learning (SSL) has attracted much attentions by reducing labeling cost. However, most of the existing SSLs neglected the prior shape and position information specialized in the medical images, leading to unsatisfactory localization and non-smooth of objects. In this paper, we propose a novel atlas-based semi-supervised segmentation network with multi-task learning for medical organs, named MTL-ABS3Net, which incorporates the anatomical priors and makes full use of unlabeled data in a self-training and multi-task learning manner. The MTL-ABS3Net consists of two components: an Atlas-Based Semi-Supervised Segmentation Network (ABS3Net) and Reconstruction-Assisted Module (RAM). Specifically, the ABS3Net improves the existing SSLs by utilizing atlas prior, which generates credible pseudo labels in a self-training manner; while the RAM further assists the segmentation network by capturing the anatomical structures from the original images in a multi-task learning manner. Better reconstruction quality is achieved by using MS-SSIM loss function, which further improves the segmentation accuracy. Experimental results from the liver and spleen datasets demonstrated that the performance of our method was significantly improved compared to existing state-of-the-art methods. Huimin Huang 0002, Qingqing Chen 0001, Lanfen Lin, Qiaowei Zhang, Yutaro Iwamoto, Xianhua Han, Akira Furukawa, Shuzo Kanasaki, Yen-Wei Chen 0001, Ruofeng Tong 0001, Hongjie Hu |
IEEE J. Biomed. Health Informatics | 10 |
| 2022 | DeepRecS: From RECIST Diameters to Precise Liver Tumor SegmentationabstractLiver tumor segmentation (LiTS) is of primary importance in diagnosis and treatment of hepatocellular carcinoma. Known automated LiTS methods could not yield satisfactory results for clinical use since they were hard to model flexible tumor shapes and locations. In clinical practice, radiologists usually estimate tumor shape and size by a Response Evaluation Criteria in Solid Tumor (RECIST) mark. Inspired by this, in this paper, we explore a deep learning (DL) based interactive LiTS method, which incorporates guidance from user-provided RECIST marks. Our method takes a three-step framework to predict liver tumor boundaries. Under this architecture, we develop a RECIST mark propagation network (RMP-Net) to estimate RECIST-like marks in off-RECIST slices. We also devise a context-guided boundary-sensitive network (CGBS-Net) to distill tumors' contextual and boundary information from corresponding RECIST(-like) marks, and then predict tumor maps. To further refine the segmentation results, we process the tumor maps using a 3D conditional random field (CRF) algorithm and a morphology hole-filling operation. Verified on two clinical contrast-enhanced abdomen computed tomography (CT) image datasets, our proposed approach can produce promising segmentation results, and outperforms the state-of-the-art interactive segmentation methods. Yue Zhang 0042, Chengtao Peng, Liying Peng, Lanfen Lin, Ruofeng Tong 0001, Zhiyi Peng, Xiongwei Mao, Hongjie Hu, Yen-Wei Chen 0001, Jingsong Li 0001 |
IEEE J. Biomed. Health Informatics | 10 |
| 2022 | Hyperspectral Image Reconstruction Using Multi-scale Fusion LearningabstractHyperspectral imaging is a promising imaging modality that simultaneously captures several images for the same scene on narrow spectral bands, and it has made considerable progress in different fields, such as agriculture, astronomy, and surveillance. However, the existing hyperspectral (HS) cameras sacrifice the spatial resolution for providing the detail spectral distribution of the imaged scene, which leads to low-resolution (LR) HS images compared with the common red-green-blue (RGB) images. Generating a high-resolution HS (HR-HS) image via fusing an observed LR-HS image with the corresponding HR-RGB image has been actively studied. Existing methods for this fusing task generally investigate hand-crafted priors to model the inherent structure of the latent HR-HS image, and they employ optimization approaches for solving it. However, proper priors for different scenes can possibly be diverse, and to figure it out for a specific scene is difficult. This study investigates a deep convolutional neural network (DCNN)-based method for automatic prior learning, and it proposes a novel fusion DCNN model with multi-scale spatial and spectral learning for effectively merging an HR-RGB and LR-HS images. Specifically, we construct an U-shape network architecture for gradually reducing the feature sizes of the HR-RGB image (Encoder-side) and increasing the feature sizes of the LR-HS image (Decoder-side), and we fuse the HR spatial structure and the detail spectral attribute in multiple scales for tackling the large resolution difference in spatial domain of the observed HR-RGB and LR-HS images. Then, we employ multi-level cost functions for the proposed multi-scale learning network to alleviate the gradient vanish problem in long-propagation procedure. In addition, for further improving the reconstruction performance of the HR-HS image, we refine the predicted HR-HS image using an alternating back-projection method for minimizing the reconstruction errors of the observed LR-HS and HR-RGB images. Experiments on three benchmark HS image datasets demonstrate the superiority of the proposed method in both quantitative values and visual qualities. Xianhua Han, Yinqiang Zheng, Yen-Wei Chen 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2021 | Graph-Based Pyramid Global Context Reasoning With a Saliency- Aware Projection for Covid-19 Lung Infections SegmentationabstractCoronavirus Disease 2019 (COVID-19) has rapidly spread in 2020, emerging a mass of studies for lung infection segmentation from CT images. Though many methods have been proposed for this issue, it is a challenging task because of infections of various size appearing in different lobe zones. To tackle these issues, we propose a Graph-based Pyramid Global Context Reasoning (Graph-PGCR) module, which is capable of modeling long-range dependencies among disjoint infections as well as adapt size variation. We first incorporate graph convolution to exploit long-term contextual information from multiple lobe zones. Different from previous average pooling or maximum object probability, we propose a saliency-aware projection mechanism to pick up infection-related pixels as a set of graph nodes. After graph reasoning, the relation-aware features are reversed back to the original coordinate space for the down-stream tasks. We further construct multiple graphs with different sampling rates to handle the size variation problem. To this end, distinct multi-scale long-range contextual patterns can be captured. Our Graph- PGCR module is plug-and-play, which can be integrated into any architecture to improve its performance. Experiments demonstrated that the proposed method consistently boost the performance of state-of-the-art backbone architectures on both of public and our private COVID-19 datasets. Huimin Huang 0002, Lanfen Lin, Xiongwei Mao, Xiaohan Qian, Zhiyi Peng, Jianying Zhou 0006, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Ruofeng Tong 0001 |
ICASSP | 11 |
| 2021 | Graph-BAS3Net: Boundary-Aware Semi-Supervised Segmentation Network with Bilateral Graph ConvolutionabstractSemi-supervised learning (SSL) algorithms have attracted much attentions in medical image segmentation by leveraging unlabeled data, which challenge in acquiring massive pixel-wise annotated samples. However, most of the existing SSLs neglected the geometric shape constraint in object, leading to unsatisfactory boundary and non-smooth of object. In this paper, we propose a novel boundary-aware semi-supervised medical image segmentation network, named Graph-BAS3Net, which incorporates the boundary information and learns duality constraints between semantics and geometrics in the graph domain. Specifically, the proposed method consists of two components: a multi-task learning framework BAS3Net and a graph-based cross-task module BGCM. The BAS3Net improves the existing GAN-based SSL by adding a boundary detection task, which encodes richer features of object shape and surface. Moreover, the BGCM further explores the co-occurrence relations between the semantics segmentation and boundary detection task, so that the network learns stronger semantic and geometric correspondences from both labeled and unlabeled data. Experimental results on the LiTS dataset and COVID-19 dataset confirm that our proposed Graph-BAS3Net outperforms the state-of-the-art methods in semi-supervised segmentation task. Huimin Huang 0002, Lanfen Lin, Yue Zhang 0042, Xiongwei Mao, Xiaohan Qian, Zhiyi Peng, Jianying Zhou 0006, Yen-Wei Chen 0001, Ruofeng Tong 0001 |
ICCV | 10 |
| 2021 | A Teacher-Student Learning Based On Composed Ground-Truth Images For Accurate Cephalometric Landmark DetectionabstractComputer-aided automatic cephalometric landmark localization has been a hot topic since last century. Recent proposed deep learning-based methods have made great contributions to this research topic. Among them, convolutional neural networks (CNN)-based regression is widely used, where ground-truth (GT) information is mainly used in the calculation of loss function, thus, mimics the difference between the predicted landmarks ' locations and the ground-truth locations through backpropagation. However, considering the limited number of annotated cephalometric data, we believe the performance can be better improved by better utilizing ground-truth information. In this paper, we propose a teacher-student learning method using GT images for accurate cephalometric detection. We first use images composed with GT landmarks as input images to train a detection model, which is treated as a teacher model. Then the teacher model is used to guide a student model, which is trained by original images, by transferring useful features. We believe the features between GT images and original images have similar domain distribution since they both represent same structure. We validate our method on public grand challenge dataset. Our method achieves better performance compared with state-of-the-art methods. Yu Song 0008, Xu Qiao, Yutaro Iwamoto, Yen-Wei Chen 0001 |
ICIP | 4 |
| 2021 | 3D Graph-S2Net: Shape-Aware Self-ensembling Network for Semi-supervised Segmentation with Bilateral Graph Convolution
Huimin Huang 0002, Lanfen Lin, Hongjie Hu, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Ruofeng Tong 0001 |
MICCAI (2) | 7 |
| 2021 | Patch-Free 3D Medical Image Segmentation Driven by Super-Resolution Technique and Self-Supervised Guidance
Hongyi Wang 0002, Lanfen Lin, Hongjie Hu, Qingqing Chen 0001, Yinhao Li 0002, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Ruofeng Tong 0001 |
MICCAI (1) | 8 |
| 2021 | Multi-phase Liver Tumor Segmentation with Spatial Aggregation and Uncertain Region Inpainting
Yue Zhang 0042, Chengtao Peng, Liying Peng, Huimin Huang 0002, Ruofeng Tong 0001, Lanfen Lin, Jingsong Li 0001, Yen-Wei Chen 0001, Qingqing Chen 0001, Hongjie Hu, Zhiyi Peng |
MICCAI (1) | 8 |
| 2021 | A Tensor Sparse Representation-Based CBMIR System for Computer-Aided Diagnosis of Focal Liver Lesions and its Pilot TrialabstractClinicians refer to diagnosed medical cases in order to make correct diagnosis and take appropriate treatments, due to the complexity of focal liver lesions. It's a heavy burden, however, for medical doctors to find out similar and meaningful cases from the accumulated extreme large medical datasets. Content based medical image retrieval (CBMIR) that searches for similar images in a large database has been attracting increasing research interest recently. A CBMIR system provides doctors the diagnosed cases to improve the diagnosis accuracy and confidence. This paper proposed a tensor sparse representation method to extract temporal and spatial features of multi-phase CT images, so as to provide doctors medical cases more relevant to the query one. The proposed tensor sparse representation method is applied to the retrieval of focal liver lesions (FLLs). Experiments show that the proposed method achieved better retrieval performance than conventional methods. Pilot trial was conducted and results show that diagnosis accuracy and confidence was improved significantly by the developed CBMIR system based on the proposed method. Jian Wang 0004, Xianhua Han, Lanfen Lin, Hongjie Hu, Yen-Wei Chen 0001 |
ICMR | 5 |
| 2021 | M-DFNet: Multi-phase Discriminative Feature Network for Retrieval of Focal Liver LesionsabstractContent based medical image retrieval (CBMIR) plays a great role in computer aided diagnosis for assisting radiologists to detect and characterize focal liver lesions (FLLs). Deep learning has gained exciting performance on CBMIR. While the features generated by deep learning models trained using softmax loss are always separable but not discriminative enough, which is insufficient for retrieval task. In this paper, we propose a multi-phase discriminative feature network (M-DFNet) with a DeepExtracter and a feature refine module (FRModule) to learn discriminative and separable features under a joint supervision of center loss and softmax loss. The hybrid loss enables to minimize intra-class variations and enlarge inter-class differences as much as possible. The FRModule is proposed to recalibrate the deep features based on the learned class centers to tackle the complex imaging manifestations of FLLs and further enhance both the feature discrimination and generalization. Multi-phase computed tomography (CT) images contain pivotal information for diagnosis of FLLs. Thus the M-DFNet is designed to cope with multi-phase information and we explore an appropriate and effective method for multi-phase feature integration on limited data. Experimental results clearly demonstrate strong performance superiority by our proposed method. Jing Liu 0041, Lanfen Lin, Hongjie Hu, Ruofeng Tong 0001, Jingsong Li 0001, Yen-Wei Chen 0001 |
ICMR | 7 |
| 2021 | Accurate and fast mitotic detection using an anchor-free method based on full-scale connection with recurrent deep layer aggregation in 4D microscopy imagesabstractBACKGROUND: To effectively detect and investigate various cell-related diseases, it is essential to understand cell behaviour. The ability to detection mitotic cells is a fundamental step in diagnosing cell-related diseases. Convolutional neural networks (CNNs) have been successfully applied to object detection tasks, however, when applied to mitotic cell detection, most existing methods generate high false-positive rates due to the complex characteristics that differentiate normal cells from mitotic cells. Cell size and orientation variations in each stage make detecting mitotic cells difficult in 2D approaches. Therefore, effective extraction of the spatial and temporal features from mitotic data is an important and challenging task. The computational time required for detection is another major concern for mitotic detection in 4D microscopic images. RESULTS: In this paper, we propose a backbone feature extraction network named full scale connected recurrent deep layer aggregation (RDLA++) for anchor-free mitotic detection. We utilize a 2.5D method that includes 3D spatial information extracted from several 2D images from neighbouring slices that form a multi-stream input. CONCLUSIONS: Our proposed technique addresses the scale variation problem and can efficiently extract spatial and temporal features from 4D microscopic images, resulting in improved detection accuracy and reduced computation time compared with those of other state-of-the-art methods. Titinunt Kitrungrotsakul, Yutaro Iwamoto, Satoko Takemoto, Hideo Yokota, Sari Ipponjima, Tomomi Nemoto, Lanfen Lin, Ruofeng Tong 0001, Jingsong Li 0001, Yen-Wei Chen 0001 |
BMC Bioinform. | 10 |
| 2021 | A Cascade of 2.5D CNN and Bidirectional CLSTM Network for Mitotic Cell Detection in 4D Microscopy ImageabstractMitosis detection is one of the challenging steps in biomedical imaging research, which can be used to observe the cell behavior. Most of the already existing methods that are applied in detecting mitosis usually contain many nonmitotic events (normal cell and background) in the result (false positives, FPs). In order to address such a problem, in this study, we propose to apply 2.5-dimensional (2.5D) networks called CasDetNet_CLSTM, which can accurately detect mitotic events in 4D microscopic images. This CasDetNet_CLSTM involves a 2.5D faster region-based convolutional neural network (Faster R-CNN) as the first network, and a convolutional long short-term memory (CLSTM) network as the second network. The first network is used to select candidate cells using the information from nearby slices, whereas the second network uses temporal information to eliminate FPs and refine the result of the first network. Our experiment shows that the precision and recall of our networks yield better results than those of other state-of-the-art methods. Titinunt Kitrungrotsakul, Xianhua Han, Yutaro Iwamoto, Satoko Takemoto, Hideo Yokota, Sari Ipponjima, Tomomi Nemoto, Wei Xiong 0001, Yen-Wei Chen 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 9 |
| 2021 | VolumeNet: A Lightweight Parallel Network for Super-Resolution of MR and CT Volumetric DataabstractDeep learning-based super-resolution (SR) techniques have generally achieved excellent performance in the computer vision field. Recently, it has been proven that three-dimensional (3D) SR for medical volumetric data delivers better visual results than conventional two-dimensional (2D) processing. However, deepening and widening 3D networks increases training difficulty significantly due to the large number of parameters and small number of training samples. Thus, we propose a 3D convolutional neural network (CNN) for SR of magnetic resonance (MR) and computer tomography (CT) volumetric data called ParallelNet using parallel connections. We construct a parallel connection structure based on the group convolution and feature aggregation to build a 3D CNN that is as wide as possible with a few parameters. As a result, the model thoroughly learns more feature maps with larger receptive fields. In addition, to further improve accuracy, we present an efficient version of ParallelNet (called VolumeNet), which reduces the number of parameters and deepens ParallelNet using a proposed lightweight building block module called the Queue module. Unlike most lightweight CNNs based on depthwise convolutions, the Queue module is primarily constructed using separable 2D cross-channel convolutions. As a result, the number of network parameters and computational complexity can be reduced significantly while maintaining accuracy due to full channel fusion. Experimental results demonstrate that the proposed VolumeNet significantly reduces the number of model parameters and achieves high precision results compared to state-of-the-art methods in tasks of brain MR image SR, abdomen CT image SR, and reconstruction of super-resolution 7T-like images from their 3T counterparts. Yinhao Li 0002, Yutaro Iwamoto, Lanfen Lin, Rui Xu 0002, Ruofeng Tong 0001, Yen-Wei Chen 0001 |
IEEE Trans. Image Process. | 6 |
| 2021 | Attention-RefNet: Interactive Attention Refinement Network for Infected Area Segmentation of COVID-19abstractCOVID-19 pneumonia is a disease that causes an existential health crisis in many people by directly affecting and damaging lung cells. The segmentation of infected areas from computed tomography (CT) images can be used to assist and provide useful information for COVID-19 diagnosis. Although several deep learning-based segmentation methods have been proposed for COVID-19 segmentation and have achieved state-of-the-art results, the segmentation accuracy is still not high enough (approximately 85%) due to the variations of COVID-19 infected areas (such as shape and size variations) and the similarities between COVID-19 and non-COVID-infected areas. To improve the segmentation accuracy of COVID-19 infected areas, we propose an interactive attention refinement network (Attention RefNet). The interactive attention refinement network can be connected with any segmentation network and trained with the segmentation network in an end-to-end fashion. We propose a skip connection attention module to improve the important features in both segmentation and refinement networks and a seed point module to enhance the important seeds (positions) for interactive refinement. The effectiveness of the proposed method was demonstrated on public datasets (COVID-19CTSeg and MICCAI) and our private multicenter dataset. The segmentation accuracy was improved to more than 90%. We also confirmed the generalizability of the proposed network on our multicenter dataset. The proposed method can still achieve high segmentation accuracy. Titinunt Kitrungrotsakul, Qingqing Chen 0001, Huitao Wu, Yutaro Iwamoto, Hongjie Hu, Wenchao Zhu, Fangyi Xu, Lanfen Lin, Ruofeng Tong 0001, Jingsong Li 0001, Yen-Wei Chen 0001 |
IEEE J. Biomed. Health Informatics | 13 |
| 2021 | Integration of CNN, CBMIR, and Visualization Techniques for Diagnosis and Quantification of Covid-19 DiseaseabstractDiagnosis techniques based on medical image modalities have higher sensitivities compared to conventional RT-PCT tests. We propose two methods for diagnosing COVID-19 disease using X-ray images and differentiating it from viral pneumonia. The diagnosis section is based on deep neural networks, and the discriminating uses an image retrieval approach. Both units were trained by healthy, pneumonia, and COVID-19 images. In COVID-19 patients, the maximum intensity projection of the lung CT is visualized to a physician, and the CT Involvement Score is calculated. The performance of the CNN and image retrieval algorithms were improved by transfer learning and hashing functions. We achieved an accuracy of 97% and an overall prec@10 of 87%, respectively, concerning the CNN and the retrieval methods. Saeed Mohagheghi, Mehdi Alizadeh, Seyed Mahdi Safavi, Amir Hossein Foruzan, Yen-Wei Chen 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2021 | Joint Extraction of Retinal Vessels and Centerlines Based on Deep Semantics and Multi-Scaled Cross-Task AggregationabstractRetinal vessel segmentation and centerline extraction are crucial steps in building a computer-aided diagnosis system on retinal images. Previous works treat them as two isolated tasks, while ignoring their tight association. In this paper, we propose a deep semantics and multi-scaled cross-task aggregation network that takes advantage of the association to jointly improve their performances. Our network is featured by two sub-networks. The forepart is a deep semantics aggregation sub-network that aggregates strong semantic information to produce more powerful features for both tasks, and the tail is a multi-scaled cross-task aggregation sub-network that explores complementary information to refine the results. We evaluate the proposed method on three public databases, which are DRIVE, STARE and CHASE_DB1. Experimental results show that our method can not only simultaneously extract retinal vessels and their centerlines but also achieve the state-of-the-art performances on both tasks. Rui Xu 0002, Xinchen Ye, Lin Lin 0008, Liang Li 0002, Yen-Wei Chen 0001 |
IEEE J. Biomed. Health Informatics | 8 |
| 2021 | Medical Image Segmentation With Deep Atlas PriorabstractOrgan segmentation from medical images is one of the most important pre-processing steps in computer-aided diagnosis, but it is a challenging task because of limited annotated data, low-contrast and non-homogenous textures. Compared with natural images, organs in the medical images have obvious anatomical prior knowledge (e.g., organ shape and position), which can be used to improve the segmentation accuracy. In this paper, we propose a novel segmentation framework which integrates the medical image anatomical prior through loss into the deep learning models. The proposed prior loss function is based on probabilistic atlas, which is called as deep atlas prior (DAP). It includes prior location and shape information of organs, which are important prior information for accurate organ segmentation. Further, we combine the proposed deep atlas prior loss with the conventional likelihood losses such as Dice loss and focal loss into an adaptive Bayesian loss in a Bayesian framework, which consists of a prior and a likelihood. The adaptive Bayesian loss dynamically adjusts the ratio of the DAP loss and the likelihood loss in the training epoch for better learning. The proposed loss function is universal and can be combined with a wide variety of existing deep segmentation models to further enhance their performance. We verify the significance of our proposed framework with some state-of-the-art models, including fully-supervised and semi-supervised segmentation models on a public dataset (ISBI LiTS 2017 Challenge) for liver segmentation and a private dataset for spleen segmentation. Huimin Huang 0002, Lanfen Lin, Hongjie Hu, Qiaowei Zhang, Qingqing Chen 0001, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Ruofeng Tong 0001 |
IEEE Trans. Medical Imaging | 10 |
| 2020 | UNet 3+: A Full-Scale Connected UNet for Medical Image SegmentationabstractRecently, a growing interest has been seen in deep learning-based semantic segmentation. UNet, which is one of deep learning networks with an encoder-decoder architecture, is widely used in medical image segmentation. Combining multi-scale features is one of important factors for accurate segmentation. UNet++ was developed as a modified Unet by designing an architecture with nested and dense skip connections. However, it does not explore sufficient information from full scales and there is still a large room for improvement. In this paper, we propose a novel UNet 3+, which takes advantage of full-scale skip connections and deep supervisions. The full-scale skip connections incorporate low-level details with high-level semantics from feature maps in different scales; while the deep supervision learns hierarchical representations from the full-scale aggregated feature maps. The proposed method is especially benefiting for organs that appear at varying scales. In addition to accuracy improvements, the proposed UNet 3+ can reduce the network parameters to improve the computation efficiency. We further propose a hybrid loss function and devise a classification-guided module to enhance the organ boundary and reduce the over-segmentation in a non-organ image, yielding more accurate segmentation results. The effectiveness of the proposed method is demonstrated on two datasets. The code is available at: github.com/ZJUGiveLab/UNet-Version. Huimin Huang 0002, Lanfen Lin, Ruofeng Tong 0001, Hongjie Hu, Qiaowei Zhang, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Jian Wu 0001 |
ICASSP | 8 |
| 2020 | Unsupervised Detection of Pulmonary Opacities for Computer-Aided Diagnosis of COVID-19 on CT ImagesabstractCOVID-19 emerged towards the end of 2019 which was identified as a global pandemic by the world heath organization (WHO). With the rapid spread of COVID-19, the number of infected and suspected patients has increased dramatically. Chest computed tomography (CT) has been recognized as an efficient tool for the diagnosis of COVID-19. However, the huge CT data make it difficult for radiologist to fully exploit them on the diagnosis. In this paper, we propose a computer-aided diagnosis system that can automatically analyze CT images to distinguish the COVID-19 against to community-acquired pneumonia (CAP). The proposed system is based on an unsupervised pulmonary opacity detection method that locates opacity regions by a detector unsupervisedly trained from CT images with normal lung tissues. Radiomics based features are extracted insides the opacity regions, and fed into classifiers for classification. We evaluate the proposed CAD system by using 200 CT images collected from different patients in several hospitals. The accuracy, precision, recall, f1-score and AUC achieved are 95.5%, 100%, 91%, 95.1% and 95.9% respectively, exhibiting the promising capacity on the differential diagnosis of COVID-19 from CT images. Rui Xu 0002, Xiao Cao, Yen-Wei Chen 0001, Xinchen Ye, Lin Lin 0008, Wenchao Zhu, Fangyi Xu, Hongjie Hu, Shoji Kido, Noriyuki Tomiyama |
ICPR | 4 |
| 2020 | BG-Net: Boundary-Guided Network for Lung Segmentation on Clinical CT ImagesabstractLung segmentation on CT images is a crucial step for a computer-aided diagnosis system of lung diseases. The existing deep learning based lung segmentation methods are less efficient to segment lungs on clinical CT images, especially that the segmentation on lung boundaries is not accurate enough due to complex pulmonary opacities in practical clinics. In this paper, we propose a boundary-guided network (BG-Net) to address this problem. It contains two auxiliary branches that seperately segment lungs and extract the lung boundaries, and an aggregation branch that efficiently exploits lung boundary cues to guide the network for more accurate lung segmentation on clinical CT images. We evaluate the proposed method on a private dataset collected from the Osaka university hospital and four public datasets including StructSeg [1], HUG [2], VESSEL12 [3], and a Novel Coronavirus 2019 (COVID-19) dataset [4]. Experimental results show that the proposed method can segment lungs more accurately and outperform several other deep learning based methods. Rui Xu 0002, Yi Wang 0037, Xinchen Ye, Lin Lin 0008, Yen-Wei Chen 0001, Shoji Kido, Noriyuki Tomiyama |
ICPR | 6 |
| 2020 | Multimodal Priors Guided Segmentation of Liver Lesions in MRI Using Mutual Information Based Graph Co-Attention Networks
Shaocong Mo, Lanfen Lin, Ruofeng Tong 0001, Qingqing Chen 0001, Fang Wang 0030, Hongjie Hu, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001 |
MICCAI (4) | 10 |
| 2020 | Boosting Connectivity in Retinal Vessel Segmentation via a Recursive Semantics-Guided Network
Rui Xu 0002, Xinchen Ye, Lin Lin 0008, Yen-Wei Chen 0001 |
MICCAI (5) | 5 |
| 2020 | Novel image restoration method based on multi-frame super-resolution for atmospherically distorted imagesabstractIn this study, the authors propose a novel multi‐frame super‐resolution method using frame selection and multiple fusions for atmospherically distorted, zoomed‐in, image‐quality enhancement. When a small part of the image captured by placing a target several kilometres away from the fixed camera is enlarged, the quality of the part becomes poor owing to low resolution, spatial deformations and noise that are mainly caused by long distance and atmospheric turbulence. Thus, the authors propose an adaptive frame selection method that selects only a few frames with small blur based on the corresponding images with relatively clear edges. Further, they propose multiple fusion schemes to reconstruct the selected frames, thereby suppressing the influence of deformation. By converting all the frames into high‐resolution based on each frame and integrating them, deformation and noise are effectively removed without high computation cost using the multiple fusion scheme. The proposed method, which enhances the quality of atmospherically distorted zoomed‐in images, exhibits superior performance than the state‐of‐the‐art image super‐resolution methods with regard to high accuracy, efficiency and ease of implementation, ensuring that the proposed method is suitable for enhancing the quality of an image captured using a general digital camera or a smartphone. Yinhao Li 0002, Katsuhisa Ogawa, Yutaro Iwamoto, Yen-Wei Chen 0001 |
IET Image Process. | 4 |
| 2020 | An end-to-end CNN and LSTM network with 3D anchors for mitotic cell detection in 4D microscopic images and its parallel implementation on multiple GPUs
Titinunt Kitrungrotsakul, Xianhua Han, Yutaro Iwamoto, Satoko Takemoto, Hideo Yokota, Sari Ipponjima, Tomomi Nemoto, Wei Xiong 0001, Yen-Wei Chen 0001 |
Neural Comput. Appl. | 9 |
| 2020 | Tensor-based sparse representations of multi-phase medical images for classification of focal liver lesions
Jian Wang 0004, Jing Li 0046, Xianhua Han, Lanfen Lin, Hongjie Hu, Qingqing Chen 0001, Yutaro Iwamoto, Yen-Wei Chen 0001 |
Pattern Recognit. Lett. | 9 |
| 2020 | Semi-Supervised Learning for Semantic Segmentation of Emphysema With Partial AnnotationsabstractSegmentation and quantification of each subtype of emphysema is helpful to monitor chronic obstructive pulmonary disease. Due to the nature of emphysema (diffuse pulmonary disease), it is very difficult for experts to allocate semantic labels to every pixel in the CT images. In practice, partially annotating is a better choice for the radiologists to reduce their workloads. In this paper, we propose a new end-to-end trainable semi-supervised framework for semantic segmentation of emphysema with partial annotations, in which a segmentation network is trained from both annotated and unannotated areas. In addition, we present a new loss function, referred to as Fisher loss, to enhance the discriminative power of the model and successfully integrate it into our proposed framework. Our experimental results show that the proposed methods have superior performance over the baseline supervised approach (trained with only annotated areas) and outperform the state-of-the-art methods for emphysema segmentation. Liying Peng, Lanfen Lin, Hongjie Hu, Yue Zhang 0042, Huali Li, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001 |
IEEE J. Biomed. Health Informatics | 8 |
| 2020 | Hyperspectral Reconstruction with Redundant Camera Spectral Sensitivity FunctionsabstractHigh-resolution hyperspectral (HS) reconstruction has recently achieved significantly progress, among which the method based on the fusion of the RGB and HS images of the same scene can greatly improve the reconstruction performance compared with those based on the individually spectral or spatial enhancement. It is well known that the HS image is obtained only via the costly hypersoectral sensor, whereas the RGB images can be provided by low-price RGB cameras and the spectral sensitivity (SS) functions of RGB cameras are usually different. Thus, this study proposes a HS reconstruction, which fuses merely two RGB images with redundant spectral responses. In this work, we design a new RGB camera via shifting the SS of an existed RGB camera, which can provide similar strength of spectral response with different spectral centers of SS, and fuse the new achieved color image with an existed RGB image by a deep ResNet. Experiments validate that fusion of two existed RGB images can provide impressive HS reconstruction performance and further improvement can be achieved by integrating the color image of the simulated SS with the RGB image. Xianhua Han, Yinqiang Zheng, Jiande Sun 0001, Yen-Wei Chen 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2019 | A Cascade of CNN and LSTM Network with 3D Anchors for Mitotic Cell Detection in 4D Microscopic ImageabstractMitotic event detection is a fundamental step in investigating of cell behaviors. The event can be used to analyze various diseases, but most mitotic event detections performed previously focused only on two-dimensional (2D) images with time information. Owing to the complex background (normal cells) and mitotic event orientations, the 2D detection methods yield many false positive and false negative results. To solve this problem, we proposed a 2.5 dimensional (2.5D) cascaded end-to-end network combined with 3D anchors for accurate detection of mitotic events in 4D microscopic images. Our proposed network uses a convolutional long short-term memory to handle issues relating to time sequence; this helps to improve the detection accuracy (reduction of false positives). Furthermore, it uses 3D anchors to capture volume information used to address the orientation problem (reduction of false negatives). The experimental results show that the proposed method can achieve higher precision and recall compared with state-of-the-art methods. Titinunt Kitrungrotsakul, Yutaro Iwamoto, Xianhua Han, Satoko Takemoto, Hideo Yokota, Sari Ipponjima, Tomomi Nemoto, Wei Xiong 0001, Yen-Wei Chen 0001 |
ICASSP | 9 |
| 2019 | A Dual-Attention Dilated Residual Network for Liver Lesion Classification and Localization on CT ImagesabstractAutomatic liver lesion classification on computed tomography images is of great importance to early cancer diagnosis and remains a challenging task. State-of-the-art liver lesion classification algorithms are currently based on manually selected regions of interest (ROIs) or automatically detected ROIs. However, liver lesions usually vary in size and shape, which makes the ROI selection process labor-intensive and also poses an obstacle to automatic lesion detection. In this paper, we propose a dual-attention dilated residual network (DADRN) as a potential solution to lesion classification task without manual ROI selection or automatic lesion detection. We incorporated a novel dual-attention module in order to capture the non-local feature dependencies and help the deep neural network focus on the lesion area by enlarging the difference between the lesion area and nonlesion area. To the best of our knowledge, we are the first to employ the self-attention mechanism to address liver lesion classification task. In addition, the well-trained DADRN can be used for weakly-supervised lesion localization without any architectural change or retraining. Experiment results show that DADRN could achieve a lesion classification accuracy comparable to that of the state-of-the-art ROI-based method and outperformed state-of-the-art attention-based approaches in both liver lesion classification and localization tasks. Xiao Chen 0016, Jian Wu 0001, Lanfen Lin, Hongjie Hu, Qiaowei Zhang, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Ruofeng Tong 0001 |
ICIP | 9 |
| 2019 | Multi-Stream Scale-Insensitive Convolutional and Recurrent Neural Networks for Liver Tumor Detection in Dynamic Ct ImagesabstractConvolutional neural networks (CNNs) have achieved great success in numerous challenging vision tasks, and have great potential for object detection in natural images. Compared with the natural images, medical images exhibit some unique characteristics. Therefore, substantial challenges still remain in this field. The first challenge is to develop a method for effectively distilling enhancement patterns from the dynamic CT images. Moreover, since tumor sizes vary greatly and small lesions are important for early liver tumor detection, lesion detection with a widely variable scale is another challenge. In this paper, we propose a multi-stream scale-insensitive convolutional and recurrent neural network (MSCR) for liver tumor detection. Specifically, we propose the use of grouped convolutional long short-term memory (GCLSTM) to extract enhancement patterns, which is developed as a plug-and-play module. Experiments show that the MSCR framework exhibits superior performance over state-of-the-art approaches, achieving an average precision of 77.06% for detection of focal liver lesions. We have released the code of MSCR in1. Ruofeng Tong 0001, Jian Wu 0001, Lanfen Lin, Xiao Chen 0016, Hongjie Hu, Qiaowei Zhang, Qingqing Chen 0001, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001 |
ICIP | 11 |
| 2019 | An Improved Hand Gesture Recognition with Two-Stage Convolution Neural Networks Using a Hand Color Image and its Pseudo-Depth ImageabstractRobust hand gesture recognition has been playing a significant role in the field of human-computer interaction for a long time, but it is still full of challenges due to many accept such as cluttered backgrounds and hand self-occlusion. With the help of depth information, depth-based methods have better performance, but the depth cameras are not as widely used and affordable as color cameras. Therefore, in this paper, we propose a two-stage deep convolutional neural network (CNN) architecture for accurate color-based hand gesture recognition. The first stage performs generation of pseudo-depth hand images from color images and the second stage recognizes hand gesture classes using both the color image and its pseudo-depth hand image. The generation stage architecture is based on an image-to-image translation network. In the recognition stage, a two-stream CNN architecture with color image and its pseudo depth image is proposed to improve the color image-based recognition performance. We also propose two strategies in two-stream fusion: feature fusion and committee fusion. To validate our approach, we construct a new dataset called MaHG-RGBD dataset. Experiments demonstrate that our approach significantly improves the performance in RGB-only recognition for hand gestures. Jiaqing Liu, Kotaro Furusawa, Tomoko Tateyama, Yutaro Iwamoto, Yen-Wei Chen 0001 |
ICIP | 5 |
| 2019 | Semi-supervised Segmentation of Liver Using Adversarial Learning with Deep Atlas Prior
Lanfen Lin, Hongjie Hu, Qiaowei Zhang, Qingqing Chen 0001, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Ruofeng Tong 0001, Jian Wu 0001 |
MICCAI (6) | 8 |
| 2019 | Robust manifold broad learning system for large-scale noisy chaotic time series prediction: A perturbation perspective
Shoubo Feng, Min Han 0001, Yen-Wei Chen 0001 |
Neural Networks | 4 |
| 2019 | Classification and Quantification of Emphysema Using a Multi-Scale Residual NetworkabstractAutomated tissue classification is an essential step for quantitative analysis and treatment of emphysema. Although many studies have been conducted in this area, there still remain two major challenges. First, different emphysematous tissue appears in different scales, which we call "inter-class variations." Second, the intensities of CT images acquired from different patients, scanners or scanning protocols may vary, which we call "intra-class variations". In this paper, we present a novel multi-scale residual network with two channels of raw CT image and its differential excitation component. We incorporate multi-scale information into our networks to address the challenge of inter-class variations. In addition to the conventional raw CT image, we use its differential excitation component as a pair of inputs to handle intra-class variations. Experimental results show that our approach has superior performance over the state-of-the- art methods, achieving a classification accuracy of 93.74% on our original emphysema database. Based on the classification results, we also perform the quantitative analysis of emphysema in 50 subjects by correlating the quantitative results (the area percentage of each class) with pulmonary functions. We show that centrilobular emphysema (CLE) and panlobular emphysema (PLE) have strong correlation with the pulmonary functions and the sum of CLE and PLE can be used as a new and accurate measure of emphysema severity instead of the conventional measure (sum of all subtypes of emphysema). The correlations between the new measure and various pulmonary functions are up to |r| = 0.922 (r is correlation coefficient). Liying Peng, Yen-Wei Chen 0001, Lanfen Lin, Hongjie Hu, Huali Li, Qingqing Chen 0001, Xiaoli Ling, Xianhua Han, Yutaro Iwamoto |
IEEE J. Biomed. Health Informatics | 2 |
| 2018 | Classification of Pulmonary Emphysema in CT Images Based on Multi-Scale Deep Convolutional Neural NetworksabstractIn this work, we aim at classifying emphysema in computed tomography (CT) images of lungs. Most previous works are limited to extracting low-level features or mid-level features without enough high-level information. Moreover, these approaches do not take the characteristics (scales) of different emphysema into account, which are crucial for feature extraction. In contrast to previous works, we propose a novel deep learning method based on multiscale deep convolutional neural networks. There are three contributions for this paper. First, we propose to use a base residual network with 20 layers to extract more high-level information. To the best of our knowledge, this is the first deep learning method for classification of emphysema. Second, we incorporate multi-scale information into our deep neural networks so as to take full consideration of the characteristics of different emphysema. Finally, we established a high-quality emphysema dataset which contains 91 high-resolution computed tomography (HRCT) volumes, annotated manually by two experienced radiologists and checked by one experienced chest radiologist. A 92.68% classification accuracy is achieved on this dataset. The results show that (1) the multi-scale method is highly effective in comparison to the single scale setting; (2) the proposed approach is superior to the state-of-the-art techniques. Liying Peng, Lanfen Lin, Hongjie Hu, Huali Li, Xiaoli Ling, Xianhua Han, Yutaro Iwamoto, Yen-Wei Chen 0001 |
ICIP | 9 |
| 2018 | Comprehensive Study of Multiple CNNs Fusion for Fine-Grained Dog Breed CategorizationabstractFine-grained visual categorization aims to distinguish objects in subordinate classes instead of basic class, and is a challenge visual task due to the high correlation between subordinated classes and large intra-class variation (e.g. different object poses). Although, deep convolutional neural network (DCNN) has brought dramatic success on generic object classification, detection and segmentation with the availability of the large-scale training samples, direct application of DCNN on fine-grained visual categorization, where only decades or at most hundreds of training samples for each subordinate class are available in most public finegrained image datasets, cannot lead to satisfactory classification results due to small number of training samples. This study explores the transfer learning strategy for finegrained dog breed categorization based on the learned CNN models with the large-scale image dataset: ImageNet, and prove promising performance with two DCNN models: AlexNet and VGG-16. Furthermore, we argue that different DCNN architecture may extract the representation of different image aspects due to the previously defined CNN kernel sizes, number and various operations in the model learning procedure, and thus result in different performance for visual categorization. This study proposes to fusion multiple CNN architectures for combining different aspect representations to give more accurate performance. We compressively study the fusion of different layers such as Fc6 and Fc7 in AlexNet and VGG-16, and manifest 2.88% improvement of the fusion architecture over the best performance of the only one DCNN model: VGG-16 from 81.2% to 84.08%. Minori Uno, Xianhua Han, Yen-Wei Chen 0001 |
ISM | 3 |
| 2018 | Combining Convolutional and Recurrent Neural Networks for Classification of Focal Liver Lesions in Multi-phase CT Images
Lanfen Lin, Hongjie Hu, Qiaowei Zhang, Qingqing Chen 0001, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001 |
MICCAI (2) | 8 |
| 2018 | Residual Convolutional Neural Networks with Global and Local Pathways for Classification of Focal Liver Lesions
Lanfen Lin, Hongjie Hu, Qiaowei Zhang, Qingqing Chen 0001, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001 |
PRICAI (1) | 8 |
| 2018 | Generic and Specific Impressions Estimation and Their Application to KANSEI-Based Clothing Fabric Image RetrievalabstractCurrent image retrieval techniques are mainly based on text or visual contents. However, both text-based and contents-based methods lack the capability of utilizing human intuition and KANSEI (impression). In this paper, we proposed an impression-based image retrieval method in order to realize the image retrieval according to our impression presented by impression keywords. We first propose a generic and specific impressions estimation method based on machine learning and then apply it to impression-based clothing fabric image retrieval. We use a semantic differential (SD) method to measure the user’s impressions such as brightness and warmth while they view a cloth fabric image. We also extract both global and local features of cloth fabric images such as color and texture using computer vision techniques. Then we use support vector regression to model the mapping functions between the generic impression (or specific impression) and image features. The learnt mapping functions are used to estimate the generic and specific impressions of cloth fabric images. The retrieval is done by comparing the query impression with the estimated impression of images in the database. Yen-Wei Chen 0001, Xinyin Huang, Dingye Chen, Xianhua Han |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2017 | Joint weber-based rotation invariant uniform local ternary pattern for classification of pulmonary emphysema in CT imagesabstractIn this paper, we present a novel image representation approach for classifying emphysema in computed tomography (CT) images of the lung. Our proposed method extends rotation invariant uniform local binary pattern (RIULBP) and local ternary pattern (LTP), which are extensively used in a variety of computer vision applications, into rotation invariant uniform local ternary pattern (RIULTP) with a human perception principle: Weber's law. In addition, by integrating the upper pattern and the lower pattern of the Weber-based RIULTP (WRIULTP), we further put forward the joint Weber-based rotation invariant uniform local ternary pattern (JWRIULTP), which allows for a much richer representation and also takes the comprehensive information of the image into account. The proposed methods are tested on the Outex database (texture database) and the Bruijne and Srensen database (emphysema database). The results show the superiority of the proposed approaches to the state-of-the-art techniques for emphysema classification including rotation invariant local binary pattern (RILBP) and texton-based approach. Liying Peng, Lanfen Lin, Hongjie Hu, Xiaoli Ling, Xianhua Han, Yen-Wei Chen 0001 |
ICIP | 7 |
| 2017 | Incorporating a locally estimated appearance model in the graphcuts algorithm to extract small hepatic vesselsabstractIn this paper, we incorporate a locally estimated appearance model to enrich the data term of the graph-cuts algorithm. It balances between data term and smoothing terms in order to extract small liver vessels. We estimate stochastic parameters of vessel and liver tissues using the MIP image. Medial axes of the vessels are then enhanced by a multi-scale filter. The skeleton of the axes are used to prepare the appearance model which is employed to prepare s-link weights in the graph-cuts algorithm. We evaluated the proposed method quantitatively using public synthetic data and qualitatively using clinical images. We obtained an average Dice measure of 0.93 which was comparable to recent researches. We achieved segmentation of liver vessels up to the fourth order too. Neda Sangsefidi, Amir Hossein Foruzan, Ardeshir Dolati, Yen-Wei Chen 0001 |
ICIP | 4 |
| 2017 | Hyper-spectral Image Super-resolution Using Non-negative Spectral Representation with Data-Guided SparsityabstractHyperspectral imaging has great potential for understanding the characteristics of different materials in many applications ranging from remote sensing to medical imaging. However, due to various hardware limitations, only low-resolution hyperspectral and high-resolution multi-spectral images can be available using existing imaging techniques. This study aims to generate a high-resolution hyperspectral image via fusion of the available LR-HS and HR-MS images. We propose a novel hyperspectral image superresolution method via non-negative sparse representation of reflectance spectral with adaptive sparsity constraint. By analyzing local content similarity of a focused pixel in the available high-resolution multi-spectral image, which can measure pixel material purity according to surrounding pixels, we generate a sparsity map for guiding non-negative sparse coding optimization procedure of the spectral representation called non-negative spectral representation with data-guided sparsity. Since the proposed method adaptively adjust the sparsity in the spectral representation based on the local content of the available high-resolution multi-spectral image, it can produce more robust spectral representation for recovering the target high-resolution hyper-spectral image. Comprehensive experiments on two public hyperspectral datasets validate that the proposed method achieves promising performances compared with the existing state of the art methods. Xianhua Han, Jan Wang, Boxin Shi, Yinqiang Zheng, Yen-Wei Chen 0001 |
ISM | 5 |
| 2017 | Tensor Sparse Representation of Temporal Features for Content-Based Retrieval of Focal Liver Lesions Using Multi-phase Medical ImagesabstractContent Based Image Retrieval (CBIR) systems that search similar images in a large database are attracting more and more research interests recently, and have been applied to medical image characterization for expert's experience sharing. One challenging task in CBIR is how to extract features for effective image representation. Therein sparse coding technique has been proven to be an effective way to learn inherent structure features for image analysis. However, it is necessary to first vectorize the 2- or 3-dimensional spatial structure for analysis with sparse coding, and then destroy the spatial relation of nearby voxels. In this study, we propose a multilinear sparse coding method to learn features from multi-dimensional medical images. We regard high dimensional local structures as tensors and propose a K-CP (CANDECOMP/PARAFAC) algorithm to learn a tensor dictionary in an iterative way. With the learned tensor dictionary, sparse coefficients of tensor local structures are calculated by multilinear orthogonal matching pursuit (MOMP) algorithm, which is an extended multilinear version of the conventional linear OMP. The proposed multilinear sparse coding method is prospected to be more efficient and effective for inherent feature extraction compared with conventional linear methods. The proposed method is applied to a CBIR system for retrieval of focal liver lesions (FLLs) using a medical database consisting of contrast-enhanced multi-phase computer-tomography (CT) images. Experiments show that the constructed CBIR with multilinear sparse coding method can achieve promising retrieval performance. Jian Wang 0004, Xianhua Han, Lanfen Lin, Hongjie Hu, Chongwu Jin, Yen-Wei Chen 0001 |
ISM | 7 |
| 2017 | Multi-dimensional data representation using linear tensor codingabstractLinear coding is widely used to concisely represent data sets by discovering basis functions of capturing high‐level features. However, the efficient identification of linear codes for representing multi‐dimensional data remains very challenging. In this study, the authors address the problem by proposing a linear tensor coding algorithm to represent multi‐dimensional data succinctly via a linear combination of tensor‐formed bases without data expansion. Motivated by the amalgamation of linear image coding and multi‐linear algebra, each basis function in the authors’ algorithm captures some specific variabilities. The basis‐associated coefficients can be used for data representation, compression and classification. When the authors apply the algorithm on both simulated phantom data and real facial data, the experimental results demonstrate their algorithm not only preserves the original information of input data, but also produces localised bases with concrete physical meanings. Xu Qiao, Yen-Wei Chen 0001, Zhi-Ping Liu |
IET Image Process. | 3 |
| 2017 | HEp-2 staining pattern recognition using stacked fisher network for encoding weber local descriptor
Xianhua Han, Yen-Wei Chen 0001 |
Pattern Recognit. | 2 |
| 2016 | Quantitative analysis of facial paralysis based on three-dimensional featuresabstractObjective evaluation of disease is one of the desirable goals in medicine. This paper presents a technique for the objective evaluation of facial paralysis, in which features are extracted based on landmark's positions in three-dimensional space (3D-landmarks). The landmarks are initialized manually in the first frontal frame and are tracked in the subsequent frontal frames. Then, the landmark's positions are reconstructed in 3D-space using multiview images and a camera self-calibration technique. From the 3D-landmarks, the features are extracted in 3D-space (called 3D-features) and used for classification. These 3D-features may contain enhanced information such as depth information and, therefore, may help improve the accuracy rates of predicted scores. In addition, our method uses the camera self-calibration technique for estimating the camera's parameters, and does not use laser scanning for 3D reconstruction, so it is more flexible to set up and safer for the patient. For overall evaluation, experiments showed that our technique achieved superior results to other methods. Truc Hung Ngo, Yen-Wei Chen 0001, Masataka Seo, Naoki Matsushiro, Wei Xiong 0001 |
ICIP | 2 |
| 2016 | Quantitative analysis of facial paralysis based on limited-orientation modified circular Gabor filtersabstractThe diagnosis of disease with the aid of computer programs has been developing more and more in recent years. This paper presents an approach which is based on frequency technique for the objective quantitative analysis of facial paralysis. In this method, limited-orientation modified circular Gabor filters (LO-MCGFs) are used to enhance the desirable frequencies in images. Then, features are extracted from the filtered images for classification. The first advantage of the LO-MCGF is that its inner passbands are uniform, so it helps remove noise and control frequencies more effectively. The second benefit is that the LO-MCGF utilizes the existing robust characteristics of circular Gabor filter for rotation invariant texture regions. Hence, the LO-MCGF-based technique improves remarkably the accuracies of score estimation for some expressions whose local textures are invariant in rotation. Finally, the limited filtered regions, or limited propagation orientations, help the LO-MCGF focus on only some specific spaces. Therefore, the LO-MCGF can avoid the influences of irrelevant regions. In other words, it improves the spatial localization. For overall evaluation, experiments show that our proposed method is superior to other contemporary techniques tested on a dynamic facial expression database. Truc Hung Ngo, Masataka Seo, Naoki Matsushiro, Wei Xiong 0001, Yen-Wei Chen 0001 |
ICPR | 5 |
| 2016 | Bag of temporal co-occurrence words for retrieval of focal liver lesions using 3D multiphase contrast-enhanced CT imagesabstractComputer-aided diagnosis (CAD) systems have been verified to have the potential to assist radiologists in clinical diagnosis to detect and characterize focal liver lesions (FLLs) based on single- or multiphase contrast-enhanced computed tomography (CT) images. Features extracted from multiphase contrast-enhanced CT images carry more important diagnostic information i.e. enhancement pattern and demonstrate much stronger discriminative ability compared to those of single-phase CT images. In this paper, we propose a new method for multiphase image feature generation called the bag of temporal co-occurrence words (BoTCoW). A temporal co-occurrence image connecting intensity from multiphase images is constructed. Then the bag of visual word (BoVW) model is employed on the temporal co-occurrence images to extract temporal features. The proposed method effectively captures temporal enhancement information and demonstrates the distribution of the evolution patterns. The effectiveness of this method is validated in a retrieval system using 132 FLLs with confirmed pathology type. The preliminary results show that the proposed BoTCoW method outperforms the previously proposed temporal features and multiphase features based on the BoVW model. Lanfen Lin, Hongjie Hu, Yitao Liu, Jian Wang 0004, Xianhua Han, Yen-Wei Chen 0001 |
ICPR | 8 |
| 2016 | A novel and fast connected component count algorithm based on graph theoryabstractA fast component-counting algorithm is proposed based on graph theory in this paper. We derived a formulation to count faces in a plane given only the vertices based on Euler polyhedron formula. Vertices with degree no more than two are ineffective in counting components. With the derived formula, a graph component counting algorithm is constructed based only on searching cross points whose degree are no less than three. When applied to a two-dimensional binary image, the proposed method divides an image into patches of same size and decides which of them will be used in counting by searching the circumferential pixels of each patch. If the number of component edges within the circumferential pixels of a patch is no less than three, then the patch will be used in counting. After determining all the vertices with degree no less than three, the number of components can be calculated by the formula. Because only a small number of pixels are investigated in the process, the computational time is very fast. The disconnection of edges in an image is one of the main reasons which causes miscounts for scanning-based algorithms. The difficulty, however, can be naturally overcome by the proposed algorithm because disconnected points will be identified as futile pixels in the algorithm. Experimental results show the algorithm is more efficient than existing methods. When applied to images with disconnected edges, the counted number given by scanning-based algorithms is much smaller than the correct number whereas the proposed algorithm obtains satisfactory results. Sihai Yang, Duansheng Chen, Xianhua Han, Yen-Wei Chen 0001 |
SNPD | 4 |
| 2016 | Integration of spatial and orientation contexts in local ternary patterns for HEp-2 cell classification
Xianhua Han, Yen-Wei Chen 0001 |
Pattern Recognit. Lett. | 2 |
| 2015 | Liver segmentation using superpixel-based graph cuts and restricted regions of shape constrainsabstractLiver segmentation is one of the most fundamental and challenging tasks in computer aided diagnosis (CAD) system for liver diseases. Graph cut algorithms have been successfully applied to medical image segmentation of different organs for 3D volume data, which not only leads to very large-scale graph due to the same node number as voxel number, but also completely ignore some available organ shape priors. Thus, a slice by slice liver segmentation method by combining shape constraints according to previously slice segmentation has been proposed based on graph cut. However, the constructed graph scale is still large, and the computation of distance map from all voxel to the segmented shape leads to high cost. In order to explore an efficient and effective slice by slice segmentation method for liver, this paper proposes to apply clustering algorithm to firstly group slice pixels into superpixels as nodes for constructing graph, which not only greatly reduce the graph scale but also significantly speed up the optimization procedure of the graph. Furthermore, we restrict the regions near organ boundary as shape constraints, which can further reduce computational time. To validate effectiveness and efficiency of our proposed method, we conduct experiments on 10 CT volumes, most of which have tumors inside liver, and abnormal deformed shape of liver. Our method can yield an average dice coefficient: 0.94, about 659.22 second in computation, and take only 1.5GB in memory usage. Titinunt Kitrungrotsakul, Xianhua Han, Yen-Wei Chen 0001 |
ICIP | 3 |
| 2015 | High-Order Statistics of Weber Local Descriptors for Image RepresentationabstractHighly discriminant visual features play a key role in different image classification applications. This study aims to realize a method for extracting highly-discriminant features from images by exploring a robust local descriptor inspired by Weber's law. The investigated local descriptor is based on the fact that human perception for distinguishing a pattern depends not only on the absolute intensity of the stimulus but also on the relative variance of the stimulus. Therefore, we firstly transform the original stimulus (the images in our study) into a differential excitation-domain according to Weber's law, and then explore a local patch, called micro-Texton, in the transformed domain as Weber local descriptor (WLD). Furthermore, we propose to employ a parametric probability process to model the Weber local descriptors, and extract the higher-order statistics to the model parameters for image representation. The proposed strategy can adaptively characterize the WLD space using generative probability model, and then learn the parameters for better fitting the training space, which would lead to more discriminant representation for images. In order to validate the efficiency of the proposed strategy, we apply three different image classification applications including texture, food images and HEp-2 cell pattern recognition, which validates that our proposed strategy has advantages over the state-of-the-art approaches. Xianhua Han, Yen-Wei Chen 0001 |
IEEE Trans. Cybern. | 2 |
| 2015 | Penrose DemosaickingabstractThe Penrose pixel layout, an aperiodic pixel layout in rhombus Penrose tiling, has been shown to substantially outperform the existing square pixel layout in super-resolution. However, it was tested only on grayscale images. To study its performance on color images, we have to reconstruct regular color images from Penrose raw images, i.e., images with only one color component at each Penrose pixel, resulting in the problem of demosaicking from Penrose pixels. Penrose demosaicking is more difficult than regular demosaicking, because none of the color components of the reconstructed regular color images are available. Therefore, most of the traditional demosaicking methods do not apply. We develop a sparse representation-based method for Penrose demosaicking. Extensive experiments show that Penrose pixel layout outperforms regular pixel layouts in terms of both perceptual evaluation and S-CIELAB. The Penrose pixel layout is unique among all irregular layouts because it is uniformly three-colorable and it has only two pixel shapes, thick and thin rhombi, making its manufacturing relatively easy. Chen-Yan Bai, Jia Li 0002, Zhouchen Lin, Jian Yu 0001, Yen-Wei Chen 0001 |
IEEE Trans. Image Process. | 5 |
| 2014 | Sparse and Low Rank Matrix Decomposition Based Local Morphological Analysis and Its Application to Diagnosis of Cirrhosis LiversabstractCirrhosis liver is a terrible disease which is threatening our lives. Meanwhile, cirrhosis will cause significant hepatic morphological changes. While it is well known that the livers from different subjects have similar global shape structure which means liver shape ensemble should be low-rank. However the deformation which caused by cirrhosis can be considered as sparse compared with the whole liver. Therefore, in this study, we proposed to apply spare and low-rank matrix decomposition to partition the local deformation part (sparse error matrix E) from the global similar structure (low-rank matrix A) using the input liver shape D, which is the landmark coordinates of liver shapes and already have been aligned by the current rigid registration methods firstly. And then sparse matrix E is used for diagnosis. In common sense, the normal liver should have less local deformation than that of abnormal liver, which means that the norm of sparse matrix E for normal liver is smaller than the norm for abnormal one. Thus, we can simply use a threshold classify normal and abnormal livers using the norm of E for these two categories. The proposed method is evaluated by a liver database which includes 30 normal livers and 30 abnormal livers. The experimental results of proposed method is better than those of state of the art statistical shape model(SSM) based methods. Junping Deng, Xianhua Han, Yen-Wei Chen 0001 |
ICPR | 4 |
| 2014 | Hybrid Aggregation of Sparse Coded Descriptors for Food RecognitionabstractRecent year, with the increasing of unhealthy diets which will threaten people's life due to the various resulted risks such as heart stroke, liver trouble and so on, the maintaining for healthy life has attracted much attention and then how to manage the dietary life is becoming more and more important. In this research, we aim to construct an auto-recognition system of food images and keep the daily food-log records which will contribute to manage dietary life. With the easily available food images taken by mobile phone, it prospects to give the insight about the daily dietary of users with our constructed food recognition system. In order to achieve the acceptable recognition performance of the food images, we propose to apply a sparse model for coding local descriptors extracted from the food images and various pooling methods for aggregating the xoded descriptors. Sparse coding: an extension of vector quantization for local descriptors, which is popularly used in Bag-of-Features (BoF) for image representation, can reconstruct the local descriptors more effective, and then obtain more discriminated feature for food image representation. However, in order to emphasize the strongest activated pattern, the widely applied aggregation strategy of the sparse coded vector is only to retain the maximum coefficient in all (named as Max-pooling), which would completely ignore the frequency: an important signature for identifying different types of images, of the activated patterns. Therefore, we explore a hybrid aggregation strategy named as top-ranked average pooling (TRAP), which integrates not only the maximum activated magnitude but also the stronger activated number for image representation. Experiments validate that the proposed hybrid aggregation strategy combined with sparse model can greatly improve the recognition rates compared with the conventional BOF model and the state-of-the-art methods on two databases: our constructed RFID and the public PFID. Riko Kusumoto, Xianhua Han, Yen-Wei Chen 0001 |
ICPR | 3 |
| 2014 | Automatic optical phase identification of micro-drill bits based on improved ASM and bag of shape segment in PCB production
Guifang Duan, Hongcui Wang, Zhenyu Liu 0005, Jianrong Tan, Yen-Wei Chen 0001 |
Mach. Vis. Appl. | 5 |
| 2013 | Pilot study of applying shape analysis to liver cirrhosis diagnosisabstractThis paper explores the potential of applying shape analysis to classify normal/cirrhotic liver and in addition estimate the severity of abnormal cases. Conventional Computer-Aided Diagnosis (CAD) systems are developed for automatically providing a binary output as a second opinion to assist radiologists to draw conclusions about the condition of the pathology (normal or abnormal). After the disease is diagnosed, grasping the proceeding stage of the abnormal degree is essential for adopting the appropriate strength of treatment. However, none of existing CAD system is well established for such a challenging task. Liver cirrhosis has an important feature: morphological changes of the liver and the spleen occur during the clinical course of liver cirrhosis. In this study we constructed liver, spleen and their joint Statistical Shape Models (SSMs) to quantitatively assess the global shape variation and selected several modes from the SSMs. Then we learnt a mapping function between coefficients of selected modes and the ground truth staging label by Support Vector Regression (SVR). Using this mapping function, the proceeding stage of new input data can be estimated. Experimental results have validated the potential of our method on assisting the cirrhosis diagnosis. Yen-Wei Chen 0001, Xianhua Han, Tomoko Tateyama, Akira Furukawa, Shuzo Kanasaki |
ICIP | 2 |
| 2013 | Reconstruction of 3D dynamic expressions from single facial imageabstractRecently automatic facial expression analysis and recognition is rapidly gaining more and more interest in the field of computer vision. The capture and construction of 3D dynamic expressions often take large time and need specialized hardware, which limits its possible applications. In this paper, we try to reconstruct 3D dynamic expression images from single 2D facial image. The proposed method is based on statistical learning, where multiple subspaces are learned and support for 3D dynamic expression generation. The results show that the proposed method can effectively generate 3D dynamic expressions using only one input 2D facial image. Shunya Osawa, Guifang Duan, Masataka Seo, Takanori Igarashi, Yen-Wei Chen 0001 |
ICIP | 5 |
| 2013 | Residual Image Compensations for Enhancement of High-Frequency Components in Face Hallucination
Yen-Wei Chen 0001, So Sasatani, Xianhua Han |
ISNN (1) | 1 |
| 2013 | SOR Based Fuzzy K-Means Clustering Algorithm for Classification of Remotely Sensed Images
Dong-jun Xin, Yen-Wei Chen 0001 |
ISNN (1) | 2 |
| 2013 | Utilizing Disease-Specific Organ Shape Components for Disease Discrimination: Application to Discrimination of Chronic Liver Disease from CT Data
Dipti Prasad Mukherjee, Keisuke Higashiura, Toshiyuki Okada, Masatoshi Hori, Yen-Wei Chen 0001, Noriyuki Tomiyama, Yoshinobu Sato |
MICCAI (1) | 5 |
| 2013 | Generalized N-dimensional independent component analysis and its application to multiple feature selection and fusion for image classification
Danni Ai, Guifang Duan, Xianhua Han, Yen-Wei Chen 0001 |
Neurocomputing | 4 |
| 2013 | Fast and effective color-based object tracking by boosted color distribution
Dong Wang 0004, Huchuan Lu, Ziyang Xiao, Yen-Wei Chen 0001 |
Pattern Anal. Appl. | 4 |
| 2012 | Multiple feature selection and fusion based on generalized N-dimensional independent component analysis
Danni Ai, Guifang Duan, Xianhua Han, Yen-Wei Chen 0001 |
ICPR | 4 |
| 2012 | K-CPD: Learning of overcomplete dictionaries for tensor sparse coding
Guifang Duan, Hongcui Wang, Zhenyu Liu 0005, Junping Deng, Yen-Wei Chen 0001 |
ICPR | 5 |
| 2012 | Group sparse representation of adaptive sub-domain selection for image classification
Xianhua Han, Xu Qiao, Yen-Wei Chen 0001 |
ICPR | 3 |
| 2012 | Super-resolution of MR volumetric images using sparse representation and self-similarity
Yutaro Iwamoto, Xianhua Han, So Sasatani, Kazuki Taniguchi, Wei Xiong 0001, Yen-Wei Chen 0001 |
ICPR | 6 |
| 2012 | Image super-resolution based on locality-constrained linear coding
Kazuki Taniguchi, Xianhua Han, Yutaro Iwamoto, So Sasatani, Yen-Wei Chen 0001 |
ICPR | 5 |
| 2012 | Application of ICA to X-ray coronary digital subtraction angiography
Songyuan Tang, Yongtian Wang, Yen-Wei Chen 0001 |
Neurocomputing | 3 |
| 2012 | Human body segmentation based on deformable models and two-scale superpixel
Shifeng Li, Huchuan Lu, Xiang Ruan, Yen-Wei Chen 0001 |
Pattern Anal. Appl. | 4 |
| 2012 | Multilinear Supervised Neighborhood Embedding of a Local Descriptor Tensor for Scene/Object RecognitionabstractIn this paper, we propose to represent an image as a local descriptor tensor and use a multilinear supervised neighborhood embedding (MSNE) for discriminant feature extraction, which is able to be used for subject or scene recognition. The contributions of this paper include: 1) a novel feature extraction approach denoted as the histogram of orientation weighted with a normalized gradient (NHOG) for local region representation, which is robust to large illumination variation in an image; 2) an image representation framework denoted as the local descriptor tensor, which can effectively combine a moderate amount of local features together for image representation and be more efficient than the popular existing bag-of-feature model; and 3) an MSNE analysis algorithm, which can directly deal with the local descriptor tensor for extracting discriminant and compact features and, at the same time, preserve neighborhood structure in tensor-feature space for subject/scene recognition. We demonstrate the performance advantages of our proposed approach over existing techniques on different types of benchmark database such as a scene data set (i.e., OT8), face data sets (i.e., YALE and PIE), and view-based object data sets (COIL-100 and ETH-80). Xianhua Han, Yen-Wei Chen 0001, Xiang Ruan |
IEEE Trans. Image Process. | 2 |
| 2012 | A Machine Learning-Based Framework for Automatic Visual Inspection of Microdrill Bits in PCB ProductionabstractIn this paper, an automatic visual inspection scheme with phase identification of microdrill bits in printed circuit board (PCB) production is proposed. Different from conventional methods in which the geometric quantities of microdrill bits are measured to compare with the prior standards, the proposed method adopts a strategy of machine learning. Thus, it lowers the requirement for the enlargement of lens and the resolution of charge-coupled device; therefore, the cost of inspecting instrument can be relatively reduced. Our method mainly includes two procedures: First, the statistical shape models of microdrill bit are built to get the shape subspace, and then the phase identification is performed in the shape subspace using some pattern recognition techniques. In this paper, we compared the performance of two statistical model methods, principal component analysis (PCA) and linear discriminate analysis, together with three classifiers, support vector machines (SVMs), neural networks, andk-nearest neighbors, respectively, for phase identification of microdrill bits. The experimental results demonstrate that using low enlargement and resolution microdrill bit images the proposed method can measure up to high inspection accuracy, and provide a conclusion that the highest identification rates are obtained by PCA-SVMs, which are higher than that of the conventional method. Guifang Duan, Hongcui Wang, Zhenyu Liu 0005, Yen-Wei Chen 0001 |
IEEE Trans. Syst. Man Cybern. Part C | 4 |
| 2011 | Batch-incremental principal component analysis with exact mean updateabstractIncremental principal component analysis (IPCA) has been of great interest in computer vision and machine learning. In this paper, we introduce a new incremental learning procedure for principal component analysis (PCA). The proposed method can keep an accurate track of the mean of the data, and can deal with a set of new observed data in batch each time in subspace updating. Furthermore, a weighting function is proposed for contribution balance of the current data and the new observed data to the new subspace. The performance of our method is illustrated in the experiments on face modeling and face recognition. Guifang Duan, Yen-Wei Chen 0001 |
ICIP | 2 |
| 2011 | Canonical correlation analysis of local feature set for view-based object recognitionabstractIn this paper, we propose to use local feature set for image representation, which can represent variations in an object's appearance due to changing viewpoint or camera pose. It was evidenced that usually only a part of the object are appeared in common when taking a photo of an object in different view points. With comparison of local features set extracted from different positions of images, an object can be recognized when common part is appeared in two images, which take photos of one object in different view points. In this paper, we use Canonical Correlation (also known as principle or canonical angles), which can be thought of as the angles between two d-dimensional subspace, as similarity measure of local feature sets. The proposed approach is evaluated in various view-based object datasets (Coil-100 and ETH80) for object and object category recognition. Experiments show that the performance advantages of our proposed approach can be achieved over existing techniques. Xianhua Han, Yen-Wei Chen 0001, Xiang Ruan |
ICIP | 2 |
| 2011 | Preliminary study on statistical shape model applied to diagnosis of liver cirrhosisabstractIn computational anatomy, statistical shape model (SSM) is used for the quantitative evaluation of variations in the shapes of different organs. This paper focuses on the construction of a SSM of the liver and its application to computer-assisted diagnosis of cirrhosis. We prove the potential application of SSMs in the classification of normal and cirrhotic livers. In constructing a SSM of the liver, we first normalize volume data followed by the construction of the model using principal component analysis. The coefficients of the model are used as indicators of liver pathology. The effectiveness of the constructed model is evaluated by the classification accuracy of both normal and abnormal data. Shinya Kohara, Tomoko Tateyama, Amir Hossein Foruzan, Akira Furukawa, Shuzo Kanasaki, Makoto Wakamiya, Wei Xiong 0001, Yen-Wei Chen 0001 |
ICIP | 8 |
| 2011 | Pose estimation and body segmentation based on hierarchical searching treeabstractIn this paper, we propose a novel method for pose estimation and body segmentation. We estimate the partial configuration of adjacent parts instead of detecting each single part, which makes our method more robust and accurate. Further, we develop a general model to calculate the partial configuration. Besides, we present a tree-based hierarchical probabilistic method to derive the global optimal pose. Additionally, the coarse-to-fine strategy is employed to speed up the pose estimation in the whole procedure. After finishing pose estimation, the estimated pose is used to guide body segmentation. Experiments suggest that our method is efficient and effective for pose estimation and body segmentation simultaneously. Shifeng Li, Huchuan Lu, Xiang Ruan, Yen-Wei Chen 0001 |
ICIP | 4 |
| 2011 | Liver tumor detection in CT images by adaptive contrast enhancement and the EM/MPM algorithmabstractAutomatic tumor detection and segmentation is essential for the computer-aided diagnosis of live tumors in CT images. However, it is a challenging task in low-contrast images as the low-level images are too weak to detect. In this paper, we propose a new method for the automatic detection of liver tumors. We first adaptively enhance the intensity contrast of CT images by probability density function estimation. Then, to detect tumorous regions, we use the expectation maximization/maximization of the posterior marginal (EM/MPM) algorithm, which utilizes both the intensity and label information of the adjacent regions. Finally, a shape constraint is applied to reduce noise and identify focal tumors. Quantitative evaluation experiments show that our method can accurately and effectively detect tumors even in poor-contrast CT images. Yu Masuda, Tomoko Tateyama, Wei Xiong 0001, Jiayin Zhou, Makoto Wakamiya, Syuzo Kanasaki, Akira Furukawa, Yen-Wei Chen 0001 |
ICIP | 8 |
| 2011 | High frequency compensated face hallucinationabstractFace Hallucination is, one of a learning-based super-resolution technique that can reconstruct a high-resolution image using only one low-resolution image. However, there are often some detailed high-frequency components of the reconstructed image that cannot be recovered using this method. In this study, we proposed a high-frequency compensated face hallucination method for enhancing reconstruction performance. The proposed method can be divided into three steps: 1)high-resolution image reconstruction using a conventional hallucination method; 2)residual (high-frequency components) image recovery by “training” a residual image pair; 3)compensation of the reconstructed high-resolution image obtained in step 1 with the reconstructed residual image. Experimental results show that the high-resolution images obtained using our proposed approach are much better than those obtained by conventional hallucination. So Sasatani, Xianhua Han, Takanori Igarashi, Motonori Ohashi, Yutaro Iwamoto, Yen-Wei Chen 0001 |
ICIP | 6 |
| 2011 | PCA Based Regional Mutual Information for Robust Medical Image Registration
Yen-Wei Chen 0001, Chen-Lun Lin |
ISNN (3) | 1 |
| 2011 | Analysis of cypriot icon faces using ICA-enhanced active shape model representationabstractReligious iconography is an integral component of the cultural heritage of Cyprus, which was once a part of the great Byzantine empire. On one hand, icons exhibit strict adherence to conventional symbols, poses and apparel. On the other hand, there is a great variety in the style of depiction that can be attributed to different schools and periods. This paper proposes an active shape model (ASM) based technique for icon face representation that can be used for style comparison and attribution. For centuries-old icons suffering from loss of paint, cracks and added noise from digitization artifacts, we apply an independent component analysis (ICA) technique to enhance the paintings' original work. The experimental results show that our method can effectively characterize Cypriot icons. Guifang Duan, Neela Sawant, James Z. Wang 0001, Dean R. Snow, Danni Ai, Yen-Wei Chen 0001 |
ACM Multimedia | 6 |
| 2010 | Robust Tracking Based on Pixel-Wise Spatial Pyramid and Biased Fusion
Huchuan Lu, Shipeng Lu, Yen-Wei Chen 0001 |
ACCV (4) | 3 |
| 2010 | On Feature Combination and Multiple Kernel Learning for Object Tracking
Huchuan Lu, Wenling Zhang, Yen-Wei Chen 0001 |
ACCV (3) | 3 |
| 2010 | Human Tracking by Multiple Kernel Boosting with Locality Affinity Constraints
Fan Yang 0016, Huchuan Lu, Yen-Wei Chen 0001 |
ACCV (4) | 3 |
| 2010 | Image recognition by learned linear subspace of combined bag-of-features and low-level featuresabstractImage category recognition is important to access visual information on the level of objects and scene types. This paper combines different feature representations of images and learn a compact subspace of different features for the automatic recognition of object and scene classes. Compact visual-words and low-level-features object class subspaces are automatically learned from a set of training images by a Regularized Linear Discriminant analysis (RLDA) algorithm, and the extracted RLDA-domain features are used for Support Vector Machine (SVM) classifier. The main contribution of this paper is two folds: i) Different features (bag-of-features and low-level features)is fused for image representation. ii) The compact feature subspaces (low-dimension features) of different features are learned for rendering to SVM classifier, which is computationally efficient for image category. High classification accuracy is demonstrated on object recognition database (Caltech). We confirm that the proposed strategy cam improve accuracy rate compared with state-of-the-art methods for object recognition databases. Xianhua Han, Yen-Wei Chen 0001, Xiang Ruan |
ICIP | 2 |
| 2010 | Object tracking by multi-cues spatial pyramid matchingabstractIn this paper, we propose a novel tracking framework, multi-cues spatial pyramid matching (MSPM). Different cues are used to generate a set of probability maps, where the value of each pixel indicates the probability that it belongs to the foreground. Then those probability maps are combined into a single probability map by a weighted linear function. There exist two main contributions. First, a generic probability maps fusion mechanism is proposed. The weights of different probability maps are updated dynamically to maintain local discriminative power, which is achieved by solving a regression problem efficiently. Second, spatial pyramid matching kernel is adopted as a likelihood function, which considers spatial information of object and is able to cope with occlusions naturally. Experiments performed on several challenging public video sequences demonstrate that our proposed framework achieves considerable performance, compared to algorithms with individual cues or equal weights combination, and other state-of-the-art ones. Dong Wang 0004, Huchuan Lu, Yen-Wei Chen 0001 |
ICIP | 3 |
| 2010 | Robust tracking based on Boosted Color Soft Segmentation and ICA-RabstractIn this paper, we propose a novel approach for robust visual tracking. To separate the foreground from the background, we propose a novel Boosted Color Soft Segmentation (BCSS) algorithm and incorporate Independent Component Analysis with Reference (ICA-R) into the tracking framework. In addition, we design a scheme to fuse and update BCSS and ICA-R. We also propose adaptive scale of tracking window to handle objects' scale changes. Experiments shows that our approach is more robust than some popular tracking systems. Fan Yang 0016, Huchuan Lu, Yen-Wei Chen 0001 |
ICIP | 3 |
| 2010 | Adaptive Color Independent Components Based SIFT Descriptors for Image ClassificationabstractThis paper proposes an adaptive color independent components based SIFT descriptor (termed CIC-SIFT) for image classification. Our motivation is to seek an adaptive and efficient color space for color SIFT feature extraction. Our work has two key contributions. First, based on independent component analysis (ICA), an adaptive and efficient color space is proposed for color image representation. Second, in this ICA-based color space, a discriminative CIC-SIFT descriptor is calculated for image classification. The experiment results indicate that (1) contrast between objects and background can be enhanced on the ICA-based color space and (2) the CIC-SIFT descriptor outperforms other conventional color SIFT descriptors on image classification. Danni Ai, Xianhua Han, Xiang Ruan, Yen-Wei Chen 0001 |
ICPR | 4 |
| 2010 | Image Categorization by Learned Nonlinear Subspace of Combined Visual-Words and Low-Level FeaturesabstractImage category recognition is important to access visual information on the level of objects and scene types. This paper presents a new algorithm for the automatic recognition of object and scene classes. Compact and yet discriminative visual-words and low-level-features object class subspaces are automatically learned from a set of training images by a Supervised Nonlinear Neighborhood Embedding (SNNE) algorithm, which can learn an adaptive nonlinear subspace by preserving the neighborhood structure of the visual feature space. The main contribution of this paper is two fold: i) an optimally compact and discriminative feature subspace is learned by the proposed SNNE algorithm for different feature space (visual-word and low-level features). ii) An effective merge of different feature subspace can be implemented simply. High classification accuracy is demonstrated on different database including the scene database (Simplicity) and object recognition database (Caltech). We confirm that the proposed strategy is much better than state-of-the-art methods for different databases. Xianhua Han, Yen-Wei Chen 0001, Xiang Ruan |
ICPR | 2 |
| 2010 | Semi-supervised and Interactive Semantic Concept Learning for Scene RecognitionabstractIn this paper, we present a novel semi-supervised and interactive concept learning algorithm for scene recognition by local semantic description. Our work is motivated by the continuing effort in content-based image retrieval to extract and to model the semantic content of images. The basic idea of the semantic modeling is to classify local image regions into semantic concept classes such as water, sunset, or sky. However, labeling concept sampling manually for training semantic model is fairly expensive, and the labeling results is, to some extent, subjective to the operators. In this paper, by using the proposed semi-supervised and interactive learning algorithm, training samples and new concepts can be obtained accurately and efficiently. Through extensive experiments, we demonstrate that the image concept representation is well suited for modeling the semantic content of heterogenous scene categories, and thus for recognition and retrieval. Furthermore, higher recognition accuracy can be achieved by updating new training samples and concepts, which are obtained by the novel proposed algorithm. Xianhua Han, Yen-Wei Chen 0001, Xiang Ruan |
ICPR | 2 |
| 2010 | Statistical Texture Modeling for Medical Volume Using Generalized N-Dimensional Principal Component Analysis Method and 3D Volume MorphingabstractIn this paper, a statistical texture modeling method is proposed for medical volumes. As the shapes of the human organ are very different from one case to another, 3D volume morphing is applied to normalize all the volume datasets to a same shape for removing shape variations. In order to deal with the problems of high-dimension and small number of medial samples, we propose an effective image compression method named Generalized N-dimensional Principal Component Analysis (GND-PCA) to construct a statistical model. Experiments applied on liver volumes show good performance on generalization using our method. A simple experiment is employed to show that the features extracted by the statistical texture model have capability of discrimination for different types of data, such as normal and abnormal. Xu Qiao, Yen-Wei Chen 0001 |
ICPR | 2 |
| 2010 | Incremental MPCA for Color Object TrackingabstractThe task of visual tracking is to deal with dynamic image streams that change over time. For color object tracking, although a color object is a 3-order tensor in essence, little attention has been focused on this attribute. In this paper, we propose a novel Incremental Multiple Principal Component Analysis (IMPCA) method for online learning dynamic tensor streams. When newly added tensor set arrives, the mean tenor and the covariance matrices of different modes can be updated easily, and then projection matrices can be effectively calculated based on covariance matrices. Finally, we apply our IMPCA method to color object tracking using Bayes inference framework. Experiments are performed on some changeling public and our own video sequences. The experimental results demonstrate that the proposed method achieves considerable performance. Dong Wang 0004, Huchuan Lu, Yen-Wei Chen 0001 |
ICPR | 3 |
| 2010 | Bag of Features TrackingabstractIn this paper, we propose a visual tracking approach based on "bag of features" (BoF) algorithm. We randomly sample image patches within the object region in training frames for constructing two codebooks using RGB and LBP features, instead of only one codebook in traditional BoF. Tracking is accomplished by searching for the highest similarity between candidates and codebooks. Besides, updating mechanism and result refinement scheme are included in BoF tracking. We fuse patch-based approach and global template-based approach into a unified framework. Experiments demonstrate that our approach is robust in handling occlusion, scaling and rotation. Fan Yang 0016, Huchuan Lu, Yen-Wei Chen 0001 |
ICPR | 3 |
| 2010 | Liver Segmentation from Low Contrast Open MR Scans Using K-Means Clustering and Graph-Cuts
Yen-Wei Chen 0001, Katsumi Tsubokawa, Amir Hossein Foruzan |
ISNN (2) | 1 |
| 2010 | Tensor-based subspace learning and its applications in multi-pose face synthesis
Xu Qiao, Xianhua Han, Takanori Igarashi, Keisuke Nakao, Yen-Wei Chen 0001 |
Neurocomputing | 5 |
| 2010 | Automatic optical flank wear measurement of microdrills using level set for cutting plane segmentation
Guifang Duan, Yen-Wei Chen 0001, Takeshi Sukegawa |
Mach. Vis. Appl. | 2 |
| 2010 | A novel method for gaze tracking by local pattern model and support vector regressor
Huchuan Lu, Guo-Liang Fang, Yen-Wei Chen 0001 |
Signal Process. | 4 |
| 2009 | Independent component analysis based ring artifact reduction in cone-beam CT imagesabstractCone-beam CT (CBCT) scanners are based on volumetric tomography, using a 2D extended digital array providing an area detector. Compared to traditional CT, CBCT has many advantages, such as less X-ray beam limitation, high image accuracy, rapid scan time, etc. However, In CBCT images there are always some ring artifacts that appear as rings centered on the rotation axis. Due to the data of the constructed images are corrupted by these ring artifacts, qualitative and quantitative analysis of CBCT images will be compromised. Post processing and application such as image segmentation and registration also turn more complex as the presence of such artifacts. In this paper, a method based on independent component analysis (ICA) is presented. It deals with the reconstructed CBCT image and can effectively reduce such ring artifacts. Yen-Wei Chen 0001, Guifen Duan |
ICIP | 1 |
| 2009 | Hybrid particle swarm optimization for 3-D image registrationabstractIn image guided surgery, the registration of pre-and intra-operative image data is an important issue. In registrations, we seek an estimate of the transformation that registers the reference image and test image by optimizing their metric function (similarity measure). To date, local optimization techniques, such as the gradient decent method, are frequently used for medical image registrations. But these methods need good initial values for estimation in order to avoid the local minimum. Recently several global optimization methods such as genetic algorithm (GA) and particle swarm optimization (PSO) have been proposed for medical image registration. In this paper, we propose a new approach named hybrid particle swarm optimization (HPSO) for 3-D medical image registration, which incorporates two concepts (subpopulation and crossover) of genetic algorithms into the conventional PSO. Experimental results with both mathematic test functions and medical volume data show that the proposed HPSO performs much better results than conventional gradient decent method, GA and PSO. Yen-Wei Chen 0001, Aya Mimori, Chen-Lun Lin |
ICIP | 1 |
| 2009 | Improved Active Shape Model for automatic optical phase identification of microdrill bits in Printed Circuit Board productionabstractAn improved Active Shape Model (ASM), for automatic optical phase identification of microdrill bits in Printed Circuit Board (PCB) production, is presented. To overcome the limitations of conventional ASM on fitting new instants of microdrill bits, six key landmarks are defined for the initialization and optimization of ASM, and a novel method based on projection profiles is also proposed for these key landmarks detection. In addition, local structures of landmarks are redefined according to the feature of microdrill bit images. The fitted shape points are employed for phase identification of microdrill bits with a correlation coefficient as the distance criterion. Experimental results show that our proposed method outperforms the conventional ASM and can improve the accuracy of phase identification of microdrill bits. Guifang Duan, Yen-Wei Chen 0001 |
ICIP | 2 |
| 2009 | An active contours method based on intensity and reduced Gabor features for texture segmentationabstractIn this paper, we propose a cooperative strategy for segmentation of texture images which integrates reduced Gabor features and image components. In contrast with the structure tensor method, our algorithm can extract more important features for segmentation. In this work, Gabor filters tuned to a set of orientations, scales and frequencies are used to extract texture local features, and the vector-valued active contour without edges model is employed to segment images. The main contribution of this work is the cooperation of image components and the reduced Gabor features which are extracted by principal components analysis (PCA) to represent image features. This cooperation improves the quality of the method, since the segmentation is faster and better. We demonstrate the effectiveness of our algorithm by comparing with the method proposed by Wang for segmenting synthetic and nature texture images. Huchuan Lu, Yunyun Liu, Yen-Wei Chen 0001 |
ICIP | 4 |
| 2009 | Generalized N-dimensional principal component analysis (GND-PCA) and its application on construction of statistical appearance models for medical volumes with fewer samples
Rui Xu 0002, Yen-Wei Chen 0001 |
Neurocomputing | 2 |
| 2008 | Semiautomatic non-rigid 3-D image registration for MR-Guided Liver Cancer SurgeryabstractRecently a growing interest has been seen in minimally invasive treatments with open configuration magnetic resonance (Open-MR) scanners. In this paper, we proposed a semi-automatic non-rigid 3D MR-CT image registration technique for MR-Guided Liver Cancer Surgery in which cancer tissues are coagulated by microwave ablation. Because of the lower magnetic field (0.5 T) and various different surgical conditions, sometimes tumors can not be visualized clearly on Open-MR volumes. Combining of CT volumes acquired before surgery, it is possible to identify the tumor's location by application of registration techniques. Since such a registration problem belongs to a non-rigid one considering the easy deformation of livers, free-form deformation (FFD) based registration method is applied. Similarity measurement in registration is normalized mutual information (NMI). Both phantom and clinical experiments show that the registration is accurate enough (ap1.45 mm) for liver cancer surgery given some proper processing steps. Yen-Wei Chen 0001, Katsumi Tsubokawa, Rui Xu 0002, Shigehiro Morikawa, Yoshimasa Kurumi |
ICIP | 1 |
| 2008 | A supervised nonlinear neighborhood embedding of color histogram for image indexingabstractSubspace learning techniques are widespread in pattern recognition research. They include PCA, ICA, LPP, etc. These techniques are generally linear and unsupervised. The problem of image indexing is very complicated and the processed images are usually lie on non-linear image subspaces. In this paper, we propose a supervised nonlinear neighborhood embedding algorithm which learns an adaptive nonlinear subspace by preserving the neighborhood structure of the image color space. In the proposed algorithm, we combine the idea of nonlinear kernel mapping and preserving the neighborhood structure of the samples, so it can not only gain a perfect approximation of the nonlinear image manifold, but also enhance within-class neighborhood information. Experimental results show that the proposed method outperform other linear or unsupervised subspace learning methods. Xianhua Han, Yen-Wei Chen 0001, Takeshi Sukegawa |
ICIP | 2 |
| 2008 | Multilinear analysis based on image texture for face recognitionabstractIn this paper, a multilinear approach based on image texture for face recognition is present. First, we extract the texture features of the facial images using the local binary pattern (LBP) algorithm. Then, we apply the high-order orthogonal iteration (HOOI) algorithm, the algebra of higher-order tensors, to obtain a compact and effective representation of the facial images based on the texture features. Our representation yields improved facial recognition rates relative to standard eigenface and tensorface especially when the facial images are confronted by a variety of viewpoints and illuminations. To evaluate the validity of our approach, a series of experiments are performed on the CMU PIE facial databases. Huchuan Lu, Hao Chen 0011, Yen-Wei Chen 0001 |
ICPR | 3 |
| 2008 | Gaze tracking by Binocular Vision and LBP featuresabstractIn this paper, a new method for eye gaze tracking is proposed under natural head movement. In this method, Local-Binary-Pattern Texture Feature (LBP) is adopted to calculate the eye features according to the characteristic of the eye, and a precise Binocular Vision approach is used to detect the space coordinate of the eye. The combined features of space coordinates and LBP features of the eyes are fed into Support Vector Regression (SVR) to match the gaze mapping function, in the hope of tracking gaze direction under natural head movement. The experimental results prove that the proposed method can determine the gaze direction accurately. Huchuan Lu, Yen-Wei Chen 0001 |
ICPR | 3 |
| 2008 | Classification of High-Resolution Satellite Images Using Supervised Locality Preserving Projections
Yen-Wei Chen 0001, Xianhua Han |
KES (2) | 1 |
| 2008 | Application of Interactive Genetic Algorithms to Boid Model Based Artificial Fish Schools
Yen-Wei Chen 0001, Kanami Kobayashi, Hitoshi Kawabayashi, Xinyin Huang |
KES (2) | 1 |
| 2007 | Appearance Models for Medical Volumes with Few Samples by Generalized 3D-PCA
Rui Xu 0002, Yen-Wei Chen 0001 |
ICONIP (1) | 2 |
| 2007 | The Application of ICA to the X-Ray Digital Subtraction Angiography
Songyuan Tang, Yongtian Wang, Yen-Wei Chen 0001 |
ISNN (2) | 3 |
| 2007 | Automated Segmentation of the Liver from 3D CT Images Using Probabilistic Atlas and Multi-level Statistical Shape Model
Toshiyuki Okada, Ryuji Shimada, Yoshinobu Sato, Masatoshi Hori, Keita Yokota, Masahiko Nakamoto, Yen-Wei Chen 0001, Hironobu Nakamura, Shinichi Tamura |
MICCAI (1) | 7 |
| 2006 | Detection of Moving Objects by Independent Component Analysis
Masaki Yamazaki, Yen-Wei Chen 0001 |
ACCV (2) | 3 |
| 2006 | A Robust MR Image Segmentation Technique Using Spatial Information and Principle Component Analysis
Yen-Wei Chen 0001, Yuuta Iwasaki |
ISNN (2) | 1 |
| 2006 | Genetic Algorithms for Optimization of Boids Model
Yen-Wei Chen 0001, Kanami Kobayashi, Xinyin Huang, Zensho Nakao |
KES (2) | 1 |
| 2006 | Segmentation of MR Images Using Independent Component Analysis
Yen-Wei Chen 0001, Daigo Sugiki |
KES (2) | 1 |
| 2006 | Ensemble learning for independent component analysis
Jian Cheng 0001, Qingshan Liu 0001, Hanqing Lu, Yen-Wei Chen 0001 |
Pattern Recognit. | 4 |
| 2006 | Robust multi-logo watermarking by RDWT and ICA
Thai Duy Hien, Zensho Nakao, Yen-Wei Chen 0001 |
Signal Process. | 3 |
| 2006 | Robust RDWT-ICA based information hiding
Thai Duy Hien, Zensho Nakao, Yen-Wei Chen 0001 |
Soft Comput. | 3 |
| 2005 | A Cascaded Ensemble Learning for Independent Component Analysis
Jian Cheng 0001, Kongqiao Wang, Yen-Wei Chen 0001 |
ISNN (1) | 3 |
| 2005 | Selection of ICA Features for Texture Classification
Xiang-Yan Zeng, Yen-Wei Chen 0001, Deborah van Alphen, Zensho Nakao |
ISNN (2) | 2 |
| 2005 | Supervised kernel locality preserving projections for face recognition
Jian Cheng 0001, Qingshan Liu 0001, Hanqing Lu, Yen-Wei Chen 0001 |
Neurocomputing | 4 |
| 2004 | Segmentation of High Resolution Satellite Images by Direction and Morphological FiltersabstractThis paper examines images taken from IKONOS to extract several features such as road relations automatically. We propose a new method which combines color, texture information and shape information for segmentation of high resolution satellite images. The method uses color and texture information for global segmentation, and shape information for local analysis. We propose a new direction filter which pays its attention to road features having information on specific directionality. We also propose another new morphology filter which is used as a length filter extracting length of each region more efficiently. Tomoko Tateyama, Zensho Nakao, Xiang-Yan Zeng, Yen-Wei Chen 0001 |
HIS | 4 |
| 2004 | A supervised nonlinear local embedding for face recognitionabstractMany recent works demonstrated that subspace analysis is a good method for face recognition. How to find the subspace is a key issue. In this paper, a supervised nonlinear local embedding (SNLE) method is proposed to construct a subspace for face recognition, in which we combine the idea of nonlinear kernel mapping and preserving local geometric relations of the samples belonging to same class. SNLE can not only gain a perfect approximation of the nonlinear face manifold, but also enhance within-class local information. Moreover, it is also equivalent to solving a generalized eigenvalue problem in mathematics. Our experiments are performed on two benchmarks, and experimental results show that the proposed method has an encouraging performance. Jian Cheng 0001, Qingshan Liu 0001, Hanqing Lu, Yen-Wei Chen 0001 |
ICIP | 4 |
| 2004 | Genetic Generation of High-Degree-of-Freedom Feed-Forward Neural Networks
Yen-Wei Chen 0001, Sulistiyo, Zensho Nakao |
ISNN (1) | 1 |
| 2004 | Image Feature Representation by the Subspace of Nonlinear PCA
Yen-Wei Chen 0001, Xiang-Yan Zeng |
KES | 1 |
| 2004 | Random Independent Subspace for Face Recognition
Jian Cheng 0001, Qingshan Liu 0001, Hanqing Lu, Yen-Wei Chen 0001 |
KES | 4 |
| 2004 | An RDWT Based Logo Watermark Embedding Scheme with Independent Component Analysis Detection
Thai Duy Hien, Zensho Nakao, Yen-Wei Chen 0001 |
KES | 3 |
| 2004 | Robust Digital Watermarking Based On Principal Component AnalysisabstractWe propose a robust digital watermarking technique based on Principal Component Analysis (PCA) and evaluate the effectiveness of the method against some watermark attacks. In this proposed method, watermarks are embedded in the PCA domain and the method is closely related to DCT or DWT based frequency-domain watermarking. The orthogonal basis functions, however, are determined by data and they are adaptive to the data. The presented technique has been successfully evaluated and compared with DCT and DWT based watermarking methods. Experimental results show robust performance of the PCA based method against most prominent attacks. Thai Duy Hien, Yen-Wei Chen 0001, Zensho Nakao |
Int. J. Comput. Intell. Appl. | 2 |
| 2004 | Texture representation based on pattern map
Xiang-Yan Zeng, Yen-Wei Chen 0001, Zensho Nakao, Hanqing Lu |
Signal Process. | 2 |
| 2003 | A Robust Logo Multiresolution Watermarking Based on Independent Component Analysis Extraction
Thai Duy Hien, Zensho Nakao, Yen-Wei Chen 0001 |
IWDW | 3 |
| 2003 | Face Recognition Using Overcomplete Independent Component Analysis
Jian Cheng 0001, Hanqing Lu, Yen-Wei Chen 0001, Xiang-Yan Zeng |
KES | 3 |
| 2003 | An ICA-Based Method for Poisson Noise Reduction
Xianhua Han, Yen-Wei Chen 0001, Zensho Nakao |
KES | 2 |
| 2003 | PCA Based Digital Watermarking
Thai Duy Hien, Yen-Wei Chen 0001, Zensho Nakao |
KES | 2 |
| 2003 | Image Retrieval Based on Independent Components of Color Histograms
Xiang-Yan Zeng, Yen-Wei Chen 0001, Zensho Nakao, Jian Cheng 0001, Hanqing Lu |
KES | 2 |
| 2000 | Playing the Rock-Paper-Scissors game with a genetic algorithmabstractThis paper describes a strategy to follow whilst playing the Rock-Paper-Scissors game. Instead of making a biased decision, a rule is adopted where the outcomes of the game from the last few turns are observed and then a deterministic decision is made. Such a strategy is encoded into a genetic string and a genetic algorithm works on a population of such strings. Good strings are produced in later generations. Such a strategy is found to be successful, and its efficiency is demonstrated by testing the strategy against both systematic and human strategies. F. F. Ali, Zensho Nakao, Yen-Wei Chen 0001 |
CEC | 3 |
| 2000 | Blind Separation Based on an Evolutionary Neural NetworkabstractWe propose an evolutionary neural network for blind source separation. In the proposed method, the separating matrix is used as connection weights of the network, which are updated by a genetic algorithm. A higher-order statistics of kurtosis, which is a simple and original criterion for independence, is used as a fitness function. The applicability of the proposed method for blind source separation is demonstrated by the simulation results. Yen-Wei Chen 0001, Xiang-Yan Zeng, Zensho Nakao |
ICPR | 1 |
| 2000 | Blind signal separation by an evolutionary neural network with higher-order statisticsabstractThe authors propose an evolutionary neural network for blind source separation (BSS). In the proposed method, the separating matrix is used as connection weights of the network, which are updated by a genetic algorithm (GA). A higher-order statistics of kurtosis, which is a simple and original criterion for independence, is used as a fitness function. The applicability of the proposed method for blind source separation is demonstrated by simulations. Yen-Wei Chen 0001, Xiang-Yan Zeng, Zensho Nakao |
KES | 1 |
| 2000 | Blind nonlinear channel identification based on higher order statistics using hybrid genetic algorithmabstractA method with 4th-order cumulant is proposed for nonlinear channel identification. Compared with the conventional method which uses 3rd-order cumulant, the proposed method does not need to use any constraints. Since the cost function with higher order statistics has local minima, we also propose to use a hybrid genetic algorithm (GA) to minimize the cost function. The applicability of the proposed method is demonstrated by computer simulations. Shusuke Narieda, Yen-Wei Chen 0001, Katsumi Yamashita |
SMC | 2 |
| 1999 | CT image reconstruction by stochastic relaxationabstractPresented in the paper is a stochastic relaxation algorithm for reconstruction of CT image from projection data obtained from four different directions. The basic idea of the algorithm presented is similar to that of the algorithm applied by S. Geman and D. Geman (1984) to image restoration. An initial configuration of an image is generated randomly. Each pixel in the image is represented by a unit in a reconstruction in a network system. An energy function describes the current states of the system. The algorithm works to minimize the energy of the system. Dynamics of the system involve visiting each unit in the reconstruction layer individually and setting its state to a new one stochastically according to a probability distribution, determined as the sigmoid of the output from all the units. Fath El Alem F. Ali, S. Yoyegawa, Zensho Nakao, Yen-Wei Chen 0001 |
KES | 4 |
| 1999 | A fast image restoration algorithm based on simulated annealingabstractA fast algorithm is proposed for image restoration based on simulated annealing. In the SA based method, the image restoration is modeled as an optimization problem, whose cost function is to be minimized by the SA. The advantage of the SA based method is that the complicated a priori constraints can be easily incorporated by the appropriate modification of the cost function in SA and that it can be used to solve ill-posed restoration problems. The disadvantage is it takes a large computation cost. The fast algorithm is proposed to reduce the large computation cost. Yen-Wei Chen 0001, T. Enokura, Zensho Nakao |
KES | 1 |
| 1999 | Two new neural network approaches to two-dimensional CT image reconstruction
Fath El Alem F. Ali, Zensho Nakao, Yen-Wei Chen 0001 |
Fuzzy Sets Syst. | 3 |
| 1999 | Restoration of gray images based on a genetic algorithm with Laplacian constraint
Yen-Wei Chen 0001, Zensho Nakao, Kouichi Arakaki, Xue Fang, Shinichi Tamura |
Fuzzy Sets Syst. | 1 |
| 1998 | A Hybrid Neural Network Training Approach of Backpropagation and Genetic Algorithm for Classification of Remotely Sensed Images
Yen-Wei Chen 0001, Xiang-Yan Zeng, Zensho Nakao |
ICONIP | 1 |
| 1998 | Three new soft computing approaches to two-dimensional CT image reconstruction: a comparative studyabstractWe develop three soft computing techniques for reconstructing two-dimensional CT images from a small number of projection data. They are genetic algorithms, simulated annealing, and backpropagation reconstruction techniques. The three techniques have been developed independently of each other. We present a comparative evaluation study of these techniques. We start by introducing the three new approaches respectively, and then present the simulation result. The reconstruction result obtained by the conventional algebraic reconstruction technique is also presented. A quantitative evaluation among the four reconstruction methods is presented. A pixel-wise error estimator is used to calculate the overall error in the reconstructed images. The estimator reveals the effectiveness of the proposed techniques compared to the conventional method. Fath El Alem F. Ali, Zensho Nakao, Yen-Wei Chen 0001 |
KES (3) | 3 |
| 1998 | A hybrid GA/SA approach to blind deconvolutionabstractA hybrid GA/SA (genetic algorithm/simulated annealing) approach is proposed for the blind deconvolution problem of image restoration. The blind deconvolution problem is modeled as an optimization problem whose cost function is to be minimized by the proposed hybrid approach. The approach combines the advantage of GA for global searches and the advantage of SA for local ones. The results indicate that it is possible to arrive at high quality solutions in reasonable time even for large-scale problems, such as image processing. Yen-Wei Chen 0001, T. Enokura, Zensho Nakao |
KES (3) | 1 |
| 1998 | An application of automaton neural networks to artificial agentsabstractWe present a model that transfers artificial intelligence into an intelligent neural network, which is called automaton neural network, and is composed of two algorithms: an automaton algorithm and a neural network algorithm. The model was applied to artificial agents to provide them with intelligence, and its applicability was demonstrated by computer simulation. Y. Kawano, Zensho Nakao, Yen-Wei Chen 0001 |
KES (3) | 3 |
| 1998 | Reconstruction of CT images by the backpropagation algorithmabstractA new and modified neural network model is proposed for CT image reconstruction from four projection data only. The model uses the well known backpropagation delta rule for adaptation of its weights. In addition to the error in projection data of the image being reconstructed, the proposed network makes use of errors in pixels between a filtered image and the reconstructed one. Improved reconstruction was obtained, and the proposed method was found to be very effective in CT image reconstruction when the given number of projection directions is very limited. I. Ohkawa, S. Tobaru, Zensho Nakao, Yen-Wei Chen 0001 |
KES (3) | 4 |
| 1998 | Optimization of computer-generated holograms by an artificial neural networkabstractSeveral computer-generated hologram (CGH) methods, such as the direct binary search, simulated annealing and genetic algorithm, have been proposed or used in order to decrease the quantum noise and reconstruction noise or to optimize the CGH. Since these methods are iterative approaches, they require long computation time to generate a CGH. In this paper, we propose a new method based on an artificial neural network (ANN) to reduce the high computation cost. In this scheme, we first use a couple of known optimized CGHs, which may be obtained by the traditional optimization methods, as teaching signals to train the ANN. With the trained ANN, we can easily and quickly obtain an optimized CGH without the optimization process for other input images. S. Yamauchi, Yen-Wei Chen 0001, Zensho Nakao |
KES (3) | 2 |
| 1998 | Blind deconvolution based on a hybrid GA/SA approachabstractA hybrid GA/SA (genetic algorithm/simulated annealing) approach is proposed for the blind deconvolution problem of image restoration. The blind deconvolution problem is modeled as an optimization problem, whose cost function is to be minimized by the proposed hybrid approach. The approach combines the advantage of GAs for global searches and the advantage of SA for local ones. The results indicate that it is possible to arrive at high-quality solutions in a reasonable time, even for large-scale problems such as image processing. Yen-Wei Chen 0001, T. Enokura, Zensho Nakao |
SMC | 1 |
| 1997 | CT image reconstruction by back-propagationabstractA neural network model is used in CT image reconstruction from four projections. The system is based on the backpropagation algorithm for adaptation of connection weights. Satisfactory agreement between the original and reconstructed images was obtained in simulation, and the results obtained are compared to those obtained by the well-known algebraic reconstruction technique (ART), and it was found that the neural network method is more effective than ART when the number of projection directions is very limited. Zensho Nakao, Fath El Alem F. Ali, Yen-Wei Chen 0001 |
KES (2) | 3 |
| 1997 | CT image reconstruction by the Boltzmann machineabstractThe Boltzmann machine model is used in CT image reconstruction from four projections. The system is based on Boltzmann simulated annealing for adaptation of pixel values. As the temperature is decreased, the gray level of images is increased exponentially to 256. Satisfactory agreement between the original and reconstructed images was obtained in simulation, and the results obtained are compared to those obtained by the well-known algebraic reconstruction technique (ART), and it was found that the neural network method is more effective than ART when the number of projection directions is very limited. Zensho Nakao, M. Noborikawa, Yen-Wei Chen 0001, Yasuyuki Kina |
KES (2) | 3 |
| 1997 | Evolutionary CT image reconstructionabstractAn evolutionary algorithm for reconstructing CT gray images from projections is presented; the algorithm reconstructs two-dimensional unknown images from four one-dimensional projections. A Laplacian constraint term is included in the fitness function of the genetic algorithm for handling smooth images, and the evolutionary process reconstructs images into finer ones by partitioning the images gradually thereby increasing the chromosome size exponentially as the generation proceeds. Results obtained are compared to those obtained by the well-known algebraic reconstruction technique (ART), and it was found that the evolutionary method is more effective than ART when the number of projection directions is very limited. Zensho Nakao, Midori Takashibu, Yen-Wei Chen 0001 |
KES (2) | 3 |
| 1996 | Reconstruction of neutron penumbral images by a constrained genetic algorithmabstractPenumbral imaging is a technique for imaging of neutrons or other penetrating radiations. The technique uses the facts that spatial information can be recovered from the shadow or penumbra that an unknown source casts through a simple large circular aperture. The limitation is that the straightforward image reconstruction will introduce some significant distortion for a large field of view because of nonisoplanaticity of the aperture point spread function. A genetic algorithm (GA) is proposed for reconstruction of penumbral images, and the technique allows distortion-free reconstruction over a large field of view. Furthermore, because in GA the complicated a priori constraints can be easily incorporated by the appropriate modification of the cost function, the algorithm is also tolerant of the noise. Yen-Wei Chen 0001, Zensho Nakao, Kouichi Arakaki, Ikuo Nakamura, Shinichi Tamura |
ICPR | 1 |
| 1996 | A parallel genetic algorithm for image restorationabstractA parallel genetic algorithm based on the island model for image restoration is presented. The algorithm divides a large population into smaller subpopulations and executes the main loop of the traditional genetic algorithm on each processor with its own subpopulation in parallel. Its performance is evaluated in a multi-workstation environment. The simulation results show that the algorithm achieves a linear speed-up with the number of processors. The parallel algorithm is also shown to have better performance on image restoration than the traditional genetic algorithm. Yen-Wei Chen 0001, Zensho Nakao, Xue Fang, Shinichi Tamura |
ICPR | 1 |
| 1996 | GA-Based Reconstruction of Plane Images from Projections
Zensho Nakao, Yen-Wei Chen 0001, Fath El Alem F. Ali |
IEA/AIE | 2 |