EDBT 2026 Demo / reviewers in the wild / expert
Xianhua Han
dblp:72/3760 · also Xian-Hua Han
· DBLP profile ↗
88ranked-venue papers
22as first author
35since 2021 · last 2025
0000-0002-5003-3180ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 63 · 15 first-author · 26 since 2021Artificial intelligence and machine learning · 32 · 10 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 1 first-author · 8 since 2021Computer networks · 3 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multi-Degradation Oriented Deep Unfolding Model for Hyperspectral Image ReconstructionabstractDeep unfolding framework effectively combines model-driven and data-driven approaches, which generally contains a data reconstruction error term and a prior learning network, and has made significant progresses in hyperspectral image (HSI) reconstruction. However, existing methods face challenges in generalization and representation for high-dimensional HSI data, which are evident in two main areas: 1) the reliance on a fixed sensing mask limits the ability to generalize when reconstructing compressive measurements that are out of distribution, and 2) the prior representation network struggles to accurately represent high-dimensional data in both spatial and spectral domains. To address the above challenges, this study exploits a novel deep unfolding model (DUM), named Multi-Degradation Oriented DUM (MDO-DUM), designed to enhance both inverse projection and prior learning. Our approach improves generalization by training the DUM with samples synthesized using a variety of masks and integrating a mask-aware data modeling module (MADM). The MADM module works in conjunction with both the data reconstruction term and the prior learning network to facilitate degradation-aware projection and context-aware representation. For robust prior representation, we employ a spatial-spectral transformer that models both non-local spatial and spectral dependencies, effectively capturing the 3D attributes of HSIs. Additionally, we incorporate feature interactions across stages to capture diverse contexts and use auxiliary losses at each stage to boost the recovery performance. Extensive experiments on both simulated and real-world datasets demonstrate that our method surpasses existing state-of-the-art HSI reconstruction techniques. Xianhua Han, Jian Wang 0004 |
ICASSP | 1 |
| 2025 | Deep Dual Internal Learning for Hyperspectral Image Super-Resolution
Yongqing Sun, Hong Liu 0009, Qiong Chang, Xianhua Han |
MMM (1) | 4 |
| 2025 | Deep RGB-guided generative network for unsupervised hyperspectral image super-resolution
Xianhua Han, Zhe Liu 0039 |
Appl. Intell. | 1 |
| 2024 | Dual Directional Complementary Gradient Fusion and Deep Refinement for Hyperspectral Image Super ResolutionabstractThe spatial and spectral resolution trade-off in the hyperspectral imaging is a fundamental and essential issue, and automatically generating high-resolution images in both spatial and spectral domains (HR-HS) by merging a low spatial resolution hyperspectral (LR-HS) image and a high spatial resolution RGB (HR-RGB) image, which are captured by the existing commercial sensors, has recently attracted extensive attention. Motivated by the powerful representation capability of the deep learning networks, current dominated methods have devoted to design deep and complicated network architectures, and manifested great performance progress. This study aims to exploit a simple yet effective deep model by aggregating the complementary missing information into the feature learning branches and automatically modeling the relationship between the target and observations. Specifically, we incorporate the spectral gradient of the LR-HS image with the feature learning branch of the HR-RGB image while aggregate the spatial gradient of the HR-RGB image into the learning branch of the LR-HS image to enhance the representation capability in both spatial ans spectral domains. Moreover, we reconstruct a serials of target HR-HS image from the fused features of dual branches, and employ the un-recovered residuals in the observations by automatically learning the degradation procedure to further refine the former reconstruction in an asymptotic way. Comprehensive experiments have demonstrated that our proposed deep model for HR-HS image reconstruction achieves superior SR performance over state-of-the-art methods in term of quantitative metrics and perceptive quality. YinWei Du, Jian Wang 0004, Xing Wu 0001, Xianhua Han |
ICASSP | 4 |
| 2024 | Hyperspectral Image Reconstruction Using Hierarchical Neural Architecture Search from A Snapshot ImageabstractHyperspectral imaging is a promising imaging modality, and has attracted increasing research attention by compressive sensing such as coded aperture snapshot spectral imaging (CASSI), for simultaneously capturing abundant information in spatial, spectral and temporal domains. Hyperspectral image (HSI) reconstruction in the CASSI aims to retrieve the original 3D signal upon the 2D compressed snapshot. Recently, deep learning has extensively been employed for HSI reconstruction via manually designing network architectures, and usually causes complicated and massive-computational models, which are difficult for being embedding in the real imaging systems. This study aims to leverage network architecture search to automatically design effective and efficient network architectures for HSI reconstruction. Specifically, we exploit gradient-based search strategies and prepare optional operations (cells) with adaptive receptive field such as dilate and deformable convolutional layers to construct a flexible hierarchical search space. Through sharing cells within different levels of features and utilizing an early stopping technique, we achieve a computational and memory efficient NAS strategy to automatically design an effective lightweight model for HSI reconstruction. Extensive experimental results have demonstrated that the network architecture achieved by our proposed NAS has much smaller model size and a lower computational cost while produce better or comparable HSI reconstruction performance compared with the state-of-the-art methods. Xianhua Han, Huiyan Jiang, Yen-Wei Chen 0001 |
ICASSP | 1 |
| 2024 | Deep Versatile Hyperspectral Reconstruction Model from A Snapshot Measurement with Arbitrary MasksabstractRecently, coded aperture snapshot spectral imaging (CASSI) has been actively researched to capture three-dimensional (3D) hyperspectral (HS) images for dynamic scenes, where the optical systems detect a 2D snapshot measurement while a computational algorithm performs the inverse problem for recovering the latent HS cubic data. Benefiting from the powerful modeling capability of the deep convolution neural networks (DCNN), the reconstruction performance of the HS images has been significantly improved. However, the existing deep methods usually assume a particular hardware mask to train the reconstruction models, and restrict widely applicability to the snapshots measured in different hardwares. This study exploits a novel deep versatile HS reconstruction framework for adaptively handling the snapshots with arbitrary masks. Specifically, we employ a meta-learning like training procedure using the training paired samples of different distributions to learn a highly generalized model, and further incorporate a mask structure modeling module to produce effective knowledge for modulating the spectral recovering procedure. Moreover, we configure the deep reconstruction model with the spectral transformer for modeling the long-dependence in spectral domain, which is especially critical for high fidelity spectral recovering. Experiments on two benchmark HS datasets have demonstrated the superiority of our framework over the state-of-the-art methods. Takumi Takabe, Xianhua Han, Yen-Wei Chen 0001 |
ICASSP | 2 |
| 2024 | Lesion Feature Extraction and Classification Optimization Method Using Dynamic Fusion of Global Attention and Local AttentionabstractIn tumor diagnosis, due to subtle differences in the imaging appearance of different diseases, accurately classifying lesions based on solely imaging data proves challenging. Existing machine learning and deep learning methods face limitations due to the small sample size of medical datasets and the intricate nature of disease image manifestations. This paper proposes a novel lesion classification method to fully explore distinctions among confused lesion features associated with different diseases. The proposed method comprises three key steps: Firstly, a lesion feature calculation method using dynamic fusion of global attention and local attention is proposed. The weight of global attention and local attention is dynamically allocated, and the global and local features are fused by dynamic weight. Secondly, feature dimension reduction is realized to improve the effect of distinguishable features using sparse autoencoder and polynomial constraint loss function. Finally, to improve the performance of classification, the monarch butterfly optimization algorithm based on adaptive neighborhood search radius method is used to optimize the parameters of multi-kernel support vector machine. The private PET/CT image classification dataset of lymphoma and Still’s disease was used to validate our results. The experimental results demonstrate that the method's efficacy in lymphoma and Still’ disease classification tasks, achieving an accuracy (ACC) of 82.8% and an area under the curve (AUC) of 87.1%, respectively. Xueyao Cui, Huiyan Jiang, Xianhua Han, Xuena Li, Yan Pei 0001 |
IJCNN | 4 |
| 2024 | FedEL: Federated ensemble learning for non-iid data
Xing Wu 0001, Jie Pei, Xianhua Han, Yen-Wei Chen 0001, Junfeng Yao, Yang Liu 0005, Quan Qian, Yike Guo |
Expert Syst. Appl. | 3 |
| 2024 | Coupled image and kernel prior learning for high-generalized super-resolution
Xianhua Han, Kazuhiro Yamawaki, Huiyan Jiang |
Neurocomputing | 1 |
| 2024 | Segmentation Guided Crossing Dual Decoding Generative Adversarial Network for Synthesizing Contrast-Enhanced Computed Tomography ImagesabstractAlthough contrast-enhanced computed tomography (CE-CT) images significantly improve the accuracy of diagnosing focal liver lesions (FLLs), the administration of contrast agents imposes a considerable physical burden on patients. The utilization of generative models to synthesize CE-CT images from non-contrasted CT images offers a promising solution. However, existing image synthesis models tend to overlook the importance of critical regions, inevitably reducing their effectiveness in downstream tasks. To overcome this challenge, we propose an innovative CE-CT image synthesis model called Segmentation Guided Crossing Dual Decoding Generative Adversarial Network (SGCDD-GAN). Specifically, the SGCDD-GAN involves a crossing dual decoding generator including an attention decoder and an improved transformation decoder. The attention decoder is designed to highlight some critical regions within the abdominal cavity, while the improved transformation decoder is responsible for synthesizing CE-CT images. These two decoders are interconnected using a crossing technique to enhance each other's capabilities. Furthermore, we employ a multi-task learning strategy to guide the generator to focus more on the lesion area. To evaluate the performance of proposed SGCDD-GAN, we test it on an in-house CE-CT dataset. In both CE-CT image synthesis tasks-namely, synthesizing ART images and synthesizing PV images-the proposed SGCDD-GAN demonstrates superior performance metrics across the entire image and liver region, including SSIM, PSNR, MSE, and PCC scores. Furthermore, CE-CT images synthetized from our SGCDD-GAN achieve remarkable accuracy rates of 82.68%, 94.11%, and 94.11% in a deep learning-based FLLs classification task, along with a pilot assessment conducted by two radiologists. Qingqing Chen 0001, Yinhao Li 0002, Fang Wang 0030, Xianhua Han, Yutaro Iwamoto, Jing Liu 0041, Lanfen Lin, Hongjie Hu, Yen-Wei Chen 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2023 | Local to global prior Learning for blind Unsupervised Image super ResolutionabstractDeep convolutional neural networks (DCNN) have dominated the single image super resolution (SR) field, and demonstrated significant success in generating high-resolution (HR) images from the ideal low-resolution (LR) images captured under controlled imaging conditions. However, the resolved image performance would be dramatically degraded for the LR image captured in real environment. Recently, some works attempt to automatically learn image and kernel priors by leveraging the local statistic modeling capability of the DCNNs from the LR observation only, and illustrated the feasibility to construct a specific CNN for real-world application. The basic convolution operation in the DCNNs can potentially capture the local interaction to generate plausible image appearance but are far from sufficiency to achieve global context, which has been proven to be a promising property of a transformer block, to further boost the SR performance. This study proposes a cooperative local to global prior learning (LoGPT) framework for blind unsupervised image super resolution by jointly modeling the local connectivity with convolution operations and global context with transformer block. Specifically, we elaborate a multi-scale encoder-decoder architecture configuring with the convolution blocks on the high resolution scales to learn local priors while the transformer block on the low-resolution scales to capture long-range dependencies, and then incorporate with a simple convolution-based subnet to simultaneously learning the local to global image priors and kernel priors in an unsupervised way using the observed LR image only. Extensive experiments have demonstrated that our proposed blind SR method achieves superior SR performance over both supervised and unsupervised state-of-the-art methods in term of quantitative metrics and perceptive quality. Kazuhiro Yamawaki, Xianhua Han |
ICASSP | 2 |
| 2023 | Coarse-To-Fine Pyramid Feature Mining for Wheat Head DetectionabstractAutomatic detection of wheat head attracts extensive attention for efficient and effective wheat farm management and breading study. Benefiting from the powerful learning capability of deep convolution neural networks (DCNNs), recent work have demonstrated the potential and feasibility of the detection automation for wheat head. However, because of the appearance uncertainty of wheat head and the large variability of imaging conditions, detection performance is still needed to be improved for real application. This study exploits a novel coarse-to-fine pyramid feature mining network (CFPFM-Net) for anchor-free wheat head detection. The proposed CFPFM-Net incorporates pyramid context fusion and a U-net-based refining module for coarse-to-fine feature mining, and only simply predict the centerness and the size of wheat head based on the aggregated context in a high resolution permitting accurate detection for small-size wheat head. Moreover, to capture more effective features in various scales and alleviate the gradient vanishing problem, we leverage the auxiliary supervision on several intermediate feature maps in the training phase. Experiments on the Global Wheat Head Detection (GWHD) dataset have demonstrated that the proposed framework achieves superior performance over the existing state-of-the-art methods. Sho Harada, Xianhua Han |
ICIP | 2 |
| 2023 | Coupling Spatial and Channel Transformer for Single Image DerainingabstractSingle image deraining is a fundamental low-level vision task, and has evolved remarkable progress with the deep learning technique. Recently, benefiting from the powerful modeling ability of long-range dependence, transformer as an alternative architecture of the dominant convolutional neural network has demonstrated large margin performance improvement in various high-level vision tasks, and has begun to be applied for low-level vision tasks. The benchmark transformer block captures long dependence via incorporating the self-attention among the spatial points of the learned feature map, and causes heavy computational workload and memory footprint quadratically increased with spatial resolutions, making it impossible to handle high-resolution images. This study proposes a novel spatial and channel coupled Transformer to jointly explore long-range dependence and correlation in both spatial and channel domains, and results in a lightweight deraining transformer model for potentially processing high-resolution images. Yuto Namba, Jiande Sun 0001, Xianhua Han |
ICIP | 3 |
| 2023 | A Lightweight Network for Contextual and Morphological Awareness for Hepatic Vein SegmentationabstractAccurate segmentation of the hepatic vein can improve the precision of liver disease diagnosis and treatment. Since the hepatic venous system is a small target and sparsely distributed, with various and diverse morphology, data labeling is difficult. Therefore, automatic hepatic vein segmentation is extremely challenging. We propose a lightweight contextual and morphological awareness network and design a novel morphology aware module based on attention mechanism and a 3D reconstruction module. The morphology aware module can obtain the slice similarity awareness mapping, which can enhance the continuous area of the hepatic veins in two adjacent slices through attention weighting. The 3D reconstruction module connects the 2D encoder and the 3D decoder to obtain the learning ability of 3D context with a very small amount of parameters. Compared with other SOTA methods, using the proposed method demonstrates an enhancement in the dice coefficient with few parameters on the two datasets. A small number of parameters can reduce hardware requirements and potentially have stronger generalization, which is an advantage in clinical deployment. Guoyu Tong, Huiyan Jiang, Tianyu Shi 0002, Xianhua Han, Yu-Dong Yao |
IEEE J. Biomed. Health Informatics | 4 |
| 2023 | IDH mutation status prediction by a radiomics associated modality attention network
Yutaro Iwamoto, Jingliang Cheng, Guohua Zhao, Xianhua Han, Yen-Wei Chen 0001 |
Vis. Comput. | 7 |
| 2022 | Mixed Transformer U-Net for Medical Image SegmentationabstractThough U-Net has achieved tremendous success in medical image segmentation tasks, it lacks the ability to explicitly model long-range dependencies. Therefore, Vision Transformers have emerged as alternative segmentation structures recently, for their innate ability of capturing long-range correlations through Self-Attention (SA). However, Transformers usually rely on large-scale pre-training and have high computational complexity. Furthermore, SA can only model self-affinities within a single sample, ignoring the potential correlations of the overall dataset. To address these problems, we propose a novel Transformer module named Mixed Transformer Module (MTM) for simultaneous inter- and intra- affinities learning. MTM first calculates self-affinities efficiently through our well-designed Local-Global Gaussian-Weighted Self-Attention (LGG-SA). Then, it mines inter-connections between data samples through External Attention (EA). By using MTM, we construct a U-shaped model named Mixed Transformer U-Net (MT-UNet) for accurate medical image segmentation. We test our method on two different public datasets, and the experimental results show that the proposed method achieves better performance over other state-of-the-art methods. The code is available at: https://github.com/Dootmaan/MT-UNet. Hongyi Wang 0002, Shiao Xie, Lanfen Lin, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Ruofeng Tong 0001 |
ICASSP | 5 |
| 2022 | Unsupervised Generative Network for Blind Hyperspectral Image Super-ResolutionabstractHyperspectral (HS) imaging sacrifices spatial resolution to ensure a high spectral resolution when capturing the detailed spectral signature at each spatial location of the scene. To compensate for this deficiency, fusing low-resolution HS (LR-HS) images with high-resolution RGB (HR-RGB) images to obtain high-resolution HS (HR-HS) images has attracted remarkable attention. Recently, deep learning-based fusion methods in a fully-supervised manner have been proven to make great progress in hyperspectral image super-resolution (HSI-SR) tasks. However, these methods require collecting a large number of training samples and constructing a non-blind prediction model to super-resolve the observations captured under controlled imaging conditions. This study proposes a novel unsupervised generative network (UGN) for learning network parameters using the observed LR-HS, HR-RGB only without the corresponding ground-truth, and designs the spatial and spectral degradation blocks to automatically learn the image degradation operations for constructing an end-to-end blind HSI SR framework. To verify the effectiveness of our proposed method, we con-duct experiments on two benchmark HS image datasets and demonstrate superior performance compared with the super-vised and unsupervised blind/non-blind SoTA methods. Zhe Liu 0039, Xianhua Han, Jiande Sun 0001, Yen-Wei Chen 0001 |
ICIP | 2 |
| 2022 | Generalized Deep Internal Learning for Hyperspectral Image Super ResolutionabstractRecently, deep-learning-based methods have made remarkable progress for reconstructing the high-resolution hyper-spectral (HR-HS) image through automatically learning the inherent priors from images. These methods are basically implemented in a fully-supervised learning manner with a previously prepared large-scale external dataset captured un-der controlled conditions, which would greatly restrict the wide applicability to real scenarios. Therefore, this study proposes a novel generalized deep internal learning to solve the HS image super-resolution (HSI SR) problem. Specifically, we aim to train an image-specific CNN model for an under-studying scene using the extracted triplet samples from the down-sampled LR-HS and HR-RGB images and the original LR-HS image as well as the observed HR-RGB and LR-HS images without the corresponding ground-truth for unsupervised learning. To implement the unsupervised learning, we design the degradation blocks to approximate the spatial and spectral degradation operations, and then transform the learned HR-HS target to the LR-HS and HR-RGB estimations for evaluating the network learning states. To verify the effectiveness of our proposed framework, we conduct extensive experiments on two benchmark HS datasets, and demonstrate that the proposed method achieves favorable performance over the state-of-the-art methods. Zhe Liu 0039, Xianhua Han |
ICIP | 2 |
| 2022 | Hyperspectral Reconstruction Using Auxiliary Rgb Learning From A Snapshot ImageabstractTo solve the low spatial and temporal resolution issue in the conventional hyperspectral (HS) imaging sensors, coded aperture snapshot HS imaging, which encodes the 3D HS image into a 2D compressive snapshot and then adopts computational technique to recover the latent HS image, has attracted remarkable attention in recent year. This study aims to reconstruct the latent HS image with the detail spectral distribution from its compressive snapshot using the deep convolution neural network (DCNN). Due to the ill-posed nature, the HS image reconstruction is a challenge task, and the spectral distortion is unavoidably produced even with the powerful learning capability of the DCNN. To alleviate this limitation, we leverage an auxiliary RGB learning task to reconstruct the corresponding RGB image from the snapshot image in training phase, and incorporate the learned features of the auxiliary task to assist the more difficult reconstruction of the latent HS image. Specifically, we design the DCNN architecture with two branches for both reconstruction learnings of a small number of spectral image (such as RGB) and the full-spectral HS image, and then integrate the intermediate features in the auxiliary RGB branch to the HS reconstruction branch for augment the spectral learning capability. Experimental results demonstrate our proposed method with the auxiliary learning can achieve comparable performance with the state-of-the-art methods while enable the reduction of the model size. Kazuhiro Yamawaki, Yorimoto Kohei, Xianhua Han |
ICIP | 3 |
| 2022 | ScaleFormer: Revisiting the Transformer-based Backbones from a Scale-wise Perspective for Medical Image SegmentationabstractRecently, a variety of vision transformers have been developed as their capability of modeling long-range dependency. In current transformer-based backbones for medical image segmentation, convolutional layers were replaced with pure transformers, or transformers were added to the deepest encoder to learn global context. However, there are mainly two challenges in a scale-wise perspective: (1) intra-scale problem: the existing methods lacked in extracting local-global cues in each scale, which may impact the signal propagation of small objects; (2) inter-scale problem: the existing methods failed to explore distinctive information from multiple scales, which may hinder the representation learning from objects with widely variable size, shape and location. To address these limitations, we propose a novel backbone, namely ScaleFormer, with two appealing designs: (1) A scale-wise intra-scale transformer is designed to couple the CNN-based local features with the transformer-based global cues in each scale, where the row-wise and column-wise global dependencies can be extracted by a lightweight Dual-Axis MSA. (2) A simple and effective spatial-aware inter-scale transformer is designed to interact among consensual regions in multiple scales, which can highlight the cross-scale dependency and resolve the complex scale variations. Experimental results on different benchmarks demonstrate that our Scale-Former outperforms the current state-of-the-art methods. The code is publicly available at: https://github.com/ZJUGiveLab/ScaleFormer. Huimin Huang 0002, Shiao Xie, Lanfen Lin, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Ruofeng Tong 0001 |
IJCAI | 5 |
| 2022 | An Accurate Unsupervised Liver Lesion Detection Method Using Pseudo-lesions
He Li 0042, Yutaro Iwamoto, Xianhua Han, Lanfen Lin, Hongjie Hu, Yen-Wei Chen 0001 |
MICCAI (8) | 3 |
| 2022 | Multi-Scale Channel Transformer Network for Single Image DerainingabstractSingle image deraining is a very challenging task, as it requires not only restoring the spatial details and high contextual structures of the images, but also removing multiple layers of rain with varying degrees of blurring and resolutions. Recently, due to the powerful modeling capability of long-dependency, transformer-based models have manifested superior performance for high-level vision tasks, and have begun to be applied for low-level vision tasks such as various image restoration applications. However, its computational complexity increases quadratically with spatial resolutions, making it impossible to apply it to high-resolution images. In this study, we propose a novel Channel Transformer, which performs self-attention in the channel direction instead of the spatial direction. Specifically, we first incorporate multiple channel transformer blocks into a multi-scale architecture to extract multi-scale contexts and exploit channel long-dependence, and then learn a coarse estimation of the rain-free image. Finally, an original-resolution CNN-based module is employed to refine the coarse estimation via leveraging the previously learned multi-scale contexts. Experiments on several benchmark datasets demonstrate its superiority over the state-of-the-art methods. Yuto Namba, Xianhua Han |
MMAsia | 2 |
| 2022 | Deep Image and Kernel Prior Learning for Blind Super-ResolutionabstractRecently, single image super-resolution (SR) has witnessed significant progress due to the powerful modeling capability of the deep learning networks. However, conventional deep learning-based super-resolution methods predict high-resolution (HR) images under the assumption of ideal degradation model such as the simulated bicubic down-sampling, and then unavoidably deteriorate the SR performance under un-controlled imaging conditions, such as real-world LR images. This study proposes an universal blind SR framework for adaptively and simultaneously predicting the underlying HR image and the counterpart blurring kernel from the observed LR image only. Specifically, we employ an encoder-decoder-based generative network to learn the inherent statistic prior of the HR image from a noise input while adopt a shallow convolution subnet with several stacked layers to estimate the blurring kernel from the observed LR image. Then, a convolution-based degradation module by setting the estimated blurring kernel as its weights is incorporated to obtain the approximated version of the LR image for formulating the loss function. In addition, a pre-trained discriminator is adopted to integrate the perceptual loss for recovering more accurate and natural HR image. We demonstrate the effectiveness of the proposed deep image and kernel prior learning framework using extensive experiments on both synthetic and real images, showing superiority over the state-of-the-art blind SR performance. Kazuhiro Yamawaki, Xianhua Han |
MMAsia | 2 |
| 2022 | Mutual Information-Based Graph Co-Attention Networks for Multimodal Prior-Guided Magnetic Resonance Imaging SegmentationabstractMultimodal magnetic resonance imaging (MRI) provides complementary information about targets, and the segmentation of multimodal MRI is widely used as an essential preprocessing step for initial diagnosis, stage differentiation, and post-treatment efficacy evaluation in clinical situations. For the main modality or each of the modalities, it is important to enhance the visual information by modeling the connection and effectively fusing the features among them. However, the existing methods for multimodal segmentation have a drawback; they coincidentally drop information of individual modality during the fusion process. Recently, graph learning-based methods have been applied in segmentation, and these methods have achieved considerable improvements by modeling the relationships across feature regions and reasoning using global information. In this paper, we propose a graph learning-based approach to efficiently extract modality-specific features and establish regional correspondence effectively among all modalities. In detail, after projecting features into a graph domain and employing graph convolution to propagate information across all regions for learning global modality-specific features, we propose a mutual information-based graph co-attention module to learn the weight coefficients of one bipartite graph constructed by the fully connected graphs having different modalities in the graph domain and by selectively fusing the node features. Based on the deformation diagram between the spatial-graph space and our proposed graph co-attention module, we present a multimodal prior-guided segmentation framework, which uses two strategies for two clinical situations:Modality-Specific Learning StrategyandCo-Modality Learning Strategy. Besides, the improvedCo-Modality Learning Strategyis used with trainable weights in the multi-task loss for the optimization of the proposed framework. We validated our proposed modules and frameworks on two multimodal MRI datasets: our private liver lesion dataset and a public prostate zone dataset. Our experimental results on both datasets prove the superiority of our proposed approaches. Shaocong Mo, Lanfen Lin, Ruofeng Tong 0001, Qingqing Chen 0001, Fang Wang 0030, Hongjie Hu, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 9 |
| 2022 | MTL-ABS3Net: Atlas-Based Semi-Supervised Organ Segmentation Network With Multi-Task Learning for Medical ImagesabstractOrgan segmentation is one of the most important step for various medical image analysis tasks. Recently, semi-supervised learning (SSL) has attracted much attentions by reducing labeling cost. However, most of the existing SSLs neglected the prior shape and position information specialized in the medical images, leading to unsatisfactory localization and non-smooth of objects. In this paper, we propose a novel atlas-based semi-supervised segmentation network with multi-task learning for medical organs, named MTL-ABS3Net, which incorporates the anatomical priors and makes full use of unlabeled data in a self-training and multi-task learning manner. The MTL-ABS3Net consists of two components: an Atlas-Based Semi-Supervised Segmentation Network (ABS3Net) and Reconstruction-Assisted Module (RAM). Specifically, the ABS3Net improves the existing SSLs by utilizing atlas prior, which generates credible pseudo labels in a self-training manner; while the RAM further assists the segmentation network by capturing the anatomical structures from the original images in a multi-task learning manner. Better reconstruction quality is achieved by using MS-SSIM loss function, which further improves the segmentation accuracy. Experimental results from the liver and spleen datasets demonstrated that the performance of our method was significantly improved compared to existing state-of-the-art methods. Huimin Huang 0002, Qingqing Chen 0001, Lanfen Lin, Qiaowei Zhang, Yutaro Iwamoto, Xianhua Han, Akira Furukawa, Shuzo Kanasaki, Yen-Wei Chen 0001, Ruofeng Tong 0001, Hongjie Hu |
IEEE J. Biomed. Health Informatics | 7 |
| 2022 | Deep Self-Supervised Hyperspectral Image ReconstructionabstractReconstructing a high-resolution hyperspectral (HR-HS) image via merging a low-resolution hyperspectral (LR-HS) image and a high-resolution RGB (HR-RGB) image has become a hot research topic, and can greatly benefit for different subsequent high-level vision tasks. Recently, deep learning–based approaches have evolved for HS image reconstruction and validated impressive performance. However, to learn a good reconstruction model in the deep learning–based methods, it is mandatory to previously collect large-scale training triplets consisting of the LR-HS, HR-RGB, and HR-HS images, which is difficult to be collected in real applications. This study proposes a deep self-supervised HS image reconstruction framework (DSSH), which does not have to depend on any handcrafted prior and previously collected training triplets at all. The proposed DSSH method leverages the designed network architecture itself for capturing the prior of the underlying structure in the latent HR-HS image and employs the observed LR-HS and HR-RGB images only for network parameter learning. Experiments on two benchmark HS image datasets validated that the proposed DSSH method manifests very impressive reconstruction performance, and is even better than some state-of-the-art supervised learning approaches. Zhe Liu 0039, Xianhua Han |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2022 | Hyperspectral Image Reconstruction Using Multi-scale Fusion LearningabstractHyperspectral imaging is a promising imaging modality that simultaneously captures several images for the same scene on narrow spectral bands, and it has made considerable progress in different fields, such as agriculture, astronomy, and surveillance. However, the existing hyperspectral (HS) cameras sacrifice the spatial resolution for providing the detail spectral distribution of the imaged scene, which leads to low-resolution (LR) HS images compared with the common red-green-blue (RGB) images. Generating a high-resolution HS (HR-HS) image via fusing an observed LR-HS image with the corresponding HR-RGB image has been actively studied. Existing methods for this fusing task generally investigate hand-crafted priors to model the inherent structure of the latent HR-HS image, and they employ optimization approaches for solving it. However, proper priors for different scenes can possibly be diverse, and to figure it out for a specific scene is difficult. This study investigates a deep convolutional neural network (DCNN)-based method for automatic prior learning, and it proposes a novel fusion DCNN model with multi-scale spatial and spectral learning for effectively merging an HR-RGB and LR-HS images. Specifically, we construct an U-shape network architecture for gradually reducing the feature sizes of the HR-RGB image (Encoder-side) and increasing the feature sizes of the LR-HS image (Decoder-side), and we fuse the HR spatial structure and the detail spectral attribute in multiple scales for tackling the large resolution difference in spatial domain of the observed HR-RGB and LR-HS images. Then, we employ multi-level cost functions for the proposed multi-scale learning network to alleviate the gradient vanish problem in long-propagation procedure. In addition, for further improving the reconstruction performance of the HR-HS image, we refine the predicted HR-HS image using an alternating back-projection method for minimizing the reconstruction errors of the observed LR-HS and HR-RGB images. Experiments on three benchmark HS image datasets demonstrate the superiority of the proposed method in both quantitative values and visual qualities. Xianhua Han, Yinqiang Zheng, Yen-Wei Chen 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2021 | Multi-scale Residual Aggregation Deraining Network with Spatial Context-aware Pooling and Activation
Kohei Yamamichi, Xianhua Han |
BMVC | 2 |
| 2021 | Graph-Based Pyramid Global Context Reasoning With a Saliency- Aware Projection for Covid-19 Lung Infections SegmentationabstractCoronavirus Disease 2019 (COVID-19) has rapidly spread in 2020, emerging a mass of studies for lung infection segmentation from CT images. Though many methods have been proposed for this issue, it is a challenging task because of infections of various size appearing in different lobe zones. To tackle these issues, we propose a Graph-based Pyramid Global Context Reasoning (Graph-PGCR) module, which is capable of modeling long-range dependencies among disjoint infections as well as adapt size variation. We first incorporate graph convolution to exploit long-term contextual information from multiple lobe zones. Different from previous average pooling or maximum object probability, we propose a saliency-aware projection mechanism to pick up infection-related pixels as a set of graph nodes. After graph reasoning, the relation-aware features are reversed back to the original coordinate space for the down-stream tasks. We further construct multiple graphs with different sampling rates to handle the size variation problem. To this end, distinct multi-scale long-range contextual patterns can be captured. Our Graph- PGCR module is plug-and-play, which can be integrated into any architecture to improve its performance. Experiments demonstrated that the proposed method consistently boost the performance of state-of-the-art backbone architectures on both of public and our private COVID-19 datasets. Huimin Huang 0002, Lanfen Lin, Xiongwei Mao, Xiaohan Qian, Zhiyi Peng, Jianying Zhou 0006, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Ruofeng Tong 0001 |
ICASSP | 10 |
| 2021 | Deep Blind Un-Supervised Learning Network for Single Image Super ResolutionabstractDeep learning based methods have recently made significant progress in image super resolution (SR) field and lead to great performance gain in terms of both effectiveness and efficiency. Most of the current methods have struggled to design more complicated and deeper network architectures and aimed to learn a good LR-to-HR mapping with the previously prepared training sample pairs under a fixed degradation model (Blurring and down-sampling operations) such as bicubic dawn-sampling. However, these methods are generally implemented in a fully-supervised way with largescale training dataset, and are hardly generalized to most real scenarios with unknown and complicated degradation model. This study proposes a blind un-supervised learning network for automatically estimating the degradation operations in single SR problem, where the blurring kernel (operation) is unknown. Motivated by the considerable possessed image priors in the network architectures, we construct a generative network for simultaneously learning the inherent priors of the latent high resolution (HR) image and the degradation operations with the under-studying low-resolution (LR) observation only. Specifically, we exploit a general depth-wise convolutional layer for both approximating a special degradation and automatically learning any complicated blurring kernel in a general SR framework, and propose an end-to-end HR image learning network from its LR observation. Experimental results on two benchmark datasets validate that our proposed method achieve promising performance under the unknown degradation model. Kazuhiro Yamawaki, Xianhua Han |
ICIP | 2 |
| 2021 | 3D Graph-S2Net: Shape-Aware Self-ensembling Network for Semi-supervised Segmentation with Bilateral Graph Convolution
Huimin Huang 0002, Lanfen Lin, Hongjie Hu, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Ruofeng Tong 0001 |
MICCAI (2) | 6 |
| 2021 | Patch-Free 3D Medical Image Segmentation Driven by Super-Resolution Technique and Self-Supervised Guidance
Hongyi Wang 0002, Lanfen Lin, Hongjie Hu, Qingqing Chen 0001, Yinhao Li 0002, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Ruofeng Tong 0001 |
MICCAI (1) | 7 |
| 2021 | A Tensor Sparse Representation-Based CBMIR System for Computer-Aided Diagnosis of Focal Liver Lesions and its Pilot TrialabstractClinicians refer to diagnosed medical cases in order to make correct diagnosis and take appropriate treatments, due to the complexity of focal liver lesions. It's a heavy burden, however, for medical doctors to find out similar and meaningful cases from the accumulated extreme large medical datasets. Content based medical image retrieval (CBMIR) that searches for similar images in a large database has been attracting increasing research interest recently. A CBMIR system provides doctors the diagnosed cases to improve the diagnosis accuracy and confidence. This paper proposed a tensor sparse representation method to extract temporal and spatial features of multi-phase CT images, so as to provide doctors medical cases more relevant to the query one. The proposed tensor sparse representation method is applied to the retrieval of focal liver lesions (FLLs). Experiments show that the proposed method achieved better retrieval performance than conventional methods. Pilot trial was conducted and results show that diagnosis accuracy and confidence was improved significantly by the developed CBMIR system based on the proposed method. Jian Wang 0004, Xianhua Han, Lanfen Lin, Hongjie Hu, Yen-Wei Chen 0001 |
ICMR | 2 |
| 2021 | A Cascade of 2.5D CNN and Bidirectional CLSTM Network for Mitotic Cell Detection in 4D Microscopy ImageabstractMitosis detection is one of the challenging steps in biomedical imaging research, which can be used to observe the cell behavior. Most of the already existing methods that are applied in detecting mitosis usually contain many nonmitotic events (normal cell and background) in the result (false positives, FPs). In order to address such a problem, in this study, we propose to apply 2.5-dimensional (2.5D) networks called CasDetNet_CLSTM, which can accurately detect mitotic events in 4D microscopic images. This CasDetNet_CLSTM involves a 2.5D faster region-based convolutional neural network (Faster R-CNN) as the first network, and a convolutional long short-term memory (CLSTM) network as the second network. The first network is used to select candidate cells using the information from nearby slices, whereas the second network uses temporal information to eliminate FPs and refine the result of the first network. Our experiment shows that the precision and recall of our networks yield better results than those of other state-of-the-art methods. Titinunt Kitrungrotsakul, Xianhua Han, Yutaro Iwamoto, Satoko Takemoto, Hideo Yokota, Sari Ipponjima, Tomomi Nemoto, Wei Xiong 0001, Yen-Wei Chen 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2021 | Medical Image Segmentation With Deep Atlas PriorabstractOrgan segmentation from medical images is one of the most important pre-processing steps in computer-aided diagnosis, but it is a challenging task because of limited annotated data, low-contrast and non-homogenous textures. Compared with natural images, organs in the medical images have obvious anatomical prior knowledge (e.g., organ shape and position), which can be used to improve the segmentation accuracy. In this paper, we propose a novel segmentation framework which integrates the medical image anatomical prior through loss into the deep learning models. The proposed prior loss function is based on probabilistic atlas, which is called as deep atlas prior (DAP). It includes prior location and shape information of organs, which are important prior information for accurate organ segmentation. Further, we combine the proposed deep atlas prior loss with the conventional likelihood losses such as Dice loss and focal loss into an adaptive Bayesian loss in a Bayesian framework, which consists of a prior and a likelihood. The adaptive Bayesian loss dynamically adjusts the ratio of the DAP loss and the likelihood loss in the training epoch for better learning. The proposed loss function is universal and can be combined with a wide variety of existing deep segmentation models to further enhance their performance. We verify the significance of our proposed framework with some state-of-the-art models, including fully-supervised and semi-supervised segmentation models on a public dataset (ISBI LiTS 2017 Challenge) for liver segmentation and a private dataset for spleen segmentation. Huimin Huang 0002, Lanfen Lin, Hongjie Hu, Qiaowei Zhang, Qingqing Chen 0001, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Ruofeng Tong 0001 |
IEEE Trans. Medical Imaging | 9 |
| 2020 | MCGKT-Net: Multi-level Context Gating Knowledge Transfer Network for Single Image Deraining
Kohei Yamamichi, Xianhua Han |
ACCV (2) | 2 |
| 2020 | UNet 3+: A Full-Scale Connected UNet for Medical Image SegmentationabstractRecently, a growing interest has been seen in deep learning-based semantic segmentation. UNet, which is one of deep learning networks with an encoder-decoder architecture, is widely used in medical image segmentation. Combining multi-scale features is one of important factors for accurate segmentation. UNet++ was developed as a modified Unet by designing an architecture with nested and dense skip connections. However, it does not explore sufficient information from full scales and there is still a large room for improvement. In this paper, we propose a novel UNet 3+, which takes advantage of full-scale skip connections and deep supervisions. The full-scale skip connections incorporate low-level details with high-level semantics from feature maps in different scales; while the deep supervision learns hierarchical representations from the full-scale aggregated feature maps. The proposed method is especially benefiting for organs that appear at varying scales. In addition to accuracy improvements, the proposed UNet 3+ can reduce the network parameters to improve the computation efficiency. We further propose a hybrid loss function and devise a classification-guided module to enhance the organ boundary and reduce the over-segmentation in a non-organ image, yielding more accurate segmentation results. The effectiveness of the proposed method is demonstrated on two datasets. The code is available at: github.com/ZJUGiveLab/UNet-Version. Huimin Huang 0002, Lanfen Lin, Ruofeng Tong 0001, Hongjie Hu, Qiaowei Zhang, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Jian Wu 0001 |
ICASSP | 7 |
| 2020 | Deep Residual Attention Network for Hyperspectral Image ReconstructionabstractCoded aperture snapshot spectral imaging (CASSI) captures a full frame spectral image as a single compressive image and is mandatory to reconstruct the underlying hyperspectral image (HSI) from the snapshot as the post-processing, which is a challenge inverse problem due to its ill-posed nature. Existing methods for HSI reconstruction from a snapshot usually employs optimization for solving the formulated image degradation model regularized with the empirically designed priors, and still cannot achieve enough reconstruction accuracy for real HS image analysis systems. Motivated by the recent advances of deep learning for different inverse problems, deep learning based HSI reconstruction method has attracted a lot of attention and can boost the reconstruction performance. This study proposes a novel deep convolutional neural network (DCNN) based framework for effectively learning the spatial structure and spectral attribute in the underlying HSI with the reciprocal spatial and spectral modules. Further, to adaptively leverage the useful learned feature for better HSI image reconstruction, we integrate residual attention modules into our DCNN via exploring both spatial and spectral attention maps. Experimental results on two benchmark HSI datasets show that our method outperforms state-of-the-art methods in both quantitative values and visual effects. Yorimoto Kohei, Xianhua Han |
ICPR | 2 |
| 2020 | Multimodal Priors Guided Segmentation of Liver Lesions in MRI Using Mutual Information Based Graph Co-Attention Networks
Shaocong Mo, Lanfen Lin, Ruofeng Tong 0001, Qingqing Chen 0001, Fang Wang 0030, Hongjie Hu, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001 |
MICCAI (4) | 9 |
| 2020 | Image super-resolution based on two-level residual learning CNN
Min Gao 0001, Xianhua Han, Jing Li 0046, Huaxiang Zhang 0001, Jiande Sun 0001 |
Multim. Tools Appl. | 2 |
| 2020 | An end-to-end CNN and LSTM network with 3D anchors for mitotic cell detection in 4D microscopic images and its parallel implementation on multiple GPUs
Titinunt Kitrungrotsakul, Xianhua Han, Yutaro Iwamoto, Satoko Takemoto, Hideo Yokota, Sari Ipponjima, Tomomi Nemoto, Wei Xiong 0001, Yen-Wei Chen 0001 |
Neural Comput. Appl. | 2 |
| 2020 | Tensor-based sparse representations of multi-phase medical images for classification of focal liver lesions
Jian Wang 0004, Jing Li 0046, Xianhua Han, Lanfen Lin, Hongjie Hu, Qingqing Chen 0001, Yutaro Iwamoto, Yen-Wei Chen 0001 |
Pattern Recognit. Lett. | 3 |
| 2020 | Semi-Supervised Learning for Semantic Segmentation of Emphysema With Partial AnnotationsabstractSegmentation and quantification of each subtype of emphysema is helpful to monitor chronic obstructive pulmonary disease. Due to the nature of emphysema (diffuse pulmonary disease), it is very difficult for experts to allocate semantic labels to every pixel in the CT images. In practice, partially annotating is a better choice for the radiologists to reduce their workloads. In this paper, we propose a new end-to-end trainable semi-supervised framework for semantic segmentation of emphysema with partial annotations, in which a segmentation network is trained from both annotated and unannotated areas. In addition, we present a new loss function, referred to as Fisher loss, to enhance the discriminative power of the model and successfully integrate it into our proposed framework. Our experimental results show that the proposed methods have superior performance over the baseline supervised approach (trained with only annotated areas) and outperform the state-of-the-art methods for emphysema segmentation. Liying Peng, Lanfen Lin, Hongjie Hu, Yue Zhang 0042, Huali Li, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001 |
IEEE J. Biomed. Health Informatics | 7 |
| 2020 | Hyperspectral Reconstruction with Redundant Camera Spectral Sensitivity FunctionsabstractHigh-resolution hyperspectral (HS) reconstruction has recently achieved significantly progress, among which the method based on the fusion of the RGB and HS images of the same scene can greatly improve the reconstruction performance compared with those based on the individually spectral or spatial enhancement. It is well known that the HS image is obtained only via the costly hypersoectral sensor, whereas the RGB images can be provided by low-price RGB cameras and the spectral sensitivity (SS) functions of RGB cameras are usually different. Thus, this study proposes a HS reconstruction, which fuses merely two RGB images with redundant spectral responses. In this work, we design a new RGB camera via shifting the SS of an existed RGB camera, which can provide similar strength of spectral response with different spectral centers of SS, and fuse the new achieved color image with an existed RGB image by a deep ResNet. Experiments validate that fusion of two existed RGB images can provide impressive HS reconstruction performance and further improvement can be achieved by integrating the color image of the simulated SS with the RGB image. Xianhua Han, Yinqiang Zheng, Jiande Sun 0001, Yen-Wei Chen 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2019 | A Cascade of CNN and LSTM Network with 3D Anchors for Mitotic Cell Detection in 4D Microscopic ImageabstractMitotic event detection is a fundamental step in investigating of cell behaviors. The event can be used to analyze various diseases, but most mitotic event detections performed previously focused only on two-dimensional (2D) images with time information. Owing to the complex background (normal cells) and mitotic event orientations, the 2D detection methods yield many false positive and false negative results. To solve this problem, we proposed a 2.5 dimensional (2.5D) cascaded end-to-end network combined with 3D anchors for accurate detection of mitotic events in 4D microscopic images. Our proposed network uses a convolutional long short-term memory to handle issues relating to time sequence; this helps to improve the detection accuracy (reduction of false positives). Furthermore, it uses 3D anchors to capture volume information used to address the orientation problem (reduction of false negatives). The experimental results show that the proposed method can achieve higher precision and recall compared with state-of-the-art methods. Titinunt Kitrungrotsakul, Yutaro Iwamoto, Xianhua Han, Satoko Takemoto, Hideo Yokota, Sari Ipponjima, Tomomi Nemoto, Wei Xiong 0001, Yen-Wei Chen 0001 |
ICASSP | 3 |
| 2019 | A Dual-Attention Dilated Residual Network for Liver Lesion Classification and Localization on CT ImagesabstractAutomatic liver lesion classification on computed tomography images is of great importance to early cancer diagnosis and remains a challenging task. State-of-the-art liver lesion classification algorithms are currently based on manually selected regions of interest (ROIs) or automatically detected ROIs. However, liver lesions usually vary in size and shape, which makes the ROI selection process labor-intensive and also poses an obstacle to automatic lesion detection. In this paper, we propose a dual-attention dilated residual network (DADRN) as a potential solution to lesion classification task without manual ROI selection or automatic lesion detection. We incorporated a novel dual-attention module in order to capture the non-local feature dependencies and help the deep neural network focus on the lesion area by enlarging the difference between the lesion area and nonlesion area. To the best of our knowledge, we are the first to employ the self-attention mechanism to address liver lesion classification task. In addition, the well-trained DADRN can be used for weakly-supervised lesion localization without any architectural change or retraining. Experiment results show that DADRN could achieve a lesion classification accuracy comparable to that of the state-of-the-art ROI-based method and outperformed state-of-the-art attention-based approaches in both liver lesion classification and localization tasks. Xiao Chen 0016, Jian Wu 0001, Lanfen Lin, Hongjie Hu, Qiaowei Zhang, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Ruofeng Tong 0001 |
ICIP | 8 |
| 2019 | Multi-Stream Scale-Insensitive Convolutional and Recurrent Neural Networks for Liver Tumor Detection in Dynamic Ct ImagesabstractConvolutional neural networks (CNNs) have achieved great success in numerous challenging vision tasks, and have great potential for object detection in natural images. Compared with the natural images, medical images exhibit some unique characteristics. Therefore, substantial challenges still remain in this field. The first challenge is to develop a method for effectively distilling enhancement patterns from the dynamic CT images. Moreover, since tumor sizes vary greatly and small lesions are important for early liver tumor detection, lesion detection with a widely variable scale is another challenge. In this paper, we propose a multi-stream scale-insensitive convolutional and recurrent neural network (MSCR) for liver tumor detection. Specifically, we propose the use of grouped convolutional long short-term memory (GCLSTM) to extract enhancement patterns, which is developed as a plug-and-play module. Experiments show that the MSCR framework exhibits superior performance over state-of-the-art approaches, achieving an average precision of 77.06% for detection of focal liver lesions. We have released the code of MSCR in1. Ruofeng Tong 0001, Jian Wu 0001, Lanfen Lin, Xiao Chen 0016, Hongjie Hu, Qiaowei Zhang, Qingqing Chen 0001, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001 |
ICIP | 10 |
| 2019 | Semi-supervised Segmentation of Liver Using Adversarial Learning with Deep Atlas Prior
Lanfen Lin, Hongjie Hu, Qiaowei Zhang, Qingqing Chen 0001, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Ruofeng Tong 0001, Jian Wu 0001 |
MICCAI (6) | 7 |
| 2019 | Classification and Quantification of Emphysema Using a Multi-Scale Residual NetworkabstractAutomated tissue classification is an essential step for quantitative analysis and treatment of emphysema. Although many studies have been conducted in this area, there still remain two major challenges. First, different emphysematous tissue appears in different scales, which we call "inter-class variations." Second, the intensities of CT images acquired from different patients, scanners or scanning protocols may vary, which we call "intra-class variations". In this paper, we present a novel multi-scale residual network with two channels of raw CT image and its differential excitation component. We incorporate multi-scale information into our networks to address the challenge of inter-class variations. In addition to the conventional raw CT image, we use its differential excitation component as a pair of inputs to handle intra-class variations. Experimental results show that our approach has superior performance over the state-of-the- art methods, achieving a classification accuracy of 93.74% on our original emphysema database. Based on the classification results, we also perform the quantitative analysis of emphysema in 50 subjects by correlating the quantitative results (the area percentage of each class) with pulmonary functions. We show that centrilobular emphysema (CLE) and panlobular emphysema (PLE) have strong correlation with the pulmonary functions and the sum of CLE and PLE can be used as a new and accurate measure of emphysema severity instead of the conventional measure (sum of all subtypes of emphysema). The correlations between the new measure and various pulmonary functions are up to |r| = 0.922 (r is correlation coefficient). Liying Peng, Yen-Wei Chen 0001, Lanfen Lin, Hongjie Hu, Huali Li, Qingqing Chen 0001, Xiaoli Ling, Xianhua Han, Yutaro Iwamoto |
IEEE J. Biomed. Health Informatics | 9 |
| 2019 | Adaptive Semi-Supervised Feature Selection for Cross-Modal RetrievalabstractIn order to exploit the abundant potential information of the unlabeled data and contribute to analyzing the correlation among heterogeneous data, we propose the semi-supervised model named adaptive semi-supervised feature selection for cross-modal retrieval. First, we utilize the semantic regression to strengthen the neighboring relationship between the data with the same semantic. And the correlation between heterogeneous data can be optimized via keeping the pairwise closeness when learning the common latent space. Second, we adopt the graph-based constraint to predict accurate labels for unlabeled data, and it can also keep the geometric structure consistency between the label space and the feature space of heterogeneous data in the common latent space. Finally, an efficient joint optimization algorithm is proposed to update the mapping matrices and the label matrix for unlabeled data simultaneously and iteratively. It makes samples from different classes to be far apart, while the samples from same class lie as close as possible. Meanwhile, the l2,1-norm constraint is used for feature selection and outlier reduction when the mapping matrices are learned. In addition, we propose learning different mapping matrices corresponding to different sub-tasks to emphasize the semantic and structural information of query data. Experiment results on three datasets demonstrate that our method performs better than the state-of-the-art methods. En Yu, Jiande Sun 0001, Jing Li 0046, Xiaojun Chang, Xianhua Han, Alex Hauptmann 0001 |
IEEE Trans. Multim. | 5 |
| 2018 | SSF-CNN: Spatial and Spectral Fusion with CNN for Hyperspectral Image Super-ResolutionabstractFusing a low-resolution hyperspectral image with the corresponding high-resolution RGB image to obtain a high-resolution hyperspectral image is usually solved as an optimization problem with prior-knowledge such as sparsity representation and spectral physical properties as constraints, which have limited applicability. Deep convolutional neural network extracts more comprehensive features and is proved to be effective in upsampling RGB images. However, directly applying CNNs to upsample either the spatial or spectral dimension alone may not produce pleasing results due to the neglect of complementary information from both low resolution hyper spectral and high resolution RGB images. This paper proposes two types of novel CNN architectures to take advantages of spatial and spectral fusion for hyperspectral image superresolution. Experiment results on benchmark datasets validate that the proposed spatial and spectral fusion CNNs outperforms the state-of-the-art methods and baseline CNN architectures in both quantitative values and visual qualities. Xianhua Han, Boxin Shi, Yinqiang Zheng |
ICIP | 1 |
| 2018 | Classification of Pulmonary Emphysema in CT Images Based on Multi-Scale Deep Convolutional Neural NetworksabstractIn this work, we aim at classifying emphysema in computed tomography (CT) images of lungs. Most previous works are limited to extracting low-level features or mid-level features without enough high-level information. Moreover, these approaches do not take the characteristics (scales) of different emphysema into account, which are crucial for feature extraction. In contrast to previous works, we propose a novel deep learning method based on multiscale deep convolutional neural networks. There are three contributions for this paper. First, we propose to use a base residual network with 20 layers to extract more high-level information. To the best of our knowledge, this is the first deep learning method for classification of emphysema. Second, we incorporate multi-scale information into our deep neural networks so as to take full consideration of the characteristics of different emphysema. Finally, we established a high-quality emphysema dataset which contains 91 high-resolution computed tomography (HRCT) volumes, annotated manually by two experienced radiologists and checked by one experienced chest radiologist. A 92.68% classification accuracy is achieved on this dataset. The results show that (1) the multi-scale method is highly effective in comparison to the single scale setting; (2) the proposed approach is superior to the state-of-the-art techniques. Liying Peng, Lanfen Lin, Hongjie Hu, Huali Li, Xiaoli Ling, Xianhua Han, Yutaro Iwamoto, Yen-Wei Chen 0001 |
ICIP | 7 |
| 2018 | Residual HSRCNN: Residual Hyper-Spectral Reconstruction CNN from an RGB ImageabstractHyper-spectral imaging has great potential for understanding the characteristics of different materials in many applications ranging from remote sensing to medical imaging. However, due to various hardware limitations, only low-resolution hyper-spectral and high-resolution multi-spectral or RGB images can be captured at video rate. This study aims to generate a hyper-spectral image via enhancing spectral resolution of an RGB image, which might be easily obtained by a commodity camera. Motivated by the success of deep convolutional neural network (DCNN) for spatial resolution enhancement of natural images, we explore a spectral reconstruction CNN for spectral super-resolution with an available RGB image, which predicts the high-frequency content of the fine spectral wavelength in narrow band interval. Since the lost high-frequency content can not be perfectly recovered, by leveraging on the baseline CNN, we further propose a novel residual hyper-spectral reconstruction CNN framework to estimate the non-recovered high-frequency content (Residual) from the output of the baseline CNN. Experiments on benchmark hyper-spectral datasets validate that the proposed method achieves promising performances compared with the existing state-of-the-art methods. Xianhua Han, Boxin Shi, Yinqiang Zheng |
ICPR | 1 |
| 2018 | Comprehensive Study of Multiple CNNs Fusion for Fine-Grained Dog Breed CategorizationabstractFine-grained visual categorization aims to distinguish objects in subordinate classes instead of basic class, and is a challenge visual task due to the high correlation between subordinated classes and large intra-class variation (e.g. different object poses). Although, deep convolutional neural network (DCNN) has brought dramatic success on generic object classification, detection and segmentation with the availability of the large-scale training samples, direct application of DCNN on fine-grained visual categorization, where only decades or at most hundreds of training samples for each subordinate class are available in most public finegrained image datasets, cannot lead to satisfactory classification results due to small number of training samples. This study explores the transfer learning strategy for finegrained dog breed categorization based on the learned CNN models with the large-scale image dataset: ImageNet, and prove promising performance with two DCNN models: AlexNet and VGG-16. Furthermore, we argue that different DCNN architecture may extract the representation of different image aspects due to the previously defined CNN kernel sizes, number and various operations in the model learning procedure, and thus result in different performance for visual categorization. This study proposes to fusion multiple CNN architectures for combining different aspect representations to give more accurate performance. We compressively study the fusion of different layers such as Fc6 and Fc7 in AlexNet and VGG-16, and manifest 2.88% improvement of the fusion architecture over the best performance of the only one DCNN model: VGG-16 from 81.2% to 84.08%. Minori Uno, Xianhua Han, Yen-Wei Chen 0001 |
ISM | 2 |
| 2018 | Combining Convolutional and Recurrent Neural Networks for Classification of Focal Liver Lesions in Multi-phase CT Images
Lanfen Lin, Hongjie Hu, Qiaowei Zhang, Qingqing Chen 0001, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001 |
MICCAI (2) | 7 |
| 2018 | Residual Convolutional Neural Networks with Global and Local Pathways for Classification of Focal Liver Lesions
Lanfen Lin, Hongjie Hu, Qiaowei Zhang, Qingqing Chen 0001, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001 |
PRICAI (1) | 7 |
| 2018 | Generic and Specific Impressions Estimation and Their Application to KANSEI-Based Clothing Fabric Image RetrievalabstractCurrent image retrieval techniques are mainly based on text or visual contents. However, both text-based and contents-based methods lack the capability of utilizing human intuition and KANSEI (impression). In this paper, we proposed an impression-based image retrieval method in order to realize the image retrieval according to our impression presented by impression keywords. We first propose a generic and specific impressions estimation method based on machine learning and then apply it to impression-based clothing fabric image retrieval. We use a semantic differential (SD) method to measure the user’s impressions such as brightness and warmth while they view a cloth fabric image. We also extract both global and local features of cloth fabric images such as color and texture using computer vision techniques. Then we use support vector regression to model the mapping functions between the generic impression (or specific impression) and image features. The learnt mapping functions are used to estimate the generic and specific impressions of cloth fabric images. The retrieval is done by comparing the query impression with the estimated impression of images in the database. Yen-Wei Chen 0001, Xinyin Huang, Dingye Chen, Xianhua Han |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2018 | Self-Similarity Constrained Sparse Representation for Hyperspectral Image Super-ResolutionabstractFusing a low-resolution hyperspectral image with the corresponding high-resolution multispectral image to obtain a high-resolution hyperspectral image is an important technique for capturing comprehensive scene information in both spatial and spectral domains. Existing approaches adopt sparsity promoting strategy, and encode the spectral information of each pixel independently, which results in noisy sparse representation. We propose a novel hyperspectral image super-resolution method via a self-similarity constrained sparse representation. We explore the similar patch structures across the whole image and the pixels with close appearance in local regions to create globalstructure groups and local-spectral super-pixels. By forcing the similarity of the sparse representations for pixels belonging to the same group and super-pixel, we alleviate the effect of the outliers in the learned sparse coding. Experiment results on benchmark datasets validate that the proposed method outperforms the stateof- the-art methods in both quantitative metrics and visual effect. Xianhua Han, Boxin Shi, Yinqiang Zheng |
IEEE Trans. Image Process. | 1 |
| 2017 | Joint weber-based rotation invariant uniform local ternary pattern for classification of pulmonary emphysema in CT imagesabstractIn this paper, we present a novel image representation approach for classifying emphysema in computed tomography (CT) images of the lung. Our proposed method extends rotation invariant uniform local binary pattern (RIULBP) and local ternary pattern (LTP), which are extensively used in a variety of computer vision applications, into rotation invariant uniform local ternary pattern (RIULTP) with a human perception principle: Weber's law. In addition, by integrating the upper pattern and the lower pattern of the Weber-based RIULTP (WRIULTP), we further put forward the joint Weber-based rotation invariant uniform local ternary pattern (JWRIULTP), which allows for a much richer representation and also takes the comprehensive information of the image into account. The proposed methods are tested on the Outex database (texture database) and the Bruijne and Srensen database (emphysema database). The results show the superiority of the proposed approaches to the state-of-the-art techniques for emphysema classification including rotation invariant local binary pattern (RILBP) and texton-based approach. Liying Peng, Lanfen Lin, Hongjie Hu, Xiaoli Ling, Xianhua Han, Yen-Wei Chen 0001 |
ICIP | 6 |
| 2017 | Hyper-spectral Image Super-resolution Using Non-negative Spectral Representation with Data-Guided SparsityabstractHyperspectral imaging has great potential for understanding the characteristics of different materials in many applications ranging from remote sensing to medical imaging. However, due to various hardware limitations, only low-resolution hyperspectral and high-resolution multi-spectral images can be available using existing imaging techniques. This study aims to generate a high-resolution hyperspectral image via fusion of the available LR-HS and HR-MS images. We propose a novel hyperspectral image superresolution method via non-negative sparse representation of reflectance spectral with adaptive sparsity constraint. By analyzing local content similarity of a focused pixel in the available high-resolution multi-spectral image, which can measure pixel material purity according to surrounding pixels, we generate a sparsity map for guiding non-negative sparse coding optimization procedure of the spectral representation called non-negative spectral representation with data-guided sparsity. Since the proposed method adaptively adjust the sparsity in the spectral representation based on the local content of the available high-resolution multi-spectral image, it can produce more robust spectral representation for recovering the target high-resolution hyper-spectral image. Comprehensive experiments on two public hyperspectral datasets validate that the proposed method achieves promising performances compared with the existing state of the art methods. Xianhua Han, Jan Wang, Boxin Shi, Yinqiang Zheng, Yen-Wei Chen 0001 |
ISM | 1 |
| 2017 | Tensor Sparse Representation of Temporal Features for Content-Based Retrieval of Focal Liver Lesions Using Multi-phase Medical ImagesabstractContent Based Image Retrieval (CBIR) systems that search similar images in a large database are attracting more and more research interests recently, and have been applied to medical image characterization for expert's experience sharing. One challenging task in CBIR is how to extract features for effective image representation. Therein sparse coding technique has been proven to be an effective way to learn inherent structure features for image analysis. However, it is necessary to first vectorize the 2- or 3-dimensional spatial structure for analysis with sparse coding, and then destroy the spatial relation of nearby voxels. In this study, we propose a multilinear sparse coding method to learn features from multi-dimensional medical images. We regard high dimensional local structures as tensors and propose a K-CP (CANDECOMP/PARAFAC) algorithm to learn a tensor dictionary in an iterative way. With the learned tensor dictionary, sparse coefficients of tensor local structures are calculated by multilinear orthogonal matching pursuit (MOMP) algorithm, which is an extended multilinear version of the conventional linear OMP. The proposed multilinear sparse coding method is prospected to be more efficient and effective for inherent feature extraction compared with conventional linear methods. The proposed method is applied to a CBIR system for retrieval of focal liver lesions (FLLs) using a medical database consisting of contrast-enhanced multi-phase computer-tomography (CT) images. Experiments show that the constructed CBIR with multilinear sparse coding method can achieve promising retrieval performance. Jian Wang 0004, Xianhua Han, Lanfen Lin, Hongjie Hu, Chongwu Jin, Yen-Wei Chen 0001 |
ISM | 2 |
| 2017 | HEp-2 staining pattern recognition using stacked fisher network for encoding weber local descriptor
Xianhua Han, Yen-Wei Chen 0001 |
Pattern Recognit. | 1 |
| 2016 | Bag of temporal co-occurrence words for retrieval of focal liver lesions using 3D multiphase contrast-enhanced CT imagesabstractComputer-aided diagnosis (CAD) systems have been verified to have the potential to assist radiologists in clinical diagnosis to detect and characterize focal liver lesions (FLLs) based on single- or multiphase contrast-enhanced computed tomography (CT) images. Features extracted from multiphase contrast-enhanced CT images carry more important diagnostic information i.e. enhancement pattern and demonstrate much stronger discriminative ability compared to those of single-phase CT images. In this paper, we propose a new method for multiphase image feature generation called the bag of temporal co-occurrence words (BoTCoW). A temporal co-occurrence image connecting intensity from multiphase images is constructed. Then the bag of visual word (BoVW) model is employed on the temporal co-occurrence images to extract temporal features. The proposed method effectively captures temporal enhancement information and demonstrates the distribution of the evolution patterns. The effectiveness of this method is validated in a retrieval system using 132 FLLs with confirmed pathology type. The preliminary results show that the proposed BoTCoW method outperforms the previously proposed temporal features and multiphase features based on the BoVW model. Lanfen Lin, Hongjie Hu, Yitao Liu, Jian Wang 0004, Xianhua Han, Yen-Wei Chen 0001 |
ICPR | 7 |
| 2016 | A novel and fast connected component count algorithm based on graph theoryabstractA fast component-counting algorithm is proposed based on graph theory in this paper. We derived a formulation to count faces in a plane given only the vertices based on Euler polyhedron formula. Vertices with degree no more than two are ineffective in counting components. With the derived formula, a graph component counting algorithm is constructed based only on searching cross points whose degree are no less than three. When applied to a two-dimensional binary image, the proposed method divides an image into patches of same size and decides which of them will be used in counting by searching the circumferential pixels of each patch. If the number of component edges within the circumferential pixels of a patch is no less than three, then the patch will be used in counting. After determining all the vertices with degree no less than three, the number of components can be calculated by the formula. Because only a small number of pixels are investigated in the process, the computational time is very fast. The disconnection of edges in an image is one of the main reasons which causes miscounts for scanning-based algorithms. The difficulty, however, can be naturally overcome by the proposed algorithm because disconnected points will be identified as futile pixels in the algorithm. Experimental results show the algorithm is more efficient than existing methods. When applied to images with disconnected edges, the counted number given by scanning-based algorithms is much smaller than the correct number whereas the proposed algorithm obtains satisfactory results. Sihai Yang, Duansheng Chen, Xianhua Han, Yen-Wei Chen 0001 |
SNPD | 3 |
| 2016 | Integration of spatial and orientation contexts in local ternary patterns for HEp-2 cell classification
Xianhua Han, Yen-Wei Chen 0001 |
Pattern Recognit. Lett. | 1 |
| 2015 | Liver segmentation using superpixel-based graph cuts and restricted regions of shape constrainsabstractLiver segmentation is one of the most fundamental and challenging tasks in computer aided diagnosis (CAD) system for liver diseases. Graph cut algorithms have been successfully applied to medical image segmentation of different organs for 3D volume data, which not only leads to very large-scale graph due to the same node number as voxel number, but also completely ignore some available organ shape priors. Thus, a slice by slice liver segmentation method by combining shape constraints according to previously slice segmentation has been proposed based on graph cut. However, the constructed graph scale is still large, and the computation of distance map from all voxel to the segmented shape leads to high cost. In order to explore an efficient and effective slice by slice segmentation method for liver, this paper proposes to apply clustering algorithm to firstly group slice pixels into superpixels as nodes for constructing graph, which not only greatly reduce the graph scale but also significantly speed up the optimization procedure of the graph. Furthermore, we restrict the regions near organ boundary as shape constraints, which can further reduce computational time. To validate effectiveness and efficiency of our proposed method, we conduct experiments on 10 CT volumes, most of which have tumors inside liver, and abnormal deformed shape of liver. Our method can yield an average dice coefficient: 0.94, about 659.22 second in computation, and take only 1.5GB in memory usage. Titinunt Kitrungrotsakul, Xianhua Han, Yen-Wei Chen 0001 |
ICIP | 2 |
| 2015 | High-Order Statistics of Weber Local Descriptors for Image RepresentationabstractHighly discriminant visual features play a key role in different image classification applications. This study aims to realize a method for extracting highly-discriminant features from images by exploring a robust local descriptor inspired by Weber's law. The investigated local descriptor is based on the fact that human perception for distinguishing a pattern depends not only on the absolute intensity of the stimulus but also on the relative variance of the stimulus. Therefore, we firstly transform the original stimulus (the images in our study) into a differential excitation-domain according to Weber's law, and then explore a local patch, called micro-Texton, in the transformed domain as Weber local descriptor (WLD). Furthermore, we propose to employ a parametric probability process to model the Weber local descriptors, and extract the higher-order statistics to the model parameters for image representation. The proposed strategy can adaptively characterize the WLD space using generative probability model, and then learn the parameters for better fitting the training space, which would lead to more discriminant representation for images. In order to validate the efficiency of the proposed strategy, we apply three different image classification applications including texture, food images and HEp-2 cell pattern recognition, which validates that our proposed strategy has advantages over the state-of-the-art approaches. Xianhua Han, Yen-Wei Chen 0001 |
IEEE Trans. Cybern. | 1 |
| 2014 | Sparse and Low Rank Matrix Decomposition Based Local Morphological Analysis and Its Application to Diagnosis of Cirrhosis LiversabstractCirrhosis liver is a terrible disease which is threatening our lives. Meanwhile, cirrhosis will cause significant hepatic morphological changes. While it is well known that the livers from different subjects have similar global shape structure which means liver shape ensemble should be low-rank. However the deformation which caused by cirrhosis can be considered as sparse compared with the whole liver. Therefore, in this study, we proposed to apply spare and low-rank matrix decomposition to partition the local deformation part (sparse error matrix E) from the global similar structure (low-rank matrix A) using the input liver shape D, which is the landmark coordinates of liver shapes and already have been aligned by the current rigid registration methods firstly. And then sparse matrix E is used for diagnosis. In common sense, the normal liver should have less local deformation than that of abnormal liver, which means that the norm of sparse matrix E for normal liver is smaller than the norm for abnormal one. Thus, we can simply use a threshold classify normal and abnormal livers using the norm of E for these two categories. The proposed method is evaluated by a liver database which includes 30 normal livers and 30 abnormal livers. The experimental results of proposed method is better than those of state of the art statistical shape model(SSM) based methods. Junping Deng, Xianhua Han, Yen-Wei Chen 0001 |
ICPR | 2 |
| 2014 | Hybrid Aggregation of Sparse Coded Descriptors for Food RecognitionabstractRecent year, with the increasing of unhealthy diets which will threaten people's life due to the various resulted risks such as heart stroke, liver trouble and so on, the maintaining for healthy life has attracted much attention and then how to manage the dietary life is becoming more and more important. In this research, we aim to construct an auto-recognition system of food images and keep the daily food-log records which will contribute to manage dietary life. With the easily available food images taken by mobile phone, it prospects to give the insight about the daily dietary of users with our constructed food recognition system. In order to achieve the acceptable recognition performance of the food images, we propose to apply a sparse model for coding local descriptors extracted from the food images and various pooling methods for aggregating the xoded descriptors. Sparse coding: an extension of vector quantization for local descriptors, which is popularly used in Bag-of-Features (BoF) for image representation, can reconstruct the local descriptors more effective, and then obtain more discriminated feature for food image representation. However, in order to emphasize the strongest activated pattern, the widely applied aggregation strategy of the sparse coded vector is only to retain the maximum coefficient in all (named as Max-pooling), which would completely ignore the frequency: an important signature for identifying different types of images, of the activated patterns. Therefore, we explore a hybrid aggregation strategy named as top-ranked average pooling (TRAP), which integrates not only the maximum activated magnitude but also the stronger activated number for image representation. Experiments validate that the proposed hybrid aggregation strategy combined with sparse model can greatly improve the recognition rates compared with the conventional BOF model and the state-of-the-art methods on two databases: our constructed RFID and the public PFID. Riko Kusumoto, Xianhua Han, Yen-Wei Chen 0001 |
ICPR | 2 |
| 2013 | Pilot study of applying shape analysis to liver cirrhosis diagnosisabstractThis paper explores the potential of applying shape analysis to classify normal/cirrhotic liver and in addition estimate the severity of abnormal cases. Conventional Computer-Aided Diagnosis (CAD) systems are developed for automatically providing a binary output as a second opinion to assist radiologists to draw conclusions about the condition of the pathology (normal or abnormal). After the disease is diagnosed, grasping the proceeding stage of the abnormal degree is essential for adopting the appropriate strength of treatment. However, none of existing CAD system is well established for such a challenging task. Liver cirrhosis has an important feature: morphological changes of the liver and the spleen occur during the clinical course of liver cirrhosis. In this study we constructed liver, spleen and their joint Statistical Shape Models (SSMs) to quantitatively assess the global shape variation and selected several modes from the SSMs. Then we learnt a mapping function between coefficients of selected modes and the ground truth staging label by Support Vector Regression (SVR). Using this mapping function, the proceeding stage of new input data can be estimated. Experimental results have validated the potential of our method on assisting the cirrhosis diagnosis. Yen-Wei Chen 0001, Xianhua Han, Tomoko Tateyama, Akira Furukawa, Shuzo Kanasaki |
ICIP | 3 |
| 2013 | Residual Image Compensations for Enhancement of High-Frequency Components in Face Hallucination
Yen-Wei Chen 0001, So Sasatani, Xianhua Han |
ISNN (1) | 3 |
| 2013 | Generalized N-dimensional independent component analysis and its application to multiple feature selection and fusion for image classification
Danni Ai, Guifang Duan, Xianhua Han, Yen-Wei Chen 0001 |
Neurocomputing | 3 |
| 2012 | Multiple feature selection and fusion based on generalized N-dimensional independent component analysis
Danni Ai, Guifang Duan, Xianhua Han, Yen-Wei Chen 0001 |
ICPR | 3 |
| 2012 | Group sparse representation of adaptive sub-domain selection for image classification
Xianhua Han, Xu Qiao, Yen-Wei Chen 0001 |
ICPR | 1 |
| 2012 | Super-resolution of MR volumetric images using sparse representation and self-similarity
Yutaro Iwamoto, Xianhua Han, So Sasatani, Kazuki Taniguchi, Wei Xiong 0001, Yen-Wei Chen 0001 |
ICPR | 2 |
| 2012 | Image super-resolution based on locality-constrained linear coding
Kazuki Taniguchi, Xianhua Han, Yutaro Iwamoto, So Sasatani, Yen-Wei Chen 0001 |
ICPR | 2 |
| 2012 | Multilinear Supervised Neighborhood Embedding of a Local Descriptor Tensor for Scene/Object RecognitionabstractIn this paper, we propose to represent an image as a local descriptor tensor and use a multilinear supervised neighborhood embedding (MSNE) for discriminant feature extraction, which is able to be used for subject or scene recognition. The contributions of this paper include: 1) a novel feature extraction approach denoted as the histogram of orientation weighted with a normalized gradient (NHOG) for local region representation, which is robust to large illumination variation in an image; 2) an image representation framework denoted as the local descriptor tensor, which can effectively combine a moderate amount of local features together for image representation and be more efficient than the popular existing bag-of-feature model; and 3) an MSNE analysis algorithm, which can directly deal with the local descriptor tensor for extracting discriminant and compact features and, at the same time, preserve neighborhood structure in tensor-feature space for subject/scene recognition. We demonstrate the performance advantages of our proposed approach over existing techniques on different types of benchmark database such as a scene data set (i.e., OT8), face data sets (i.e., YALE and PIE), and view-based object data sets (COIL-100 and ETH-80). Xianhua Han, Yen-Wei Chen 0001, Xiang Ruan |
IEEE Trans. Image Process. | 1 |
| 2011 | Canonical correlation analysis of local feature set for view-based object recognitionabstractIn this paper, we propose to use local feature set for image representation, which can represent variations in an object's appearance due to changing viewpoint or camera pose. It was evidenced that usually only a part of the object are appeared in common when taking a photo of an object in different view points. With comparison of local features set extracted from different positions of images, an object can be recognized when common part is appeared in two images, which take photos of one object in different view points. In this paper, we use Canonical Correlation (also known as principle or canonical angles), which can be thought of as the angles between two d-dimensional subspace, as similarity measure of local feature sets. The proposed approach is evaluated in various view-based object datasets (Coil-100 and ETH80) for object and object category recognition. Experiments show that the performance advantages of our proposed approach can be achieved over existing techniques. Xianhua Han, Yen-Wei Chen 0001, Xiang Ruan |
ICIP | 1 |
| 2011 | High frequency compensated face hallucinationabstractFace Hallucination is, one of a learning-based super-resolution technique that can reconstruct a high-resolution image using only one low-resolution image. However, there are often some detailed high-frequency components of the reconstructed image that cannot be recovered using this method. In this study, we proposed a high-frequency compensated face hallucination method for enhancing reconstruction performance. The proposed method can be divided into three steps: 1)high-resolution image reconstruction using a conventional hallucination method; 2)residual (high-frequency components) image recovery by “training” a residual image pair; 3)compensation of the reconstructed high-resolution image obtained in step 1 with the reconstructed residual image. Experimental results show that the high-resolution images obtained using our proposed approach are much better than those obtained by conventional hallucination. So Sasatani, Xianhua Han, Takanori Igarashi, Motonori Ohashi, Yutaro Iwamoto, Yen-Wei Chen 0001 |
ICIP | 2 |
| 2010 | Image recognition by learned linear subspace of combined bag-of-features and low-level featuresabstractImage category recognition is important to access visual information on the level of objects and scene types. This paper combines different feature representations of images and learn a compact subspace of different features for the automatic recognition of object and scene classes. Compact visual-words and low-level-features object class subspaces are automatically learned from a set of training images by a Regularized Linear Discriminant analysis (RLDA) algorithm, and the extracted RLDA-domain features are used for Support Vector Machine (SVM) classifier. The main contribution of this paper is two folds: i) Different features (bag-of-features and low-level features)is fused for image representation. ii) The compact feature subspaces (low-dimension features) of different features are learned for rendering to SVM classifier, which is computationally efficient for image category. High classification accuracy is demonstrated on object recognition database (Caltech). We confirm that the proposed strategy cam improve accuracy rate compared with state-of-the-art methods for object recognition databases. Xianhua Han, Yen-Wei Chen 0001, Xiang Ruan |
ICIP | 1 |
| 2010 | Adaptive Color Independent Components Based SIFT Descriptors for Image ClassificationabstractThis paper proposes an adaptive color independent components based SIFT descriptor (termed CIC-SIFT) for image classification. Our motivation is to seek an adaptive and efficient color space for color SIFT feature extraction. Our work has two key contributions. First, based on independent component analysis (ICA), an adaptive and efficient color space is proposed for color image representation. Second, in this ICA-based color space, a discriminative CIC-SIFT descriptor is calculated for image classification. The experiment results indicate that (1) contrast between objects and background can be enhanced on the ICA-based color space and (2) the CIC-SIFT descriptor outperforms other conventional color SIFT descriptors on image classification. Danni Ai, Xianhua Han, Xiang Ruan, Yen-Wei Chen 0001 |
ICPR | 2 |
| 2010 | Image Categorization by Learned Nonlinear Subspace of Combined Visual-Words and Low-Level FeaturesabstractImage category recognition is important to access visual information on the level of objects and scene types. This paper presents a new algorithm for the automatic recognition of object and scene classes. Compact and yet discriminative visual-words and low-level-features object class subspaces are automatically learned from a set of training images by a Supervised Nonlinear Neighborhood Embedding (SNNE) algorithm, which can learn an adaptive nonlinear subspace by preserving the neighborhood structure of the visual feature space. The main contribution of this paper is two fold: i) an optimally compact and discriminative feature subspace is learned by the proposed SNNE algorithm for different feature space (visual-word and low-level features). ii) An effective merge of different feature subspace can be implemented simply. High classification accuracy is demonstrated on different database including the scene database (Simplicity) and object recognition database (Caltech). We confirm that the proposed strategy is much better than state-of-the-art methods for different databases. Xianhua Han, Yen-Wei Chen 0001, Xiang Ruan |
ICPR | 1 |
| 2010 | Semi-supervised and Interactive Semantic Concept Learning for Scene RecognitionabstractIn this paper, we present a novel semi-supervised and interactive concept learning algorithm for scene recognition by local semantic description. Our work is motivated by the continuing effort in content-based image retrieval to extract and to model the semantic content of images. The basic idea of the semantic modeling is to classify local image regions into semantic concept classes such as water, sunset, or sky. However, labeling concept sampling manually for training semantic model is fairly expensive, and the labeling results is, to some extent, subjective to the operators. In this paper, by using the proposed semi-supervised and interactive learning algorithm, training samples and new concepts can be obtained accurately and efficiently. Through extensive experiments, we demonstrate that the image concept representation is well suited for modeling the semantic content of heterogenous scene categories, and thus for recognition and retrieval. Furthermore, higher recognition accuracy can be achieved by updating new training samples and concepts, which are obtained by the novel proposed algorithm. Xianhua Han, Yen-Wei Chen 0001, Xiang Ruan |
ICPR | 1 |
| 2010 | Tensor-based subspace learning and its applications in multi-pose face synthesis
Xu Qiao, Xianhua Han, Takanori Igarashi, Keisuke Nakao, Yen-Wei Chen 0001 |
Neurocomputing | 2 |
| 2008 | A supervised nonlinear neighborhood embedding of color histogram for image indexingabstractSubspace learning techniques are widespread in pattern recognition research. They include PCA, ICA, LPP, etc. These techniques are generally linear and unsupervised. The problem of image indexing is very complicated and the processed images are usually lie on non-linear image subspaces. In this paper, we propose a supervised nonlinear neighborhood embedding algorithm which learns an adaptive nonlinear subspace by preserving the neighborhood structure of the image color space. In the proposed algorithm, we combine the idea of nonlinear kernel mapping and preserving the neighborhood structure of the samples, so it can not only gain a perfect approximation of the nonlinear image manifold, but also enhance within-class neighborhood information. Experimental results show that the proposed method outperform other linear or unsupervised subspace learning methods. Xianhua Han, Yen-Wei Chen 0001, Takeshi Sukegawa |
ICIP | 1 |
| 2008 | Classification of High-Resolution Satellite Images Using Supervised Locality Preserving Projections
Yen-Wei Chen 0001, Xianhua Han |
KES (2) | 2 |
| 2008 | Edge detection algorithm based on ICA-domain shrinkage in noisy images
Xianhua Han, Shuiyan Dai, Guorong Xia |
Sci. China Ser. F Inf. Sci. | 1 |
| 2003 | An ICA-Based Method for Poisson Noise Reduction
Xianhua Han, Yen-Wei Chen 0001, Zensho Nakao |
KES | 1 |