Lei Guo 0002

dblp:64/1967-2 · DBLP profile ↗
← Back
152ranked-venue papers
4as first author
42since 2021 · last 2026
0000-0003-0728-896XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 99 · 33 since 2021Graphics, computer vision, multimedia, augmented reality and games · 65 · 8 since 2021Artificial intelligence and machine learning · 19 · 3 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-authorDatabases, data management, data science and information retrieval · 3Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 Toward Trustworthy Multi-View Representation With Fine-Grained Explainability Embeddings
abstract
Multiomics co-learning is a powerful analytical paradigm that has benefited biomedical studies substantially. However, due to the diverse information and complex relationships of multiomics data, naive multi-view learning methods usually run into spurious correlations and biased signatures irrelevant to the diseases of interest. Therefore, the learned representations and cross-omics associations cannot translate into clinical knowledge for disease prediction. This issue becomes particularly severe when clinical data are limited and scarce. To handle this issue, we propose a novel and powerful scheme, referred to as the Causality-driven Trustworthy Multi-View maPping approach (Cad-TMVP). Specifically, we design a fined multi-directional mapping module to extract co-expression patterns across different modalities and capture fine-grained interpretability factors. We also meticulously design dynamic mechanisms to facilitate adaptive loss-term reweighting and trustworthy integration of multiple modalities. Cad-TMVP enhances downstream tasks by developing a cooperative learning module that simultaneously performs automated diagnosis and result interpretation. Furthermore, we develop an efficient search strategy and support computation to reduce the high computational burden, making our approach practicable. We conduct extensive experiments on different types of multiomics data. The proposed method establishes new state-of-the-art results in various settings while maintaining excellent interpretability. Thus, it sets a potentially newparadigm in trustworthy multi-modal learning and verifies its flexibility and versatility in real biomedical applications.
Jin Zhang 0023, Yan Yang 0011, Muheng Shang, Lei Guo 0002, Daoqiang Zhang, Lei Du 0001
IEEE Trans. Medical Imaging4
2025 Mutual-assistance learning for trustworthy biomarker discovery and disease prediction
abstract
Integrating and analyzing multiple omics datasets, such as genomics, environmental influences, and imaging endophenotypes, has yielded an abundance of candidate biomarkers. However, translating such findings into beneficial clinical knowledge for disease prediction remains challenging. This becomes even more challenging when studying interpretable high-order feature interactions such as gene-environment interaction (G$\times $E) to understand the etiology. To fill this gap, we draw on the idea of mutual-assistance (MA) learning and accordingly propose a fresh and powerful scheme, referred to as mutual-assistance causal biomarker discovery and stable disease prediction approach (MA-CBxDP). Specifically, we design an interpretable bi-directional mapping framework, integrated with a causal feature interaction module, to extract co-expression patterns across different modalities and identify trustworthy biomarkers including G$\times $E. A cooperative prediction module is further incorporated to ensure accurate diagnosis and identification of causal effects for pathogenesis. Importantly, biomarker discovery and disease prediction can mutually reinforce each other, helping to provide novel insights into chronic diseases. Furthermore, in light of the large computational burden incurred by the high-dimensional interactions, we devise a rapid strategy and extend it to a more practical but challenging chromosome-wide setting. We conduct extensive experiments on two databases under three tasks, i.e. multimodal correlation, disease diagnosis, and trait prediction. MA-CBxDP establishes new state-of-the-art results in predicting clinical scores and disease status classification, while maintaining exceptional interpretability, verifying its flexibility and versatility in practical applications.
Jin Zhang 0023, Yan Yang 0011, Muheng Shang, Lei Guo 0002, Daoqiang Zhang, Lei Du 0001
Briefings Bioinform.4
2025 Trustworthy causal biomarker discovery: a multiomics brain imaging genetics-based approach
abstract
MOTIVATION: Discovering genetic variations underpinning brain disorders is important to understand their pathogenesis. Indirect associations or spurious causal relationships pose a threat to the reliability of biomarker discovery for brain disorders, potentially misleading or incurring bias in subsequent decision-making. Unfortunately, the stringent selection of reliable biomarker candidates for brain disorders remains a predominantly unexplored challenge. RESULTS: In this article, to fill this gap, we propose a fresh and powerful scheme, referred to as the Causality-aware Genotype intermediate Phenotype Correlation Approach (Ca-GPCA). Specifically, we design a bidirectional association learning framework, integrated with a parallel causal variable decorrelation module and sparse variable regularizer module, to identify trustworthy causal biomarkers. A disease diagnosis module is further incorporated to ensure accurate diagnosis and identification of causal effects for pathogenesis. Additionally, considering the large computational burden incurred by high-dimensional genotype-phenotype covariances, we develop a fast and efficient strategy to reduce the runtime and prompt practical availability and applicability. Extensive experimental results on four simulation data and real neuroimaging genetic data clearly show that Ca-GPCA outperforms state-of-the-art methods with excellent built-in interpretability. This can provide novel and reliable insights into the underlying pathogenic mechanisms of brain disorders. AVAILABILITY AND IMPLEMENTATION: The software is publicly available at https://github.com/ZJ-Techie/Ca-GPCA.
Jin Zhang 0023, Yan Yang 0011, Muheng Shang, Lei Guo 0002, Daoqiang Zhang, Lei Du 0001
Bioinform.4
2025 Contrastive machine learning reveals species -shared and -specific brain functional architecture
Guannan Cao, Songyao Zhang, Weihan Zhang, Yusong Sun, Jingchao Zhou, Tianyang Zhong, Yixuan Yuan, Tao Liu 0044, Tianming Liu 0001, Lei Guo 0002, Yongchun Yu, Xi Jiang 0001, Gang Li 0001, Junwei Han 0001
Medical Image Anal.11
2025 A Foundational fMRI Model for Representing Continuous Brain States
abstract
Foundational models have significant potential to advance brain function research, particularly in understanding the dynamics of brain states. However, most existing models process brain signals within fixed time windows, restricting their ability to capture the full temporal complexity of brain activity. In this study, we propose BrainSN (Brain States Network), a novel fMRI foundational model designed to represent continuous brain state information and support diverse downstream tasks. First, leveraging a transformer-based architecture, BrainSN reconstructs input brain states across multiple time scales and predicts future brain activity, effectively capturing both short-term and long-term dependencies. Second, through multiple embeddings and a channel gating module, the model integrates brain state information and applies an attention mechanism to extract critical features. Additionally, we train BrainSN on 1,256 hours of resting-state and naturalistic stimulus fMRI data, enabling it to learn large-scale brain dynamics without relying on task-based paradigms. Without fine-tuning, BrainSN achieves 75.23% and 75.82% accuracy in autism and attention disorder diagnosis tasks, respectively, matching the performance of leading models pretrained on disease-specific data. After fine-tuning, it surpasses these models. In mental state decoding, BrainSN attains 95.31% accuracy without fine-tuning, outperforming the best models trained on large-scale task-based fMRI data. Furthermore, by analyzing BrainSN's embeddings in relation to movie stimuli, we demonstrate that the model effectively captures the semantic content of movie scenes embedded in fMRI signals and is highly sensitive to sequence. These results highlight BrainSN's ability to model brain state dynamics and underscore its potential advantages for clinical diagnosis, treatment evaluation, and cognitive neuroscience research.
Lei Guo 0002, Yixuan Yuan, Junwei Han 0001, Xintao Hu
IEEE J. Biomed. Health Informatics2
2025 Dual Stream Relation Learning Network for Image-Text Retrieval
abstract
Image-text retrieval has made remarkable achievements through the development of feature extraction networks and model architectures. However, almost all region feature-based methods face two serious problems when modeling modality interactions. First, region features are prone to feature entanglement in the feature extraction stage, making it difficult to accurately reason complex intra-model relations between visual objects. Second, region features lack rich contextual information, background, and object details, making it difficult to achieve precise inter-modal alignment with textual information. In this paper, we propose a novel Dual Stream Relation Learning Network (DSRLN) to jointly solve these issues with two key components: a Geometry-sensitive Interactive Self-Attention (GISA) module and a Dual Information Fusion (DIF) module. Specifically, GISA extends the vanilla self-attention network from two aspects to better model the intrinsic relationships between different regions, thereby improving high-level visual-semantic reasoning ability. DIF uses grid features as an additional visual information source, and achieves deeper and complex fusion between the two types of features through a masked cross-attention module and an adaptive gate fusion module, which can capture comprehensive visual information to learn more precise inter-modal alignment. Besides, our method also learns a more comprehensive hierarchical correspondence between images and sentences through local and global alignment. Experimental results on two public datasets, i.e., Flickr30K and MS-COCO, fully demonstrate the superiority and effectiveness of our model.
Dongqing Wu, Cang Gu, Lei Guo 0002
IEEE Trans. Multim.4
2024 Identification of disease-related genetic variants and imaging factors leveraging summary statistics
abstract
Brain imaging genetics offers insights into the genetic basis of brain structure and function by exploring the relationships between genetic variations and neuroimaging features, with canonical correlation association learning as a vital and effective tool. However, imaging large cohorts affected by specific brain diseases entails significant costs. To tackle this challenge, we introduced a novel bi-multivariate sparse canonical correlation association method based on summary statistics from large GWAS (S-SCCA). S-SCCA leverages effect sizes obtained from these datasets to identify genetic variants associated with complex traits, including those influenced by pleiotropy, while simultaneously identifying imaging factors related to the disease under study. Moreover, we have implemented a rapid optimization strategy to circumvent computational burdens while identifying disease-associated risk factors within genetic variations across the entire chromosome. We assessed S-SCCA against conventional SCCA using a neuroimaging genetic dataset from the Alzheimer’s Disease Neuroimaging Initiative. Results showed that S-SCCA demonstrated comparable or superior modeling performance and feature selection capabilities. Furthermore, we applied S-SCCA to two summary statistics datasets from two large GWAS, where original imaging and genetic data were inaccessible. S-SCCA replicated the genetic loci identified by GWAS and additional meaningful variants. Additionally, it revealed bi-multivariate relationships between imaging QTs and SNPs, indicating its powerful modeling capability. These findings highlight the promise of S-SCCA as a practical bi-multivariate learning technique in brain imaging genetics, circumventing the need for sensitive individual-level imaging and genetic data, thereby enhancing its potential for broader applicability and accessibility in biomedical studies.
Duo Xi, Dingnan Cui, Minjianan Zhang, Jin Zhang 0023, Muheng Shang, Lei Guo 0002, Lei Du 0001, Junwei Han 0001
BIBM6
2024 Disentangling Disease-sensitive Multimodal Neuroimaging Phenotypes and Related Genetic Factors: A Multimodal Study of ADNI Cohort
abstract
Understanding neurological manifestations and their genetic architectures are important for exploring the etiology and pathology of brain disorders. Multimodal neuroimaging data carry complementary information and are known to exhibit shared and specific characteristics from different perspectives. Hence, exploring modality-shared and modality-specific imaging features as well as their genetic underpinnings is a challenging but beneficial task. Unfortunately, this issue has been largely unexploited. In this paper, to fill this gap, we propose a fresh and straightforward insight, referred as Multimodality-Disentangled Phenotype-Genotype Correlation approach (MDPGC). Specifically, we design a unified framework for exploring the multimodality-disentangled characteristics of image-based phenotypes, and further detect genetic variants associated with the disorder using modality-shared and modality-specific biomarkers as intermediate phenotypes. Extensive experimental results on Alzheimer’s Disease Neuroimaging Initiative (ADNI) dataset reveal that our method attains superior correlation coefficients compared to state-of-the-art methods, and at the same time provided excellent interpretability. In addition, the subsequent analysis demonstrates that MDPGC successfully identifies different types of characteristics of imaging phenotypes and reveals relevant genetic variations. These findings not only contribute to AD diagnosis but also help better understand the pathological and pathogenic mechanisms of brain disorders.
Jin Zhang 0023, Minjianan Zhang, Lei Guo 0002, Daoqiang Zhang, Lei Du 0001
BIBM3
2024 Disease Progression Prediction Incorporating Genotype-Environment Interactions: A Longitudinal Neurodegenerative Disorder Study
Jin Zhang 0023, Muheng Shang, Yan Yang 0011, Lei Guo 0002, Junwei Han 0001, Lei Du 0001
MICCAI (3)4
2024 Eye-gaze Guided Multi-modal Alignment for Medical Representation Learning
abstract
In the medical multi-modal frameworks, the alignment of cross-modality features presents a significant challenge. However, existing works have learned features that are implicitly aligned from the data, without considering the explicit relationships in the medical context. This data-reliance may lead to low generalization of the learned alignment relationships. In this work, we propose the Eye-gaze Guided Multi-modal Alignment (EGMA) framework to harness eye-gaze data for better alignment of medical visual and textual features. We explore the natural auxiliary role of radiologists' eye-gaze data in aligning medical images and text, and introduce a novel approach by using eye-gaze data, collected synchronously by radiologists during diagnostic evaluations. We conduct downstream tasks of image classification and image-text retrieval on four medical datasets, where EGMA achieved state-of-the-art performance and stronger generalization across different datasets. Additionally, we explore the impact of varying amounts of eye-gaze data on model performance, highlighting the feasibility and utility of integrating this auxiliary data into multi-modal alignment framework.
Chong Ma 0004, Hanqi Jiang, Wenting Chen, Yiwei Li 0002, Zihao Wu 0001, Xiaowei Yu 0001, Zhengliang Liu, Lei Guo 0002, Dajiang Zhu, Dinggang Shen, Tianming Liu 0001, Xiang Li 0001
NeurIPS8
2024 Modeling genotype-protein interaction and correlation for Alzheimer's disease: a multi-omics imaging genetics study
abstract
Integrating and analyzing multiple omics data sets, including genomics, proteomics and radiomics, can significantly advance researchers' comprehensive understanding of Alzheimer's disease (AD). However, current methodologies primarily focus on the main effects of genetic variation and protein, overlooking non-additive effects such as genotype-protein interaction (GPI) and correlation patterns in brain imaging genetics studies. Importantly, these non-additive effects could contribute to intermediate imaging phenotypes, finally leading to disease occurrence. In general, the interaction between genetic variations and proteins, and their correlations are two distinct biological effects, and thus disentangling the two effects for heritable imaging phenotypes is of great interest and need. Unfortunately, this issue has been largely unexploited. In this paper, to fill this gap, we propose $\textbf{M}$ulti-$\textbf{T}$ask $\textbf{G}$enotype-$\textbf{P}$rotein $\textbf{I}$nteraction and $\textbf{C}$orrelation disentangling method ($\textbf{MT-GPIC}$) to identify GPI and extract correlation patterns between them. To ensure stability and interpretability, we use novel and off-the-shelf penalties to identify meaningful genetic risk factors, as well as exploit the interconnectedness of different brain regions. Additionally, since computing GPI poses a high computational burden, we develop a fast optimization strategy for solving MT-GPIC, which is guaranteed to converge. Experimental results on the Alzheimer's Disease Neuroimaging Initiative data set show that MT-GPIC achieves higher correlation coefficients and classification accuracy than state-of-the-art methods. Moreover, our approach could effectively identify interpretable phenotype-related GPI and correlation patterns in high-dimensional omics data sets. These findings not only enhance the diagnostic accuracy but also contribute valuable insights into the underlying pathogenic mechanisms of AD.
Jin Zhang 0023, Zikang Ma, Yan Yang 0011, Lei Guo 0002, Lei Du 0001
Briefings Bioinform.4
2024 A Multi-Task Deep Feature Selection Method for Brain Imaging Genetics
abstract
Using brain imaging quantitative traits (QTs) for identifying genetic risk factors is an important research topic in brain imaging genetics. Many efforts have been made for this task via building linear models between imaging QTs and genetic factors such as single nucleotide polymorphisms (SNPs). To the best of our knowledge, linear models could not fully uncover the complicated relationship due to the loci's elusive and diverse influences on imaging QTs. In this paper, we propose a novel multi-task deep feature selection (MTDFS) method for brain imaging genetics. MTDFS first builds a multi-task deep neural network to model the complicated associations between imaging QTs and SNPs. And then designs a multi-task one-to-one layer and imposes a combined penalty to identify SNPs that make significant contributions. MTDFS can not only extract the nonlinear relationship but also arms the deep neural network with feature selection. We compared MTDFS to multi-task linear regression (MTLR) and single-task DFS (DFS) methods on the real neuroimaging genetic data. The experimental results showed that MTDFS performed better than MTLR and DFS on the QT-SNP relationship identification and feature selection. Thus, MTDFS is powerful for identifying risk loci and could be a great supplement to brain imaging genetics.
Shu Zhang 0006, Muheng Shang, Lei Guo 0002, Junwei Han 0001, Lei Du 0001
IEEE ACM Trans. Comput. Biol. Bioinform.4
2024 Spatial-Channel Attention Transformer With Pseudo Regions for Remote Sensing Image-Text Retrieval
abstract
Recently, remote sensing image-text retrieval (RSITR) has received significant attention due to its flexible query form and effective management of remote sensing images. However, prior work often relies on compact global features and ignores local features that can reflect salient objects in the images. Moreover, these methods primarily model interactions between features in the spatial domain, which is insufficient for mining the rich semantic information presented in remote sensing images. In this article, we propose a novel spatial-channel attention transformer (SCAT) with pseudo regions to address these issues. Concretely, in order to acquire the fine-grained perception of local objects, we introduce a pseudo region generation (PRG) module that adaptively aggregates grid features with similar semantic information into multiple clusters through a clustering algorithm. These generated cluster centers are able to flexibly and efficiently represent local objects in remote sensing images without relying on sophisticated object detectors. Furthermore, in order to achieve a comprehensive understanding of image semantics information, we carefully construct a novel SCAT. By exploiting spatial and channel attention to explore the dependencies between features at both spatial and channel domains, the proposed SCAT enhances the model’s ability to identify both “where to look” and “what it is,” thereby obtaining a more powerful representation. In addition, SCAT incorporates two novel designs that alleviate the high overhead caused by attention modeling. Extensive experiments on two benchmark datasets, RSICD and RSITMD, fully demonstrate the effectiveness and superiority of our proposed method.
Dongqing Wu, Yinxuan Hou, Cuili Xu, Gong Cheng 0003, Lei Guo 0002
IEEE Trans. Geosci. Remote. Sens.6
2024 Feature First: Advancing Image-Text Retrieval Through Improved Visual Features
abstract
Current image-text retrieval methods mainly utilize region features that provide object-level information to represent images, making the retrieval results more accurate and interpretable. However, there are several issues with region features, such as lack of rich contextual information, loss of object details and risk of detection redundancy. The ideal visual features in image-text retrieval should have three characteristics: object-level, semantically-rich, and language-aligned. To this end, we propose a novel visual representation framework to capture more comprehensive and powerful visual features. Specifically, since these region feature disadvantages are the grid feature advantages, we first build a two-step interaction model to explore the complex relationship between them from the spatial and semantic perspectives to integrate their complementary information, making the fused visual features both object-level and semantic-rich. Then, we design a text-integrated visual embedding module that utilizes textual information as guidance to filter redundant regions, further endowing visual features with language-aligned capabilities. Finally, we develop a multi-attention pooling module to better aggregate these enhanced visual features in a more fine-grained manner. Extensive experiments demonstrate that our proposed model achieves state-of-the-art performance on the benchmark datasets Flickr30K and MS-COCO.
Dongqing Wu, Cang Gu, Cuili Xu, Yinxuan Hou, Lei Guo 0002
IEEE Trans. Multim.7
2024 Anatomy-Guided Spatio-Temporal Graph Convolutional Networks (AG-STGCNs) for Modeling Functional Connectivity Between Gyri and Sulci Across Multiple Task Domains
abstract
The cerebral cortex is folded as gyri and sulci, which provide the foundation to unveil anatomo-functional relationship of brain. Previous studies have extensively demonstrated that gyri and sulci exhibit intrinsic functional difference, which is further supported by morphological, genetic, and structural evidences. Therefore, systematically investigating the gyro-sulcal (G-S) functional difference can help deeply understand the functional mechanism of brain. By integrating functional magnetic resonance imaging (fMRI) with advanced deep learning models, recent studies have unveiled the temporal difference in functional activity between gyri and sulci. However, the potential difference of functional connectivity, which represents functional dependency between gyri and sulci, is much unknown. Moreover, the regularity and variability of the G-S functional connectivity difference across multiple task domains remains to be explored. To address the two concerns, this study developed new anatomy-guided spatio-temporal graph convolutional networks (AG-STGCNs) to investigate the regularity and variability of functional connectivity differences between gyri and sulci across multiple task domains. Based on 830 subjects with seven different task-based and one resting state fMRI (rs-fMRI) datasets from the public Human Connectome Project (HCP), we consistently found that there are significant differences of functional connectivity between gyral and sulcal regions within task domains compared with resting state (RS). Furthermore, there is considerable variability of such functional connectivity and information flow between gyri and sulci across different task domains, which are correlated with individual cognitive behaviors. Our study helps better understand the functional segregation of gyri and sulci within task domains as well as the anatomo-functional-behavioral relationship of the human brain.
Mingxin Jiang, Yuzhong Chen 0002, Jiadong Yan, Zhenxiang Xiao, Shimin Yang, Zhongbo Zhao, Lei Guo 0002, Benjamin Becker, Dezhong Yao 0001, Keith M. Kendrick, Xi Jiang 0001
IEEE Trans. Neural Networks Learn. Syst.10
2024 Rectify ViT Shortcut Learning by Visual Saliency
abstract
Shortcut learning in deep learning models occurs when unintended features are prioritized, resulting in degenerated feature representations and reduced generalizability and interpretability. However, shortcut learning in the widely used vision transformer (ViT) framework is largely unknown. Meanwhile, introducing domain-specific knowledge is a major approach to rectifying the shortcuts that are predominated by background-related factors. For example, eye-gaze data from radiologists are effective human visual prior knowledge that has the great potential to guide the deep learning models to focus on meaningful foreground regions. However, obtaining eye-gaze data can still sometimes be time-consuming, labor-intensive, and even impractical. In this work, we propose a novel and effective saliency-guided ViT (SGT) model to rectify shortcut learning in ViT with the absence of eye-gaze data. Specifically, a computational visual saliency model (either pretrained or fine-tuned) is adopted to predict saliency maps for input image samples. Then, the saliency maps are used to filter the most informative image patches. Considering that this filter operation may lead to global information loss, we further introduce a residual connection that calculates the self-attention across all the image patches. The experiment results on natural and medical image datasets show that our SGT framework can effectively learn and leverage human prior knowledge without eye-gaze data and achieves much better performance than baselines. Meanwhile, it successfully rectifies the harmful shortcut learning and significantly improves the interpretability of the ViT model, demonstrating the promise of transferring human prior knowledge derived visual saliency in rectifying shortcut learning.
Chong Ma 0004, Lin Zhao 0004, Yuzhong Chen 0002, Lei Guo 0002, Xintao Hu, Dinggang Shen, Xi Jiang 0001, Tianming Liu 0001
IEEE Trans. Neural Networks Learn. Syst.4
2023 Identifying Disease-related Brain Imaging Quantitative Traits and Related Genetic Variations via A Bidirectional Association Learning Method
abstract
Discovering critical genetic biomarkers of Alzheimer’s disease (AD) by detecting the complex associations between genotypes (i.e. single nucleotide polymorphism, SNP) and phenotypes (i.e. quantitative trait, QT) is a long-standing and beneficial task for the diagnosis and the follow-up treatment of patients. The function of genes and their relationships with phenotypes are extremely complex. A lot of imaging genetic methods have been designed to uncover the association between brain imaging QTs and SNPs. However, most of them are focused on the effect of a single SNP, which may have limited ability due to the oligogenic or polygenic characteristic of AD. In this paper, we propose a deep reconstruction bidirectional association with feature selection (DRBA-FS) method to explore the multi-SNPmulti-QT associations. In this method, the co-effect of multiple AD-related genetic variations is identified and aggregated, and their high-level genetic associations to brain imaging QTs are jointly modeled. Experiment results on real neuroimaging genetic data from Alzheimer’s Disease Neuroimaging Initiative (ADNI) show that the identified biomarkers are all related to AD. Interestingly, our method can learn the joint effect of multiple AD-related genetic variations across the genome, and thus has significant potential in understanding the genetic mechanism of AD.
Muheng Shang, Yan Yang 0011, Minjianan Zhang, Jin Zhang 0023, Duo Xi, Lei Guo 0002, Lei Du 0001
BIBM6
2023 Identifying Main and Epistasis Effects of Genetic Variations on Neuroimaging Phenotypes Using Effective Feature Interaction Learning
abstract
Brain imaging genetics investigates the complex relationships between genetic variations and brain imaging quantitative traits (QTs). However, existing approaches primarily focus on the main effects of genetic variations, potentially neglecting the crucial role of epistasis that explains the missing heritability of brain disorders. Unfortunately, this issue has been largely unexploited. In this paper, to fill this gap, we present Multi-Task feature interaction-aware Sparse Canonical Correlation Analysis (MTfiSCCA) to identify disease-related main effect and epistasis of risk genetic factors on multimodal neuroimaging phenotypes simultaneously. To ensure stability and interpretation, we use innovative sparsity-inducing penalties to identify biomarkers that make significant contributions. Additionally, we develop an efficient optimization algorithm to solve the proposed method, which converges to a local optimum. Experimental results on the Alzheimer’s disease neuroimaging initiative (ADNI) dataset show that our MTfiSCCA method achieves higher canonical correlation coefficients (CCC) and better feature selection subsets such as disease-related biomarkers compared to the state-of-the-art methods. Furthermore, MTfiSCCA reveals interpretable epistasis among genetic variations implicated in AD, offering novel insights into the underlying pathogenic mechanisms of brain disorders such as Alzheimer’s disease (AD).
Jin Zhang 0023, Muheng Shang, Duo Xi, Minjianan Zhang, Lei Guo 0002, Lei Du 0001
BIBM6
2023 Prediction of Cognitive Scores by Joint Use of Movie-Watching fMRI Connectivity and Eye Tracking via Attention-CensNet
Jiaxing Gao, Lin Zhao 0004, Tianyang Zhong, Changhe Li, Yaonai Wei, Shu Zhang 0001, Lei Guo 0002, Tianming Liu 0001, Junwei Han 0001
MICCAI (2)8
2023 Identification of Disease-Sensitive Brain Imaging Phenotypes and Genetic Factors Using GWAS Summary Statistics
Duo Xi, Dingnan Cui, Jin Zhang 0023, Muheng Shang, Minjianan Zhang, Lei Guo 0002, Junwei Han 0001, Lei Du 0001
MICCAI (5)6
2023 A Small-Sample Method with EEG Signals Based on Abductive Learning for Motor Imagery Decoding
Tianyang Zhong, Xiaozheng Wei, Enze Shi, Jiaxing Gao, Chong Ma 0004, Yaonai Wei, Songyao Zhang, Lei Guo 0002, Junwei Han 0001, Tianming Liu 0001
MICCAI (1)8
2023 Adaptive structured sparse multiview canonical correlation analysis for multimodal brain imaging association identification
Lei Du 0001, Huiai Wang, Jin Zhang 0023, Shu Zhang 0001, Lei Guo 0002, Junwei Han 0001
Sci. China Inf. Sci.5
2023 Bel: Batch Equalization Loss for scene graph generation
Baorong Liu, Dongqing Wu, Lei Guo 0002
Pattern Anal. Appl.5
2023 Eye-Gaze-Guided Vision Transformer for Rectifying Shortcut Learning
abstract
Learning harmful shortcuts such as spurious correlations and biases prevents deep neural networks from learning meaningful and useful representations, thus jeopardizing the generalizability and interpretability of the learned representation. The situation becomes even more serious in medical image analysis, where the clinical data are limited and scarce while the reliability, generalizability and transparency of the learned model are highly required. To rectify the harmful shortcuts in medical imaging applications, in this paper, we propose a novel eye-gaze-guided vision transformer (EG-ViT) model which infuses the visual attention from radiologists to proactively guide the vision transformer (ViT) model to focus on regions with potential pathology rather than spurious correlations. To do so, the EG-ViT model takes the masked image patches that are within the radiologists' interest as input while has an additional residual connection to the last encoder layer to maintain the interactions of all patches. The experiments on two medical imaging datasets demonstrate that the proposed EG-ViT model can effectively rectify the harmful shortcut learning and improve the interpretability of the model. Meanwhile, infusing the experts' domain knowledge can also improve the large-scale ViT model's performance over all compared baseline methods with limited samples available. In general, EG-ViT takes the advantages of powerful deep neural networks while rectifies the harmful shortcut learning with human expert's prior knowledge. This work also opens new avenues for advancing current artificial intelligence paradigms by infusing human intelligence.
Chong Ma 0004, Lin Zhao 0004, Yuzhong Chen 0002, Sheng Wang 0014, Lei Guo 0002, Dinggang Shen, Xi Jiang 0001, Tianming Liu 0001
IEEE Trans. Medical Imaging5
2022 A Sparse Multi-task Contrastive and Discriminative Learning Method with Feature Selection for Brain Imaging Genetics
abstract
Alzheimer’s disease (AD) is a very complex neurodegenerative disease. Generally, different diagnostic groups could exhibit discriminative and specific patterns, including the single nucleotide polymorphisms (SNPs), brain imaging quantitative traits (QTs), as well as their associations, which may facilitate the comprehensive understanding of AD. However, most existing methods cannot guarantee to identify discriminative or class-specific biomarkers or both of them. To overcome this shortcoming, we propose a sparse multi-task contrastive and discriminative learning approach (MTCDA) to jointly learn the discriminative and specific patterns for multiple diagnostic groups. MTCDA can identify the class-relevant and discriminative SNP-QTs associations, and relevant SNPs, imaging QTs underpinning this relationship. We introduce an efficient algorithm to solve the proposed method which converges to a local optimum. The experimental results on Alzheimer’s Disease Neuroimaging Initiative (ADNI) show that MTCDA can obtain higher canonical correlation coefficients, classification accuracy and better feature selection results than state-of-the-art methods, which demonstrates the potential of our method for multi-class brain imaging genetics.
Jin Zhang 0023, Muheng Shang, Minjianan Zhang, Duo Xi, Lei Guo 0002, Junwei Han 0001, Lei Du 0001
BIBM6
2022 Improving Fusion of Region Features and Grid Features via Two-Step Interaction for Image-Text Retrieval
abstract
In recent years, region features extracted from object detection networks have been widely used in the image-text retrieval task. However, they lack rich background and contextual information, which makes it difficult to match words describing global concepts in sentences. Meanwhile, the region features also lose the details of objects in the image. Fortunately, these disadvantages of region features are the advantages of grid features. In this paper, we propose a novel framework, which fuses the region features and grid features through a two-step interaction strategy, thus extracting a more comprehensive image representation for image-text retrieval. Concretely, in the first step, a joint graph with spatial information constraints is constructed, where all region features and grid features are represented as graph nodes. By modeling the relationships using the joint graph, the information can be passed edge-wise. In the second step, we propose a Cross-attention Gated Fusion module, which further explores the complex interactions between region features and grid features, and then adaptively fuses different types of features. With these two steps, our model can fully realize the complementary advantages of region features and grid features. In addition, we propose a Multi-Attention Pooling module to better aggregate the fused region features and grid features. Extensive experiments on two public datasets, including Flickr30K and MS-COCO, demonstrate that our model achieves the state-of-the-art and pushes the performance of image-text retrieval to a new height.
Dongqing Wu, Cang Gu, Lei Guo 0002
ACM Multimedia4
2022 Identification of multimodal brain imaging association via a parameter decomposition based sparse multi-view canonical correlation analysis method
abstract
BACKGROUND: With the development of noninvasive imaging technology, collecting different imaging measurements of the same brain has become more and more easy. These multimodal imaging data carry complementary information of the same brain, with both specific and shared information being intertwined. Within these multimodal data, it is essential to discriminate the specific information from the shared information since it is of benefit to comprehensively characterize brain diseases. While most existing methods are unqualified, in this paper, we propose a parameter decomposition based sparse multi-view canonical correlation analysis (PDSMCCA) method. PDSMCCA could identify both modality-shared and -specific information of multimodal data, leading to an in-depth understanding of complex pathology of brain disease. RESULTS: Compared with the SMCCA method, our method obtains higher correlation coefficients and better canonical weights on both synthetic data and real neuroimaging data. This indicates that, coupled with modality-shared and -specific feature selection, PDSMCCA improves the multi-view association identification and shows meaningful feature selection capability with desirable interpretation. CONCLUSIONS: The novel PDSMCCA confirms that the parameter decomposition is a suitable strategy to identify both modality-shared and -specific imaging features. The multimodal association and the diverse information of multimodal imaging data enable us to better understand the brain disease such as Alzheimer's disease.
Jin Zhang 0023, Huiai Wang, Ying Zhao 0015, Lei Guo 0002, Lei Du 0001
BMC Bioinform.4
2022 Global-Guided Asymmetric Attention Network for Image-Text Matching
Dongqing Wu, Yinge Tang, Lei Guo 0002
Neurocomputing4
2022 Guiding Clean Features for Object Detection in Remote Sensing Images
abstract
Recently, object detection has gained significant progress in remote sensing images. Nevertheless, we conclude two defects in remote sensing image object detection. At first, most methods rely on feature pyramid, but the features of different levels would influence each other when we use top-down operation. Second, the traditional label assignment strategy cannot assign suitable labels, as it adopts the fixed intersection over union (IoU) threshold to divide positive samples and negative samples during training. According to the problems we pointed out, a simple yet effective framework is employed to eliminate these two limitations. It integrates two novel components: aware feature pyramid network (AFPN) and group assignment strategy (GAS). AFPN is to mitigate the adverse effects caused by the first problem. Specifically, it learns a vector for the higher level features in the feature pyramid to obtain clean features. As for the second limitation, we recommend a new label assignment strategy named GAS to address this problem. Samples will be grouped according to their overlaps with ground truth, and then, they are assigned to positive or negative labels in each group. Extensive experiments are conducted on the large-scale object detection dataset DIOR and DOTA. With the newly introduced two key components, our model significantly improves the detection accuracy. Without bells and whistles, our proposed method achieves 2.0% and 1.9% higher mean average precision (mAP) than Faster R-CNN with FPN when using ResNet50 and ResNet101 as the backbones, respectively. Finally, we obtain 73.3% mAP on the DIOR dataset without any tricks. Our code is available athttps://github.com/hm-better/dior_detect.
Gong Cheng 0003, Hailong Hong, Xiwen Yao, Xiaoliang Qian, Lei Guo 0002
IEEE Geosci. Remote. Sens. Lett.6
2022 Gumbel-Softmax based Neural Architecture Search for Hierarchical Brain Networks Decomposition
Tianji Pang, Shijie Zhao 0001, Junwei Han 0001, Shu Zhang 0001, Lei Guo 0002, Tianming Liu 0001
Medical Image Anal.5
2022 SPNet: Siamese-Prototype Network for Few-Shot Remote Sensing Image Scene Classification
abstract
Few-shot image classification has attracted extensive attention, which aims to recognize unseen classes given only a few labeled samples. Due to the large intraclass variances and interclass similarity of remote sensing scenes, the task under such circumstance is much more challenging than general few-shot image classification. Most existing prototype-based few-shot algorithms usually calculate prototypes directly from support samples and ignore the validity of prototypes, which results in a decline in the accuracy of subsequent inferences based on prototypes. To tackle this problem, we propose a Siamese-prototype network (SPNet) with prototype self-calibration (SC) and intercalibration (IC). First, to acquire more accurate prototypes, we utilize the supervision information from support labels to calibrate the prototypes generated from support features. This process is called SC. Second, we propose to consider the confidence scores of the query samples as another type of prototypes, which are then used to predict the support samples in the same way. Thus, the information interaction between support and query samples is implicitly a further calibration for prototypes (so-called IC). Our model is optimized with three losses, of which two additional losses help the model to learn more representative prototypes and make more accurate predictions. With no additional parameters to be learned, our model is very lightweight and convenient to employ. The experiments on three public remote sensing image datasets demonstrate competitive performance compared with other advanced few-shot image classification approaches. The source code is available athttps://github.com/zoraup/SPNet.
Gong Cheng 0003, Liming Cai, Chunbo Lang, Xiwen Yao, Lei Guo 0002, Junwei Han 0001
IEEE Trans. Geosci. Remote. Sens.6
2022 Perturbation-Seeking Generative Adversarial Networks: A Defense Framework for Remote Sensing Image Scene Classification
abstract
The methods for remote sensing image (RSI) scene classification based on deep convolutional neural networks (DCNNs) have achieved prominent success. However, confronted with adversarial examples obtained by adding imperceptible perturbations to clean images, the great vulnerability of DCNNs makes it worth exploring effective defense methods. To date, numerous countermeasures for adversarial examples have been proposed, but how to improve the defensive ability for unknown attacks still to be answered. To address this issue, in this article, we propose an effective defense framework specified for RSI scene classification, named perturbation-seeking generative adversarial networks (PSGANs). In brief, a new training framework is designed to train the classifier by introducing the examples generated during the image reconstruction process, in addition to clean examples and adversarial ones. These generated examples can be random kinds of unknown attacks during training and thus are utilized to eliminate the blind spots of a classifier. To assist the proposed training framework, a reconstruction method is developed. First, instead of modeling the distribution of clean examples, we model the distributions of the perturbations added in adversarial examples. Second, to make a tradeoff between the diversity of the reconstructed examples and the optimization of PSGAN, a scale factor named seeking radius is introduced to scale the generated perturbations before they are subtracted by the given adversarial examples. Comprehensive and extensive experimental results on three widely used benchmarks for RSI scene classification demonstrate the great effectiveness of PSGAN when faced with both known and unknown attacks. Our source code is available athttps://github.com/xuxiangsun/PSGAN.
Gong Cheng 0003, Xuxiang Sun 0001, Ke Li 0005, Lei Guo 0002, Junwei Han 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 Prototype-CNN for Few-Shot Object Detection in Remote Sensing Images
abstract
Recently, due to the excellent representation ability of convolutional neural networks (CNNs), object detection in remote sensing images has undergone remarkable development. However, when trained with a small number of samples, the performance of the object detectors drops sharply. In this article, we focus on the following three main challenges of few-shot object detection in remote sensing images: 1) since the sample number of novel classes is far less than base classes, object detectors would fail to quickly adapt to the features of novel classes, which would result in overfitting; 2) the scarcity of samples in novel classes leads to a sparse orientation space, while the objects in remote sensing images usually have arbitrary orientations; and 3) the distribution of object instances in remote sensing images is scattered and, therefore, it is hard to identify foreground objects from the complex background. To tackle these problems, we propose a simple yet effective method named prototype-CNN (P-CNN), which mainly consists of three parts: a prototype learning network (PLN) converting support images to class-aware prototypes, a prototype-guided region proposal network (P-G RPN) for better generation of region proposals, and a detector head extending the head of Faster region-based CNN (R-CNN) to further boost the performance. Comprehensive evaluations on the large-scale DIOR dataset demonstrate the effectiveness of our P-CNN. The source code is available athttps://github.com/Ybowei/P-CNN.
Gong Cheng 0003, Bowei Yan, Peizhen Shi, Ke Li 0005, Xiwen Yao, Lei Guo 0002, Junwei Han 0001
IEEE Trans. Geosci. Remote. Sens.6
2022 Multiscale Generative Adversarial Network Based on Wavelet Feature Learning for SAR-to-Optical Image Translation
abstract
The synthetic aperture radar (SAR) system is a kind of active remote sensing, which can be carried on a variety of flight platforms and can observe the Earth under all-day and all-weather conditions, so it has a wide range of applications. However, the interpretation of SAR images is quite challenging and not suitable for nonexperts. In order to enhance the visual effect of SAR images, this article proposes a multiscale generative adversarial network based on wavelet feature learning (WFLM-GAN) to implement the translation from SAR images to optical images; the translated images not only retain the key content of SAR images but also have the style of optical images. The main advantages of this method over the previous SAR-to-optical image translation (S2OIT) methods are given as follows. First, the generator does not learn the mapping from SAR images to optical images directly but learns the mapping from SAR images to wavelet features and then reconstructs the gray-scale images to optimize the content, increasing the mapping relationships and helping to learn more effective features. Second, a multiscale coloring network based on detail learning and style learning is designed to further translate the gray-scale images into optical images, which makes the generated images have an excellent visual effect with details closer to real images. Extensive experiments on SAR image datasets in different regions and seasons demonstrate the superior performance of WFLM-GAN over the baseline algorithms in terms of structural similarity (SSIM), the peak signal-to-noise ratio (PSNR), the Frechet inception distance (FID), and the kernel inception distance (KID). Comprehensive ablation studies are also carried out to isolate the validity of each proposed component. Our codes will be available athttps://github.com/G2022G/WFLM-GAN.
Cang Gu, Dongqing Wu, Gong Cheng 0003, Lei Guo 0002
IEEE Trans. Geosci. Remote. Sens.5
2021 Cross-Scale Feature Fusion for Object Detection in Optical Remote Sensing Images
abstract
For the time being, there are many groundbreaking object detection frameworks used in natural scene images. These algorithms have good detection performance on the data sets of open natural scenes. However, applying these frameworks to remote sensing images directly is not very effective. The existing deep-learning-based object detection algorithms still face some challenges when dealing with remote sensing images because these images usually contain a number of targets with large variations of object sizes as well as interclass similarity. Aiming at the challenges of object detection in optical remote sensing images, we propose an end-to-end cross-scale feature fusion (CSFF) framework, which can effectively improve the object detection accuracy. Specifically, we first use a feature pyramid network (FPN) to obtain multilevel feature maps and then insert a squeeze and excitation (SE) block into the top layer to model the relationship between different feature channels. Next, we use the CSFF module to obtain powerful and discriminative multilevel feature representations. Finally, we implement our work in the framework of Faster region-based CNN (R-CNN). In the experiment, we evaluate our method on a publicly available large-scale data set, named DIOR, and obtain an improvement of 3.0% measured in terms of mAP compared with Faster R-CNN with FPN.
Gong Cheng 0003, Yongjie Si, Hailong Hong, Xiwen Yao, Lei Guo 0002
IEEE Geosci. Remote. Sens. Lett.5
2021 Corrigendum to Identifying associations among genomic, proteomic and imaging biomarkers via adaptive sparse multi-view canonical correlation analysis [Medical Image Analysis 70 (2021) 1-12/102003]
Lei Du 0001, Jin Zhang 0023, Huiai Wang, Lei Guo 0002, Junwei Han 0001
Medical Image Anal.5
2021 Identifying associations among genomic, proteomic and imaging biomarkers via adaptive sparse multi-view canonical correlation analysis
Lei Du 0001, Jin Zhang 0023, Huiai Wang, Lei Guo 0002, Junwei Han 0001
Medical Image Anal.5
2021 Salient object detection using feature clustering and compactness prior
Yanbang Zhang, Fen Zhang, Lei Guo 0002, Henry Han
Multim. Tools Appl.3
2021 Multi-Task Sparse Canonical Correlation Analysis with Application to Multi-Modal Brain Imaging Genetics
abstract
Brain imaging genetics studies the genetic basis of brain structures and functionalities via integrating genotypic data such as single nucleotide polymorphisms (SNPs) and imaging quantitative traits (QTs). In this area, both multi-task learning (MTL) and sparse canonical correlation analysis (SCCA) methods are widely used since they are superior to those independent and pairwise univariate analysis. MTL methods generally incorporate a few of QTs and could not select features from multiple QTs; while SCCA methods typically employ one modality of QTs to study its association with SNPs. Both MTL and SCCA are computational expensive as the number of SNPs increases. In this paper, we propose a novel multi-task SCCA (MTSCCA) method to identify bi-multivariate associations between SNPs and multi-modal imaging QTs. MTSCCA could make use of the complementary information carried by different imaging modalities. MTSCCA enforces sparsity at the group level via the${\mathrm G}_{2,1}$-norm, and jointly selects features across multiple tasks for SNPs and QTs via the$\ell _{2,1}$-norm. A fast optimization algorithm is proposed using the grouping information of SNPs. Compared with conventional SCCA methods, MTSCCA obtains better correlation coefficients and canonical weights patterns. In addition, MTSCCA runs very fast and easy-to-implement, indicating its potential power in genome-wide brain-wide imaging genetics.
Lei Du 0001, Kefei Liu 0001, Xiaohui Yao, Shannon L. Risacher, Junwei Han 0001, Andrew J. Saykin, Lei Guo 0002, Li Shen 0001
IEEE ACM Trans. Comput. Biol. Bioinform.7
2021 DLA-MatchNet for Few-Shot Remote Sensing Image Scene Classification
abstract
Few-shot scene classification aims to recognize unseen scene concepts from few labeled samples. However, most existing works are generally inclined to learn metalearners or transfer knowledge while ignoring the importance to learn discriminative representations and a proper metric for remote sensing images. To address these challenges, in this article, we propose an end-to-end network for boosting a few-shot remote sensing image scene classification, called discriminative learning of adaptive match network (DLA-MatchNet). Specifically, we first adopt the attention technique to delve into the interchannel and interspatial relationships to automatically discover discriminative regions. Then, the channel attention and spatial attention modules can be incorporated with the feature network by using different feature fusion schemes, achieving “discriminative learning.” Afterward, considering the issues of the large intraclass variances and interclass similarity of remote sensing images, instead of simply computing the distances between the support samples and query samples, we concatenate the support and query discriminative features in depth and utilize a matcher to “adaptively” select the semantically relevant sample pairs to assign similarity scores. Our method leverages an episode-based strategy to train the model. Once trained, our model can predict the category of query image without further fine-tuning. Experimental results on three public remote sensing image data sets demonstrate the effectiveness of our model in the few-shot scene classification task.
Lingjun Li, Junwei Han 0001, Xiwen Yao, Gong Cheng 0003, Lei Guo 0002
IEEE Trans. Geosci. Remote. Sens.5
2021 Automatic Weakly Supervised Object Detection From High Spatial Resolution Remote Sensing Images via Dynamic Curriculum Learning
abstract
In this article, we focus on tackling the problem of weakly supervised object detection from high spatial resolution remote sensing images, which aims to learn detectors with only image-level annotations, i.e., without object location information during the training stage. Although promising results have been achieved, most approaches often fail to provide high-quality initial samples and thus are difficult to obtain optimal object detectors. To address this challenge, a dynamic curriculum learning strategy is proposed to progressively learn the object detectors by feeding training images with increasing difficulty that matches current detection ability. To this end, an entropy-based criterion is firstly designed to evaluate the difficulty for localizing objects in images. Then, an initial curriculum that ranks training images in ascending order of difficulty is generated, in which easy images are selected to provide reliable instances for learning object detectors. With the gained stronger detection ability, the subsequent order in the curriculum for retraining detectors is accordingly adjusted by promoting difficult images as easy ones. In such way, the detectors can be well prepared by training on easy images for learning from more difficult ones and thus gradually improve their detection ability more effectively. Moreover, an effective instance-aware focal loss function for detector learning is developed to alleviate the influence of positive instances of bad quality and meanwhile enhance the discriminative information of class-specific hard negative instances. Comprehensive experiments and comparisons with state-of-the-art methods on two publicly available data sets demonstrate the superiority of our proposed method.
Xiwen Yao, Xiaoxu Feng, Junwei Han 0001, Gong Cheng 0003, Lei Guo 0002
IEEE Trans. Geosci. Remote. Sens.5
2021 Eliminating Indefiniteness of Clinical Spectrum for Better Screening COVID-19
abstract
The coronavirus disease 2019 (COVID-19) has swept all over the world. Due to the limited detection facilities, especially in developing countries, a large number of suspected cases can only receive common clinical diagnosis rather than more effective detections like Reverse Transcription Polymerase Chain Reaction (RT-PCR) tests or CT scans. This motivates us to develop a quick screening method via common clinical diagnosis results. However, the diagnostic items of different patients may vary greatly, and there is a huge variation in the dimension of the diagnosis data among different suspected patients, it is hard to process these indefinite dimension data via classical classification algorithms. To resolve this problem, we propose an Indefiniteness Elimination Network (IE-Net) to eliminate the influence of the varied dimensions and make predictions about the COVID-19 cases. The IE-Net is in an encoder-decoder framework fashion, and an indefiniteness elimination operation is proposed to transfer the indefinite dimension feature into a fixed dimension feature. Comprehensive experiments were conducted on the public available COVID-19 Clinical Spectrum dataset. Experimental results show that the proposed indefiniteness elimination operation greatly improves the classification performance, the IE-Net achieves 94.80% accuracy, 92.79% recall, 92.97% precision and 94.93% AUC for distinguishing COVID-19 cases from non-COVID-19 cases with only common clinical diagnose data. We further compared our methods with 3 classical classification algorithms: random forest, gradient boosting and multi-layer perceptron (MLP). To explore each clinical test item's specificity, we further analyzed the possible relationship between each clinical test item and COVID-19.
Guangyu Guo 0001, Zhuoyan Liu, Shijie Zhao 0001, Lei Guo 0002, Tianming Liu 0001
IEEE J. Biomed. Health Informatics4
2020 Mining High-order Multimodal Brain Image Associations via Sparse Tensor Canonical Correlation Analysis
abstract
Neuroimaging techniques have shown increasing power to understand the neuropathology of brain disorders. Multimodal brain imaging data carry distinct but complementary information and thus could depict brain disorders comprehensively. To deepen our understanding, it is essential to investigate the intrinsic associations among multiple modalities. To date, the pairwise correlations between imaging data captured by different imaging modalities have been well studied, leaving formidable challenges to identify high-order associations. In this paper, we first propose a new sparse tensor canonical correlation analysis (STCCA) with feature selection to analyze the complex high-order relationships among multimodal brain imaging data. In addition, we find that methods for identifying pairwise associations and high-order associations have complementary advantages, providing a sound reason to fuse them. Therefore, we further propose an improved STCCA (STCCA+) which integrates STCCA and sparse multiple CCA (SMCCA) to fully uncover associations among multiple imaging modalities. The proposed STCCA+detects equivalent association levels among multimodal imaging data compared to SMCCA. Most importantly, both STCCA and STCCA+yield modality-consistent imaging markers and modality-specific ones, assuring a better and meaningful feature selection capability. Finally, the identified imaging markers and their high-order correlations could form a comprehensive indication of brain disorders, showing their promise in high-order multimodal brain imaging analysis.
Lei Du 0001, Jin Zhang 0023, Minjianan Zhang, Huiai Wang, Lei Guo 0002, Junwei Han 0001
BIBM6
2020 Species-Shared and -Specific Structural Connections Revealed by Dirty Multi-task Regression
Xi Jiang 0001, Lei Guo 0002, Xiaoping Hu 0001, Tianming Liu 0001, Lei Du 0001
MICCAI (7)4
2020 Identifying diagnosis-specific genotype-phenotype associations via joint multitask sparse canonical correlation analysis and classification
abstract
MOTIVATION: Brain imaging genetics studies the complex associations between genotypic data such as single nucleotide polymorphisms (SNPs) and imaging quantitative traits (QTs). The neurodegenerative disorders usually exhibit the diversity and heterogeneity, originating from which different diagnostic groups might carry distinct imaging QTs, SNPs and their interactions. Sparse canonical correlation analysis (SCCA) is widely used to identify bi-multivariate genotype-phenotype associations. However, most existing SCCA methods are unsupervised, leading to an inability to identify diagnosis-specific genotype-phenotype associations. RESULTS: In this article, we propose a new joint multitask learning method, named MT-SCCALR, which absorbs the merits of both SCCA and logistic regression. MT-SCCALR learns genotype-phenotype associations of multiple tasks jointly, with each task focusing on identifying one diagnosis-specific genotype-phenotype pattern. Meanwhile, MT-SCCALR cannot only select relevant SNPs and imaging QTs for each diagnostic group alone, but also allows the selection of those shared by multiple diagnostic groups. We derive an efficient optimization algorithm whose convergence to a local optimum is guaranteed. Compared with two state-of-the-art methods, MT-SCCALR yields better or similar canonical correlation coefficients and classification performances. In addition, it owns much better discriminative canonical weight patterns of great interest than competitors. This demonstrates the power and capability of MTSCCAR in identifying diagnostically heterogeneous genotype-phenotype patterns, which would be helpful to understand the pathophysiology of brain disorders. AVAILABILITY AND IMPLEMENTATION: The software is publicly available at https://github.com/dulei323/MTSCCALR. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Lei Du 0001, Kefei Liu 0001, Xiaohui Yao, Shannon L. Risacher, Junwei Han 0001, Lei Guo 0002, Andrew J. Saykin, Li Shen 0001
Bioinform.7
2020 Detecting genetic associations with brain imaging phenotypes in Alzheimer's disease via a novel structured SCCA approach
Lei Du 0001, Kefei Liu 0001, Xiaohui Yao, Shannon L. Risacher, Junwei Han 0001, Andrew J. Saykin, Lei Guo 0002, Li Shen 0001
Medical Image Anal.7
2020 Identifying Cross-individual Correspondences of 3-hinge Gyri
Ying Huang 0007, Lin Zhao 0004, Xi Jiang 0001, Lei Guo 0002, Xiaoping Hu 0001, Tianming Liu 0001
Medical Image Anal.6
2020 High-Quality Proposals for Weakly Supervised Object Detection
abstract
Despite significant efforts made so far for Weakly Supervised Object Detection (WSOD), proposal generation and proposal selection are still two major challenges. In this paper, we focus on addressing the two challenges by generating and selecting high-quality proposals. To be specific, for proposal generation, we combine selective search and a Gradient-weighted Class Activation Mapping (Grad-CAM) based technique to generate more proposals having higher Intersection-Over-Union (IOU) with ground truth boxes than those obtained by greedy search approaches, which can better envelop the entire objects. As regards proposal selection, for each object class, we choose as many confident positive proposals as possible and meanwhile only select class-specific hard negatives to focus training on more discriminative negative proposals by up-weighting their losses, which can make training more effective. The proposed proposal generation and proposal selection approaches are generic and thus can be broadly applied to many WSOD methods. In this work, we unify them into the framework of Online Instance Classifier Refinement (OICR). Experimental results on the PASCAL VOC 2007 and 2012 datasets and MS COCO dataset demonstrate that our method significantly improves the baseline method OICR by large margins (13.4% mAP and 11.6% CorLoc gains on the VOC 2007 dataset, 15.0% mAP and 8.9% CorLoc gains on the VOC 2012 dataset, and 6.4% mAP and 5.0% CorLoc gains on the COCO dataset) and achieves the state-of-the-art results compared with existing methods.
Gong Cheng 0003, Junyu Yang, Decheng Gao, Lei Guo 0002, Junwei Han 0001
IEEE Trans. Image Process.4
2019 Performance Comparison of Two Pooling Strategies for Remote Sensing Image Scene Classification
abstract
With the advances of convolutional neural networks (CNNs), the accuracy of remote sensing image scene classification has been greatly boosted thanks to the powerful features extracted through CNNs. Although significant success has been achieved, most of existing methods are dominated by the use of fully-connected CNN features. This paper focuses on the performance comparison of two kinds of novel pooling strategies, including generalized max pooling (GMP) and taskdriven pooling (TDP), for remote sensing image scene classification. To this end, an off-the-shelf CNN model is used as backbone network to extract multi-scale convolutional features. Then, GMP and TDP are respectively adopted to obtain globally pooled features. Finally, scene classification is performed with support vector machine (SVM). In the experiment, we evaluate the performance of these two kinds of pooling schemes on a widely-used scene classification benchmark data set. The experimental results show that (i) using pooled CNN convolutional features can obtain better results than using fully-connected CNN features and (ii) TDP is slightly better than GMP.
Maoxiong Wu, Gong Cheng 0003, Xiwen Yao, Xiaoliang Qian, Junwei Han 0001, Lei Guo 0002
IGARSS6
2019 Learning Region Response Ranking Features for Remote Sensing Image Scene Classification
abstract
Recently, deep learning especially convolutional neural networks (CNNs) has huge great success for remote sensing image scene classification. However, global CNN features still lack geometric invariance for addressing the problem of large intra-class variations and so are not optimal for scene classification. In this paper, we introduce a new feature representation for scene classification, named region response ranking (3R) feature representations by using off-the-shelf CNN models. Specifically, by considering each cube pixel of a certain convolutional feature map as one image region, we jointly train a class-specific support vector machine (SVM) base classifier and a decision function for each scene class. The base classifier is used to generate 3R feature by reordering the SVM responses of all image regions in descending order and the decision function is used for classification with 3R feature representations. Comprehensive evaluations on the publicly available NWPU-RESISC45 data set and comparisons with state-of-the-art methods demonstrate that the proposed 3R feature is effective for remote sensing image scene classification.1
Junyu Yang, Gong Cheng 0003, Xiwen Yao, Junwei Han 0001, Lei Guo 0002
IGARSS5
2019 Rotation-Invariant Latent Semantic Representation Learning for Object Detection in VHR Optical Remote Sensing Images
abstract
Object detection in very high resolution (VHR) optical remote sensing images is a fundamental yet challenging problem for the field of remote sensing image analysis. The detection performance is heavily dependent on the representation capability of the extracted features. Recently, convolutional neural networks (CNNs) have made a breakthrough for various applications in nature images. However, it is problematic to directly apply CNN to perform object detection in VHR optical remote sensing images due to the problem of object rotation variations. To address this issue, a novel rotation invariant probabilistic Latent Semantic Analysis (RI-pLSA) model is proposed to learn latent semantic representations for object detection. This is achieved by imposing a rotation-invariant regularization term on the objective function of pLSA to enforce the learned representation from all rotations of the same sample to be as consistent as possible. Additionally, the proposed RI-pLSA model takes the CNN features as input, which generates more powerful semantic representation for object detection. Comprehensive experiments on a publicly available ten-class object detection dataset demonstrate the superiority and effectiveness of our method compared with state-of-the-arts.
Xiwen Yao, Xiaoxu Feng, Gong Cheng 0003, Junwei Han 0001, Lei Guo 0002
IGARSS5
2019 Scene Classification of High Resolution Remote Sensing Images Via Self-Paced Deep Learning
abstract
Scene classification of high resolution remote sensing (HRRS) images is a fundamental yet challenging problem for remote sensing image analysis. In this paper, we focus on tackling the problem of HRSS scene classification using a small pool of unlabeled images and only a few labeled images per category, namely, few-shot scene classification (FSSC), which is more challenging than common scene classification task. The key challenge arises from selecting trustworthy samples from the pool of unlabeled images that have high confidence. To address this challenge, a novel local manifold constrained self-paced deep learning method is proposed. Specifically, the model is learned by gradually selecting easy samples from the pool of unlabeled images, assigning them with pseudo-labels and further adopting them with labeled images as the new training set. In addition, a local manifold constraint is introduced to enforce that the pseudo-labels assigned by the initial model should be consistent with the local manifold of the labeled samples. In such way, the confidence of the selecting samples is increased and is beneficial to train more robust classifier. Experimental results on a publicly available large scale NWPU-RESISC45 data set demonstrated the effectiveness of our method in achieving competitive performance while significantly reducing manually labeled cost.
Xiwen Yao, Gong Cheng 0003, Junwei Han 0001, Lei Guo 0002
IGARSS5
2019 A Dirty Multi-task Learning Method for Multi-modal Brain Imaging Genetics
Lei Du 0001, Kefei Liu 0001, Xiaohui Yao, Shannon L. Risacher, Junwei Han 0001, Lei Guo 0002, Andrew J. Saykin, Li Shen 0001
MICCAI (4)7
2019 Multi-view Graph Matching of Cortical Landmarks
Ying Huang 0007, Lei Guo 0002, Tianming Liu 0001
MICCAI (4)3
2019 Group-Wise Graph Matching of Cortical Gyral Hinges
Xiao Li 0024, Lin Zhao 0004, Ying Huang 0007, Lei Guo 0002, Tianming Liu 0001
MICCAI (4)6
2019 Identifying progressive imaging genetic patterns via multi-task sparse canonical correlation analysis: a longitudinal study of the ADNI cohort
abstract
MOTIVATION: Identifying the genetic basis of the brain structure, function and disorder by using the imaging quantitative traits (QTs) as endophenotypes is an important task in brain science. Brain QTs often change over time while the disorder progresses and thus understanding how the genetic factors play roles on the progressive brain QT changes is of great importance and meaning. Most existing imaging genetics methods only analyze the baseline neuroimaging data, and thus those longitudinal imaging data across multiple time points containing important disease progression information are omitted. RESULTS: We propose a novel temporal imaging genetic model which performs the multi-task sparse canonical correlation analysis (T-MTSCCA). Our model uses longitudinal neuroimaging data to uncover that how single nucleotide polymorphisms (SNPs) play roles on affecting brain QTs over the time. Incorporating the relationship of the longitudinal imaging data and that within SNPs, T-MTSCCA could identify a trajectory of progressive imaging genetic patterns over the time. We propose an efficient algorithm to solve the problem and show its convergence. We evaluate T-MTSCCA on 408 subjects from the Alzheimer's Disease Neuroimaging Initiative database with longitudinal magnetic resonance imaging data and genetic data available. The experimental results show that T-MTSCCA performs either better than or equally to the state-of-the-art methods. In particular, T-MTSCCA could identify higher canonical correlation coefficients and capture clearer canonical weight patterns. This suggests that T-MTSCCA identifies time-consistent and time-dependent SNPs and imaging QTs, which further help understand the genetic basis of the brain QT changes over the time during the disease progression. AVAILABILITY AND IMPLEMENTATION: The software and simulation data are publicly available at https://github.com/dulei323/TMTSCCA. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Lei Du 0001, Kefei Liu 0001, Lei Zhu 0011, Xiaohui Yao, Shannon L. Risacher, Lei Guo 0002, Andrew J. Saykin, Li Shen 0001
Bioinform.6
2019 Identifying Brain Networks at Multiple Time Scales via Deep Recurrent Neural Network
abstract
For decades, task functional magnetic resonance imaging has been a powerful noninvasive tool to explore the organizational architecture of human brain function. Researchers have developed a variety of brain network analysis methods for task fMRI data, including the general linear model, independent component analysis, and sparse representation methods. However, these shallow models are limited in faithful reconstruction and modeling of the hierarchical and temporal structures of brain networks, as demonstrated in more and more studies. Recently, recurrent neural networks (RNNs) exhibit great ability of modeling hierarchical and temporal dependence features in the machine learning field, which might be suitable for task fMRI data modeling. To explore such possible advantages of RNNs for task fMRI data, we propose a novel framework of a deep recurrent neural network (DRNN) to model the functional brain networks from task fMRI data. Experimental results on the motor task fMRI data of Human Connectome Project 900 subjects release demonstrated that the proposed DRNN can not only faithfully reconstruct functional brain networks, but also identify more meaningful brain networks with multiple time scales which are overlooked by traditional shallow models. In general, this work provides an effective and powerful approach to identifying functional brain networks at multiple time scales from task fMRI data.
Yan Cui 0005, Shijie Zhao 0001, Han Wang 0012, Yaowu Chen, Junwei Han 0001, Lei Guo 0002, Fan Zhou 0007, Tianming Liu 0001
IEEE J. Biomed. Health Informatics7
2018 Fast Multi-Task SCCA Learning with Feature Selection for Multi-Modal Brain Imaging Genetics
Lei Du 0001, Kefei Liu 0001, Xiaohui Yao, Shannon L. Risacher, Junwei Han 0001, Lei Guo 0002, Andrew J. Saykin, Li Shen 0001
BIBM6
2018 Identifying Brain Networks of Multiple Time Scales via Deep Recurrent Neural Network
Yan Cui 0005, Shijie Zhao 0001, Han Wang 0012, Yaowu Chen, Junwei Han 0001, Lei Guo 0002, Fan Zhou 0007, Tianming Liu 0001
MICCAI (3)7
2018 Identification of Species-Preserved Cortical Landmarks
Xiao Li 0024, Lin Zhao 0004, Ying Huang 0007, Lei Guo 0002, Tianming Liu 0001
MICCAI (3)5
2018 A novel SCCA approach via truncated ℓ1-norm and truncated group lasso for brain imaging genetics
abstract
MOTIVATION: Brain imaging genetics, which studies the linkage between genetic variations and structural or functional measures of the human brain, has become increasingly important in recent years. Discovering the bi-multivariate relationship between genetic markers such as single-nucleotide polymorphisms (SNPs) and neuroimaging quantitative traits (QTs) is one major task in imaging genetics. Sparse Canonical Correlation Analysis (SCCA) has been a popular technique in this area for its powerful capability in identifying bi-multivariate relationships coupled with feature selection. The existing SCCA methods impose either the ℓ1-norm or its variants to induce sparsity. The ℓ0-norm penalty is a perfect sparsity-inducing tool which, however, is an NP-hard problem. RESULTS: In this paper, we propose the truncated ℓ1-norm penalized SCCA to improve the performance and effectiveness of the ℓ1-norm based SCCA methods. Besides, we propose an efficient optimization algorithms to solve this novel SCCA problem. The proposed method is an adaptive shrinkage method via tuning τ. It can avoid the time intensive parameter tuning if given a reasonable small τ. Furthermore, we extend it to the truncated group-lasso (TGL), and propose TGL-SCCA model to improve the group-lasso-based SCCA methods. The experimental results, compared with four benchmark methods, show that our SCCA methods identify better or similar correlation coefficients, and better canonical loading profiles than the competing methods. This demonstrates the effectiveness and efficiency of our methods in discovering interesting imaging genetic associations. AVAILABILITY AND IMPLEMENTATION: The Matlab code and sample data are freely available at http://www.iu.edu/∼shenlab/tools/tlpscca/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Lei Du 0001, Kefei Liu 0001, Xiaohui Yao, Shannon L. Risacher, Junwei Han 0001, Lei Guo 0002, Andrew J. Saykin, Li Shen 0001
Bioinform.8
2018 Identifying affective levels on music video via completing the missing modality
Gong Cheng 0003, Lei Guo 0002
Multim. Tools Appl.3
2018 Discriminative Joint-Feature Topic Model With Dual Constraints for WCE Classification
abstract
Wireless capsule endoscopy (WCE) enables clinicians to examine the digestive tract without any surgical operations, at the cost of a large amount of images to be analyzed. The main challenge for automatic computer-aided diagnosis arises from the difficulty of robust characterization of these images. To tackle this problem, a novel discriminative joint-feature topic model (DJTM) with dual constraints is proposed to classify multiple abnormalities in WCE images. We first propose a joint-feature probabilistic latent semantic analysis (PLSA) model, where color and texture descriptors extracted from same image patches are jointly modeled with their conditional distributions. Then the proposed dual constraints: visual words importance and local image manifold are embedded into the joint-feature PLSA model simultaneously to obtain discriminative latent semantic topics. The visual word importance is proposed in our DJTM to guarantee that visual words with similar importance come from close latent topics while the local image manifold constraint enforces that images within the same category share similar latent topics. Finally, each image is characterized by distribution of latent semantic topics instead of low level features. Our proposed DJTM showed an excellent overall recognition accuracy 90.78%. Comprehensive comparison results demonstrate that our method outperforms existing multiple abnormalities classification methods for WCE images.
Yixuan Yuan, Xiwen Yao, Junwei Han 0001, Lei Guo 0002, Max Q.-H. Meng
IEEE Trans. Cybern.4
2018 Exploring Hierarchical Convolutional Features for Hyperspectral Image Classification
abstract
Hyperspectral image (HSI) classification is an active and important research task driven by many practical applications. To leverage deep learning models especially convolutional neural networks (CNNs) for HSI classification, this paper proposes a simple yet effective method to extract hierarchical deep spatial feature for HSI classification by exploring the power of off-the-shelf CNN models, without any additional retraining or fine-tuning on the target data set. To obtain better classification accuracy, we further propose a unified metric learning-based framework to alternately learn discriminative spectral-spatial features, which have better representation capability and train support vector machine (SVM) classifiers. To this end, we design a new objective function that explicitly embeds a metric learning regularization term into SVM training. The metric learning regularization term is used to learn a powerful spectral-spatial feature representation by fusing spectral feature and deep spatial feature, which has small intraclass scatter but big between class separation. By transforming HSI data into new spectral-spatial feature space through CNN and metric learning, we can pull the pixels from the same class closer, while pushing the different class pixels farther away. In the experiments, we comprehensively evaluate the proposed method on three commonly used HSI benchmark data sets. State-of-the-art results are achieved when compared with the existing HSI classification methods.
Gong Cheng 0003, Junwei Han 0001, Xiwen Yao, Lei Guo 0002
IEEE Trans. Geosci. Remote. Sens.5
2018 When Deep Learning Meets Metric Learning: Remote Sensing Image Scene Classification via Learning Discriminative CNNs
abstract
Remote sensing image scene classification is an active and challenging task driven by many applications. More recently, with the advances of deep learning models especially convolutional neural networks (CNNs), the performance of remote sensing image scene classification has been significantly improved due to the powerful feature representations learnt through CNNs. Although great success has been obtained so far, the problems of within-class diversity and between-class similarity are still two big challenges. To address these problems, in this paper, we propose a simple but effective method to learn discriminative CNNs (D-CNNs) to boost the performance of remote sensing image scene classification. Different from the traditional CNN models that minimize only the cross entropy loss, our proposed D-CNN models are trained by optimizing a new discriminative objective function. To this end, apart from minimizing the classification error, we also explicitly impose a metric learning regularization term on the CNN features. The metric learning regularization enforces the D-CNN models to be more discriminative so that, in the new D-CNN feature spaces, the images from the same scene class are mapped closely to each other and the images of different classes are mapped as farther apart as possible. In the experiments, we comprehensively evaluate the proposed method on three publicly available benchmark data sets using three off-the-shelf CNN models. Experimental results demonstrate that our proposed D-CNN methods outperform the existing baseline methods and achieve state-of-the-art results on all three data sets.
Gong Cheng 0003, Ceyuan Yang, Xiwen Yao, Lei Guo 0002, Junwei Han 0001
IEEE Trans. Geosci. Remote. Sens.4
2018 Modeling Task fMRI Data Via Deep Convolutional Autoencoder
abstract
Task-based functional magnetic resonance imaging (tfMRI) has been widely used to study functional brain networks under task performance. Modeling tfMRI data is challenging due to at least two problems: the lack of the ground truth of underlying neural activity and the highly complex intrinsic structure of tfMRI data. To better understand brain networks based on fMRI data, data-driven approaches have been proposed, for instance, independent component analysis (ICA) and sparse dictionary learning (SDL). However, both ICA and SDL only build shallow models, and they are under the strong assumption that original fMRI signal could be linearly decomposed into time series components with their corresponding spatial maps. As growing evidence shows that human brain function is hierarchically organized, new approaches that can infer and model the hierarchical structure of brain networks are widely called for. Recently, deep convolutional neural network (CNN) has drawn much attention, in that deep CNN has proven to be a powerful method for learning high-level and mid-level abstractions from low-level raw data. Inspired by the power of deep CNN, in this paper, we developed a new neural network structure based on CNN, called deep convolutional auto-encoder (DCAE), in order to take the advantages of both data-driven approach and CNN's hierarchical feature abstraction ability for the purpose of learning mid-level and high-level features from complex, large-scale tfMRI time series in an unsupervised manner. The DCAE has been applied and tested on the publicly available human connectome project tfMRI data sets, and promising results are achieved.
Heng Huang 0001, Xintao Hu, Yu Zhao 0007, Milad Makkie, Qinglin Dong, Shijie Zhao 0001, Lei Guo 0002, Tianming Liu 0001
IEEE Trans. Medical Imaging7
2017 Multi-way Regression Reveals Backbone of Macaque Structural Brain Connectivity in Longitudinal Datasets
Xiao Li 0024, Lin Zhao 0004, Xintao Hu, Tianming Liu 0001, Lei Guo 0002
MICCAI (1)6
2017 Remote Sensing Image Scene Classification Using Bag of Convolutional Features
abstract
More recently, remote sensing image classification has been moving from pixel-level interpretation to scene-level semantic understanding, which aims to label each scene image with a specific semantic class. While significant efforts have been made in developing various methods for remote sensing image scene classification, most of them rely on handcrafted features. In this letter, we propose a novel feature representation method for scene classification, named bag of convolutional features (BoCF). Different from the traditional bag of visual words-based methods in which the visual words are usually obtained by using handcrafted feature descriptors, the proposed BoCF generates visual words from deep convolutional features using off-the-shelf convolutional neural networks. Extensive evaluations on a publicly available remote sensing image scene classification benchmark and comparison with the state-of-the-art methods demonstrate the effectiveness of the proposed BoCF method for remote sensing image scene classification.
Gong Cheng 0003, Xiwen Yao, Lei Guo 0002, Zhongliang Wei
IEEE Geosci. Remote. Sens. Lett.4
2017 Task fMRI data analysis based on supervised stochastic coordinate coding
Jinglei Lv, Qingyang Li 0001, Wei Zhang 0090, Yu Zhao 0007, Xi Jiang 0001, Lei Guo 0002, Junwei Han 0001, Xintao Hu, Christine Cong Guo, Jieping Ye, Tianming Liu 0001
Medical Image Anal.7
2016 Sparse Canonical Correlation Analysis via truncated ℓ1-norm with application to brain imaging genetics
abstract
Discovering bi-multivariate associations between genetic markers and neuroimaging quantitative traits is a major task in brain imaging genetics. Sparse Canonical Correlation Analysis (SCCA) is a popular technique in this area for its powerful capability in identifying bi-multivariate relationships coupled with feature selection. The existing SCCA methods impose either the ℓ1-norm or its variants. The ℓ0-norm is more desirable, which however remains unexplored since the ℓ0-norm minimization is NP-hard. In this paper, we impose the truncated ℓ1-norm to improve the performance of the ℓ1-norm based SCCA methods. Besides, we propose two efficient optimization algorithms and prove their convergence. The experimental results, compared with two benchmark methods, show that our method identifies better and meaningful canonical loading patterns in both simulated and real imaging genetic analyse.
Lei Du 0001, Kefei Liu 0001, Xiaohui Yao, Shannon L. Risacher, Lei Guo 0002, Andrew J. Saykin, Li Shen 0001
BIBM7
2016 Exploring auditory network composition during free listening to audio excerpts via group-wise sparse representation
abstract
With the growing number of audio excerpts through various media and distribution channels, advanced audio analysis approaches have received significant interest in the multimedia field. However, current audio analysis approaches are still far from satisfactory due to the semantic gaps between the low-level acoustic features and high-level semantics perceived by human brain. In order to alleviate the problem, this paper propose a novel computational framework to bridge acoustic features with high-level semantic features derived from functional magnetic resonance imaging (fMRI) signals which record the brain's response during free listening to music/speech excerpts, and to explore the brain auditory network composition of acoustic features for different types of music/speech excerpts. Specifically, we identify meaningful brain networks and corresponding brain activities representing high-level semantic features via a novel group-wise sparse representation of whole brain fMRI signals. Then we associate the brain activities with specific low-level acoustic features and analyze the auditory network composition of acoustic features for different types of music/speech excerpts. Experimental results demonstrate that multiple acoustic features are involved in the brain auditory networks during free listening to music/speech excerpts. Meanwhile, there is considerable variability of auditory network composition of acoustic features for different types of music/speech. Our results provide new insights of how to narrow the semantic gaps in audio content analysis.
Shijie Zhao 0001, Junwei Han 0001, Xi Jiang 0001, Xintao Hu, Jinglei Lv, Shu Zhang 0001, Bao Ge, Lei Guo 0002, Tianming Liu 0001
ICME8
2016 Semantic annotation of satellite images via joint multi-feature learning with diversity constraint
abstract
Automatic semantic annotation of high-resolution optical satellite images is a task to assign one or several predefined semantic concepts to an image according to its content. The fundamental challenge arises from the difficulty of characterizing complex and ambiguous contents of the satellite images. To address this challenge, a diversity constrained joint multi-feature learning method is proposed to learn robust feature representations for annotating satellite images. The key motivation of our method is to make full use of the complementarity diversity information among the heterogeneous features in the learning process. Comprehensive experiments on an annotation dataset demonstrate the superiority and effectiveness of our method compared with baseline multi-feature learning method.
Xiwen Yao, Junwei Han 0001, Gong Cheng 0003, Peicheng Zhou, Lei Guo 0002
IGARSS5
2016 Species Preserved and Exclusive Structural Connections Revealed by Sparse CCA
Xiao Li 0024, Lei Du 0001, Xintao Hu, Xi Jiang 0001, Lei Guo 0002, Tianming Liu 0001
MICCAI (1)6
2016 Temporal Concatenated Sparse Coding of Resting State fMRI Data Reveal Network Interaction Changes in mTBI
Jinglei Lv, Armin Iraji, Fangfei Ge, Shijie Zhao 0001, Xintao Hu, Junwei Han 0001, Lei Guo 0002, Zhifeng Kou, Tianming Liu 0001
MICCAI (1)8
2016 A Multi-stage Sparse Coding Framework to Explore the Effects of Prenatal Alcohol Exposure
Shijie Zhao 0001, Junwei Han 0001, Jinglei Lv, Xi Jiang 0001, Xintao Hu, Shu Zhang 0001, Mary Ellen Lynch, Claire Coles, Lei Guo 0002, Xiaoping Hu 0001, Tianming Liu 0001
MICCAI (1)9
2016 Saliency detection by selective color features
Yanbang Zhang, Fen Zhang, Lei Guo 0002
Neurocomputing3
2016 Group-wise consistent cortical parcellation based on connectional profiles
Dajiang Zhu, Xi Jiang 0001, Shu Zhang 0001, Zhifeng Kou, Lei Guo 0002, Tianming Liu 0001
Medical Image Anal.6
2016 Predicting Movie Trailer Viewer's "Like/Dislike" via Learned Shot Editing Patterns
abstract
Nowadays, there are many movie trailers publicly available on social media website such as YouTube, and many thousands of users have independently indicated whether they like or dislike those trailers. Although it is understandable that there are multiple factors that could influence viewers' like or dislike of the trailer, we aim to address a preference question in this work: Can subjective multimedia features be developed to predict the viewer's preference presented by like (by thumbs-up) or dislike (by thumbs-down) during and after watching movie trailers? We designed and implemented a computational framework that is composed of low-level multimedia feature extraction, feature screening and selection, and classification, and applied it to a collection of 725 movie trailers. Experimental results demonstrated that, among dozens of multimedia features, the single low-level multimedia feature of shot length variance is highly predictive of a viewer's “like/dislike” for a large portion of movie trailers. We interpret these findings such that variable shot lengths in a trailer tend to produce a rhythm that is likely to stimulate a viewer's positive preference. This conclusion was also proved by the repeatability experiments results using another 600 trailer videos and it was further interpreted by viewers'eye-tracking data.
Shu Zhang 0001, Xi Jiang 0001, Xiang Li 0001, Xintao Hu, Junwei Han 0001, Lei Guo 0002, L. Stephen Miller, Richard Neupert, Tianming Liu 0001
IEEE Trans. Affect. Comput.8
2016 Two-Stage Learning to Predict Human Eye Fixations via SDAEs
abstract
Saliency detection models aiming to quantitatively predict human eye-attended locations in the visual field have been receiving increasing research interest in recent years. Unlike traditional methods that rely on hand-designed features and contrast inference mechanisms, this paper proposes a novel framework to learn saliency detection models from raw image data using deep networks. The proposed framework mainly consists of two learning stages. At the first learning stage, we develop a stacked denoising autoencoder (SDAE) model to learn robust, representative features from raw image data under an unsupervised manner. The second learning stage aims to jointly learn optimal mechanisms to capture the intrinsic mutual patterns as the feature contrast and to integrate them for final saliency prediction. Given the input of pairs of a center patch and its surrounding patches represented by the features learned at the first stage, a SDAE network is trained under the supervision of eye fixation labels, which achieves both contrast inference and contrast integration simultaneously. Experiments on three publically available eye tracking benchmarks and the comparisons with 16 state-of-the-art approaches demonstrate the effectiveness of the proposed framework.
Junwei Han 0001, Dingwen Zhang, Shifeng Wen, Lei Guo 0002, Tianming Liu 0001, Xuelong Li 0001
IEEE Trans. Cybern.4
2016 Semantic Annotation of High-Resolution Satellite Images via Weakly Supervised Learning
abstract
In this paper, we focus on tackling the problem of automatic semantic annotation of high resolution (HR) optical satellite images, which aims to assign one or several predefined semantic concepts to an image according to its content. The main challenges arise from the difficulty of characterizing complex and ambiguous contents of the satellite images and the high human labor cost caused by preparing a large amount of training examples with high-quality pixel-level labels in fully supervised annotation methods. To address these challenges, we propose a unified annotation framework by combining discriminative high-level feature learning and weakly supervised feature transferring. Specifically, an efficient stacked discriminative sparse autoencoder (SDSAE) is first proposed to learn high-level features on an auxiliary satellite image data set for the land-use classification task. Inspired by the motivation that the encoder of the prelearned SDSAE can be regarded as a generic high-level feature extractor for HR optical satellite images, we then transfer the learned high-level features to semantic annotation. To compensate the difference between the auxiliary data set and the annotation data set, the transferred high-level features are further fine-tuned in a weakly supervised scheme by using the tile-level annotated training data. Finally, the fine-tuning process is formulated as an ultimate optimization problem, which can be solved efficiently with our proposed alternate iterative optimization method. Comprehensive experiments on a publicly available land-use classification data set and an annotation data set demonstrate the superiority of our SDSAE-based high-level feature learning method and the effectiveness of our weakly supervised semantic annotation framework compared with state-of-the-art fully supervised annotation methods.
Xiwen Yao, Junwei Han 0001, Gong Cheng 0003, Xueming Qian, Lei Guo 0002
IEEE Trans. Geosci. Remote. Sens.5
2015 Identifying valence and arousal levels via connectivity between EEG channels
abstract
Implicit emotion tagging is a central theme in the area of affective computing. To this end, Several physiological signals acquired from subjects can be employed, for example, electroencephalography (EEG) and functional magnetic resonance imaging (fMRI) from brain, electrocardiography (ECG) from cardiac activities, and other peripheral physiological signals, such as galvanic skin resistance, electromyogram (EMG), blood volume pressure etc. Brain is regarded as the place where emotional activities evoke. Determining affective states by observing brain activities directly is of therefore great interest. There are several published works that use EEG signals to identify affective states in different aspects with various stimuli, e.s., images, musics and videos. In this paper, we propose to adopt EEG connectivity between electrodes to identify subjects' affective levels in both valence and arousal space during video stimuli presentation. Three catagories of connectivity are adopted in magnitude and phase domains. One open accessed affective database, DEAP, is used as benchmark. We will show that with the proposed connectivity-based representation, the accuracy of affective levels identification tasks are higher than the same tasks in existing works based on same database.
Junwei Han 0001, Lei Guo 0002, Ioannis Patras
ACII3
2015 Learning coarse-to-fine sparselets for efficient object detection and scene classification
abstract
Part model-based methods have been successfully applied to object detection and scene classification and have achieved state-of-the-art results. More recently the “sparselets” work [1-3] were introduced to serve as a universal set of shared basis learned from a large number of part detectors, resulting in notable speedup. Inspired by this framework, in this paper, we propose a novel scheme to train more effective sparselets with a coarse-to-fine framework. Specifically, we first train coarse sparselets to exploit the redundancy existing among part detectors by using an unsupervised single-hidden-layer auto-encoder. Then, we simultaneously train fine sparselets and activation vectors using a supervised single-hidden-layer neural network, in which sparselets training and discriminative activation vectors learning are jointly embedded into a unified framework. In order to adequately explore the discriminative information hidden in the part detectors and to achieve sparsity, we propose to optimize a new discriminative objective function by imposing L0-norm sparsity constraint on the activation vectors. By using the proposed framework, promising results for multi-class object detection and scene classification are achieved on PASCAL VOC 2007, MIT Scene-67, and UC Merced Land Use datasets, compared with the existing sparselets baseline methods.
Gong Cheng 0003, Junwei Han 0001, Lei Guo 0002, Tianming Liu 0001
CVPR3
2015 Fiber Connection Pattern-Guided Structured Sparse Representation of Whole-Brain fMRI Signals for Functional Network Inference
Xi Jiang 0001, Jianfeng Lu 0003, Lei Guo 0002, Tianming Liu 0001
MICCAI (1)5
2015 Modeling Task FMRI Data via Supervised Stochastic Coordinate Coding
Jinglei Lv, Wei Zhang 0090, Xi Jiang 0001, Xintao Hu, Junwei Han 0001, Lei Guo 0002, Jieping Ye, Tianming Liu 0001
MICCAI (1)7
2015 Multi-scale and Multimodal Fusion of Tract-Tracing, Myelin Stain and DTI-derived Fibers in Macaque Brains
Ke Jing, Hanbo Chen, Xi Jiang 0001, Longchuan Li, Lei Guo 0002, Jianfeng Lu 0003, Xiaoping Hu 0001, Tianming Liu 0001
MICCAI (2)7
2015 Semantic Segmentation based on Stacked Discriminative Autoencoders and Context-Constrained Weakly Supervised Learning
abstract
In this paper, we focus on tacking the problem of weakly supervised semantic segmentation. The aim is to predict the class label of image regions under weakly supervised settings, where training images are only provided with image-level labels indicating the classes they contain. The main difficulty of weakly supervised semantic segmentation arises from the complex diversity of visual classes and the lack of supervision information for learning a multi-classes classifier. To conquer the challenge, we propose a novel discriminative deep feature learning framework based on stacked autoencoders (SAE) by integrating pairwise constraints to serve as a discriminative term. Furthermore, to mine effective supervision information, global context about co-occurrence of visual classes as well as local context around each image region is exploited as constraints for training a multi-class classifier. Finally, the classifier training is formulated as an ultimate optimization problem, which can be solved efficiently by an alternate iterative optimization method. Comprehensive experiments on the MSRC 21 dataset demonstrate the superior performance compared with several state-of-the-art weakly supervised image segmentation methods.
Xiwen Yao, Junwei Han 0001, Gong Cheng 0003, Lei Guo 0002
ACM Multimedia4
2015 Auto-encoder-based shared mid-level visual dictionary learning for scene classification using very high resolution remote sensing images
abstract
Effective representation and classification of scenes using very high resolution (VHR) remote sensing images cover a wide range of applications. Although robust low‐level image features have been proven to be effective for scene classification, they are not semantically meaningful and thus have difficulty to deal with challenging visual recognition tasks. In this study, the authors propose a new and effective auto‐encoder‐based method to learn a shared mid‐level visual dictionary. This dictionary serves as a shared and universal basis to discover mid‐level visual elements. On the one hand, the mid‐level visual dictionary learnt using machine learning technique is more discriminative and contains rich semantic information, compared with the traditional low‐level visual words. On the other hand, the mid‐level visual dictionary is more robust to occlusions and image clutters. In the authors' scene‐classification scheme, they use discriminative mid‐level visual elements, rather than individual pixels or low‐level image features, to represent images. This new image representation is able to capture much of the high‐level meaning and contents of the image, facilitating challenging remote sensing image scene‐classification tasks. Comprehensive evaluations on a challenging VHR remote sensing images data set and comparisons with state‐of‐the‐art approaches demonstrate the effectiveness and superiority of their study.
Gong Cheng 0003, Peicheng Zhou, Junwei Han 0001, Lei Guo 0002, Jungong Han
IET Comput. Vis.4
2015 A coarse-to-fine model for airport detection from remote sensing images using target-oriented visual saliency and CRF
Xiwen Yao, Junwei Han 0001, Lei Guo 0002, Shuhui Bu, Zhenbao Liu
Neurocomputing3
2015 Analysis of music/speech via integration of audio content and functional brain response
Junwei Han 0001, Xi Jiang 0001, Xintao Hu, Lei Guo 0002, Jungong Han, Ling Shao 0001, Tianming Liu 0001
Inf. Sci.5
2015 Weakly Supervised Learning for Target Detection in Remote Sensing Images
abstract
In this letter, we develop a novel framework of leveraging weakly supervised learning techniques to efficiently detect targets from remote sensing images, which enables us to reduce the tedious manual annotation for collecting training data while maintaining the detection accuracy to large extent. The proposed framework consists of a weakly supervised training procedure to yield the detectors and an effective scheme to detect targets from testing images. Comprehensive evaluations on three benchmarks which have different spatial resolutions and contain different types of targets as well as the comparisons with traditional supervised learning schemes demonstrate the efficiency and effectiveness of the proposed framework.
Dingwen Zhang, Junwei Han 0001, Gong Cheng 0003, Zhenbao Liu, Shuhui Bu, Lei Guo 0002
IEEE Geosci. Remote. Sens. Lett.6
2015 Sparse representation of whole-brain fMRI signals for identification of functional networks
Jinglei Lv, Xi Jiang 0001, Xiang Li 0001, Dajiang Zhu, Hanbo Chen, Shu Zhang 0001, Xintao Hu, Junwei Han 0001, Heng Huang 0001, Jing Zhang 0010, Lei Guo 0002, Tianming Liu 0001
Medical Image Anal.12
2015 Arousal Recognition Using Audio-Visual Features and FMRI-Based Brain Response
abstract
As the indicator of emotion intensity, arousal is a significant clue for users to find their interested content. Hence, effective techniques for video arousal recognition are highly required. In this paper, we propose a novel framework for recognizing arousal levels by integrating low-level audio-visual features derived from video content and human brain's functional activity in response to videos measured by functional magnetic resonance imaging (fMRI). At first, a set of audio-visual features which have been demonstrated to be correlated with video arousal are extracted. Then, the fMRI-derived features that convey the brain activity of comprehending videos are extracted based on a number of brain regions of interests (ROIs) identified by a universal brain reference system. Finally, these two sets of features are integrated to learn a joint representation by using a multimodal deep Boltzmann machine (DBM). The learned joint representation can be utilized as the feature for training classifiers. Due to the fact that fMRI scanning is expensive and time-consuming, our DBM fusion model has the ability to predict the joint representation of the videos without fMRI scans. The experimental results on a video benchmark demonstrated the effectiveness of our framework and the superiority of integrated features.
Junwei Han 0001, Xintao Hu, Lei Guo 0002, Tianming Liu 0001
IEEE Trans. Affect. Comput.4
2015 Background Prior-Based Salient Object Detection via Deep Reconstruction Residual
abstract
Detection of salient objects from images is gaining increasing research interest in recent years as it can substantially facilitate a wide range of content-based multimedia applications. Based on the assumption that foreground salient regions are distinctive within a certain context, most conventional approaches rely on a number of hand-designed features and their distinctiveness is measured using local or global contrast. Although these approaches have been shown to be effective in dealing with simple images, their limited capability may cause difficulties when dealing with more complicated images. This paper proposes a novel framework for saliency detection by first modeling the background and then separating salient objects from the background. We develop stacked denoising autoencoders with deep learning architectures to model the background where latent patterns are explored and more powerful representations of data are learned in an unsupervised and bottom-up manner. Afterward, we formulate the separation of salient objects from the background as a problem of measuring reconstruction residuals of deep autoencoders. Comprehensive evaluations of three benchmark datasets and comparisons with nine state-of-the-art algorithms demonstrate the superiority of this paper.
Junwei Han 0001, Dingwen Zhang, Xintao Hu, Lei Guo 0002, Jinchang Ren
IEEE Trans. Circuits Syst. Video Technol.4
2015 Effective and Efficient Midlevel Visual Elements-Oriented Land-Use Classification Using VHR Remote Sensing Images
abstract
Land-use classification using remote sensing images covers a wide range of applications. With more detailed spatial and textural information provided in very high resolution (VHR) remote sensing images, a greater range of objects and spatial patterns can be observed than ever before. This offers us a new opportunity for advancing the performance of land-use classification. In this paper, we first introduce an effective midlevel visual elementsoriented land-use classification method based on “partlets,” which are a library of pretrained part detectors used for midlevel visual elements discovery. Taking advantage of midlevel visual elements rather than low-level image features, a partlets-based method represents images by computing their responses to a large number of part detectors. As the number of part detectors grows, a main obstacle to the broader application of this method is its computational cost. To address this problem, we next propose a novel framework to train coarse-to-fine shared intermediate representations, which are termed “sparselets,” from a large number of pretrained part detectors. This is achieved by building a single-hidden-layer autoencoder and a single-hidden-layer neural network with an L0-norm sparsity constraint, respectively. Comprehensive evaluations on a publicly available 21-class VHR landuse data set and comparisons with state-of-the-art approaches demonstrate the effectiveness and superiority of this paper.
Gong Cheng 0003, Junwei Han 0001, Lei Guo 0002, Zhenbao Liu, Shuhui Bu, Jinchang Ren
IEEE Trans. Geosci. Remote. Sens.3
2015 Object Detection in Optical Remote Sensing Images Based on Weakly Supervised Learning and High-Level Feature Learning
abstract
The abundant spatial and contextual information provided by the advanced remote sensing technology has facilitated subsequent automatic interpretation of the optical remote sensing images (RSIs). In this paper, a novel and effective geospatial object detection framework is proposed by combining the weakly supervised learning (WSL) and high-level feature learning. First, deep Boltzmann machine is adopted to infer the spatial and structural information encoded in the low-level and middle-level features to effectively describe objects in optical RSIs. Then, a novel WSL approach is presented to object detection where the training sets require only binary labels indicating whether an image contains the target object or not. Based on the learnt high-level features, it jointly integrates saliency, intraclass compactness, and interclass separability in a Bayesian framework to initialize a set of training examples from weakly labeled images and start iterative learning of the object detector. A novel evaluation criterion is also developed to detect model drift and cease the iterative learning. Comprehensive experiments on three optical RSI data sets have demonstrated the efficacy of the proposed approach in benchmarking with several state-of-the-art supervised-learning-based object detection approaches.
Junwei Han 0001, Dingwen Zhang, Gong Cheng 0003, Lei Guo 0002, Jinchang Ren
IEEE Trans. Geosci. Remote. Sens.4
2015 Supervised Dictionary Learning for Inferring Concurrent Brain Networks
abstract
Task-based fMRI (tfMRI) has been widely used to explore functional brain networks via predefined stimulus paradigm in the fMRI scan. Traditionally, the general linear model (GLM) has been a dominant approach to detect task-evoked networks. However, GLM focuses on task-evoked or event-evoked brain responses and possibly ignores the intrinsic brain functions. In comparison, dictionary learning and sparse coding methods have attracted much attention recently, and these methods have shown the promise of automatically and systematically decomposing fMRI signals into meaningful task-evoked and intrinsic concurrent networks. Nevertheless, two notable limitations of current data-driven dictionary learning method are that the prior knowledge of task paradigm is not sufficiently utilized and that the establishment of correspondences among dictionary atoms in different brains have been challenging. In this paper, we propose a novel supervised dictionary learning and sparse coding method for inferring functional networks from tfMRI data, which takes both of the advantages of model-driven method and data-driven method. The basic idea is to fix the task stimulus curves as predefined model-driven dictionary atoms and only optimize the other portion of data-driven dictionary atoms. Application of this novel methodology on the publicly available human connectome project (HCP) tfMRI datasets has achieved promising results.
Shijie Zhao 0001, Junwei Han 0001, Jinglei Lv, Xi Jiang 0001, Xintao Hu, Yu Zhao 0007, Bao Ge, Lei Guo 0002, Tianming Liu 0001
IEEE Trans. Medical Imaging8
2014 Salient region detection using background contrast
abstract
In this paper, the salient region detection problem is investigated by using background contrast. Since background colors usually appear near image border, and all background colors can be mainly represented by the colors in the image boundary, the boundary-based model is established via computing the different between the intern colors and the boundary colors. As we know, the nearer the patches are close to center, the more they affect other patches. Based on this, a new distribution-based model is proposed. Because of the fact that pixels in the small neighborhood usually have the very similar color components, and computing region-based contrast can reduce the computational complexity, the superpixel algorithm is used in the pretreatment process. Finally, experimental results demonstrate that the proposed method outperforms the state-of-the-art approaches.
Yanbang Zhang, Junwei Han 0001, Lei Guo 0002
ICIP3
2014 Visual attention computation in video of driving environment
abstract
We here study the problem of visual attention computation in video of driving environment via the learning from eye movements. We collect a large-scale database of eye movements from 28 subjects on 30 videos of road scenes, which simulate the driving environment. The analysis on this eye movement database reveals that visual attention in driving environment is directed by high-level cognitive factors such as objects. We then present a new high-level representation called Traffic Object Bank (TOB), which is comprised of many individual road object detectors trained comprehensively in semantic space as well as viewpoint space. TOB provides semantically rich object-level features. Finally, we develop a computational model to predict where drivers look via the mapping from TOB-based representation and to gaze data. Experimental results on our traffic scene video benchmark indicate high accordance with human eye movement and show great promise for further applications.
Junwei Han 0001, Liye Sun, Dingwen Zhang, Xintao Hu, Gong Cheng 0003, Lei Guo 0002
ICME6
2014 Saliency detection based on feature learning using Deep Boltzmann Machines
abstract
Saliency detection has been a very active research area in recent years. Most traditional methods suffer from the problem that existing visual features are not discriminative or not robust enough to predict salient locations. As a result, the experimental results of these previous methods are still far from satisfactory. In this paper, we propose to utilize a two-layer Deep Boltzmann Machine (DBM) to learn enhanced features from existing contrast-based low-level features, which are more discriminative and reliable. A saliency computation model is then trained to build a mapping from those enhanced features to eye fixation data. The proposed work is amongst the earliest efforts of examining the feasibility of applying deep learning algorithms to saliency detection. Comprehensive evaluations on two publically available benchmark datasets and comparisons with a number of state-of-the-art approaches demonstrate the effectiveness of the proposed work.
Shifeng Wen, Junwei Han 0001, Dingwen Zhang, Lei Guo 0002
ICME4
2014 Scalable multi-class geospatial object detection in high-spatial-resolution remote sensing images
abstract
In this paper we present a conceptually simple but surprisingly effective multi-class geospatial object detection method based on Collection of Part Detectors (COPD), which can be easily scaled to a larger number of object classes. The presented COPD is composed of a set of representative and discriminative part detectors, where each part detector is a linear support vector machine (SVM) classifier trained using a weakly supervised learning method that only requires image labels indicating the presence of objects for the training data. Here, each part detector corresponds to a particular viewpoint of an object class, so the collection of them provides a feasible solution for rotation-invariant and simultaneous detection of multi-class geospatial objects. Comprehensive evaluations on high-spatial-resolution remote sensing images and comparisons with a number of state-of-the-art approaches demonstrate the effectiveness and superiority of the presented method.
Gong Cheng 0003, Junwei Han 0001, Peicheng Zhou, Lei Guo 0002
IGARSS4
2014 Decoding Auditory Saliency from FMRI Brain Imaging
abstract
Given the growing number of available audio streams through a variety of sources and distribution channels, effective and advanced computational audio analysis has received increasing interest in the multimedia field. However, the effectiveness of current audio analysis strategies might be hampered due to the lack of effective representation of high-level semantics perceived by the human and the lack of effective approaches to bridging the gaps between most low-level acoustic features and high-level semantic features. This semantic gap has become the 'bottleneck' problem in audio analysis. In this paper, we propose a computational framework to decode biologically-plausible auditory saliency using high-level features derived from functional magnetic resonance imaging (fMRI) which monitors the human brain's response under the natural stimulus of audio listening. Specifically, we identify meaningful intrinsic brain networks which are involved in audio listening via effective online dictionary learning and sparse representation of whole-brain fMRI signals, reconstruct auditory saliency features using those identified brain network components, and perform group-wise analysis to identify consistent 'brain decoders' of the saliency features across different excerpts and participants. Experimental results demonstrate that the auditory saliency features are effectively decoded via our methods, which potentially provide opportunities for various applications in the multimedia field.
Shijie Zhao 0001, Xi Jiang 0001, Junwei Han 0001, Xintao Hu, Dajiang Zhu, Jinglei Lv, Lei Guo 0002, Tianming Liu 0001
ACM Multimedia8
2014 Video abstraction based on fMRI-driven visual attention model
Junwei Han 0001, Kaiming Li, Ling Shao 0001, Xintao Hu, Lei Guo 0002, Jungong Han, Tianming Liu 0001
Inf. Sci.6
2014 Characterization of U-shape streamline fibers: Methods and applications
Hanbo Chen, Lei Guo 0002, Kaiming Li, Longchuan Li, Shu Zhang 0001, Dinggang Shen, Xiaoping Hu 0001, Tianming Liu 0001
Medical Image Anal.3
2014 Interactive object-based image retrieval and annotation on iPad
Junwei Han 0001, Lei Guo 0002, Tianming Liu 0001
Multim. Tools Appl.4
2014 Merging Neuroimaging and Multimedia: Methods, Opportunities, and Challenges
abstract
Neuroimaging and brain mapping can provide meaningful guidance to multimedia analyses. and advanced computational multimedia analysis can be used to better understand the functional mechanisms of the human brain. Essentially, brain imaging and brain mapping techniques can serve as a bridge that links the digital representation of multimedia and the perception and comprehension of its content. This paper summarizes methods that integrate brain imaging with multimedia analysis and discusses the opportunities and challenges in this interdisciplinary field. In general, quantitative modeling of brain responses during multimedia comprehension has advanced content-based multimedia studies such as image and video classification and tagging. Multimedia analysis has promoted functional brain mapping by using naturalistic multimedia as stimuli during neuroimaging. Challenges and opportunities in merging neuroimaging and multimedia include the quantification of the brain's responses, the quantification of multimedia, and the mapping between brain responses and computational multimedia features.
Tianming Liu 0001, Xintao Hu, Junwei Han 0001, Lei Guo 0002
IEEE Trans. Hum. Mach. Syst.6
2013 Assessing Graph Properties and Dynamics of the Functional Brain Networks in Alzheimer's Disease
abstract
The human brain is the most complex system in nature. It is intrinsically organized into networked system. Theoretical graphic analysis of human brain networks not only sheds new light into the understanding how the human brain works, but also provides information for exploring into neurological and psychiatric disorders. This paper presents our work of MR imaging data and resting-state functional magnetic resonance imaging (R-fMRI) data, characterizing the graph properties related to clustering coefficient, average degree percentage, characteristic path, global efficiency and small-world ness and network dynamics property on the constructed functional brain networks, then, inferring the discrepancies of those properties between patients with Alzheimer's disease and normal controls. Our experimental results demonstrate that functional brain network of normal controls has higher clustering coefficient, average degree percentage and global efficiency and lower characteristic path. Moreover, it also has stronger small-world ness and propensity for synchronization, compared those of functional brain networks of patients with Alzheimer's disease.
Lei Guo 0002
ICIG2
2013 Predictive Models of Resting State Networks for Assessment of Altered Functional Connectivity in MCI
Xi Jiang 0001, Dajiang Zhu, Kaiming Li, Dinggang Shen, Lei Guo 0002, Tianming Liu 0001
MICCAI (2)6
2013 Anatomy-Guided Discovery of Large-Scale Consistent Connectivity-Based Cortical Landmarks
Xi Jiang 0001, Dajiang Zhu, Kaiming Li, Jinglei Lv, Lei Guo 0002, Tianming Liu 0001
MICCAI (3)6
2013 Modeling Dynamic Functional Information Flows on Large-Scale Brain Networks
Peili Lv, Lei Guo 0002, Xintao Hu, Xiang Li 0001, Changfeng Jin, Junwei Han 0001, Lingjiang Li, Tianming Liu 0001
MICCAI (2)2
2013 Sparse Representation of Group-Wise FMRI Signals
Jinglei Lv, Xiang Li 0001, Dajiang Zhu, Xi Jiang 0001, Xin Zhang 0151, Xintao Hu, Lei Guo 0002, Tianming Liu 0001
MICCAI (3)8
2013 Group-Wise FMRI Activation Detection on Corresponding Cortical Landmarks
Jinglei Lv, Dajiang Zhu, Xintao Hu, Xin Zhang 0151, Junwei Han 0001, Lei Guo 0002, Tianming Liu 0001
MICCAI (2)7
2013 Sparse Representation of Higher-Order Functional Interaction Patterns in Task-Based FMRI Data
Shu Zhang 0001, Xiang Li 0001, Jinglei Lv, Xi Jiang 0001, Dajiang Zhu, Hanbo Chen, Lei Guo 0002, Tianming Liu 0001
MICCAI (3)8
2013 Characterization of task-free and task-performance brain states via functional connectome patterns
Xin Zhang 0151, Lei Guo 0002, Xiang Li 0001, Dajiang Zhu, Kaiming Li, Hanbo Chen, Jinglei Lv, Changfeng Jin, Lingjiang Li, Tianming Liu 0001
Medical Image Anal.2
2013 Predicting cortical ROIs via joint modeling of anatomical and connectional profiles
Dajiang Zhu, Xi Jiang 0001, Bao Ge, Xintao Hu, Junwei Han 0001, Lei Guo 0002, Tianming Liu 0001
Medical Image Anal.7
2013 Optimal contrast based saliency detection
Xiaoliang Qian, Junwei Han 0001, Gong Cheng 0003, Lei Guo 0002
Pattern Recognit. Lett.4
2013 An Object-Oriented Visual Saliency Detection Framework Based on Sparse Coding Representations
abstract
Saliency detection aims at quantitatively predicting attended locations in an image. It may mimic the selection mechanism of the human vision system, which processes a small subset of a massive amount of visual input while the redundant information is ignored. Motivated by the biological evidence that the receptive fields of simple cells in V1 of the vision system are similar to sparse codes learned from natural images, this paper proposes a novel framework for saliency detection by using image sparse coding representations as features. Unlike many previous approaches dedicated to examining the local or global contrast of each individual location, this paper develops a probabilistic computational algorithm by integrating objectness likelihood with appearance rarity. In the proposed framework, image sparse coding representations are yielded through learning on a large amount of eye-fixation patches from an eye-tracking dataset. The objectness likelihood is measured by three generic cues called compactness, continuity, and center bias. The appearance rarity is inferred by using a Gaussian mixture model. The proposed paper can serve as a basis for many techniques such as image/video segmentation, retrieval, retargeting, and compression. Extensive evaluations on benchmark databases and comparisons with a number of up-to-date algorithms demonstrate its effectiveness.
Junwei Han 0001, Xiaoliang Qian, Lei Guo 0002, Tianming Liu 0001
IEEE Trans. Circuits Syst. Video Technol.5
2013 Representing and Retrieving Video Shots in Human-Centric Brain Imaging Space
abstract
Meaningful representation and effective retrieval of video shots in a large-scale database has been a profound challenge for the image/video processing and computer vision communities. A great deal of effort has been devoted to the extraction of low-level visual features, such as color, shape, texture, and motion for characterizing and retrieving video shots. However, the accuracy of these feature descriptors is still far from satisfaction due to the well-known semantic gap. In order to alleviate the problem, this paper investigates a novel methodology of representing and retrieving video shots using human-centric high-level features derived in brain imaging space (BIS) where brain responses to natural stimulus of video watching can be explored and interpreted. At first, our recently developed dense individualized and common connectivity-based cortical landmarks (DICCCOL) system is employed to locate large-scale functional brain networks and their regions of interests (ROIs) that are involved in the comprehension of video stimulus. Then, functional connectivities between various functional ROI pairs are utilized as BIS features to characterize the brain's comprehension of video semantics. Then an effective feature selection procedure is applied to learn the most relevant features while removing redundancy, which results in the formation of the final BIS features. Afterwards, a mapping from low-level visual features to high-level semantic features in the BIS is built via the Gaussian process regression (GPR) algorithm, and a manifold structure is then inferred, in which video key frames are represented by the mapped feature vectors in the BIS. Finally, the manifold-ranking algorithm concerning the relationship among all data is applied to measure the similarity between key frames of video shots. Experimental results on the TRECVID 2005 dataset demonstrate the superiority of the proposed work in comparison with traditional methods.
Junwei Han 0001, Xintao Hu, Dajiang Zhu, Kaiming Li, Xi Jiang 0001, Guangbin Cui, Lei Guo 0002, Tianming Liu 0001
IEEE Trans. Image Process.8
2013 Inferring Group-Wise Consistent Multimodal Brain Networks via Multi-View Spectral Clustering
abstract
Quantitative modeling and analysis of structural and functional brain networks based on diffusion tensor imaging (DTI) and functional magnetic resonance imaging (fMRI) data have received extensive interest recently. However, the regularity of these structural and functional brain networks across multiple neuroimaging modalities and also across different individuals is largely unknown. This paper presents a novel approach to inferring group-wise consistent brain subnetworks from multimodal DTI/resting-state fMRI datasets via multi-view spectral clustering of cortical networks, which were constructed upon our recently developed and validated large-scale cortical landmarks-DICCCOL (dense individualized and common connectivity-based cortical landmarks). We applied the algorithms on DTI data of 100 healthy young females and 50 healthy young males, obtained consistent multimodal brain networks within and across multiple groups, and further examined the functional roles of these networks. Our experimental results demonstrated that the derived brain networks have substantially improved inter-modality and inter-subject consistency.
Hanbo Chen, Kaiming Li, Dajiang Zhu, Xi Jiang 0001, Yixuan Yuan, Peili Lv, Lei Guo 0002, Dinggang Shen, Tianming Liu 0001
IEEE Trans. Medical Imaging8
2012 Inferring Group-Wise Consistent Multimodal Brain Networks via Multi-view Spectral Clustering
Hanbo Chen, Kaiming Li, Dajiang Zhu, Changfeng Jin, Lei Guo 0002, Lingjiang Li, Tianming Liu 0001
MICCAI (3)6
2012 Group-Wise Consistent Fiber Clustering Based on Multimodal Connectional and Functional Profiles
Bao Ge, Lei Guo 0002, Dajiang Zhu, Kaiming Li, Xintao Hu, Junwei Han 0001, Tianming Liu 0001
MICCAI (3)2
2012 Characterization of Task-Free/Task-Performance Brain States
Xin Zhang 0151, Lei Guo 0002, Xiang Li 0001, Dajiang Zhu, Kaiming Li, Zhenqiang Sun, Changfeng Jin, Xintao Hu, Junwei Han 0001, Lingjiang Li, Tianming Liu 0001
MICCAI (2)2
2012 Music/speech classification using high-level features derived from fmri brain imaging
abstract
With the availability of large amount of audio tracks through a variety of sources and distribution channels, automatic music/speech classification becomes an indispensable tool in social audio websites and online audio communities. However, the accuracy of current acoustic-based low-level feature classification methods is still rather far from satisfaction. The discrepancy between the limited descriptive power of low-level features and the richness of high-level semantics perceived by the human brain has become the 'bottleneck' problem in audio signal analysis. In this paper, functional magnetic resonance imaging (fMRI) which monitors the human brain's response under the natural stimulus of music/speech listening is used as high-level features in the brain imaging space (BIS). We developed a computational framework to model the relationships between BIS features and low-level features in the training dataset with fMRI scans, predict BIS features of testing dataset without fMRI scans, and use the predicted BIS features for music/speech classification in the application stage. Experimental results demonstrated the significantly improved performance of music/speech classification via predicted BIS features than that via the original low-level features.
Xi Jiang 0001, Xintao Hu, Lie Lu, Junwei Han 0001, Lei Guo 0002, Tianming Liu 0001
ACM Multimedia6
2012 Bridging the Semantic Gap via Functional Brain Imaging
abstract
The multimedia content analysis community has made significant efforts to bridge the gaps between low-level features and high-level semantics perceived by humans. Recent advances in brain imaging and neuroscience in exploring the human brain's responses during multimedia comprehension demonstrated the possibility of leveraging cognitive neuroscience knowledge to bridge the semantic gaps. This paper presents our initial effort in this direction by using functional magnetic resonance imaging (fMRI). Specifically, task-based fMRI (T-fMRI) was performed to accurately localize the brain regions involved in video comprehension. Then, natural stimulus fMRI (N-fMRI) data were acquired when subjects watched the multimedia clips selected from the TRECVID datasets. The responses in the localized brain regions were measured and used to extract high-level features as the representation of the brain's comprehension of semantics in the videos. A novel computational framework was developed to learn the most relevant low-level feature sets that best correlate the fMRI-derived semantic features based on the training videos with fMRI scans, and then the learned model was applied to larger scale TRECVID video datasets without fMRI scans for category classification. Our experimental results demonstrate: 1) there are meaningful couplings between brain's fMRI-derived responses and video stimuli, suggesting the validity of linking semantics and low-level features via fMRI and 2) the computationally learned low-level features can significantly (p <; 0.01) improve video classification in comparison with original low-level features and extracted low-level features resulted from well-known feature projection algorithms.
Xintao Hu, Kaiming Li, Junwei Han 0001, Xian-Sheng Hua 0001, Lei Guo 0002, Tianming Liu 0001
IEEE Trans. Multim.5
2011 Retrieving video shots in semantic brain imaging space using manifold-ranking
abstract
In recent two decades, a large amount of effort has been devoted to content-based video retrieval (CBVR), which aims to manage large-scale video databases in an effective way based on visual features such as color, shape, texture, and motion. However, the performance of CBVR systems is still far from satisfaction due to the well-known semantic gap. In order to alleviate the problem, this paper proposes a novel retrieval methodology using semantic features derived from brain imaging space (BIS) that reflects brain responses and interactions under natural stimulus of video watching. A mapping from visual features to semantic features in BIS is built through Gaussian process regression. A manifold structure is then inferred where video key frames are represented by mapped feature vectors in BIS. Finally, the manifold-ranking algorithm concerning the relationship among all data is applied to measure the similarity between key frames. Preliminary experimental results on the TRECVID 2005 dataset demonstrate the superiority of the proposed work in comparison with traditional methods.
Junwei Han 0001, Xintao Hu, Kaiming Li, Fan Deng 0001, Lei Guo 0002, Tianming Liu 0001
ICIP7
2011 Assessing Regularity and Variability of Cortical Folding Patterns of Working Memory ROIs
Hanbo Chen, Kaiming Li, Xintao Hu, Lei Guo 0002, Tianming Liu 0001
MICCAI (2)5
2011 Resting State fMRI-Guided Fiber Clustering
Bao Ge, Lei Guo 0002, Jinglei Lv, Xintao Hu, Junwei Han 0001, Tianming Liu 0001
MICCAI (2)2
2011 Fiber-Centered Granger Causality Analysis
Xiang Li 0001, Kaiming Li, Lei Guo 0002, Chulwoo Lim, Tianming Liu 0001
MICCAI (2)3
2011 Robust Deformable-Surface-Based Skull-Stripping for Large-Scale Studies
Jingxin Nie, Pew-Thian Yap, Feng Shi 0001, Lei Guo 0002, Dinggang Shen
MICCAI (3)5
2011 Predicting Functional Brain ROIs via Fiber Shape Models
Lei Guo 0002, Kaiming Li, Dajiang Zhu, Guangbin Cui, Tianming Liu 0001
MICCAI (2)2
2011 A biologically inspired computational model for image saliency detection
abstract
Image saliency detection provides a powerful tool for predicting where human tends to look at in an image, which has been a long attempt for the computer vision community. In this paper, we propose a biologically-inspired model for computing image saliency. At first, a set of basis functions that accords with visual responses to natural stimuli is learned by using eye-fixation patches from an eye-tracking dataset. Three features are then derived based on the learned basis functions including continuity, clutter contrast, and local contrast. Finally, these three features are combined into the saliency map. The proposed approach is easy to implement and can be used in many image and video content analysis applications. Experiments on a large-scale benchmark dataset and comparisons with a number of the state-of-the-art approaches demonstrate its superiority.
Junwei Han 0001, Xintao Hu, Lei Guo 0002, Tianming Liu 0001
ACM Multimedia5
2010 A Dynamic Skull Model for Simulation of Cerebral Cortex Folding
Hanbo Chen, Lei Guo 0002, Jingxin Nie, Xintao Hu, Tianming Liu 0001
MICCAI (2)2
2010 Fiber-Centered Analysis of Brain Connectivities Using DTI and Resting State FMRI Data
Jinglei Lv, Lei Guo 0002, Xintao Hu, Kaiming Li, Degang Zhang, Tianming Liu 0001
MICCAI (2)2
2010 Bridging low-level features and high-level semantics via fMRI brain imaging for video classification
abstract
The multimedia content analysis community has made significant effort to bridge the gap between low-level features and high-level semantics perceived by human cognitive systems such as real-world objects and concepts. In the two fields of multimedia analysis and brain imaging, both topics of low-level features and high level semantics are extensively studied. For instance, in the multimedia analysis field, many algorithms are available for multimedia feature extraction, and benchmark datasets are available such as the TRECVID. In the brain imaging field, brain regions that are responsible for vision, auditory perception, language, and working memory are well studied via functional magnetic resonance imaging (fMRI). This paper presents our initial effort in marrying these two fields in order to bridge the gaps between low-level features and high-level semantics via fMRI brain imaging. Our experimental paradigm is that we performed fMRI brain imaging when university student subjects watched the video clips selected from the TRECVID datasets. At current stage, we focus on the three concepts of sports, weather, and commercial-/advertisement specified in the TRECVID 2005. Meanwhile, the brain regions in vision, auditory, language, and working memory networks are quantitatively localized and mapped via task-based paradigm fMRI, and the fMRI responses in these regions are used to extract features as the representation of the brain's comprehension of semantics. Our computational framework aims to learn the most relevant low-level feature sets that best correlate the fMRI-derived semantics based on the training videos with fMRI scans, and then the learned models are applied to larger scale test datasets without fMRI scans for category classifications. Our result shows that: 1) there are meaningful couplings between brain's fMRI responses and video stimuli, suggesting the validity of linking semantics and low-level features via fMRI; 2) The computationally learned low-level feature sets from fMRI-derived semantic features can significantly improve the classification of video categories in comparison with that based on original low-level features.
Xintao Hu, Fan Deng 0001, Kaiming Li, Hanbo Chen, Xi Jiang 0001, Jinglei Lv, Dajiang Zhu, Carlos Faraco, Degang Zhang, Arsham Mesbah, Junwei Han 0001, Xian-Sheng Hua 0001, L. Stephen Miller, Lei Guo 0002, Tianming Liu 0001
ACM Multimedia16
2010 Individualized ROI Optimization via Maximization of Group-wise Consistency of Structural and Functional Profiles
abstract
Functional segregation and integration are fundamental characteristics of the human brain. Studying the connectivity among segregated regions and the dynamics of integrated brain networks has drawn increasing interest. A very controversial, yet fundamental issue in these studies is how to determine the best functional brain regions or ROIs (regions of interests) for individuals. Essentially, the computed connectivity patterns and dynamics of brain networks are very sensitive to the locations, sizes, and shapes of the ROIs. This paper presents a novel methodology to optimize the locations of an individual's ROIs in the working memory system. Our strategy is to formulate the individual ROI optimization as a group variance minimization problem, in which group-wise functional and structural connectivity patterns, and anatomic profiles are defined as optimization constraints. The optimization problem is solved via the simulated annealing approach. Our experimental results show that the optimized ROIs have significantly improved consistency in structural and functional profiles across subjects, and have more reasonable localizations and more consistent morphological and anatomic profiles.
Kaiming Li, Lei Guo 0002, Carlos Faraco, Dajiang Zhu, Fan Deng 0001, Xi Jiang 0001, Degang Zhang, Hanbo Chen, Xintao Hu, L. Stephen Miller, Tianming Liu 0001
NIPS2
2010 An automated pipeline for cortical sulcal fundi extraction
Gang Li 0001, Lei Guo 0002, Jingxin Nie, Tianming Liu 0001
Medical Image Anal.2
2009 Comparative Analysis of Fingerprint Orientation Field Algorithms
abstract
Fingerprint images are textural images consisting of ridges and valleys. The orientation of textures can be determined by orientation field computation. Fingerprint orientation field is the critical basis for fingerprint image segmentation, filtering enhancement and matching processes, and the fingerprint orientation field algorithm plays a very important role in the applied Automated Fingerprint Identification Systems (AFIS). The available fingerprint orientation field algorithms include mainly the mask algorithm and the gradient algorithm. They are used in spatial and frequency domains, respectively, and both can offer satisfactory fingerprint orientation matrixes. Based on an experimental comparison of their effectiveness in fingerprint preprocessing, this paper analyzes in detail the performance of these two algorithms.
Dahai Chen, Feng Fan, Lei Guo 0002, Weihua Meng
ICIG5
2009 Grouping of Brain MR Images via Affinity Propagation
abstract
The human brain anatomy is extremely variable across individuals in terms of its size, shape, and structure patterning. In this paper, a novel method is proposed for grouping brain MR images into different patterns. This method adopts the affinity propagation methodology to partition a population of brain images into different clusters. In the affinity propagation method, the tissue-segmented and anatomically-parcellated images are used to define the similarity between brain images, in contrast to intensity-based similarity measurement used in previous methods. After clustering, in each cluster (called a sub-group) a representative exemplar image is identified as the single subject atlas for the sub-group. Meanwhile, all the subject images belonging to the same sub-group are identified. This method has been applied to the publicly available OASIS neuroimaging dataset that includes 414 subject brain MRI images. Experiments show that the method is able to group brain MR images into different patterns effectively.
Gang Li 0001, Lei Guo 0002, Tianming Liu 0001
ISCAS2
2009 Gyral Folding Pattern Analysis via Surface Profiling
Kaiming Li, Lei Guo 0002, Gang Li 0001, Jingxin Nie, Carlos Faraco, L. Stephen Miller, Tianming Liu 0001
MICCAI (1)2
2009 A Computational Model of Cerebral Cortex Folding
Jingxin Nie, Gang Li 0001, Lei Guo 0002, Tianming Liu 0001
MICCAI (1)3
2009 Parametric Representation of Cortical Surface Folding Based on Polynomials
Lei Guo 0002, Gang Li 0001, Jingxin Nie, Tianming Liu 0001
MICCAI (1)2
2008 Ontology clarification by using semantic disambiguation
abstract
Semantic Web technology highly depends on the quality of ontology. In order to enhance quality of ontology, a vast amount of research has focused on concept modeling task, but there is one major problem with lexical representation of ontology. Current lexical representation is term which may have different meanings, so as to result in frustrating misunderstanding and ambiguity during the application of ontology. To solve this problem, sense is used to replace term as the lexical representation of concepts and properties for its unique meaning. Ontology clarification is the process of disambiguating terms in ontology by using its surrounding ontology elements and its nearby terms in annotated documents using this ontology. The right sense is assigned to a target term by maximizing the relatedness between the target and its neighbors for semantic relatedness between them. Experiments show our ontology clarification method is valid. Comparing with the best word sense disambiguation method, the concept precision is almost 2 times than the precision of noun, and the property precision is almost 3 times than the precision of verb. Another experiment proves that our method is also effective in a semi-automatic process.
Lei Guo 0002, Xiaodong Wang 0004
CSCWD1
2008 Automated ontology selection based on description logic
abstract
With the widespread use of Ontology and extension of Ontology repositories, there come deficiencies on Ontology Selection. The deficiencies are mainly low automated level and absence of Ontological knowledge in the selection criteria. To fill the gaps, we propose a Description Logic based Automated Ontology Selection Framework (DL-AOSF), which consists of automated components and is designed to be suitable for various application scenarios. The core algorithm of DL-AOSF adopts dual criteria, viz. topic coverage and knowledge richness, to comprehensively assess the satisfaction of candidates with concrete application information need. Distinguishing from other Ontology selections, DL- AOSF implements the knowledge richness criteria on semantic level, by using knowledge-driven Ontology modularization technique and DL reasoning. The preliminary experimental results indicate DL-AOSF is valid and promising.
Xiaodong Wang 0004, Lei Guo 0002
CSCWD2
2008 A Novel Method for Cortical Sulcal Fundi Extraction
Gang Li 0001, Tianming Liu 0001, Jingxin Nie, Lei Guo 0002, Stephen T. C. Wong
MICCAI (1)4
2008 ZFIQ: a software package for zebrafish biology
abstract
Abstract Summary: Rapid development, transparency and small size are the outstanding features of zebrafish that make it as an increasingly important vertebrate system for developmental biology, functional genomics, disease modeling and drug discovery. Zebrafish has been regarded as ideal animal specie for studying the relationship between genotype and phenotype, for pathway analysis and systems biology. However, the tremendous amount of data generated from large numbers of embryos has led to the bottleneck of data analysis and modeling. The zebrafish image quantitator (ZFIQ) software provides streamlined data processing and analysis capability for developmental biology and disease modeling using zebrafish model. Availability: ZFIQ is available for download at http://www.cbi-platform.net Contact: [email protected] Supplementary information: Additional documentation for this software package is referred to http://www.cbi-platform.net/document.htm. Application examples of this software are referred to http://www.cbi-platform.net/download.htm
Tianming Liu 0001, Jingxin Nie, Gang Li 0001, Lei Guo 0002, Stephen T. C. Wong
Bioinform.4
2006 Importance of Entities in Knowledge
abstract
There is a growing need for managing importance of entities in knowledge system in order to realize the full potential of knowledge. How to calculate the importance of entities automatically is the primary issue. We argue that importance of entities in knowledge is dynamic, the importance is changed along with the using of ontology; and different groups of user have different criteria of importance. In this paper, a novel weight assignment method which takes usage and structure properties of ontology into account is proposed. When considering the usage information of ontology, we analyze paths that are used to respond to queries; and use the frequency of entities included in the paths to produce the optimal weight assignment for the assumption of high importance of entities which included in paths that respond to queries. After get the initial weight of entities, a pervasion algorithm which considers the structure of ontology is used to compute the final weight of entities. Weight of a node is high if the node has many incoming links and the incoming links and nodes which these links are from have high scores. Experiment show effectiveness of this weight assignment method.
Lei Guo 0002, Xiaodong Wang 0004, Ning Yang 0003, Weili Yang
Web Intelligence2
2003 Automatic attention object extraction from images
abstract
With the development of content-based multimedia systems, there is a need for automatic extraction of objects from natural images. However, the objects extracted by most existing approaches are often inconsistent with human perception since these approaches totally neglect the viewer's attentions. To address this issue, a method is presented in this paper to automatically extract the viewer's attended objects from an image. Without fully understanding of the semantic content of an image, this method takes advantage of computational attention mechanisms and the seeded region growing technique. It may further facilitate the content-based image/video coding, indexing, and retrieval. Preliminary experimental evaluations on 200 real images demonstrate the effectiveness of this method.
Junwei Han 0001, Mingjing Li, HongJiang Zhang, Lei Guo 0002
ICIP (2)4
2003 A memorization learning model for image retrieval
abstract
Current image retrieval systems still have major difficulties in bridging the gap between high-level concept and low-level image representation. To overcome these difficulties, a memorization learning model is proposed in this paper. It memorizes the semantic knowledge of images in a database by simply accumulating the user-provided relevance feedback information. From the memorized knowledge, it then learns some hidden semantic information of images. Image retrieval is finally based on a seamless combination of low-level features, memorized semantic information, and estimated hidden semantic information. The model is easy to implement and can be efficiently applied to an image retrieval system. Preliminary experimental results on 10,000 images demonstrate the effectiveness of the proposed model.
Junwei Han 0001, Mingjing Li, HongJiang Zhang, Lei Guo 0002
ICIP (3)4
2003 A shape-based image retrieval method using salient edges
Junwei Han 0001, Lei Guo 0002
Signal Process. Image Commun.2
2002 A new image retrieval system supporting query by semantics and example
abstract
We propose a new image retrieval system that provides users with both semantics based query and visual features based query. Our system has several advantages. First, it integrates visual features and semantics seamlessly. Second, it uses some effective techniques, such as image classification and relevance feedback, to bridge the gap between visual features and semantics. Third, it proposes several ways to obtain the semantic information of the image, which reduces manual labor and reduces the "subjectivity" of semantics by human. Fourth, it can update the semantics of the image by human intervention, which makes the image retrieval more flexible. We have implemented an image retrieval system based on our proposed approach. Experiments on an image database containing 22,000 items show that our scheme can achieve high efficiency.
Junwei Han 0001, Lei Guo 0002
ICIP (3)2
2002 Adaptive self-excitation groups in visual curve integration
Lei Guo 0002, Tianming Liu 0001, Junwei Han 0001
Neurocomputing1
2000 Parallel double-conversion spiking neural network for world-centered recognition
Lei Guo 0002
Neurocomputing1
1999 Random time-division operation for salience of visual contours
abstract
The mutual excitation among the local stimuli satisfying curve distribution (position and orientation continuity) called self-excitation of curves here is an effective method for the discovery and enhancement of visual curves. This article presents a new method using dynamic time-division curve searches and self-excitation. The searches realized by random walks of active particle are guided by inputs, limited by the rules of curve distribution performed repetitively, and temporally divided for different curve candidates. The time-division operations play an extremely important role in both the structure division used previously and the global memory of various search routes.
Lei Guo 0002, Tianming Liu 0001
IJCNN1