Xi Jiang 0001

dblp:31/3804-1 · DBLP profile ↗
← Back
53ranked-venue papers
7as first author
22since 2021 · last 2025
0000-0003-3711-0847ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 31 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 5 first-author · 4 since 2021Artificial intelligence and machine learning · 14 · 2 first-author · 12 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 Adaptive critical subgraph mining for cognitive impairment conversion prediction with T1-MRI-based brain network
Yilin Leng, Wenju Cui, Xi Jiang 0001, Yunsong Peng, Jian Zheng 0001
Expert Syst. Appl.4
2025 Contrastive machine learning reveals species -shared and -specific brain functional architecture
Guannan Cao, Songyao Zhang, Weihan Zhang, Yusong Sun, Jingchao Zhou, Tianyang Zhong, Yixuan Yuan, Tao Liu 0044, Tianming Liu 0001, Lei Guo 0002, Yongchun Yu, Xi Jiang 0001, Gang Li 0001, Junwei Han 0001
Medical Image Anal.13
2025 A Unified and Biologically Plausible Relational Graph Representation of Vision Transformers
abstract
Vision transformer (ViT) and its variants have achieved remarkable success in various tasks. The key characteristic of these ViT models is to adopt different aggregation strategies of spatial patch information within the artificial neural networks (ANNs). However, there is still a key lack of unified representation of different ViT architectures for systematic understanding and assessment of model representation performance. Moreover, how those well-performing ViT ANNs are similar to real biological neural networks (BNNs) is largely unexplored. To answer these fundamental questions, we, for the first time, propose a unified and biologically plausible relational graph representation of ViT models. Specifically, the proposed relational graph representation consists of two key subgraphs: an aggregation graph and an affine graph. The former considers ViT tokens as nodes and describes their spatial interaction, while the latter regards network channels as nodes and reflects the information communication between channels. Using this unified relational graph representation, we found that: 1) model performance was closely related to graph measures; 2) the proposed relational graph representation of ViT has high similarity with real BNNs; and 3) there was a further improvement in model performance when training with a superior model to constrain the aggregation graph.
Yuzhong Chen 0002, Zhenxiang Xiao, Lin Zhao 0004, Lu Zhang 0050, Zihao Wu 0001, Dajiang Zhu, Dezhong Yao 0001, Xintao Hu, Tianming Liu 0001, Xi Jiang 0001
IEEE Trans. Neural Networks Learn. Syst.12
2025 Mask-Guided Vision Transformer for Few-Shot Learning
abstract
Learning with little data is challenging but often inevitable in various application scenarios where the labeled data are limited and costly. Recently, few-shot learning (FSL) gained increasing attention because of its generalizability of prior knowledge to new tasks that contain only a few samples. However, for data-intensive models such as vision transformer (ViT), current fine-tuning-based FSL approaches are inefficient in knowledge generalization and, thus, degenerate the downstream task performances. In this article, we propose a novel mask-guided ViT (MG-ViT) to achieve an effective and efficient FSL on the ViT model. The key idea is to apply a mask on image patches to screen out the task-irrelevant ones and to guide the ViT focusing on task-relevant and discriminative patches during FSL. Particularly, MG-ViT only introduces an additional mask operation and a residual connection, enabling the inheritance of parameters from pretrained ViT without any other cost. To optimally select representative few-shot samples, we also include an active learning-based sample selection method to further improve the generalizability of MG-ViT-based FSL. We evaluate the proposed MG-ViT on classification, object detection, and segmentation tasks using gradient-weighted class activation mapping (Grad-CAM) to generate masks. The experimental results show that the MG-ViT model significantly improves the performance and efficiency compared with general fine-tuning-based ViT and ResNet models, providing novel insights and a concrete approach toward generalizing data-intensive and large-scale deep learning models for FSL.
Yuzhong Chen 0002, Zhenxiang Xiao, Yi Pan 0001, Lin Zhao 0004, Haixing Dai, Zihao Wu 0001, Changhe Li, Changying Li, Dajiang Zhu, Tianming Liu 0001, Xi Jiang 0001
IEEE Trans. Neural Networks Learn. Syst.12
2025 A Novel Dynamic Neural Network for Heterogeneity-Aware Structural Brain Network Exploration and Alzheimer's Disease Diagnosis
abstract
Heterogeneity is a fundamental characteristic of brain diseases, distinguished by variability not only in brain atrophy but also in the complexity of neural connectivity and brain networks. However, existing data-driven methods fail to provide a comprehensive analysis of brain heterogeneity. Recently, dynamic neural networks (DNNs) have shown significant advantages in capturing sample-wise heterogeneity. Therefore, in this article, we first propose a novel dynamic heterogeneity-aware network (DHANet) to identify critical heterogeneous brain regions, explore heterogeneous connectivity between them, and construct a heterogeneous-aware structural brain network (HGA-SBN) using structural magnetic resonance imaging (sMRI). Specifically, we develop a 3-D dynamic convmixer to extract abundant heterogeneous features from sMRI first. Subsequently, the critical brain atrophy regions are identified by dynamic prototype learning with embedding the hierarchical brain semantic structure. Finally, we employ a joint dynamic edge-correlation (JDE) modeling approach to construct the heterogeneous connectivity between these regions and analyze the HGA-SBN. To evaluate the effectiveness of the DHANet, we conduct elaborate experiments on three public datasets and the method achieves state-of-the-art (SOTA) performance on two classification tasks.
Wenju Cui, Yilin Leng, Yunsong Peng, Lei Li 0058, Xi Jiang 0001, Jian Zheng 0001
IEEE Trans. Neural Networks Learn. Syst.6
2025 ChatABL: Abductive Learning via Natural Language Interaction With ChatGPT
abstract
Large language models (LLMs) such as ChatGPT have recently demonstrated significant potential in mathematical abilities, providing a valuable reasoning paradigm consistent with human natural language. However, LLMs currently have difficulty in bridging perception, language understanding, and reasoning (PLR) capabilities due to incompatibility of the underlying information flow among them, making their reasoning ability not fully elicited and challenging to accomplish complicated reasoning tasks autonomously. To resolve the above problem, a novel method called ChatABL is proposed by integrating LLMs into an abductive learning (ABL) framework, capable of unifying the three abilities effectively in a more user-friendly and understandable manner. Initially, the proposed method uses LLMs to correct the incomplete logical facts for optimizing the perception module, by summarizing and reorganizing domain knowledge represented in natural language format. Then, the perception module also provides necessary logical reasoning materials for feeding LLMs. Finally, these parts are integrated into a dynamic closed-loop system by introducing the feedback form and automatic learning strategies to mutually promote their performance. As a testbed, the variable-length handwritten equation decipherment (HED), an abstract expression of the Mayan calendar decoding, is used to demonstrate that ChatABL has reasoning ability beyond most existing state-of-the-art methods, which has been well-supported by comparative studies. To the best of authors' knowledge, the proposed ChatABL is the first attempt to explore a possible and novel avenue to approaching human-level cognitive ability via natural language interaction by means of ChatGPT.
Tianyang Zhong, Yi Pan 0001, Yutong Zhang 0019, Yaonai Wei, Zhengliang Liu, Xiaozheng Wei, Wenjun Li 0001, Chong Ma 0004, Xi Jiang 0001, Dinggang Shen, Junwei Han 0001
IEEE Trans. Neural Networks Learn. Syst.11
2024 Mask-guided BERT for few-shot text classification
Wenxiong Liao, Zhengliang Liu, Haixing Dai, Zihao Wu 0001, Yiyang Zhang 0003, Yuzhong Chen 0002, Xi Jiang 0001, Dajiang Zhu, Sheng Li 0001, Wei Liu 0146, Tianming Liu 0001, Quanzheng Li, Hongmin Cai, Xiang Li 0001
Neurocomputing8
2024 Fusing multi-scale functional connectivity patterns via Multi-Branch Vision Transformer (MB-ViT) for macaque brain age prediction
Jingchao Zhou, Yuzhong Chen 0002, Xuewei Jin, Zhenxiang Xiao, Songyao Zhang, Tianming Liu 0001, Keith M. Kendrick, Xi Jiang 0001
Neural Networks10
2024 Brain Structural Connectivity Guided Vision Transformers for Identification of Functional Connectivity Characteristics in Preterm Neonates
abstract
Preterm birth is the leading cause of death in children under five years old, and is associated with a wide sequence of complications in both short and long term. In view of rapid neurodevelopment during the neonatal period, preterm neonates may exhibit considerable functional alterations compared to term ones. However, the identified functional alterations in previous studies merely achieve moderate classification performance, while more accurate functional characteristics with satisfying discrimination ability for better diagnosis and therapeutic treatment is underexplored. To address this problem, we propose a novel brain structural connectivity (SC) guided Vision Transformer (SCG-ViT) to identify functional connectivity (FC) differences among three neonatal groups: preterm, preterm with early postnatal experience, and term. Particularly, inspired by the neuroscience-derived information, a novel patch token of SC/FC matrix is defined, and the SC matrix is then adopted as an effective mask into the ViT model to screen out input FC patch embeddings with weaker SC, and to focus on stronger ones for better classification and identification of FC differences among the three groups. The experimental results on multi-modal MRI data of 437 neonatal brains from publicly released Developing Human Connectome Project (dHCP) demonstrate that SCG-ViT achieves superior classification ability compared to baseline models, and successfully identifies holistically different FC patterns among the three groups. Moreover, these different FCs are significantly correlated with the differential gene expressions of the three groups. In summary, SCG-ViT provides a powerfully brain-guided pipeline of adopting large-scale and data-intensive deep learning models for medical imaging-based diagnosis.
Yuzhong Chen 0002, Zhenxiang Xiao, Yusong Sun, Jingchao Zhou, Weitong Guo, Chong Ma 0004, Lin Zhao 0004, Keith M. Kendrick, Benjamin Becker, Tianming Liu 0001, Xi Jiang 0001
IEEE J. Biomed. Health Informatics17
2024 MGIML: Cancer Grading With Incomplete Radiology-Pathology Data via Memory Learning and Gradient Homogenization
abstract
Taking advantage of multi-modal radiology-pathology data with complementary clinical information for cancer grading is helpful for doctors to improve diagnosis efficiency and accuracy. However, radiology and pathology data have distinct acquisition difficulties and costs, which leads to incomplete-modality data being common in applications. In this work, we propose a Memory- and Gradient-guided Incomplete Modal-modal Learning (MGIML) framework for cancer grading with incomplete radiology-pathology data. Firstly, to remedy missing-modality information, we propose a Memory-driven Hetero-modality Complement (MH-Complete) scheme, which constructs modal-specific memory banks constrained by a coarse-grained memory boosting (CMB) loss to record generic radiology and pathology feature patterns, and develops a cross-modal memory reading strategy enhanced by a fine-grained memory consistency (FMC) loss to take missing-modality information from well-stored memories. Secondly, as gradient conflicts exist between missing-modality situations, we propose a Rotation-driven Gradient Homogenization (RG-Homogenize) scheme, which estimates instance-specific rotation matrices to smoothly change the feature-level gradient directions, and computes confidence-guided homogenization weights to dynamically balance gradient magnitudes. By simultaneously mitigating gradient direction and magnitude conflicts, this scheme well avoids the negative transfer and optimization imbalance problems. Extensive experiments on CPTAC-UCEC and CPTAC-PDA datasets show that the proposed MGIML framework performs favorably against state-of-the-art multi-modal methods on missing-modality situations.
Pengyu Wang 0005, Huaqi Zhang, Meilu Zhu, Xi Jiang 0001, Harry Qin, Yixuan Yuan
IEEE Trans. Medical Imaging4
2024 Adversarial Learning Based Node-Edge Graph Attention Networks for Autism Spectrum Disorder Identification
abstract
Graph neural networks (GNNs) have received increasing interest in the medical imaging field given their powerful graph embedding ability to characterize the non-Euclidean structure of brain networks based on magnetic resonance imaging (MRI) data. However, previous studies are largely node-centralized and ignore edge features for graph classification tasks, resulting in moderate performance of graph classification accuracy. Moreover, the generalizability of GNN model is still far from satisfactory in brain disorder [e.g., autism spectrum disorder (ASD)] identification due to considerable individual differences in symptoms among patients as well as data heterogeneity among different sites. In order to address the above limitations, this study proposes a novel adversarial learning-based node-edge graph attention network (AL-NEGAT) for ASD identification based on multimodal MRI data. First, both node and edge features are modeled based on structural and functional MRI data to leverage complementary brain information and preserved in the constructed weighted adjacent matrix for individuals through the attention mechanism in the proposed NEGAT. Second, two AL methods are employed to improve the generalizability of NEGAT. Finally, a gradient-based saliency map strategy is utilized for model interpretation to identify important brain regions and connections contributing to the classification. Experimental results based on the public Autism Brain Imaging Data Exchange I (ABIDE I) data demonstrate that the proposed framework achieves a classification accuracy of 74.7% between ASD and typical developing (TD) groups based on 1007 subjects across 17 different sites and outperforms the state-of-the-art methods, indicating satisfying classification ability and generalizability of the proposed AL-NEGAT model. Our work provides a powerful tool for brain disorder identification.
Yuzhong Chen 0002, Jiadong Yan, Mingxin Jiang, Zhongbo Zhao, Weihua Zhao, Jian Zheng 0001, Dezhong Yao 0001, Keith M. Kendrick, Xi Jiang 0001
IEEE Trans. Neural Networks Learn. Syst.11
2024 Anatomy-Guided Spatio-Temporal Graph Convolutional Networks (AG-STGCNs) for Modeling Functional Connectivity Between Gyri and Sulci Across Multiple Task Domains
abstract
The cerebral cortex is folded as gyri and sulci, which provide the foundation to unveil anatomo-functional relationship of brain. Previous studies have extensively demonstrated that gyri and sulci exhibit intrinsic functional difference, which is further supported by morphological, genetic, and structural evidences. Therefore, systematically investigating the gyro-sulcal (G-S) functional difference can help deeply understand the functional mechanism of brain. By integrating functional magnetic resonance imaging (fMRI) with advanced deep learning models, recent studies have unveiled the temporal difference in functional activity between gyri and sulci. However, the potential difference of functional connectivity, which represents functional dependency between gyri and sulci, is much unknown. Moreover, the regularity and variability of the G-S functional connectivity difference across multiple task domains remains to be explored. To address the two concerns, this study developed new anatomy-guided spatio-temporal graph convolutional networks (AG-STGCNs) to investigate the regularity and variability of functional connectivity differences between gyri and sulci across multiple task domains. Based on 830 subjects with seven different task-based and one resting state fMRI (rs-fMRI) datasets from the public Human Connectome Project (HCP), we consistently found that there are significant differences of functional connectivity between gyral and sulcal regions within task domains compared with resting state (RS). Furthermore, there is considerable variability of such functional connectivity and information flow between gyri and sulci across different task domains, which are correlated with individual cognitive behaviors. Our study helps better understand the functional segregation of gyri and sulci within task domains as well as the anatomo-functional-behavioral relationship of the human brain.
Mingxin Jiang, Yuzhong Chen 0002, Jiadong Yan, Zhenxiang Xiao, Shimin Yang, Zhongbo Zhao, Lei Guo 0002, Benjamin Becker, Dezhong Yao 0001, Keith M. Kendrick, Xi Jiang 0001
IEEE Trans. Neural Networks Learn. Syst.14
2024 Deep Metric Learning Based on Meta-Mining Strategy With Semiglobal Information
abstract
Recently, deep metric learning (DML) has achieved great success. Some existing DML methods propose adaptive sample mining strategies, which learn to weight the samples, leading to interesting performance. However, these methods suffer from a small memory (e.g., one training batch), limiting their efficacy. In this work, we introduce a data-driven method, meta-mining strategy with semiglobal information (MMSI), to apply meta-learning to learn to weight samples during the whole training, leading to an adaptive mining strategy. To introduce richer information than one training batch only, we elaborately take advantage of the validation set of meta-learning by implicitly adding additional validation sample information to training. Furthermore, motivated by the latest self-supervised learning, we introduce a dictionary (memory) that maintains very large and diverse information. Together with the validation set, this dictionary presents much richer information to the training, leading to promising performance. In addition, we propose a new theoretical framework that can formulate pairwise and tripletwise metric learning loss functions in a unified framework. This framework brings new insights to society and facilitates us to generalize our MMSI to many existing DML methods. We conduct extensive experiments on three public datasets, CUB200-2011, Cars-196, and Stanford Online Products (SOP). Results show that our method can achieve the state of the art or very competitive performance. Our source codes have been made available at https://github.com/NUST-Machine-Intelligence-Laboratory/MMSI.
Xi Jiang 0001, Sheng Liu 0009, Xili Dai, Guosheng Hu, Xingguo Huang, Yazhou Yao, Guosen Xie, Ling Shao 0001
IEEE Trans. Neural Networks Learn. Syst.1
2024 Rectify ViT Shortcut Learning by Visual Saliency
abstract
Shortcut learning in deep learning models occurs when unintended features are prioritized, resulting in degenerated feature representations and reduced generalizability and interpretability. However, shortcut learning in the widely used vision transformer (ViT) framework is largely unknown. Meanwhile, introducing domain-specific knowledge is a major approach to rectifying the shortcuts that are predominated by background-related factors. For example, eye-gaze data from radiologists are effective human visual prior knowledge that has the great potential to guide the deep learning models to focus on meaningful foreground regions. However, obtaining eye-gaze data can still sometimes be time-consuming, labor-intensive, and even impractical. In this work, we propose a novel and effective saliency-guided ViT (SGT) model to rectify shortcut learning in ViT with the absence of eye-gaze data. Specifically, a computational visual saliency model (either pretrained or fine-tuned) is adopted to predict saliency maps for input image samples. Then, the saliency maps are used to filter the most informative image patches. Considering that this filter operation may lead to global information loss, we further introduce a residual connection that calculates the self-attention across all the image patches. The experiment results on natural and medical image datasets show that our SGT framework can effectively learn and leverage human prior knowledge without eye-gaze data and achieves much better performance than baselines. Meanwhile, it successfully rectifies the harmful shortcut learning and significantly improves the interpretability of the ViT model, demonstrating the promise of transferring human prior knowledge derived visual saliency in rectifying shortcut learning.
Chong Ma 0004, Lin Zhao 0004, Yuzhong Chen 0002, Lei Guo 0002, Xintao Hu, Dinggang Shen, Xi Jiang 0001, Tianming Liu 0001
IEEE Trans. Neural Networks Learn. Syst.8
2023 A-GCL: Adversarial graph contrastive learning for fMRI analysis to diagnose neurodevelopmental disorders
Xiang Chen 0031, Bohan Ren, Haibo Yang 0002, Xi Jiang 0001, Dinggang Shen, Yuan Zhou 0004, Xiao-Yong Zhang
Medical Image Anal.7
2023 Characterizing functional brain networks via Spatio-Temporal Attention 4D Convolutional Neural Networks (STA-4DCNNs)
Xi Jiang 0001, Jiadong Yan, Yu Zhao 0007, Mingxin Jiang, Yuzhong Chen 0002, Jingchao Zhou, Zhenxiang Xiao, Benjamin Becker, Dajiang Zhu, Keith M. Kendrick, Tianming Liu 0001
Neural Networks1
2023 Eye-Gaze-Guided Vision Transformer for Rectifying Shortcut Learning
abstract
Learning harmful shortcuts such as spurious correlations and biases prevents deep neural networks from learning meaningful and useful representations, thus jeopardizing the generalizability and interpretability of the learned representation. The situation becomes even more serious in medical image analysis, where the clinical data are limited and scarce while the reliability, generalizability and transparency of the learned model are highly required. To rectify the harmful shortcuts in medical imaging applications, in this paper, we propose a novel eye-gaze-guided vision transformer (EG-ViT) model which infuses the visual attention from radiologists to proactively guide the vision transformer (ViT) model to focus on regions with potential pathology rather than spurious correlations. To do so, the EG-ViT model takes the masked image patches that are within the radiologists' interest as input while has an additional residual connection to the last encoder layer to maintain the interactions of all patches. The experiments on two medical imaging datasets demonstrate that the proposed EG-ViT model can effectively rectify the harmful shortcut learning and improve the interpretability of the model. Meanwhile, infusing the experts' domain knowledge can also improve the large-scale ViT model's performance over all compared baseline methods with limited samples available. In general, EG-ViT takes the advantages of powerful deep neural networks while rectifies the harmful shortcut learning with human expert's prior knowledge. This work also opens new avenues for advancing current artificial intelligence paradigms by infusing human intelligence.
Chong Ma 0004, Lin Zhao 0004, Yuzhong Chen 0002, Sheng Wang 0014, Lei Guo 0002, Dinggang Shen, Xi Jiang 0001, Tianming Liu 0001
IEEE Trans. Medical Imaging8
2022 Modeling spatio-temporal patterns of holistic functional brain networks via multi-head guided attention graph neural networks (Multi-Head GAGNNs)
Jiadong Yan, Yuzhong Chen 0002, Zhenxiang Xiao, Shu Zhang 0001, Mingxin Jiang, Jinglei Lv, Benjamin Becker, Dajiang Zhu, Junwei Han 0001, Dezhong Yao 0001, Keith M. Kendrick, Tianming Liu 0001, Xi Jiang 0001
Medical Image Anal.16
2021 Multi-head GAGNN: A Multi-head Guided Attention Graph Neural Network for Modeling Spatio-temporal Patterns of Holistic Brain Functional Networks
Jiadong Yan, Yuzhong Chen 0002, Shimin Yang, Shu Zhang 0001, Mingxin Jiang, Zhongbo Zhao, Yu Zhao 0007, Benjamin Becker, Tianming Liu 0001, Keith M. Kendrick, Xi Jiang 0001
MICCAI (7)12
2021 Exploring the Functional Difference of Gyri/Sulci via Hierarchical Interpretable Autoencoder
Lin Zhao 0004, Haixing Dai, Xi Jiang 0001, Dajiang Zhu, Tianming Liu 0001
MICCAI (7)3
2021 Attention-Based Node-Edge Graph Convolutional Networks for Identification of Autism Spectrum Disorder Using Multi-Modal MRI Data
Yuzhong Chen 0002, Jiadong Yan, Mingxin Jiang, Zhongbo Zhao, Weihua Zhao, Keith M. Kendrick, Xi Jiang 0001
PRCV (3)8
2021 A Guided Attention 4D Convolutional Neural Network for Modeling Spatio-Temporal Patterns of Functional Brain Networks
Jiadong Yan, Yu Zhao 0007, Mingxin Jiang, Shu Zhang 0001, Shimin Yang, Yuzhong Chen 0002, Zhongbo Zhao, Benjamin Becker, Tianming Liu 0001, Keith M. Kendrick, Xi Jiang 0001
PRCV (3)13
2020 Species-Shared and -Specific Structural Connections Revealed by Dirty Multi-task Regression
Xi Jiang 0001, Lei Guo 0002, Xiaoping Hu 0001, Tianming Liu 0001, Lei Du 0001
MICCAI (7)3
2020 Identifying Cross-individual Correspondences of 3-hinge Gyri
Ying Huang 0007, Lin Zhao 0004, Xi Jiang 0001, Lei Guo 0002, Xiaoping Hu 0001, Tianming Liu 0001
Medical Image Anal.5
2019 Identify Hierarchical Structures from Task-Based fMRI Data via Hybrid Spatiotemporal Neural Architecture Search Net
Wei Zhang 0090, Lin Zhao 0004, Qing Li 0027, Shijie Zhao 0001, Qinglin Dong, Xi Jiang 0001, Tianming Liu 0001
MICCAI (3)6
2017 Joint Representation of Connectome-Scale Structural and Functional Profiles for Identification of Consistent Cortical Landmarks in Human Brains
Shu Zhang 0001, Xi Jiang 0001, Tianming Liu 0001
MICCAI (1)2
2017 Task fMRI data analysis based on supervised stochastic coordinate coding
Jinglei Lv, Qingyang Li 0001, Wei Zhang 0090, Yu Zhao 0007, Xi Jiang 0001, Lei Guo 0002, Junwei Han 0001, Xintao Hu, Christine Cong Guo, Jieping Ye, Tianming Liu 0001
Medical Image Anal.6
2016 Exploring auditory network composition during free listening to audio excerpts via group-wise sparse representation
abstract
With the growing number of audio excerpts through various media and distribution channels, advanced audio analysis approaches have received significant interest in the multimedia field. However, current audio analysis approaches are still far from satisfactory due to the semantic gaps between the low-level acoustic features and high-level semantics perceived by human brain. In order to alleviate the problem, this paper propose a novel computational framework to bridge acoustic features with high-level semantic features derived from functional magnetic resonance imaging (fMRI) signals which record the brain's response during free listening to music/speech excerpts, and to explore the brain auditory network composition of acoustic features for different types of music/speech excerpts. Specifically, we identify meaningful brain networks and corresponding brain activities representing high-level semantic features via a novel group-wise sparse representation of whole brain fMRI signals. Then we associate the brain activities with specific low-level acoustic features and analyze the auditory network composition of acoustic features for different types of music/speech excerpts. Experimental results demonstrate that multiple acoustic features are involved in the brain auditory networks during free listening to music/speech excerpts. Meanwhile, there is considerable variability of auditory network composition of acoustic features for different types of music/speech. Our results provide new insights of how to narrow the semantic gaps in audio content analysis.
Shijie Zhao 0001, Junwei Han 0001, Xi Jiang 0001, Xintao Hu, Jinglei Lv, Shu Zhang 0001, Bao Ge, Lei Guo 0002, Tianming Liu 0001
ICME3
2016 Modeling Functional Dynamics of Cortical Gyri and Sulci
Xi Jiang 0001, Xiang Li 0001, Jinglei Lv, Shijie Zhao 0001, Shu Zhang 0001, Wei Zhang 0090, Tianming Liu 0001
MICCAI (1)1
2016 Discover Mouse Gene Coexpression Landscape Using Dictionary Learning and Sparse Coding
Yujie Li 0004, Hanbo Chen, Xi Jiang 0001, Xiang Li 0001, Jinglei Lv, Hanchuan Peng, Joe Z. Tsien, Tianming Liu 0001
MICCAI (1)3
2016 Species Preserved and Exclusive Structural Connections Revealed by Sparse CCA
Xiao Li 0024, Lei Du 0001, Xintao Hu, Xi Jiang 0001, Lei Guo 0002, Tianming Liu 0001
MICCAI (1)5
2016 A Multi-stage Sparse Coding Framework to Explore the Effects of Prenatal Alcohol Exposure
Shijie Zhao 0001, Junwei Han 0001, Jinglei Lv, Xi Jiang 0001, Xintao Hu, Shu Zhang 0001, Mary Ellen Lynch, Claire Coles, Lei Guo 0002, Xiaoping Hu 0001, Tianming Liu 0001
MICCAI (1)4
2016 Exploring Brain Networks via Structured Sparse Representation of fMRI Data
Jianfeng Lu 0003, Jinglei Lv, Xi Jiang 0001, Shijie Zhao 0001, Tianming Liu 0001
MICCAI (1)4
2016 Group-wise consistent cortical parcellation based on connectional profiles
Dajiang Zhu, Xi Jiang 0001, Shu Zhang 0001, Zhifeng Kou, Lei Guo 0002, Tianming Liu 0001
Medical Image Anal.3
2016 Predicting Movie Trailer Viewer's "Like/Dislike" via Learned Shot Editing Patterns
abstract
Nowadays, there are many movie trailers publicly available on social media website such as YouTube, and many thousands of users have independently indicated whether they like or dislike those trailers. Although it is understandable that there are multiple factors that could influence viewers' like or dislike of the trailer, we aim to address a preference question in this work: Can subjective multimedia features be developed to predict the viewer's preference presented by like (by thumbs-up) or dislike (by thumbs-down) during and after watching movie trailers? We designed and implemented a computational framework that is composed of low-level multimedia feature extraction, feature screening and selection, and classification, and applied it to a collection of 725 movie trailers. Experimental results demonstrated that, among dozens of multimedia features, the single low-level multimedia feature of shot length variance is highly predictive of a viewer's “like/dislike” for a large portion of movie trailers. We interpret these findings such that variable shot lengths in a trailer tend to produce a rhythm that is likely to stimulate a viewer's positive preference. This conclusion was also proved by the repeatability experiments results using another 600 trailer videos and it was further interpreted by viewers'eye-tracking data.
Shu Zhang 0001, Xi Jiang 0001, Xiang Li 0001, Xintao Hu, Junwei Han 0001, Lei Guo 0002, L. Stephen Miller, Richard Neupert, Tianming Liu 0001
IEEE Trans. Affect. Comput.4
2015 Longitudinal Analysis of Brain Recovery after Mild Traumatic Brain Injury Based on Groupwise Consistent Brain Network Clusters
Hanbo Chen, Armin Iraji, Xi Jiang 0001, Jinglei Lv, Zhifeng Kou, Tianming Liu 0001
MICCAI (2)3
2015 Fiber Connection Pattern-Guided Structured Sparse Representation of Whole-Brain fMRI Signals for Functional Network Inference
Xi Jiang 0001, Jianfeng Lu 0003, Lei Guo 0002, Tianming Liu 0001
MICCAI (1)1
2015 Modeling Task FMRI Data via Supervised Stochastic Coordinate Coding
Jinglei Lv, Wei Zhang 0090, Xi Jiang 0001, Xintao Hu, Junwei Han 0001, Lei Guo 0002, Jieping Ye, Tianming Liu 0001
MICCAI (1)4
2015 Multi-scale and Multimodal Fusion of Tract-Tracing, Myelin Stain and DTI-derived Fibers in Macaque Brains
Ke Jing, Hanbo Chen, Xi Jiang 0001, Longchuan Li, Lei Guo 0002, Jianfeng Lu 0003, Xiaoping Hu 0001, Tianming Liu 0001
MICCAI (2)5
2015 Analysis of music/speech via integration of audio content and functional brain response
Junwei Han 0001, Xi Jiang 0001, Xintao Hu, Lei Guo 0002, Jungong Han, Ling Shao 0001, Tianming Liu 0001
Inf. Sci.3
2015 Sparse representation of whole-brain fMRI signals for identification of functional networks
Jinglei Lv, Xi Jiang 0001, Xiang Li 0001, Dajiang Zhu, Hanbo Chen, Shu Zhang 0001, Xintao Hu, Junwei Han 0001, Heng Huang 0001, Jing Zhang 0010, Lei Guo 0002, Tianming Liu 0001
Medical Image Anal.2
2015 Supervised Dictionary Learning for Inferring Concurrent Brain Networks
abstract
Task-based fMRI (tfMRI) has been widely used to explore functional brain networks via predefined stimulus paradigm in the fMRI scan. Traditionally, the general linear model (GLM) has been a dominant approach to detect task-evoked networks. However, GLM focuses on task-evoked or event-evoked brain responses and possibly ignores the intrinsic brain functions. In comparison, dictionary learning and sparse coding methods have attracted much attention recently, and these methods have shown the promise of automatically and systematically decomposing fMRI signals into meaningful task-evoked and intrinsic concurrent networks. Nevertheless, two notable limitations of current data-driven dictionary learning method are that the prior knowledge of task paradigm is not sufficiently utilized and that the establishment of correspondences among dictionary atoms in different brains have been challenging. In this paper, we propose a novel supervised dictionary learning and sparse coding method for inferring functional networks from tfMRI data, which takes both of the advantages of model-driven method and data-driven method. The basic idea is to fix the task stimulus curves as predefined model-driven dictionary atoms and only optimize the other portion of data-driven dictionary atoms. Application of this novel methodology on the publicly available human connectome project (HCP) tfMRI datasets has achieved promising results.
Shijie Zhao 0001, Junwei Han 0001, Jinglei Lv, Xi Jiang 0001, Xintao Hu, Yu Zhao 0007, Bao Ge, Lei Guo 0002, Tianming Liu 0001
IEEE Trans. Medical Imaging4
2014 Decoding Auditory Saliency from FMRI Brain Imaging
abstract
Given the growing number of available audio streams through a variety of sources and distribution channels, effective and advanced computational audio analysis has received increasing interest in the multimedia field. However, the effectiveness of current audio analysis strategies might be hampered due to the lack of effective representation of high-level semantics perceived by the human and the lack of effective approaches to bridging the gaps between most low-level acoustic features and high-level semantic features. This semantic gap has become the 'bottleneck' problem in audio analysis. In this paper, we propose a computational framework to decode biologically-plausible auditory saliency using high-level features derived from functional magnetic resonance imaging (fMRI) which monitors the human brain's response under the natural stimulus of audio listening. Specifically, we identify meaningful intrinsic brain networks which are involved in audio listening via effective online dictionary learning and sparse representation of whole-brain fMRI signals, reconstruct auditory saliency features using those identified brain network components, and perform group-wise analysis to identify consistent 'brain decoders' of the saliency features across different excerpts and participants. Experimental results demonstrate that the auditory saliency features are effectively decoded via our methods, which potentially provide opportunities for various applications in the multimedia field.
Shijie Zhao 0001, Xi Jiang 0001, Junwei Han 0001, Xintao Hu, Dajiang Zhu, Jinglei Lv, Lei Guo 0002, Tianming Liu 0001
ACM Multimedia2
2013 Predictive Models of Resting State Networks for Assessment of Altered Functional Connectivity in MCI
Xi Jiang 0001, Dajiang Zhu, Kaiming Li, Dinggang Shen, Lei Guo 0002, Tianming Liu 0001
MICCAI (2)1
2013 Anatomy-Guided Discovery of Large-Scale Consistent Connectivity-Based Cortical Landmarks
Xi Jiang 0001, Dajiang Zhu, Kaiming Li, Jinglei Lv, Lei Guo 0002, Tianming Liu 0001
MICCAI (3)1
2013 Sparse Representation of Group-Wise FMRI Signals
Jinglei Lv, Xiang Li 0001, Dajiang Zhu, Xi Jiang 0001, Xin Zhang 0151, Xintao Hu, Lei Guo 0002, Tianming Liu 0001
MICCAI (3)4
2013 Sparse Representation of Higher-Order Functional Interaction Patterns in Task-Based FMRI Data
Shu Zhang 0001, Xiang Li 0001, Jinglei Lv, Xi Jiang 0001, Dajiang Zhu, Hanbo Chen, Lei Guo 0002, Tianming Liu 0001
MICCAI (3)4
2013 Predicting cortical ROIs via joint modeling of anatomical and connectional profiles
Dajiang Zhu, Xi Jiang 0001, Bao Ge, Xintao Hu, Junwei Han 0001, Lei Guo 0002, Tianming Liu 0001
Medical Image Anal.3
2013 Representing and Retrieving Video Shots in Human-Centric Brain Imaging Space
abstract
Meaningful representation and effective retrieval of video shots in a large-scale database has been a profound challenge for the image/video processing and computer vision communities. A great deal of effort has been devoted to the extraction of low-level visual features, such as color, shape, texture, and motion for characterizing and retrieving video shots. However, the accuracy of these feature descriptors is still far from satisfaction due to the well-known semantic gap. In order to alleviate the problem, this paper investigates a novel methodology of representing and retrieving video shots using human-centric high-level features derived in brain imaging space (BIS) where brain responses to natural stimulus of video watching can be explored and interpreted. At first, our recently developed dense individualized and common connectivity-based cortical landmarks (DICCCOL) system is employed to locate large-scale functional brain networks and their regions of interests (ROIs) that are involved in the comprehension of video stimulus. Then, functional connectivities between various functional ROI pairs are utilized as BIS features to characterize the brain's comprehension of video semantics. Then an effective feature selection procedure is applied to learn the most relevant features while removing redundancy, which results in the formation of the final BIS features. Afterwards, a mapping from low-level visual features to high-level semantic features in the BIS is built via the Gaussian process regression (GPR) algorithm, and a manifold structure is then inferred, in which video key frames are represented by the mapped feature vectors in the BIS. Finally, the manifold-ranking algorithm concerning the relationship among all data is applied to measure the similarity between key frames of video shots. Experimental results on the TRECVID 2005 dataset demonstrate the superiority of the proposed work in comparison with traditional methods.
Junwei Han 0001, Xintao Hu, Dajiang Zhu, Kaiming Li, Xi Jiang 0001, Guangbin Cui, Lei Guo 0002, Tianming Liu 0001
IEEE Trans. Image Process.6
2013 Inferring Group-Wise Consistent Multimodal Brain Networks via Multi-View Spectral Clustering
abstract
Quantitative modeling and analysis of structural and functional brain networks based on diffusion tensor imaging (DTI) and functional magnetic resonance imaging (fMRI) data have received extensive interest recently. However, the regularity of these structural and functional brain networks across multiple neuroimaging modalities and also across different individuals is largely unknown. This paper presents a novel approach to inferring group-wise consistent brain subnetworks from multimodal DTI/resting-state fMRI datasets via multi-view spectral clustering of cortical networks, which were constructed upon our recently developed and validated large-scale cortical landmarks-DICCCOL (dense individualized and common connectivity-based cortical landmarks). We applied the algorithms on DTI data of 100 healthy young females and 50 healthy young males, obtained consistent multimodal brain networks within and across multiple groups, and further examined the functional roles of these networks. Our experimental results demonstrated that the derived brain networks have substantially improved inter-modality and inter-subject consistency.
Hanbo Chen, Kaiming Li, Dajiang Zhu, Xi Jiang 0001, Yixuan Yuan, Peili Lv, Lei Guo 0002, Dinggang Shen, Tianming Liu 0001
IEEE Trans. Medical Imaging4
2012 Music/speech classification using high-level features derived from fmri brain imaging
abstract
With the availability of large amount of audio tracks through a variety of sources and distribution channels, automatic music/speech classification becomes an indispensable tool in social audio websites and online audio communities. However, the accuracy of current acoustic-based low-level feature classification methods is still rather far from satisfaction. The discrepancy between the limited descriptive power of low-level features and the richness of high-level semantics perceived by the human brain has become the 'bottleneck' problem in audio signal analysis. In this paper, functional magnetic resonance imaging (fMRI) which monitors the human brain's response under the natural stimulus of music/speech listening is used as high-level features in the brain imaging space (BIS). We developed a computational framework to model the relationships between BIS features and low-level features in the training dataset with fMRI scans, predict BIS features of testing dataset without fMRI scans, and use the predicted BIS features for music/speech classification in the application stage. Experimental results demonstrated the significantly improved performance of music/speech classification via predicted BIS features than that via the original low-level features.
Xi Jiang 0001, Xintao Hu, Lie Lu, Junwei Han 0001, Lei Guo 0002, Tianming Liu 0001
ACM Multimedia1
2010 Bridging low-level features and high-level semantics via fMRI brain imaging for video classification
abstract
The multimedia content analysis community has made significant effort to bridge the gap between low-level features and high-level semantics perceived by human cognitive systems such as real-world objects and concepts. In the two fields of multimedia analysis and brain imaging, both topics of low-level features and high level semantics are extensively studied. For instance, in the multimedia analysis field, many algorithms are available for multimedia feature extraction, and benchmark datasets are available such as the TRECVID. In the brain imaging field, brain regions that are responsible for vision, auditory perception, language, and working memory are well studied via functional magnetic resonance imaging (fMRI). This paper presents our initial effort in marrying these two fields in order to bridge the gaps between low-level features and high-level semantics via fMRI brain imaging. Our experimental paradigm is that we performed fMRI brain imaging when university student subjects watched the video clips selected from the TRECVID datasets. At current stage, we focus on the three concepts of sports, weather, and commercial-/advertisement specified in the TRECVID 2005. Meanwhile, the brain regions in vision, auditory, language, and working memory networks are quantitatively localized and mapped via task-based paradigm fMRI, and the fMRI responses in these regions are used to extract features as the representation of the brain's comprehension of semantics. Our computational framework aims to learn the most relevant low-level feature sets that best correlate the fMRI-derived semantics based on the training videos with fMRI scans, and then the learned models are applied to larger scale test datasets without fMRI scans for category classifications. Our result shows that: 1) there are meaningful couplings between brain's fMRI responses and video stimuli, suggesting the validity of linking semantics and low-level features via fMRI; 2) The computationally learned low-level feature sets from fMRI-derived semantic features can significantly improve the classification of video categories in comparison with that based on original low-level features.
Xintao Hu, Fan Deng 0001, Kaiming Li, Hanbo Chen, Xi Jiang 0001, Jinglei Lv, Dajiang Zhu, Carlos Faraco, Degang Zhang, Arsham Mesbah, Junwei Han 0001, Xian-Sheng Hua 0001, L. Stephen Miller, Lei Guo 0002, Tianming Liu 0001
ACM Multimedia6
2010 Individualized ROI Optimization via Maximization of Group-wise Consistency of Structural and Functional Profiles
abstract
Functional segregation and integration are fundamental characteristics of the human brain. Studying the connectivity among segregated regions and the dynamics of integrated brain networks has drawn increasing interest. A very controversial, yet fundamental issue in these studies is how to determine the best functional brain regions or ROIs (regions of interests) for individuals. Essentially, the computed connectivity patterns and dynamics of brain networks are very sensitive to the locations, sizes, and shapes of the ROIs. This paper presents a novel methodology to optimize the locations of an individual's ROIs in the working memory system. Our strategy is to formulate the individual ROI optimization as a group variance minimization problem, in which group-wise functional and structural connectivity patterns, and anatomic profiles are defined as optimization constraints. The optimization problem is solved via the simulated annealing approach. Our experimental results show that the optimized ROIs have significantly improved consistency in structural and functional profiles across subjects, and have more reasonable localizations and more consistent morphological and anatomic profiles.
Kaiming Li, Lei Guo 0002, Carlos Faraco, Dajiang Zhu, Fan Deng 0001, Xi Jiang 0001, Degang Zhang, Hanbo Chen, Xintao Hu, L. Stephen Miller, Tianming Liu 0001
NIPS7