VLDB 2026 Research / reviewers in the wild / expert
Xintao Hu
dblp:97/8526
· DBLP profile ↗
49ranked-venue papers
3as first author
11since 2021 · last 2025
0000-0001-5633-3806ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 28 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 22 · 3 since 2021Artificial intelligence and machine learning · 10 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Language models and text generation · 35% Image recognition and object detection · 24% Trustworthy machine learning · 20% | |
| Computer graphics and multimedia
5 papers |
Multimedia analysis and retrieval · 75% Audio and music processing · 25% | |
| Interdisciplinary, comprehensive, and emerging computing
5 papers |
Medical and health informatics · 64% Bioinformatics and computational biology · 36% |
Topics — the 18 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Image recognition and object detection
saliency prediction |
0.9 | 2 | 2024 | BI-AVAN: A Brain-Inspired Adversarial Visual Attention Network for Characterizing Human Visual Attention From Neural Activity · IEEE Trans. Multim. 2024 A biologically inspired computational model for image saliency detection · ACM Multimedia 2011 |
Machine learning › Trustworthy machine learning
interpretability |
0.8 | 1 | 2024 | BI-AVAN: A Brain-Inspired Adversarial Visual Attention Network for Characterizing Human Visual Attention From Neural Activity · IEEE Trans. Multim. 2024 |
Natural language and speech › Language models and text generation › pre-trained language model › pretrained language model analysis
BERT analysis |
0.7 | 1 | 2023 | Coupling Artificial Neurons in BERT and Biological Neurons in the Human Brain · AAAI 2023 |
Natural language and speech › Language models and text generation › language modeling › language model architecture
transformer language model |
0.7 | 1 | 2023 | Coupling Artificial Neurons in BERT and Biological Neurons in the Human Brain · AAAI 2023 |
Medical and health informatics
neuroimaging |
0.3 | 3 | 2014 | Decoding Auditory Saliency from FMRI Brain Imaging · ACM Multimedia 2014 Music/speech classification using high-level features derived from fmri brain imaging · ACM Multimedia 2012 Bridging low-level features and high-level semantics via fMRI brain imaging for video classification · ACM Multimedia 2010 |
Multimedia analysis and retrieval
video classification |
0.3 | 2 | 2012 | Bridging the Semantic Gap via Functional Brain Imaging · IEEE Trans. Multim. 2012 Bridging low-level features and high-level semantics via fMRI brain imaging for video classification · ACM Multimedia 2010 |
Bioinformatics and computational biology › neuroscience › neuroinformatics › neural data analysis
neural activity analysis |
0.2 | 1 | 2024 | BI-AVAN: A Brain-Inspired Adversarial Visual Attention Network for Characterizing Human Visual Attention From Neural Activity · IEEE Trans. Multim. 2024 |
Medical and health informatics › neuroimaging
fMRI decoding |
0.2 | 1 | 2014 | Decoding Auditory Saliency from FMRI Brain Imaging · ACM Multimedia 2014 |
Audio and music processing
audio analysis |
0.2 | 1 | 2014 | Decoding Auditory Saliency from FMRI Brain Imaging · ACM Multimedia 2014 |
Multimedia analysis and retrieval
cross-modal retrieval |
0.2 | 1 | 2013 | Representing and Retrieving Video Shots in Human-Centric Brain Imaging Space · IEEE Trans. Image Process. 2013 |
Multimedia analysis and retrieval
video retrieval |
0.2 | 1 | 2013 | Representing and Retrieving Video Shots in Human-Centric Brain Imaging Space · IEEE Trans. Image Process. 2013 |
Multimedia analysis and retrieval › video retrieval
video shot retrieval |
0.2 | 1 | 2013 | Representing and Retrieving Video Shots in Human-Centric Brain Imaging Space · IEEE Trans. Image Process. 2013 |
Audio and music processing › audio classification
speech/music discrimination |
0.1 | 1 | 2012 | Music/speech classification using high-level features derived from fmri brain imaging · ACM Multimedia 2012 |
Multimedia analysis and retrieval
video content analysis |
0.1 | 1 | 2012 | Bridging the Semantic Gap via Functional Brain Imaging · IEEE Trans. Multim. 2012 |
Computer vision › Segmentation and scene understanding
saliency detection |
0.1 | 1 | 2011 | A biologically inspired computational model for image saliency detection · ACM Multimedia 2011 |
Bioinformatics and computational biology › neuroscience › neuroinformatics
brain network analysis |
0.1 | 1 | 2010 | Individualized ROI Optimization via Maximization of Group-wise Consistency of Structural and Functional Profiles · NIPS 2010 |
Medical and health informatics › neuroimaging › neuroimaging analysis
fMRI analysis |
0.0 | 1 | 2010 | Bridging low-level features and high-level semantics via fMRI brain imaging for video classification · ACM Multimedia 2010 |
Mathematical optimization
combinatorial optimization |
0.0 | 1 | 2010 | Individualized ROI Optimization via Maximization of Group-wise Consistency of Structural and Functional Profiles · NIPS 2010 |
Methods — techniques the papers use, named apart from their topics
unsupervised learning · 1.5adversarial learning · 1.5fMRI · 1.3eye-tracking · 0.9eye tracking · 0.8synchronization analysis · 0.7functional brain network · 0.7feature selection · 0.5sparse representation · 0.4dictionary learning · 0.4brain imaging space features · 0.3simulated annealing · 0.2group variance minimization · 0.2manifold ranking · 0.2gaussian process regression · 0.2basis function learning · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Understanding LLMs: A comprehensive overview from training to inference
Tianle Han, Jiaming Tian, Yutong Zhang 0019, Jiaqi Wang 0010, Xiaohui Gao, Tianyang Zhong, Yi Pan 0001, Shaochen Xu, Zihao Wu 0001, Zhengliang Liu, Xin Zhang 0151, Shu Zhang 0001, Xintao Hu, Ning Qiang, Tianming Liu 0001, Bao Ge |
Neurocomputing | 17 |
| 2025 | A Foundational fMRI Model for Representing Continuous Brain StatesabstractFoundational models have significant potential to advance brain function research, particularly in understanding the dynamics of brain states. However, most existing models process brain signals within fixed time windows, restricting their ability to capture the full temporal complexity of brain activity. In this study, we propose BrainSN (Brain States Network), a novel fMRI foundational model designed to represent continuous brain state information and support diverse downstream tasks. First, leveraging a transformer-based architecture, BrainSN reconstructs input brain states across multiple time scales and predicts future brain activity, effectively capturing both short-term and long-term dependencies. Second, through multiple embeddings and a channel gating module, the model integrates brain state information and applies an attention mechanism to extract critical features. Additionally, we train BrainSN on 1,256 hours of resting-state and naturalistic stimulus fMRI data, enabling it to learn large-scale brain dynamics without relying on task-based paradigms. Without fine-tuning, BrainSN achieves 75.23% and 75.82% accuracy in autism and attention disorder diagnosis tasks, respectively, matching the performance of leading models pretrained on disease-specific data. After fine-tuning, it surpasses these models. In mental state decoding, BrainSN attains 95.31% accuracy without fine-tuning, outperforming the best models trained on large-scale task-based fMRI data. Furthermore, by analyzing BrainSN's embeddings in relation to movie stimuli, we demonstrate that the model effectively captures the semantic content of movie scenes embedded in fMRI signals and is highly sensitive to sequence. These results highlight BrainSN's ability to model brain state dynamics and underscore its potential advantages for clinical diagnosis, treatment evaluation, and cognitive neuroscience research. Lei Guo 0002, Yixuan Yuan, Junwei Han 0001, Xintao Hu |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | A Unified and Biologically Plausible Relational Graph Representation of Vision TransformersabstractVision transformer (ViT) and its variants have achieved remarkable success in various tasks. The key characteristic of these ViT models is to adopt different aggregation strategies of spatial patch information within the artificial neural networks (ANNs). However, there is still a key lack of unified representation of different ViT architectures for systematic understanding and assessment of model representation performance. Moreover, how those well-performing ViT ANNs are similar to real biological neural networks (BNNs) is largely unexplored. To answer these fundamental questions, we, for the first time, propose a unified and biologically plausible relational graph representation of ViT models. Specifically, the proposed relational graph representation consists of two key subgraphs: an aggregation graph and an affine graph. The former considers ViT tokens as nodes and describes their spatial interaction, while the latter regards network channels as nodes and reflects the information communication between channels. Using this unified relational graph representation, we found that: 1) model performance was closely related to graph measures; 2) the proposed relational graph representation of ViT has high similarity with real BNNs; and 3) there was a further improvement in model performance when training with a superior model to constrain the aggregation graph. Yuzhong Chen 0002, Zhenxiang Xiao, Lin Zhao 0004, Lu Zhang 0050, Zihao Wu 0001, Dajiang Zhu, Dezhong Yao 0001, Xintao Hu, Tianming Liu 0001, Xi Jiang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 10 |
| 2024 | Task sub-type states decoding via group deep bidirectional recurrent neural network
Shijie Zhao 0001, Long Fang, Yang Yang 0009, Guochang Tang, Guoxin Luo, Junwei Han 0001, Tianming Liu 0001, Xintao Hu |
Medical Image Anal. | 8 |
| 2024 | BI-AVAN: A Brain-Inspired Adversarial Visual Attention Network for Characterizing Human Visual Attention From Neural ActivityabstractVisual attention is a fundamental mechanism in the human brain, and it inspires the design of attention mechanisms in deep neural networks. However, most of the visual attention studies adopted eye-tracking data rather than the direct measurement of brain activity to characterize human visual attention. In addition, the adversarial relationship between the attention-related objects and attention-neglected background in the human visual system was not fully exploited. To bridge these gaps, we propose a novel brain-inspired adversarial visual attention network (BI-AVAN) to characterize human visual attention directly from functional brain activity. Our BI-AVAN model imitates the biased competition process between attention-related/neglected objects to identify and locate the visual objects in a movie frame the human brain focuses on in an unsupervised manner. We use independent eye-tracking data as ground truth for validation and experimental results show that our model achieves robust and promising results when inferring meaningful human visual attention and mapping the relationship between brain activities and visual stimuli. Our BI-AVAN model contributes to the emerging field of leveraging the brain's functional architecture to inspire and guide the model design in artificial intelligence (AI), e.g., deep neural networks. Heng Huang 0003, Lin Zhao 0004, Haixing Dai, Lu Zhang 0050, Xintao Hu, Dajiang Zhu, Tianming Liu 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | Rectify ViT Shortcut Learning by Visual SaliencyabstractShortcut learning in deep learning models occurs when unintended features are prioritized, resulting in degenerated feature representations and reduced generalizability and interpretability. However, shortcut learning in the widely used vision transformer (ViT) framework is largely unknown. Meanwhile, introducing domain-specific knowledge is a major approach to rectifying the shortcuts that are predominated by background-related factors. For example, eye-gaze data from radiologists are effective human visual prior knowledge that has the great potential to guide the deep learning models to focus on meaningful foreground regions. However, obtaining eye-gaze data can still sometimes be time-consuming, labor-intensive, and even impractical. In this work, we propose a novel and effective saliency-guided ViT (SGT) model to rectify shortcut learning in ViT with the absence of eye-gaze data. Specifically, a computational visual saliency model (either pretrained or fine-tuned) is adopted to predict saliency maps for input image samples. Then, the saliency maps are used to filter the most informative image patches. Considering that this filter operation may lead to global information loss, we further introduce a residual connection that calculates the self-attention across all the image patches. The experiment results on natural and medical image datasets show that our SGT framework can effectively learn and leverage human prior knowledge without eye-gaze data and achieves much better performance than baselines. Meanwhile, it successfully rectifies the harmful shortcut learning and significantly improves the interpretability of the ViT model, demonstrating the promise of transferring human prior knowledge derived visual saliency in rectifying shortcut learning. Chong Ma 0004, Lin Zhao 0004, Yuzhong Chen 0002, Lei Guo 0002, Xintao Hu, Dinggang Shen, Xi Jiang 0001, Tianming Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2023 | Coupling Artificial Neurons in BERT and Biological Neurons in the Human BrainabstractLinking computational natural language processing (NLP) models and neural responses to language in the human brain on the one hand facilitates the effort towards disentangling the neural representations underpinning language perception, on the other hand provides neurolinguistics evidence to evaluate and improve NLP models. Mappings of an NLP model’s representations of and the brain activities evoked by linguistic input are typically deployed to reveal this symbiosis. However, two critical problems limit its advancement: 1) The model’s representations (artificial neurons, ANs) rely on layer-level embeddings and thus lack fine-granularity; 2) The brain activities (biological neurons, BNs) are limited to neural recordings of isolated cortical unit (i.e., voxel/region) and thus lack integrations and interactions among brain functions. To address those problems, in this study, we 1) define ANs with fine-granularity in transformer-based NLP models (BERT in this study) and measure their temporal activations to input text sequences; 2) define BNs as functional brain networks (FBNs) extracted from functional magnetic resonance imaging (fMRI) data to capture functional interactions in the brain; 3) couple ANs and BNs by maximizing the synchronization of their temporal activations. Our experimental results demonstrate 1) The activations of ANs and BNs are significantly synchronized; 2) the ANs carry meaningful linguistic/semantic information and anchor to their BN signatures; 3) the anchored BNs are interpretable in a neurolinguistic context. Overall, our study introduces a novel, general, and effective framework to link transformer-based NLP models and neural activities in response to language and may provide novel insights for future studies such as brain-inspired evaluation and development of NLP models. Mengyue Zhou, Gaosheng Shi, Lin Zhao 0004, Zihao Wu 0001, Tianming Liu 0001, Xintao Hu |
AAAI | 9 |
| 2023 | A Novel Contactless Prediction Algorithm of Indoor Thermal Comfort Based on Posture Estimation
Shuchang Chu, Xiaogang Cheng, Xintao Hu, Caoxin Xu |
ICIG (4) | 4 |
| 2023 | A Novel Attention-DeblurGAN-Based Defogging Algorithm
Xintao Hu, Xiaogang Cheng, Zhaobin Wang, Jie Ni, Limin Song |
ICIG (2) | 1 |
| 2023 | Hyperspectral anomaly detection using ensemble and robust collaborative representation
Shaoxi Wang, Xintao Hu, Jialong Sun, Jinzhuo Liu |
Inf. Sci. | 2 |
| 2023 | A generic framework for embedding human brain function with temporally correlated autoencoder
Lin Zhao 0004, Zihao Wu 0001, Haixing Dai, Zhengliang Liu, Xintao Hu, Dajiang Zhu, Tianming Liu 0001 |
Medical Image Anal. | 5 |
| 2018 | Modeling Task fMRI Data Via Deep Convolutional AutoencoderabstractTask-based functional magnetic resonance imaging (tfMRI) has been widely used to study functional brain networks under task performance. Modeling tfMRI data is challenging due to at least two problems: the lack of the ground truth of underlying neural activity and the highly complex intrinsic structure of tfMRI data. To better understand brain networks based on fMRI data, data-driven approaches have been proposed, for instance, independent component analysis (ICA) and sparse dictionary learning (SDL). However, both ICA and SDL only build shallow models, and they are under the strong assumption that original fMRI signal could be linearly decomposed into time series components with their corresponding spatial maps. As growing evidence shows that human brain function is hierarchically organized, new approaches that can infer and model the hierarchical structure of brain networks are widely called for. Recently, deep convolutional neural network (CNN) has drawn much attention, in that deep CNN has proven to be a powerful method for learning high-level and mid-level abstractions from low-level raw data. Inspired by the power of deep CNN, in this paper, we developed a new neural network structure based on CNN, called deep convolutional auto-encoder (DCAE), in order to take the advantages of both data-driven approach and CNN's hierarchical feature abstraction ability for the purpose of learning mid-level and high-level features from complex, large-scale tfMRI time series in an unsupervised manner. The DCAE has been applied and tested on the publicly available human connectome project tfMRI data sets, and promising results are achieved. Heng Huang 0001, Xintao Hu, Yu Zhao 0007, Milad Makkie, Qinglin Dong, Shijie Zhao 0001, Lei Guo 0002, Tianming Liu 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2017 | Multi-way Regression Reveals Backbone of Macaque Structural Brain Connectivity in Longitudinal Datasets
Xiao Li 0024, Lin Zhao 0004, Xintao Hu, Tianming Liu 0001, Lei Guo 0002 |
MICCAI (1) | 4 |
| 2017 | Task fMRI data analysis based on supervised stochastic coordinate coding
Jinglei Lv, Qingyang Li 0001, Wei Zhang 0090, Yu Zhao 0007, Xi Jiang 0001, Lei Guo 0002, Junwei Han 0001, Xintao Hu, Christine Cong Guo, Jieping Ye, Tianming Liu 0001 |
Medical Image Anal. | 9 |
| 2016 | Exploring auditory network composition during free listening to audio excerpts via group-wise sparse representationabstractWith the growing number of audio excerpts through various media and distribution channels, advanced audio analysis approaches have received significant interest in the multimedia field. However, current audio analysis approaches are still far from satisfactory due to the semantic gaps between the low-level acoustic features and high-level semantics perceived by human brain. In order to alleviate the problem, this paper propose a novel computational framework to bridge acoustic features with high-level semantic features derived from functional magnetic resonance imaging (fMRI) signals which record the brain's response during free listening to music/speech excerpts, and to explore the brain auditory network composition of acoustic features for different types of music/speech excerpts. Specifically, we identify meaningful brain networks and corresponding brain activities representing high-level semantic features via a novel group-wise sparse representation of whole brain fMRI signals. Then we associate the brain activities with specific low-level acoustic features and analyze the auditory network composition of acoustic features for different types of music/speech excerpts. Experimental results demonstrate that multiple acoustic features are involved in the brain auditory networks during free listening to music/speech excerpts. Meanwhile, there is considerable variability of auditory network composition of acoustic features for different types of music/speech. Our results provide new insights of how to narrow the semantic gaps in audio content analysis. Shijie Zhao 0001, Junwei Han 0001, Xi Jiang 0001, Xintao Hu, Jinglei Lv, Shu Zhang 0001, Bao Ge, Lei Guo 0002, Tianming Liu 0001 |
ICME | 4 |
| 2016 | Species Preserved and Exclusive Structural Connections Revealed by Sparse CCA
Xiao Li 0024, Lei Du 0001, Xintao Hu, Xi Jiang 0001, Lei Guo 0002, Tianming Liu 0001 |
MICCAI (1) | 4 |
| 2016 | Temporal Concatenated Sparse Coding of Resting State fMRI Data Reveal Network Interaction Changes in mTBI
Jinglei Lv, Armin Iraji, Fangfei Ge, Shijie Zhao 0001, Xintao Hu, Junwei Han 0001, Lei Guo 0002, Zhifeng Kou, Tianming Liu 0001 |
MICCAI (1) | 5 |
| 2016 | A Multi-stage Sparse Coding Framework to Explore the Effects of Prenatal Alcohol Exposure
Shijie Zhao 0001, Junwei Han 0001, Jinglei Lv, Xi Jiang 0001, Xintao Hu, Shu Zhang 0001, Mary Ellen Lynch, Claire Coles, Lei Guo 0002, Xiaoping Hu 0001, Tianming Liu 0001 |
MICCAI (1) | 5 |
| 2016 | Predicting Movie Trailer Viewer's "Like/Dislike" via Learned Shot Editing PatternsabstractNowadays, there are many movie trailers publicly available on social media website such as YouTube, and many thousands of users have independently indicated whether they like or dislike those trailers. Although it is understandable that there are multiple factors that could influence viewers' like or dislike of the trailer, we aim to address a preference question in this work: Can subjective multimedia features be developed to predict the viewer's preference presented by like (by thumbs-up) or dislike (by thumbs-down) during and after watching movie trailers? We designed and implemented a computational framework that is composed of low-level multimedia feature extraction, feature screening and selection, and classification, and applied it to a collection of 725 movie trailers. Experimental results demonstrated that, among dozens of multimedia features, the single low-level multimedia feature of shot length variance is highly predictive of a viewer's “like/dislike” for a large portion of movie trailers. We interpret these findings such that variable shot lengths in a trailer tend to produce a rhythm that is likely to stimulate a viewer's positive preference. This conclusion was also proved by the repeatability experiments results using another 600 trailer videos and it was further interpreted by viewers'eye-tracking data. Shu Zhang 0001, Xi Jiang 0001, Xiang Li 0001, Xintao Hu, Junwei Han 0001, Lei Guo 0002, L. Stephen Miller, Richard Neupert, Tianming Liu 0001 |
IEEE Trans. Affect. Comput. | 6 |
| 2015 | Modeling Task FMRI Data via Supervised Stochastic Coordinate Coding
Jinglei Lv, Wei Zhang 0090, Xi Jiang 0001, Xintao Hu, Junwei Han 0001, Lei Guo 0002, Jieping Ye, Tianming Liu 0001 |
MICCAI (1) | 5 |
| 2015 | Analysis of music/speech via integration of audio content and functional brain response
Junwei Han 0001, Xi Jiang 0001, Xintao Hu, Lei Guo 0002, Jungong Han, Ling Shao 0001, Tianming Liu 0001 |
Inf. Sci. | 4 |
| 2015 | Sparse representation of whole-brain fMRI signals for identification of functional networks
Jinglei Lv, Xi Jiang 0001, Xiang Li 0001, Dajiang Zhu, Hanbo Chen, Shu Zhang 0001, Xintao Hu, Junwei Han 0001, Heng Huang 0001, Jing Zhang 0010, Lei Guo 0002, Tianming Liu 0001 |
Medical Image Anal. | 8 |
| 2015 | Arousal Recognition Using Audio-Visual Features and FMRI-Based Brain ResponseabstractAs the indicator of emotion intensity, arousal is a significant clue for users to find their interested content. Hence, effective techniques for video arousal recognition are highly required. In this paper, we propose a novel framework for recognizing arousal levels by integrating low-level audio-visual features derived from video content and human brain's functional activity in response to videos measured by functional magnetic resonance imaging (fMRI). At first, a set of audio-visual features which have been demonstrated to be correlated with video arousal are extracted. Then, the fMRI-derived features that convey the brain activity of comprehending videos are extracted based on a number of brain regions of interests (ROIs) identified by a universal brain reference system. Finally, these two sets of features are integrated to learn a joint representation by using a multimodal deep Boltzmann machine (DBM). The learned joint representation can be utilized as the feature for training classifiers. Due to the fact that fMRI scanning is expensive and time-consuming, our DBM fusion model has the ability to predict the joint representation of the videos without fMRI scans. The experimental results on a video benchmark demonstrated the effectiveness of our framework and the superiority of integrated features. Junwei Han 0001, Xintao Hu, Lei Guo 0002, Tianming Liu 0001 |
IEEE Trans. Affect. Comput. | 3 |
| 2015 | Background Prior-Based Salient Object Detection via Deep Reconstruction ResidualabstractDetection of salient objects from images is gaining increasing research interest in recent years as it can substantially facilitate a wide range of content-based multimedia applications. Based on the assumption that foreground salient regions are distinctive within a certain context, most conventional approaches rely on a number of hand-designed features and their distinctiveness is measured using local or global contrast. Although these approaches have been shown to be effective in dealing with simple images, their limited capability may cause difficulties when dealing with more complicated images. This paper proposes a novel framework for saliency detection by first modeling the background and then separating salient objects from the background. We develop stacked denoising autoencoders with deep learning architectures to model the background where latent patterns are explored and more powerful representations of data are learned in an unsupervised and bottom-up manner. Afterward, we formulate the separation of salient objects from the background as a problem of measuring reconstruction residuals of deep autoencoders. Comprehensive evaluations of three benchmark datasets and comparisons with nine state-of-the-art algorithms demonstrate the superiority of this paper. Junwei Han 0001, Dingwen Zhang, Xintao Hu, Lei Guo 0002, Jinchang Ren |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2015 | Learning Computational Models of Video Memorability from fMRI Brain ImagingabstractGenerally, various visual media are unequally memorable by the human brain. This paper looks into a new direction of modeling the memorability of video clips and automatically predicting how memorable they are by learning from brain functional magnetic resonance imaging (fMRI). We propose a novel computational framework by integrating the power of low-level audiovisual features and brain activity decoding via fMRI. Initially, a user study experiment is performed to create a ground truth database for measuring video memorability and a set of effective low-level audiovisual features is examined in this database. Then, human subjects' brain fMRI data are obtained when they are watching the video clips. The fMRI-derived features that convey the brain activity of memorizing videos are extracted using a universal brain reference system. Finally, due to the fact that fMRI scanning is expensive and time-consuming, a computational model is learned on our benchmark dataset with the objective of maximizing the correlation between the low-level audiovisual features and the fMRI-derived features using joint subspace learning. The learned model can then automatically predict the memorability of videos without fMRI scans. Evaluations on publically available image and video databases demonstrate the effectiveness of the proposed framework. Junwei Han 0001, Changyuan Chen, Ling Shao 0001, Xintao Hu, Jungong Han, Tianming Liu 0001 |
IEEE Trans. Cybern. | 4 |
| 2015 | Supervised Dictionary Learning for Inferring Concurrent Brain NetworksabstractTask-based fMRI (tfMRI) has been widely used to explore functional brain networks via predefined stimulus paradigm in the fMRI scan. Traditionally, the general linear model (GLM) has been a dominant approach to detect task-evoked networks. However, GLM focuses on task-evoked or event-evoked brain responses and possibly ignores the intrinsic brain functions. In comparison, dictionary learning and sparse coding methods have attracted much attention recently, and these methods have shown the promise of automatically and systematically decomposing fMRI signals into meaningful task-evoked and intrinsic concurrent networks. Nevertheless, two notable limitations of current data-driven dictionary learning method are that the prior knowledge of task paradigm is not sufficiently utilized and that the establishment of correspondences among dictionary atoms in different brains have been challenging. In this paper, we propose a novel supervised dictionary learning and sparse coding method for inferring functional networks from tfMRI data, which takes both of the advantages of model-driven method and data-driven method. The basic idea is to fix the task stimulus curves as predefined model-driven dictionary atoms and only optimize the other portion of data-driven dictionary atoms. Application of this novel methodology on the publicly available human connectome project (HCP) tfMRI datasets has achieved promising results. Shijie Zhao 0001, Junwei Han 0001, Jinglei Lv, Xi Jiang 0001, Xintao Hu, Yu Zhao 0007, Bao Ge, Lei Guo 0002, Tianming Liu 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2014 | Visual attention computation in video of driving environmentabstractWe here study the problem of visual attention computation in video of driving environment via the learning from eye movements. We collect a large-scale database of eye movements from 28 subjects on 30 videos of road scenes, which simulate the driving environment. The analysis on this eye movement database reveals that visual attention in driving environment is directed by high-level cognitive factors such as objects. We then present a new high-level representation called Traffic Object Bank (TOB), which is comprised of many individual road object detectors trained comprehensively in semantic space as well as viewpoint space. TOB provides semantically rich object-level features. Finally, we develop a computational model to predict where drivers look via the mapping from TOB-based representation and to gaze data. Experimental results on our traffic scene video benchmark indicate high accordance with human eye movement and show great promise for further applications. Junwei Han 0001, Liye Sun, Dingwen Zhang, Xintao Hu, Gong Cheng 0003, Lei Guo 0002 |
ICME | 4 |
| 2014 | Decoding Auditory Saliency from FMRI Brain ImagingabstractGiven the growing number of available audio streams through a variety of sources and distribution channels, effective and advanced computational audio analysis has received increasing interest in the multimedia field. However, the effectiveness of current audio analysis strategies might be hampered due to the lack of effective representation of high-level semantics perceived by the human and the lack of effective approaches to bridging the gaps between most low-level acoustic features and high-level semantic features. This semantic gap has become the 'bottleneck' problem in audio analysis. In this paper, we propose a computational framework to decode biologically-plausible auditory saliency using high-level features derived from functional magnetic resonance imaging (fMRI) which monitors the human brain's response under the natural stimulus of audio listening. Specifically, we identify meaningful intrinsic brain networks which are involved in audio listening via effective online dictionary learning and sparse representation of whole-brain fMRI signals, reconstruct auditory saliency features using those identified brain network components, and perform group-wise analysis to identify consistent 'brain decoders' of the saliency features across different excerpts and participants. Experimental results demonstrate that the auditory saliency features are effectively decoded via our methods, which potentially provide opportunities for various applications in the multimedia field. Shijie Zhao 0001, Xi Jiang 0001, Junwei Han 0001, Xintao Hu, Dajiang Zhu, Jinglei Lv, Lei Guo 0002, Tianming Liu 0001 |
ACM Multimedia | 4 |
| 2014 | Clustering and retrieval of video shots based on natural stimulus fMRI
Junwei Han 0001, Xintao Hu, Jungong Han, Tianming Liu 0001 |
Neurocomputing | 3 |
| 2014 | Spatial and temporal visual attention prediction in videos using eye movement data
Junwei Han 0001, Liye Sun, Xintao Hu, Jungong Han, Ling Shao 0001 |
Neurocomputing | 3 |
| 2014 | Video abstraction based on fMRI-driven visual attention model
Junwei Han 0001, Kaiming Li, Ling Shao 0001, Xintao Hu, Lei Guo 0002, Jungong Han, Tianming Liu 0001 |
Inf. Sci. | 4 |
| 2014 | Merging Neuroimaging and Multimedia: Methods, Opportunities, and ChallengesabstractNeuroimaging and brain mapping can provide meaningful guidance to multimedia analyses. and advanced computational multimedia analysis can be used to better understand the functional mechanisms of the human brain. Essentially, brain imaging and brain mapping techniques can serve as a bridge that links the digital representation of multimedia and the perception and comprehension of its content. This paper summarizes methods that integrate brain imaging with multimedia analysis and discusses the opportunities and challenges in this interdisciplinary field. In general, quantitative modeling of brain responses during multimedia comprehension has advanced content-based multimedia studies such as image and video classification and tagging. Multimedia analysis has promoted functional brain mapping by using naturalistic multimedia as stimuli during neuroimaging. Challenges and opportunities in merging neuroimaging and multimedia include the quantification of the brain's responses, the quantification of multimedia, and the mapping between brain responses and computational multimedia features. Tianming Liu 0001, Xintao Hu, Junwei Han 0001, Lei Guo 0002 |
IEEE Trans. Hum. Mach. Syst. | 2 |
| 2013 | Modeling Dynamic Functional Information Flows on Large-Scale Brain Networks
Peili Lv, Lei Guo 0002, Xintao Hu, Xiang Li 0001, Changfeng Jin, Junwei Han 0001, Lingjiang Li, Tianming Liu 0001 |
MICCAI (2) | 3 |
| 2013 | Sparse Representation of Group-Wise FMRI Signals
Jinglei Lv, Xiang Li 0001, Dajiang Zhu, Xi Jiang 0001, Xin Zhang 0151, Xintao Hu, Lei Guo 0002, Tianming Liu 0001 |
MICCAI (3) | 6 |
| 2013 | Group-Wise FMRI Activation Detection on Corresponding Cortical Landmarks
Jinglei Lv, Dajiang Zhu, Xintao Hu, Xin Zhang 0151, Junwei Han 0001, Lei Guo 0002, Tianming Liu 0001 |
MICCAI (2) | 3 |
| 2013 | Predicting cortical ROIs via joint modeling of anatomical and connectional profiles
Dajiang Zhu, Xi Jiang 0001, Bao Ge, Xintao Hu, Junwei Han 0001, Lei Guo 0002, Tianming Liu 0001 |
Medical Image Anal. | 5 |
| 2013 | Representing and Retrieving Video Shots in Human-Centric Brain Imaging SpaceabstractMeaningful representation and effective retrieval of video shots in a large-scale database has been a profound challenge for the image/video processing and computer vision communities. A great deal of effort has been devoted to the extraction of low-level visual features, such as color, shape, texture, and motion for characterizing and retrieving video shots. However, the accuracy of these feature descriptors is still far from satisfaction due to the well-known semantic gap. In order to alleviate the problem, this paper investigates a novel methodology of representing and retrieving video shots using human-centric high-level features derived in brain imaging space (BIS) where brain responses to natural stimulus of video watching can be explored and interpreted. At first, our recently developed dense individualized and common connectivity-based cortical landmarks (DICCCOL) system is employed to locate large-scale functional brain networks and their regions of interests (ROIs) that are involved in the comprehension of video stimulus. Then, functional connectivities between various functional ROI pairs are utilized as BIS features to characterize the brain's comprehension of video semantics. Then an effective feature selection procedure is applied to learn the most relevant features while removing redundancy, which results in the formation of the final BIS features. Afterwards, a mapping from low-level visual features to high-level semantic features in the BIS is built via the Gaussian process regression (GPR) algorithm, and a manifold structure is then inferred, in which video key frames are represented by the mapped feature vectors in the BIS. Finally, the manifold-ranking algorithm concerning the relationship among all data is applied to measure the similarity between key frames of video shots. Experimental results on the TRECVID 2005 dataset demonstrate the superiority of the proposed work in comparison with traditional methods. Junwei Han 0001, Xintao Hu, Dajiang Zhu, Kaiming Li, Xi Jiang 0001, Guangbin Cui, Lei Guo 0002, Tianming Liu 0001 |
IEEE Trans. Image Process. | 3 |
| 2012 | Group-Wise Consistent Fiber Clustering Based on Multimodal Connectional and Functional Profiles
Bao Ge, Lei Guo 0002, Dajiang Zhu, Kaiming Li, Xintao Hu, Junwei Han 0001, Tianming Liu 0001 |
MICCAI (3) | 6 |
| 2012 | Characterization of Task-Free/Task-Performance Brain States
Xin Zhang 0151, Lei Guo 0002, Xiang Li 0001, Dajiang Zhu, Kaiming Li, Zhenqiang Sun, Changfeng Jin, Xintao Hu, Junwei Han 0001, Lingjiang Li, Tianming Liu 0001 |
MICCAI (2) | 8 |
| 2012 | Music/speech classification using high-level features derived from fmri brain imagingabstractWith the availability of large amount of audio tracks through a variety of sources and distribution channels, automatic music/speech classification becomes an indispensable tool in social audio websites and online audio communities. However, the accuracy of current acoustic-based low-level feature classification methods is still rather far from satisfaction. The discrepancy between the limited descriptive power of low-level features and the richness of high-level semantics perceived by the human brain has become the 'bottleneck' problem in audio signal analysis. In this paper, functional magnetic resonance imaging (fMRI) which monitors the human brain's response under the natural stimulus of music/speech listening is used as high-level features in the brain imaging space (BIS). We developed a computational framework to model the relationships between BIS features and low-level features in the training dataset with fMRI scans, predict BIS features of testing dataset without fMRI scans, and use the predicted BIS features for music/speech classification in the application stage. Experimental results demonstrated the significantly improved performance of music/speech classification via predicted BIS features than that via the original low-level features. Xi Jiang 0001, Xintao Hu, Lie Lu, Junwei Han 0001, Lei Guo 0002, Tianming Liu 0001 |
ACM Multimedia | 3 |
| 2012 | Bridging the Semantic Gap via Functional Brain ImagingabstractThe multimedia content analysis community has made significant efforts to bridge the gaps between low-level features and high-level semantics perceived by humans. Recent advances in brain imaging and neuroscience in exploring the human brain's responses during multimedia comprehension demonstrated the possibility of leveraging cognitive neuroscience knowledge to bridge the semantic gaps. This paper presents our initial effort in this direction by using functional magnetic resonance imaging (fMRI). Specifically, task-based fMRI (T-fMRI) was performed to accurately localize the brain regions involved in video comprehension. Then, natural stimulus fMRI (N-fMRI) data were acquired when subjects watched the multimedia clips selected from the TRECVID datasets. The responses in the localized brain regions were measured and used to extract high-level features as the representation of the brain's comprehension of semantics in the videos. A novel computational framework was developed to learn the most relevant low-level feature sets that best correlate the fMRI-derived semantic features based on the training videos with fMRI scans, and then the learned model was applied to larger scale TRECVID video datasets without fMRI scans for category classification. Our experimental results demonstrate: 1) there are meaningful couplings between brain's fMRI-derived responses and video stimuli, suggesting the validity of linking semantics and low-level features via fMRI and 2) the computationally learned low-level features can significantly (p <; 0.01) improve video classification in comparison with original low-level features and extracted low-level features resulted from well-known feature projection algorithms. Xintao Hu, Kaiming Li, Junwei Han 0001, Xian-Sheng Hua 0001, Lei Guo 0002, Tianming Liu 0001 |
IEEE Trans. Multim. | 1 |
| 2011 | Retrieving video shots in semantic brain imaging space using manifold-rankingabstractIn recent two decades, a large amount of effort has been devoted to content-based video retrieval (CBVR), which aims to manage large-scale video databases in an effective way based on visual features such as color, shape, texture, and motion. However, the performance of CBVR systems is still far from satisfaction due to the well-known semantic gap. In order to alleviate the problem, this paper proposes a novel retrieval methodology using semantic features derived from brain imaging space (BIS) that reflects brain responses and interactions under natural stimulus of video watching. A mapping from visual features to semantic features in BIS is built through Gaussian process regression. A manifold structure is then inferred where video key frames are represented by mapped feature vectors in BIS. Finally, the manifold-ranking algorithm concerning the relationship among all data is applied to measure the similarity between key frames. Preliminary experimental results on the TRECVID 2005 dataset demonstrate the superiority of the proposed work in comparison with traditional methods. Junwei Han 0001, Xintao Hu, Kaiming Li, Fan Deng 0001, Lei Guo 0002, Tianming Liu 0001 |
ICIP | 3 |
| 2011 | Assessing Regularity and Variability of Cortical Folding Patterns of Working Memory ROIs
Hanbo Chen, Kaiming Li, Xintao Hu, Lei Guo 0002, Tianming Liu 0001 |
MICCAI (2) | 4 |
| 2011 | Resting State fMRI-Guided Fiber Clustering
Bao Ge, Lei Guo 0002, Jinglei Lv, Xintao Hu, Junwei Han 0001, Tianming Liu 0001 |
MICCAI (2) | 4 |
| 2011 | A biologically inspired computational model for image saliency detectionabstractImage saliency detection provides a powerful tool for predicting where human tends to look at in an image, which has been a long attempt for the computer vision community. In this paper, we propose a biologically-inspired model for computing image saliency. At first, a set of basis functions that accords with visual responses to natural stimuli is learned by using eye-fixation patches from an eye-tracking dataset. Three features are then derived based on the learned basis functions including continuity, clutter contrast, and local contrast. Finally, these three features are combined into the saliency map. The proposed approach is easy to implement and can be used in many image and video content analysis applications. Experiments on a large-scale benchmark dataset and comparisons with a number of the state-of-the-art approaches demonstrate its superiority. Junwei Han 0001, Xintao Hu, Lei Guo 0002, Tianming Liu 0001 |
ACM Multimedia | 3 |
| 2010 | A Dynamic Skull Model for Simulation of Cerebral Cortex Folding
Hanbo Chen, Lei Guo 0002, Jingxin Nie, Xintao Hu, Tianming Liu 0001 |
MICCAI (2) | 5 |
| 2010 | Fiber-Centered Analysis of Brain Connectivities Using DTI and Resting State FMRI Data
Jinglei Lv, Lei Guo 0002, Xintao Hu, Kaiming Li, Degang Zhang, Tianming Liu 0001 |
MICCAI (2) | 3 |
| 2010 | Bridging low-level features and high-level semantics via fMRI brain imaging for video classificationabstractThe multimedia content analysis community has made significant effort to bridge the gap between low-level features and high-level semantics perceived by human cognitive systems such as real-world objects and concepts. In the two fields of multimedia analysis and brain imaging, both topics of low-level features and high level semantics are extensively studied. For instance, in the multimedia analysis field, many algorithms are available for multimedia feature extraction, and benchmark datasets are available such as the TRECVID. In the brain imaging field, brain regions that are responsible for vision, auditory perception, language, and working memory are well studied via functional magnetic resonance imaging (fMRI). This paper presents our initial effort in marrying these two fields in order to bridge the gaps between low-level features and high-level semantics via fMRI brain imaging. Our experimental paradigm is that we performed fMRI brain imaging when university student subjects watched the video clips selected from the TRECVID datasets. At current stage, we focus on the three concepts of sports, weather, and commercial-/advertisement specified in the TRECVID 2005. Meanwhile, the brain regions in vision, auditory, language, and working memory networks are quantitatively localized and mapped via task-based paradigm fMRI, and the fMRI responses in these regions are used to extract features as the representation of the brain's comprehension of semantics. Our computational framework aims to learn the most relevant low-level feature sets that best correlate the fMRI-derived semantics based on the training videos with fMRI scans, and then the learned models are applied to larger scale test datasets without fMRI scans for category classifications. Our result shows that: 1) there are meaningful couplings between brain's fMRI responses and video stimuli, suggesting the validity of linking semantics and low-level features via fMRI; 2) The computationally learned low-level feature sets from fMRI-derived semantic features can significantly improve the classification of video categories in comparison with that based on original low-level features. Xintao Hu, Fan Deng 0001, Kaiming Li, Hanbo Chen, Xi Jiang 0001, Jinglei Lv, Dajiang Zhu, Carlos Faraco, Degang Zhang, Arsham Mesbah, Junwei Han 0001, Xian-Sheng Hua 0001, L. Stephen Miller, Lei Guo 0002, Tianming Liu 0001 |
ACM Multimedia | 1 |
| 2010 | Individualized ROI Optimization via Maximization of Group-wise Consistency of Structural and Functional ProfilesabstractFunctional segregation and integration are fundamental characteristics of the human brain. Studying the connectivity among segregated regions and the dynamics of integrated brain networks has drawn increasing interest. A very controversial, yet fundamental issue in these studies is how to determine the best functional brain regions or ROIs (regions of interests) for individuals. Essentially, the computed connectivity patterns and dynamics of brain networks are very sensitive to the locations, sizes, and shapes of the ROIs. This paper presents a novel methodology to optimize the locations of an individual's ROIs in the working memory system. Our strategy is to formulate the individual ROI optimization as a group variance minimization problem, in which group-wise functional and structural connectivity patterns, and anatomic profiles are defined as optimization constraints. The optimization problem is solved via the simulated annealing approach. Our experimental results show that the optimized ROIs have significantly improved consistency in structural and functional profiles across subjects, and have more reasonable localizations and more consistent morphological and anatomic profiles. Kaiming Li, Lei Guo 0002, Carlos Faraco, Dajiang Zhu, Fan Deng 0001, Xi Jiang 0001, Degang Zhang, Hanbo Chen, Xintao Hu, L. Stephen Miller, Tianming Liu 0001 |
NIPS | 10 |