EDBT 2026 Demo / reviewers in the wild / expert
Tianming Liu 0001
dblp:96/5013-1
· DBLP profile ↗
209ranked-venue papers
12as first author
95since 2021 · last 2026
0000-0002-8132-9048ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 138 · 5 first-author · 60 since 2021Graphics, computer vision, multimedia, augmented reality and games · 102 · 8 first-author · 29 since 2021Artificial intelligence and machine learning · 46 · 1 first-author · 28 since 2021Databases, data management, data science and information retrieval · 7 · 2 since 2021Systems, architecture and hardware · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DCMM-Transformer: Degree-Corrected Mixed-Membership Attention for Medical ImagingabstractMedical images exhibit latent anatomical groupings, such as organs, tissues, and pathological regions, that standard Vision Transformers (ViTs) fail to exploit. While recent work like SBM-Transformer attempts to incorporate such structures through stochastic binary masking, they suffer from non-differentiability, training instability, and the inability to model complex community structure. We present DCMM-Transformer, a novel ViT architecture for medical image analysis that incorporates a Degree-Corrected Mixed-Membership (DCMM) model as an additive bias in self-attention. Unlike prior approaches that rely on multiplicative masking and binary sampling, our method introduces community structure and degree heterogeneity in a fully differentiable and interpretable manner. Comprehensive experiments across diverse medical imaging datasets, including brain, chest, breast, and ocular modalities, demonstrate the superior performance and generalizability of the proposed approach. Furthermore, the learned group structure and structured attention modulation substantially enhance interpretability by yielding attention maps that are anatomically meaningful and semantically coherent. Huimin Cheng, Xiaowei Yu 0001, Shushan Wu, Luyang Fang, Jing Zhang 0010, Tianming Liu 0001, Dajiang Zhu, Wenxuan Zhong, Ping Ma 0001 |
AAAI | 7 |
| 2026 | Beyond Adapter Retrieval: Latent Geometry-Preserving Composition via Sparse Task ProjectionabstractRecent advances in parameter-efficient transfer learning have demonstrated the utility of composing LoRA adapters from libraries of pretrained modules. However, most existing approaches rely on simple retrieval heuristics or uniform averaging, which overlook the latent structure of task relationships in representation space. We propose a new framework for adapter reuse that moves beyond retrieval, formulating adapter composition as a geometry-aware sparse reconstruction problem. Specifically, we represent each task by a latent prototype vector derived from the base model’s encoder and aim to approximate the target task prototype as a sparse linear combination of retrieved reference prototypes, under an L1-regularized optimization objective. The resulting combination weights are then used to blend the corresponding LoRA adapters, yielding a composite adapter tailored to the target task. This formulation not only preserves the local geometric structure of the task representation manifold, but also promotes interpretability and efficient reuse by selecting a minimal set of relevant adapters. We demonstrate the effectiveness of our approach across multiple domains—including medical image segmentation, medical report generation and image synthesis. Our results highlight the benefit of coupling retrieval with latent geometry-aware optimization for improved zero-shot generalization. Pengfei Jin, Peng Shu, Sifan Song, Sekeun Kim, Qing Xiao 0003, Cheng Chen 0013, Tianming Liu 0001, Xiang Li 0001, Quanzheng Li |
AAAI | 7 |
| 2026 | ADLGen: Synthesizing Symbolic, Event-Triggered Sensor Sequences for Smart-Home Human Activity ModelingabstractSmart homes equipped with ambient sensors enable privacy conscious monitoring of Activities of Daily Living (ADLs), but produce event-triggered sensor logs that differ fundamentally from regularly sampled time series. These data are discrete, symbolic, irregular, and spatially grounded, exhibiting strong structural dependencies across sensor states and locations. Collecting sufficiently diverse labeled data with adequate structural coverage remains challenging, motivating the need for realistic synthetic data generation. Existing time-series generation methods are designed for continuous or regularly sampled temporal signals and are poorly aligned with smart-home sensor data, where realism fundamentally requires jointly preserving coherent statistical patterns, physical feasibility, and activity-level semantic consistency. We propose ADLGen, a unified framework for synthesizing symbolic, event-triggered sensor sequences. ADLGen integrates a representation tailored to symbolic event data, a generation process that enforces spatial constraints while balancing diversity and coherence, and an LLM-based semantic evaluation and refinement stage for identifying and correcting behavioral inconsistencies, with efficient offline deployment through compact local models. Experiments show that ADLGen synthesizes sequences that closely match real data while improving intrinsic realism, downstream activity recognition, rare-activity learning, and cross-home generalization over strong baselines. Weihang You, Hanqi Jiang, Zishuai Liu, Tianming Liu 0001, Jin Lu 0001, Fei Dou |
SenSys | 5 |
| 2026 | The arts and crafts of android adware across a decade
Chao Wang 0097, Tianming Liu 0001, Yanjie Zhao 0001, Lin Zhang 0062, Xiaoning Du 0001, Li Li 0029, Haoyu Wang 0001 |
Autom. Softw. Eng. | 2 |
| 2026 | GAGM: Geometry-aware graph matching framework for weakly supervised gyral hinge correspondence
Wuyang Li, Tianming Liu 0001, Xiang Li 0001, Junwei Han 0001, Yixuan Yuan |
Medical Image Anal. | 3 |
| 2026 | Alzheimer's disease risk prediction via perceptual deformable attention generative adversarial network with large foundation models
Zhao-Xu Xing, Zhengliang Liu, Da-Fang Zhang 0001, Kun Xie 0001, Jinxiong Fang, Xia-an Bi, Tianming Liu 0001 |
Medical Image Anal. | 7 |
| 2026 | Usable XAI: 10 Strategies Towards Exploiting Explainability in the LLM EraabstractExplainable AI (XAI) refers to techniques that provide human-understandable insights into the workings of AI models. Recently, the focus of XAI has been extended toward explaining Large Language Models (LLMs). This extension calls for a significant transformation in the XAI methodologies for two reasons. First, many existing XAI methods cannot be directly applied to LLMs due to their complexity and advanced capabilities. Second, as LLMs are increasingly deployed in diverse applications, the role of XAI shifts from merely opening the “black box” to actively enhancing the productivity and applicability of LLMs in real-world settings. Meanwhile, the conversation and generation abilities of LLMs can reciprocally enhance XAI. Therefore, in this article, we introduce Usable XAI in the context of LLMs by analyzing (1) how XAI can explain and improve LLM-based AI systems and (2) how XAI techniques can be improved by using LLMs. We introduce 10 strategies, introducing the key techniques for each and discussing their associated challenges. We also provide case studies to demonstrate how to obtain and leverage explanations. Xuansheng Wu, Haiyan Zhao 0003, Yaochen Zhu, Fan Yang 0023, Lijie Hu, Tianming Liu 0001, Xiaoming Zhai, Wenlin Yao, Jundong Li, Mengnan Du, Ninghao Liu 0001 |
ACM Trans. Knowl. Discov. Data | 7 |
| 2026 | Incomplete Multi-Modal Disentanglement Learning With Application to Alzheimer's Disease DiagnosisabstractMulti-modal neuroimaging data, including magnetic resonance imaging (MRI) and fluorodeoxyglucose positron emission tomography (PET), have greatly advanced the computer-aided diagnosis of Alzheimer's disease (AD) by providing shared and complementary information. However, the problem of incomplete multi-modal data remains inevitable and challenging. Conventional strategies that exclude subjects with missing data or synthesize missing scans either result in substantial sample reduction or introduce unwanted noise. To address this issue, we propose an Incomplete Multi-modal Disentanglement Learning method (IMDL) for AD diagnosis without missing scan synthesis, a novel model that employs a tiny Transformer to fuse incomplete multi-modal features extracted by modality-wise variational autoencoders adaptively. Specifically, we first design a cross-modality contrastive learning module to encourage modality-wise variational autoencoders to disentangle shared and complementary representations of each modality. Then, to alleviate the potential information gap between the representations obtained from complete and incomplete multi-modal neuroimages, we leverage the technique of adversarial learning to harmonize these representations with two discriminators. Furthermore, we develop a local attention rectification module comprising local attention alignment and multi-instance attention rectification to enhance the localization of atrophic areas associated with AD. This module aligns inter-modality and intra-modality attention within the Transformer, thus making attention weights more explainable. Extensive experiments conducted on ADNI and AIBL datasets demonstrated the superior performance of the proposed IMDL in AD diagnosis, and a further validation on the HABS-HD dataset highlighted its effectiveness for dementia diagnosis using different multi-modal neuroimaging data (i.e., T1-weighted MRI and diffusion tensor imaging). Kangfu Han, Dan Hu 0004, Fenqiang Zhao, Tianming Liu 0001, Feng Yang 0012, Gang Li 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2026 | FiR-Rad: Fine-Grained Reinforcement With Structured Reasoning for Chest X-Ray Report GenerationabstractAutomated chest X-ray report generation requires not only clinical accuracy but also transparent and interpretable diagnostic reasoning. In this work, we propose FiR-Rad, a two-stage framework that combines explicit structured reasoning with targeted fine-grained optimization. In the first stage, a supervised chain-of-thought approach guides the model to sequentially analyze and describe a comprehensive range of clinically significant thoracic abnormalities, ensuring clinically meaningful coverage. In the second stage, we introduce a segment-level reinforcement learning strategy based on Group Relative Policy Optimization (GRPO), which assigns precise rewards to each disease-specific reasoning step by evaluating the accuracy of corresponding findings in the synthesized report. This design provides direct feedback for intermediate reasoning and encourages consistency between detailed abnormality analysis and final diagnostic conclusions. Experimental results on the MIMIC-CXR and IU-Xray datasets demonstrate that our framework achieves state-of-the-art performance across clinical and linguistic metrics, with strong zero-shot generalization on IU-Xray. The proposed method significantly enhances interpretability and clinical accuracy, effectively addressing key limitations in automated radiology report generation. Xin Mei, Libin Yang, Dehong Gao, Xiaoyan Cai, Junwei Han 0001, Tianming Liu 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2025 | SearchRAG: Can Search Engines Be Helpful for LLM-Based Medical Question Answering?abstractLarge Language Models (LLMs) have shown remarkable capabilities in general domains but often struggle with tasks requiring specialized knowledge. Conventional Retrieval-Augmented Generation (RAG) techniques typically retrieve external information from static knowledge bases, which can be outdated or incomplete, missing fine-grained clinical details essential for accurate medical question answering. In this work, we propose SearchRAG, a novel framework that overcomes these limitations by leveraging real-time search engines. Our method employs synthetic query generation to convert complex medical questions into search-engine-friendly queries and utilizes uncertainty-based knowledge selection to filter and incorporate the most relevant and informative medical knowledge into the LLM's input. Experimental results demonstrate that our method significantly improves response accuracy in medical question answering tasks, particularly for complex questions requiring detailed and up-to-date knowledge. We provide our code here11https://github.com/sycny/SearchRAG. Tianze Yang, Canyu Chen, Quanzheng Li, Tianming Liu 0001, Xiang Li 0001, Ninghao Liu 0001 |
BIBM | 5 |
| 2025 | HELENE: Hessian Layer-wise Clipping and Gradient Annealing for Accelerating Fine-tuning LLM with Zeroth-order OptimizationabstractFine-tuning large language models (LLMs) faces significant memory challenges due to the high cost of back-propagation.MeZO addresses this issue using zeroth-order (ZO) optimization, matching memory usage to inference but suffering from slow convergence due to varying curvatures across model parameters.To overcome this limitation, we propose HELENE, a scalable and memoryefficient optimizer that integrates annealed A-GNB gradients with diagonal Hessian estimation and layer-wise clipping as a second-order pre-conditioner.HELENE provably accelerates and stabilizes convergence by reducing dependence on total parameter space and scaling with the larger layer dimension.Experiments on RoBERTa-large and OPT-1.3Bdemonstrate superior performances, achieving up to 20× speedup over MeZO with an average accuracy improvement of 1.5%.HELENE also supports full and parameter-efficient fine-tuning methods, outperforming several state-of-the-art optimizers. Huaqin Zhao, Jiaxi Li 0002, Yi Pan 0001, Shizhe Liang, Xiaofeng Yang 0005, Fei Dou, Tianming Liu 0001, Jin Lu 0001 |
EMNLP | 7 |
| 2025 | ECHOPulse: ECG Controlled Echocardio-gram Video GenerationabstractEchocardiography (ECHO) is essential for cardiac assessments, but its video quality and interpretation heavily relies on manual expertise, leading to inconsistent results from clinical and portable devices. ECHO video generation offers a solution by improving automated monitoring through synthetic data and generating high-quality videos from routine health data. However, existing models often face high computational costs, slow inference, and rely on complex conditional prompts that require experts' annotations. To address these challenges, we propose ECHOPulse, an ECG-conditioned ECHO video generation model. ECHOPulse introduces two key advancements: (1) it accelerates ECHO video generation by leveraging VQ-VAE tokenization and masked visual token modeling for fast decoding, and (2) it conditions on readily accessible ECG signals, which are highly coherent with ECHO videos, bypassing complex conditional prompts. To the best of our knowledge, this is the first work to use time-series prompts like ECG signals for ECHO video generation. ECHOPulse not only enables controllable synthetic ECHO data generation but also provides updated cardiac function information for disease monitoring and prediction beyond ECG alone. Evaluations on three public and private datasets demonstrate state-of-the-art performance in ECHO video generation across both qualitative and quantitative measures. Additionally, ECHOPulse can be easily generalized to other modality generation tasks, such as cardiac MRI, fMRI, and 3D CT generation. We will make the synthetic ECHO dataset, along with the code and model, publicly available upon acceptance. Yiwei Li 0002, Sekeun Kim, Zihao Wu 0001, Hanqi Jiang, Yi Pan 0001, Pengfei Jin, Sifan Song, Xiaowei Yu 0001, Tianze Yang, Tianming Liu 0001, Quanzheng Li, Xiang Li 0001 |
ICLR | 11 |
| 2025 | HARP: Human-Assisted Regrouping With Permutation Invariant Critic for Multi-Agent Reinforcement LearningabstractHuman-in-the-loop reinforcement learning integrates human expertise to accelerate agent learning and provide critical guidance and feedback in complex fields. However, many existing approaches focus on single-agent tasks and require continuous human involvement during the training process, significantly increasing the human workload and limiting scalability. In this paper, we propose HARP (HumanAssisted Regrouping with Permutation Invariant Critic), a multi-agent reinforcement learning framework designed for group-oriented tasks. HARP integrates automatic agent regrouping with strategic human assistance during deployment, enabling and allowing non-experts to offer effective guidance with minimal intervention. During training, agents dynamically adjust their groupings to optimize collaborative task completion. When deployed, they actively seek human assistance and utilize the Permutation Invariant Group Critic to evaluate and refine human-proposed groupings, allowing non-expert users to contribute valuable suggestions. In multiple collaboration scenarios, our approach is able to leverage limited guidance from non-experts and enhance performance. The project can be found at https://github.com/huawen-hu/HARP. Huawen Hu, Enze Shi, Chenxi Yue, Shuocun Yang, Zihao Wu 0001, Yiwei Li 0002, Tianyang Zhong, Tianming Liu 0001, Shu Zhang 0006 |
ICRA | 9 |
| 2025 | 3D Plant Root Skeleton Detection and ExtractionabstractPlant roots typically exhibit a highly complex and dense architecture, incorporating numerous slender lateral roots and branches, which significantly hinders the precise capture and modeling of the entire root system. Additionally, roots often lack sufficient texture and color information, making it difficult to identify and track root traits using visual methods. Previous research on roots has been largely confined to 2D studies; however, exploring the 3D architecture of roots is crucial in botany. Since roots grow in real 3D space, 3D phenotypic information is more critical for studying genetic traits and their impact on root development. We have introduced a 3D root skeleton extraction method that efficiently derives the 3D architecture of plant roots from a few images. This method includes the detection and matching of lateral roots, triangulation to extract the skeletal structure of lateral roots, and the integration of lateral and primary roots. We developed a highly complex root dataset and tested our method on it. The extracted 3D root skeletons showed considerable similarity to the ground truth, validating the effectiveness of the model. This method can play a significant role in automated breeding robots. Through precise 3D root structure analysis, breeding robots can better identify plant phenotypic traits, especially root structure and growth patterns, helping practitioners select seeds with superior root systems. This automated approach not only improves breeding efficiency but also reduces manual intervention, making the breeding process more intelligent and efficient, thus advancing modern agriculture. Jiakai Lin, Jinchang Zhang, Wen-Zhan Song 0001, Tianming Liu 0001, Guoyu Lu 0001 |
IROS | 5 |
| 2025 | A Unified Continuous Staging Framework for Alzheimer's Disease and Lewy Body Dementia via Hierarchical Anatomical Features
Minheng Chen, Jing Zhang 0010, Xiaowei Yu 0001, Yanjun Lyu, Lu Zhang 0050, Tianming Liu 0001, Dajiang Zhu |
MICCAI (3) | 10 |
| 2025 | Core-Periphery Principle Guided State Space Model for Functional Connectome Classification
Minheng Chen, Xiaowei Yu 0001, Jing Zhang 0010, Yanjun Lyu, Lu Zhang 0050, Tianming Liu 0001, Dajiang Zhu |
MICCAI (12) | 9 |
| 2025 | EndoGen: Conditional Autoregressive Endoscopic Video Generation
Xinyu Liu 0001, Hengyu Liu 0007, Cheng Wang 0043, Tianming Liu 0001, Yixuan Yuan |
MICCAI (10) | 4 |
| 2025 | Oblique Genomics Mixture of Experts: Prediction of Brain Disorder with Aging-Related Changes of Brain's Structural Connectivity Under Genomic Influences
Yanjun Lyu, Jing Zhang 0010, Lu Zhang 0050, Tianming Liu 0001, Dajiang Zhu |
MICCAI (4) | 5 |
| 2025 | SAMed-2: Selective Memory Enhanced Medical Segment Anything Model
Zhiling Yan, Sifan Song, Dingjie Song, Yiwei Li 0002, Rong Zhou 0007, Weixiang Sun, Zhennong Chen, Sekeun Kim, Hui Ren 0001, Tianming Liu 0001, Quanzheng Li, Xiang Li 0001, Lifang He 0001, Lichao Sun 0001 |
MICCAI (13) | 10 |
| 2025 | Domain-Adaptive Diagnosis of Lewy Body Disease with Transferability Aware Transformer
Xiaowei Yu 0001, Jing Zhang 0010, Minheng Chen, Yanjun Lyu, Lu Zhang 0050, Tianming Liu 0001, Dajiang Zhu |
MICCAI (7) | 10 |
| 2025 | Memory Injection Attacks on LLM Agents via Query-Only InteractionabstractAgents powered by large language models (LLMs) have demonstrated strong capabilities in a wide range of complex, real-world applications.
However, LLM agents with a compromised memory bank may easily produce harmful outputs when the past records retrieved for demonstration are malicious.
In this paper, we propose a novel Memory INJection Attack, MINJA, without assuming that the attacker can directly modify the memory bank of the agent.
The attacker injects malicious records into the memory bank by only **interacting with the agent via queries and output observations**.
These malicious records are designed to elicit a sequence of malicious reasoning steps corresponding to a different target query during the agent's execution of the victim user's query.
Specifically, we introduce a sequence of *bridging steps* to link victim queries to the malicious reasoning steps.
During the memory injection, we propose an *indication prompt* that guides the agent to autonomously generate similar bridging steps, with a *progressive shortening strategy* that gradually removes the indication prompt, such that the malicious record will be easily retrieved when processing later victim queries.
Our extensive experiments across diverse agents demonstrate the effectiveness of MINJA in compromising agent memory.
With minimal requirements for execution, MINJA enables any user to influence agent memory, highlighting the risk. Shen Dong, Shaochen Xu, Yige Li, Jiliang Tang, Tianming Liu 0001, Hui Liu 0031, Zhen Xiang |
NeurIPS | 6 |
| 2025 | Bi-VLGM: Bi-Level Class-Severity-Aware Vision-Language Graph Matching for Text Guided Medical Image SegmentationabstractAbstract Medical reports containing specific diagnostic results and additional information not present in medical images can be effectively employed to assist image understanding tasks, and the modality gap between vision and language can be bridged by vision-language matching (VLM). However, current vision-language models distort the intra-model relation and only include class information in reports that is insufficient for segmentation task. In this paper, we introduce a novel Bi-level class-severity-aware Vision-Language Graph Matching (Bi-VLGM) for text guided medical image segmentation, composed of a word-level VLGM module and a sentence-level VLGM module, to exploit the class-severity-aware relation among visual-textual features. In word-level VLGM, to mitigate the distorted intra-modal relation during VLM, we reformulate VLM as graph matching problem and introduce a vision-language graph matching (VLGM) to exploit the high-order relation among visual-textual features. Then, we perform VLGM between the local features for each class region and class-aware prompts to bridge their gap. In sentence-level VLGM, to provide disease severity information for segmentation task, we introduce a severity-aware prompting to quantify the severity level of disease lesion, and perform VLGM between the global features and the severity-aware prompts. By exploiting the relation between the local (global) and class (severity) features, the segmentation model can include the class-aware and severity-aware information to promote segmentation performance. Extensive experiments proved the effectiveness of our method and its superiority to existing methods. The source code will be released. Wenting Chen, Jie Liu 0044, Tianming Liu 0001, Yixuan Yuan |
Int. J. Comput. Vis. | 3 |
| 2025 | Understanding LLMs: A comprehensive overview from training to inference
Tianle Han, Jiaming Tian, Yutong Zhang 0019, Jiaqi Wang 0010, Xiaohui Gao, Tianyang Zhong, Yi Pan 0001, Shaochen Xu, Zihao Wu 0001, Zhengliang Liu, Xin Zhang 0151, Shu Zhang 0001, Xintao Hu, Ning Qiang, Tianming Liu 0001, Bao Ge |
Neurocomputing | 20 |
| 2025 | Contrastive machine learning reveals species -shared and -specific brain functional architecture
Guannan Cao, Songyao Zhang, Weihan Zhang, Yusong Sun, Jingchao Zhou, Tianyang Zhong, Yixuan Yuan, Tao Liu 0044, Tianming Liu 0001, Lei Guo 0002, Yongchun Yu, Xi Jiang 0001, Gang Li 0001, Junwei Han 0001 |
Medical Image Anal. | 10 |
| 2025 | Learning lifespan brain anatomical correspondence via cortical developmental continuity transfer
Lu Zhang 0050, Zhengwang Wu, Xiaowei Yu 0001, Yanjun Lyu, Zihao Wu 0001, Haixing Dai, Lin Zhao 0004, Li Wang 0026, Gang Li 0001, Xianqiao Wang, Tianming Liu 0001, Dajiang Zhu |
Medical Image Anal. | 11 |
| 2025 | Learning better contrastive view from radiologist's gaze
Sheng Wang 0014, Zihao Zhao 0002, Zixu Zhuang, Xi Ouyang, Lichi Zhang, Zheren Li, Chong Ma 0004, Tianming Liu 0001, Dinggang Shen, Qian Wang 0001 |
Pattern Recognit. | 8 |
| 2025 | AugGPT: Leveraging ChatGPT for Text Data AugmentationabstractText data augmentation is an effective strategy for overcoming the challenge of limited sample sizes in many natural language processing (NLP) tasks. This challenge is especially prominent in the few-shot learning (FSL) scenario, where the data in the target domain is generally much scarcer and of lowered quality. A natural and widely used strategy to mitigate such challenges is to perform data augmentation to better capture data invariance and increase the sample size. However, current text data augmentation methods either can’t ensure the correct labeling of the generated data (lacking faithfulness), or can’t ensure sufficient diversity in the generated data (lacking compactness), or both. Inspired by the recent success of large language models (LLM), especially the development of ChatGPT, we propose a text data augmentation approach based on ChatGPT (named ”AugGPT”). AugGPT rephrases each sentence in the training samples into multiple conceptually similar but semantically different samples. The augmented samples can then be used in downstream model training. Experiment results on multiple few-shot learning text classification tasks show the superior performance of the proposed AugGPT approach over state-of-the-art text data augmentation methods in terms of testing accuracy and distribution of the augmented samples. Haixing Dai, Zhengliang Liu, Wenxiong Liao, Zihao Wu 0001, Lin Zhao 0004, Shaochen Xu, Fang Zeng, Wei Liu 0146, Ninghao Liu 0001, Sheng Li 0001, Dajiang Zhu, Hongmin Cai, Lichao Sun 0001, Quanzheng Li, Dinggang Shen, Tianming Liu 0001, Xiang Li 0001 |
IEEE Trans. Big Data | 18 |
| 2025 | Exploring New Frontiers in Agricultural NLP: Investigating the Potential of Large Language Models for Food ApplicationsabstractThis paper explores new frontiers in agricultural natural language processing (NLP) by investigating the effectiveness of food-related text corpora for pretraining transformer-based language models. Specifically, we focus on semantic matching, establishing mappings between food descriptions and nutrition data through fine-tuning AgriBERT with the FoodOn ontology. Our work introduces an expanded comparison with state-of-the-art language models such as GPT-4, Mistral-large, Claude 3 Sonnet, and Gemini 1.0 Ultra. This exploratory investigation, rather than a direct comparison, aims to understand how AgriBERT, a domain-specific, fine-tuned, open-source model, complements the broad knowledge and generative abilities of these advanced LLMs in addressing the unique challenges of the agricultural sector. We also experiment with other applications, such as cuisine prediction from ingredients, expanding our research to include various NLP tasks beyond semantic matching. Overall, this paper underscores the potential of integrating domain-specific models like AgriBERT with advanced LLMs to enhance the performance and applicability of agricultural NLP applications. Saed Rezayi, Zhengliang Liu, Zihao Wu 0001, Chandra Dhakal, Bao Ge, Haixing Dai, Gengchen Mai, Ninghao Liu 0001, Chen Zhen, Tianming Liu 0001, Sheng Li 0001 |
IEEE Trans. Big Data | 10 |
| 2025 | Exploring the Trade-Offs: Unified Large Language Models vs Local Fine-Tuned Models for Highly-Specific Radiology NLI TaskabstractRecently, ChatGPT and GPT-4 have emerged and gained immense global attention due to their unparalleled performance in language processing. Despite demonstrating impressive capability in various open-domain tasks, their adequacy in highly specific fields like radiology remains untested. Radiology presents unique linguistic phenomena distinct from open-domain data due to its specificity and complexity. Assessing the performance of large language models (LLMs) in such specific domains is crucial not only for a thorough evaluation of their overall performance but also for providing valuable insights into future model design directions: whether model design should be generic or domain-specific. To this end, in this study, we evaluate the performance of ChatGPT/GPT-4 on a radiology natural language inference (NLI) task and compare it to other models fine-tuned specifically on task-related data samples. We also conduct a comprehensive investigation on ChatGPT/GPT-4’s reasoning ability by introducing varying levels of inference difficulty. Our results show that 1) ChatGPT and GPT-4 outperform other LLMs in the radiology NLI task and 2) other specifically fine-tuned Bert-based models require significant amounts of data samples to achieve comparable performance to ChatGPT/GPT-4. These findings not only demonstrate the feasibility and promise of constructing a generic model capable of addressing various tasks across different domains, but also highlight several key factors crucial for developing a unified model, particularly in a medical context, paving the way for future artificial general intelligence (AGI) systems. We release our code and data to the research community. Zihao Wu 0001, Lu Zhang 0050, Xiaowei Yu 0001, Zhengliang Liu, Lin Zhao 0004, Yiwei Li 0002, Haixing Dai, Chong Ma 0004, Gang Li 0001, Wei Liu 0146, Quanzheng Li, Dinggang Shen, Xiang Li 0001, Dajiang Zhu, Tianming Liu 0001 |
IEEE Trans. Big Data | 16 |
| 2025 | MediViSTA: Medical Video Segmentation Via Temporal Fusion SAM Adaptation for EchocardiographyabstractDespite achieving impressive results in general-purpose semantic segmentation with strong generalization on natural images, the Segment Anything Model (SAM) has shown less precision and stability in medical image segmentation. In particular, the original SAM architecture is designed for 2D natural images and is therefore not support to handle three-dimensional information, which is particularly important for medical imaging modalities that are often volumetric or video data. In this paper, we introduce MediViSTA, a parameter-efficient fine-tuning method designed to adapt the vision foundation model for medical video, with a specific focus on echocardiography segmentation. To achieve spatial adaptation, we propose a frequency feature fusion technique that injects spatial frequency information from a CNN branch. For temporal adaptation, we integrate temporal adapters within the transformer blocks of the image encoder. Using a fine-tuning strategy, only a small subset of pre-trained parameters is updated, allowing efficient adaptation to echocardiography data. The effectiveness of our method has been comprehensively evaluated on three datasets, comprising two public datasets and one multi-center in-house dataset. Our method consistently outperforms various state-of-the-art approaches without using any prompts. Furthermore, our model exhibits strong generalization capabilities on unseen datasets, surpassing the second-best approach by 2.15% in Dice and 0.09 in temporal consistency. The results demonstrate the potential of MediViSTA to significantly advance echocardiography video segmentation, offering improved accuracy and robustness in cardiac assessment applications. Sekeun Kim, Pengfei Jin, Cheng Chen 0013, Kyung Sang Kim, Zhiliang Lyu, Hui Ren 0001, Zhengliang Liu, Aoxiao Zhong, Tianming Liu 0001, Xiang Li 0001, Quanzheng Li |
IEEE J. Biomed. Health Informatics | 10 |
| 2025 | Voxel-Level Brain States Prediction Using Swin TransformerabstractUnderstanding brain dynamics is important for neuroscience and mental health. Functional magnetic resonance imaging (fMRI) enables the measurement of neural activities through blood-oxygen-level-dependent (BOLD) signals, which represent brain states. In this study, we aim to predict future human resting brain states with fMRI. Due to the 3D voxel-wise spatial organization and temporal dependencies of the fMRI data, we propose a novel architecture which employs a 4D Shifted Window (Swin) Transformer as encoder to efficiently learn spatio-temporal information and a convolutional decoder to enable brain state prediction at the same spatial and temporal resolution as the input fMRI data. We used 100 unrelated subjects from the Human Connectome Project (HCP) for model training and testing. Our novel model has shown high accuracy when predicting 7.2s resting-state brain activities based on the prior 23.04s fMRI time series. The predicted brain states highly resemble BOLD contrast and dynamics. This work shows promising evidence that the spatiotemporal organization of the human brain can be learned by a Swin Transformer model, at high resolution, which provides a potential for reducing the fMRI scan time and the development of brain-computer interfaces in the future. Yifei Sun 0013, Daniel Chahine, Qinghao Wen, Tianming Liu 0001, Xiang Li 0001, Yixuan Yuan, Fernando Calamante, Jinglei Lv |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | EchoFM: Foundation Model for Generalizable Echocardiogram AnalysisabstractEchocardiography is the first-line non-invasive cardiac imaging modality, providing rich spatio-temporal information on cardiac anatomy and physiology. Recently, foundation model trained on extensive and diverse datasets has shown strong performance in various downstream tasks. However, translating foundation models into the medical imaging domain remains challenging due to domain differences between medical and natural images, the lack of diverse patient and disease datasets. In this paper, we introduce EchoFM, a general-purpose vision foundation model for echocardiography trained on a large-scale dataset of over 20 million echocardiographic images from 6,500 patients. To enable effective learning of rich spatio-temporal representations from periodic videos, we propose a novel self-supervised learning framework based on a masked autoencoder with a spatio-temporal consistent masking strategy and periodic-driven contrastive learning. The learned cardiac representations can be readily adapted and fine-tuned for a wide range of downstream tasks, serving as a strong and flexible backbone model. We validate EchoFM through experiments across key downstream tasks in the clinical echocardiography workflow, leveraging public and multi-center internal datasets. EchoFM consistently outperforms SOTA methods, demonstrating superior generalization capabilities and flexibility. The code and checkpoints are available at: https://github.com/SekeunKim/EchoFM.git. Sekeun Kim, Pengfei Jin, Sifan Song, Cheng Chen 0013, Yiwei Li 0002, Hui Ren 0001, Xiang Li 0001, Tianming Liu 0001, Quanzheng Li |
IEEE Trans. Medical Imaging | 8 |
| 2025 | Guest Editorial Special Issue on Advancements in Foundation Models for Medical ImagingabstractPretrained on massive datasets, Foundation Models (FMs) are revolutionizing medical imaging by offering scalable and generalizable solutions to longstanding challenges. This Special Issue on Advancements in Foundation Models for Medical Imaging presents FM-related works that explore the potential of FMs to address data scarcity, domain shifts, and multimodal integration across a wide range of medical imaging tasks, including segmentation, diagnosis, reconstruction, and prognosis. The included papers also examine critical concerns such as interpretability, efficiency, benchmarking, and ethics in the adoption of FMs for medical imaging. Collectively, the articles in this Special Issue mark a significant step toward establishing FMs as a cornerstone of next-generation medical imaging AI. Tianming Liu 0001, Dinggang Shen, Jong Chul Ye, Marleen de Bruijne |
IEEE Trans. Medical Imaging | 1 |
| 2025 | LLM-Guided Decoupled Probabilistic Prompt for Continual Learning in Medical Image DiagnosisabstractDeep learning-based traditional diagnostic models typically exhibit limitations when applied to dynamic clinical environments that require handling the emergence of new diseases. Continual learning (CL) offers a promising solution, aiming to learn new knowledge while preserving previously learned knowledge. Though recent rehearsal-free CL methods employing prompt tuning (PT) have shown promise, they rely on deterministic prompts that struggle to handle diverse fine-grained knowledge. Moreover, existing PT methods utilize randomly initialized prompts that are trained under standard classification constraints, impeding expert knowledge integration and optimal performance acquisition. In this paper, we propose an LLM-guided Decoupled Probabilistic Prompt (LDPP) for Continual Learning in medical image diagnosis. Specifically, we develop an Expert Knowledge Generation (EKG) module that leverages LLM to acquire decoupled expert knowledge and comprehensive category descriptions. Then, we introduce a Decoupled Probabilistic Prompt pool (DePP) to construct a shared decoupled probabilistic prompt pool, which constructs a shared prompt pool with probabilistic prompts derived from the expert knowledge set. These prompts dynamically provide diverse and flexible descriptions for input images. Finally, We design a Steering Prompt Pool (SPP) to enhance intra-class compactness and promote model performance by learning non-shared prompts. With extensive experimental validation, LDPP consistently sets state-of-the-art performance under the challenging class-incremental setting in CL. Code is available at: https://github.com/CUHK-AIM-Group/LDPP. Yiwen Luo, Wuyang Li, Xiang Li 0001, Tianming Liu 0001, Tianye Niu, Yixuan Yuan |
IEEE Trans. Medical Imaging | 5 |
| 2025 | Adaptive Medical Topic Learning for Enhanced Fine-Grained Cross-Modal Alignment in Medical Report GenerationabstractMedical report generation refers to the automatic creation of accurate and coherent diagnostic reports for medical images. This task can alleviate the workload of radiologists, enhance the efficiency of disease diagnosis, and therefore holds significant value and challenges. Considering the feature differences between different modalities, existing methods primarily focus on facilitating medical report generation through cross-modal alignment of images and texts. However, since medical images are very similar to each other, it is difficult to tag obvious objects, making most methods limited to coarse-grained image-text global alignment. In this paper, we propose a medical report generation model based on adaptive topic learning and fine-grained cross-modal alignment, which aligns images and texts from medical topic perspective and token perspective. From the medical topic perspective, a global-local contrastive loss is introduced to adaptively learn efficient medical topic features, and medical topics are utilized to map images and texts to the same semantic space for fine-grained alignment. From the token perspective, a token prediction module is designed to enable the model to focus on important local information by predicting the key tokens contained in the report. Experimental results on the two public datasets (i.e. IU-Xray and MIMIC-CXR) demonstrate that our proposed model outperforms state-of-the-art baselines. Xin Mei, Libin Yang, Dehong Gao, Xiaoyan Cai, Junwei Han 0001, Tianming Liu 0001 |
IEEE Trans. Multim. | 6 |
| 2025 | Community Graph Convolution Neural Network for Alzheimer's Disease Classification and Pathogenetic Factors IdentificationabstractAs a complex neural network system, the brain regions and genes collaborate to effectively store and transmit information. We abstract the collaboration correlations as the brain region gene community network (BG-CN) and present a new deep learning approach, such as the community graph convolutional neural network (Com-GCN), for investigating the transmission of information within and between communities. The results can be used for diagnosing and extracting causal factors for Alzheimer's disease (AD). First, an affinity aggregation model for BG-CN is developed to describe intercommunity and intracommunity information transmission. Second, we design the Com-GCN architecture with intercommunity convolution and intracommunity convolution operations based on the affinity aggregation model. Through sufficient experimental validation on the AD neuroimaging initiative (ADNI) dataset, the design of Com-GCN matches the physiological mechanism better and improves the interpretability and classification performance. Furthermore, Com-GCN can identify lesioned brain regions and disease-causing genes, which may assist precision medicine and drug design in AD and serve as a valuable reference for other neurological disorders. Xia-an Bi, Siyu Jiang, Wenyan Zhou, Zhao-Xu Xing, Luyun Xu, Zhengliang Liu, Tianming Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 9 |
| 2025 | A Unified and Biologically Plausible Relational Graph Representation of Vision TransformersabstractVision transformer (ViT) and its variants have achieved remarkable success in various tasks. The key characteristic of these ViT models is to adopt different aggregation strategies of spatial patch information within the artificial neural networks (ANNs). However, there is still a key lack of unified representation of different ViT architectures for systematic understanding and assessment of model representation performance. Moreover, how those well-performing ViT ANNs are similar to real biological neural networks (BNNs) is largely unexplored. To answer these fundamental questions, we, for the first time, propose a unified and biologically plausible relational graph representation of ViT models. Specifically, the proposed relational graph representation consists of two key subgraphs: an aggregation graph and an affine graph. The former considers ViT tokens as nodes and describes their spatial interaction, while the latter regards network channels as nodes and reflects the information communication between channels. Using this unified relational graph representation, we found that: 1) model performance was closely related to graph measures; 2) the proposed relational graph representation of ViT has high similarity with real BNNs; and 3) there was a further improvement in model performance when training with a superior model to constrain the aggregation graph. Yuzhong Chen 0002, Zhenxiang Xiao, Lin Zhao 0004, Lu Zhang 0050, Zihao Wu 0001, Dajiang Zhu, Dezhong Yao 0001, Xintao Hu, Tianming Liu 0001, Xi Jiang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 11 |
| 2025 | Mask-Guided Vision Transformer for Few-Shot LearningabstractLearning with little data is challenging but often inevitable in various application scenarios where the labeled data are limited and costly. Recently, few-shot learning (FSL) gained increasing attention because of its generalizability of prior knowledge to new tasks that contain only a few samples. However, for data-intensive models such as vision transformer (ViT), current fine-tuning-based FSL approaches are inefficient in knowledge generalization and, thus, degenerate the downstream task performances. In this article, we propose a novel mask-guided ViT (MG-ViT) to achieve an effective and efficient FSL on the ViT model. The key idea is to apply a mask on image patches to screen out the task-irrelevant ones and to guide the ViT focusing on task-relevant and discriminative patches during FSL. Particularly, MG-ViT only introduces an additional mask operation and a residual connection, enabling the inheritance of parameters from pretrained ViT without any other cost. To optimally select representative few-shot samples, we also include an active learning-based sample selection method to further improve the generalizability of MG-ViT-based FSL. We evaluate the proposed MG-ViT on classification, object detection, and segmentation tasks using gradient-weighted class activation mapping (Grad-CAM) to generate masks. The experimental results show that the MG-ViT model significantly improves the performance and efficiency compared with general fine-tuning-based ViT and ResNet models, providing novel insights and a concrete approach toward generalizing data-intensive and large-scale deep learning models for FSL. Yuzhong Chen 0002, Zhenxiang Xiao, Yi Pan 0001, Lin Zhao 0004, Haixing Dai, Zihao Wu 0001, Changhe Li, Changying Li, Dajiang Zhu, Tianming Liu 0001, Xi Jiang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 11 |
| 2024 | RadChat: A Radiology Chatbot Incorporating Clinical Context for Radiological Reports SummarizationabstractRadiological Report Summarization (RRS) involves automated summarization of key impressions derived from identified findings, intending to alleviate the workload and stress experienced by radiologists. Many existing RRS methods predominantly concentrate on summarizing findings, neglecting crucial clinical context, such as the patient’s previous medical examinations. This context, which is a focal point for radiologists, plays a critical role in producing comprehensive and accurate impressions. This paper endeavors to emulate the workflows of radiologists by incorporating the patient’s clinical context alongside current findings. To achieve this, we reconceptualize RRS as a conversational question-answering task, generating temporal radiological conversations. These conversations are subsequently employed to fine-tune a large chat model. The resulting radiology chatbot, RadChat, demonstrates superior performance in RRS task, showcasing the potential of integrating clinical context for more accurate impressions. Experimental results conducted on the MIMIC-CXR dataset validate the superiority of RadChat in comparison to state-of-the-art baselines. Xin Mei, Libin Yang, Dehong Gao, Xiaoyan Cai, Tianming Liu 0001, Junwei Han 0001 |
BIBM | 5 |
| 2024 | Position: TrustLLM: Trustworthiness in Large Language ModelsabstractLarge language models (LLMs) have gained considerable attention for their excellent natural language processing capabilities. Nonetheless, these LLMs present many challenges, particularly in the realm of trustworthiness. This paper introduces TrustLLM, a comprehensive study of trustworthiness in LLMs, including principles for different dimensions of trustworthiness, established benchmark, evaluation, and analysis of trustworthiness for mainstream LLMs, and discussion of open challenges and future directions. Specifically, we first propose a set of principles for trustworthy LLMs that span eight different dimensions. Based on these principles, we further establish a benchmark across six dimensions including truthfulness, safety, fairness, robustness, privacy, and machine ethics. We then present a study evaluating 16 mainstream LLMs in TrustLLM, consisting of over 30 datasets. Our findings firstly show that in general trustworthiness and capability (i.e., functional effectiveness) are positively related. Secondly, our observations reveal that proprietary LLMs generally outperform most open-source counterparts in terms of trustworthiness, raising concerns about the potential risks of widely accessible open-source LLMs. However, a few open-source LLMs come very close to proprietary ones, suggesting that open-source models can achieve high levels of trustworthiness without additional mechanisms like moderator, offering valuable insights for developers in this field. Thirdly, it is important to note that some LLMs may be overly calibrated towards exhibiting trustworthiness, to the extent that they compromise their utility by mistakenly treating benign prompts as harmful and consequently not responding. Besides these observations, we’ve uncovered key insights into the multifaceted trustworthiness in LLMs. We emphasize the importance of ensuring transparency not only in the models themselves but also in the technologies that underpin trustworthiness. We advocate that the establishment of an AI alliance between industry, academia, the open-source community to foster collaboration is imperative to advance the trustworthiness of LLMs. Yue Huang 0001, Lichao Sun 0001, Haoran Wang 0005, Siyuan Wu 0001, Qihui Zhang, Chujie Gao, Wenhan Lyu, Yixuan Zhang 0001, Xiner Li, Hanchi Sun, Zhengliang Liu, Yixin Liu 0002, Yijue Wang, Bertie Vidgen, Bhavya Kailkhura, Caiming Xiong, Chaowei Xiao, Chunyuan Li, Eric P. Xing, Furong Huang, Heng Ji 0001, Hongyi Wang 0001, Huan Zhang 0001, Huaxiu Yao, Manolis Kellis, Marinka Zitnik, Meng Jiang 0001, Mohit Bansal, James Zou 0001, Jian Pei 0001, Jianfeng Gao 0001, Jiawei Han 0001, Jieyu Zhao 0001, Jiliang Tang, Jindong Wang 0001, Joaquin Vanschoren, John C. Mitchell, Kai Shu, Kaidi Xu, Kai-Wei Chang 0001, Lifang He 0001, Lifu Huang, Michael Backes 0001, Neil Zhenqiang Gong, Philip S. Yu, Quanquan Gu, Ran Xu 0001, Rex Ying, Shuiwang Ji, Suman Jana, Tianlong Chen 0001, Tianming Liu 0001, Tianyi Zhou 0001, William Yang Wang, Xiang Li 0001, Xiangliang Zhang 0001, Xiao Wang 0012, Xing Xie 0001, Xuyu Wang, Yan Liu 0002, Yanfang Ye 0001, Yinzhi Cao, Yong Chen 0016, Yue Zhao 0016 |
ICML | 58 |
| 2024 | F2TNet: FMRI to T1w MRI Knowledge Transfer Network for Brain Multi-phenotype Prediction
Wuyang Li, Yu Jiang 0013, Zhihao Peng 0002, Pengyu Wang 0005, Xiang Li 0001, Tianming Liu 0001, Junwei Han 0001, Yixuan Yuan |
MICCAI (11) | 7 |
| 2024 | Epileptic Seizure Detection in SEEG Signals Using a Unified Multi-Scale Temporal-Spatial-Spectral Transformer Model
Zhuoyi Li, Wenjun Li 0001, Ning Zhu 0006, Junwei Han 0001, Tianming Liu 0001 |
MICCAI (11) | 5 |
| 2024 | Conditional Score-Based Diffusion Model for Cortical Thickness Trajectory Prediction
Qing Xiao 0003, Siyeop Yoon, Hui Ren 0001, Matthew Tivnan, Lichao Sun 0001, Quanzheng Li, Tianming Liu 0001, Yu Zhang 0064, Xiang Li 0001 |
MICCAI (2) | 7 |
| 2024 | Brain Cortical Functional Gradients Predict Cortical Folding Patterns via Attention Mesh Convolution
Tianyang Zhong, Changhe Li, Dajiang Zhu, Junwei Han 0001, Tianming Liu 0001 |
MICCAI (7) | 7 |
| 2024 | Gyri vs. Sulci: Core-Periphery Organization in Functional Brain Networks
Xiaowei Yu 0001, Lu Zhang 0050, Yanjun Lyu, Jing Zhang 0010, Tianming Liu 0001, Dajiang Zhu |
MICCAI (12) | 7 |
| 2024 | Eye-gaze Guided Multi-modal Alignment for Medical Representation LearningabstractIn the medical multi-modal frameworks, the alignment of cross-modality features presents a significant challenge. However, existing works have learned features that are implicitly aligned from the data, without considering the explicit relationships in the medical context. This data-reliance may lead to low generalization of the learned alignment relationships. In this work, we propose the Eye-gaze Guided Multi-modal Alignment (EGMA) framework to harness eye-gaze data for better alignment of medical visual and textual features. We explore the natural auxiliary role of radiologists' eye-gaze data in aligning medical images and text, and introduce a novel approach by using eye-gaze data, collected synchronously by radiologists during diagnostic evaluations. We conduct downstream tasks of image classification and image-text retrieval on four medical datasets, where EGMA achieved state-of-the-art performance and stronger generalization across different datasets. Additionally, we explore the impact of varying amounts of eye-gaze data on model performance, highlighting the feasibility and utility of integrating this auxiliary data into multi-modal alignment framework. Chong Ma 0004, Hanqi Jiang, Wenting Chen, Yiwei Li 0002, Zihao Wu 0001, Xiaowei Yu 0001, Zhengliang Liu, Lei Guo 0002, Dajiang Zhu, Dinggang Shen, Tianming Liu 0001, Xiang Li 0001 |
NeurIPS | 12 |
| 2024 | Real-time Core-Periphery Guided ViT with Smart Data Layout Selection on Mobile DevicesabstractMobile devices have become essential enablers for AI applications, particularly in scenarios that require real-time performance. Vision Transformer (ViT) has become a fundamental cornerstone in this regard due to its high accuracy. Recent efforts have been dedicated to developing various transformer architectures that offer im- proved accuracy while reducing the computational requirements. However, existing research primarily focuses on reducing the theoretical computational complexity through methods such as local attention and model pruning, rather than considering realistic performance on mobile hardware. Although these optimizations reduce computational demands, they either introduce additional overheads related to data transformation (e.g., Reshape and Transpose) or irregular computation/data-access patterns. These result in significant overhead on mobile devices due to their limited bandwidth, which even makes the latency worse than vanilla ViT on mobile. In this paper, we present ECP-ViT, a real-time framework that employs the core-periphery principle inspired by the brain functional networks to guide self-attention in ViTs and enable the deployment of ViT models on smartphones. We identify the main bottleneck in transformer structures caused by data transformation and propose a hardware-friendly core-periphery guided self-attention to decrease computation demands. Additionally, we design the system optimizations for intensive data transformation in pruned models. ECP-ViT, with the proposed algorithm-system co-optimizations, achieves a speedup of 4.6× to 26.9× on mobile GPUs across four datasets: STL-10, CIFAR100, TinyImageNet, and ImageNet. Zhihao Shu, Xiaowei Yu 0001, Zihao Wu 0001, Wenqi Jia 0003, Yinchen Shi, Miao Yin, Tianming Liu 0001, Dajiang Zhu, Wei Niu 0002 |
NeurIPS | 7 |
| 2024 | Mask-guided BERT for few-shot text classification
Wenxiong Liao, Zhengliang Liu, Haixing Dai, Zihao Wu 0001, Yiyang Zhang 0003, Yuzhong Chen 0002, Xi Jiang 0001, Dajiang Zhu, Sheng Li 0001, Wei Liu 0146, Tianming Liu 0001, Quanzheng Li, Hongmin Cai, Xiang Li 0001 |
Neurocomputing | 13 |
| 2024 | Zero-shot relation triplet extraction as Next-Sentence Prediction
Wenxiong Liao, Zhengliang Liu, Yiyang Zhang 0003, Ninghao Liu 0001, Tianming Liu 0001, Quanzheng Li, Xiang Li 0001, Hongmin Cai |
Knowl. Based Syst. | 6 |
| 2024 | MA-SAM: Modality-agnostic SAM adaptation for 3D medical image segmentation
Cheng Chen 0013, Juzheng Miao, Dufan Wu, Aoxiao Zhong, Zhiling Yan, Sekeun Kim, Zhengliang Liu, Lichao Sun 0001, Xiang Li 0001, Tianming Liu 0001, Pheng-Ann Heng, Quanzheng Li |
Medical Image Anal. | 11 |
| 2024 | Mask-aware transformer with structure invariant loss for CT translation
Wenting Chen, Wei Zhao 0040, Zhen Chen 0013, Tianming Liu 0001, Li Liu 0017, Jun Liu 0007, Yixuan Yuan |
Medical Image Anal. | 4 |
| 2024 | Task sub-type states decoding via group deep bidirectional recurrent neural network
Shijie Zhao 0001, Long Fang, Yang Yang 0009, Guochang Tang, Guoxin Luo, Junwei Han 0001, Tianming Liu 0001, Xintao Hu |
Medical Image Anal. | 7 |
| 2024 | Fusing multi-scale functional connectivity patterns via Multi-Branch Vision Transformer (MB-ViT) for macaque brain age prediction
Jingchao Zhou, Yuzhong Chen 0002, Xuewei Jin, Zhenxiang Xiao, Songyao Zhang, Tianming Liu 0001, Keith M. Kendrick, Xi Jiang 0001 |
Neural Networks | 8 |
| 2024 | Structure Mapping Generative Adversarial Network for Multi-View Information Mapping Pattern MiningabstractMulti-view learning is dedicated to integrating information from different views and improving the generalization performance of models. However, in most current works, learning under different views has significant independency, overlooking common information mapping patterns that exist between these views. This paper proposes a Structure Mapping Generative adversarial network (SM-GAN) framework, which utilizes the consistency and complementarity of multi-view data from the innovative perspective of information mapping. Specifically, based on network-structured multi-view data, a structural information mapping model is proposed to capture hierarchical interaction patterns among views. Subsequently, three different types of graph convolutional operations are designed in SM-GAN based on the model. Compared with regular GAN, we add a structural information mapping module between the encoder and decoder wthin the generator, completing the structural information mapping from the micro-view to the macro-view. This paper conducted sufficient validation experiments using public imaging genetics data in Alzheimer's Disease Neuroimaging Initiative (ADNI) dataset. It is shown that SM-GAN outperforms baseline and advanced methods in multi-label classification and evolution prediction tasks. Xia-an Bi, YangJun Huang, Zicheng Yang, Zhao-Xu Xing, Luyun Xu, Xiang Li 0001, Zhengliang Liu, Tianming Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 2024 | Brain Structural Connectivity Guided Vision Transformers for Identification of Functional Connectivity Characteristics in Preterm NeonatesabstractPreterm birth is the leading cause of death in children under five years old, and is associated with a wide sequence of complications in both short and long term. In view of rapid neurodevelopment during the neonatal period, preterm neonates may exhibit considerable functional alterations compared to term ones. However, the identified functional alterations in previous studies merely achieve moderate classification performance, while more accurate functional characteristics with satisfying discrimination ability for better diagnosis and therapeutic treatment is underexplored. To address this problem, we propose a novel brain structural connectivity (SC) guided Vision Transformer (SCG-ViT) to identify functional connectivity (FC) differences among three neonatal groups: preterm, preterm with early postnatal experience, and term. Particularly, inspired by the neuroscience-derived information, a novel patch token of SC/FC matrix is defined, and the SC matrix is then adopted as an effective mask into the ViT model to screen out input FC patch embeddings with weaker SC, and to focus on stronger ones for better classification and identification of FC differences among the three groups. The experimental results on multi-modal MRI data of 437 neonatal brains from publicly released Developing Human Connectome Project (dHCP) demonstrate that SCG-ViT achieves superior classification ability compared to baseline models, and successfully identifies holistically different FC patterns among the three groups. Moreover, these different FCs are significantly correlated with the differential gene expressions of the three groups. In summary, SCG-ViT provides a powerfully brain-guided pipeline of adopting large-scale and data-intensive deep learning models for medical imaging-based diagnosis. Yuzhong Chen 0002, Zhenxiang Xiao, Yusong Sun, Jingchao Zhou, Weitong Guo, Chong Ma 0004, Lin Zhao 0004, Keith M. Kendrick, Benjamin Becker, Tianming Liu 0001, Xi Jiang 0001 |
IEEE J. Biomed. Health Informatics | 15 |
| 2024 | Guest Editorial Computational Mathematics Modeling in Cancer AnalysisabstractCancer is a complex disease that can affect any body part. One key feature of cancer is the rapid production of abnormal cells that grow beyond their usual borders and can invade adjoining parts of the body and spread/metastasized to other organs. The process of metastasis is the crucial cause of cancer death. Environmental factors are a significant contributor to cancer initiation [1]. Numerous studies investigate various aspects of cancer, including pathogenesis, prevention, diagnosis, and treatment methods, with the goal of improving patient quality of life and increasing survival rates. Despite significant advances in the field, cancer continues to represent a global challenge for prevention and treatment. Wenjian Qin, Tianming Liu 0001, Fa Zhang 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2024 | CE-GAN: Community Evolutionary Generative Adversarial Network for Alzheimer's Disease Risk PredictionabstractIn the studies of neurodegenerative diseases such as Alzheimer's Disease (AD), researchers often focus on the associations among multi-omics pathogeny based on imaging genetics data. However, current studies overlook the communities in brain networks, leading to inaccurate models of disease development. This paper explores the developmental patterns of AD from the perspective of community evolution. We first establish a mathematical model to describe functional degeneration in the brain as the community evolution driven by entropy information propagation. Next, we propose an interpretable Community Evolutionary Generative Adversarial Network (CE-GAN) to predict disease risk. In the generator of CE-GAN, community evolutionary convolutions are designed to capture the evolutionary patterns of AD. The experiments are conducted using functional magnetic resonance imaging (fMRI) data and single nucleotide polymorphism (SNP) data. CE-GAN achieves 91.67% accuracy and 91.83% area under curve (AUC) in AD risk prediction tasks, surpassing advanced methods on the same dataset. In addition, we validated the effectiveness of CE-GAN for pathogeny extraction. The source code of this work is available at https://github.com/fmri123456/CE-GAN. Xia-an Bi, Zicheng Yang, YangJun Huang, Zhao-Xu Xing, Luyun Xu, Zihao Wu 0001, Zhengliang Liu, Xiang Li 0001, Tianming Liu 0001 |
IEEE Trans. Medical Imaging | 9 |
| 2024 | PhraseAug: An Augmented Medical Report Generation Model With PhrasebookabstractMedical report generation is a valuable and challenging task, which automatically generates accurate and fluent diagnostic reports for medical images, reducing workload of radiologists and improving efficiency of disease diagnosis. Fine-grained alignment of medical images and reports facilitates the exploration of close correlations between images and texts, which is crucial for cross-modal generation. However, visual and linguistic biases caused by radiologists' writing styles make cross-modal image-text alignment difficult. To alleviate visual-linguistic bias, this paper discretizes medical reports and introduces an intermediate modality, i.e. phrasebook, consisting of key noun phrases. As discretized representation of medical reports, phrasebook contains both disease-related medical terms, and synonymous phrases representing different writing styles which can identify synonymous sentences, thereby promoting fine-grained alignment between images and reports. In this paper, an augmented two-stage medical report generation model with phrasebook (PhraseAug) is developed, which combines medical images, clinical histories and writing styles to generate diagnostic reports. In the first stage, phrasebook is used to extract semantically relevant important features and predict key phrases contained in the report. In the second stage, medical reports are generated according to the predicted key phrases which contain synonymous phrases, promoting our model to adapt to different writing styles and generating diverse medical reports. Experimental results on two public datasets, IU-Xray and MIMIC-CXR, demonstrate that our proposed PhraseAug outperforms state-of-the-art baselines. Xin Mei, Libin Yang, Denghong Gao, Xiaoyan Cai, Junwei Han 0001, Tianming Liu 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2024 | BI-AVAN: A Brain-Inspired Adversarial Visual Attention Network for Characterizing Human Visual Attention From Neural ActivityabstractVisual attention is a fundamental mechanism in the human brain, and it inspires the design of attention mechanisms in deep neural networks. However, most of the visual attention studies adopted eye-tracking data rather than the direct measurement of brain activity to characterize human visual attention. In addition, the adversarial relationship between the attention-related objects and attention-neglected background in the human visual system was not fully exploited. To bridge these gaps, we propose a novel brain-inspired adversarial visual attention network (BI-AVAN) to characterize human visual attention directly from functional brain activity. Our BI-AVAN model imitates the biased competition process between attention-related/neglected objects to identify and locate the visual objects in a movie frame the human brain focuses on in an unsupervised manner. We use independent eye-tracking data as ground truth for validation and experimental results show that our model achieves robust and promising results when inferring meaningful human visual attention and mapping the relationship between brain activities and visual stimuli. Our BI-AVAN model contributes to the emerging field of leveraging the brain's functional architecture to inspire and guide the model design in artificial intelligence (AI), e.g., deep neural networks. Heng Huang 0003, Lin Zhao 0004, Haixing Dai, Lu Zhang 0050, Xintao Hu, Dajiang Zhu, Tianming Liu 0001 |
IEEE Trans. Multim. | 7 |
| 2024 | Editorial Special Issue on Explainable and Generalizable Deep Learning for Medical ImagingabstractThe rapid advancements in deep learning technologies have profoundly influenced the field of medical image analysis, yet their full integration into clinical radiology practices has not progressed as quickly as expected. A significant hurdle to their widespread adoption among radiologists and clinicians is the prevailing lack of trust and confidence in the outcomes produced by these technologies. This concern primarily stems from concerns regarding the explainability and generalizability of deep learning models within the realm of medical imaging. As part of the responses from the Medical Image Analysis Community to address these critical issues, we organized the IEEE Transactions on Neural Networks and Learning Systems (TNNLS) Special Issue on explainable and generalizable deep learning for medical imaging. This IEEE TNNLS Special Issue calls for original and innovative methodological contributions that aim to address the key challenges on explainability and generalizability of deep learning for medical imaging. This IEEE TNNLS Special Issue emphasizes the research and advanced development of the technical aspects of new image analysis methodologies, and all the developed new methods should also be evaluated or validated on real and large-scale medical imaging data. Tianming Liu 0001, Dajiang Zhu, Fei Wang 0001, Islem Rekik, Xia Ben Hu, Dinggang Shen |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Rectify ViT Shortcut Learning by Visual SaliencyabstractShortcut learning in deep learning models occurs when unintended features are prioritized, resulting in degenerated feature representations and reduced generalizability and interpretability. However, shortcut learning in the widely used vision transformer (ViT) framework is largely unknown. Meanwhile, introducing domain-specific knowledge is a major approach to rectifying the shortcuts that are predominated by background-related factors. For example, eye-gaze data from radiologists are effective human visual prior knowledge that has the great potential to guide the deep learning models to focus on meaningful foreground regions. However, obtaining eye-gaze data can still sometimes be time-consuming, labor-intensive, and even impractical. In this work, we propose a novel and effective saliency-guided ViT (SGT) model to rectify shortcut learning in ViT with the absence of eye-gaze data. Specifically, a computational visual saliency model (either pretrained or fine-tuned) is adopted to predict saliency maps for input image samples. Then, the saliency maps are used to filter the most informative image patches. Considering that this filter operation may lead to global information loss, we further introduce a residual connection that calculates the self-attention across all the image patches. The experiment results on natural and medical image datasets show that our SGT framework can effectively learn and leverage human prior knowledge without eye-gaze data and achieves much better performance than baselines. Meanwhile, it successfully rectifies the harmful shortcut learning and significantly improves the interpretability of the ViT model, demonstrating the promise of transferring human prior knowledge derived visual saliency in rectifying shortcut learning. Chong Ma 0004, Lin Zhao 0004, Yuzhong Chen 0002, Lei Guo 0002, Xintao Hu, Dinggang Shen, Xi Jiang 0001, Tianming Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 9 |
| 2023 | Coupling Artificial Neurons in BERT and Biological Neurons in the Human BrainabstractLinking computational natural language processing (NLP) models and neural responses to language in the human brain on the one hand facilitates the effort towards disentangling the neural representations underpinning language perception, on the other hand provides neurolinguistics evidence to evaluate and improve NLP models. Mappings of an NLP model’s representations of and the brain activities evoked by linguistic input are typically deployed to reveal this symbiosis. However, two critical problems limit its advancement: 1) The model’s representations (artificial neurons, ANs) rely on layer-level embeddings and thus lack fine-granularity; 2) The brain activities (biological neurons, BNs) are limited to neural recordings of isolated cortical unit (i.e., voxel/region) and thus lack integrations and interactions among brain functions. To address those problems, in this study, we 1) define ANs with fine-granularity in transformer-based NLP models (BERT in this study) and measure their temporal activations to input text sequences; 2) define BNs as functional brain networks (FBNs) extracted from functional magnetic resonance imaging (fMRI) data to capture functional interactions in the brain; 3) couple ANs and BNs by maximizing the synchronization of their temporal activations. Our experimental results demonstrate 1) The activations of ANs and BNs are significantly synchronized; 2) the ANs carry meaningful linguistic/semantic information and anchor to their BN signatures; 3) the anchored BNs are interpretable in a neurolinguistic context. Overall, our study introduces a novel, general, and effective framework to link transformer-based NLP models and neural activities in response to language and may provide novel insights for future studies such as brain-inspired evaluation and development of NLP models. Mengyue Zhou, Gaosheng Shi, Lin Zhao 0004, Zihao Wu 0001, Tianming Liu 0001, Xintao Hu |
AAAI | 8 |
| 2023 | Individual Functional Network Abnormalities Mapping via Graph Representation-Based Neural Architecture Search
Qing Li 0027, Haixing Dai, Jinglei Lv, Lin Zhao 0004, Zhengliang Liu, Zihao Wu 0001, Xia Wu 0001, Claire Coles, Xiaoping Hu 0001, Tianming Liu 0001, Dajiang Zhu |
ADMA (3) | 10 |
| 2023 | Matching Exemplar as Next Sentence Prediction (MeNSP): Zero-Shot Prompt Learning for Automatic Scoring in Science Education
Xuansheng Wu, Tianming Liu 0001, Ninghao Liu 0001, Xiaoming Zhai |
AIED | 3 |
| 2023 | Coarse-to-fine Knowledge Graph Domain Adaptation based on Distantly-supervised Iterative TrainingabstractThe knowledge graph (KG) is a highly needed basis to support the high-fidelity and high-interpretability modeling of various tasks in healthcare artificial intelligence. In this work, we focus on constructing an oncology knowledge graph that will be used in downstream cancer research and solution development. Modern supervised learning for knowledge graph construction requires a large amount of manually labeled data, which makes the process time-consuming and labor-intensive. Although there exists multiple research on named entity recognition and relation extraction based on distantly supervised learning, constructing a domain-specific knowledge graph from large collections of textual data without manual annotations is still an urgent problem to be solved. In response, we propose an integrated framework for adapting and re-learning knowledge graphs from a general domain (biomedical in our case) to a fine-defined domain (oncology). In this framework, we apply distant-supervision on cross-domain knowledge graph adaptation. Consequently, no manual data annotation is required to train the model. We introduce a novel iterative training strategy to facilitate the discovery of domain-specific named entities and triplets. Experimental results indicate that the proposed framework can perform domain adaptation and construction of knowledge graphs efficiently. Wenxiong Liao, Zhengliang Liu, Yiyang Zhang 0003, Fei Qi 0007, Siqi Ding, Hui Ren 0001, Zihao Wu 0001, Haixing Dai, Sheng Li 0001, Lingfei Wu 0001, Ninghao Liu 0001, Quanzheng Li, Tianming Liu 0001, Xiang Li 0001, Hongmin Cai |
BIBM | 14 |
| 2023 | FMRI-Guided Time-Symmetric Joint Model for Visual Attention PredictionabstractVisual attention prediction is linked to brain activity, cognition, and behavior. Despite the availability of brain activity features, previous studies have not fully utilized them, resulting in saliency maps predicted by models primarily based on image features that do not accurately reflect visual attention in the human brain. This inspires us to use functional Magnetic Resonance Imaging (fMRI) signals as a "brain observer" to supervise the training of developing models that integrate top-down image attention-dependent cues and supervise information from saliency maps generated from gaze movement patterns under natural stimuli. Hence, this paper presents an FMRI-Guided Time-Symmetric Joint Model to predict saliency maps from movie clips, which captures the dynamic aspects of human brain cognition and attention, enabling the combination of image features with brain features. Furthermore, we generalize the model to the MS-COCO challenge, evaluating its performance on non-movie data. Our model outperforms other brain-feature-free methods in focusing on visual attention regions of humans in both movie and non-movie datasets. Additionally, incorporating brain features improves model performance, indicating their ability to bridge the semantic gap between human cognition and visual images, allowing for more accurate capture of visual attention regions. Yaonai Wei, Chong Ma 0004, Tianyang Zhong, Lei Du 0001, Songyao Zhang, Tianming Liu 0001, Muheng Shang, Junwei Han 0001 |
BIBM | 8 |
| 2023 | Chat2Brain: A Method for Mapping Open-Ended Semantic Queries to Brain Activation MapsabstractOver decades, neuroscience has accumulated a wealth of research results in the text modality that can be used to explore cognitive processes. Meta-analysis is a typical method that successfully establishes a link from text queries to brain activation maps using these research results, but it still relies on an ideal query environment. In practical applications, text queries used for meta-analyses may encounter issues such as semantic redundancy and ambiguity, resulting in an inaccurate mapping to brain images. On the other hand, large language models (LLMs) like ChatGPT have shown great potential in tasks such as context understanding and reasoning, displaying a high degree of consistency with human natural language. Hence, LLMs could improve the connection between text modality and neuroscience, resolving existing challenges of meta-analyses. In this study, we propose a method called Chat2Brain that combines LLMs to basic text-2-image model, known as Text2Brain, to map open-ended semantic queries to brain activation maps in data-scarce and complex query environments. By utilizing the understanding and reasoning capabilities of LLMs, the performance of the mapping model is optimized by transferring text queries to semantic queries. We demonstrate that Chat2Brain can synthesize anatomically plausible neural activation patterns for more complex tasks of text queries. Yaonai Wei, Tianyang Zhong, Songyao Zhang, Xiao Li 0024, Lin Zhao 0004, Zhengliang Liu, Muheng Shang, Tianming Liu 0001, Chong Ma 0004, Lei Du 0001, Junwei Han 0001 |
BIBM | 9 |
| 2023 | Prediction of Cognitive Scores by Joint Use of Movie-Watching fMRI Connectivity and Eye Tracking via Attention-CensNet
Jiaxing Gao, Lin Zhao 0004, Tianyang Zhong, Changhe Li, Yaonai Wei, Shu Zhang 0001, Lei Guo 0002, Tianming Liu 0001, Junwei Han 0001 |
MICCAI (2) | 9 |
| 2023 | Weakly Supervised Cerebellar Cortical Surface Parcellation with Self-Visual Representation Learning
Zhengwang Wu, Fenqiang Zhao, Yue Sun 0001, Dajiang Zhu, Tianming Liu 0001, Valerie Jewells, Weili Lin, Li Wang 0026, Gang Li 0001 |
MICCAI (8) | 7 |
| 2023 | Multimodal Deep Fusion in Hyperbolic Space for Mild Cognitive Impairment Study
Lu Zhang 0050, Saiyang Na, Tianming Liu 0001, Dajiang Zhu, Junzhou Huang |
MICCAI (5) | 3 |
| 2023 | Disentangling Site Effects with Cycle-Consistent Adversarial Autoencoder for Multi-site Cortical Data Harmonization
Fenqiang Zhao, Zhengwang Wu, Dajiang Zhu, Tianming Liu 0001, John H. Gilmore, Weili Lin, Li Wang 0026, Gang Li 0001 |
MICCAI (8) | 4 |
| 2023 | A Small-Sample Method with EEG Signals Based on Abductive Learning for Motor Imagery Decoding
Tianyang Zhong, Xiaozheng Wei, Enze Shi, Jiaxing Gao, Chong Ma 0004, Yaonai Wei, Songyao Zhang, Lei Guo 0002, Junwei Han 0001, Tianming Liu 0001 |
MICCAI (1) | 10 |
| 2023 | An automatic classifier for monitoring applied behaviors of cage-free laying hens with deep learning
Xiao Yang 0027, Ramesh Bahadur Bist, Sachin Subedi, Zihao Wu 0001, Tianming Liu 0001, Lilong Chai |
Eng. Appl. Artif. Intell. | 5 |
| 2023 | An explainable deep learning framework for characterizing and interpreting human brain states
Shu Zhang 0006, Junxin Wang, Sigang Yu, Ruoyang Wang, Junwei Han 0001, Shijie Zhao 0001, Tianming Liu 0001, Jinglei Lv |
Medical Image Anal. | 7 |
| 2023 | A generic framework for embedding human brain function with temporally correlated autoencoder
Lin Zhao 0004, Zihao Wu 0001, Haixing Dai, Zhengliang Liu, Xintao Hu, Dajiang Zhu, Tianming Liu 0001 |
Medical Image Anal. | 8 |
| 2023 | Characterizing functional brain networks via Spatio-Temporal Attention 4D Convolutional Neural Networks (STA-4DCNNs)
Xi Jiang 0001, Jiadong Yan, Yu Zhao 0007, Mingxin Jiang, Yuzhong Chen 0002, Jingchao Zhou, Zhenxiang Xiao, Benjamin Becker, Dajiang Zhu, Keith M. Kendrick, Tianming Liu 0001 |
Neural Networks | 13 |
| 2023 | Differentiating brain states via multi-clip random fragment strategy-based interactive bidirectional recurrent neural network
Shu Zhang 0001, Enze Shi, Ruoyang Wang, Sigang Yu, Zhengliang Liu, Shaochen Xu, Tianming Liu 0001, Shijie Zhao 0001 |
Neural Networks | 8 |
| 2023 | Eye-Gaze-Guided Vision Transformer for Rectifying Shortcut LearningabstractLearning harmful shortcuts such as spurious correlations and biases prevents deep neural networks from learning meaningful and useful representations, thus jeopardizing the generalizability and interpretability of the learned representation. The situation becomes even more serious in medical image analysis, where the clinical data are limited and scarce while the reliability, generalizability and transparency of the learned model are highly required. To rectify the harmful shortcuts in medical imaging applications, in this paper, we propose a novel eye-gaze-guided vision transformer (EG-ViT) model which infuses the visual attention from radiologists to proactively guide the vision transformer (ViT) model to focus on regions with potential pathology rather than spurious correlations. To do so, the EG-ViT model takes the masked image patches that are within the radiologists' interest as input while has an additional residual connection to the last encoder layer to maintain the interactions of all patches. The experiments on two medical imaging datasets demonstrate that the proposed EG-ViT model can effectively rectify the harmful shortcut learning and improve the interpretability of the model. Meanwhile, infusing the experts' domain knowledge can also improve the large-scale ViT model's performance over all compared baseline methods with limited samples available. In general, EG-ViT takes the advantages of powerful deep neural networks while rectifies the harmful shortcut learning with human expert's prior knowledge. This work also opens new avenues for advancing current artificial intelligence paradigms by infusing human intelligence. Chong Ma 0004, Lin Zhao 0004, Yuzhong Chen 0002, Sheng Wang 0014, Lei Guo 0002, Dinggang Shen, Xi Jiang 0001, Tianming Liu 0001 |
IEEE Trans. Medical Imaging | 9 |
| 2023 | ChestXRayBERT: A Pretrained Language Model for Chest Radiology Report SummarizationabstractAutomatically generating the “impression” section of a radiology report given the “findings” section can summarize as much salient information of the “findings” section as possible, thus promoting more effective communication between radiologists and referring physicians. To significantly reduce the workload of radiologists, we develop and evaluate a novel framework of abstractive summarization methods to automatically generate the “impression” section of chest radiology reports. Despite recent advancements in natural language process (NLP) field such as BERT and its variants, existing abstractive summarization models and methods could not be directly applied to radiology reports, partly due to domain-specific radiology terminology. In response, we develop a pre-trained language model in the chest radiology domain, named ChestXRayBERT, to solve the problem of automatically summarizing chest radiology reports. Specifically, we first collect radiology-related scientific papers as pre-training corpus and pre-train a ChestXRayBERT on it. Then, an abstractive summarization model is proposed, which consists of the pre-trained ChestXRayBERT and a Transformer decoder. Finally, the model is fine-tuned on chest X-ray reports for the abstractive summarization task. When evaluated on the publicly available OPEN-I and MIMIC-CXR datasets, the performance of our proposed model achieves significant improvement compared with other neural networks-based abstractive summarization models. In general, the proposed ChestXRayBERT demonstrates the feasibility and promise of tailoring and extending advanced NLP techniques to the domain of medical imaging and radiology, as well as in the broader biomedicine and healthcare fields in the future. Xiaoyan Cai, Sen Liu 0004, Junwei Han 0001, Libin Yang, Tianming Liu 0001 |
IEEE Trans. Multim. | 6 |
| 2022 | AgriBERT: Knowledge-Infused Agricultural Language Models for Matching Food and NutritionabstractPretraining domain-specific language models remains an important challenge which limits their applicability in various areas such as agriculture. This paper investigates the effectiveness of leveraging food related text corpora (e.g., food and agricultural literature) in pretraining transformer-based language models. We evaluate our trained language model, called AgriBERT, on the task of semantic matching, i.e., establishing mapping between food descriptions and nutrition data, which is a long-standing challenge in the agricultural domain. In particular, we formulate the task as an answer selection problem, fine-tune the trained language model with the help of an external source of knowledge (e.g., FoodOn ontology), and establish a baseline for this task. The experimental results reveal that our language model substantially outperforms other language models and baselines in the task of matching food description and nutrition. Saed Rezayi, Zhengliang Liu, Zihao Wu 0001, Chandra Dhakal, Bao Ge, Chen Zhen, Tianming Liu 0001, Sheng Li 0001 |
IJCAI | 7 |
| 2022 | Hierarchical Brain Networks Decomposition via Prior Knowledge Guided Deep Belief Network
Tianji Pang, Dajiang Zhu, Tianming Liu 0001, Junwei Han 0001, Shijie Zhao 0001 |
MICCAI (1) | 3 |
| 2022 | Longitudinal Infant Functional Connectivity Prediction via Conditional Intensive Triplet Network
Xiaowei Yu 0001, Dan Hu 0004, Lu Zhang 0050, Ying Huang 0007, Zhengwang Wu, Tianming Liu 0001, Li Wang 0026, Weili Lin, Dajiang Zhu, Gang Li 0001 |
MICCAI (8) | 6 |
| 2022 | Embedding Human Brain Function via Transformer
Lin Zhao 0004, Zihao Wu 0001, Haixing Dai, Zhengliang Liu, Dajiang Zhu, Tianming Liu 0001 |
MICCAI (1) | 7 |
| 2022 | COVIDSum: A linguistically enriched SciBERT-based summarization model for COVID-19 scientific papers
Xiaoyan Cai, Sen Liu 0004, Libin Yang, Jintao Zhao, Dinggang Shen, Tianming Liu 0001 |
J. Biomed. Informatics | 7 |
| 2022 | Auto-DenseUNet: Searchable neural network architecture for mass segmentation in 3D automated breast ultrasound
Houjin Chen, Yanfeng Li 0001, Yahui Peng, Yue Zhou 0006, Tianming Liu 0001, Dinggang Shen |
Medical Image Anal. | 7 |
| 2022 | NAS-optimized topology-preserving transfer learning for differentiating cortical folding patterns
Shengfeng Liu, Fangfei Ge, Lin Zhao 0004, Tianfu Wang 0001, Dong Ni 0001, Tianming Liu 0001 |
Medical Image Anal. | 6 |
| 2022 | Gumbel-Softmax based Neural Architecture Search for Hierarchical Brain Networks Decomposition
Tianji Pang, Shijie Zhao 0001, Junwei Han 0001, Shu Zhang 0001, Lei Guo 0002, Tianming Liu 0001 |
Medical Image Anal. | 6 |
| 2022 | Modeling spatio-temporal patterns of holistic functional brain networks via multi-head guided attention graph neural networks (Multi-Head GAGNNs)
Jiadong Yan, Yuzhong Chen 0002, Zhenxiang Xiao, Shu Zhang 0001, Mingxin Jiang, Jinglei Lv, Benjamin Becker, Dajiang Zhu, Junwei Han 0001, Dezhong Yao 0001, Keith M. Kendrick, Tianming Liu 0001, Xi Jiang 0001 |
Medical Image Anal. | 15 |
| 2022 | Follow My Eye: Using Gaze to Supervise Computer-Aided DiagnosisabstractWhen deep neural network (DNN) was first introduced to the medical image analysis community, researchers were impressed by its performance. However, it is evident now that a large number of manually labeled data is often a must to train a properly functioning DNN. This demand for supervision data and labels is a major bottleneck in current medical image analysis, since collecting a large number of annotations from experienced experts can be time-consuming and expensive. In this paper, we demonstrate that the eye movement of radiologists reading medical images can be a new form of supervision to train the DNN-based computer-aided diagnosis (CAD) system. Particularly, we record the tracks of the radiologists' gaze when they are reading images. The gaze information is processed and then used to supervise the DNN's attention via an Attention Consistency module. To the best of our knowledge, the above pipeline is among the earliest efforts to leverage expert eye movement for deep-learning-based CAD. We have conducted extensive experiments on knee X-ray images for osteoarthritis assessment. The results show that our method can achieve considerable improvement in diagnosis performance, with the help of gaze supervision. Sheng Wang 0014, Xi Ouyang, Tianming Liu 0001, Qian Wang 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 3 |
| 2021 | Multi-head GAGNN: A Multi-head Guided Attention Graph Neural Network for Modeling Spatio-temporal Patterns of Holistic Brain Functional Networks
Jiadong Yan, Yuzhong Chen 0002, Shimin Yang, Shu Zhang 0001, Mingxin Jiang, Zhongbo Zhao, Yu Zhao 0007, Benjamin Becker, Tianming Liu 0001, Keith M. Kendrick, Xi Jiang 0001 |
MICCAI (7) | 10 |
| 2021 | Exploring the Functional Difference of Gyri/Sulci via Hierarchical Interpretable Autoencoder
Lin Zhao 0004, Haixing Dai, Xi Jiang 0001, Dajiang Zhu, Tianming Liu 0001 |
MICCAI (7) | 6 |
| 2021 | A Guided Attention 4D Convolutional Neural Network for Modeling Spatio-Temporal Patterns of Functional Brain Networks
Jiadong Yan, Yu Zhao 0007, Mingxin Jiang, Shu Zhang 0001, Shimin Yang, Yuzhong Chen 0002, Zhongbo Zhao, Benjamin Becker, Tianming Liu 0001, Keith M. Kendrick, Xi Jiang 0001 |
PRCV (3) | 11 |
| 2021 | Differentiable neural architecture search for optimal spatial/temporal brain function network decomposition
Qing Li 0027, Xia Wu 0001, Tianming Liu 0001 |
Medical Image Anal. | 3 |
| 2021 | Deep Fusion of Brain Structure-Function in Mild Cognitive Impairment
Lu Zhang 0050, Li Wang 0033, Jean Gao, Shannon L. Risacher, Gang Li 0001, Tianming Liu 0001, Dajiang Zhu |
Medical Image Anal. | 7 |
| 2021 | Eliminating Indefiniteness of Clinical Spectrum for Better Screening COVID-19abstractThe coronavirus disease 2019 (COVID-19) has swept all over the world. Due to the limited detection facilities, especially in developing countries, a large number of suspected cases can only receive common clinical diagnosis rather than more effective detections like Reverse Transcription Polymerase Chain Reaction (RT-PCR) tests or CT scans. This motivates us to develop a quick screening method via common clinical diagnosis results. However, the diagnostic items of different patients may vary greatly, and there is a huge variation in the dimension of the diagnosis data among different suspected patients, it is hard to process these indefinite dimension data via classical classification algorithms. To resolve this problem, we propose an Indefiniteness Elimination Network (IE-Net) to eliminate the influence of the varied dimensions and make predictions about the COVID-19 cases. The IE-Net is in an encoder-decoder framework fashion, and an indefiniteness elimination operation is proposed to transfer the indefinite dimension feature into a fixed dimension feature. Comprehensive experiments were conducted on the public available COVID-19 Clinical Spectrum dataset. Experimental results show that the proposed indefiniteness elimination operation greatly improves the classification performance, the IE-Net achieves 94.80% accuracy, 92.79% recall, 92.97% precision and 94.93% AUC for distinguishing COVID-19 cases from non-COVID-19 cases with only common clinical diagnose data. We further compared our methods with 3 classical classification algorithms: random forest, gradient boosting and multi-layer perceptron (MLP). To explore each clinical test item's specificity, we further analyzed the possible relationship between each clinical test item and COVID-19. Guangyu Guo 0001, Zhuoyan Liu, Shijie Zhao 0001, Lei Guo 0002, Tianming Liu 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2020 | Discovering Functional Brain Networks with 3D Residual Autoencoder (ResAE)
Qinglin Dong, Ning Qiang, Jinglei Lv, Xiang Li 0001, Tianming Liu 0001, Quanzheng Li |
MICCAI (7) | 5 |
| 2020 | Spatiotemporal Attention Autoencoder (STAAE) for ADHD Classification
Qinglin Dong, Ning Qiang, Jinglei Lv, Xiang Li 0001, Tianming Liu 0001, Quanzheng Li |
MICCAI (7) | 5 |
| 2020 | Neural Architecture Search for Optimization of Spatial-Temporal Brain Network Decomposition
Qing Li 0027, Wei Zhang 0090, Jinglei Lv, Xia Wu 0001, Tianming Liu 0001 |
MICCAI (7) | 5 |
| 2020 | Species-Shared and -Specific Structural Connections Revealed by Dirty Multi-task Regression
Xi Jiang 0001, Lei Guo 0002, Xiaoping Hu 0001, Tianming Liu 0001, Lei Du 0001 |
MICCAI (7) | 6 |
| 2020 | Identifying Cross-individual Correspondences of 3-hinge Gyri
Ying Huang 0007, Lin Zhao 0004, Xi Jiang 0001, Lei Guo 0002, Xiaoping Hu 0001, Tianming Liu 0001 |
Medical Image Anal. | 8 |
| 2020 | Deep Neural Networks for In Situ Hybridization Grid Completion and ClusteringabstractTranscriptome in brain plays a crucial role in understanding the cortical organization and the development of brain structure and function. Two challenges, incomplete data and high dimensionality of transcriptome, remain unsolved. Here, we present a novel training scheme that successfully adapts the U-net architecture to the problem of volume recovery. By analogy to denoising autoencoder, we hide a portion of each training sample so that the network can learn to recover missing voxels from context. Then on the completed volumes, we show that Restricted Boltzmann Machines (RBMs) can be used to infer co-occurrences among voxels, providing foundations for dividing the cortex into discrete subregions. As we stack multiple RBMs to form a deep belief network (DBN), we progressively map the high-dimensional raw input into abstract representations and create a hierarchy of transcriptome architecture. A coarse to fine organization emerges from the network layers. This organization incidentally corresponds to the anatomical structures, suggesting a close link between structures and the genetic underpinnings. Thus, we demonstrate a new way of learning transcriptome-based hierarchical organization using RBM and DBN. Yujie Li 0004, Heng Huang 0001, Hanbo Chen, Tianming Liu 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2019 | Identify Hierarchical Structures from Task-Based fMRI Data via Hybrid Spatiotemporal Neural Architecture Search Net
Wei Zhang 0090, Lin Zhao 0004, Qing Li 0027, Shijie Zhao 0001, Qinglin Dong, Xi Jiang 0001, Tianming Liu 0001 |
MICCAI (3) | 8 |
| 2019 | Multi-view Graph Matching of Cortical Landmarks
Ying Huang 0007, Lei Guo 0002, Tianming Liu 0001 |
MICCAI (4) | 4 |
| 2019 | Group-Wise Graph Matching of Cortical Gyral Hinges
Xiao Li 0024, Lin Zhao 0004, Ying Huang 0007, Lei Guo 0002, Tianming Liu 0001 |
MICCAI (4) | 7 |
| 2019 | Fast and scalable distributed deep convolutional autoencoder for fMRI big data analytics
Milad Makkie, Heng Huang 0001, Yu Zhao 0007, Athanasios V. Vasilakos, Tianming Liu 0001 |
Neurocomputing | 5 |
| 2019 | Discovering hierarchical common brain networks via multimodal deep belief network
Shu Zhang 0001, Qinglin Dong, Wei Zhang 0090, Heng Huang 0001, Dajiang Zhu, Tianming Liu 0001 |
Medical Image Anal. | 6 |
| 2019 | A Distributed Computing Platform for fMRI Big Data AnalyticsabstractSince the BRAIN Initiative and Human Brain Project began, a few efforts have been made to address the computational challenges of neuroscience Big Data. The promises of these two projects were to model the complex interaction of brain and behavior and to understand and diagnose brain diseases by collecting and analyzing large quanitites of data. Archiving, analyzing, and sharing the growing neuroimaging datasets posed major challenges. New computational methods and technologies have emerged in the domain of Big Data but have not been fully adapted for use in neuroimaging. In this work, we introduce the current challenges of neuroimaging in a big data context. We review our efforts toward creating a data management system to organize the large-scale fMRI datasets, and present our novel algorithms/methods for the distributed fMRI data processing that employs Hadoop and Spark. Finally, we demonstrate the significant performance gains of our algorithms/methods to perform distributed dictionary learning. Milad Makkie, Xiang Li 0001, Shannon Quinn, Jieping Ye, Geoffrey Mon, Tianming Liu 0001 |
IEEE Trans. Big Data | 7 |
| 2019 | Identifying Brain Networks at Multiple Time Scales via Deep Recurrent Neural NetworkabstractFor decades, task functional magnetic resonance imaging has been a powerful noninvasive tool to explore the organizational architecture of human brain function. Researchers have developed a variety of brain network analysis methods for task fMRI data, including the general linear model, independent component analysis, and sparse representation methods. However, these shallow models are limited in faithful reconstruction and modeling of the hierarchical and temporal structures of brain networks, as demonstrated in more and more studies. Recently, recurrent neural networks (RNNs) exhibit great ability of modeling hierarchical and temporal dependence features in the machine learning field, which might be suitable for task fMRI data modeling. To explore such possible advantages of RNNs for task fMRI data, we propose a novel framework of a deep recurrent neural network (DRNN) to model the functional brain networks from task fMRI data. Experimental results on the motor task fMRI data of Human Connectome Project 900 subjects release demonstrated that the proposed DRNN can not only faithfully reconstruct functional brain networks, but also identify more meaningful brain networks with multiple time scales which are overlooked by traditional shallow models. In general, this work provides an effective and powerful approach to identifying functional brain networks at multiple time scales from task fMRI data. Yan Cui 0005, Shijie Zhao 0001, Han Wang 0012, Yaowu Chen, Junwei Han 0001, Lei Guo 0002, Fan Zhou 0007, Tianming Liu 0001 |
IEEE J. Biomed. Health Informatics | 9 |
| 2019 | Recognizing Brain States Using Deep Sparse Recurrent Neural NetworkabstractBrain activity is a dynamic combination of different sensory responses and thus brain activity/state is continuously changing over time. However, the brain's dynamical functional states recognition at fast time-scales in task fMRI data have been rarely explored. In this paper, we propose a novel 5-layer deep sparse recurrent neural network (DSRNN) model to accurately recognize the brain states across the whole scan session. Specifically, the DSRNN model includes an input layer, one fully-connected layer, two recurrent layers, and a softmax output layer. The proposed framework has been tested on seven task fMRI data sets of Human Connectome Project. Extensive experiment results demonstrate that the proposed DSRNN model can accurately identify the brain's state in different task fMRI data sets and significantly outperforms other auto-correlation methods or non-temporal approaches in the dynamic brain state recognition accuracy. In general, the proposed DSRNN offers a new methodology for basic neuroscience and clinical research. Han Wang 0012, Shijie Zhao 0001, Qinglin Dong, Yan Cui 0005, Yaowu Chen, Junwei Han 0001, Tianming Liu 0001 |
IEEE Trans. Medical Imaging | 8 |
| 2018 | 3D Deep Convolutional Neural Network Revealed the Value of Brain Network Overlap in Differentiating Autism Spectrum Disorder from Healthy Controls
Yu Zhao 0007, Fangfei Ge, Shu Zhang 0001, Tianming Liu 0001 |
MICCAI (3) | 4 |
| 2018 | Identifying Brain Networks of Multiple Time Scales via Deep Recurrent Neural Network
Yan Cui 0005, Shijie Zhao 0001, Han Wang 0012, Yaowu Chen, Junwei Han 0001, Lei Guo 0002, Fan Zhou 0007, Tianming Liu 0001 |
MICCAI (3) | 9 |
| 2018 | Exploring Fiber Skeletons via Joint Representation of Functional Networks and Structural Connectivity
Shu Zhang 0001, Tianming Liu 0001, Dajiang Zhu |
MICCAI (3) | 2 |
| 2018 | Identification of Species-Preserved Cortical Landmarks
Xiao Li 0024, Lin Zhao 0004, Ying Huang 0007, Lei Guo 0002, Tianming Liu 0001 |
MICCAI (3) | 6 |
| 2018 | Towards MR-Only Radiotherapy Treatment Planning: Synthetic CT Generation Using Multi-view Deep Convolutional Neural Networks
Yu Zhao 0007, Shu Liao, Yimo Guo, Liang Zhao 0018, Zhennan Yan, Sungmin Hong, Gerardo Hermosillo, Tianming Liu 0001, Xiang Sean Zhou, Yiqiang Zhan |
MICCAI (1) | 8 |
| 2018 | Modeling 4D fMRI Data via Spatio-Temporal Convolutional Neural Networks (ST-CNN)
Yu Zhao 0007, Xiang Li 0001, Wei Zhang 0090, Shijie Zhao 0001, Milad Makkie, Mo Zhang, Quanzheng Li, Tianming Liu 0001 |
MICCAI (3) | 8 |
| 2018 | Automatic recognition of holistic functional brain networks using iteratively optimized convolutional neural networks (IO-CNN) with weak label initialization
Yu Zhao 0007, Fangfei Ge, Tianming Liu 0001 |
Medical Image Anal. | 3 |
| 2018 | Probabilistic Methods in Computational NeuroscienceabstractProbabilistic models have been successfully adopted in computational biology and bioinformatics. Recently, a number of powerful probabilistic models and methods have been developed in the field of computational neuroscience, and these effective models have significantly advanced this field. This special section aims to capture some snapshots of recent developments of probabilistic methods in the synergistic combinations of cognitive brain science, brain imaging, and neuroscience. It aims to report the latest advances in these fields to the research community working on probabilistic methods in brain imaging analysis and computational neuroscience. This special section includes five contributed articles. Jing Zhang 0010, Tianming Liu 0001, Gopikrishna Deshpande |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2018 | Modeling Task fMRI Data Via Deep Convolutional AutoencoderabstractTask-based functional magnetic resonance imaging (tfMRI) has been widely used to study functional brain networks under task performance. Modeling tfMRI data is challenging due to at least two problems: the lack of the ground truth of underlying neural activity and the highly complex intrinsic structure of tfMRI data. To better understand brain networks based on fMRI data, data-driven approaches have been proposed, for instance, independent component analysis (ICA) and sparse dictionary learning (SDL). However, both ICA and SDL only build shallow models, and they are under the strong assumption that original fMRI signal could be linearly decomposed into time series components with their corresponding spatial maps. As growing evidence shows that human brain function is hierarchically organized, new approaches that can infer and model the hierarchical structure of brain networks are widely called for. Recently, deep convolutional neural network (CNN) has drawn much attention, in that deep CNN has proven to be a powerful method for learning high-level and mid-level abstractions from low-level raw data. Inspired by the power of deep CNN, in this paper, we developed a new neural network structure based on CNN, called deep convolutional auto-encoder (DCAE), in order to take the advantages of both data-driven approach and CNN's hierarchical feature abstraction ability for the purpose of learning mid-level and high-level features from complex, large-scale tfMRI time series in an unsupervised manner. The DCAE has been applied and tested on the publicly available human connectome project tfMRI data sets, and promising results are achieved. Heng Huang 0001, Xintao Hu, Yu Zhao 0007, Milad Makkie, Qinglin Dong, Shijie Zhao 0001, Lei Guo 0002, Tianming Liu 0001 |
IEEE Trans. Medical Imaging | 8 |
| 2018 | Learning to Predict Eye Fixations via Multiresolution Convolutional Neural NetworksabstractEye movements in the case of freely viewing natural scenes are believed to be guided by local contrast, global contrast, and top-down visual factors. Although a lot of previous works have explored these three saliency cues for several years, there still exists much room for improvement on how to model them and integrate them effectively. This paper proposes a novel computation model to predict eye fixations, which adopts a multiresolution convolutional neural network (Mr-CNN) to infer these three types of saliency cues from raw image data simultaneously. The proposed Mr-CNN is trained directly from fixation and nonfixation pixels with multiresolution input image regions with different contexts. It utilizes image pixels as inputs and eye fixation points as labels. Then, both the local and global contrasts are learned by fusing information in multiple contexts. Meanwhile, various top-down factors are learned in higher layers. Finally, optimal combination of top-down factors and bottom-up contrasts can be learned to predict eye fixations. The proposed approach significantly outperforms the state-of-the-art methods on several publically available benchmark databases, demonstrating the superiority of Mr-CNN. We also apply our method to the RGB-D image saliency detection problem. Through learning saliency cues induced by depth and RGB information on pixel level jointly and their interactions, our model achieves better performance on predicting eye fixations in RGB-D images. Nian Liu 0002, Junwei Han 0001, Tianming Liu 0001, Xuelong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2017 | Joint Representation of Connectome-Scale Structural and Functional Profiles for Identification of Consistent Cortical Landmarks in Human Brains
Shu Zhang 0001, Xi Jiang 0001, Tianming Liu 0001 |
MICCAI (1) | 3 |
| 2017 | Multi-way Regression Reveals Backbone of Macaque Structural Brain Connectivity in Longitudinal Datasets
Xiao Li 0024, Lin Zhao 0004, Xintao Hu, Tianming Liu 0001, Lei Guo 0002 |
MICCAI (1) | 5 |
| 2017 | Gyral net: A new representation of cortical folding organization
Hanbo Chen, Yujie Li 0004, Fangfei Ge, Gang Li 0001, Dinggang Shen, Tianming Liu 0001 |
Medical Image Anal. | 6 |
| 2017 | Task fMRI data analysis based on supervised stochastic coordinate coding
Jinglei Lv, Qingyang Li 0001, Wei Zhang 0090, Yu Zhao 0007, Xi Jiang 0001, Lei Guo 0002, Junwei Han 0001, Xintao Hu, Christine Cong Guo, Jieping Ye, Tianming Liu 0001 |
Medical Image Anal. | 12 |
| 2017 | Constructing fine-granularity functional brain network atlases via deep convolutional autoencoder
Yu Zhao 0007, Qinglin Dong, Hanbo Chen, Armin Iraji, Yujie Li 0004, Milad Makkie, Zhifeng Kou, Tianming Liu 0001 |
Medical Image Anal. | 8 |
| 2016 | Distributed rank-1 dictionary learning: Towards fast and scalable solutions for fMRI big data analyticsabstractThe use of functional brain imaging for research and diagnosis has benefitted greatly from the recent advancements in neuroimaging technologies, as well as the explosive growth in size and availability of fMRI data. While it has been shown in literature that using multiple and large scale fMRI datasets can improve reproducibility and lead to new discoveries, the computational and informatics systems supporting the analysis and visualization of such fMRI big data are extremely limited and largely under-discussed. We propose to address these shortcomings in this work, based on previous success in using dictionary learning method for functional network decomposition studies on fMRI data. We presented a distributed dictionary learning framework based on rank-1 matrix decomposition with sparseness constraint (D-r1DL framework). The framework was implemented using the Spark distributed computing engine and deployed on three different processing units: an in-house server, in-house high performance clusters, and the Amazon Elastic Compute Cloud (EC2) service. The whole analysis pipeline was integrated with our neuroinformatics system for data management, user input/output, and real-time visualization. Performance and accuracy of D-r1DL on both individual and group-wise fMRI Human Connectome Project (HCP) dataset shows that the proposed framework is highly scalable. The resulting group-wise functional network decompositions are highly accurate, and the fast processing time confirm this claim. In addition, D-r1DL can provide real-time user feedback and results visualization which are vital for large-scale data analysis. Milad Makkie, Xiang Li 0001, Tianming Liu 0001, Shannon Quinn, Jieping Ye |
IEEE BigData | 3 |
| 2016 | Implementing dictionary learning in Apache Flink, Or: How I learned to relax and love iterationsabstractThe authors evaluate the use of Apache Flink, a novel data analysis framework offering optimizations over competitors such as Apache Spark, in order to use a rank-1 dictionary learning (r1DL) algorithm to decompose fMRI data. We first expand the functionality of the Flink Python API in order to accommodate the implementation of rank-1 dictionary learning, a model for decomposing a large matrix. Iterative algorithms, aggregators, and other features are added to the incomplete Python API, and the experiences and lessons learned are described. Using these features, we port an existing implementation of r1DL from using the Python API of Apache Spark to using the Python API of Apache Flink. In preliminary testing, this implementation suggests performance boosts over Spark for large input files, meriting further research. We conclude that Flink is likely a feasible tool for the application of dictionary learning to decompose fMRI data, and we continue to evaluate and apply it. Geoffrey Mon, Milad Makkie, Xiang Li 0001, Tianming Liu 0001, Shannon Quinn |
IEEE BigData | 4 |
| 2016 | Exploring auditory network composition during free listening to audio excerpts via group-wise sparse representationabstractWith the growing number of audio excerpts through various media and distribution channels, advanced audio analysis approaches have received significant interest in the multimedia field. However, current audio analysis approaches are still far from satisfactory due to the semantic gaps between the low-level acoustic features and high-level semantics perceived by human brain. In order to alleviate the problem, this paper propose a novel computational framework to bridge acoustic features with high-level semantic features derived from functional magnetic resonance imaging (fMRI) signals which record the brain's response during free listening to music/speech excerpts, and to explore the brain auditory network composition of acoustic features for different types of music/speech excerpts. Specifically, we identify meaningful brain networks and corresponding brain activities representing high-level semantic features via a novel group-wise sparse representation of whole brain fMRI signals. Then we associate the brain activities with specific low-level acoustic features and analyze the auditory network composition of acoustic features for different types of music/speech excerpts. Experimental results demonstrate that multiple acoustic features are involved in the brain auditory networks during free listening to music/speech excerpts. Meanwhile, there is considerable variability of auditory network composition of acoustic features for different types of music/speech. Our results provide new insights of how to narrow the semantic gaps in audio content analysis. Shijie Zhao 0001, Junwei Han 0001, Xi Jiang 0001, Xintao Hu, Jinglei Lv, Shu Zhang 0001, Bao Ge, Lei Guo 0002, Tianming Liu 0001 |
ICME | 9 |
| 2016 | Scalable Fast Rank-1 Dictionary Learning for fMRI Big Data AnalysisabstractIt has been shown from various functional neuroimaging studies that sparsity-regularized dictionary learning could achieve superior performance in decomposing comprehensive and neuroscientifically meaningful functional networks from massive fMRI signals. However, the computational cost for solving the dictionary learning problem has been known to be very demanding, especially when dealing with large-scale data sets. Thus in this work, we propose a novel distributed rank-1 dictionary learning (D-r1DL) model and apply it for fMRI big data analysis. The model estimates one rank-1 basis vector with sparsity constraint on its loading coefficient from the input data at each learning step through alternating least squares updates. By iteratively learning the rank-1 basis and deflating the input data at each step, the model is then capable of decomposing the whole set of functional networks. We implement and parallelize the rank-1 dictionary learning algorithm using Spark engine and deployed the resilient distributed dataset (RDDs) abstracts for the data distribution and operations. Experimental results from applying the model on the Human Connectome Project (HCP) data show that the proposed D-r1DL model is efficient and scalable towards fMRI big data analytics, thus enabling data-driven neuroscientific discovery from massive fMRI big data in the future. Xiang Li 0001, Milad Makkie, Mojtaba Sedigh Fazli, Ian Davidson, Jieping Ye, Tianming Liu 0001, Shannon Quinn |
KDD | 7 |
| 2016 | Modeling Functional Dynamics of Cortical Gyri and Sulci
Xi Jiang 0001, Xiang Li 0001, Jinglei Lv, Shijie Zhao 0001, Shu Zhang 0001, Wei Zhang 0090, Tianming Liu 0001 |
MICCAI (1) | 8 |
| 2016 | Discover Mouse Gene Coexpression Landscape Using Dictionary Learning and Sparse Coding
Yujie Li 0004, Hanbo Chen, Xi Jiang 0001, Xiang Li 0001, Jinglei Lv, Hanchuan Peng, Joe Z. Tsien, Tianming Liu 0001 |
MICCAI (1) | 8 |
| 2016 | Species Preserved and Exclusive Structural Connections Revealed by Sparse CCA
Xiao Li 0024, Lei Du 0001, Xintao Hu, Xi Jiang 0001, Lei Guo 0002, Tianming Liu 0001 |
MICCAI (1) | 7 |
| 2016 | Temporal Concatenated Sparse Coding of Resting State fMRI Data Reveal Network Interaction Changes in mTBI
Jinglei Lv, Armin Iraji, Fangfei Ge, Shijie Zhao 0001, Xintao Hu, Junwei Han 0001, Lei Guo 0002, Zhifeng Kou, Tianming Liu 0001 |
MICCAI (1) | 10 |
| 2016 | A Multi-stage Sparse Coding Framework to Explore the Effects of Prenatal Alcohol Exposure
Shijie Zhao 0001, Junwei Han 0001, Jinglei Lv, Xi Jiang 0001, Xintao Hu, Shu Zhang 0001, Mary Ellen Lynch, Claire Coles, Lei Guo 0002, Xiaoping Hu 0001, Tianming Liu 0001 |
MICCAI (1) | 11 |
| 2016 | Exploring Brain Networks via Structured Sparse Representation of fMRI Data
Jianfeng Lu 0003, Jinglei Lv, Xi Jiang 0001, Shijie Zhao 0001, Tianming Liu 0001 |
MICCAI (1) | 6 |
| 2016 | What Makes a Good Movie Trailer?: Interpretation from Simultaneous EEG and Eyetracker RecordingabstractWhat makes a good movie trailer? It's a big challenge to answer this question because of the complexity of multimedia in both low level sensory features and high level semantic features. However, human perception and reactivity could be straightforward evidence for evaluation. Modern Electro-encephalography (EEG) technology provides measurement of consequential brain neural activity to external stimuli. Meanwhile, visual perception and attention could be captured and interpreted by Eye Tracking technology. Intuitively, simultaneous EEG and Eye Tracker recording of human audience with multimedia stimuli could bridge the gap between human comprehension and multimedia analysis, and provide a new way for movie trailer evaluation. In this paper, we propose a novel platform to simultaneously record EEG and eye movement for participants with video stimuli by integrating 256-channel EEG, Eye Tracker and video display device as a system. Based on the proposed system a novel experiment has been designed, in which independent and joint features of EEG and Eye tracking data were mined to evaluate the movie trailer. Our analysis has shown interesting features that are corresponding with trailer quality and video shoot changes. Sidi Liu, Jinglei Lv, Ting Shoemaker, Qinglin Dong, Kaiming Li, Tianming Liu 0001 |
ACM Multimedia | 7 |
| 2016 | Group-wise consistent cortical parcellation based on connectional profiles
Dajiang Zhu, Xi Jiang 0001, Shu Zhang 0001, Zhifeng Kou, Lei Guo 0002, Tianming Liu 0001 |
Medical Image Anal. | 7 |
| 2016 | Predicting Movie Trailer Viewer's "Like/Dislike" via Learned Shot Editing PatternsabstractNowadays, there are many movie trailers publicly available on social media website such as YouTube, and many thousands of users have independently indicated whether they like or dislike those trailers. Although it is understandable that there are multiple factors that could influence viewers' like or dislike of the trailer, we aim to address a preference question in this work: Can subjective multimedia features be developed to predict the viewer's preference presented by like (by thumbs-up) or dislike (by thumbs-down) during and after watching movie trailers? We designed and implemented a computational framework that is composed of low-level multimedia feature extraction, feature screening and selection, and classification, and applied it to a collection of 725 movie trailers. Experimental results demonstrated that, among dozens of multimedia features, the single low-level multimedia feature of shot length variance is highly predictive of a viewer's “like/dislike” for a large portion of movie trailers. We interpret these findings such that variable shot lengths in a trailer tend to produce a rhythm that is likely to stimulate a viewer's positive preference. This conclusion was also proved by the repeatability experiments results using another 600 trailer videos and it was further interpreted by viewers'eye-tracking data. Shu Zhang 0001, Xi Jiang 0001, Xiang Li 0001, Xintao Hu, Junwei Han 0001, Lei Guo 0002, L. Stephen Miller, Richard Neupert, Tianming Liu 0001 |
IEEE Trans. Affect. Comput. | 11 |
| 2016 | Two-Stage Learning to Predict Human Eye Fixations via SDAEsabstractSaliency detection models aiming to quantitatively predict human eye-attended locations in the visual field have been receiving increasing research interest in recent years. Unlike traditional methods that rely on hand-designed features and contrast inference mechanisms, this paper proposes a novel framework to learn saliency detection models from raw image data using deep networks. The proposed framework mainly consists of two learning stages. At the first learning stage, we develop a stacked denoising autoencoder (SDAE) model to learn robust, representative features from raw image data under an unsupervised manner. The second learning stage aims to jointly learn optimal mechanisms to capture the intrinsic mutual patterns as the feature contrast and to integrate them for final saliency prediction. Given the input of pairs of a center patch and its surrounding patches represented by the features learned at the first stage, a SDAE network is trained under the supervision of eye fixation labels, which achieves both contrast inference and contrast integration simultaneously. Experiments on three publically available eye tracking benchmarks and the comparisons with 16 state-of-the-art approaches demonstrate the effectiveness of the proposed framework. Junwei Han 0001, Dingwen Zhang, Shifeng Wen, Lei Guo 0002, Tianming Liu 0001, Xuelong Li 0001 |
IEEE Trans. Cybern. | 5 |
| 2015 | Learning coarse-to-fine sparselets for efficient object detection and scene classificationabstractPart model-based methods have been successfully applied to object detection and scene classification and have achieved state-of-the-art results. More recently the “sparselets” work [1-3] were introduced to serve as a universal set of shared basis learned from a large number of part detectors, resulting in notable speedup. Inspired by this framework, in this paper, we propose a novel scheme to train more effective sparselets with a coarse-to-fine framework. Specifically, we first train coarse sparselets to exploit the redundancy existing among part detectors by using an unsupervised single-hidden-layer auto-encoder. Then, we simultaneously train fine sparselets and activation vectors using a supervised single-hidden-layer neural network, in which sparselets training and discriminative activation vectors learning are jointly embedded into a unified framework. In order to adequately explore the discriminative information hidden in the part detectors and to achieve sparsity, we propose to optimize a new discriminative objective function by imposing L0-norm sparsity constraint on the activation vectors. By using the proposed framework, promising results for multi-class object detection and scene classification are achieved on PASCAL VOC 2007, MIT Scene-67, and UC Merced Land Use datasets, compared with the existing sparselets baseline methods. Gong Cheng 0003, Junwei Han 0001, Lei Guo 0002, Tianming Liu 0001 |
CVPR | 4 |
| 2015 | Predicting eye fixations using convolutional neural networksabstractIt is believed that eye movements in free-viewing of natural scenes are directed by both bottom-up visual saliency and top-down visual factors. In this paper, we propose a novel computational framework to simultaneously learn these two types of visual features from raw image data using a multiresolution convolutional neural network (Mr-CNN) for predicting eye fixations. The Mr-CNN is directly trained from image regions centered on fixation and non-fixation locations over multiple resolutions, using raw image pixels as inputs and eye fixation attributes as labels. Diverse top-down visual features can be learned in higher layers. Meanwhile bottom-up visual saliency can also be inferred via combining information over multiple resolutions. Finally, optimal integration of bottom-up and top-down cues can be learned in the last logistic regression layer to predict eye fixations. The proposed approach achieves state-of-the-art results over four publically available benchmark datasets, demonstrating the superiority of our work. Nian Liu 0002, Junwei Han 0001, Dingwen Zhang, Shifeng Wen, Tianming Liu 0001 |
CVPR | 5 |
| 2015 | Longitudinal Analysis of Brain Recovery after Mild Traumatic Brain Injury Based on Groupwise Consistent Brain Network Clusters
Hanbo Chen, Armin Iraji, Xi Jiang 0001, Jinglei Lv, Zhifeng Kou, Tianming Liu 0001 |
MICCAI (2) | 6 |
| 2015 | Fiber Connection Pattern-Guided Structured Sparse Representation of Whole-Brain fMRI Signals for Functional Network Inference
Xi Jiang 0001, Jianfeng Lu 0003, Lei Guo 0002, Tianming Liu 0001 |
MICCAI (1) | 6 |
| 2015 | Modeling Task FMRI Data via Supervised Stochastic Coordinate Coding
Jinglei Lv, Wei Zhang 0090, Xi Jiang 0001, Xintao Hu, Junwei Han 0001, Lei Guo 0002, Jieping Ye, Tianming Liu 0001 |
MICCAI (1) | 9 |
| 2015 | Distance Networks for Morphological Profiling and Characterization of DICCCOL Landmarks
Hanbo Chen, Jianfeng Lu 0003, Tianming Liu 0001 |
MICCAI (2) | 5 |
| 2015 | Multi-scale and Multimodal Fusion of Tract-Tracing, Myelin Stain and DTI-derived Fibers in Macaque Brains
Ke Jing, Hanbo Chen, Xi Jiang 0001, Longchuan Li, Lei Guo 0002, Jianfeng Lu 0003, Xiaoping Hu 0001, Tianming Liu 0001 |
MICCAI (2) | 10 |
| 2015 | Analysis of music/speech via integration of audio content and functional brain response
Junwei Han 0001, Xi Jiang 0001, Xintao Hu, Lei Guo 0002, Jungong Han, Ling Shao 0001, Tianming Liu 0001 |
Inf. Sci. | 8 |
| 2015 | Sparse representation of whole-brain fMRI signals for identification of functional networks
Jinglei Lv, Xi Jiang 0001, Xiang Li 0001, Dajiang Zhu, Hanbo Chen, Shu Zhang 0001, Xintao Hu, Junwei Han 0001, Heng Huang 0001, Jing Zhang 0010, Lei Guo 0002, Tianming Liu 0001 |
Medical Image Anal. | 13 |
| 2015 | Arousal Recognition Using Audio-Visual Features and FMRI-Based Brain ResponseabstractAs the indicator of emotion intensity, arousal is a significant clue for users to find their interested content. Hence, effective techniques for video arousal recognition are highly required. In this paper, we propose a novel framework for recognizing arousal levels by integrating low-level audio-visual features derived from video content and human brain's functional activity in response to videos measured by functional magnetic resonance imaging (fMRI). At first, a set of audio-visual features which have been demonstrated to be correlated with video arousal are extracted. Then, the fMRI-derived features that convey the brain activity of comprehending videos are extracted based on a number of brain regions of interests (ROIs) identified by a universal brain reference system. Finally, these two sets of features are integrated to learn a joint representation by using a multimodal deep Boltzmann machine (DBM). The learned joint representation can be utilized as the feature for training classifiers. Due to the fact that fMRI scanning is expensive and time-consuming, our DBM fusion model has the ability to predict the joint representation of the videos without fMRI scans. The experimental results on a video benchmark demonstrated the effectiveness of our framework and the superiority of integrated features. Junwei Han 0001, Xintao Hu, Lei Guo 0002, Tianming Liu 0001 |
IEEE Trans. Affect. Comput. | 5 |
| 2015 | Learning Computational Models of Video Memorability from fMRI Brain ImagingabstractGenerally, various visual media are unequally memorable by the human brain. This paper looks into a new direction of modeling the memorability of video clips and automatically predicting how memorable they are by learning from brain functional magnetic resonance imaging (fMRI). We propose a novel computational framework by integrating the power of low-level audiovisual features and brain activity decoding via fMRI. Initially, a user study experiment is performed to create a ground truth database for measuring video memorability and a set of effective low-level audiovisual features is examined in this database. Then, human subjects' brain fMRI data are obtained when they are watching the video clips. The fMRI-derived features that convey the brain activity of memorizing videos are extracted using a universal brain reference system. Finally, due to the fact that fMRI scanning is expensive and time-consuming, a computational model is learned on our benchmark dataset with the objective of maximizing the correlation between the low-level audiovisual features and the fMRI-derived features using joint subspace learning. The learned model can then automatically predict the memorability of videos without fMRI scans. Evaluations on publically available image and video databases demonstrate the effectiveness of the proposed framework. Junwei Han 0001, Changyuan Chen, Ling Shao 0001, Xintao Hu, Jungong Han, Tianming Liu 0001 |
IEEE Trans. Cybern. | 6 |
| 2015 | Morphological Analysis of the Left Ventricular Endocardial Surface Using a Bag-of-Features DescriptorabstractThe limitations of conventional imaging techniques have hitherto precluded a thorough and formal investigation of the complex morphology of the left ventricular (LV) endocardial surface and its relation to the severity of coronary artery disease (CAD). However, recent developments in high-resolution multirow-detector computed tomography (MDCT) scanner technology have enabled the imaging of the complex LV endocardial surface morphology in a single heartbeat. Analysis of high-resolution computed tomography images from a 320-MDCT scanner allows for the noninvasive study of the relationship between the percent diameter stenosis (DS) values of the major coronary arteries and localization of the cardiac segments affected by coronary arterial stenosis. In this paper, a novel approach for the analysis of the nonrigid LV endocardial surface from MDCT images, using a combination of rigid body transformation-invariant shape descriptors and a more generalized isometry-invariant Bag-of-Features descriptor, is proposed and implemented. The proposed approach is shown to be successful in identifying, localizing, and quantifying the incidence and extent of CAD and, thus, is seen to have a potentially significant clinical impact. Specifically, the association between the incidence and extent of CAD, determined via the percent DS measurements of the major coronary arteries, and the alterations in the endocardial surface morphology is formally quantified. The results of the proposed approach on 16 normal datasets and 16 abnormal datasets exhibiting CAD with varying levels of severity are presented. A multivariable regression test is employed to test the effectiveness of the proposed morphological analysis approach. Experiments performed on a strictly leave-one-out basis are shown to exhibit a distinct and interesting pattern in terms of the correlation coefficient values within the cardiac segments, where the incidence of coronary arterial stenosis is localized. Anirban Mukhopadhyay 0003, Suchendra M. Bhandarkar, Tianming Liu 0001, Szilard Voros, Sarah Rinehart |
IEEE J. Biomed. Health Informatics | 4 |
| 2015 | Supervised Dictionary Learning for Inferring Concurrent Brain NetworksabstractTask-based fMRI (tfMRI) has been widely used to explore functional brain networks via predefined stimulus paradigm in the fMRI scan. Traditionally, the general linear model (GLM) has been a dominant approach to detect task-evoked networks. However, GLM focuses on task-evoked or event-evoked brain responses and possibly ignores the intrinsic brain functions. In comparison, dictionary learning and sparse coding methods have attracted much attention recently, and these methods have shown the promise of automatically and systematically decomposing fMRI signals into meaningful task-evoked and intrinsic concurrent networks. Nevertheless, two notable limitations of current data-driven dictionary learning method are that the prior knowledge of task paradigm is not sufficiently utilized and that the establishment of correspondences among dictionary atoms in different brains have been challenging. In this paper, we propose a novel supervised dictionary learning and sparse coding method for inferring functional networks from tfMRI data, which takes both of the advantages of model-driven method and data-driven method. The basic idea is to fix the task stimulus curves as predefined model-driven dictionary atoms and only optimize the other portion of data-driven dictionary atoms. Application of this novel methodology on the publicly available human connectome project (HCP) tfMRI datasets has achieved promising results. Shijie Zhao 0001, Junwei Han 0001, Jinglei Lv, Xi Jiang 0001, Xintao Hu, Yu Zhao 0007, Bao Ge, Lei Guo 0002, Tianming Liu 0001 |
IEEE Trans. Medical Imaging | 9 |
| 2014 | Construct and Assess Multimodal Mouse Brain Connectomes via Joint Modeling of Multi-scale DTI and Neuron Tracer Data
Hanbo Chen, Yu Zhao 0007, Hongmiao Zhang, Hui Kuang, Joe Z. Tsien, Tianming Liu 0001 |
MICCAI (3) | 8 |
| 2014 | Group-Wise Optimization of Common Brain Landmarks with Joint Structural and Functional Regulations
Dajiang Zhu, Jinglei Lv, Hanbo Chen, Tianming Liu 0001 |
MICCAI (2) | 4 |
| 2014 | Decoding Auditory Saliency from FMRI Brain ImagingabstractGiven the growing number of available audio streams through a variety of sources and distribution channels, effective and advanced computational audio analysis has received increasing interest in the multimedia field. However, the effectiveness of current audio analysis strategies might be hampered due to the lack of effective representation of high-level semantics perceived by the human and the lack of effective approaches to bridging the gaps between most low-level acoustic features and high-level semantic features. This semantic gap has become the 'bottleneck' problem in audio analysis. In this paper, we propose a computational framework to decode biologically-plausible auditory saliency using high-level features derived from functional magnetic resonance imaging (fMRI) which monitors the human brain's response under the natural stimulus of audio listening. Specifically, we identify meaningful intrinsic brain networks which are involved in audio listening via effective online dictionary learning and sparse representation of whole-brain fMRI signals, reconstruct auditory saliency features using those identified brain network components, and perform group-wise analysis to identify consistent 'brain decoders' of the saliency features across different excerpts and participants. Experimental results demonstrate that the auditory saliency features are effectively decoded via our methods, which potentially provide opportunities for various applications in the multimedia field. Shijie Zhao 0001, Xi Jiang 0001, Junwei Han 0001, Xintao Hu, Dajiang Zhu, Jinglei Lv, Lei Guo 0002, Tianming Liu 0001 |
ACM Multimedia | 9 |
| 2014 | Clustering and retrieval of video shots based on natural stimulus fMRI
Junwei Han 0001, Xintao Hu, Jungong Han, Tianming Liu 0001 |
Neurocomputing | 5 |
| 2014 | Video abstraction based on fMRI-driven visual attention model
Junwei Han 0001, Kaiming Li, Ling Shao 0001, Xintao Hu, Lei Guo 0002, Jungong Han, Tianming Liu 0001 |
Inf. Sci. | 8 |
| 2014 | Characterization of U-shape streamline fibers: Methods and applications
Hanbo Chen, Lei Guo 0002, Kaiming Li, Longchuan Li, Shu Zhang 0001, Dinggang Shen, Xiaoping Hu 0001, Tianming Liu 0001 |
Medical Image Anal. | 9 |
| 2014 | Interactive object-based image retrieval and annotation on iPad
Junwei Han 0001, Lei Guo 0002, Tianming Liu 0001 |
Multim. Tools Appl. | 5 |
| 2014 | Merging Neuroimaging and Multimedia: Methods, Opportunities, and ChallengesabstractNeuroimaging and brain mapping can provide meaningful guidance to multimedia analyses. and advanced computational multimedia analysis can be used to better understand the functional mechanisms of the human brain. Essentially, brain imaging and brain mapping techniques can serve as a bridge that links the digital representation of multimedia and the perception and comprehension of its content. This paper summarizes methods that integrate brain imaging with multimedia analysis and discusses the opportunities and challenges in this interdisciplinary field. In general, quantitative modeling of brain responses during multimedia comprehension has advanced content-based multimedia studies such as image and video classification and tagging. Multimedia analysis has promoted functional brain mapping by using naturalistic multimedia as stimuli during neuroimaging. Challenges and opportunities in merging neuroimaging and multimedia include the quantification of the brain's responses, the quantification of multimedia, and the mapping between brain responses and computational multimedia features. Tianming Liu 0001, Xintao Hu, Junwei Han 0001, Lei Guo 0002 |
IEEE Trans. Hum. Mach. Syst. | 1 |
| 2013 | Identifying Group-Wise Consistent White Matter Landmarks via Novel Fiber Shape Descriptor
Hanbo Chen, Tianming Liu 0001 |
MICCAI (1) | 3 |
| 2013 | Predictive Models of Resting State Networks for Assessment of Altered Functional Connectivity in MCI
Xi Jiang 0001, Dajiang Zhu, Kaiming Li, Dinggang Shen, Lei Guo 0002, Tianming Liu 0001 |
MICCAI (2) | 7 |
| 2013 | Anatomy-Guided Discovery of Large-Scale Consistent Connectivity-Based Cortical Landmarks
Xi Jiang 0001, Dajiang Zhu, Kaiming Li, Jinglei Lv, Lei Guo 0002, Tianming Liu 0001 |
MICCAI (3) | 7 |
| 2013 | Modeling Dynamic Functional Information Flows on Large-Scale Brain Networks
Peili Lv, Lei Guo 0002, Xintao Hu, Xiang Li 0001, Changfeng Jin, Junwei Han 0001, Lingjiang Li, Tianming Liu 0001 |
MICCAI (2) | 8 |
| 2013 | Sparse Representation of Group-Wise FMRI Signals
Jinglei Lv, Xiang Li 0001, Dajiang Zhu, Xi Jiang 0001, Xin Zhang 0151, Xintao Hu, Lei Guo 0002, Tianming Liu 0001 |
MICCAI (3) | 9 |
| 2013 | Group-Wise FMRI Activation Detection on Corresponding Cortical Landmarks
Jinglei Lv, Dajiang Zhu, Xintao Hu, Xin Zhang 0151, Junwei Han 0001, Lei Guo 0002, Tianming Liu 0001 |
MICCAI (2) | 8 |
| 2013 | Sparse Representation of Higher-Order Functional Interaction Patterns in Task-Based FMRI Data
Shu Zhang 0001, Xiang Li 0001, Jinglei Lv, Xi Jiang 0001, Dajiang Zhu, Hanbo Chen, Lei Guo 0002, Tianming Liu 0001 |
MICCAI (3) | 9 |
| 2013 | Characterization of task-free and task-performance brain states via functional connectome patterns
Xin Zhang 0151, Lei Guo 0002, Xiang Li 0001, Dajiang Zhu, Kaiming Li, Hanbo Chen, Jinglei Lv, Changfeng Jin, Lingjiang Li, Tianming Liu 0001 |
Medical Image Anal. | 12 |
| 2013 | Predicting cortical ROIs via joint modeling of anatomical and connectional profiles
Dajiang Zhu, Xi Jiang 0001, Bao Ge, Xintao Hu, Junwei Han 0001, Lei Guo 0002, Tianming Liu 0001 |
Medical Image Anal. | 8 |
| 2013 | An Object-Oriented Visual Saliency Detection Framework Based on Sparse Coding RepresentationsabstractSaliency detection aims at quantitatively predicting attended locations in an image. It may mimic the selection mechanism of the human vision system, which processes a small subset of a massive amount of visual input while the redundant information is ignored. Motivated by the biological evidence that the receptive fields of simple cells in V1 of the vision system are similar to sparse codes learned from natural images, this paper proposes a novel framework for saliency detection by using image sparse coding representations as features. Unlike many previous approaches dedicated to examining the local or global contrast of each individual location, this paper develops a probabilistic computational algorithm by integrating objectness likelihood with appearance rarity. In the proposed framework, image sparse coding representations are yielded through learning on a large amount of eye-fixation patches from an eye-tracking dataset. The objectness likelihood is measured by three generic cues called compactness, continuity, and center bias. The appearance rarity is inferred by using a Gaussian mixture model. The proposed paper can serve as a basis for many techniques such as image/video segmentation, retrieval, retargeting, and compression. Extensive evaluations on benchmark databases and comparisons with a number of up-to-date algorithms demonstrate its effectiveness. Junwei Han 0001, Xiaoliang Qian, Lei Guo 0002, Tianming Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2013 | Representing and Retrieving Video Shots in Human-Centric Brain Imaging SpaceabstractMeaningful representation and effective retrieval of video shots in a large-scale database has been a profound challenge for the image/video processing and computer vision communities. A great deal of effort has been devoted to the extraction of low-level visual features, such as color, shape, texture, and motion for characterizing and retrieving video shots. However, the accuracy of these feature descriptors is still far from satisfaction due to the well-known semantic gap. In order to alleviate the problem, this paper investigates a novel methodology of representing and retrieving video shots using human-centric high-level features derived in brain imaging space (BIS) where brain responses to natural stimulus of video watching can be explored and interpreted. At first, our recently developed dense individualized and common connectivity-based cortical landmarks (DICCCOL) system is employed to locate large-scale functional brain networks and their regions of interests (ROIs) that are involved in the comprehension of video stimulus. Then, functional connectivities between various functional ROI pairs are utilized as BIS features to characterize the brain's comprehension of video semantics. Then an effective feature selection procedure is applied to learn the most relevant features while removing redundancy, which results in the formation of the final BIS features. Afterwards, a mapping from low-level visual features to high-level semantic features in the BIS is built via the Gaussian process regression (GPR) algorithm, and a manifold structure is then inferred, in which video key frames are represented by the mapped feature vectors in the BIS. Finally, the manifold-ranking algorithm concerning the relationship among all data is applied to measure the similarity between key frames of video shots. Experimental results on the TRECVID 2005 dataset demonstrate the superiority of the proposed work in comparison with traditional methods. Junwei Han 0001, Xintao Hu, Dajiang Zhu, Kaiming Li, Xi Jiang 0001, Guangbin Cui, Lei Guo 0002, Tianming Liu 0001 |
IEEE Trans. Image Process. | 9 |
| 2013 | Inferring Group-Wise Consistent Multimodal Brain Networks via Multi-View Spectral ClusteringabstractQuantitative modeling and analysis of structural and functional brain networks based on diffusion tensor imaging (DTI) and functional magnetic resonance imaging (fMRI) data have received extensive interest recently. However, the regularity of these structural and functional brain networks across multiple neuroimaging modalities and also across different individuals is largely unknown. This paper presents a novel approach to inferring group-wise consistent brain subnetworks from multimodal DTI/resting-state fMRI datasets via multi-view spectral clustering of cortical networks, which were constructed upon our recently developed and validated large-scale cortical landmarks-DICCCOL (dense individualized and common connectivity-based cortical landmarks). We applied the algorithms on DTI data of 100 healthy young females and 50 healthy young males, obtained consistent multimodal brain networks within and across multiple groups, and further examined the functional roles of these networks. Our experimental results demonstrated that the derived brain networks have substantially improved inter-modality and inter-subject consistency. Hanbo Chen, Kaiming Li, Dajiang Zhu, Xi Jiang 0001, Yixuan Yuan, Peili Lv, Lei Guo 0002, Dinggang Shen, Tianming Liu 0001 |
IEEE Trans. Medical Imaging | 10 |
| 2012 | Group-Wise Consistent Parcellation of Gyri via Adaptive Multi-view Spectral Clustering of Fiber Shapes
Hanbo Chen, Dajiang Zhu, Feiping Nie 0001, Tianming Liu 0001, Heng Huang 0001 |
MICCAI (2) | 5 |
| 2012 | Inferring Group-Wise Consistent Multimodal Brain Networks via Multi-view Spectral Clustering
Hanbo Chen, Kaiming Li, Dajiang Zhu, Changfeng Jin, Lei Guo 0002, Lingjiang Li, Tianming Liu 0001 |
MICCAI (3) | 8 |
| 2012 | Optimization of fMRI-Derived ROIs Based on Coherent Functional Interaction Patterns
Fan Deng 0001, Dajiang Zhu, Tianming Liu 0001 |
MICCAI (3) | 3 |
| 2012 | Group-Wise Consistent Fiber Clustering Based on Multimodal Connectional and Functional Profiles
Bao Ge, Lei Guo 0002, Dajiang Zhu, Kaiming Li, Xintao Hu, Junwei Han 0001, Tianming Liu 0001 |
MICCAI (3) | 8 |
| 2012 | Morphological Analysis of the Left Ventricular Endocardial Surface and Its Clinical Implications
Anirban Mukhopadhyay 0003, Suchendra M. Bhandarkar, Tianming Liu 0001, Sarah Rinehart, Szilard Voros |
MICCAI (2) | 4 |
| 2012 | Characterization of Task-Free/Task-Performance Brain States
Xin Zhang 0151, Lei Guo 0002, Xiang Li 0001, Dajiang Zhu, Kaiming Li, Zhenqiang Sun, Changfeng Jin, Xintao Hu, Junwei Han 0001, Lingjiang Li, Tianming Liu 0001 |
MICCAI (2) | 12 |
| 2012 | Music/speech classification using high-level features derived from fmri brain imagingabstractWith the availability of large amount of audio tracks through a variety of sources and distribution channels, automatic music/speech classification becomes an indispensable tool in social audio websites and online audio communities. However, the accuracy of current acoustic-based low-level feature classification methods is still rather far from satisfaction. The discrepancy between the limited descriptive power of low-level features and the richness of high-level semantics perceived by the human brain has become the 'bottleneck' problem in audio signal analysis. In this paper, functional magnetic resonance imaging (fMRI) which monitors the human brain's response under the natural stimulus of music/speech listening is used as high-level features in the brain imaging space (BIS). We developed a computational framework to model the relationships between BIS features and low-level features in the training dataset with fMRI scans, predict BIS features of testing dataset without fMRI scans, and use the predicted BIS features for music/speech classification in the application stage. Experimental results demonstrated the significantly improved performance of music/speech classification via predicted BIS features than that via the original low-level features. Xi Jiang 0001, Xintao Hu, Lie Lu, Junwei Han 0001, Lei Guo 0002, Tianming Liu 0001 |
ACM Multimedia | 7 |
| 2012 | Bridging the Semantic Gap via Functional Brain ImagingabstractThe multimedia content analysis community has made significant efforts to bridge the gaps between low-level features and high-level semantics perceived by humans. Recent advances in brain imaging and neuroscience in exploring the human brain's responses during multimedia comprehension demonstrated the possibility of leveraging cognitive neuroscience knowledge to bridge the semantic gaps. This paper presents our initial effort in this direction by using functional magnetic resonance imaging (fMRI). Specifically, task-based fMRI (T-fMRI) was performed to accurately localize the brain regions involved in video comprehension. Then, natural stimulus fMRI (N-fMRI) data were acquired when subjects watched the multimedia clips selected from the TRECVID datasets. The responses in the localized brain regions were measured and used to extract high-level features as the representation of the brain's comprehension of semantics in the videos. A novel computational framework was developed to learn the most relevant low-level feature sets that best correlate the fMRI-derived semantic features based on the training videos with fMRI scans, and then the learned model was applied to larger scale TRECVID video datasets without fMRI scans for category classification. Our experimental results demonstrate: 1) there are meaningful couplings between brain's fMRI-derived responses and video stimuli, suggesting the validity of linking semantics and low-level features via fMRI and 2) the computationally learned low-level features can significantly (p <; 0.01) improve video classification in comparison with original low-level features and extracted low-level features resulted from well-known feature projection algorithms. Xintao Hu, Kaiming Li, Junwei Han 0001, Xian-Sheng Hua 0001, Lei Guo 0002, Tianming Liu 0001 |
IEEE Trans. Multim. | 6 |
| 2011 | Retrieving video shots in semantic brain imaging space using manifold-rankingabstractIn recent two decades, a large amount of effort has been devoted to content-based video retrieval (CBVR), which aims to manage large-scale video databases in an effective way based on visual features such as color, shape, texture, and motion. However, the performance of CBVR systems is still far from satisfaction due to the well-known semantic gap. In order to alleviate the problem, this paper proposes a novel retrieval methodology using semantic features derived from brain imaging space (BIS) that reflects brain responses and interactions under natural stimulus of video watching. A mapping from visual features to semantic features in BIS is built through Gaussian process regression. A manifold structure is then inferred where video key frames are represented by mapped feature vectors in BIS. Finally, the manifold-ranking algorithm concerning the relationship among all data is applied to measure the similarity between key frames. Preliminary experimental results on the TRECVID 2005 dataset demonstrate the superiority of the proposed work in comparison with traditional methods. Junwei Han 0001, Xintao Hu, Kaiming Li, Fan Deng 0001, Lei Guo 0002, Tianming Liu 0001 |
ICIP | 8 |
| 2011 | Assessing Regularity and Variability of Cortical Folding Patterns of Working Memory ROIs
Hanbo Chen, Kaiming Li, Xintao Hu, Lei Guo 0002, Tianming Liu 0001 |
MICCAI (2) | 6 |
| 2011 | Resting State fMRI-Guided Fiber Clustering
Bao Ge, Lei Guo 0002, Jinglei Lv, Xintao Hu, Junwei Han 0001, Tianming Liu 0001 |
MICCAI (2) | 7 |
| 2011 | Fiber-Centered Granger Causality Analysis
Xiang Li 0001, Kaiming Li, Lei Guo 0002, Chulwoo Lim, Tianming Liu 0001 |
MICCAI (2) | 5 |
| 2011 | Predicting Functional Brain ROIs via Fiber Shape Models
Lei Guo 0002, Kaiming Li, Dajiang Zhu, Guangbin Cui, Tianming Liu 0001 |
MICCAI (2) | 6 |
| 2011 | A biologically inspired computational model for image saliency detectionabstractImage saliency detection provides a powerful tool for predicting where human tends to look at in an image, which has been a long attempt for the computer vision community. In this paper, we propose a biologically-inspired model for computing image saliency. At first, a set of basis functions that accords with visual responses to natural stimuli is learned by using eye-fixation patches from an eye-tracking dataset. Three features are then derived based on the learned basis functions including continuity, clutter contrast, and local contrast. Finally, these three features are combined into the saliency map. The proposed approach is easy to implement and can be used in many image and video content analysis applications. Experiments on a large-scale benchmark dataset and comparisons with a number of the state-of-the-art approaches demonstrate its superiority. Junwei Han 0001, Xintao Hu, Lei Guo 0002, Tianming Liu 0001 |
ACM Multimedia | 6 |
| 2010 | Automated Cell Phase Classification for Zebrafish Fluorescence Microscope ImagesabstractAutomated cell phenotype image classification is an interesting bioinformatics problem. In this paper, an automated cell phase classification framework is investigated for zebra fish presomitic mesoderm (PSM) images. Low image resolution, gradual transitions between adjacent categories and irregularity of real cell images make this classification task tough but intriguing. The proposed framework first segments zebra fish image into cell patches by a two-stage segmentation procedure, then extracts feature set NF9, which designed especially for this low resolution image set, on each cell patch, and finally employs support vector machine (SVM) as cell classifier. At present, the total accuracy by NF9 is 75%. Yanting Lu, Jianfeng Lu 0003, Tianming Liu 0001, Jing-Yu Yang 0001 |
ICPR | 3 |
| 2010 | A Dynamic Skull Model for Simulation of Cerebral Cortex Folding
Hanbo Chen, Lei Guo 0002, Jingxin Nie, Xintao Hu, Tianming Liu 0001 |
MICCAI (2) | 6 |
| 2010 | Fiber-Centered Analysis of Brain Connectivities Using DTI and Resting State FMRI Data
Jinglei Lv, Lei Guo 0002, Xintao Hu, Kaiming Li, Degang Zhang, Tianming Liu 0001 |
MICCAI (2) | 8 |
| 2010 | Bridging low-level features and high-level semantics via fMRI brain imaging for video classificationabstractThe multimedia content analysis community has made significant effort to bridge the gap between low-level features and high-level semantics perceived by human cognitive systems such as real-world objects and concepts. In the two fields of multimedia analysis and brain imaging, both topics of low-level features and high level semantics are extensively studied. For instance, in the multimedia analysis field, many algorithms are available for multimedia feature extraction, and benchmark datasets are available such as the TRECVID. In the brain imaging field, brain regions that are responsible for vision, auditory perception, language, and working memory are well studied via functional magnetic resonance imaging (fMRI). This paper presents our initial effort in marrying these two fields in order to bridge the gaps between low-level features and high-level semantics via fMRI brain imaging. Our experimental paradigm is that we performed fMRI brain imaging when university student subjects watched the video clips selected from the TRECVID datasets. At current stage, we focus on the three concepts of sports, weather, and commercial-/advertisement specified in the TRECVID 2005. Meanwhile, the brain regions in vision, auditory, language, and working memory networks are quantitatively localized and mapped via task-based paradigm fMRI, and the fMRI responses in these regions are used to extract features as the representation of the brain's comprehension of semantics. Our computational framework aims to learn the most relevant low-level feature sets that best correlate the fMRI-derived semantics based on the training videos with fMRI scans, and then the learned models are applied to larger scale test datasets without fMRI scans for category classifications. Our result shows that: 1) there are meaningful couplings between brain's fMRI responses and video stimuli, suggesting the validity of linking semantics and low-level features via fMRI; 2) The computationally learned low-level feature sets from fMRI-derived semantic features can significantly improve the classification of video categories in comparison with that based on original low-level features. Xintao Hu, Fan Deng 0001, Kaiming Li, Hanbo Chen, Xi Jiang 0001, Jinglei Lv, Dajiang Zhu, Carlos Faraco, Degang Zhang, Arsham Mesbah, Junwei Han 0001, Xian-Sheng Hua 0001, L. Stephen Miller, Lei Guo 0002, Tianming Liu 0001 |
ACM Multimedia | 17 |
| 2010 | Individualized ROI Optimization via Maximization of Group-wise Consistency of Structural and Functional ProfilesabstractFunctional segregation and integration are fundamental characteristics of the human brain. Studying the connectivity among segregated regions and the dynamics of integrated brain networks has drawn increasing interest. A very controversial, yet fundamental issue in these studies is how to determine the best functional brain regions or ROIs (regions of interests) for individuals. Essentially, the computed connectivity patterns and dynamics of brain networks are very sensitive to the locations, sizes, and shapes of the ROIs. This paper presents a novel methodology to optimize the locations of an individual's ROIs in the working memory system. Our strategy is to formulate the individual ROI optimization as a group variance minimization problem, in which group-wise functional and structural connectivity patterns, and anatomic profiles are defined as optimization constraints. The optimization problem is solved via the simulated annealing approach. Our experimental results show that the optimized ROIs have significantly improved consistency in structural and functional profiles across subjects, and have more reasonable localizations and more consistent morphological and anatomic profiles. Kaiming Li, Lei Guo 0002, Carlos Faraco, Dajiang Zhu, Fan Deng 0001, Xi Jiang 0001, Degang Zhang, Hanbo Chen, Xintao Hu, L. Stephen Miller, Tianming Liu 0001 |
NIPS | 12 |
| 2010 | An automated pipeline for cortical sulcal fundi extraction
Gang Li 0001, Lei Guo 0002, Jingxin Nie, Tianming Liu 0001 |
Medical Image Anal. | 4 |
| 2009 | Grouping of Brain MR Images via Affinity PropagationabstractThe human brain anatomy is extremely variable across individuals in terms of its size, shape, and structure patterning. In this paper, a novel method is proposed for grouping brain MR images into different patterns. This method adopts the affinity propagation methodology to partition a population of brain images into different clusters. In the affinity propagation method, the tissue-segmented and anatomically-parcellated images are used to define the similarity between brain images, in contrast to intensity-based similarity measurement used in previous methods. After clustering, in each cluster (called a sub-group) a representative exemplar image is identified as the single subject atlas for the sub-group. Meanwhile, all the subject images belonging to the same sub-group are identified. This method has been applied to the publicly available OASIS neuroimaging dataset that includes 414 subject brain MRI images. Experiments show that the method is able to group brain MR images into different patterns effectively. Gang Li 0001, Lei Guo 0002, Tianming Liu 0001 |
ISCAS | 3 |
| 2009 | Gyral Folding Pattern Analysis via Surface Profiling
Kaiming Li, Lei Guo 0002, Gang Li 0001, Jingxin Nie, Carlos Faraco, L. Stephen Miller, Tianming Liu 0001 |
MICCAI (1) | 8 |
| 2009 | A Computational Model of Cerebral Cortex Folding
Jingxin Nie, Gang Li 0001, Lei Guo 0002, Tianming Liu 0001 |
MICCAI (1) | 4 |
| 2009 | Parametric Representation of Cortical Surface Folding Based on Polynomials
Lei Guo 0002, Gang Li 0001, Jingxin Nie, Tianming Liu 0001 |
MICCAI (1) | 5 |
| 2008 | A Novel Method for Cortical Sulcal Fundi Extraction
Gang Li 0001, Tianming Liu 0001, Jingxin Nie, Lei Guo 0002, Stephen T. C. Wong |
MICCAI (1) | 2 |
| 2008 | ZFIQ: a software package for zebrafish biologyabstractAbstract Summary: Rapid development, transparency and small size are the outstanding features of zebrafish that make it as an increasingly important vertebrate system for developmental biology, functional genomics, disease modeling and drug discovery. Zebrafish has been regarded as ideal animal specie for studying the relationship between genotype and phenotype, for pathway analysis and systems biology. However, the tremendous amount of data generated from large numbers of embryos has led to the bottleneck of data analysis and modeling. The zebrafish image quantitator (ZFIQ) software provides streamlined data processing and analysis capability for developmental biology and disease modeling using zebrafish model. Availability: ZFIQ is available for download at http://www.cbi-platform.net Contact: [email protected] Supplementary information: Additional documentation for this software package is referred to http://www.cbi-platform.net/document.htm. Application examples of this software are referred to http://www.cbi-platform.net/download.htm Tianming Liu 0001, Jingxin Nie, Gang Li 0001, Lei Guo 0002, Stephen T. C. Wong |
Bioinform. | 1 |
| 2006 | Stochastic Robust Stability Analysis for Markovian Jump Discrete-Time Delayed Neural Networks with Multiplicative Nonlinear Perturbations
Tianming Liu 0001, Guodong Lu, Jilin Liu, Stephen T. C. Wong |
ISNN (1) | 2 |
| 2005 | Robust Stability for Delayed Neural Networks with Nonlinear Perturbation
Tianming Liu 0001, Jilin Liu, WeiKang Gu, Stephen T. C. Wong |
ISNN (1) | 2 |
| 2005 | 76-Space Analysis of Grey Matter Diffusivity: Methods and Applications
Tianming Liu 0001, Geoffrey S. Young, Nankuei Chen, Stephen T. C. Wong |
MICCAI | 1 |
| 2004 | Deformable Registration of Tumor-Diseased Brain Images
Tianming Liu 0001, Dinggang Shen, Christos Davatzikos |
MICCAI (1) | 1 |
| 2003 | Deformable Registration of Cortical Structures via Hybrid Volumetric and Surface Warping
Tianming Liu 0001, Dinggang Shen, Christos Davatzikos |
MICCAI (2) | 1 |
| 2003 | A novel video key-frame-extraction algorithm based on perceived motion energy modelabstractThe key frame is a simple yet effective form of summarizing a long video sequence. The number of key frames used to abstract a shot should be compliant to visual content complexity within the shot and the placement of key frames should represent most salient visual content. Motion is the more salient feature in presenting actions or events in video and, thus, should be the feature to determine key frames. We propose a triangle model of perceived motion energy (PME) to model motion patterns in video and a scheme to extract key frames based on this model. The frames at the turning point of the motion acceleration and motion deceleration are selected as key frames. The key-frame selection process is threshold free and fast and the extracted key frames are representative. Tianming Liu 0001, HongJiang Zhang, Feihu Qi |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2002 | Adaptive self-excitation groups in visual curve integration
Lei Guo 0002, Tianming Liu 0001, Junwei Han 0001 |
Neurocomputing | 2 |
| 2002 | A systematic rate controller for MPEG-4 FGS video streaming
Tianming Liu 0001, HongJiang Zhang, Feihu Qi |
Multim. Syst. | 1 |
| 2001 | Systematic rate controller for MPEG-4 FGS video streamingabstractThis paper proposes a systematic rate controller (SRC) for streamed delivery of MPEG-4 FGS video over the Internet. It provides both coarse-grain and fine-grain scalability of rate adaptation to various user access rates, long-term or short-term bandwidth fluctuations, and bit-rate variations of streamed video in a general scope of time-scale. A state machine (SM) is developed to implement the SRC. The adaptation is realized through mode and state transitions of the SM. Tianming Liu 0001, HongJiang Zhang, Feihu Qi |
ICIP (2) | 1 |
| 2001 | A Content-Aware Rate Controller For Streamed Delivery Of MPEG-4 Fgs VideoabstractA content-aware rate controller (CRC) for MPEG-4 FGS stored video streaming over the Internet is proposed. The CRC optimally uses available bandwidth and client buffer to produce smooth video quality and provide protection to video segments with important content. The FGS video stream is forward-shifted to client buffer by actively dropping high enhancement layers. When bandwidth decreases sharply, the FS buffer is used as bandwidth so that the sharp bandwidth drop is hidden from the decoder. The active dropping and forward-shifting process is restarted when bandwidth recovers. Thus, higher layers of less importance content are dropped and lower layers of more importance are protected. A priori information about the stream content is used in the CRC such that video segments with important content are protected by allocating them more bandwidth. Tianming Liu 0001, HongJiang Zhang, Feihu Qi |
ICME | 1 |
| 2000 | Novel approach of combining temporal segmentation results to the region-binding process for separating moving objects from still background
Tianming Liu 0001, Feihu Qi, Yiqiang Zhan |
VCIP | 1 |
| 1999 | Random time-division operation for salience of visual contoursabstractThe mutual excitation among the local stimuli satisfying curve distribution (position and orientation continuity) called self-excitation of curves here is an effective method for the discovery and enhancement of visual curves. This article presents a new method using dynamic time-division curve searches and self-excitation. The searches realized by random walks of active particle are guided by inputs, limited by the rules of curve distribution performed repetitively, and temporally divided for different curve candidates. The time-division operations play an extremely important role in both the structure division used previously and the global memory of various search routes. Lei Guo 0002, Tianming Liu 0001 |
IJCNN | 2 |