Jiaxing Xu

dblp:118/5482 · DBLP profile ↗
← Back
27ranked-venue papers
10as first author
27since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 4 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 5 first-author · 9 since 2021Databases, data management, data science and information retrieval · 8 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Bridge Breaking for Graph Neural Networks
abstract
Graph Neural Networks (GNNs) have demonstrated remarkable success in modeling graph-structured data, particularly under the assumption of homophily, where connected nodes share similar attributes or class labels. However, many real-world networks exhibit heterophily, leading to the suboptimal performance of conventional GNNs. While existing heterophily-aware models primarily address feature and class differences between central and neighboring nodes, we identify a critical yet underexplored challenge: class disparities in the neighborhoods of bridge nodes—nodes that connect disparate classes. Through theoretical analysis, we demonstrate how neighborhood differences around bridge nodes increase classification difficulty. To tackle this, we propose Bridge Breaking Graph Neural Network (BBGNN), a novel approach that explicitly mitigates performance degradation in these critical regions. We introduce a bridge ratio metric to identify bridge nodes without requiring label information and design a bridge-breaking aggregation mechanism to counteract excessive smoothing in these regions. Extensive experiments across multiple benchmark datasets validate the effectiveness of BBGNN, significantly improving GNN performance in bridge node regions.
Wei Li 0231, Jiaxing Xu, Xia Dong, Yiping Ke
WSDM2
2026 Text4Seg++: Advancing Image Segmentation via Generative Language Modeling
abstract
Multimodal Large Language Models (MLLMs) have shown exceptional capabilities in vision-language tasks. However, effectively integrating image segmentation into these models remains a significant challenge.In this work, we propose a novel text-as-mask paradigm that casts image segmentation as a text generation problem, eliminating the need for additional decoders and significantly simplifying the segmentation process. Our key innovation is semantic descriptors, a new textual representation of segmentation masks where each image patch is mapped to its corresponding text label. We first introduce image-wise semantic descriptors, a patch-aligned textual representation of segmentation masks that integrates naturally into the language modeling pipeline. To enhance efficiency, we introduce the Row-wise Run-Length Encoding (R-RLE), which compresses redundant text sequences, reducing the length of semantic descriptorsby 74% and accelerating inference by $3\times$3×, without compromising performance. Building upon this, our initial framework Text4Segachieves strong segmentation performance across a wide range of vision tasks. To further improve granularity and compactness, we propose box-wise semantic descriptors, which localizes regions of interest using bounding boxes and represents region masks via structured mask tokens called semantic bricks. This leads to our refined model, Text4Seg++, which formulates segmentation as a next-brick prediction task, combining precision, scalability, and generative efficiency. Comprehensive experiments on natural and remote sensing datasets show that Text4Seg++consistently outperforms state-of-the-art models across diverse benchmarks without any task-specific fine-tuning, while remaining compatible with existing MLLM backbones. Our work highlights the effectiveness, scalability, and generalizability of text-driven image segmentation within the MLLM framework.
Mengcheng Lan, Chaofeng Chen, Jiaxing Xu, Zongrui Li 0001, Yiping Ke, Xudong Jiang 0001, Yingchen Yu, Yunqing Zhao, Song Bai 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 External Retrievals or Internal Priors? From RAG to Epitome-Augmented Generation by Fuzzy Selection
abstract
Retrieval-Augmented Generation (RAG) offers a promising solution to the limitations of static knowledge and hallucinations in Large Language Models (LLMs). While prior research has introduced numerous enhancements to RAG systems, a significant challenge remains under-explored: the potential conflict between external retrievals and LLMs' internal priors, which can undermine the quality of generated outputs. To tackle this issue, we present theEpitome-AugmentedGeneration (EAG) framework, which strategically aligns queries, external retrievals, and internal priors to produce high-quality LLM generations by selecting fuzzy inputs. EAG employs two novel lightweight modules, Criticism and Distillation, allowing traditional RAGs to be upgraded to EAGs without the need for specialized training data. Extensive experiments on five datasets across general and medical domains, including both open-ended and closed-ended tasks, validate the effectiveness of EAG. Our framework achieves substantial F1 score improvements: 7.03%, 23.35%, and 21.58% over baseline RAGs in medical QA tasks, 11.80% in law domain, 7.95% in finance domain, and 4.13% and 5.16% in general domain. Beyond performance gains, our study delves into the interplay between LLMs' internal priors and external retrievals, uncovering key principles that govern generation quality and providing valuable insights for future retrieval-augmented frameworks.
Kai He 0001, Jiaxing Xu, Qika Lin, Zeyu Gao 0001, Jialun Wu, Mengling Feng
IEEE Trans. Fuzzy Syst.2
2026 Multi-Atlas Brain Network Classification Through Consistency Distillation and Complementary Information Fusion
abstract
Brain network analysis plays a crucial role in identifying distinctive patterns associated with neurological disorders. Functional magnetic resonance imaging (fMRI) enables the construction of brain networks by analyzing correlations in blood-oxygen-level-dependent (BOLD) signals across different brain regions, known as regions of interest (ROIs). These networks are typically constructed using atlases that parcellate the brain based on various hypotheses of functional and anatomical divisions. However, there is no standard atlas for brain network classification, leading to limitations in detecting abnormalities in disorders. Recent methods leveraging multiple atlases fail to ensure consistency across atlases and lack effective ROI-level information exchange, limiting their efficacy. To address these challenges, we propose the Atlas-Integrated Distillation and Fusion network (AIDFusion), a novel framework designed to enhance brain network classification using fMRI data. AIDFusion introduces a disentangle Transformer to filter out inconsistent atlas-specific information and distill meaningful cross-atlas connections. Additionally, it enforces subject- and population-level consistency constraints to improve cross-atlas coherence. To further enhance feature integration, AIDFusion incorporates an inter-atlas message-passing mechanism that facilitates the fusion of complementary information across brain regions. We evaluate AIDFusion on four resting-state fMRI datasets encompassing different neurological disorders. Experimental results demonstrate its superior classification performance and computational efficiency compared to state-of-the-art methods. Furthermore, a case study highlights AIDFusion's ability to extract interpretable patterns that align with established neuroscience findings, reinforcing its potential as a robust tool for multi-atlas brain network analysis.
Jiaxing Xu, Mengcheng Lan, Xia Dong, Kai He 0001, Wayne Zhang 0001, Qingtian Bian, Yiping Ke
IEEE J. Biomed. Health Informatics1
2026 BrainPrompt+: Multi-Level Brain Prompt Learning for Knowledge-Guided Neurological Disorder Identification
abstract
Accurate identification of neurological disorders such as Alzheimer's disease (AD), Parkinson's disease (PD), and Autism Spectrum Disorder (ASD) is challenging due to subtle early-stage symptoms and heterogeneous brain dynamics. Resting-state functional MRI (rs-fMRI) enables the construction of functional brain networks, where Graph Neural Networks (GNNs) have shown promise for disease classification. However, existing GNN-based methods face three key limitations: correlation-based graph construction introduces noise and negative edges; domain knowledge about brain regions is ignored; and demographic or clinical metadata are fused through simplistic encodings. To overcome these limitations, we propose BrainPrompt+, a knowledge-guided framework that integrates Large Language Models (LLMs) with multi-level natural language prompts. Five types of prompts are introduced: spectral (frequency-domain BOLD features), spatial (inter-ROI connectivity), ROI (anatomical and functional knowledge), disease (progression stages), and subject (demographic context). These prompts are encoded by a frozen LLM and incorporated into a GNN pipeline, unifying imaging, clinical, and external knowledge in a semantically enriched and interpretable manner. Experiments on three rs-fMRI datasets show that BrainPrompt+ consistently outperforms state-of-the-art baselines, achieving accuracy gains of up to 8.93%. Biomarker analysis further demonstrates that the highlighted ROIs align with established neuroscience findings, confirming the interpretability of the model. BrainPrompt+ thus establishes a flexible and generalizable paradigm for knowledge-guided brain network analysis. The source code is available at https://github.com/AngusMonroe/BrainPromptPlus.
Jiaxing Xu, Kai He 0001, Wei Li 0231, Mengcheng Lan, Yue Xun, Qika Lin, Peifan Ran, Yiping Ke, Mengling Feng
IEEE Trans. Medical Imaging1
2025 Crab: A Novel Configurable Role-Playing LLM with Assessing Benchmark
abstract
Kai He, Yucheng Huang, Wenqing Wang, Delong Ran, Dongming Sheng, Junxuan Huang, Qika Lin, Jiaxing Xu, Wenqiang Liu, Mengling Feng. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Kai He 0001, Delong Ran, Dongming Sheng, Junxuan Huang, Qika Lin, Jiaxing Xu, Mengling Feng
ACL (1)8
2025 Neuro-Symbolic AI in Healthcare
abstract
Medical AI has achieved strong predictive performance, yet most systems remain limited by shallow reasoning, poor transparency, and weak generalisation in safety-critical settings. Neurosymbolic AI offers a path beyond these constraints by combining neural models' ability to learn from complex clinical data with the explicit structure, logic, and domain knowledge of symbolic methods. This article examines how neurosymbolic approaches can address core challenges in healthcare AI through five key areas: hybrid reasoning that unifies learning and logic; symbol grounding that links internal representations to clinically meaningful concepts; clinical interpretability that exposes reasoning steps; human-integrated decision-making that keeps clinicians in control; and knowledge-driven diagnosis that incorporates guidelines, ontologies, and causal understanding. Together, these elements outline how neurosymbolic AI can support systems that are not only accurate but also transparent, clinically aligned, and robust in complex or data-sparse scenarios. Advancing this paradigm will require collaboration across AI research, clinical practice, and knowledge engineering, as well as governance mechanisms that ensure fairness and accountability. Neurosymbolic AI thus represents a promising direction for building trustworthy, knowledge-rich intelligence in healthcare.
Jialun Wu, Xin Mei, Kai He 0001, Jiaxing Xu, Qika Lin, Zeyu Gao 0001, Rui Mao 0010
BIBM4
2025 Text4Seg: Reimagining Image Segmentation as Text Generation
abstract
Multimodal Large Language Models (MLLMs) have shown exceptional capabilities in vision-language tasks; however, effectively integrating image segmentation into these models remains a significant challenge. In this paper, we introduce Text4Seg, a novel text-as-mask paradigm that casts image segmentation as a text generation problem, eliminating the need for additional decoders and significantly simplifying the segmentation process. Our key innovation is semantic descriptors, a new textual representation of segmentation masks where each image patch is mapped to its corresponding text label. This unified representation allows seamless integration into the auto-regressive training pipeline of MLLMs for easier optimization. We demonstrate that representing an image with $16\times16$ semantic descriptors yields competitive segmentation performance. To enhance efficiency, we introduce the Row-wise Run-Length Encoding (R-RLE), which compresses redundant text sequences, reducing the length of semantic descriptors by 74\% and accelerating inference by $3\times$, without compromising performance. Extensive experiments across various vision tasks, such as referring expression segmentation and comprehension, show that Text4Seg achieves state-of-the-art performance on multiple datasets by fine-tuning different MLLM backbones. Our approach provides an efficient, scalable solution for vision-centric tasks within the MLLM framework.
Mengcheng Lan, Chaofeng Chen, Yue Zhou 0005, Jiaxing Xu, Yiping Ke, Xinjiang Wang, Litong Feng, Wayne Zhang 0001
ICLR4
2025 BrainOOD: Out-of-distribution Generalizable Brain Network Analysis
abstract
In neuroscience, identifying distinct patterns linked to neurological disorders, such as Alzheimer's and Autism, is critical for early diagnosis and effective intervention. Graph Neural Networks (GNNs) have shown promising in analyzing brain networks, but there are two major challenges in using GNNs: (1) distribution shifts in multi-site brain network data, leading to poor Out-of-Distribution (OOD) generalization, and (2) limited interpretability in identifying key brain regions critical to neurological disorders. Existing graph OOD methods, while effective in other domains, struggle with the unique characteristics of brain networks. To bridge these gaps, we introduce BrainOOD, a novel framework tailored for brain networks that enhances GNNs' OOD generalization and interpretability. BrainOOD framework consists of a feature selector and a structure extractor, which incorporates various auxiliary losses including an improved Graph Information Bottleneck (GIB) objective to recover causal subgraphs. By aligning structure selection across brain networks and filtering noisy features, BrainOOD offers reliable interpretations of critical brain regions. Our approach outperforms 16 existing methods and improves generalization to OOD subjects by up to 8.5%. Case studies highlight the scientific validity of the patterns extracted, which aligns with the findings in known neuroscience literature. We also propose the first OOD brain network benchmark, which provides a foundation for future research in this field. Our code is available at https://github.com/AngusMonroe/BrainOOD.
Jiaxing Xu, Yongqiang Chen 0002, Xia Dong, Mengcheng Lan, Qingtian Bian, James Cheng, Yiping Ke
ICLR1
2025 Divergent Paths: Separating Homophilic and Heterophilic Learning for Enhanced Graph-level Representations
abstract
Graph Convolutional Networks (GCNs) are predominantly tailored for graphs displaying homophily, where similar nodes connect, but often fail on heterophilic graphs. The strategy of adopting distinct approaches to learn from homophilic and heterophilic components in node-level tasks has been widely discussed and proven effective both theoretically and experimentally. However, in graph-level tasks, research on this topic remains notably scarce. Addressing this gap, our research conducts an analysis on graphs with nodes' category ID available, distinguishing intra-category and inter-category components as embodiment of homophily and heterophily, respectively. We find while GCNs excel at extracting information within categories, they frequently capture noise from inter-category components. Consequently, it is crucial to employ distinct learning strategies for intra- and inter-category elements. To alleviate this problem, we separately learn the intra- and inter-category parts by a combination of an intra-category convolution (IntraNet) and an inter-category high-pass graph convolution (InterNet). Our IntraNet is supported by sophisticated graph preprocessing steps and a novel category-based graph readout function. For the InterNet, we utilize a high-pass filter to amplify the node disparities, enhancing the recognition of details in the high-frequency components. The proposed approach, DivGNN, combines the IntraNet and InterNet with a gated mechanism and substantially improves classification performance on graph-level tasks, surpassing traditional GNN baselines in effectiveness.
Han Lei, Jiaxing Xu, Xia Dong, Yiping Ke
KDD (2)2
2025 BrainPrompt: Multi-level Brain Prompt Enhancement for Neurological Condition Identification
Jiaxing Xu, Kai He 0001, Wei Li 0231, Mengcheng Lan, Xia Dong, Yiping Ke, Mengling Feng
MICCAI (12)1
2025 Ada-FCN: Adaptive Frequency-Coupled Network for fMRI-Based Brain Disorder Classification
Yue Xun, Jiaxing Xu
MICCAI (12)2
2025 Multi-Domain Enhancement via Residual Interwoven Transfer in Cross-Domain Sequential Recommendation
abstract
To mitigate data sparsity in Sequential Recommendation, Cross-Domain Sequential Recommendation (CDSR) exploits dynamic knowledge transfer across domains. Traditional CDSR approaches merge specific-domain sequences into mixed-domain sequences to reconnect users' dispersed interests. However, most methods rely on unidirectional transfer between mixed and specific domains on each domain task, overlooking the complex interplay between mixed-domain and domain-specific dynamics. Moreover, token-level transfer between coinciding domain sequences fails to consider inherent sequential dynamics. To address these limitations, we propose Multi-Domain Enhancement via Residual Interwoven Transfer (MERIT). Specifically, MERIT enhances domain representations along multiple domain-to-domain paths, leveraging the proposed extended cross-attention fusion compatible with partially overlapping sequences. To facilitate such transfers, MERIT further employs MoE networks in encoders to generate both intra-domain and inter-domain representations. In addition, by integrating stopped-gradient mixed-domain representations into specific-domain representations, MERIT enables the model to learn the residual signal of the mixed-domain information, better aligning with downstream specific-domain tasks. Extensive experiments on three real-world datasets demonstrate that MERIT consistently outperforms state-of-the-art CDSR counterparts with statistical significance.
Qingtian Bian, Tieying Li, Marcus Vinícius de Carvalho, Jiaxing Xu, Hui Fang 0002, Yiping Ke
ACM Multimedia4
2025 ABXI: Invariant Interest Adaptation for Task-Guided Cross-Domain Sequential Recommendation
abstract
Cross-Domain Sequential Recommendation (CDSR) has recently gained attention for countering data sparsity by transferring knowledge across domains.A common approach merges domain-specific sequences into cross-domain sequences, serving as bridges to connect domains.One key challenge is to correctly extract the shared knowledge among these sequences and appropriately transfer it.Most existing works directly transfer unfiltered cross-domain knowledge rather than extracting domain-invariant components and adaptively integrating them into domain-specific modelings.Another challenge lies in aligning the domain-specific and cross-domain sequences.Existing methods align these sequences based on timestamps, but this approach can cause prediction mismatches when the current tokens and their targets belong to different domains.In such cases, the domain-specific knowledge carried by the current tokens may degrade performance.To address these challenges, we propose the A-B-Cross-to-Invariant Learning Recommender (ABXI).Specifically, leveraging LoRA's effectiveness for efficient adaptation, ABXI incorporates two types of LoRAs to facilitate knowledge adaptation.First, all sequences are processed through a shared encoder that employs a domain LoRA for each sequence, thereby preserving unique domain characteristics.Next, we introduce an invariant projector that extracts domain-invariant interests from cross-domain representations, utilizing an invariant LoRA to adapt these interests into modeling each specific domain.Besides, to avoid prediction mismatches, all domain-specific sequences are aligned to match the domains of the cross-domain ground truths.
Qingtian Bian, Marcus Vinícius de Carvalho, Tieying Li, Jiaxing Xu, Hui Fang 0002, Yiping Ke
WWW4
2025 Rethinking the message passing for graph-level classification tasks in a category-based view
Jiaxing Xu, Jinjie Ni, Yiping Ke
Eng. Appl. Artif. Intell.2
2024 Union Subgraph Neural Networks
abstract
Graph Neural Networks (GNNs) are widely used for graph representation learning in many application domains. The expressiveness of vanilla GNNs is upper-bounded by 1-dimensional Weisfeiler-Leman (1-WL) test as they operate on rooted subtrees through iterative message passing. In this paper, we empower GNNs by injecting neighbor-connectivity information extracted from a new type of substructure. We first investigate different kinds of connectivities existing in a local neighborhood and identify a substructure called union subgraph, which is able to capture the complete picture of the 1-hop neighborhood of an edge. We then design a shortest-path-based substructure descriptor that possesses three nice properties and can effectively encode the high-order connectivities in union subgraphs. By infusing the encoded neighbor connectivities, we propose a novel model, namely Union Subgraph Neural Network (UnionSNN), which is proven to be strictly more powerful than 1-WL in distinguishing non-isomorphic graphs. Additionally, the local encoding from union subgraphs can also be injected into arbitrary message-passing neural networks (MPNNs) and Transformer-based models as a plugin. Extensive experiments on 18 benchmarks of both graph-level and node-level tasks demonstrate that UnionSNN outperforms state-of-the-art baseline models, with competitive computational efficiency. The injection of our local encoding to existing models is able to boost the performance by up to 11.09%. Our code is available at https://github.com/AngusMonroe/UnionSNN.
Jiaxing Xu, Aihu Zhang, Qingtian Bian, Vijay Prakash Dwivedi, Yiping Ke
AAAI1
2024 Contrasformer: A Brain Network Contrastive Transformer for Neurodegenerative Condition Identification
abstract
Understanding neurological disorder is a fundamental problem in neuroscience, which often requires the analysis of brain networks derived from functional magnetic resonance imaging (fMRI) data. Despite the prevalence of Graph Neural Networks (GNNs) and Graph Transformers in various domains, applying them to brain networks faces challenges. Specifically, the datasets are severely impacted by the noises caused by distribution shifts across sub- populations and the neglect of node identities, both obstruct the identification of disease-specific patterns. To tackle these challenges, we propose Contrasformer, a novel contrastive brain network Transformer. It generates a prior-knowledge-enhanced contrast graph to address the distribution shifts across sub-populations by a two-stream attention mechanism. A cross attention with identity embedding highlights the identity of nodes, and three auxiliary losses ensure group consistency. Evaluated on 4 functional brain network datasets over 4 different diseases, Contrasformer outperforms the state-of-the-art methods for brain networks by achieving up to 10.8% improvement in accuracy, which demonstrates its efficacy in neurological disorder identification. Case studies illustrate its interpretability, especially in the context of neuroscience. This paper provides a solution for analyzing brain networks, offering valuable insights into neurological disorders. Our code is available at https://github.com/AngusMonroe/Contrasformer.
Jiaxing Xu, Kai He 0001, Mengcheng Lan, Qingtian Bian, Wei Li 0231, Tieying Li, Yiping Ke, Miao Qiao
CIKM1
2024 Alleviating the Inconsistency of Multimodal Data in Cross-Modal Retrieval
abstract
With the explosive growth of multimodal Internet data, cross-modal hashing retrieval has become crucial for semantically searching instances across different modalities. However, existing cross-modal retrieval methods rely on assumptions of perfect consistency between modalities and between modalities and labels, which often do not hold in real-world data. We introduce two types of inconsistency: Modality-Modality (M-M) and Modality-Label (M-L) inconsistencies. We further validate the prevalent existence of inconsistent data in multimodal datasets and highlight it will reduce the accuracy of existing Cross-Modal retrieval methods. In this paper, we propose a novel framework called Inconsistency Alleviated Cross-Modal Retrieval (IA-CMR), addressing challenges posed by these inconsistencies. We first utilize two forms of contrastive learning loss and a mutual exclusion constraint to effectively disentangle modal information into modality-common hash codes and modality-unique hash codes. Our dedicated design in modality disentanglement is capable of alleviating the M-M inconsistency. Subsequently, we refine common labels through a label refinement loss and employ a Cross-modal Common Semantic Alignment module for effective alignment. The label refinement process and the CCSA module collectively handle the M-L inconsistency issue. IA-CMR outperforms 9 comparison baselines on two benchmark multimodal datasets, achieving an improvement in retrieval accuracy of up to 25.13%. The results confirm the effectiveness of IA-CMR in alleviating inconsistency and enhancing cross-modal retrieval performance.
Tieying Li, Xiaochun Yang 0001, Yiping Ke, Bin Wang 0015, Yinan Liu 0001, Jiaxing Xu
ICDE6
2024 Benchmarking Complex Instruction-Following with Multiple Constraints Composition
abstract
Instruction following is one of the fundamental capabilities of large language models (LLMs). As the ability of LLMs is constantly improving, they have been increasingly applied to deal with complex human instructions in real-world scenarios. Therefore, how to evaluate the ability of complex instruction-following of LLMs has become a critical research problem. Existing benchmarks mainly focus on modeling different types of constraints in human instructions while neglecting the composition of different constraints, which is an indispensable constituent in complex instructions. To this end, we propose ComplexBench, a benchmark for comprehensively evaluating the ability of LLMs to follow complex instructions composed of multiple constraints. We propose a hierarchical taxonomy for complex instructions, including 4 constraint types, 19 constraint dimensions, and 4 composition types, and manually collect a high-quality dataset accordingly. To make the evaluation reliable, we augment LLM-based evaluators with rules to effectively verify whether generated texts can satisfy each constraint and composition. Furthermore, we obtain the final evaluation score based on the dependency structure determined by different composition types. ComplexBench identifies significant deficiencies in existing LLMs when dealing with complex instructions with multiple constraints composition.
Bosi Wen, Pei Ke, Xiaotao Gu, Lindong Wu, Jinfeng Zhou, Wenchuang Li, Binxin Hu, Wendy Gao, Jiaxing Xu, Jie Tang 0001, Hongning Wang, Minlie Huang
NeurIPS10
2024 A class-aware representation refinement framework for graph classification
Jiaxing Xu, Jinjie Ni, Yiping Ke
Inf. Sci.1
2024 Contrastive Graph Pooling for Explainable Classification of Brain Networks
abstract
Functional magnetic resonance imaging (fMRI) is a commonly used technique to measure neural activation. Its application has been particularly important in identifying underlying neurodegenerative conditions such as Parkinson's, Alzheimer's, and Autism. Recent analysis of fMRI data models the brain as a graph and extracts features by graph neural networks (GNNs). However, the unique characteristics of fMRI data require a special design of GNN. Tailoring GNN to generate effective and domain-explainable features remains challenging. In this paper, we propose a contrastive dual-attention block and a differentiable graph pooling method called ContrastPool to better utilize GNN for brain networks, meeting fMRI-specific requirements. We apply our method to 5 resting-state fMRI brain network datasets of 3 diseases and demonstrate its superiority over state-of-the-art baselines. Our case study confirms that the patterns extracted by our method match the domain knowledge in neuroscience literature, and disclose direct and interesting insights. Our contributions underscore the potential of ContrastPool for advancing the understanding of brain networks and neurodegenerative conditions. The source code is available at https://github.com/AngusMonroe/ContrastPool.
Jiaxing Xu, Qingtian Bian, Xinhang Li 0001, Aihu Zhang, Yiping Ke, Miao Qiao, Wei Zhang 0266, Wei Khang Jeremy Sim, Balázs Gulyás
IEEE Trans. Medical Imaging1
2024 Corrections to "Contrastive Graph Pooling for Explainable Classification of Brain Networks"
Jiaxing Xu, Qingtian Bian, Xinhang Li 0001, Aihu Zhang, Yiping Ke, Miao Qiao, Wei Zhang 0266, Wei Khang Jeremy Sim, Balázs Gulyás
IEEE Trans. Medical Imaging1
2023 CPMR: Context-Aware Incremental Sequential Recommendation with Pseudo-Multi-Task Learning
abstract
The motivations of users to make interactions can be divided into static preference and dynamic interest. To accurately model user representations over time, recent studies in sequential recommendation utilize information propagation and evolution to mine from batches of arriving interactions. However, they ignore the fact that people are easily influenced by the recent actions of other users in the contextual scenario, and applying evolution across all historical interactions dilutes the importance of recent ones, thus failing to model the evolution of dynamic interest accurately. To address this issue, we propose a Context-Aware Pseudo-Multi-Task Recommender System (CPMR) to model the evolution in both historical and contextual scenarios by creating three representations for each user and item under different dynamics: static embedding, historical temporal states, and contextual temporal states. To dually improve the performance of temporal states evolution and incremental recommendation, we design a Pseudo-Multi-Task Learning (PMTL) paradigm by stacking the incremental single-target recommendations into one multi-target task for joint optimization. Within the PMTL paradigm, CPMR employs a shared-bottom network to conduct the evolution of temporal states across historical and contextual scenarios, as well as the fusion of them at the user-item level. In addition, CPMR incorporates one real tower for incremental predictions, and two pseudo towers dedicated to updating the respective temporal states based on new batches of interactions. Experimental results on four benchmark recommendation datasets show that CPMR consistently outperforms state-of-the-art baselines and achieves significant gains on three of them. The source code is available at https://github.com/DiMarzioBian/CPMR.
Qingtian Bian, Jiaxing Xu, Hui Fang 0002, Yiping Ke
CIKM2
2023 SmooSeg: Smoothness Prior for Unsupervised Semantic Segmentation
abstract
Unsupervised semantic segmentation is a challenging task that segments images into semantic groups without manual annotation. Prior works have primarily focused on leveraging prior knowledge of semantic consistency or priori concepts from self-supervised learning methods, which often overlook the coherence property of image segments. In this paper, we demonstrate that the smoothness prior, asserting that close features in a metric space share the same semantics, can significantly simplify segmentation by casting unsupervised semantic segmentation as an energy minimization problem. Under this paradigm, we propose a novel approach called SmooSeg that harnesses self-supervised learning methods to model the closeness relationships among observations as smoothness signals. To effectively discover coherent semantic segments, we introduce a novel smoothness loss that promotes piecewise smoothness within segments while preserving discontinuities across different segments. Additionally, to further enhance segmentation quality, we design an asymmetric teacher-student style predictor that generates smoothly updated pseudo labels, facilitating an optimal fit between observations and labeling outputs. Thanks to the rich supervision cues of the smoothness prior, our SmooSeg significantly outperforms STEGO in terms of pixel accuracy on three datasets: COCOStuff (+14.9\%), Cityscapes (+13.0\%), and Potsdam-3 (+5.7\%).
Mengcheng Lan, Xinjiang Wang, Yiping Ke, Jiaxing Xu, Litong Feng, Wayne Zhang 0001
NeurIPS4
2023 Data-Driven Network Neuroscience: On Data Collection and Benchmark
abstract
This paper presents a comprehensive and quality collection of functional human brain network data for potential research in the intersection of neuroscience, machine learning, and graph analytics. Anatomical and functional MRI images have been used to understand the functional connectivity of the human brain and are particularly important in identifying underlying neurodegenerative conditions such as Alzheimer's, Parkinson's, and Autism. Recently, the study of the brain in the form of brain networks using machine learning and graph analytics has become increasingly popular, especially to predict the early onset of these conditions. A brain network, represented as a graph, retains rich structural and positional information that traditional examination methods are unable to capture. However, the lack of publicly accessible brain network data prevents researchers from data-driven explorations. One of the main difficulties lies in the complicated domain-specific preprocessing steps and the exhaustive computation required to convert the data from MRI images into brain networks. We bridge this gap by collecting a large amount of MRI images from public databases and a private source, working with domain experts to make sensible design choices, and preprocessing the MRI images to produce a collection of brain network datasets. The datasets originate from 6 different sources, cover 4 brain conditions, and consist of a total of 2,702 subjects. We test our graph datasets on 12 machine learning models to provide baselines and validate the data quality on a recent graph analysis model. To lower the barrier to entry and promote the research in this interdisciplinary field, we release our brain network data and complete preprocessing details including codes at https://doi.org/10.17608/k6.auckland.21397377 and https://github.com/brainnetuoa/datadrivennetwork_neuroscience.
Jiaxing Xu, Yunhan Yang, David Tse Jung Huang, Sophi Shilpa Gururajapathy, Yiping Ke, Miao Qiao, Haribalan Kumar, Josh McGeown, Eryn Kwon
NeurIPS1
2023 IMF: Interactive Multimodal Fusion Model for Link Prediction
abstract
Link prediction aims to identify potential missing triples in knowledge graphs. To get better results, some recent studies have introduced multimodal information to link prediction. However, these methods utilize multimodal information separately and neglect the complicated interaction between different modalities. In this paper, we aim at better modeling the inter-modality information and thus introduce a novel Interactive Multimodal Fusion (IMF) model to integrate knowledge from different modalities. To this end, we propose a two-stage multimodal fusion framework to preserve modality-specific knowledge as well as take advantage of the complementarity between different modalities. Instead of directly projecting different modalities into a unified space, our multimodal fusion module limits the representations of different modalities independent while leverages bilinear pooling for fusion and incorporates contrastive learning as additional constraints. Furthermore, the decision fusion module delivers the learned weighted average over the predictions of all modalities to better incorporate the complementarity of different modalities. Our approach has been demonstrated to be effective through empirical evaluations on several real-world datasets. The implementation code is available online at https://github.com/HestiaSky/IMF-Pytorch.
Xinhang Li 0001, Xiangyu Zhao 0001, Jiaxing Xu, Yong Zhang 0002, Chunxiao Xing
WWW3
2023 Graph over-parameterization: Why the graph helps the training of deep graph convolutional network
Yucong Lin, Silu Li, Jiaxing Xu, Wendi Zheng
Neurocomputing3