Xiao-Xin He

dblp:72/5872 · also Xiaoxin He · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
9since 2021 · last 2025
0000-0002-8281-8070ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Computer networks · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
10 papers
Graph learning · 56% Language models and text generation · 9% Deep learning architectures and training · 8%
Network and information security
1 paper
Security and privacy of machine learning · 100%

Topics — the 29 heaviest of 30, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Graph learning
graph neural network
1.832025
A General Graph Spectral Wavelet Convolution via Chebyshev Order Decomposition · ICML 2025
A Generalization of ViT/MLP-Mixer to Graphs · ICML 2023
G-Retriever: Retrieval-Augmented Generation for Textual Graph Understanding and Question Answering · NeurIPS 2024
Machine learning › Graph learning
graph foundation model
1.722025
UniGraph2: Learning a Unified Embedding Space to Bind Multimodal Graphs · WWW 2025
UniGraph: Learning a Unified Cross-Domain Foundation Model for Text-Attributed Graphs · KDD (1) 2025
Machine learning › Transfer learning and domain adaptation
cross-domain transfer
1.122025
UniGraph: Learning a Unified Cross-Domain Foundation Model for Text-Attributed Graphs · KDD (1) 2025
UniGraph2: Learning a Unified Embedding Space to Bind Multimodal Graphs · WWW 2025
Natural language and speech › Language models and text generation
large language model
1.122025
FlipAttack: Jailbreak LLMs via Flipping · ICML 2025
Harnessing Explanations: LLM-to-LM Interpreter for Enhanced Text-Attributed Graph Representation Learning · ICLR 2024
Machine learning › Graph learning
graph self-supervised learning
0.912025
UniGraph: Learning a Unified Cross-Domain Foundation Model for Text-Attributed Graphs · KDD (1) 2025
Machine learning › Graph learning › graph signal processing
graph wavelet
0.912025
A General Graph Spectral Wavelet Convolution via Chebyshev Order Decomposition · ICML 2025
Machine learning › Graph learning › graph self-supervised learning
masked graph modeling
0.912025
UniGraph: Learning a Unified Cross-Domain Foundation Model for Text-Attributed Graphs · KDD (1) 2025
Machine learning › Graph learning
multimodal graph learning
0.912025
UniGraph2: Learning a Unified Embedding Space to Bind Multimodal Graphs · WWW 2025
Machine learning › Graph learning › graph neural network › spectral graph neural network
spectral graph convolution
0.912025
A General Graph Spectral Wavelet Convolution via Chebyshev Order Decomposition · ICML 2025
Machine learning › Representation and self-supervised learning › representation learning › embedding learning
unified embedding
0.912025
UniGraph2: Learning a Unified Embedding Space to Bind Multimodal Graphs · WWW 2025
Security and privacy of machine learning
adversarial attack
0.912025
FlipAttack: Jailbreak LLMs via Flipping · ICML 2025
Security and privacy of machine learning › adversarial attack
jailbreak attack
0.912025
FlipAttack: Jailbreak LLMs via Flipping · ICML 2025
Machine learning › Graph learning
graph clustering
0.812024
ProCom: A Few-shot Targeted Community Detection Algorithm · KDD 2024
Machine learning › Graph learning › graph neural network › graph neural network generalization
graph few-shot learning
0.812024
ProCom: A Few-shot Targeted Community Detection Algorithm · KDD 2024
Natural language and speech › Question answering and dialogue systems
knowledge base question answering
0.812024
G-Retriever: Retrieval-Augmented Generation for Textual Graph Understanding and Question Answering · NeurIPS 2024
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge graph reasoning
0.812024
G-Retriever: Retrieval-Augmented Generation for Textual Graph Understanding and Question Answering · NeurIPS 2024
Natural language and speech › Language models and text generation
retrieval-augmented generation
0.812024
G-Retriever: Retrieval-Augmented Generation for Textual Graph Understanding and Question Answering · NeurIPS 2024
Machine learning › Graph learning › graph representation learning
text-attributed graph learning
0.812024
Harnessing Explanations: LLM-to-LM Interpreter for Enhanced Text-Attributed Graph Representation Learning · ICLR 2024
Machine learning › Graph learning
graph representation learning
0.712023
A Generalization of ViT/MLP-Mixer to Graphs · ICML 2023
Machine learning › Graph learning › graph neural network
graph transformer
0.712023
A Generalization of ViT/MLP-Mixer to Graphs · ICML 2023
Machine learning › Deep learning architectures and training › sequence modeling
long-range dependency modeling
0.712023
A Generalization of ViT/MLP-Mixer to Graphs · ICML 2023
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
masked autoencoder
0.712023
A Study on Transformer Configuration and Training Objective · ICML 2023
Machine learning › Deep learning architectures and training
transformer
0.712023
A Study on Transformer Configuration and Training Objective · ICML 2023
Machine learning › Efficient and distributed learning
federated learning
0.412020
Federated learning on wearable devices: demo abstract · SenSys 2020
Computer vision › Video understanding and tracking › activity recognition
human activity recognition
0.412020
Federated learning on wearable devices: demo abstract · SenSys 2020
Machine learning › Trustworthy machine learning
robustness
0.312025
FlipAttack: Jailbreak LLMs via Flipping · ICML 2025
Machine learning › Deep learning architectures and training › feedforward neural network › MLP-based architecture
MLP-Mixer
0.212023
A Generalization of ViT/MLP-Mixer to Graphs · ICML 2023
Machine learning › Deep learning architectures and training › transformer
vision transformer
0.212023
A Generalization of ViT/MLP-Mixer to Graphs · ICML 2023
Ubiquitous computing and smart environments › context recognition
activity recognition
0.112020
Federated learning on wearable devices: demo abstract · SenSys 2020

Methods — techniques the papers use, named apart from their topics

graph neural network · 3.3text flipping · 1.7left-side perturbation · 1.7large language model · 1.5wavelet admissibility · 0.9multiresolution analysis · 0.9masked graph modeling · 0.9language model · 0.9instruction tuning · 0.9chebyshev polynomials · 0.9federated learning · 0.4
YearPublicationVenuePosition
2025 A General Graph Spectral Wavelet Convolution via Chebyshev Order Decomposition
abstract
Spectral graph convolution, an important tool of data filtering on graphs, relies on two essential decisions: selecting spectral bases for signal transformation and parameterizing the kernel for frequency analysis. While recent techniques mainly focus on standard Fourier transform and vector-valued spectral functions, they fall short in flexibility to model signal distributions over large spatial ranges, and capacity of spectral function. In this paper, we present a novel wavelet-based graph convolution network, namely WaveGC, which integrates multi-resolution spectral bases and a matrix-valued filter kernel. Theoretically, we establish that WaveGC can effectively capture and decouple short-range and long-range information, providing superior filtering flexibility, surpassing existing graph wavelet neural networks. To instantiate WaveGC, we introduce a novel technique for learning general graph wavelets by separately combining odd and even terms of Chebyshev polynomials. This approach strictly satisfies wavelet admissibility criteria. Our numerical experiments showcase the consistent improvements in both short-range and long-range tasks. This underscores the effectiveness of the proposed model in handling different scenarios.
Nian Liu 0001, Xiao-Xin He, Thomas Laurent 0001, Francesco Di Giovanni, Michael M. Bronstein, Xavier Bresson
ICML2
2025 FlipAttack: Jailbreak LLMs via Flipping
abstract
This paper proposes a simple yet effective jailbreak attack named FlipAttack against black-box LLMs. First, from the autoregressive nature, we reveal that LLMs tend to understand the text from left to right and find that they struggle to comprehend the text when the perturbation is added to the left side. Motivated by these insights, we propose to disguise the harmful prompt by constructing a left-side perturbation merely based on the prompt itself, then generalize this idea to 4 flipping modes. Second, we verify the strong ability of LLMs to perform the text-flipping task and then develop 4 variants to guide LLMs to understand and execute harmful behaviors accurately. These designs keep FlipAttack universal, stealthy, and simple, allowing it to jailbreak black-box LLMs within only 1 query. Experiments on 8 LLMs demonstrate the superiority of FlipAttack. Remarkably, it achieves $\sim$78.97% attack success rate across 8 LLMs on average and $\sim$98% bypass rate against 5 guard models on average.
Yue Liu 0008, Xiao-Xin He, Miao Xiong, Jinlan Fu, Shumin Deng, Yingwei Ma, Jiaheng Zhang, Bryan Hooi
ICML2
2025 UniGraph: Learning a Unified Cross-Domain Foundation Model for Text-Attributed Graphs
abstract
Foundation models like ChatGPT and GPT-4 have revolutionized artificial intelligence, exhibiting remarkable abilities to generalize across a wide array of tasks and applications beyond their initial training objectives. However, graph learning has predominantly focused on single-graph models, tailored to specific tasks or datasets, lacking the ability to transfer learned knowledge to different domains. This limitation stems from the inherent complexity and diversity of graph structures, along with the different feature and label spaces specific to graph data. In this paper, we recognize text as an effective unifying medium and employ Text-Attributed Graphs (TAGs) to leverage this potential. We present our UniGraph framework, designed to learn a foundation model for TAGs, which is capable of generalizing to unseen graphs and tasks across diverse domains. Unlike single-graph models that use pre-computed node features of varying dimensions as input, our approach leverages textual features for unifying node representations, even for graphs such as molecular graphs that do not naturally have textual features. We propose a novel cascaded architecture of Language Models (LMs) and Graph Neural Networks (GNNs) as backbone networks. Additionally, we propose the first pre-training algorithm specifically designed for large-scale self-supervised learning on TAGs, based on Masked Graph Modeling. We introduce graph instruction tuning using Large Language Models (LLMs) to enable zero-shot prediction ability. Our comprehensive experiments across various graph learning tasks and domains demonstrate the model's effectiveness in self-supervised representation learning on unseen graphs, few-shot in-context transfer, and zero-shot transfer, even surpassing or matching the performance of GNNs that have undergone supervised training on target datasets.
Yuan Sui 0001, Xiao-Xin He, Bryan Hooi
KDD (1)3
2025 UniGraph2: Learning a Unified Embedding Space to Bind Multimodal Graphs
abstract
Existing foundation models, such as CLIP, aim to learn a unified embedding space for multimodal data, enabling a wide range of downstream web-based applications like search, recommendation, and content classification. However, these models often overlook the inherent graph structures in multimodal datasets, where entities and their relationships are crucial. Multimodal graphs (MMGs) represent such graphs where each node is associated with features from different modalities, while the edges capture the relationships between these entities.On the other hand, existing graph foundation models primarily focus on text-attributed graphs (TAGs) and are not designed to handle the complexities of MMGs. To address these limitations, we propose UniGraph2, a novel cross-domain graph foundation model that enables general representation learning on MMGs, providing a unified embedding space. UniGraph2 employs modality-specific encoders alongside a graph neural network (GNN) to learn a unified low-dimensional embedding space that captures both the multimodal information and the underlying graph structure. We propose a new cross-domain multi-graph pre-training algorithm at scale to ensure effective transfer learning across diverse graph domains and modalities. Additionally, we adopt a Mixture of Experts (MoE) component to align features from different domains and modalities, ensuring coherent and robust embeddings that unify the information across modalities. Extensive experiments on a variety of multimodal graph tasks demonstrate that UniGraph2 significantly outperforms state-of-the-art models in tasks such as representation learning, transfer learning, and multimodal generative tasks, offering a scalable and flexible solution for learning on MMGs.
Yuan Sui 0001, Xiao-Xin He, Yue Liu 0008, Yifei Sun 0002, Bryan Hooi
WWW3
2024 Harnessing Explanations: LLM-to-LM Interpreter for Enhanced Text-Attributed Graph Representation Learning
abstract
Representation learning on text-attributed graphs (TAGs) has become a critical research problem in recent years. A typical example of a TAG is a paper citation graph, where the text of each paper serves as node attributes. Initial graph neural network (GNN) pipelines handled these text attributes by transforming them into shallow or hand-crafted features, such as skip-gram or bag-of-words features. Recent efforts have focused on enhancing these pipelines with language models (LMs), which typically demand intricate designs and substantial computational resources. With the advent of powerful large language models (LLMs) such as GPT or Llama2, which demonstrate an ability to reason and to utilize general knowledge, there is a growing need for techniques which combine the textual modelling abilities of LLMs with the structural learning capabilities of GNNs. Hence, in this work, we focus on leveraging LLMs to capture textual information as features, which can be used to boost GNN performance on downstream tasks. A key innovation is our use of \emph{explanations as features}: we prompt an LLM to perform zero-shot classification, request textual explanations for its decision-making process, and design an \emph{LLM-to-LM interpreter} to translate these explanations into informative features for downstream GNNs. Our experiments demonstrate that our method achieves state-of-the-art results on well-established TAG datasets, including \texttt{Cora}, \texttt{PubMed}, \texttt{ogbn-arxiv}, as well as our newly introduced dataset, \texttt{tape-arxiv23}. Furthermore, our method significantly speeds up training, achieving a 2.88 times improvement over the closest baseline on \texttt{ogbn-arxiv}. Lastly, we believe the versatility of the proposed method extends beyond TAGs and holds the potential to enhance other tasks involving graph-text data~\footnote{Our codes and datasets are available at: \url{https://github.com/XiaoxinHe/TAPE}}.
Xiao-Xin He, Xavier Bresson, Thomas Laurent 0001, Adam Perold, Yann LeCun, Bryan Hooi
ICLR1
2024 ProCom: A Few-shot Targeted Community Detection Algorithm
abstract
Targeted community detection aims to distinguish a particular type of community in the network. This is an important task with a lot of real-world applications, e.g., identifying fraud groups in transaction networks. Traditional community detection methods fail to capture the specific features of the targeted community and detect all types of communities indiscriminately. Semi-supervised community detection algorithms, emerged as a feasible alternative, are inherently constrained by their limited adaptability and substantial reliance on a large amount of labeled data, which demands extensive domain knowledge and manual effort.
Xixi Wu, Kaiyu Xiong, Yun Xiong, Xiao-Xin He, Yao Zhang 0009, Yizhu Jiao, Jiawei Zhang 0001
KDD4
2024 G-Retriever: Retrieval-Augmented Generation for Textual Graph Understanding and Question Answering
abstract
Given a graph with textual attributes, we enable users to `chat with their graph': that is, to ask questions about the graph using a conversational interface. In response to a user's questions, our method provides textual replies and highlights the relevant parts of the graph. While existing works integrate large language models (LLMs) and graph neural networks (GNNs) in various ways, they mostly focus on either conventional graph tasks (such as node, edge, and graph classification), or on answering simple graph queries on small or synthetic graphs. In contrast, we develop a flexible question-answering framework targeting real-world textual graphs, applicable to multiple applications including scene graph understanding, common sense reasoning, and knowledge graph reasoning. Toward this goal, we first develop a Graph Question Answering (GraphQA) benchmark with data collected from different tasks. Then, we propose our \textit{G-Retriever} method, introducing the first retrieval-augmented generation (RAG) approach for general textual graphs, which can be fine-tuned to enhance graph understanding via soft prompting. To resist hallucination and to allow for textual graphs that greatly exceed the LLM's context window size, \textit{G-Retriever} performs RAG over a graph by formulating this task as a Prize-Collecting Steiner Tree optimization problem. Empirical evaluations show that our method outperforms baselines on textual graph tasks from multiple domains, scales well with larger graph sizes, and mitigates hallucination.~\footnote{Our codes and datasets are available at: \url{https://github.com/XiaoxinHe/G-Retriever}}
Xiao-Xin He, Yijun Tian 0001, Yifei Sun 0002, Nitesh V. Chawla, Thomas Laurent 0001, Yann LeCun, Xavier Bresson, Bryan Hooi
NeurIPS1
2023 A Generalization of ViT/MLP-Mixer to Graphs
abstract
Graph Neural Networks (GNNs) have shown great potential in the field of graph representation learning. Standard GNNs define a local message-passing mechanism which propagates information over the whole graph domain by stacking multiple layers. This paradigm suffers from two major limitations, over-squashing and poor long-range dependencies, that can be solved using global attention but significantly increases the computational cost to quadratic complexity. In this work, we propose an alternative approach to overcome these structural limitations by leveraging the ViT/MLP-Mixer architectures introduced in computer vision. We introduce a new class of GNNs, called Graph ViT/MLP-Mixer, that holds three key properties. First, they capture long-range dependency and mitigate the issue of over-squashing as demonstrated on Long Range Graph Benchmark and TreeNeighbourMatch datasets. Second, they offer better speed and memory efficiency with a complexity linear to the number of nodes and edges, surpassing the related Graph Transformer and expressive GNN models. Third, they show high expressivity in terms of graph isomorphism as they can distinguish at least 3-WL non-isomorphic graphs. We test our architecture on 4 simulated datasets and 7 real-world benchmarks, and show highly competitive results on all of them. The source code is available for reproducibility at: https://github.com/XiaoxinHe/Graph-ViT-MLPMixer.
Xiao-Xin He, Bryan Hooi, Thomas Laurent 0001, Adam Perold, Yann LeCun, Xavier Bresson
ICML1
2023 A Study on Transformer Configuration and Training Objective
abstract
Transformer-based models have delivered impressive results on many tasks, particularly vision and language tasks. In many model training situations, conventional configurations are often adopted. For example, we usually set the base model with hidden size (i.e. model width) to be 768 and the number of transformer layers (i.e. model depth) to be 12. In this paper, we revisit these conventional configurations by studying the the relationship between transformer configuration and training objective. We show that the optimal transformer configuration is closely related to the training objective. Specifically, compared with the simple classification objective, the masked autoencoder is effective in alleviating the over-smoothing issue in deep transformer training. Based on this finding, we propose “Bamboo”, a notion of using deeper and narrower transformer configurations, for masked autoencoder training. On ImageNet, with such a simple change in configuration, the re-designed Base-level transformer achieves 84.2% top-1 accuracy and outperforms SoTA models like MAE by $0.9%$. On language tasks, re-designed model outperforms BERT with the default setting by 1.1 points on average, on GLUE benchmark with 8 datasets.
Fuzhao Xue, Jianghai Chen, Aixin Sun, Xiaozhe Ren, Zangwei Zheng, Xiao-Xin He, Yongming Chen, Xin Jiang 0002, Yang You 0001
ICML6
2020 Federated learning on wearable devices: demo abstract
abstract
Wearable devices collect user information about their activities and provide insights to improve their daily lifestyles. Smart health applications have achieved great success by training Machine Learning (ML) models on a large quantity of user data from wearables. However, user privacy and scalability are becoming critical challenges for training ML models in a centralized way. Federated learning (FL) is a novel ML paradigm with the goal of training high quality models while distributing training data over a large number of devices. In this demo, we present FL4W, a FL system with wearable devices enabling training a human activity recognition classifier. We also perform preliminary analytics to investigate the model performance with increasing computation of clients.
Xiao-Xin He, Xiang Su 0001, Yang Chen 0001, Pan Hui 0001
SenSys1
2010 Tree-Based Service Discovery in Mobile Ad Hoc Networks
abstract
Utilizing service oriented architectures to enhance seamless collaborations and information sharing among nodes in mobile ad hoc networks with limited communication capability is a challenging task. In this paper a novel tree-based service discovery mechanism, able to achieve high accuracy in the process of service discovery, is proposed, which affords wireless communication overheads nearly linear with the number of property-changing services offered by the whole network. The mechanism is suitable and customized for typical mobile networks. As properties such as quality and capability of the aimed services are always changing rather than remaining invariable, a service matching process is also provided, where an efficient method is implemented for searching for a target service in a local service repository with given conditions and policies.
Mingxue Liao, He Jing, Zhu Rongfu, Wang Xianqing, Xiao-Xin He
APSCC5