Zhongtian Sun

dblp:266/3353 · DBLP profile ↗
← Back
14ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0003-0489-5203ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Heterophily-Agnostic Hypergraph Neural Networks with Riemannian Local Exchanger
abstract
Hypergraphs are the natural description of higher-order interactions among objects, widely applied in social network analysis, cross-modal retrieval, etc. Hypergraph Neural Networks (HGNNs) have become the dominant solution for learning on hypergraphs. Traditional HGNNs are extended from message passing graph neural networks, following the homophily assumption, and thus struggle with the prevalent heterophilic hypergraphs that call for long-range dependence modeling. Existing solutions enlarge the message flow through the hypergraph bottleneck, mitigating the oversquashing issue and capturing long-range dependence. However, they often accelerate the loss of representation distinguishability in the repeated aggregations, leading to oversmoothing. This dilemma motivates an interesting question: Can we develop a unified mechanism that is agnostic to both homophilic and heterophilic hypergraphs? In this paper, we achieve the best of both worlds through the lens of Riemannian geometry, which provides the potential to adjust the message passing behavior in different regions. The key insight lies in the connection between oversquashing and hypergraph bottleneck within the framework of Riemannian manifold heat flow. Building on this, we propose the novel idea of locally adapting the bottlenecks of different subhypergraphs. The core innovation of the proposed mechanism is the design of an adaptive local (heat) exchanger. Specifically, it captures the rich long-range dependencies via the Robin condition, and preserves the representation distinguishability via source terms, thereby enabling heterophily-agnostic message passing with theoretical guarantees. Based on this theoretical foundation, we present a novel Heat-Exchanger with Adaptive Locality for Hypergraph Neural Network (HealHGNN), designed as a node-hyperedge bidirectional systems with linear complexity in the number of nodes and hyperedges. Extensive experiments on both homophilic and heterophilic cases show that HealHGNN achieves the state-of-the-art performance.
Li Sun 0008, Ming Zhang 0034, Wenxin Jin, Zhongtian Sun, Zhenhao Huang 0001, Hao Peng 0001, Sen Su, Philip S. Yu
WWW4
2026 Relation Extraction from the Perspective of the Frequency Domain: Frequency-Domain Aware Gated Graph Attention Network
abstract
In relation extraction task, graph attention network, as the dominant model, often faces the challenge of attention bias caused by complex semantic environment. Existing approaches ignore decoupling and build fine-grained models to filter and directly interact the multiple levels of information overlapping in word vectors (word, phrase, clause, and sentence level), but using methods that focus only on context (such as additional knowledge or structure). Ignoring overlapping multilevel information leads to limited performance improvement of the model for attention bias, but also increases the processing cost. To overcome this core limitation, we propose a model to decouple and process multiple levels of semantic information from the spectrum domain: Frequency-domain aware Gated Graph Attention Network (FD-GGAN-RE). The network first uses spectral decomposition to decouple contextual word vectors into spectral domain vectors containing different levels of semantic information. Then, use Frequency Feature Selective Gate layer to realize adaptive semantic filtering, reducing the influence of irrelevant semantics on the subsequent graph attention calculation. Final the Frequency-domain graph attention layer realizes the direct interaction of multiple levels of semantic information in the spectrum domain, avoiding the attention bias caused by the context graph attention mechanism interacting with word vectors containing multiple levels of semantic overlap. SemEval and KBP37 scored 90.33 and 69.06 respectively for F1, which was 27% faster than GATs while F1 scored 0.15 and 0.84 higher, respectively. Frequency graph attention visualization further demonstrates the model’s capability to capture complex key semantics, while presenting a frequency-based approach that holds potential for application in other natural language processing tasks.
Zhan'ao Yao, Wenli Geng, Zhongtian Sun
Neural Process. Lett.4
2025 RicciFlowRec: A Geometric Root Cause Recommender Using Ricci Curvature on Financial Graphs
abstract
We propose RicciFlowRec, a geometric recommendation framework that performs root cause attribution via Ricci curvature and flow on dynamic financial graphs.By modelling evolving interactions among stocks, macroeconomic indicators, and news, we quantify local stress using discrete Ricci curvature and trace shock propagation via Ricci flow.Curvature gradients reveal causal substructures, informing a structural risk-aware ranking function.Preliminary results on S&P 500 data with FinBERT-based sentiment show improved robustness and interpretability under synthetic perturbations.This ongoing work supports curvature-based attribution and early-stage risk-aware ranking, with plans for portfolio optimization and return forecasting.To our knowledge, RicciFlowRec is the first recommender to apply geometric flow-based reasoning in financial decision support.
Zhongtian Sun, Anoushka Harit
RecSys1
2025 Three trustworthiness challenges in large language model-based financial systems: real-world examples and mitigation strategies
abstract
大语言模型 (LLM) 在金融应用中的集成展现出显著潜力, 可提升决策流程、实现操作自动化并提供个性化服务。 然而, 金融系统的高风险特性要求极高的可信度, 而当前LLM往往难以满足这一要求。 本研究识别并探讨了基于LLM的金融系统中的3大可信度挑战: (1) 逃逸式提示——利用模型对齐漏洞生成有害或违规响应; (2) 幻觉现象——模型产出事实错误的输出误导金融决策; (3) 偏见与公平性问题——LLM内嵌的人口统计或制度偏见可能导致个体或区域遭受不公平对待。 为具体呈现这些风险, 我们设计了3项金融相关测试, 并对涵盖专有与开源家族的主流LLM进行评估。 在所有模型中, 每项测试至少出现一次风险行为。 基于这些发现, 系统性地总结了现有风险缓解策略。 我们认为, 解决这些问题不仅对确保金融领域人工智能的负责任使用至关重要, 更是实现其安全可扩展部署的关键所在。
Shurui Xu, Shuyan Li, Mengzhen Fan, Zhongtian Sun
Frontiers Inf. Technol. Electron. Eng.5
2023 A Rewiring Contrastive Patch PerformerMixer Framework for Graph Representation Learning
abstract
Integrating transformers with graph representation learning has emerged as a research focal point. However, recent studies showed that positional encoding in Transformers does not capture enough structural information between nodes. Additionally, existing graph neural network (GNN) models face the oversquashing issue, impeding information retention from distant nodes. To address, we transform graphs into regular structures, such as tokens, to enhance positional understanding and leverage transformer strengths. Inspired by the visual transformer (ViT) model, we propose partitioning graphs into patches and apply GNN models obtain fixed size vectors. Notably, our approach adopts contrastive learning for in-depth graph structure and incorporate more topological information via Ricci curvature to alleviate over-squashing problem by attenuating the effects of negatively curved edges while preserving the original graph structure. Unlike existing graph rewiring methods that directly modify graph structure by adding or removing edges, this approach is potentially more suitable for applications such as molecular learning where structural preservation is important. Our innovative pipeline subsequently introduces the PerformerMixer, a transformer variant with linear complexity, ensuring efficient computation. Evaluations on real-world benchmarks demonstrate our framework’s superior performance, like Peptides-func and achieve 3-WL expressiveness.
Zhongtian Sun, Anoushka Harit, Alexandra I. Cristea, Jingyun Wang 0003, Pietro Liò
IEEE Big Data1
2022 Is Unimodal Bias Always Bad for Visual Question Answering? A Medical Domain Study with Dynamic Attention
abstract
Medical visual question answering (Med-VQA) is to answer medical questions based on clinical images provided. This field is still in its infancy due to the complexity of the trio formed of questions, multimodal features and expert knowledge. In this paper, we tackle, a ’myth’ in the Natural Language Processing area - that unimodal bias is always considered undesirable in learning models. Additionally, we study the effect of integrating a novel dynamic attention mechanism into such models, inspired by a recent graph deep learning study.Unlike traditional attention, dynamic attention scores are conditioned on different query words in a question and thus enhance the representation learning ability of texts. We propose that some questions are answered more accurately with a reinforcement of question embedding after fusing multimodal features. Extensive experiments have been implemented on the VQA-RAD datasets and demonstrate that our proposed model, reinforCe unimOdal dynamiC Attention (COCA), outperforms the state-of-the-art methods overall and performs competitively at open-ended question answering.
Zhongtian Sun, Anoushka Harit, Alexandra I. Cristea, Jialin Yu 0001, Noura Al Moubayed, Lei Shi 0003
IEEE Big Data1
2022 Contrastive Learning with Heterogeneous Graph Attention Networks on Short Text Classification
abstract
Graph neural networks (GNNs) have attracted extensive interest in text classification tasks due to their expected superior performance in representation learning. However, most existing studies adopted the same semi-supervised learning setting as the vanilla Graph Convolution Network (GCN), which requires a large amount of labelled data during training and thus is less robust when dealing with large-scale graph data with fewer labels. Additionally, graph structure information is normally captured by direct information aggregation via network schema and is highly dependent on correct adjacency information. Therefore, any missing adjacency knowledge may hinder the performance. Addressing these problems, this paper thus proposes a novel method to learn a graph structure, NC-HGAT, by expanding a state-of-the-art self-supervised heterogeneous graph neural network model (HGAT) with simple neighbour contrastive learning. The new NC-HGAT considers the graph structure information from heterogeneous graphs with multilayer perceptrons (MLPs) and delivers consistent results, despite the corrupted neighbouring connections. Extensive experiments have been implemented on four benchmark short-text datasets. The results demonstrate that our proposed model NC-HGAT significantly outperforms state-of-the-art methods on three datasets and achieves competitive performance on the remaining dataset.
Zhongtian Sun, Anoushka Harit, Alexandra I. Cristea, Jialin Yu 0001, Lei Shi 0003, Noura Al Moubayed
IJCNN1
2022 INTERACTION: A Generative XAI Framework for Natural Language Inference Explanations
abstract
XAI with natural language processing aims to produce human-readable explanations as evidence for AI decision-making, which addresses explainability and transparency. However, from an HCI perspective, the current approaches only focus on delivering a single explanation, which fails to account for the diversity of human thoughts and experiences in language. This paper thus addresses this gap, by proposing a generative XAI framework, INTERACTION (explain aNd predicT thEn queRy with contextuAl CondiTional varIational autO-eNcoder). Our novel framework presents explanation in two steps: (step one) Explanation and Label Prediction; and (step two) Diverse Evidence Generation. We conduct intensive experiments with the Transformer architecture on a benchmark dataset, e-SNLI [1]. Our method achieves competitive or better performance against state-of-the-art baseline models on explanation generation (up to 4.7% gain in BLEU) and prediction (up to 4.4% gain in accuracy) in step one; it can also generate multiple diverse explanations in step two.
Jialin Yu 0001, Alexandra I. Cristea, Anoushka Harit, Zhongtian Sun, Olanrewaju Tahir Aduragba, Lei Shi 0003, Noura Al Moubayed
IJCNN4
2022 Efficient Uncertainty Quantification for Multilabel Text Classification
abstract
Despite rapid advances of modern artificial intelligence (AI), there is a growing concern regarding its capacity to be explainable, transparent, and accountable. One crucial step towards such AI systems involves reliable and efficient uncertainty quantification methods. Existing approaches to uncertainty quantification in natural language processing (NLP) take a Bayesian Deep Learning approach. However, the latter is known to not be computationally efficient in testing time, thus hindering its applicability in real-life scenarios. This paper proposes a new focus on the efficiency of uncertainty quantification methods, evaluating them on four multi-label text classification tasks. Our novel methods of representing epistemic and aleatoric uncertainties enable efficient uncertainty quantification (around 13 to 45 times faster than existing approaches, depending on architecture) with posterior analysis in the (approximated) latent- and data space. We conduct extensive experiments and studies on diverse neural network architectures (LSTM, CNN and Transformer) to analyse their power. Our results prove the benefits of explicitly modelling uncertainty in neural networks.
Jialin Yu 0001, Alexandra I. Cristea, Anoushka Harit, Zhongtian Sun, Olanrewaju Tahir Aduragba, Lei Shi 0003, Noura Al Moubayed
IJCNN4
2021 A Generative Bayesian Graph Attention Network for Semi-Supervised Classification on Scarce Data
abstract
This research focuses on semi-supervised classification tasks, specifically for graph-structured data under data-scarce situations. It is known that the performance of conventional supervised graph convolutional models is mediocre at classification tasks, when only a small fraction of the labeled nodes are given. Additionally, most existing graph neural network models often ignore the noise in graph generation and consider all the relations between objects as genuine ground-truth. Hence, the missing edges may not be considered, while other spurious edges are included. Addressing those challenges, we propose a Bayesian Graph Attention model which utilizes a generative model to randomly generate the observed graph. The method infers the joint posterior distribution of node labels and graph structure, by combining the Mixed-Membership Stochastic Block Model with the Graph Attention Model. We adopt a variety of approximation methods to estimate the Bayesian posterior distribution of the missing labels. The proposed method is comprehensively evaluated on three graph-based deep learning benchmark data sets. The experimental results demonstrate a competitive performance of our proposed model BGAT against the current state of the art models when there are few labels available (the highest improvement is 5%), for semi-supervised node classification tasks.
Zhongtian Sun, Anoushka Harit, Jialin Yu 0001, Alexandra I. Cristea, Noura Al Moubayed
IJCNN1
2021 MOOC Next Week Dropout Prediction: Weekly Assessing Time and Learning Patterns
Ahmed Alamri, Zhongtian Sun, Alexandra I. Cristea, Craig D. Stewart, Filipe D. Pereira
ITS2
2021 A Brief Survey of Deep Learning Approaches for Learning Analytics on MOOCs
Zhongtian Sun, Anoushka Harit, Jialin Yu 0001, Alexandra I. Cristea, Lei Shi 0003
ITS1
2021 Exploring Bayesian Deep Learning for Urgent Instructor Intervention Need in MOOC Forums
Jialin Yu 0001, Laila Alrajhi, Anoushka Harit, Zhongtian Sun, Alexandra I. Cristea, Lei Shi 0003
ITS4
2020 Is MOOC Learning Different for Dropouts? A Visually-Driven, Multi-granularity Explanatory ML Approach
Ahmed Alamri, Zhongtian Sun, Alexandra I. Cristea, Gautham Senthilnathan, Lei Shi 0003, Craig D. Stewart
ITS2