VLDB 2026 Research / reviewers in the wild / expert
Xiaozhuang Song
dblp:283/0298
· DBLP profile ↗
15ranked-venue papers
6as first author
15since 2021 · last 2026
0000-0002-7861-8957ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 5 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Physically-Informed Flow Matching with Graph Neural Networks for Complex Fluid DynamicsabstractComputational fluid dynamics (CFD) simulations traditionally require extensive computational resources, limiting their utility in many scientific and engineering applications at scale. We introduce Physically-Informed Flow Matching Graph Networks (PIFM-GN), a novel generative framework that directly samples fluid states under specified physical conditions without requiring expensive time-stepping simulations. The key innovation of our approach is the incorporation of incompressibility constraints directly into the flow matching transport process by parameterizing velocity fields through vector potentials, with graph-based curl operators ensuring divergence-free predictions without requiring global pressure-Poisson solves. Experiments on diverse fluid dynamics problems -- ranging from two-dimensional surface pressure distributions and complete flow fields, to complex three-dimensional airflow fields -- demonstrate that PIFM-GN generates high-fidelity samples with significantly fewer sampling steps than diffusion-based alternatives. Most notably, our model maintains competitive performance even with a single sampling step, a regime where diffusion models completely fail. Our generated samples accurately reproduce the statistical characteristics of target flows, successfully capturing multi-modal pressure distributions across various flow conditions, while achieving significant computational speedups compared to diffusion-based methods. PIFM-GN thus enables efficient generation of fluid states for downstream analysis and design tasks in scientific and engineering applications. Xiaozhuang Song, Tianshu Yu 0001 |
AAAI | 1 |
| 2026 | Deep Tabular Representation CorrectorabstractTabular data have been playing a mostly important role in diverse real-world fields, such as healthcare, engineering, finance, etc. The recent success of deep learning has fostered many deep networks (e.g., Transformer, ResNet) based tabular learning methods. Generally, existing deep tabular machine learning methods are along with the two paradigms, i.e., in-learning and pre-learning. In-learning methods need to train networks from scratch or impose extra constraints to regulate the representations which nonetheless train multiple tasks simultaneously and make learning more difficult, while pre-learning methods design several pretext tasks for pre-training and then conduct task-specific fine-tuning, which however need much extra training effort with prior knowledge. In this paper, we introduce a novel deep Tabular Representation Corrector, TRC, to enhance any trained deep tabular model's representations without altering its parameters in a model-agnostic manner. Specifically, targeting the representation shift and representation redundancy that hinder prediction, we propose two tasks, i.e., (i) Tabular Representation Re-estimation, that involves training a shift estimator to calculate the inherent shift of tabular representations to subsequently mitigate it, thereby re-estimating the representations and (ii) Tabular Space Mapping, that transforms the above re-estimated representations into a light-embedding vector space via a coordinate estimator while preserves crucial predictive information to minimize redundancy. The two tasks jointly enhance the representations of deep tabular models without touching on the original models thus enjoying high efficiency. Finally, we conduct extensive experiments on state-of-the-art deep tabular machine learning models coupled with TRC on various tabular benchmarks which have shown consistent superiority. Hangting Ye, Wei Fan 0010, Xiaozhuang Song, He Zhao 0001, Dandan Guo, Yi Chang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Enhancing Generalizability in Molecular Conformation Generation with METRIZATION-Informed Geometric Diffusion PretrainingabstractDiffusion-based generative models have recently excelled in generating molecular conformations but struggled with the generalization issue -- models trained on one dataset may produce meaningless conformations on out-of-distribution molecules. On the other hand, distance geometry serves as a generalizable tool for the traditional computational chemistry methods of molecular conformation, which is predicated on the assumption that it is possible to adequately define the set of all potential conformations of any non-rigid molecular system using purely geometric constraints. In this work, we for the first time explicitly incorporate distance geometry constraints into pretraining phase of diffusion-based molecular generation models to improve the generalizability. Inspired by the classical distance geometry solution designed for solving the molecular distance geometry problem, we propose MiGDiff, a Metrization-Informed Geometric Diffusion framework. MiGDiff injects distance geometry constraints by pretraining the deep geometric diffusion backbone within the Metrization sampling approach, yielding a "Metrization-driven pretraining + Data-driven finetuning" paradigm. Experimental results demonstrate that MiGDiff outperforms state-of-the-art methods and possesses strong generalization capabilities, particularly on generating previously unseen molecules, revealing the vast untapped potential of combining traditional computational methods with deep generative models for 3D molecular generation. Xiaozhuang Song, Yuzhao Tu, Hangting Ye, Wei Fan 0010, Tianshu Yu 0001 |
AAAI | 1 |
| 2025 | The Underappreciated Power of Vision Models for Graph Structural UnderstandingabstractGraph Neural Networks operate through bottom-up message-passing, fundamentally differing from human visual perception, which intuitively captures global structures first. We investigate the underappreciated potential of vision models for graph understanding, finding they achieve performance comparable to GNNs on established benchmarks while exhibiting distinctly different learning patterns.
These divergent behaviors, combined with limitations of existing benchmarks that conflate domain features with topological understanding, motivate our introduction of GraphAbstract. This benchmark evaluates models' ability to perceive global graph properties as humans do: recognizing organizational archetypes, detecting symmetry, sensing connectivity strength, and identifying critical elements. Our results reveal that vision models significantly outperform GNNs on tasks requiring holistic structural understanding and
maintain generalizability across varying graph scales, while GNNs struggle with global pattern abstraction and degrade with increasing graph size. This work demonstrates that vision models possess remarkable yet underutilized capabilities for graph structural understanding, particularly for problems requiring global topological awareness and scale-invariant reasoning. These findings open new avenues to leverage this underappreciated potential for developing more effective graph foundation models for tasks dominated by holistic pattern recognition. Xinjian Zhao, Zhongkai Xue, Xiangru Jian, Yaoyao Xu, Xiaozhuang Song, Tianshu Yu 0001 |
NeurIPS | 7 |
| 2025 | Towards Multi-resolution Spatiotemporal Graph Learning for Medical Time Series ClassificationabstractMedical time series has been playing a vital role in real-world healthcare systems as valuable information in monitoring health conditions of patients. Traditional methods towards medical time series classification rely on handcrafted feature extraction and statistical methods; with the recent advancement of artificial intelligence, the machine learning and deep learning methods have become more popular. However, existing methods often fail to fully model the complex spatial dynamics under different scales, which ignore the dynamic multi-resolution spatial and temporal joint inter-dependencies. Moreover, they are less likely to consider the special baseline wander problem as well as the multi-view characteristics of medical time series, which largely hinders their prediction performance. To address these limitations, we propose a Multi-resolution Spatiotemporal Graph Learning framework, MedGNN, for medical time series classification. Specifically, we first propose to construct multi-resolution adaptive graph structures to learn dynamic multi-scale embeddings. Then, to address the baseline wander problem, we propose Difference Attention Networks to operate self-attention mechanisms on the finite difference for temporal modeling. Moreover, to learn the multi-view characteristics, we utilize the Frequency Convolution Networks to capture complementary information of medical time series from the frequency domain. In addition, we introduce the Multi-resolution Graph Transformer architecture to model the dynamic dependencies and fuse the information from different resolutions. Finally, we have conducted extensive experiments on multiple medical real-world datasets that demonstrate the superior performance of our method. Our Code is available at this repository: https://github.com/aikunyi/MedGNN. Wei Fan 0010, Jingru Fei, Dingyu Guo, Kun Yi 0001, Xiaozhuang Song, Haolong Xiang, Hangting Ye, Min Li 0007 |
WWW | 5 |
| 2024 | Single Cell Gene Expression Prediction via Prototype-based Proximal Neural FactorizationabstractIn the realm of single-cell analysis, accurately predicting gene expressions is crucial for understanding cellular functions and interactions. Traditional approaches often face significant challenges due to intrinsic noise, high dimensionality, and limited data availability in single-cell datasets. On the other hand, deep learning methods are prone to overfitting and perform poorly with limited data. This paper introduces a novel framework, Prototype-based Proximal Neural Factorization (PPNF), which harnesses the power of prototype learning and neural factorization to address these issues. Our method leverages a robust learning paradigm that identifies representative prototypes from single-cell data, facilitating a more resilient and interpretable data representations. We validate our approach using a diverse set of single-cell datasets, demonstrating that our method significantly outperforms existing techniques in terms of both robustness and accuracy. PPNF shows its effectiveness even with limited data, thereby reducing the financial and computational burden associated with high-throughput technologies. By enhancing the robustness and generalizability of single-cell gene expression predictions, our framework provides significant benefits for advancing the analysis and interpretation of single-cell gene expression data, particularly in data-limited scenarios, demonstrating its potential for more cost-effective applications. Xiaozhuang Song, Hangting Ye, Yaoyao Xu, Wei Fan 0010, Tianshu Yu 0001 |
BIBM | 1 |
| 2024 | PTaRL: Prototype-based Tabular Representation Learning via Space CalibrationabstractTabular data have been playing a mostly important role in diverse real-world fields, such as healthcare, engineering, finance, etc.
With the recent success of deep learning, many tabular machine learning (ML) methods based on deep networks (e.g., Transformer, ResNet) have achieved competitive performance on tabular benchmarks. However, existing deep tabular ML methods suffer from the representation entanglement and localization, which largely hinders their prediction performance and leads to performance inconsistency on tabular tasks.
To overcome these problems, we explore a novel direction of applying prototype learning for tabular ML and propose a prototype-based tabular representation learning framework, PTaRL, for tabular prediction tasks. The core idea of PTaRL is to construct prototype-based projection space (P-Space) and learn the disentangled representation around global data prototypes. Specifically, PTaRL mainly involves two stages: (i) Prototype Generating, that constructs global prototypes as the basis vectors of P-Space for representation, and (ii) Prototype Projecting, that projects the data samples into P-Space and keeps the core global data information via Optimal Transport. Then, to further acquire the disentangled representations, we constrain PTaRL with two strategies: (i) to diversify the coordinates towards global prototypes of different representations within P-Space, we bring up a diversifying constraint for representation calibration; (ii) to avoid prototype entanglement in P-Space, we introduce a matrix orthogonalization constraint to ensure the independence of global prototypes.
Finally, we conduct extensive experiments in PTaRL coupled with state-of-the-art deep tabular ML models on various tabular benchmarks and the results have shown our consistent superiority. Hangting Ye, Wei Fan 0010, Xiaozhuang Song, Shun Zheng 0001, He Zhao 0001, Dandan Guo, Yi Chang 0001 |
ICLR | 3 |
| 2024 | Boosting Protein Language Models with Negative Sample Mining
Yaoyao Xu, Xinjian Zhao, Xiaozhuang Song, Benyou Wang, Tianshu Yu 0001 |
ECML/PKDD (10) | 3 |
| 2024 | GT-TTE: Modeling Trajectories as Graphs for Travel Time EstimationabstractTravel time estimation (TTE) aims to predict travel duration and provide reliable planning for residential travel schedules. Trajectories naturally contain sequential features in form of GPS points with temporal precedence, which can be leveraged to improve prediction performance. Besides, the spatial information, i.e., the graph structure of the road network, can well represent the road highly and is commonly used to capture spatial information in traffic networks. However, extracting regional spatial information from trajectory data, in addition to its latitude and longitude information, poses a significant challenge due to the inherent format in which the trajectory data is recorded. In light of this, we propose a graph-transformer for TTE (GT-TTE) to utilize a Graph Transformer to adapt effectively to trajectories’ sequential and spatial characteristics for improved TTE performance. By traversing the trajectory nodes with GT-TTE, we construct a graph structure for all trajectory points, thereby obtaining the relative spatial information of each point. Further, we obtain a region adjacency empirically more feature-rich over the sequential data. We evaluate GT-TTE on three real-world representative data sets and observe improvement by approximately 17% compared to the state-of-the-art baselines. Yunjie Huang, Xiaozhuang Song, Shiyao Zhang 0001, Lei Li 0003, James Jian Qiao Yu |
IEEE Internet Things J. | 2 |
| 2023 | Traffic Prediction With Transfer Learning: A Mutual Information-Based ApproachabstractIn modern traffic management, one of the most essential yet challenging tasks is accurately and timely predicting traffic. It has been well investigated and examined that deep learning-based Spatio-temporal models have an edge when exploiting Spatio-temporal relationships in traffic data. Typically, data-driven models require vast volumes of data, but gathering data in small cities can be difficult owing to constraints such as equipment deployment and maintenance costs. To resolve this problem, we propose TrafficTL, a cross-city traffic prediction approach that uses big data from other cities to aid data-scarce cities in traffic prediction. Utilizing a periodicity-based transfer paradigm, it identifies data similarity and reduces negative transfer caused by the disparity between two data distributions from distant cities. In addition, the suggested method employs graph reconstruction techniques to rectify defects in data from small data cities. TrafficTL is evaluated by comprehensive case studies on three real-world datasets and outperforms the state-of-the-art baseline by around 8 to 25 percent. Yunjie Huang, Xiaozhuang Song, Yuanshao Zhu, Shiyao Zhang 0001, James Jian Qiao Yu |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | Traffic Prediction With Missing Data: A Multi-Task Learning ApproachabstractTraffic speed prediction based on real-world traffic data is a classical problem in intelligent transportation systems (ITS). Most existing traffic speed prediction models are proposed based on the hypothesis that traffic data are complete or have rare missing values. However, such data collected in real-world scenarios are often incomplete due to various human and natural factors. Although this problem can be solved by first estimating the missing values with an imputation model and then applying a prediction model, the former potentially breaks critical latent features and further leads to the error accumulation issues. To tackle this problem, we propose a graph-based spatio-temporal autoencoder that follows an encoder-decoder structure for spatio-temporal traffic speed prediction with missing values. Specifically, we regard the imputation and prediction as two parallel tasks and train them sequentially to eliminate the negative impact of imputation on raw data for prediction and accelerate the model training process. Furthermore, we utilize graph convolutional layers with a self-adaptive adjacency matrix for spatial dependencies modeling and apply gated recurrent units for temporal learning. To evaluate the proposed model, we conduct comprehensive case studies on two real-world traffic datasets with two different missing patterns and a wide and practical missing rate range from 20% to 80%. Experimental results demonstrate that the model consistently outperforms the state-of-the-art traffic prediction with missing values methods and achieves steady performance in the investigated missing scenarios and prediction horizons. Yongchao Ye, Xiaozhuang Song, Shiyao Zhang 0001, James Jian Qiao Yu |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | Efficient and Effective Multi-task Grouping via Meta Learning on Task CombinationsabstractAs a longstanding learning paradigm, multi-task learning has been widely applied into a variety of machine learning applications. Nonetheless, identifying which tasks should be learned together is still a challenging fundamental problem because the possible task combinations grow exponentially with the number of tasks, and existing solutions heavily relying on heuristics may probably lead to ineffective groupings with severe performance degradation. To bridge this gap, we develop a systematic multi-task grouping framework with a new meta-learning problem on task combinations, which is to predict the per-task performance gains of multi-task learning over single-task learning for any combination. Our underlying assumption is that no matter how large the space of task combinations is, the relationships between task combinations and performance gains lie in some low-dimensional manifolds and thus can be learnable. Accordingly, we develop a neural meta learner, MTG-Net, to capture these relationships, and design an active learning strategy to progressively select meta-training samples. In this way, even with limited meta samples, MTG-Net holds the potential to produce reasonable gain estimations on arbitrary task combinations. Extensive experiments on diversified multi-task scenarios demonstrate the efficiency and effectiveness of our method. Specifically, in a large-scale evaluation with $27$ tasks, which produce over one hundred million task combinations, our method almost doubles the performance obtained by the existing best solution given roughly the same computational cost. Data and code are available at https://github.com/ShawnKS/MTG-Net. Xiaozhuang Song, Shun Zheng 0001, Wei Cao 0007, James Jian Qiao Yu, Jiang Bian 0002 |
NeurIPS | 1 |
| 2021 | TSTNet: A Sequence to Sequence Transformer Network for Spatial-Temporal Traffic Prediction
Xiaozhuang Song, Chenhan Zhang |
ICANN (1) | 1 |
| 2021 | TINet: Multi-dimensional Traffic Data Imputation via Transformer Network
Xiaozhuang Song, Yongchao Ye, James Jian Qiao Yu |
ICANN (1) | 1 |
| 2021 | Complicating the Social Networks for Better Storytelling: An Empirical Study of Chinese Historical Text and NovelabstractDigital humanities is an important subject because it enables developments in history, literature, and films. In this article, we perform an empirical study of a Chinese historical text, Records of the Three Kingdoms (Records), and a historical novel of the same story, Romance of the Three Kingdoms (Romance). We employ deep-learning-based natural language processing (NLP) techniques to extract characters and their relationships. The adopted NLP approach can extract 93% and 91% characters that appeared in the two books, respectively. Then, we characterize the social networks and sentiments of the main characters in the historical text and the historical novel. We find that the social network in Romance is more complex and dynamic than that of Records, and the influence of the main characters differs. These findings shed light on the different styles of storytelling in the two literary genres and how the historical novel complicates the social networks of characters to enrich the literariness of the story. Chenhan Zhang, Qingpeng Zhang, Shui Yu 0001, James Jian Qiao Yu, Xiaozhuang Song |
IEEE Trans. Comput. Soc. Syst. | 5 |