VLDB 2026 Research / reviewers in the wild / expert
Jingyi Ding
dblp:177/6967
· DBLP profile ↗
12ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0001-8953-2510ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 5 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Feature selection via anchor weight graph guided minimizing between-class similarity
Jiarui Kong, Jingyi Ding, Ronghua Shang, Yangyang Li 0001 |
Pattern Recognit. | 2 |
| 2025 | HGphormer: Heterophilic Graph TransformerabstractGraph neural networks (GNNs) have been widely used in various node-level tasks on graphs due to their powerful representation learning ability. Traditional GNNs rely on the homophily assumption that nodes with the same label in a graph tend to be connected to each other. But there are a large number of heterophily graphs in the real world, where most proximal nodes have different labels. So heterophily GNNs has been proposed, which tried to improve the performance of GNNs on heterophily graph by obtaining information from multi-hop neighbor nodes. A promising way for heterophily graphs is Graph Transformers (GTs). Without relying on the homophily assumption, GTs aggregate information of nodes depending on their similarity, thus is suitable for both homophily and heterophily graphs. Since the quadratic time complexity of GTs, most of existing GTs focus on how to reduce the complexity and make it applicable for nodes classification, their performance is still unsatisfied in heterophily graphs. To solve the above problem, Heterophilic Graph Transformer (HGphormer) is proposed in this paper. In order to reduce the interference between attribute embedding and structure embedding, a parallel architecture of Transformer is proposed. HGphormer also decouples the aggregated information into homophily and heterophily information and uses them adaptively to further improve the accuracy. A sample technique is proposed to sample neighbors from multiple hops and reduce the time complexity. Experiments show that the proposed HGphormer outperforms the state of the art methods on both homophily graph and heterophily graph datasets. Jianshe Wu, Yaolin Liu, Lingjie Zhang, Jingyi Ding |
Knowl. Based Syst. | 5 |
| 2025 | Magnus: A Holistic Approach to Data Management for Large-Scale Machine Learning WorkloadsabstractMachine learning (ML) has become a cornerstone of key applications at ByteDance. As model complexity and data volumes surge, data management for large-scale ML workloads faces substantial challenges, particularly with recent advances in large recommendation models (LRMs) and large multimodal models (LMMs). Traditional approaches exhibit limitations in storage efficiency, metadata scalability, update mechanisms, and integration with ML frameworks. To address these challenges, we propose Magnus, a holistic data management system built upon Apache Iceberg. Magnus integrates innovative optimizations across resource-efficient storage formats optimized for large wide tables and multimodal data, built-in support for vector and inverted indexes to accelerate data retrieval, scalable metadata planning with Git-like branching and tagging capabilities, and high-performance update/upsert based on lightweight merge-on-read (MOR) strategies. Additionally, Magnus provides native support and specialized enhancement for LRM and LMM training workloads. Experimental results demonstrate significant performance gains in real-world ML scenarios. Magnus has been deployed at ByteDance for over five years, enabling robust and efficient data infrastructure for large-scale ML workloads. Jingyi Ding, Irshad Kandy, Yanghao Lin, Zhongjia Wei, Zhiwei Peng, Jixi Shan, Hongyue Mao, Xiuqi Huang, Xun Song, Yanjia Li, Tianhao Yang, Xiaohong Dong, Kang Lei, Pengwei Zhao, Wei Chen 0001 |
Proc. VLDB Endow. | 2 |
| 2024 | Heterogeneous Graph CondensationabstractGraph neural networks greatly facilitate data processing in homogeneous and heterogeneous graphs. However, training GNNs on large-scale graphs poses a significant challenge to computing resources. It is especially prominent on heterogeneous graphs, which contain multiple types of nodes and edges, and heterogeneous GNNs are also several times more complex than the ordinary GNNs. Recently, Graph condensation (GCond) is proposed to address the challenge by condensing large-scale homogeneous graphs into small-scale informative graphs. Its label-based feature initialization and fully-connected design perform well on homogeneous graphs. While in heterogeneous graphs, label information generally only exists in specific types of nodes, making it difficult to be applied directly to heterogeneous graphs. In this paper, we propose heterogeneous graph condensation (HGCond). HGCond uses clustering information instead of label information for feature initialization, and constructs a sparse connection scheme accordingly. In addition, we found that the simple parameter exploration strategy in GCond leads to insufficient optimization on heterogeneous graphs. This paper proposes an exploration strategy based on orthogonal parameter sequences to address the problem. We experimentally demonstrate that the novel feature initialization and parameter exploration strategy is effective. Experiments show that HGCond significantly outperforms baselines on multiple datasets. On the dataset DBLP, HGCond can condense DBLP to 0.5% of its original scale to obtain DBLP-0.005. GNNs trained on DBLP-0.005 can retain nearly 99% accuracy compared to the GNNs trained on full-scale DBLP. Jian Gao 0010, Jianshe Wu, Jingyi Ding |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2023 | Time-Aware POI Recommendation Based on Multi-Grained Location GroupingabstractThe task of point-of-interest (POI) recommendation aims to recommend locations to users in location-based applications. Among them, the task of time-aware POI recommendation aims to capture the user’s preferences that change dynamically over time, so as to make more accurate recommendations to users at a specific time. While existing works take into account the spatial, temporal and category context of POIs, they cannot capture user preferences that are more fine-grained than the category granularity. Additionally, RNN-based methods suffer from the problem of long-term dependency when capturing a user’s check-in patterns. To address these challenges, we propose a novel model with POI multi-grained grouping method which captures the user’s co-visit patterns and weekly patterns, to obtain finer-grained POI groups. The model also utilizes the transformer model to capture the user’s check-in preference patterns. We evaluate our model on two real-world datasets, and the experimental results demonstrate the effectiveness of our proposed model. Haoxiang Zhang 0003, Wenchao Bai, Jingyi Ding, Jiahui Jin 0001 |
CSCWD | 3 |
| 2023 | Dependency-Aware Core Column Discovery for Table Understanding
Jingyi Qiu, Aibo Song, Jiahui Jin 0001, Tianbo Zhang, Jingyi Ding, Xiaolin Fang 0001, Jianguo Qian |
ISWC | 5 |
| 2023 | Constructing negative samples via entity prediction for multi-task knowledge representation learning
Guihai Chen, Jianshe Wu, Wenyun Luo, Jingyi Ding |
Knowl. Based Syst. | 4 |
| 2023 | Community evolution prediction based on a self-adaptive timeframe in social networks
Jingyi Ding, Tiwen Wang, Ruohui Cheng, Licheng Jiao, Jianshe Wu, Jing Bai 0003 |
Knowl. Based Syst. | 1 |
| 2022 | Graph label prediction based on local structure characteristics representation
Jingyi Ding, Ruohui Cheng, Jian Song 0003, Xiangrong Zhang, Licheng Jiao, Jianshe Wu |
Pattern Recognit. | 1 |
| 2021 | SGBMN: Symplectic Group Bayesian Manifold Network for Few-shot ClassificationabstractThe field of meta-learning, or learning-to-learn, has seen a dramatic rise in interest in recent years. Particularly, employing meta-learning for few-shot classification has achieved remarkable advances. However, the uncertainty problem triggered by the noisy data or modeling assumptions hinders the performance of existing approaches to be further improved. To tackle the above issue, this paper proposes a novel meta-learning approach called Symplectic Group Bayesian Manifold Network (SGBMN) for a more accurate classification prediction. Specifically, we adopt symplectic group bayesian matrix to represent each input data point, such that the uncertainty problem caused by complex data could be alleviated. Then, a manifold space is constructed by combing the above symplectic group bayesian matrices, in which the optimization can be performed by natural gradient descent. We conduct extensive experiments on two real-world image datasets and the results demonstrate that our proposed method outperforms several baseline approaches. Jingyi Ding, Fanzhang Li |
IJCNN | 1 |
| 2020 | Influence maximization based on the realistic independent cascade model
Jingyi Ding, Jianshe Wu, Yuwei Guo 0001 |
Knowl. Based Syst. | 1 |
| 2016 | Prediction of missing links based on community relevance and ruler inference
Jingyi Ding, Licheng Jiao, Jianshe Wu, Fang Liu 0001 |
Knowl. Based Syst. | 1 |