Xinyue Feng

dblp:224/1602 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2025 NeighSqueeze: Compact Neighborhood Grouping for Efficient Billion-Scale Heterogeneous Graph Learning
abstract
The rapid growth of online shopping has intensified competition among logistics companies, highlighting the importance of customer expansion, i.e., identifying customers willing to establish long-term contracts. Although existing approaches frame customer expansion as a node classification task using heterogeneous graph learning to capture complex interactions between a customer and other items, it is computationally infeasible to utilize all neighboring interactions on large-scale logistics graphs. Current sub-sampling methods reduce computational load by sampling a small part of neighborhood for training. However, they introduce substantial information loss, particularly affecting high-degree nodes and decreasing predictive accuracy. To address this, we introduce NeighSqueeze, a novel approach that groups structurally and semantically similar nodes, substantially reducing the neighbors count and facilitating full-neighbor learning. NeighSqueeze consists of three modules designed to efficiently and effectively enable node grouping on billion-scale heterogeneous graphs: (1) Structure-tightness-based neighbor filtering reduces the high redundancy and complexity in similarity computations. (2) Hybrid similarity graph construction addresses the difficulty of measuring node similarity at scale; and (3) A two-level grouping strategy resolves the label dominance issue within groups. We evaluate NeighSqueeze on JD Logistics, one of the largest logistics companies in China. Compared with sub-sampling methods, our NeighSqueeze exhibits lower runtime and memory usage with full-neighbor training on the compressed graph, while simultaneously improving average precision over 28.9% in offline evaluation and increase new customer exploration rate by 18.6% in online A/B testing.
Xinyue Feng, Shuxin Zhong, Jinquan Hang, Yuequn Zhang, Guang Yang 0028, Haotian Wang 0008, Desheng Zhang 0002, Guang Wang 0001
CIKM1
2025 Hierarchical Structure Sharing Empowers Multi-task Heterogeneous GNNs for Customer Expansion
abstract
Customer expansion, i.e., growing a business's existing customer base by acquiring new customers, is critical for scaling operations and sustaining the long-term profitability of logistics companies. Although state-of-the-art works model this task as a single-node classification problem under a heterogeneous graph learning framework and achieve good performance, they struggle with extremely positive label sparsity issues in our scenario. Multi-task learning (MTL) offers a promising solution by introducing a correlated, label-rich task to enhance the label-sparse task prediction through knowledge sharing. However, existing MTL methods result in performance degradation because they fail to discriminate task-shared and task-specific structural patterns across tasks. This issue arises from their limited consideration of the inherently complex structure learning process of heterogeneous graph neural networks, which involves the multi-layer aggregation of multi-type relations. To address the challenge, we propose a Structure-Aware Hierarchical Information Sharing Framework (SrucHIS), which explicitly regulates structural information sharing across tasks in logistics customer expansion. SrucHIS breaks down the structure learning phase into multiple stages and introduces sharing mechanisms at each stage, effectively mitigating the influence of task-specific structural patterns during each stage. We evaluate StrucHIS on both private and public datasets, achieving a 51.41% average precision improvement on the private dataset and a 10.52% macro F1 gain on the public dataset. StrucHIS is further deployed at one of the largest logistics companies in China and demonstrates a 41.67% improvement in the success contract-signing rate over existing strategies, generating over 453K new orders within just two months.
Xinyue Feng, Shuxin Zhong, Jinquan Hang, Wenjun Lyu, Yuequn Zhang, Guang Yang 0028, Haotian Wang 0008, Desheng Zhang 0002, Guang Wang 0001
KDD (2)1
2025 Web Crawling Algorithm Fusing TF-IDF and Word2Vec Feature Extraction
abstract
Current research focuses on how to efficiently extract and crawl network information because, with the growth of the Internet, network information is becoming more and more diverse. To address the problem of incorrect data extraction and topic judgment of web crawlers, this study proposes a novel approach based on a file inverse frequency algorithm and Word2Vec feature extraction. The new method improves the retrieval capability of web crawlers by using the file inverse frequency algorithm and uses Word2Vec to extract data features, which improves the data extraction capability of current crawlers. The results showed that the F1 values of the research use model were 25.8% and 26.2% higher than those of the digital filtering algorithm, respectively. The total number of localization resources for the research use strategy was 2800 and the network coverage was 81%, which was 12% higher than the optimal strategy. The research use strategy had a shorter retrieval time and the model could recognize the vocabulary of the keywords. Finally, the model used by the research also had a good model processing capability when compared to other models. In summary, the new model built by the research can improve the data retrieval ability and data extraction ability of the web crawler, which provides new research ideas for future web information extraction.
Xinyue Feng
J. Web Eng.1
2024 Paths2Pair: Meta-path Based Link Prediction in Billion-Scale Commercial Heterogeneous Graphs
abstract
Link prediction, determining if a relation exists between two entities, is an essential task in the analysis of heterogeneous graphs with diverse entities and relations. Despite extensive research in link prediction, most existing works focus on predicting the relation type between given pairs of entities. However, it is almost impractical to check every entity pair when trying to find most hidden relations in a billion-scale heterogeneous graph due to the billion squared number of possible pairs. Meanwhile, most methods aggregate information at the node level, potentially leading to the loss of direct connection information between the two nodes. In this paper, we introduce Paths2Pair, a novel framework to address these limitations for link prediction in billion-scale commercial heterogeneous graphs. (i) First, it selects a subset of reliable entity pairs for prediction based on relevant meta-paths. (ii) Then, it utilizes various types of content information from the meta-paths between each selected entity pair to predict whether a target relation exists. We first evaluate our Paths2Pair based on a large-scale dataset, and results show Paths2Pair outperforms state-of-the-art baselines significantly. We then deploy our Paths2Pair on JD Logistics, one of the largest logistics companies in the world, for business expansion. The uncovered relations by Paths2Pair have helped JD Logistics identify 108,709 contacts to attract new company customers, resulting in an 84% increase in the success rate compared to the state-of-the-practice solution, demonstrating the practical value of our framework. We have released the code of our framework at https://github.com/JQHang/Paths2Pair.
Jinquan Hang, Zhiqing Hong, Xinyue Feng, Guang Wang 0001, Guang Yang 0028, Xining Song, Desheng Zhang 0002
KDD3
2024 Complex-Path: Effective and Efficient Node Ranking with Paths in Billion-Scale Heterogeneous Graphs
abstract
Node ranking in heterogeneous graphs, which quantifies the relative importance of nodes, can often be improved by incorporating information from relevant paths. Graph database and heterogeneous graph neural network (HGNN) are two main approaches to better solve this problem. Graph databases support efficient path queries for flexible path types but require manual design to combine results for node ranking. Conversely, current HGNNs can automatically integrate semantic information from multiple linear path types for accurate node ranking. However, our experiments show that they fail to outperform a multi-layer perceptron model that utilizes features extracted from multiple nonlinear conditional paths, which can be handled by graph databases. Therefore, we aim to enable HGNN to take advantage of these path types for better performance. However, HGNNs require a generalized path schema to define the structure of input paths, and incorporating each additional path type will significantly increase the required system memory and sampling time for HGNNs. To address these limitations, we introduce CompNode, a novel framework based on a new unified path schema definition called Complex-path, which is used to describe all the required path types, including nonlinear conditional path types. Then, we design a pre-aggregation method to reduce the required system memory and sampling time by pre-aggregating the same type of complex-path. Furthermore, we develop a model that combines semantic information from all aggregated complex-paths for accurate node ranking. Real-world experiments on identifying top potential high-value customers show CompNode outperforms state-of-the-art HGNNs by 20% in average precision and the previously deployed graph database method by 252% in success rate.
Jinquan Hang, Zhiqing Hong, Xinyue Feng, Guang Wang 0001, Dongjiang Cao, Jiayang Qiao, Haotian Wang 0008, Desheng Zhang 0002
Proc. VLDB Endow.3
2023 CARPG: Cross-City Knowledge Transfer for Traffic Accident Prediction via Attentive Region-Level Parameter Generation
abstract
Traffic accident prediction is a crucial problem for public safety, emergency treatment, and urban management. Existing works leverage extensive data collected from city infrastructures to achieve encouraging performance based on various machine learning techniques but cannot achieve a good performance in situations with limited data (i.e., data scarcity). Recent developments in transfer learning bring a new opportunity to solve the data scarcity problem. In this paper, we design a novel cross-city transfer learning framework named CARPG for predicting traffic accidents in data-scarce cities. We address the unique challenge of predicting traffic accidents caused by its two fundamental characteristics, i.e., spatial heterogeneity and inherent rareness, which result in the biased performance of the state-of-the-art transfer learning methods. Specifically, we build cross-city region connections by jointly learning the spatial region representations for both source and target cities with an inter-city global graph knowledge transfer process. Further, we design an efficient attention-based parameter-generating mechanism to learn region-specific traffic accident patterns, while controlling the total number of parameters. Built upon that, we ensure that only relevant patterns are transferred to each target region during the knowledge transfer process and further to be fine-tuned. We conduct extensive experiments on three real-world datasets, and the evaluation results demonstrate the superiority of our framework compared with state-of-the-art baseline models.
Guang Yang 0028, Yuequn Zhang, Jinquan Hang, Xinyue Feng, Zejun Xie, Desheng Zhang 0002, Yu Yang 0010
CIKM4
2018 BehaviorKI: Behavior Pattern Based Runtime Integrity Checking for Operating System Kernel
abstract
Kernel rootkits pose a serious threat to system security by tampering with the state of operating system inconspicuously. To ensure operating system kernel integrity, Virtual Machine Monitor (VMM) based approaches have been proposed. Most of these approaches use snapshot-based or event-triggered techniques. However, snapshot-based techniques have been suffering from missing transient attacks or significant performance overhead, while event-triggered methods are facing with heavy workload as integrity checking might be triggered by any suspicious actions. In this paper, we propose a novel solution which is a behavior-triggered integrity checking approach named BehaviorKI. By analyzing attacking processes, BehaviorKI can extract a set of behavior patterns which characterize malicious behaviors. BehaviorKI will trigger integrity checking with kernel invariants when a malicious behavior pattern detected. In this way, our approach can alleviate the performance burden by reducing the frequent kernel integrity checking. The experiment results show that Be-haviorKI outperforms existing snapshot-based and event-triggered approaches.
Xinyue Feng, Qiusong Yang, Lin Shi 0006, Qing Wang 0001
QRS1