EDBT 2026 Demo / reviewers in the wild / expert
Yuandong Wang 0002
dblp:47/8988-2
· DBLP profile ↗
15ranked-venue papers
6as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 13 · 5 first-author · 9 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MathSight: A Benchmark Exploring Have Vision-Language Models Really Seen in University-Level Mathematical Reasoning?abstractRecent advances in Vision-Language Models (VLMs) have achieved impressive progress in multimodal mathematical reasoning.Yet, how much visual information truly contributes to reasoning remains unclear.Existing benchmarks report strong overall performance but seldom isolate the role of the image modality, leaving open whether VLMs genuinely leverage visual understanding or merely depend on linguistic priors.To address this, we present MathSight, a university-level multimodal mathematical reasoning benchmark designed to disentangle and quantify the effect of visual input.Each problem includes multiple visual variants-original, hand-drawn, photocaptured-and a text-only condition for controlled comparison.Experiments on state-ofthe-art VLMs reveal a consistent trend: the contribution of visual information diminishes with increasing problem difficulty.Remarkably, Qwen3-VL without any image input surpasses both its multimodal variants and GPT-5, underscoring the need for benchmarks like MathSight to advance genuine vision-grounded reasoning in future models.The project page is available at https Yuandong Wang 0002, Yao Cui, Zhen Yang 0034, Yangfu Zhu, Zhenzhou Shao |
ACL (1) | 1 |
| 2026 | Debiased Multimodal Personality Understanding through Dual Causal InterventionabstractMultimodal personality understanding plays a critical role in human-centered artificial intelligence. Previous work mainly focus on learning rich multimodal representations for video personality understanding. However, they often suffer from potential harm caused by subject bias (e.g., observable age and unobservable mental states), as subjects originate from diverse demographic backgrounds. Learning such spurious associations between multimodal features and traits may lead to unfair personality understanding. In this work, we construct a Structural Causal Model (SCM) to analyze the impact of these biases from a causal perspective, and propose a novel Dual Causal Adjustment Network (DCAN) to mitigate the interference of subject attributes on personality understanding. Specifically, we design a Back-door Adjustment Causal Learning (BACL) module to block spurious correlations from observable demographic factors via a prototype-based confounder dictionary, and subsequently apply a Front-door Adjustment Causal Learning (FACL) module to address latent and unobservable biases through a learned mediator dictionary intervention, thereby achieving causal disentanglement of representations for deconfounded reasoning. Importantly, we construct a Demographic-annotated Multimodal Student Personality (DMSP) dataset to support the analysis and discussion of fairness-related factors. Extensive experiments on the benchmark dataset CFI-V2 and our DMSP dataset demonstrate that DCAN consistently improves prediction accuracy, reaching 92.11% and 92.90%, respectively. Meanwhile, the improvements in the fairness metrics of equal opportunity and demographic parity are 6.57% and 7.97% on CFI-V2, and 15.38% and 20.06% on the DMSP dataset. Our code and DMSP dataset are available at https://github.com/Sabrina-han/DCAN Yangfu Zhu, Zitong Han, Nianwen Ning, Yuandong Wang 0002, Hang Feng, Zhenzhou Shao |
SIGIR | 5 |
| 2025 | LGB: Language Model and Graph Neural Network-Driven Social Bot DetectionabstractMalicious social bots achieve their malicious purposes by spreading misinformation and inciting social public opinion, seriously endangering social security, making their detection a critical concern. Recently, graph-based bot detection methods have achieved state-of-the-art (SOTA) performance. However, our research finds many isolated and poorly linked nodes in social networks, as shown in Fig. 1, which graphbased methods cannot effectively detect. To address this problem, our research focuses on effectively utilizing node semantics and network structure to jointly detect sparsely linked nodes. Given the excellent performance of language models (LMs) in natural language understanding (NLU), we propose a novel social bot detection framework LGB, which consists of two main components: language model (LM) and graph neural network (GNN). Specifically, the social account information is first extracted into unified user textual sequences, which is then used to perform supervised fine-tuning (SFT) of the language model to improve its ability to understand social account semantics. Next, the semantically enriched node representation is fed into the pretrained GNN to further enhance the node representation by aggregating information from neighbors. Finally, LGB fuses the information from both modalities to improve the detection performance of sparsely linked nodes. Extensive experiments on two real-world datasets demonstrate that LGB consistently outperforms state-of-the-art baseline models by up to 10.95%. LGB is already online: https://botdetection.aminer.cn/robotmain Ming Zhou 0004, Yuandong Wang 0002, Yuxiao Dong, Jie Tang 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2024 | MultiSPANS: A Multi-range Spatial-Temporal Transformer Network for Traffic Forecast via Structural Entropy OptimizationabstractTraffic forecasting is a complex multivariate time-series regression task of paramount importance for traffic management and planning. However, existing approaches often struggle to model complex multi-range dependencies using local spatiotemporal features and road network hierarchical knowledge. To address this, we propose MultiSPANS. First, considering that an individual recording point cannot reflect critical spatiotemporal local patterns, we design multi-filter convolution modules for generating informative ST-token embeddings to facilitate attention computation. Then, based on ST-token and spatial-temporal position encoding, we employ the Transformers to capture long-range temporal and spatial dependencies. Furthermore, we introduce structural entropy theory to optimize the spatial attention mechanism. Specifically, The structural entropy minimization algorithm is used to generate optimal road network hierarchies, i.e., encoding trees. Based on this, we propose a relative structural entropy-based position encoding and a multi-head attention masking scheme based on multi-layer encoding trees. Extensive experiments demonstrate the superiority of the presented framework over several state-of-the-art methods in real-world traffic datasets, and the longer historical windows are effectively utilized. The code is available at https://github.com/SELGroup/MultiSPANS. Dongcheng Zou, Senzhang Wang, Xuefeng Li 0003, Hao Peng 0001, Yuandong Wang 0002, Kehua Sheng, Bo Zhang 0106 |
WSDM | 5 |
| 2024 | DropConn: Dropout Connection Based Random GNNs for Molecular Property PredictionabstractRecently, molecular data mining has attracted a lot of attention owing to its great application potential in material and drug discovery. However, this mining task faces a challenge posed by the scarcity of labeled molecular graphs. To overcome this challenge, we introduce a novel data augmentation and a semi-supervised confidence-aware consistency regularization training framework for molecular property prediction. The core of our framework is a data augmentation strategy on molecular graphs, named DropConn (Dropout Connection). DropConn generates pseudo molecular graphs by softening the hard connections of chemical bonds (as edges), where the soft weights are calculated from edge features so that the adaptive interactions between different atoms can be incorporated. Besides, to enhance the model's generalization ability, a consistency regularization training strategy is proposed to take full advantage of massive unlabeled data. Furthermore, DropConn can serve as a plugin that can be seamlessly added to many existing models. Extensive experiments under both non-pre-training setting and fine-tuning setting demonstrate that DropConn can obtain superior performance (up to 8.22%) over state-of-the-art methods on molecular property prediction tasks. The code is available athttps://github.com/THUDM/DropConn. Wenzheng Feng, Yuandong Wang 0002, Zhongang Qi, Ying Shan, Jie Tang 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2023 | Detecting Social Bot on the Fly using Contrastive LearningabstractSocial bot detection is becoming a task of wide concern in social security. All along, the development of social bot detection technology is hindered by the lack of high-quality annotated data. Besides, the rapid development of AI Generated Content (AIGC) technology is dramatically improving the creative ability of social bots. For example, the recently released ChatGPT [2] can fool the state-of-the-art AI-text-detection method with a probability of 74%, bringing a large challenge to content-based bot detection methods. To address the above drawbacks, we propose a Contrastive Learning-driven Social Bot Detection framework (CBD). The core of CBD is characterized by a two-stage model learning strategy: a contrastive pre-training stage to mine generalization patterns from massive unlabeled social graphs, followed by a semi-supervised fine-tuning stage to model task-specific knowledge latent in social graphs with a few annotations. The above strategy endows our model with promising detection performance under an extreme scarcity of labeled data. In terms of system architecture, we propose a smart feedback mechanism to further improve detection performance. Comprehensive experiments on a real bot detection dataset show that CBD consistently outperforms 10 state-of-the-art baselines by a large margin for few-shot bot detection using very little (5-shot) labeled data. CBD has been deployed online. Ming Zhou 0004, Yuandong Wang 0002, Jie Tang 0001 |
CIKM | 3 |
| 2023 | ApeGNN: Node-Wise Adaptive Aggregation in GNNs for RecommendationabstractIn recent years, graph neural networks (GNNs) have made great progress in recommendation. The core mechanism of GNNs-based recommender system is to iteratively aggregate neighboring information on the user-item interaction graph. However, existing GNNs treat users and items equally and cannot distinguish diverse local patterns of each node, which makes them suboptimal in the recommendation scenario. To resolve this challenge, we present a node-wise adaptive graph neural network framework ApeGNN. ApeGNN develops a node-wise adaptive diffusion mechanism for information aggregation, in which each node is enabled to adaptively decide its diffusion weights based on the local structure (e.g., degree). We perform experiments on six widely-used recommendation datasets. The experimental results show that the proposed ApeGNN is superior to the most advanced GNN-based recommender methods (up to 48.94%), demonstrating the effectiveness of node-wise adaptive aggregation. Yifan Zhu 0001, Yuxiao Dong, Yuandong Wang 0002, Wenzheng Feng, Evgeny Kharlamov, Jie Tang 0001 |
WWW | 4 |
| 2023 | Synthesizing Realistic Trajectory Data With Differential PrivacyabstractVehicle trajectory data is critical for traffic management and location-based services. However, the released trajectories raise serious privacy concerns because they contain sensitive information such as homes and workplaces. Based on differential privacy, this problem can be addressed by generating synthetic trajectories from the original sensitive data while guaranteeing personal privacy. Unfortunately, existing methods focus on synthesizing trajectory datasets that preserve summary-level statistics (e.g., the overall distribution of user movements), making these synthetic trajectories lose individual-level mobility patterns. As shown in our experiment, this results in the low performance of their synthetic datasets in real-world applications. To address these limitations, we propose a novel solution for Synthesizing Private and Realistic Trajectories, namely SPRT, whose key idea is to integrate the public geography structures of the target area into the process of private trajectory synthesis. This enables us to capture more accurate mobility patterns to synthesize realistic trajectories, which can preserve both summary-level statistics and individual-level mobility behaviors. Consequently, the synthetic trajectories generated by SPRT are more similar to real trajectories and therefore more practical. We evaluate the performance of SPRT in real-world applications by applying its synthetic data to a series of trajectory analytic tasks. The results demonstrate that our solution improves data utility by at least 37% over state-of-the-art approaches. Xinyue Sun, Qingqing Ye 0001, Haibo Hu 0001, Yuandong Wang 0002, Kai Huang 0011, Tianyu Wo, Jie Xu 0007 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2023 | Secure Your Ride: Real-Time Matching Success Rate Prediction for Passenger-Driver PairsabstractIn recent years, online ride-hailing platforms, such as Uber and Didi, have become an indispensable part of urban transportation and make our lives more convenient. After a passenger is matched up with a driver by the platform, both the passenger and the driver have the freedom to simply accept or cancel a ride with one click. Hence, accurately predicting whether a passenger-driver pair is a good match, i.e., its matching success rate (MSR), turns out to be crucial for ride-hailing platforms to devise instant strategies such as order assignment. However, since the users of ride-hailing platforms consist of two parties, decision-making needs to simultaneously account for the dynamics from both the driver and the passenger sides. This makes it more challenging than traditional online advertising tasks that predict a user's response towards an object, e.g., click-through rate prediction for advertisements. Moreover, the amount of available data is severely imbalanced across different cities, creating difficulties for training an accurate model for smaller cities with scarce data. Though a sophisticated neural network architecture can help improve the prediction accuracy under data scarcity, the overly complex design will impede the model's capacity of delivering timely predictions in a production environment. In the paper, to accurately predict the MSR of passenger-driver, we propose theMulti-View model (MV) which comprehensively learns the interactions among the dynamic features of the passenger, driver, trip order, as well as the context. Regarding the data imbalance problem, we further design theKnowledgeDistillation framework (KD) to supplement the model's predictive power for smaller cities using the knowledge from cities with denser data, and also generate a simple model to support efficient deployment. Finally, we conduct extensive experiments on real-world datasets from several different cities, which demonstrates the superiority of our solution. Yuandong Wang 0002, Hongzhi Yin, Lian Wu, Tong Chen 0005 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2022 | Passenger Mobility Prediction via Representation Learning for Dynamic Directed and Weighted GraphsabstractIn recent years, ride-hailing services have been increasingly prevalent, as they provide huge convenience for passengers. As a fundamental problem, the timely prediction of passenger demands in different regions is vital for effective traffic flow control and route planning. As both spatial and temporal patterns are indispensable passenger demand prediction, relevant research has evolved from pure time series to graph-structured data for modeling historical passenger demand data, where a snapshot graph is constructed for each time slot by connecting region nodes via different relational edges (origin-destination relationship, geographical distance, etc.). Consequently, the spatiotemporal passenger demand records naturally carry dynamic patterns in the constructed graphs, where the edges also encode important information about the directions and volume (i.e., weights) of passenger demands between two connected regions. aspects in the graph-structure data. representation for DDW is the key to solve the prediction problem. However, existing graph-based solutions fail to simultaneously consider those three crucial aspects of dynamic, directed, and weighted graphs, leading to limited expressiveness when learning graph representations for passenger demand prediction. Therefore, we propose a novel spatiotemporal graph attention network, namely Gallat ( G raph prediction with all at tention) as a solution. In Gallat, by comprehensively incorporating those three intrinsic properties of dynamic directed and weighted graphs, we build three attention layers to fully capture the spatiotemporal dependencies among different regions across all historical time slots. Moreover, the model employs a subtask to conduct pretraining so that it can obtain accurate results more quickly. We evaluate the proposed model on real-world datasets, and our experimental results demonstrate that Gallat outperforms the state-of-the-art approaches. Yuandong Wang 0002, Hongzhi Yin, Tong Chen 0005, Tianyu Wo, Jie Xu 0007 |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2021 | Gallat: A Spatiotemporal Graph Attention Network for Passenger Demand PredictionabstractOnline ride-hailing services have become an important component of urban transportation in recent years. As a fundamental research problem for such services, the timely prediction of passenger demands in different regions is vital for effective traffic flow control. As both spatial and temporal patterns are indispensable passenger demand prediction, relevant research has evolved from pure time series to graph-structured data for modelling historical passenger demand data, where a snapshot graph is constructed for each time slot by connecting region nodes via different relational edges. Consequently, the spatiotemporal passenger demand records naturally carry dynamic patterns in the constructed graphs, where the edges also encode important information about the directions and volume (i.e., weights) of passenger demands between two connected regions. However, existing graph-based solutions fail to simultaneously consider those three crucial aspects of dynamic, directed and weighted (DDW) graphs, leading to limited expressiveness when learning graph representations for passenger demand prediction. Therefore, we propose a novel spatiotemporal graph attention network, namely Gallat (Graph prediction with all attention) as a solution. In Gallat, by comprehensively incorporating those three intrinsic properties of DDW graphs, we build three attention layers to fully capture the spatiotemporal dependencies among different regions across all historical time slots. Our experimental results on real-world datasets demonstrate that Gallat outperforms the state-of-the-art approaches. Yuandong Wang 0002, Hongzhi Yin, Tong Chen 0005, Tianyu Wo, Jie Xu 0007 |
ICDE | 1 |
| 2019 | Origin-Destination Matrix Prediction via Graph Convolution: a New Perspective of Passenger Demand ModelingabstractRide-hailing applications are becoming more and more popular for providing drivers and passengers with convenient ride services, especially in metropolises like Beijing or New York. To obtain the passengers' mobility patterns, the online platforms of ride services need to predict the number of passenger demands from one region to another in advance. We formulate this problem as an Origin-Destination Matrix Prediction (ODMP) problem. Though this problem is essential to large-scale providers of ride services for helping them make decisions and some providers have already put it forward in public, existing studies have not solved this problem well. One of the main reasons is that the ODMP problem is more challenging than the common demand prediction. Besides the number of demands in a region, it also requires the model to predict the destinations of them. In addition, data sparsity is a severe issue. To solve the problem effectively, we propose a unified model, Grid-Embedding based Multi-task Learning (GEML) which consists of two components focusing on spatial and temporal information respectively. The Grid-Embedding part is designed to model the spatial mobility patterns of passengers and neighboring relationships of different areas, the pre-weighted aggregator of which aims to sense the sparsity and range of data. The Multi-task Learning framework focuses on modeling temporal attributes and capturing several objectives of the ODMP problem. The evaluation of our model is conducted on real operational datasets from UCAR and Didi. The experimental results demonstrate the superiority of our GEML against the state-of-the-art approaches. Yuandong Wang 0002, Hongzhi Yin, Hongxu Chen 0002, Tianyu Wo, Jie Xu 0007, Kai Zheng 0001 |
KDD | 1 |
| 2019 | A Unified Framework with Multi-source Data for Predicting Passenger Demands of Ride ServicesabstractRide-hailing applications have been offering convenient ride services for people in need. However, such applications still suffer from the issue of supply-demand disequilibrium, which is a typical problem for traditional taxi services. With effective predictions on passenger demands, we can alleviate the disequilibrium by pre-dispatching, dynamic pricing or avoiding dispatching cars to zero-demand areas. Existing studies of demand predictions mainly utilize limited data sources, trajectory data, or orders of ride services or both of them, which also lacks a multi-perspective consideration. In this article, we present a unified framework with a new combined model and a road-network-based spatial partition to leverage multi-source data and model the passenger demands from temporal, spatial, and zero-demand-area perspectives. In addition, our framework realizes offline training and online predicting, which can satisfy the real-time requirement more easily. We analyze and evaluate the performance of our combined model using the actual operational data from UCAR. The experimental results indicate that our model outperforms baselines on both Mean Absolute Error and Root Mean Square Error on average. Yuandong Wang 0002, Xuelian Lin, Hua Wei 0001, Tianyu Wo, Jie Xu 0007 |
ACM Trans. Knowl. Discov. Data | 1 |
| 2018 | Context-Aware Location Annotation on Mobility Records Through User Grouping
Hua Wei 0001, Xuelian Lin, Fei Wu 0007, Zhenhui Li, Kaiheng Chen, Yuandong Wang 0002, Jie Xu 0007 |
PAKDD (3) | 7 |
| 2016 | ZEST: A Hybrid Model on Predicting Passenger Demand for Chauffeured Car ServiceabstractChauffeured car service based on mobile applications like Uber or Didi suffers from supply-demand disequilibrium, which can be alleviated by proper prediction on the distribution of passenger demand. In this paper, we propose a Zero-Grid Ensemble Spatio Temporal model (ZEST) to predict passenger demand with four predictors: a temporal predictor and a spatial predictor to model the influences of local and spatial factors separately, an ensemble predictor to combine the results of former two predictors comprehensively and a Zero-Grid predictor to predict zero demand areas specifically since any cruising within these areas costs extra waste on energy and time of driver. We demonstrate the performance of ZEST on actual operational data from ride-hailing applications with more than 6 million order records and 500 million GPS points. Experimental results indicate our model outperforms 5 other baseline models by over 10% both in MAE and sMAPE on the three-month datasets. Hua Wei 0001, Yuandong Wang 0002, Tianyu Wo, Yaxiao Liu, Jie Xu 0007 |
CIKM | 2 |