EDBT 2026 Demo / reviewers in the wild / expert
Shuxin Zhong
dblp:203/9651
· DBLP profile ↗
16ranked-venue papers in the field
5as first author
16since 2021 · last 2026
0009-0006-1758-2870ORCID · corroborated
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 10 (5 first)Data Mining & Knowledge Discovery · 4Database Systems & Data Management · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | StormMind: Disentangled Layerwise Modeling for Convective Weather SystemsabstractTimely nowcasting is critical for public safety during fast-evolving storms, where even short delays can trigger cascading failures—as in the October 2024 Spain flash flood that claimed over 90 lives within minutes. While radar offers reliable real-time sensing of atmospheric structure, models that collapse 3D volumes into 2D slices inevitably discard vertical information essential for capturing storm growth, phase transitions, and collapse. We introduce StormMind, a physically grounded framework that forecasts convective evolution by modeling causal interactions across stratified atmospheric layers. StormMind addresses two fundamental challenges:(1) the nonlinear, asynchronous coupling between low-, mid-, and high-level processes; and (2) reflectivity uncertainty, where storms with distinct vertical structures may appear deceptively similar on radar, masking their true phase and intensity. To tackle these issues, StormMind designs: i) a Convection Dynamics Extractor that models storm evolution from two complementary perspectives—horizontal morphology, capturing the spatial organization of physical processes within individual atmospheric layers, and vertical coupling, modeling energy exchanges across layers; and ii) a Convection Manifestation Reconstructor that adaptively fuses intra- and inter-layer signals, conditioned on the evolving storm state, to infer phase transitions (e.g., initiation, intensification, dissipation). Evaluated on the large-scale 3D-NEXRAD dataset (2020–2022, U.S.), StormMind outperforms strong baselines, achieving a 14.71% gain in CSI40. In real-world deployment with the Guangzhou Meteorological Bureau (Mar–May 2025), it improves CSI40 by 9.39% and boosts early-warning accuracy (98.33%) Jun Chen 0005, Minghui Qiu, Lin Chen 0020, Shuxin Zhong, Binghong Chen, Kaishun Wu |
KDD (1) | 5 |
| 2026 | LEAP: LLM-Enhanced E-commerce Demand Prediction under Emergent Events
Shuxin Zhong, Jun Chen 0005, Kaishun Wu |
WWW | 1 |
| 2025 | NeighSqueeze: Compact Neighborhood Grouping for Efficient Billion-Scale Heterogeneous Graph LearningabstractThe rapid growth of online shopping has intensified competition among logistics companies, highlighting the importance of customer expansion, i.e., identifying customers willing to establish long-term contracts. Although existing approaches frame customer expansion as a node classification task using heterogeneous graph learning to capture complex interactions between a customer and other items, it is computationally infeasible to utilize all neighboring interactions on large-scale logistics graphs. Current sub-sampling methods reduce computational load by sampling a small part of neighborhood for training. However, they introduce substantial information loss, particularly affecting high-degree nodes and decreasing predictive accuracy. To address this, we introduce NeighSqueeze, a novel approach that groups structurally and semantically similar nodes, substantially reducing the neighbors count and facilitating full-neighbor learning. NeighSqueeze consists of three modules designed to efficiently and effectively enable node grouping on billion-scale heterogeneous graphs: (1) Structure-tightness-based neighbor filtering reduces the high redundancy and complexity in similarity computations. (2) Hybrid similarity graph construction addresses the difficulty of measuring node similarity at scale; and (3) A two-level grouping strategy resolves the label dominance issue within groups. We evaluate NeighSqueeze on JD Logistics, one of the largest logistics companies in China. Compared with sub-sampling methods, our NeighSqueeze exhibits lower runtime and memory usage with full-neighbor training on the compressed graph, while simultaneously improving average precision over 28.9% in offline evaluation and increase new customer exploration rate by 18.6% in online A/B testing. Xinyue Feng, Shuxin Zhong, Jinquan Hang, Yuequn Zhang, Guang Yang 0028, Haotian Wang 0008, Desheng Zhang 0002, Guang Wang 0001 |
CIKM | 2 |
| 2025 | Hearing the Meaning, Not the Mess: Beyond Literal Transcription for Spoken LanguageabstractWith the rise of virtual communication and smart devices, speech has become the most natural medium of interaction. Yet it remains intrinsically difficult: speech is fleeting, unstructured, and disfluent, making key information prone to loss. Conventional Speech-to-Text (STT) systems attempt to acoustically reconstruct what was said. However, their frame-level alignment and rigid token-by-token decoding break down under noise, interruptions, or fragmentation. Humans, in contrast, readily grasp what was meant by exploiting syntax, discourse, pragmatics, and prosody. We argue for a paradigm shift from acoustic reconstruction to semantic transduction: inferring meaning directly from speech, abstracted from surface distortions. This shift raises two challenges: (C1) the lack of anchors between audio and meaning, and (C2) the need to maintain compositional semantics. To address these, we introduce CogTrans, a cognitively inspired speech-to-meaning framework. CogTrans tackles C1 through a Semantic Anchor Explorer, built on I-JEPA to capture higher-order regularities, prosodic rhythms, cross-frequency coarticulation, discourse continuity-providing resilient semantic scaffolds under noise and fragmentation. For C2, it designs a Lexical-Semantic Harmonizer that dynamically integrates these anchors with lexical embeddings; thereby preserving fine-grained compositional fidelity in roles, order, and entities. Extensive experiments show that CogTrans delivers consistent and substantial gains under challenging conditions. On GigaSpeech, it achieves a 6.58% relative Word Error Rate (WER) reduction, and on the multilingual VoxPopuli benchmark, the gain climbs to 12.97% at 10 dB noise-a regime where conventional models typically collapse. Beyond literal accuracy, CogTrans also boosts semantic fidelity, with a 3.40% increase in ROUGE-L and 3.45% in USE-Sim, ensuring transcripts remain faithful not only in words but also in meaning. Together, these results underscore that CogTrans is robust in noisy, unconstrained environments-precisely the conditions where reliability matters most. Jiarong Liu, Jifan Yang, Weizheng Wang 0001, Qipeng Xie, Shuxin Zhong, Kaishun Wu |
CIKM | 8 |
| 2025 | Fraudulent Delivery Detection with Multimodal Courier Behavior Data in Last-Mile DeliveryabstractThe rapid growth of e-commerce has made last-mile delivery a critical service in daily life. Despite regulations mandating doorstep delivery, the pressure of penalties for delays can lead to fraudulent delivery behaviors, where couriers may report package receipt without actually deliver the package to assigned locations. Existing studies on fraud behavior detection focus on exploring user (courier) behaviors for fraud behavior detection. However, due to the inaccuracy of GPS positioning and the variability of user behavior patterns caused by dynamic environmental factors, relying solely on behavior data remains insufficient for detecting fraudulent deliveries. In this paper, we present a Multimodal Fraudulent Delivery Detection framework (MFDD), which integrates heterogeneous data from multiple agents (courier-side and user-side)-including couriers' physical behavior, digital behavior, and conversations containing customer feedback-for detecting fraudulent deliveries in the last-mile delivery. We employ attention mechanisms to extract features from each modality and use cross-modal fusion to capture complex and varied relationships between multimodal data. To further mitigate modality imbalance during training, we introduce a dynamic gradient-modulation strategy that balances learning across all modalities. We implement and evaluate MFDD on real-world, human-annotated data, achieving a 9.6% improvement in precision and a 5.8% increase in accuracy over the state-of-the-art methods. We also deploy the model in the production environment of JD Logistics, and results show that compared to existing methods, MFDD improves accuracy by 15.3%, reducing estimated annual costs by over 18.5 million CNY. Sijing Duan, Shuxin Zhong, Zhiqing Hong, Weijian Zuo, Desheng Zhang 0002, Yi Ding 0011 |
CIKM | 3 |
| 2025 | Hierarchical Structure Sharing Empowers Multi-task Heterogeneous GNNs for Customer ExpansionabstractCustomer expansion, i.e., growing a business's existing customer base by acquiring new customers, is critical for scaling operations and sustaining the long-term profitability of logistics companies. Although state-of-the-art works model this task as a single-node classification problem under a heterogeneous graph learning framework and achieve good performance, they struggle with extremely positive label sparsity issues in our scenario. Multi-task learning (MTL) offers a promising solution by introducing a correlated, label-rich task to enhance the label-sparse task prediction through knowledge sharing. However, existing MTL methods result in performance degradation because they fail to discriminate task-shared and task-specific structural patterns across tasks. This issue arises from their limited consideration of the inherently complex structure learning process of heterogeneous graph neural networks, which involves the multi-layer aggregation of multi-type relations. To address the challenge, we propose a Structure-Aware Hierarchical Information Sharing Framework (SrucHIS), which explicitly regulates structural information sharing across tasks in logistics customer expansion. SrucHIS breaks down the structure learning phase into multiple stages and introduces sharing mechanisms at each stage, effectively mitigating the influence of task-specific structural patterns during each stage. We evaluate StrucHIS on both private and public datasets, achieving a 51.41% average precision improvement on the private dataset and a 10.52% macro F1 gain on the public dataset. StrucHIS is further deployed at one of the largest logistics companies in China and demonstrates a 41.67% improvement in the success contract-signing rate over existing strategies, generating over 453K new orders within just two months. Xinyue Feng, Shuxin Zhong, Jinquan Hang, Wenjun Lyu, Yuequn Zhang, Guang Yang 0028, Haotian Wang 0008, Desheng Zhang 0002, Guang Wang 0001 |
KDD (2) | 2 |
| 2025 | LLM4HAR: Generalizable On-device Human Activity Recognition with Pretrained LLMsabstractA long-standing challenge for pushing sensor-based human activity recognition (HAR) to industrial usage is the distribution shift between training data and testing data: significant variations in data distribution lead to a notable decline in performance. Recently, Large Language Models (LLMs) have demonstrated exceptional generalization capability, which provides a new opportunity to mitigate the distribution shift problem of HAR. However, since LLMs are inherently designed and trained on textual data, their potential to enhance generalization in HAR applications remains an open question. In this paper, we introduce LLM4HAR, a novel LLM-based model to improve cross-domain HAR. LLM4HAR consists of three main modules: (i) the Sensor Data Adaptation module, which aligns IMU signals with LLMs via sensor embedding(ii) the Sensor Knowledge Learning module, which injects sensor knowledge into LLMs for activity recognition, and (iii) the Efficiency Enhancement module, which employs a partial training strategy and reduces the model size by more than 10 times. Extensive evaluations show that LLM4HAR outperforms the existing methods by 13.82% in average F1 score, demonstrating the feasibility and effectiveness of transferring knowledge from pretrained LLMs to enhance HAR. Further, LLM4HAR has been adopted by JD Logistics to support downstream applications such as Courier Welfare Improvement and Map Data Generation. Zhiqing Hong, Yiwei Song, Anlan Yu, Shuxin Zhong, Yi Ding 0011, Tian He 0001, Desheng Zhang 0002 |
KDD (2) | 5 |
| 2025 | A Fraudulent Blind Shipment Detection Framework in LogisticsabstractAn emerging type of fraud involves malicious senders exploiting the blind shipment and cash-on-delivery (COD) mechanisms by dispatching large volumes of unsolicited, low-cost parcels. If unsuspecting receivers accept these parcels, they pay for both shipping and goods; otherwise, logistics providers bear the round-trip shipping costs. Existing detection techniques, which rely on extensive labeled cases, struggle with this emerging fraud because receivers' unawareness and low transaction values discourage complaints, resulting in few confirmed cases. Therefore, we propose leveraging receivers' complaints, though not initially collected for fraud detection, to uncover subtle indicators of fraud patterns, while addressing three challenges: (C1) noise-rich dialogues(C2) data privacy concerns, and (C3) ever-evolving fraud patterns. To address them, we design BLOFF, a Blind shipment detection Framework for LO gistics Fraud powered by large language models (LLMs). Specifically, BLOFF includes three components: i) Sensitivity Anonymization to protect sensitive user information; ii) Dialogue Profile Distillation to transform informal dialogues into structured representation, addressing C1, and distill knowledge from a teacher LLM (GPT-4o) to a lightweight student LLM (ChatGLM4-9B), addressing C2; ii) Multi-faceted Context Augmentation to enhance the interpretation of fraud signatures and adaptation of evolving patterns, addressing C3. We evaluate BLOFF on about 56,000 complaints records collected from JD Logistics between January and November 2024. Results show that BLOFF outperforms state-of-the-art methods, achieving a 10.19% improvement in precision. Furthermore, during its real-world deployment in December 2024, BLOFF identified over 90 fraudulent parcels with a 91.4% precision. Shuxin Zhong, Zhiqing Hong, Wenjun Lyu, Qipeng Xie, Haotian Wang 0008, Lu Wang 0002, Kaishun Wu |
KDD (2) | 2 |
| 2025 | InCo: Exploring Inter-Trip Cooperation for Efficient Last-mile DeliveryabstractAn efficient last-mile delivery scheme in logistics benefits customers, couriers, and the platform. In practice, the delivery scope of a delivery station is divided into multiple areas, each of which is covered by a courier. The long distances between the delivery station and areas limit the couriers' delivery efficiency given that they need to travel back and forth multiple times a day. To solve this problem, we explore an inter-trip cooperation scheme for last-mile delivery, in which couriers traveling to the delivery station and back to corresponding areas earlier can help to take others' orders back. Coordinating the courier cooperation is challenging because we need to consider the courier's status, e.g., locations, and vehicle capacity constraint simultaneously. In this work, we design an inter-trip cooperation-based last-mile delivery system, InCo, aiming to minimize the average order delivery time. InCo includes two components: i) a time-aware spanning tree algorithm to generate the cooperation result for a group of couriers; and ii) a capacity-constrained courier grouping algorithm to optimize the courier grouping result iteratively. Extensive evaluation results with real-world order data collected from one of the largest logistics companies show that InCo improves the average saved delivery time and reduces average travel time by up to 80.2% and 28.4%, respectively, compared to baseline methods. The deployment results show InCo improves the average courier working efficiency by 21.6% to the state-of-the-practice. Wenjun Lyu, Shuxin Zhong, Guang Yang 0028, Haotian Wang 0008, Yi Ding 0011, Shuai Wang 0008, Yunhuai Liu, Tian He 0001, Desheng Zhang 0002 |
WWW | 2 |
| 2024 | AdaTrans: Adaptive Transfer Time Prediction for Multi-modal Transportation ModesabstractMulti-modal transportation leverages the advantages of various transportation modes, leading to more efficient urban traveling services. Accurately predicting transfer times between different modes provides guidance for tasks such as trip planning and transportation management. Most existing transfer time prediction works rely on strong assumptions, e.g., predetermined routes, assumed speeds, and predefined downstream transportation timetables. However, these assumptions are hard to hold in practice due to internal factors like individual preferences and external factors like dynamic traffic conditions. These factors are dynamic and vary with location and time, presenting a significant challenge. To address this, we introduce an adaptive transfer time prediction framework, AdaTrans, to forecast personalized transfer times between upstream and downstream transportation modes. Firstly, an attribute learning module is designed to model the trends of internal factors. Then a spatial-temporal adaptive learning component is designed to learn dynamic external factors. Finally, an aggregation component with a capsule network is employed to fuse the influences of these factors. The extensive evaluation results in two real-world datasets demonstrate that AdaTrans effectively harnesses insights from internal and external factors, outperforming state-of-the-art methods by ~20%. Shuxin Zhong, Hua Wei 0001, Wenjun Lyu, Guang Yang 0028, Zhiqing Hong, Guang Wang 0001, Yu Yang 0010, Desheng Zhang 0002 |
CIKM | 1 |
| 2024 | A Behavior-aware Cause Identification Framework for Order Cancellation in Logistics ServiceabstractLogistics platforms provide real-time door-to-door order pickup services to enhance customer convenience. However, a high volume of unexpected order cancellations negatively impacts both customer satisfaction and logistics profitability. Identifying whether these cancellations are due to customers' decisions or couriers' behaviors is crucial for implementing targeted operational improvements. While traditional methods directly interpret customer-courier dialogues, incorporating situational context (e.g., couriers' historical performance and current workloads) helps us to accurately understand the hidden content. The main challenges lie in dynamically correlating couriers' varying behaviors with dialogue content. To tackle this challenge, we develop COCO, a cause identification framework for order cancellation in logistics, which includes: i) Multi-modal features exploration, which analyzes dialogues and couriers' behaviors (both historical and current); ii) Multi-modal features aggregation, which uses a hierarchical attention mechanism to adaptively capture the dynamic correlations within dialogues and behaviors; iii) LLM-enhanced refinement, which leverages Large Language Models to accurately process a large number of unlabeled dialogues, significantly enhancing COCO's generalization and performance. Our extensive evaluation with JD Logistics demonstrates COCO's exceptional performance, achieving an 12.2% increase in precision and a 9.1% improvement in recall over existing methods. Furthermore, after deploying COCO at JD Logistics, it has achieved an accuracy of 89.5%, further demonstrating its practical utility. Shuxin Zhong, Yahan Gu, Wenjun Lyu, Guang Yang 0028, Guang Wang 0001, Yu Yang 0010, Desheng Zhang 0002 |
CIKM | 1 |
| 2024 | Adaptive Cross-platform Transportation Time Prediction for LogisticsabstractAccurate prediction of order transportation time is essential for customer satisfaction in logistics. Existing methods based on origin-destination (OD) pairs do not consider the diversity of road segments, while route-based methods may fail to account for real-time traffic conditions due to the infrequent dispatch schedules of logistics vehicles. In reality, e-commerce platforms have collaborated with multiple logistics companies for parcel delivery, providing a richer dataset that offers a more comprehensive view of real-time transportation conditions. The key insight is that data from one company can serve as internal capability detectors and data from others can act as external environment detectors. However, a significant challenge arises in inferring travel-time-correlated station pairs across different companies, especially without full disclosure of station information. To address this, we design an Adaptive cross-platform Transportation time prediction framework built upon a hypergraph structure, named AdaTrans, comprising: i) A spatial-temporal routing graph learner employs node-centric and edge-centric hyperedges to address the complex, non-pairwise correlations among stations and station pairs within and across companies; ii) A spatial-temporal graph-based transportation time predictor that utilizes multi-task learning to enhance overall transportation time prediction by leveraging the correlations between interconnected sub-tasks (i.e., dwell and travel times prediction) Extensive evaluation with real-world data collected from JD.com, a leading e-commerce platform in China, demonstrates that consolidating records from other companies reduces RMSE, MAE, and MAPE by 12.63%, 5.18%, and 16.67%, compared to state-of-the-art methods. Shuxin Zhong, Wenjun Lyu, Zhiqing Hong, Guang Yang 0028, Weijian Zuo, Haotian Wang 0008, Guang Wang 0001, Yu Yang 0010, Desheng Zhang 0002 |
CIKM | 1 |
| 2024 | Improving Network Robustness via Cellular Infrastructure Sharing: An Empirical Study of Infrastructure Failure with All Cellular Operators in a CityabstractIndividual cellular networks have been very robust to random cell tower failure due to redundant cell tower deployments. However, a large-scale clustered failure (e.g., due to fiber cut or cyber attacks) with multiple cell towers can lead to the loss of services of a cellular network. Recently, off-the-shelf smartphones can support multiple network standards, so cellular network infrastructure sharing is a promising direction to improve the service robustness under potential large-scale clustered cell tower failure. The existing work on cellular network robustness is usually limited to large-scale studies of individual networks or small-scale studies of multiple networks. In this work, we conduct the first investigation, to our knowledge, into the benefits of cross-network infrastructure sharing for enhancing robustness at a full cellular penetration rate. We design a new metric to quantify cellular network robustness with or without cross-network sharing under both random and clustered cell tower failures. We further study the impact of spatial dynamics on cellular network robustness. Zhihan Fang, Guang Yang 0028, Wenjun Lyu, Zhiqing Hong, Shuxin Zhong, Weijian Zuo, Yu Yang 0010, Guang Wang 0001, Desheng Zhang 0002 |
SIGSPATIAL/GIS | 5 |
| 2023 | Joint Rebalancing and Charging for Shared Electric Micromobility Vehicles with Energy-informed DemandabstractShared electric micromobility (e.g., shared electric bikes and electric scooters), as an emerging way of urban transportation, has been increasingly popular in recent years. However, managing thousands of micromobility vehicles in a city, such as rebalancing and charging vehicles to meet spatial-temporally varied demand, is challenging. Existing management frameworks generally consider demand as the number of requests without the energy consumption of these requests, which can lead to less effective management. To address this limitation, we design RECOMMEND, a rebalancing and charging framework for shared electric micromobility vehicles with energy-informed demand to improve the system revenue. Specifically, we first re-define the demand from the perspective of energy consumption and predict the future energy-informed demand based on the state-of-the-art spatial-temporal prediction method. Then we fuse the predicted energy-informed demand into different components of a rebalancing and charging framework based on reinforcement learning. We evaluate the RECOMMEND system with 2-month real-world electric micromobility system operation data. Experimental results show that our method can be easily integrated into a general RL framework and outperform state-of-the-art baselines by at least 26.89% in terms of net revenue. Heng Tan, Yukun Yuan 0001, Shuxin Zhong, Yu Yang 0010 |
CIKM | 3 |
| 2023 | RLIFE: Remaining Lifespan Prediction for E-scooters
Shuxin Zhong, William Yubeaton, Wenjun Lyu, Guang Wang 0001, Desheng Zhang 0002, Yu Yang 0010 |
CIKM | 1 |
| 2021 | Data-Driven Fairness-Aware Vehicle Displacement for Large-Scale Electric Taxi FleetsabstractWe are witnessing a rapid taxi electrification process due to the ever-increasing concern about urban air quality and energy security. A key difference between conventional gas taxis and electric taxis is their energy replenishment mechanisms, i.e., refueling or charging, which is reflected in two aspects: (i) much longer charging processes vs. short refueling processes and (ii) time-varying electricity prices vs. time-invariant gasoline prices during a day. The complicated charging issues (e.g., long charging time and dynamic charging pricing) potentially reduce electric taxis' daily operation time and profits, and also cause overcrowded charging stations during some off-peak charging pricing periods. Motivated by a set of findings obtained from a data-driven investigation, in this paper, we design a fairness-aware vehicle displacement system called FairMove to improve the overall profit efficiency and profit fairness of electric taxi fleets by considering both the passenger travel demand and taxi charging demand. We first formulate the electric taxi displacement problem as multi-agent deep reinforcement learning, and then we propose a centralized multi-agent actor-critic approach to tackle this problem. More importantly, we implement and evaluate FairMove with real-world streaming data from the Chinese city Shenzhen, including GPS data and transaction data from more than 20,100 electric taxis, coupled with the data of 123 charging stations, which constitute, to our knowledge, the largest all-electric taxi network in the world. The extensive experimental results show that our fairness-aware FairMove effectively improves the profit efficiency and profit fairness of the Shenzhen electric taxi fleet by 25.2% and 54.7%, respectively. Guang Wang 0001, Shuxin Zhong, Shuai Wang 0008, Fei Miao, Zheng Dong 0002, Desheng Zhang 0002 |
ICDE | 2 |