Shuhan Zhong

dblp:218/8083 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2025
0000-0003-4037-4288ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 MTM: A Multi-Scale Token Mixing Transformer for Irregular Multivariate Time Series Classification
abstract
Irregular multivariate time series (IMTS) is characterized by the lack of synchronized observations across its different channels.In this paper, we point out that this channel-wise asynchrony can lead to poor channel-wise modeling of existing deep learning methods.To overcome this limitation, we propose MTM, a multi-scale token mixing transformer for the classification of IMTS.We find that the channel-wise asynchrony can be alleviated by down-sampling the time series to coarser timescales, and propose to incorporate a masked concat pooling in MTM that gradually down-samples IMTS to enhance the channel-wise attention modules.Meanwhile, we propose a novel channel-wise token mixing mechanism which proactively chooses important tokens from one channel and mixes them with other channels, to further boost the channel-wise learning of our model.Through extensive experiments on real-world datasets and comparison with state-of-the-art methods, we demonstrate that MTM consistently achieves the best performance on all the benchmarks, with improvements of up to 3.8% in AUPRC for classification.
Shuhan Zhong, Weipeng Zhuo, Sizhe Song, Guanyao Li, Zhongyi Yu, Shueng-Han Gary Chan
KDD (2)1
2025 TAMI: Taming Heterogeneity in Temporal Interactions for Temporal Graph Link Prediction
abstract
Temporal graph link prediction aims to predict future interactions between nodes in a graph based on their historical interactions, which are encoded in node embeddings. We observe that heterogeneity naturally appears in temporal interactions, e.g., a few node pairs can make most interaction events, and interaction events happen at varying intervals. This leads to the problems of ineffective temporal information encoding and forgetting of past interactions for a pair of nodes that interact intermittently for their link prediction. Existing methods, however, do not consider such heterogeneity in their learning process, and thus their learned temporal node embeddings are less effective, especially when predicting the links for infrequently interacting node pairs. To cope with the heterogeneity, we propose a novel framework called TAMI, which contains two effective components, namely log time encoding function (LTE) and link history aggregation (LHA). LTE better encodes the temporal information through transforming interaction intervals into more balanced ones, and LHA prevents the historical interactions for each target node pair from being forgotten. State-of-the-art temporal graph neural networks can be seamlessly and readily integrated into TAMI to improve their effectiveness. Experiment results on 13 classic datasets and three newest temporal graph benchmark (TGB) datasets show that TAMI consistently improves the link prediction performance of the underlying models in both transductive and inductive settings. Our code is available at https://github.com/Alleinx/TAMI_temporal_graph.
Zhongyi Yu, Jianqiu Wu, Shuhan Zhong, Weifeng Su, Chul-Ho Lee, Weipeng Zhuo
NeurIPS4
2024 A Multi-Scale Decomposition MLP-Mixer for Time Series Analysis
abstract
Time series data, including univariate and multivariate ones, are characterized by unique composition and complex multi-scale temporal variations. They often require special consideration of decomposition and multi-scale modeling to analyze. Existing deep learning methods on this best fit to univariate time series only, and have not sufficiently considered sub-series modeling and decomposition completeness. To address these challenges, we propose MSD-Mixer, a M ulti- S cale D ecomposition MLP- Mixer , which learns to explicitly decompose and represent the input time series in its different layers. To handle the multi-scale temporal patterns and multivariate dependencies, we propose a novel temporal patching approach to model the time series as multi-scale patches, and employ MLPs to capture intra- and inter-patch variations and channel-wise correlations. In addition, we propose a novel loss function to constrain both the mean and the autocorrelation of the decomposition residual for better decomposition completeness. Through extensive experiments on various real-world datasets for five common time series analysis tasks, we demonstrate that MSD-Mixer consistently and significantly outperforms other state-of-the-art algorithms with better efficiency.
Shuhan Zhong, Sizhe Song, Weipeng Zhuo, Guanyao Li, Yang Liu 0278, Shueng-Han Gary Chan
Proc. VLDB Endow.1
2023 A Lightweight and Accurate Spatial-Temporal Transformer for Traffic Forecasting
abstract
We study the forecasting problem for traffic with dynamic, possibly periodical, and joint spatial-temporal dependency between regions. Given the aggregated inflow and outflow traffic of regions in a city from time slots 0 to$t - 1$, we predict the traffic at time$t$for any region. Prior arts in the area often considered the spatial and temporal dependencies in a decoupled manner, or were rather computationally intensive in training with a large number of hyper-parameters which needed tuning. We propose ST-TIS, a novel, lightweight and accurateSpatial-TemporalTransformer withinformation fusion and regionsampling for traffic forecasting. ST-TIS extends the canonical Transformer with information fusion and region sampling. The information fusion module captures the complex spatial-temporal dependency between regions. The region sampling module is to improve the efficiency and prediction accuracy, cutting the computation complexity for dependency learning from$O(n^{2})$to$O(n\sqrt{n})$, where$n$is the number of regions. With far fewer parameters than state-of-the-art deep learning models, ST-TIS's offline training is significantly faster in terms of tuning and computation (with a reduction of up to$90\%$on training time and network parameters). Notwithstanding such training efficiency, extensive experiments show that ST-TIS is substantially more accurate in online prediction than state-of-the-art approaches (with an average improvement of$9.5\%$on RMSE, and$12.4\%$on MAPE compared to STDN and DSAN).
Guanyao Li, Shuhan Zhong, Xingdong Deng, Letian Xiang, Shueng-Han Gary Chan, Yang Liu 0278, Chih-Chieh Hung, Wen-Chih Peng
IEEE Trans. Knowl. Data Eng.2
2022 A Data-Driven Spatial-Temporal Graph Neural Network for Docked Bike Prediction
abstract
Docked bike systems have been widely deployed in many cities around the world. To the service provider, predicting the demand and supply of bikes at any station is crucial to offering the best service quality. The docked bike prediction problem is highly challenging because of the complicated joint spatial-temporal (ST) dependency as bikes are picked up and dropped off, the so-called “flows”, between stations. Prior works often considered the spatial and temporal dependencies separately using sequential network models, and based on locality assumptions. Without sufficiently capturing the joint spatial and temporal features, these approaches are not optimal for attaining the best prediction accuracy. We propose STGNN-DJD, a novel data-driven Spatial-Temporal Graph Neural Network to solve the bike demand and supply prediction problem by unifiedly embedding the Dynamic and Joint ST Dependency in two novel ST graphs. Given station locations and historical rental data on bike flow over the past time slots 0 to$t-1$, we seek to predict online the bike demand and supply at any station at time$t$. To extract joint spatial-temporal dependency, STGNN-DJD employs a graph generator to construct, at the beginning of time$t$, two graphs that embed the flow relationships between stations at various time slots (flow-convoluted graph) and dynamic demand-supply pattern correlation between stations (pattern correlation graph), respectively. Given the two spatial-temporal graphs, STGNN-DJD subsequently employs a graph neural network with novel flow-based and attention-based aggregators to generate embedding of each station for docked bike prediction. We have conducted extensive experiments on two large bike-sharing datasets. Our re-sults confirm the effectiveness of STGNN-DJD as compared with other state-of-the-art approaches, with significant improvement on RMSE and MAE (by 20%-50%). We also provide a case study on dynamic dependencies between stations and demonstrate that the locality assumption does not always hold for a docked bike system.
Guanyao Li, Gunarto Sindoro Njoo, Shuhan Zhong, Shueng-Han Gary Chan, Chih-Chieh Hung, Wen-Chih Peng
ICDE4
2022 A Tree-Based Structure-Aware Transformer Decoder for Image-To-Markup Generation
abstract
Image-to-markup generation aims at translating an image into markup (structured language) that represents both the contents and the structural semantics corresponding to the image. Recent encoder-decoder based approaches typically employ string decoders to model the string representation of the target markup, which cannot effectively capture the rich embedded structural information. In this paper, we propose TSDNet, a novel Tree-based Structure-aware Transformer Decoder NETwork to directly generate the tree representation of the target markup in a structure-aware manner. Specifically, our model learns to sequentially predict the node attributes, edge attributes, and node connectivities by multi-task learning. Meanwhile, we introduce a novel tree-structured attention to our decoder such that it can directly operate on the partial tree generated in each step to fully exploit the structural information. TSDNet doesn't rely on any prior assumptions on the target tree structure, and can be jointly optimized with encoders in an end-to-end fashion. We evaluate the performance of our model on public image-to-markup generation datasets, and demonstrate its ability to learn the complicated correlation from the structural information in the target markup with significant improvement over state-of-the-art methods by up to 5.6% in mathematical expression recognition and up to 35.34% in chemical formula recognition.
Shuhan Zhong, Sizhe Song, Guanyao Li, Shueng-Han Gary Chan
ACM Multimedia1
2022 vContact: Private WiFi-Based IoT Contact Tracing With Virus Lifespan
abstract
Covid-19 is primarily spread through contact with the virus, which may survive on surfaces with a lifespan of hours or even days if not sanitized. To curb its spread, it is hence of vital importance to detect those who have been in contact with the virus for a sustained period of time, the so-calledclose contacts. Most of the existing digital approaches for contact tracing focus only on direct face-to-face contacts. There has been little work on detecting indirect environmental contact, which is to detect people coming into a contaminated area with the live virus, i.e., an area last visited by an infected person within the virus lifespan. In this work, we study automatic Internet of Things (IoT) contact tracing when the virus has a lifespan, which may depend on the disinfection frequency at a location. Leveraging the ubiquity of WiFi signals, we propose vContact, a novel, private, pervasive, and fully distributed WiFi-based IoT contact tracing approach. Users carrying an IoT device (phone, wearable, dongle, etc.) continuously scan WiFi access points (APs) and store their hashed IDs. Given a confirmed case, the signals are then uploaded to a server for other users to match in their local IoT devices for virus exposure notification. vContact is not based on device pairing, and no information of other users is stored locally. The confirmed case does not need to have the device for it to work properly. As WiFi data are sampled sporadically and asynchronously, vContact uses novel and effective signal processing approaches and a similarity metric to align and match signals at any time. We conduct extensive indoor and outdoor experiments to validate vContact performance. Our results demonstrate that vContact is effective and accurate for contact detection. The precision, recall, and F1-score of contact detection are high (up to 90%) for close contact proximity (2 m). Its performance is robust against AP numbers, AP changes, and phone heterogeneity. Having implemented vContact as an Android software development kit and installed it on phones and smart watches, we present a case study to demonstrate the validity and implementability of our design in notifying its users about their exposure to the virus with a specific lifespan.
Guanyao Li, Siyan Hu, Shuhan Zhong, Wai Lun Tsui, Shueng-Han Gary Chan
IEEE Internet Things J.3