Shuai Zhang 0026

dblp:71/208-26 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
11since 2021 · last 2025
0000-0001-8502-2927ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 MalDetectFormer: Leveraging Sparse SpatioTemporal Information for Effective Malicious Traffic Detection
abstract
Malicious traffic detection is one of the main challenges in the field of cybersecurity. Although modern deep learning methods have made progress in identifying malicious traffic, they often overlook the persistent nature of attack behaviors, making it difficult to distinguish between malicious and normal traffic at a single observation point. To address this issue, we propose MalDetectFormer, which aims to accurately capture the spatiotemporal dynamics of malicious traffic. By incorporating a sparse attention mechanism, MalDetectFormer can efficiently focus on key characteristics of traffic nodes while overcoming the challenges faced by traditional long-sequence processing. Additionally, by adopting a time-cyclic attention mechanism, the model can identify and capture persistent attack patterns of malicious traffic. Experiments conducted on benchmark datasets demonstrate the advantages of the proposed MalDetectFormer in both malicious traffic detection and malicious attack recognition tasks.
Shuai Zhang 0026, Haoyi Zhou
AAAI1
2023 Expanding the prediction capacity in long sequence time-series forecasting
Haoyi Zhou, Jianxin Li 0002, Shanghang Zhang, Shuai Zhang 0026, Mengyi Yan, Hui Xiong 0001
Artif. Intell.4
2022 Learning Music Sequence Representation From Text Supervision
abstract
Music representation learning is notoriously difficult for its complex human-related concepts contained in the sequence of numerical signals. To excavate better MUsic SEquence Representation from labeled audio, we propose a novel text-supervision pre-training method, namely MUSER. MUSER adopts an audio-spectrum-text tri-modal contrastive learning framework, where the text input could be any form of meta-data with the help of text templates while the spectrum is derived from an audio sequence. Our experiments reveal that MUSER could be more flexibly adapted to downstream tasks compared with the current data-hungry pre-training method, and it only requires 0.056% of pre-training data to achieve the state-of-the-art performance.
Tianyu Chen 0017, Shuai Zhang 0026, Shaohan Huang, Haoyi Zhou, Jianxin Li 0002
ICASSP3
2022 AIQoSer: Building the efficient Inference-QoS for AI Services
abstract
The AI inspired methods have entirely changed the network QoS landscape and brought better demand-guided experiences for the end-users. However, the increasing demands of satisfactory experiences require larger AI models, whose inference efficiency becomes the non-negligible drawback in the time-sensitive network QoS. In this work, we defined this challenge as the inference-QoS (iQoS) problem of the network QoS itself, which balances inference efficiency and performance for AI services. We design a unified iQoS metric to evaluate the AI-enhanced QoS frameworks with considerations on model performance, inference latency, and input scale. Then, we propose a two-stage pipeline as the exemplar for leveraging the iQoS metric in QoS-aware AI services: (i) enhance reconstruction ability, pretraining masked autoencoder extracts intrinsic data correlations by multi-scale masking; (ii) improve inference efficiency, forecasting masked decoder uses the data scale pruning in terms of spatial and temporal dimension for prediction. Comprehensive experiments on our method demonstrate its superior inference latency and overwhelming traffic matrix prediction performance.
Jianxin Li 0002, Tianchen Zhu, Haoyi Zhou, Qingyun Sun, Shuai Zhang 0026, Chunming Hu
IWQoS6
2022 AutoST: Towards the Universal Modeling of Spatio-temporal Sequences
abstract
The analysis of spatio-temporal sequences plays an important role in many real-world applications, demanding a high model capacity to capture the interdependence among spatial and temporal dimensions. Previous studies provided separated network design in three categories: spatial first, temporal first, and spatio-temporal synchronous. However, the manually-designed heterogeneous models can hardly meet the spatio-temporal dependency capturing priority for various tasks. To address this, we proposed a universal modeling framework with three distinctive characteristics: (i) Attention-based network backbone, including S2T Layer (spatial first), T2S Layer (temporal first), and STS Layer (spatio-temporal synchronous). (ii) The universal modeling framework, named UniST, with a unified architecture that enables flexible modeling priorities with the proposed three different modules. (iii) An automatic search strategy, named AutoST, automatically searches the optimal spatio-temporal modeling priority by network architecture search. Extensive experiments on five real-world datasets demonstrate that UniST with any single type of our three proposed modules can achieve state-of-the-art performance. Furthermore, AutoST can achieve overwhelming performance with UniST.
Jianxin Li 0002, Shuai Zhang 0026, Hui Xiong 0001, Haoyi Zhou
NeurIPS2
2022 Jump Self-attention: Capturing High-order Statistics in Transformers
abstract
The recent success of Transformer has benefited many real-world applications, with its capability of building long dependency through pairwise dot-products. However, the strong assumption that elements are directly attentive to each other limits the performance of tasks with high-order dependencies such as natural language understanding and Image captioning. To solve such problems, we are the first to define the Jump Self-attention (JAT) to build Transformers. Inspired by the pieces moving of English Draughts, we introduce the spectral convolutional technique to calculate JAT on the dot-product feature map. This technique allows JAT's propagation in each self-attention head and is interchangeable with the canonical self-attention. We further develop the higher-order variants under the multi-hop assumption to increase the generality. Moreover, the proposed architecture is compatible with the pre-trained models. With extensive experiments, we empirically show that our methods significantly increase the performance on ten different tasks.
Haoyi Zhou, Siyang Xiao, Shanghang Zhang, Jieqi Peng, Shuai Zhang 0026, Jianxin Li 0002
NeurIPS5
2021 Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting
abstract
Many real-world applications require the prediction of long sequence time-series, such as electricity consumption planning. Long sequence time-series forecasting (LSTF) demands a high prediction capacity of the model, which is the ability to capture precise long-range dependency coupling between output and input efficiently. Recent studies have shown the potential of Transformer to increase the prediction capacity. However, there are several severe issues with Transformer that prevent it from being directly applicable to LSTF, including quadratic time complexity, high memory usage, and inherent limitation of the encoder-decoder architecture. To address these issues, we design an efficient transformer-based model for LSTF, named Informer, with three distinctive characteristics: (i) a ProbSparse self-attention mechanism, which achieves O(L log L) in time complexity and memory usage, and has comparable performance on sequences' dependency alignment. (ii) the self-attention distilling highlights dominating attention by halving cascading layer input, and efficiently handles extreme long input sequences. (iii) the generative style decoder, while conceptually simple, predicts the long time-series sequences at one forward operation rather than a step-by-step way, which drastically improves the inference speed of long-sequence predictions. Extensive experiments on four large-scale datasets demonstrate that Informer significantly outperforms existing methods and provides a new solution to the LSTF problem.
Haoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang 0026, Jianxin Li 0002, Hui Xiong 0001, Wancai Zhang
AAAI4
2021 MERITS: Medication Recommendation for Chronic Disease with Irregular Time-Series
abstract
Medication recommendation for chronic diseases based on the complex historical electronic medical records (EMR) is an important and challenging research problem in medical informatics because the medical records are often irregularly sampled and contain many missing data. However, most existing approaches fail to explore the irregular time-series dependencies and ignore the consecutive correlation in dynamic prescription history. To fill this gap, we propose the MEdication Recommendation network on Irregular Time-Series (MERITS), which captures the irregular time-series dependencies with the neural ordinary differential equations (Neural ODE). Meanwhile, it leverages a drug-drug interaction knowledge graph and two learned medication relation graphs to explore the co-occurrence and sequential correlations of the medications. We further propose an attention-based encoder-decoder framework to combine the historical information of patients and medications from EMR. Besides, we collect and annotate a diabetes inpatient medication dataset and demonstrate the effectiveness of MERITS by comparing it with several state-of-the-art methods of medication recommendations.
Shuai Zhang 0026, Jianxin Li 0002, Haoyi Zhou, Qishan Zhu, Shanghang Zhang, Danding Wang
ICDM1
2021 Triplet Attention: Rethinking the Similarity in Transformers
abstract
The Transformer model has benefited various real-world applications, where the self-attention mechanism with dot-products shows superior alignment ability on building long dependency. However, the pair-wisely attended self-attention limits further performance improvement on challenging tasks. To the extent of our knowledge, this is the first work to define the Triplet Attention (A3) for Transformer, which introduces triplet connections as the complementary dependency. Specifically, we define the triplet attention based on the scalar triplet product, which may be interchangeably used with the canonical one within the multi-head attention. It allows the self-attention mechanism to attend to diverse triplets and capture complex dependency. Then, we utilize the permuted formulation and kernel tricks to establish a linear approximation to A3. The proposed architecture could be smoothly integrated into the pre-training by modifying head configurations. Extensive experiments show that our methods achieve significant performance improvement on various tasks and two benchmarks.
Haoyi Zhou, Jianxin Li 0002, Jieqi Peng, Shuai Zhang 0026, Shanghang Zhang
KDD4
2021 A Bisubmodular Approach to Event Detection and Prediction in Multivariate Social Graphs
abstract
A burst event on a social graph is usually framed as an anomalous and unexpected pattern that is characterized as a compact or correlated subset of affected vertices, which is a subgraph. Subgraph detection becomes a serious problem when social graphs involve multiple attributes (i.e., multivariate graph). Most existing methods are not capable of handling the feature selection and subgraph detection problems simultaneously on the multivariate graph. In this article, we propose multivariate anomalous subgraph scanning (MASS), a generic model that detects anomalous events on the multivariate social graph. First, we reformulate the traditional nonparametric statistics as a new statistical objective function that simultaneously measures the significance of a vertices subset and an attributes subset to generate an indicator of ongoing or upcoming events. Then, we reformulate the objective function as the difference between two bisubmodular functions and approximate it with a bisubmodular objective function, which can be optimized in linear time, with an analysis of its theoretical properties. We demonstrate the performance of our proposed method using two burst event detection and prediction tasks from the real world.
Shuai Zhang 0026, Haoyi Zhou, Feng Chen 0001, Jianxin Li 0002
IEEE Trans. Comput. Soc. Syst.1
2021 POLLA: Enhancing the Local Structure Awareness in Long Sequence Spatial-temporal Modeling
abstract
The spatial-temporal modeling on long sequences is of great importance in many real-world applications. Recent studies have shown the potential of applying the self-attention mechanism to improve capturing the complex spatial-temporal dependencies. However, the lack of underlying structure information weakens its general performance on long sequence spatial-temporal problem. To overcome this limitation, we proposed a novel method, named the Proximity-aware Long Sequence Learning framework, and apply it to the spatial-temporal forecasting task. The model substitutes the canonical self-attention by leveraging the proximity-aware attention, which enhances local structure clues in building long-range dependencies with a linear approximation of attention scores. The relief adjacency matrix technique can utilize the historical global graph information for consistent proximity learning. Meanwhile, the reduced decoder allows for fast inference in a non-autoregressive manner. Extensive experiments are conducted on five large-scale datasets, which demonstrate that our method achieves state-of-the-art performance and validates the effectiveness brought by local structure information.
Haoyi Zhou, Hao Peng 0001, Jieqi Peng, Shuai Zhang 0026, Jianxin Li 0002
ACM Trans. Intell. Syst. Technol.4
2017 An Efficient Approach to Event Detection and Forecasting in Dynamic Multivariate Social Media Networks
abstract
Anomalous subgraph detection has been successfully applied to event detection in social media. However, the subgraph detection problembecomes challenging when the social media network incorporates abundant attributes, which leads to a multivariate network. The multivariate characteristic makes most existing methods incapable to tackle this problem effectively and efficiently, as it involves joint feature selection and subgraph detection that has not been well addressed in the current literature, especially, in the dynamic multivariate networks in which attributes evolve over time.
Minglai Shao 0001, Jianxin Li 0002, Feng Chen 0001, Hongyi Huang, Shuai Zhang 0026, Xunxun Chen
WWW5