Wanlin Cai

dblp:365/3391 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Time series and sequential data · 46% Graph learning · 46% Deep learning architectures and training · 7%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Hardware accelerators and domain-specific architectures · 67% Processor architecture and microarchitecture · 33%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Graph learning › graph neural network › graph convolution
adaptive graph convolution
0.812024
MSGNet: Learning Multi-Scale Inter-series Correlations for Multivariate Time Series Forecasting · AAAI 2024
Machine learning › Graph learning › graph neural network
graph convolution
0.812024
MSGNet: Learning Multi-Scale Inter-series Correlations for Multivariate Time Series Forecasting · AAAI 2024
Machine learning › Time series and sequential data › time series analysis › time series forecasting
multivariate time series forecasting
0.812024
MSGNet: Learning Multi-Scale Inter-series Correlations for Multivariate Time Series Forecasting · AAAI 2024
Machine learning › Time series and sequential data › time series analysis
time series forecasting
0.812024
MSGNet: Learning Multi-Scale Inter-series Correlations for Multivariate Time Series Forecasting · AAAI 2024
Processor architecture and microarchitecture
SIMD
0.812024
SOPHGO BM1684X: A Commercial High Performance Terminal AI Processor with Large Model Support · MICRO 2024
Hardware accelerators and domain-specific architectures › tensor accelerator
tensor processing unit
0.812024
SOPHGO BM1684X: A Commercial High Performance Terminal AI Processor with Large Model Support · MICRO 2024
Machine learning › Deep learning architectures and training › attention mechanism
self-attention
0.212024
MSGNet: Learning Multi-Scale Inter-series Correlations for Multivariate Time Series Forecasting · AAAI 2024

Methods — techniques the papers use, named apart from their topics

self-attention · 0.8mixhop graph convolution · 0.8frequency-domain analysis · 0.8crossbar · 0.8TPU-MLIR toolchain · 0.8SIMD · 0.8
YearPublicationVenuePosition
2024 MSGNet: Learning Multi-Scale Inter-series Correlations for Multivariate Time Series Forecasting
abstract
Multivariate time series forecasting poses an ongoing challenge across various disciplines. Time series data often exhibit diverse intra-series and inter-series correlations, contributing to intricate and interwoven dependencies that have been the focus of numerous studies. Nevertheless, a significant research gap remains in comprehending the varying inter-series correlations across different time scales among multiple time series, an area that has received limited attention in the literature. To bridge this gap, this paper introduces MSGNet, an advanced deep learning model designed to capture the varying inter-series correlations across multiple time scales using frequency domain analysis and adaptive graph convolution. By leveraging frequency domain analysis, MSGNet effectively extracts salient periodic patterns and decomposes the time series into distinct time scales. The model incorporates a self-attention mechanism to capture intra-series dependencies, while introducing an adaptive mixhop graph convolution layer to autonomously learn diverse inter-series correlations within each time scale. Extensive experiments are conducted on several real-world datasets to showcase the effectiveness of MSGNet. Furthermore, MSGNet possesses the ability to automatically learn explainable multi-scale inter-series correlations, exhibiting strong generalization capabilities even when applied to out-of-distribution samples.
Wanlin Cai, Xianggen Liu, Jianshuai Feng
AAAI1
2024 SOPHGO BM1684X: A Commercial High Performance Terminal AI Processor with Large Model Support
abstract
This paper presents BM1684X, a cutting-edge AI processor from SOPHGO designed to meet the demanding requirements of broad AI applications. Firstly, we employ SIMD architecture with very large data width to design our TPU to reduce the area ratio of the instruction unit and greatly improves the computing power density. Secondly, the customization of special acceleration instructions within the EU enables the dynamic pipeline execution, leading to a reduction in the total number of instructions and execution time. This customization enhances the performance of TPU in processing RQ and DQ operations, crucial for AI computations. Thirdly, the CUBE array within the TPU implements the multiplication and addition operations of 64 pairs of INT8 operands in the channel dimension of the feature map. By utilizing an addition tree instead of a conventional adder, the implementation significantly reduces both area and power consumption, optimizing the efficiency of TPU. Additionally, the BM1684X processor incorporates a 64-input, 64-output, 8-bit crossbar within the lane, facilitating high-performance data gathering. This crossbar design enhances data gathering capabilities, enabling efficient data processing and manipulation within the TPU architecture. Furthermore, BM1684X offers three distinct memory access modes, showing the processor's versatility in addressing a wide range of AI processing needs and optimizing DRAM utilization for various tasks and workloads. Finally, we design a TPU-MLIR toolchain, highlighting its rich features such as unified processing of multiple frameworks, hierarchical design of model abstractions, correctness guarantees, and traceability of each transformation step. BM1684X excels in providing high-performance computing for a variety of AI models including large models, demonstrating its capabilities through comprehensive evaluations with industry-leading peers.
Peng Gao 0016, Yang Liu 0038, Jun Wang 0175, Wanlin Cai, Guangchong Shen, Zonghui Hong, Jiali Qu
MICRO4