Xumeng Wen

dblp:358/9194 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2024
0009-0003-1758-4911ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Deep learning architectures and training · 51% Language models and text generation · 22% Time series and sequential data · 17%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Performance modeling and evaluation · 100%

Topics — the 9 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
large language model
0.812024
From Supervised to Generative: A Novel Paradigm for Tabular Deep Learning with Large Language Models · KDD 2024
Machine learning › Deep learning architectures and training
tabular data learning
0.812024
From Supervised to Generative: A Novel Paradigm for Tabular Deep Learning with Large Language Models · KDD 2024
Machine learning › Time series and sequential data › time series analysis
time series forecasting
0.812024
ElasTST: Towards Robust Varied-Horizon Forecasting with Elastic Time-Series Transformer · NeurIPS 2024
Machine learning › Deep learning architectures and training › transformer › temporal transformer
time series transformer
0.812024
ElasTST: Towards Robust Varied-Horizon Forecasting with Elastic Time-Series Transformer · NeurIPS 2024
Machine learning › Deep learning architectures and training
transformer
0.812024
ElasTST: Towards Robust Varied-Horizon Forecasting with Elastic Time-Series Transformer · NeurIPS 2024
Performance modeling and evaluation
benchmarking
0.812024
ProbTS: Benchmarking Point and Distributional Forecasting across Diverse Prediction Horizons · NeurIPS 2024
Performance modeling and evaluation › benchmarking › machine learning benchmarking
time series forecasting benchmark
0.812024
ProbTS: Benchmarking Point and Distributional Forecasting across Diverse Prediction Horizons · NeurIPS 2024
Natural language and speech › Language models and text generation
in-context learning
0.212024
From Supervised to Generative: A Novel Paradigm for Tabular Deep Learning with Large Language Models · KDD 2024
Machine learning › Transfer learning and domain adaptation
zero-shot transfer
0.212024
From Supervised to Generative: A Novel Paradigm for Tabular Deep Learning with Large Language Models · KDD 2024

Methods — techniques the papers use, named apart from their topics

deep learning models · 1.5rotary position embedding · 0.8prompt-based learning · 0.8pre-training · 0.8multi-scale patch · 0.8in-context learning · 0.8horizon reweighting · 0.8
YearPublicationVenuePosition
2024 From Supervised to Generative: A Novel Paradigm for Tabular Deep Learning with Large Language Models
abstract
Tabular data is foundational to predictive modeling in various crucial industries, including healthcare, finance, retail, sustainability, etc. Despite the progress made in specialized models, there is an increasing demand for universal models that can transfer knowledge, generalize from limited data, and follow human instructions. These are challenges that current tabular deep learning approaches have not fully tackled. Here we introduce Generative Tabular Learning (GTL), a novel framework that integrates the advanced functionalities of large language models (LLMs)-such as prompt-based zero-shot generalization and in-context learning-into tabular deep learning. GTL capitalizes on the pre-training of LLMs on diverse tabular data, enhancing their understanding of domain-specific knowledge, numerical sequences, and statistical dependencies critical for accurate predictions. Our empirical study spans 384 public datasets, rigorously analyzing GTL's convergence and scaling behaviors and assessing the impact of varied data templates. The GTL-enhanced LLaMA-2 model demonstrates superior zero-shot and in-context learning capabilities across numerous classification and regression tasks. Notably, it achieves this without fine-tuning, outperforming traditional methods and rivaling state-of-the-art models like GPT-4 in certain cases. Through GTL, we not only foster a deeper integration of LLMs' sophisticated abilities into tabular data comprehension and application but also offer a new training resource and a test bed for LLMs to enhance their ability to comprehend tabular data. To facilitate reproducible research, we release our code, data, and model checkpoints at https://github.com/microsoft/Industrial-Foundation-Models.
Xumeng Wen, Shun Zheng 0001, Wei Xu 0005, Jiang Bian 0002
KDD1
2024 ProbTS: Benchmarking Point and Distributional Forecasting across Diverse Prediction Horizons
abstract
Delivering precise point and distributional forecasts across a spectrum of prediction horizons represents a significant and enduring challenge in the application of time-series forecasting within various industries.Prior research on developing deep learning models for time-series forecasting has often concentrated on isolated aspects, such as long-term point forecasting or short-term probabilistic estimations. This narrow focus may result in skewed methodological choices and hinder the adaptability of these models to uncharted scenarios.While there is a rising trend in developing universal forecasting models, a thorough understanding of their advantages and drawbacks, especially regarding essential forecasting needs like point and distributional forecasts across short and long horizons, is still lacking.In this paper, we present ProbTS, a benchmark tool designed as a unified platform to evaluate these fundamental forecasting needs and to conduct a rigorous comparative analysis of numerous cutting-edge studies from recent years.We dissect the distinctive data characteristics arising from disparate forecasting requirements and elucidate how these characteristics can skew methodological preferences in typical research trajectories, which often fail to fully accommodate essential forecasting needs.Building on this, we examine the latest models for universal time-series forecasting and discover that our analyses of methodological strengths and weaknesses are also applicable to these universal models.Finally, we outline the limitations inherent in current research and underscore several avenues for future exploration.
Jiawen Zhang 0001, Xumeng Wen, Shun Zheng 0001, Jia Li 0009, Jiang Bian 0002
NeurIPS2
2024 ElasTST: Towards Robust Varied-Horizon Forecasting with Elastic Time-Series Transformer
abstract
Numerous industrial sectors necessitate models capable of providing robust forecasts across various horizons. Despite the recent strides in crafting specific architectures for time-series forecasting and developing pre-trained universal models, a comprehensive examination of their capability in accommodating varied-horizon forecasting during inference is still lacking. This paper bridges this gap through the design and evaluation of the Elastic Time-Series Transformer (ElasTST). The ElasTST model incorporates a non-autoregressive design with placeholders and structured self-attention masks, warranting future outputs that are invariant to adjustments in inference horizons. A tunable version of rotary position embedding is also integrated into ElasTST to capture time-series-specific periods and enhance adaptability to different horizons. Additionally, ElasTST employs a multi-scale patch design, effectively integrating both fine-grained and coarse-grained information. During the training phase, ElasTST uses a horizon reweighting strategy that approximates the effect of random sampling across multiple horizons with a single fixed horizon setting. Through comprehensive experiments and comparisons with state-of-the-art time-series architectures and contemporary foundation models, we demonstrate the efficacy of ElasTST's unique design elements. Our findings position ElasTST as a robust solution for the practical necessity of varied-horizon forecasting.
Jiawen Zhang 0001, Shun Zheng 0001, Xumeng Wen, Xiaofang Zhou 0001, Jiang Bian 0002, Jia Li 0009
NeurIPS3