Stefan Zohren

dblp:217/3236 · DBLP profile ↗
← Back
15ranked-venue papers
0as first author
13since 2021 · last 2025
0000-0002-3392-0394ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 12 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Theory of computation · 1
YearPublicationVenuePosition
2025 Unlocking the Power of LSTM for Long Term Time Series Forecasting
abstract
Traditional recurrent neural network architectures, such as long short-term memory neural networks (LSTM), have historically held a prominent role in time series forecasting (TSF) tasks. While the recently introduced sLSTM for Natural Language Processing (NLP) introduces exponential gating and memory mixing that are beneficial for long term sequential learning, its potential short memory issue is a barrier to applying sLSTM directly in TSF. To address this, we propose a simple yet efficient algorithm named P-sLSTM, which is built upon sLSTM by incorporating patching and channel independence. These modifications substantially enhance sLSTM's performance in TSF, achieving state-of-the-art results. Furthermore, we provide theoretical justifications for our design, and conduct extensive comparative and analytical experiments to fully validate the efficiency and superior performance of our model.
Yaxuan Kong, Zepu Wang, Yuqi Nie, Tian Zhou 0004, Stefan Zohren, Yuxuan Liang 0002, Peng Sun 0007, Qingsong Wen
AAAI5
2025 Time-MQA: Time Series Multi-Task Question Answering with Context Enhancement
abstract
Yaxuan Kong, Yiyuan Yang, Yoontae Hwang, Wenjie Du, Stefan Zohren, Zhangyang Wang, Ming Jin, Qingsong Wen. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yaxuan Kong, Yiyuan Yang, Yoontae Hwang, Stefan Zohren, Zhangyang Wang, Ming Jin 0005, Qingsong Wen
ACL (1)5
2025 Advances in Financial AI: Innovations, Risk, and Responsibility in the Era of LLMs
abstract
The finance sector is seeing a rapid increase in the application of machine learning and AI, with Large Language Models (LLMs), ESG (Environmental, Social, and Governance) investing, and AI Safety significantly reshaping the field. This workshop focuses on how these advancements intersect with core financial AI applications. We will foster interdisciplinary discussion on applying LLMs to finance, addressing challenges in multilingual and non-English markets like Korea. The event will also highlight the integration of ESG signals into algorithmic decision-making and explore AI Safety, emphasizing reliability, fairness, and explainability for AI systems in regulated financial environments. By bringing together experts from academia, industry, and regulatory bodies, the workshop aims to stimulate discussions on practical issues, ethical dilemmas, and cutting-edge research shaping financial AI's future. We welcome submissions that combine technical rigor with societal relevance in AI-driven financial decisions.
Nazanin Mehrasa, Chanyeol Choi, Chung-Chi Chen 0001, Dhagash Mehta, Stefan Zohren, Chulheum Lee, Yeonhee Lee, Eunsook Oh
CIKM6
2025 LOB-Bench: Benchmarking Generative AI for Finance - an Application to Limit Order Book Data
abstract
While financial data presents one of the most challenging and interesting sequence modelling tasks due to high noise, heavy tails, and strategic interactions, progress in this area has been hindered by the lack of consensus on quantitative evaluation paradigms. To address this, we present LOB-Bench, a benchmark, implemented in python, designed to evaluate the quality and realism of generative message-by-order data for limit order books (LOB) in the LOBSTER format. Our framework measures distributional differences in conditional and unconditional statistics between generated and real LOB data, supporting flexible multivariate statistical evaluation. The benchmark also includes features commonly used LOB statistics such as spread, order book volumes, order imbalance, and message inter-arrival times, along with scores from a trained discriminator network. Lastly, LOB-Bench contains "market impact metrics", i.e. the cross-correlations and price response functions for specific events in the data. We benchmark generative autoregressive state-space models, a (C)GAN, as well as a parametric LOB model and find that the autoregressive GenAI approach beats traditional model classes.
Peer Nagy, Sascha Frey, Bidipta Sarkar, Svitlana Vyetrenko, Stefan Zohren, Ani Calinescu, Jakob N. Foerster
ICML6
2025 The 11th Mining and Learning from Time Series (MILETS): From Classical Methods to LLMs
abstract
Time series data is now pervasive across domains such as healthcare, finance, entertainment, and transportation, driven by advances in sensing technologies that enable continuous data collection. The resulting increase in data volume and complexity poses significant challenges to traditional analysis methods, calling for the development of advanced, interdisciplinary approaches to temporal data mining. This workshop aims to: (1) identify key challenges in learning from time series data, including irregular sampling, spatiotemporal dependencies, and uncertainty quantification; (2) explore recent advances in algorithmic, statistical, theoretical, and systems-based solutions-ranging from classical methods to emerging techniques involving large language models (LLMs); and (3) foster collaboration by highlighting open problems and novel research directions in time series analysis. Bridging theory and practice, the workshop provides a platform for researchers and practitioners from academia, industry, and government to exchange ideas, discuss technical challenges, and showcase practical applications. Contributions from related areas such as AI, machine learning, data science, and statistics are strongly encouraged.
Sanjay Purushotham, Dongjin Song, Qingsong Wen, Jun Huan, Yuxuan Liang 0002, Cong Shen 0001, Stefan Zohren, Yuriy Nevmyvaka
KDD (2)7
2024 Accelerating Machine Learning for Trading Using Programmable Switches
abstract
High-frequency trading (HFT) employs cutting-edge hardware for rapid decision-making and order execution but often relies on simpler algorithms that may miss deeper market trends. Conversely, lower-frequency algorithmic trading uses machine learning (ML) for better market predictions but higher latency can negate its strategic benefits. To achieve the best of both worlds, we present an in-network ML solution that embeds ML processes into programmable network devices, accelerating feature engineering and extraction as well as ML inference. In this paper, we design and develop a solution that supports both stock mid-price and volatility movement forecasting using commodity switches. Our approach achieves microsecond-scale, ultra-low latency, significantly lowering it by 64% to 97% compared to previous works, while upholding the same level of ML performance as server models. Additionally, by combining network hardware and servers, a hybrid deployment strategy can keep the misclassification rate change below 0.8% relative to the server baseline while processing 49% of the traffic directly on the switch and achieving a 45% average reduction in end-to-end latency.
Xinpeng Hong, Changgang Zheng, Stefan Zohren, Noa Zilberman
ECAI3
2024 The 10th Mining and Learning from Time Series Workshop: From Classical Methods to LLMs
abstract
Time series data has become ubiquitous across various fields such as healthcare, finance, entertainment, and transportation, driven by advancements in sensing technologies that enable continuous monitoring and recording. This growth in data size and complexity presents new challenges for traditional analysis techniques, necessitating the development of advanced, interdisciplinary temporal mining algorithms. The goals of this workshop are to: (1) highlight significant challenges in learning and mining from time series data, such as irregular sampling, spatiotemporal structures, and uncertainty quantification; (2) discuss recent developments in algorithmic, theoretical, statistical, and systems-based approaches for addressing these challenges, including both classical methods and large language models (LLMs); and (3) synergize research efforts by exploring both new and open problems in time series analysis and mining. This workshop will focus on both the theoretical and practical aspects of time series data analysis, providing a platform for researchers and practitioners from academia, government, and industry to discuss potential research directions, critical technical issues, and present solutions for practical applications. Contributions from related fields such as AI, machine learning, data science, and statistics are also included.
Sanjay Purushotham, Dongjin Song, Qingsong Wen, Jun Huan, Cong Shen 0001, Stefan Zohren, Yuriy Nevmyvaka
KDD6
2023 LOBIN: In-Network Machine Learning for Limit Order Books
abstract
Machine learning is driving the evolution of algorithmic trading, but the demands for fast execution speed remain. Although both aim to increase profitability, embedding more powerful machine learning approaches and lowering trading latencies are hard to achieve simultaneously. Offloading machine learning inference to programmable network devices, also referred to as in-network machine learning, provides a delicate balance between the two ends of this trade-off. In this paper, we present LOBIN, providing machine learning based market prediction using high-frequency market data feeds. LOBIN builds limit order books and conducts inference within programmable switches. Compared with server-based solutions, LOBIN predicts future stock price movements with lower latency, higher throughput, and a minor impact on machine learning performance.
Xinpeng Hong, Changgang Zheng, Stefan Zohren, Noa Zilberman
HPSR3
2022 Same State, Different Task: Continual Reinforcement Learning without Interference
abstract
Continual Learning (CL) considers the problem of training an agent sequentially on a set of tasks while seeking to retain performance on all previous tasks. A key challenge in CL is catastrophic forgetting, which arises when performance on a previously mastered task is reduced when learning a new task. While a variety of methods exist to combat forgetting, in some cases tasks are fundamentally incompatible with each other and thus cannot be learnt by a single policy. This can occur, in reinforcement learning (RL) when an agent may be rewarded for achieving different goals from the same observation. In this paper we formalize this "interference" as distinct from the problem of forgetting. We show that existing CL methods based on single neural network predictors with shared replay buffers fail in the presence of interference. Instead, we propose a simple method, OWL, to address this challenge. OWL learns a factorized policy, using shared feature extraction layers, but separate heads, each specializing on a new task. The separate heads in OWL are used to prevent interference. At test time, we formulate policy selection as a multi-armed bandit problem, and show it is possible to select the best policy for an unknown task using feedback from the environment. The use of bandit algorithms allows the OWL agent to constructively re-use different continually learnt policies at different times during an episode. We show in multiple RL environments that existing replay based CL methods fail, while OWL is able to achieve close to optimal performance when training sequentially.
Samuel Kessler, Jack Parker-Holder, Philip J. Ball, Stefan Zohren, Stephen J. Roberts
AAAI4
2022 Forecasting COVID-19 Caseloads Using Unsupervised Embedding Clusters of Social Media Posts
abstract
Felix Drinkall, Stefan Zohren, Janet Pierrehumbert. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Felix Drinkall, Stefan Zohren, Janet B. Pierrehumbert
NAACL-HLT2
2022 Learning Rates as a Function of Batch Size: A Random Matrix Theory Approach to Neural Network Training
abstract
We study the effect of mini-batching on the loss landscape of deep neural networks using spiked, field-dependent random matrix theory. We demonstrate that the magnitude of the extremal values of the batch Hessian are larger than those of the empirical Hessian. We also derive similar results for the Generalised Gauss-Newton matrix approximation of the Hessian. As a consequence of our theorems we derive an analytical expressions for the maximal learning rates as a function of batch size, informing practical training regimens for both stochastic gradient descent (linear scaling) and adaptive algorithms, such as Adam (square root scaling), for smooth, non-convex deep neural networks. Whilst the linear scaling for stochastic gradient descent has been derived under more restrictive conditions, which we generalise, the square root scaling rule for adaptive optimisers is, to our knowledge, completely novel. We validate our claims on the VGG/WideResNet architectures on the CIFAR-100 and ImageNet data sets. Based on our investigations of the sub-sampled Hessian we develop a stochastic Lanczos quadrature based on the fly learning rate and momentum learner, which avoids the need for expensive multiple evaluations for these key hyper-parameters and shows good preliminary results on the Pre-Residual Architecture for CIFAR-100. We further investigate the similarity between the Hessian spectrum of a multi-layer perceptron, trained on Gaussian mixture data, compared to that of deep neural networks trained on natural images. We find striking similarities, with both exhibiting rank degeneracy, a bulk spectrum and outliers to that spectrum. Furthermore, we show that ZCA whitening can remove such outliers early on in training before class separation occurs, but that outliers persist in later training.
Diego Granziol, Stefan Zohren, Stephen J. Roberts
J. Mach. Learn. Res.2
2021 Fast Agent-Based Simulation Framework with Applications to Reinforcement Learning and the Study of Trading Latency Effects
Peter Belcak, Jan-P. Calliess, Stefan Zohren
MABS3
2021 Hierarchical Indian buffet neural networks for Bayesian continual learning
abstract
We place an Indian Buffet process (IBP) prior over the structure of a Bayesian Neural Network (BNN), thus allowing the complexity of the BNN to increase and decrease automatically. We further extend this model such that the prior on the structure of each hidden layer is shared globally across all layers, using a Hierarchical-IBP (H-IBP). We apply this model to the problem of resource allocation in Continual Learning (CL) where new tasks occur and the network requires extra resources. Our model uses online variational inference with reparameterisation of the Bernoulli and Beta distributions, which constitute the IBP and H-IBP priors. As we automatically learn the number of weights in each layer of the BNN, overfitting and underfitting problems are largely overcome. We show empirically that our approach offers a competitive edge over existing methods in CL.
Samuel Kessler, Stefan Zohren, Stephen J. Roberts
UAI3
2020 Recurrent Neural Filters: Learning Independent Bayesian Filtering Steps for Time Series Prediction
abstract
Despite the recent popularity of deep generative state space models, few comparisons have been made between network architectures and the inference steps of the Bayesian filtering framework - with most models simultaneously approximating both state transition and update steps with a single recurrent neural network (RNN). In this paper, we introduce the Recurrent Neural Filter (RNF), a novel recurrent autoencoder architecture that learns distinct representations for each Bayesian filtering step, captured by a series of encoders and decoders. Testing this on three real-world time series datasets, we demonstrate that the decoupled representations learnt improve the accuracy of one-step-ahead forecasts while providing realistic uncertainty estimates, and also facilitate multistep prediction through the separation of encoder stages.
Bryan Lim, Stefan Zohren, Stephen J. Roberts
IJCNN2
2020 Quantum Codes From Classical Graphical Models
abstract
We introduce a new graphical framework for designing quantum error correction codes based on classical principles. A key feature of this graphical language, over previous approaches, is that it is closely related to that of factor graphs or graphical models in classical information theory and machine learning. It enables us to formulate the description of the recently-introduced `coherent parity check' quantum error correction codes entirely within the language of classical information theory. This makes our construction accessible without requiring background in quantum error correction or even quantum mechanics in general. More importantly, this allows for a collaborative interplay where one can design new quantum error correction codes derived from classical codes.
Joschka Roffe, Stefan Zohren, Dominic Horsman, Nicholas Chancellor
IEEE Trans. Inf. Theory2