Stefan Lessmann

dblp:07/3592 · DBLP profile ↗
← Back
28ranked-venue papers
6as first author
9since 2021 · last 2026
0000-0001-7685-262XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2026 In the beginning was the Word: LLM-VaR and LLM-ES
abstract
This study introduces LLM-VaR and LLM-ES , novel risk estimation metrics that utilize general-purpose large language models (LLMs) for the forecasting tasks of Value at Risk (VaR) and Expected Shortfall (ES) in a zero-shot setting. Building on the input encoding mechanism of the LLMTime framework, we extend its application by defining new financial risk measures and performing an empirical evaluation of three generations of GPT models, GPT-3.5, GPT-4 and GPT-4o, versus advanced benchmark models such as GARCH with Student innovations and EWMA with Dynamic Conditional Score (DCS). Financial time series are encoded as numerical strings, allowing for model-free inference without requiring retraining. Results show that LLMs perform well when short rolling windows are used, particularly in volatile markets like cryptocurrencies. GPT-3.5 frequently outperforms or matches the performance of newer models, raising questions about model complexity, alignment, and biases. In contrast, performance deteriorates with longer windows, where the econometric models prove more reliable. Our findings demonstrate the potential of general-purpose LLMs as adaptive tools for short-horizon financial risk assessment and contribute a first-of-its-kind benchmark for LLM-based VaR/ES estimation.
Daniel Traian Pele, Vlad Bolovaneanu, Min-Bin Lin, Andrei Theodor Ginavar, Bruno Spilak, Alexandru-Victor Andrei, Filip-Mihai Toma, Stefan Lessmann, Wolfgang K. Härdle
Expert Syst. Appl.9
2026 Interpretable, multidimensional evaluation framework for causal discovery from observational i.i.d. data
abstract
Nonlinear causal discovery from observational data imposes strict identifiability assumptions on the formulation of structural equations utilized in the data-generating process. However, in real-life settings, the ground-truth mechanism responsible for cause-effect transformations is unknown. Thus, it is impossible to verify its identifiability. This is the first research to assess the performance of structure learning algorithms from seven different families in non-identifiable settings with an increasing degree of nonlinearity. The evaluation of structure learning methods under assumption violations requires a rigorous and interpretable approach that quantifies both the structural similarity of the estimation with the ground truth and the capacity of the discovered graphs to be used for causal inference. Motivated by the lack of a unified performance assessment indicator, we propose an interpretable, multidimensional evaluation framework, specifically tailored to the field of causal discovery from i.i.d. data. In particular, we introduce a six-dimensional evaluation metric, called distance to the optimal solution, which aims at providing a holistic overview of the performance of structure learning techniques. Our large-scale simulation study, which incorporates seven experimental factors, shows that hybrid Bayesian networks outperform most recently introduced continuous optimization techniques under certain conditions. Additionally, causal order-based methods yield results with comparatively high proximity to the optimal solution. • Our framework evaluates 14 causal discovery models in non-identifiable settings. • Hybrid Bayesian networks outperform most continuous optimization models. • Causal order-based structure learning achieves the current SOTA performance.
Georg Velev, Stefan Lessmann
Inf. Sci.2
2024 Fx-spot predictions with state-of-the-art transformer and time embeddings
abstract
The transformer architecture with its attention mechanism is the state-of-the-art deep learning method for sequence learning tasks and has achieved superior results in many areas such as NLP. Utilizing the transformer architecture for the prediction of sequential time series such as financial time series has hardly been investigated in previous studies. In this research paper, the transformer architecture with time embeddings is used in foreign exchange (FX) trading, the world’s largest financial market, and tests its suitability. A systematic comparison is made between transformer and benchmark models. It also examined which influence multivariate, cross-sectional input data have on the forecasting performance of the various models. The goal of the paper is to contribute to the empirical literature on FX forecasting by introducing a transformer with time embeddings to the forecasting community and assessing the accuracy of corresponding models by forecasting exchange rate movements. Empirical results indicate the suitability of transformer models for FX-Spot forecasting in general but also evidence the need for transformer models for multivariate, cross-sectional input data to outperform other state-of-the-art neural networks such as LSTM.
Tizian Fischer, Marius Sterling, Stefan Lessmann
Expert Syst. Appl.3
2024 The Deep Promotion Time Cure Model
abstract
We propose a novel method for predicting time-to-event data in the presence of cure fractions based on flexible survival models integrated into a deep neural network (DNN) framework. Our approach allows for nonlinear relationships and high-dimensional interactions between covariates and survival and is suitable for large-scale applications. To ensure the identifiability of the overall predictor formed of an additive decomposition of interpretable linear and nonlinear effects and potential higher-dimensional interactions captured through a DNN, we employ an orthogonalization layer. We demonstrate the usefulness and computational efficiency of our method via simulations and apply it to a large portfolio of U.S. mortgage loans. Here, we find not only a better predictive performance of our framework but also a more realistic picture of covariate effects.
Victor Medina-Olivares, Stefan Lessmann, Nadja Klein
IEEE Trans. Neural Networks Learn. Syst.2
2022 Modeling Irregular Time Series with Continuous Recurrent Units
abstract
Recurrent neural networks (RNNs) are a popular choice for modeling sequential data. Modern RNN architectures assume constant time-intervals between observations. However, in many datasets (e.g. medical records) observation times are irregular and can carry important information. To address this challenge, we propose continuous recurrent units (CRUs) {–} a neural architecture that can naturally handle irregular intervals between observations. The CRU assumes a hidden state, which evolves according to a linear stochastic differential equation and is integrated into an encoder-decoder framework. The recursive computations of the CRU can be derived using the continuous-discrete Kalman filter and are in closed form. The resulting recurrent architecture has temporal continuity between hidden states and a gating mechanism that can optimally integrate noisy observations. We derive an efficient parameterization scheme for the CRU that leads to a fast implementation f-CRU. We empirically study the CRU on a number of challenging datasets and find that it can interpolate irregular time series better than methods based on neural ordinary differential equations.
Mona Schirmer, Mazin Eltayeb, Stefan Lessmann, Maja Rudolph
ICML3
2021 Uplift modeling with value-driven evaluation metrics
Robin Marco Gubela, Stefan Lessmann
Decis. Support Syst.2
2021 Conditional Wasserstein GAN-based oversampling of tabular data for imbalanced learning
Justin Engelmann, Stefan Lessmann
Expert Syst. Appl.2
2021 Enterprise-grade protection against e-mail tracking
Benjamin Fabian, Benedict Bender, Ben Hesseldieck, Johannes Haupt, Stefan Lessmann
Inf. Syst.5
2021 Targeting customers for profit: An ensemble learning framework to support marketing decision-making
Stefan Lessmann, Johannes Haupt, Kristof Coussement, Koen W. De Bock
Inf. Sci.1
2020 Deep learning for detecting financial statement fraud
Patricia Craja, Alisa Kim, Stefan Lessmann
Decis. Support Syst.3
2020 Antisocial online behavior detection using deep learning
Elizaveta Zinovyeva, Wolfgang K. Härdle, Stefan Lessmann
Decis. Support Syst.3
2020 Predicting online shopping behaviour from clickstream data using deep learning
abstract
Clickstream data is an important source to enhance user experience and pursue business objectives in e-commerce. The paper uses clickstream data to predict online shopping behavior and target marketing interventions in real-time. Such AI-driven targeting has proven to save huge amounts of marketing costs and raise shop revenue. Previous user behavior prediction models rely on supervised machine learning (SML). Conceptually, SML is less suitable because it cannot account for the sequential structure of clickstream data. The paper proposes a methodology capable of unlocking the full potential of clickstream data using the framework of recurrent neural networks (RNNs). An empirical evaluation based on real-world e-commerce data systematically assesses multiple RNN classifiers and compares them to SML benchmarks. To this end, the paper proposes an approach to measure the revenue impact of a targeting model. Estimates of revenue impact together with results of standard classifier performance metrics evidence the viability of RNN-based clickstream modeling and guide employing deep recurrent learners for campaign targeting. Given that the empirical analysis shows RNN-based and conventional classifiers to capture different patterns in clickstream data, a specific recommendation is to combine sequence and conventional classifiers in an ensemble. The paper shows such an ensemble to consistently outperform the alternative models considered in the study.
Dennis Koehn, Stefan Lessmann, Markus Schaal
Expert Syst. Appl.2
2019 Shallow Self-learning for Reject Inference in Credit Scoring
abstract
Credit scoring models support loan approval decisions in the financial services industry. Lenders train these models on data from previously granted credit applications, where the borrowers' repayment behavior has been observed. This approach creates sample bias. The scoring model (i.e., classifier) is trained on accepted cases only. Applying the resulting model to screen credit applications from the population of all borrowers degrades model performance. Reject inference comprises techniques to overcome sampling bias through assigning labels to rejected cases. The paper makes two contributions. First, we propose a self-learning framework for reject inference. The framework is geared toward real-world credit scoring requirements through considering distinct training regimes for iterative labeling and model training. Second, we introduce a new measure to assess the effectiveness of reject inference strategies. Our measure leverages domain knowledge to avoid artificial labeling of rejected cases during strategy evaluation. We demonstrate this approach to offer a robust and operational assessment of reject inference strategies. Experiments on a real-world credit scoring data set confirm the superiority of the adjusted self-learning framework over regular self-learning and previous reject inference strategies. We also find strong evidence in favor of the proposed evaluation measure assessing reject inference strategies more reliably, raising the performance of the eventual credit scoring model.
Nikita Kozodoi, Panagiotis Katsas, Stefan Lessmann, Luís Moreira-Matias, Konstantinos Papakonstantinou
ECML/PKDD (3)3
2019 A multi-objective approach for profit-driven feature selection in credit scoring
Nikita Kozodoi, Stefan Lessmann, Konstantinos Papakonstantinou, Yiannis Gatsoulis, Bart Baesens
Decis. Support Syst.2
2018 Improving crime count forecasts using Twitter and taxi data
Lara Vomfell, Wolfgang K. Härdle, Stefan Lessmann
Decis. Support Syst.3
2018 Changing perspectives: Using graph metrics to predict purchase probabilities
Annika Baumann, Johannes Haupt, Fabian Gebert, Stefan Lessmann
Expert Syst. Appl.4
2017 A comparative analysis of data preparation algorithms for customer churn prediction: A case study in the telecommunication industry
Kristof Coussement, Stefan Lessmann, Geert Verstraeten
Decis. Support Syst.2
2017 Extreme learning machines for credit scoring: An empirical evaluation
Artem Bequé, Stefan Lessmann
Expert Syst. Appl.2
2017 Approaches for credit scorecard calibration: An empirical analysis
Artem Bequé, Kristof Coussement, Ross W. Gayler, Stefan Lessmann
Knowl. Based Syst.4
2016 Bridging the divide in financial market forecasting: machine learners vs. financial economists
Ming-Wei Hsu, Stefan Lessmann, Ming-Chien Sung, Tiejun Ma, Johnnie E. V. Johnson
Expert Syst. Appl.2
2012 Save the best for last? The treatment of dominant predictors in financial forecasting
Ming-Chien Sung, Stefan Lessmann
Expert Syst. Appl.2
2011 Tuning metaheuristics: A data mining based approach for particle swarm optimization
Stefan Lessmann, Marco Caserta, Idel Montalvo Arango
Expert Syst. Appl.1
2009 Feature Selection in Marketing Applications
Stefan Lessmann, Stefan Voß 0001
ADMA1
2008 Benchmarking Classification Models for Software Defect Prediction: A Proposed Framework and Novel Findings
abstract
Software defect prediction strives to improve software quality and testing efficiency by constructing predictive classification models from code attributes to enable a timely identification of fault-prone modules. Several classification models have been evaluated for this task. However, due to inconsistent findings regarding the superiority of one classifier over another and the usefulness of metric-based classification in general, more research is needed to improve convergence across studies and further advance confidence in experimental results. We consider three potential sources for bias: comparing classifiers over one or a small number of proprietary data sets, relying on accuracy indicators that are conceptually inappropriate for software defect prediction and cross-study comparisons, and, finally, limited use of statistical testing procedures to secure empirical findings. To remedy these problems, a framework for comparative software defect prediction experiments is proposed and applied in a large-scale empirical comparison of 22 classifiers over 10 public domain data sets from the NASA Metrics Data repository. Overall, an appealing degree of predictive accuracy is observed, which supports the view that metric-based classification is useful. However, our results indicate that the importance of the particular classification algorithm may be less than previously assumed since no significant performance differences could be detected among the top 17 classifiers.
Stefan Lessmann, Bart Baesens, Christophe Mues, Swantje Pietsch
IEEE Trans. Software Eng.1
2006 Forecasting with Computational Intelligence - An Evaluation of Support Vector Regression and Artificial Neural Networks for Time Series Prediction
abstract
Recently, novel algorithms of support vector regression and neural networks have received increasing attention in time series prediction. While they offer attractive theoretical properties, they have demonstrated only mixed results within real world application domains of particular time series structures and patterns. Commonly, time series are composed of a combination of regular patterns such as levels, trends and seasonal variations. Thus, the capability of novel methods to predict basic time series patterns is of particular relevance in evaluating their initial contribution to forecasting. This paper investigates the accuracy of competing forecasting methods of NN and SVR through an exhaustive empirical comparison of alternatively tuned candidate models on 36 artificial time series. Results obtained show that SVR and NN provide comparative accuracy and robustly outperform statistical methods on selected time series patterns.
Sven F. Crone, Stefan Lessmann, Swantje Pietsch
IJCNN2
2006 An Evaluation of Discrete Support Vector Machines for Cost-Sensitive Learning
abstract
The problem of cost-sensitive learning involves classification analysis in scenarios where different error types are associated with asymmetric misclassification costs. Business applications and problems of medical diagnosis are prominent examples and pattern recognitions techniques are routinely used to support decision making within these fields. In particular, support vector machines (SVMs) have been successfully applied, e.g. to evaluate customer credit worthiness in credit scoring or detect tumorous cells in bio-molecular data analysis. However, ordinary SVMs minimize a continuous approximation for the classification error giving similar importance to each error type. While several modifications have been proposed to make SVMs cost-sensitive the impact of the approximate error measurement is normally not considered. Recently, Orsenigo and Vercellis introduced a discrete SVM (DSVM) formulation [1] that minimize misclassification errors directly and overcomes possible limitations of an error proxy. For example, DSVM facilitates explicit cost minimization so that this technique is a promising candidate for cost-sensitive learning. Consequently, we compare DSVM with a standard procedure for cost-sensitive SVMs and investigate to what extent improvements in terms of misclassification costs are achievable. While the standard SVM performs remarkably well DSVM is found to give yet superior results.
Stefan Lessmann, Sven F. Crone, Robert Stahlbock
IJCNN1
2006 Genetic Algorithms for Support Vector Machine Model Selection
abstract
The support vector machine is a powerful classifier that has been successfully applied to a broad range of pattern recognition problems in various domains, e.g. corporate decision making, text and image recognition or medical diagnosis. Support vector machines belong to the group of semiparametric classifiers. The selection of appropriate parameters, formally known as model selection, is crucial to obtain accurate classification results for a given task. Striving to automate model selection for support vector machines we apply a meta-strategy utilizing genetic algorithms to learn combined kernels in a data-driven manner and to determine all free kernel parameters. The model selection criterion is incorporated into a fitness function guiding the evolutionary process of classifier construction. We consider two types of criteria consisting of empirical estimators or theoretical bounds for the generalization error. We evaluate their effectiveness in an empirical study on four well known benchmark data sets to find that both are applicable fitness measures for constructing accurate classifiers and conducting model selection. However, model selection focuses on finding one best classifier while genetic algorithms are based on the idea of re-combining and mutating a large number of good candidate classifiers to realize further improvements. It is shown that the empirical estimator is the superior fitness criterion in this sense, leading to a greater number of promising models on average.
Stefan Lessmann, Robert Stahlbock, Sven F. Crone
IJCNN1
2004 Empirical comparison and evaluation of classifier performance for data mining in customer relationship management
abstract
In competitive consumer markets, data mining for customer relationship management faces the challenge of systematic knowledge discovery in large data streams to achieve operational, tactical and strategic competitive advantages. Methods from computational intelligence, most prominently artificial neural networks and support vector machines, compete with established statistical methods in the domain of classification tasks. As both methods allow extensive degrees of freedom in the model building process, we analyse their comparative performance and sensitivity towards data pre-processing in a real-world scenario.
Sven F. Crone, Stefan Lessmann, Robert Stahlbock
IJCNN2