Stefan Pasch

dblp:313/1963 · DBLP profile ↗
← Back
4ranked-venue papers in the field
2as first author
4since 2021 · last 2024
0000-0002-0197-0442ORCID · corroborated

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 4 (2 first)
YearPublicationVenuePosition
2024 Text Classification with Limited Training Data: Suicide Risk Detection on Social Media
abstract
In this work, we evaluate the effectiveness of several machine learning models for text classification on small datasets, focusing on a collection of Reddit posts labeled for suicidal behavior. Unlike with larger datasets, where fine-tuning complex models is typically effective, we demonstrate that carefully engineered prompts can achieve superior classification accuracy when training data is limited. Our findings highlight the potential of prompt-based approaches for effective use in resource-constrained scenarios, offering insights for researchers tackling similar small dataset challenges in text classification.
Stefan Pasch, Jannic Cutura
IEEE Big Data1
2023 Breaking the U: Asymmetric U-Net for Object Recognition in Muon Tomography☆
abstract
We utilize a U2 net-inspired deep learning model for object detection within a scattering muon tomography setup. Capturing muon readings from scintillators situated above and beneath the examination area, we discern the angular disparities between inbound and outbound muon paths. These angular measurements serve as inputs for a neural network, proficient in forecasting the contours of objects and, to an extent, their constituent materials within the research zone. Our model handles the inherent complexity of the input data, where the angular measurements map onto a 200 x 200 space, while producing an output image of a smaller 40 x 40 resolution, resulting in an asymmetric U-Net architecture that diverges from conventional semantic segmentation models which typically maintain the same input and output dimensions.
Jannic Cutura, Stefan Pasch
IEEE Big Data2
2023 CultureBERT: Measuring Corporate Culture With Transformer-Based Language Models
abstract
This paper introduces transformer-based language models to the literature measuring corporate culture from text documents. We compile a unique data set of employee reviews that were labeled by human evaluators with respect to the information the reviews reveal about the firms’ corporate culture. Using this data set, we fine-tune state-of-the-art transformer-based language models to perform the same classification task. In out-of-sample predictions, our language models classify 17 to 30 percentage points more of employee reviews in line with human evaluators than traditional approaches of text classification. We make our models publicly available.
Stefan Pasch
IEEE Big Data2
2022 NLP for Responsible Finance: Fine-Tuning Transformer-Based Models for ESG
abstract
Evaluating companies’ performances according to environmental, social, and governance (ESG) standards has become a central task in the financial industry. We show a novel solution to fine-tune transformer-based models for the ESG domain. By combining ESG ratings with text documents from annual reports, we were able to train an ESG sentiment model that outperforms traditional text classifiers at predicting the ESG behavior of companies by up to 11 percentage points. Moreover, we show practical applications of our ESG sentiment models by predicting individual sentences and by tracking ESG-related news coverage over time.
Stefan Pasch, Daniel Ehnes
IEEE Big Data1