VLDB 2026 Research / reviewers in the wild / expert
Matthias Hirth
dblp:99/8851
· DBLP profile ↗
18ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0002-1359-363XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 11 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Computer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | 360-FORGE: A CGI Foundry for Omnidirectional Reference Ground-Truth in 360° Annotation Studies
Florian Schlechte, Pratigya, Matthias Hirth |
QoMEX | 3 |
| 2025 | Assessing the Influence of End-to-End Request Latency and Response Length on Waiting Time Satisfaction of LLM UsersabstractThe increasing deployment and resource consumption of Large Language Model (LLM) systems necessitate understanding the influencing factors of the perceived quality of these systems. A deep understanding of factors influencing the Quality of Experience (QoE) for LLM users enables novel approaches for resource optimization while maintaining high user experience quality. Existing research on the QoE for LLM users lacks controlled experiments isolating timing effects from other potential influencing factors. To address this gap, this paper presents a controlled experimental setup for assessing how system parameters—specifically end-to-end request latency and response length—influence user waiting time satisfaction as a key component of QoE. A purpose-built application emulated a realistic interface experience, allowing us to collect 2189 ratings from 178 crowdsourced workers. The statistical analysis reveals significant non-linear relationships between system-controlled End-to-End Request Latency, response length, and user-perceived waiting time satisfaction. The methodology presented provides an empirical foundation for future investigations, including animated response delivery and Time-to-First-Token-driven optimization strategies for LLM systems. Florian Schlechte, Yusra Tasnim Farooqi, Matthias Hirth |
QoMEX | 3 |
| 2024 | Visual Transformers Meet Convolutional Neural Networks: Providing Context for Convolution Layers in Semantic Segmentation of Remote Sensing Photovoltaic Imaging
Alejandro Libreros, Muhammad Hamza Shafiq, Edwin Gamboa, Martin Cleven, Matthias Hirth |
DaWaK | 5 |
| 2024 | Impact of feedback on crowdsourced visual quality assessment with paired comparisonsabstractThis paper presents a comprehensive investigation into the effects of immediate feedback on crowdworkers’ performance in subjective image quality assessment tasks using paired comparisons. The study is motivated by the need for reliable and efficient crowdsourcing tasks for image quality assessment. A large-scale experiment involving 200 participants was conducted, where participants completed 120 paired comparisons with and without feedback. The feedback informed the workers of the correctness of their responses to comparisons. Almost all of the participants (97%) preferred receiving feedback. The results indicate that feedback reduced response time, improved user experience, and did not cause a bias in the estimation of the just noticeable difference (JND). On the other hand, feedback did not significantly affect accuracy, correlation with the ground truth, or create a learning effect. This study contributes to the field by being one of the first to examine the impact of feedback on crowdworker performance in subjective image quality assessment tasks. The dataset which includes the images and ratings can be accessed at https://database.mmsp-kn.de/feedback-study-dataset.html. Mohsen Jenadeleh, Alexander Heß, Simon Hviid Del Pin, Edwin Gamboa, Matthias Hirth, Dietmar Saupe |
QoMEX | 5 |
| 2022 | Evaluating the Robustness of Speech Evaluation Standards for the CrowdabstractSubjective assessments are a key component of speech quality research. Traditionally, the assessments are conducted in laboratories in controlled conditions and following international standards like ITU-T Rec.P.800. However, even before the current pandemic, more speech quality research used crowdsourcing-based approaches for collecting subjective ratings. Crowdsourcing allows researchers to collect data even without a dedicated test laboratory, to collect data from a huge and diverse group of participants, and to perform the assessment in various real-life settings. Still, this approach raises questions about the reliability and validity of the subjective ratings, especially when comparing the ratings with data collected in standardized procedures. One step to approach these challenges was the development of the ITU-T Rec.P.808 standard. This standard helps practitioners implement best practices from speech quality studies and crowdsourcing studies in their crowdsourced speech quality assessments. However, even with the ITU-T Rec.P.808 in action, it is unclear how much background knowledge is necessary to successfully “implement” this standard. Therefore, this paper aims to assess the data quality differences between two P.808 implementations. One implementation is from a co-author of the P.808 standard, and the other is a researcher with only a little background in crowdsourcing and speech quality assessments. Both implementations are used in a large-scale crowdsourcing study with about two hundred users from Amazon Mechanical Turk. The collected ratings are compared to gold-standard data from a certified laboratory. Also, the two implementations are compared to analyze whether they lead to the same conclusions. The results show that both implementations correlate strongly with the laboratory and with each other. Thus, suggesting that the ITU-T Rec.P.808 is robust enough to be implemented by non-experts in speech evaluation or crowdsourcing. Edwin Gamboa, Babak Naderi, Matthias Hirth, Sebastian Möller 0001 |
QoMEX | 3 |
| 2021 | On the Impact of COVID-19 on Subjective Digital Media Quality AssessmentabstractThe COVID-19 pandemic has induced dramatic effects in all areas of society worldwide. It has led to drastic restrictions on academic exchanges and caused research projects to be put on hold. The need to maintain social distancing has severely affected the research communities that rely on in-person studies. In particular, Quality of Experience (QoE) research relies heavily on user studies and in-person subjective tests as means of obtaining ground truths for system designs. In this position paper, we focus on the impact of COVID-19 on conducting subjective tests for digital media quality assessment. An overview of the related discussions that take place in associated research communities is provided. A number of challenges are posed in terms of hygiene, ethics, standardization, and research quality issues. Opportunities for consolidating QoE research under COVID-19 and beyond are put forward, including alternative experimental designs and adaptation of research methodologies. Numerous enablers are suggested and discussed, revealing various avenues that can be followed to assist in conducting subjective tests under pandemic conditions. We believe that this position paper will be helpful for researchers and practitioners relying on in-person studies to keep abreast about the potential measures that support coming out of the COVID-19 pandemic strong and being prepared for future endeavors. Hans-Jürgen Zepernick, Kerstin Pieper, Robert P. Spang, Ulrich Engelke, Matthias Hirth, Babak Naderi |
MMSP | 5 |
| 2021 | Relationship Status: It's Complicated. Using APM as a QoE-qualifying Tempo MetricabstractThe results of studies investigating the perceived quality of different video games in various environments are often not easy to compare. While one input delay value might be perfectly fine for one game, it might completely break immersion and flow in another. Typically, one would assume that video games can be categorized based on their genre to reach a certain comparability, but that has not proven true in the past. Thus, efforts are made to find new ways of categorization for this issue that are based on objective and quantifiable metrics. In this work, we focus on the merits of Actions per Minute (APM) as one such candidate metric as there is already preliminary usage of the metric for this purpose in the literature. We explore a diverse dataset of real-world DOTA 2 matches and try to understand the characteristics and stability of APM and its relationship to other game-centric and player-centric objective metrics. This is an initial attempt to understand how stable such a metric would be even for just one game in the larger quest to find a suitable cross-game categorization metric. Our exploration has uncovered player-specific influences in the APM. But we are confident that with a narrow, filtered definition of the APM this approach can be used to define APM ranges that resemble the tempo of a game to be used as a basis to classify games for Quality of Experience research. Florian Metzger, Norman Stulier, Kathrin Borchert, Matthias Hirth |
QoMEX | 4 |
| 2020 | Personal Task Design Preferences of CrowdworkersabstractToday's software offers diverse functionalities that need to be made accessible to humans in an easy to understand and quick to learn way. This is not a new challenge and already well addressed in usability research. Yet, similar challenges arise in crowdsourcing tasks. While task interfaces are often less complicated than complete software products, crowdsourcing workers have only a minimal amount of time to familiarize them with the task interface. Additionally, due to the repetitiveness of the tasks, workers are required to work with this interface to solve a few hundred tasks. Unfortunately, often little attention is paid to optimize the task interfaces. Even if recent studies have shown evidence of the negative effects of poorly designed interfaces on the work performance, there is less knowledge about the importance of the usability from the worker's point of view and their personal preferences. This work aims to fill this gap in two steps. First, we analyze the relevance of good usability and interface design in relation to other task properties like payment or joyfulness by conducting a survey on two popular crowdsourcing platforms. Second, we identify interface properties that are of importance for the crowdsourcing workers and discuss both consenting and dissenting opinions of different worker groups. The results show that all workers place an essential role in the usability of a task interface when selecting tasks, but do not agree on coherent design preferences within and between the platforms. Matthias Hirth, Kathrin Borchert, Katrien De Moor, Vanessa Borst, Tobias Hoßfeld |
QoMEX | 1 |
| 2020 | Impact of the Number of Votes on the Reliability and Validity of Subjective Speech Quality Assessment in the Crowdsourcing ApproachabstractThe subjective quality of transmitted speech is traditionally assessed in a controlled laboratory environment according to ITU-T Rec. P.800. In turn, with crowdsourcing, crowdworkers participate in a subjective online experiment using their own listening device, and in their own working environment. Despite such less controllable conditions, the increased use of crowdsourcing micro-task platforms for quality assessment tasks has pushed a high demand for standardized methods, resulting in ITU-T Rec. P.808. This work investigates the impact of the number of judgments on the reliability and the validity of quality ratings collected through crowdsourcing-based speech quality assessments, as an input to ITU-T Rec. P.808 . Three crowdsourcing experiments on different platforms were conducted to evaluate the overall quality of three different speech datasets, using the Absolute Category Rating procedure. For each dataset, the Mean Opinion Scores (MOS) are calculated using differing numbers of crowdsourcing judgements. Then the results are compared to MOS values collected in a standard laboratory experiment, to assess the validity of crowdsourcing approach as a function of number of votes. In addition, the reliability of the average scores is analyzed by checking inter-rater reliability, gain in certainty, and the confidence of the MOS. The results provide a suggestion on the required number of votes per condition, and allow to model its impact on validity and reliability. Babak Naderi, Tobias Hoßfeld, Matthias Hirth, Florian Metzger, Sebastian Möller 0001, Rafael Zequeira Jiménez |
QoMEX | 3 |
| 2019 | In Vivo or in Vitro? Influence of the Study Design on Crowdsourced Video QoEabstractEvaluating the QoE of video streaming and its influence factors has become paramount for streaming providers, as they want to maintain high satisfaction for their customers. In this context, crowdsourced user studies became a valuable tool to evaluate different factors which can affect the perceived user experience on a large scale.In general, we observed that most of these crowdsourcing studies either use an in vivo or an in vitro design. In vivo design means that the study participant has to rate the QoE of a video that is embedded in an application similar to a real streaming service, e.g., YouTube or Netflix. In vitro design refers to a setting, in which the video stream is separated from a specific service and thus, the video plays on a plain background. Although these designs vary widely, the results are often compared and generalized.Therefore, in this work, we investigate the influence of these two study design alternatives on the perceived QoE. In crowdsourced user studies, participants rate the video streaming with respect to different stalling patterns (no stalling, different positions) and study designs (in vivo or in vitro). Contrary to our expectations, the results indicate that there is statistically no significant influence of the study design on the perceived video QoE and acceptance. In addition, we found that the in vivo design does not reduce the test takers' attentiveness. Kathrin Borchert, Anika Seufert, Matthias Hirth, Tobias Hoßfeld |
QoMEX | 3 |
| 2018 | Identification of Delay Thresholds Representing the Perceived Quality of Enterprise ApplicationsabstractModern enterprise applications are often designed as distributed architectures, e.g., thin client computing and thus degradations in network related Quality of Service (QoS) parameters may also negatively impact the user-perceived Quality of Experience (QoE) of the application. In this work, we create a model to predict the perceived application quality based on measurements of objective technical parameters. For this, we gathered a data set in a cooperating enterprise over a timespan of nearly three months. As the obtained data set is subject to bias that originates from seasonal effects as well as a limited and predefined set of technical parameters, we further evaluate how to identify segments of the data that lead to misclassification. Last, we quantify the trade-off between the gain in the QoE prediction accuracy and the amount of filtered data. Kathrin Borchert, Stanislav Lange, Thomas Zinner, Matthias Hirth |
QoMEX | 4 |
| 2017 | Collecting subjective ratings in enterprise environmentsabstractSimilar to many modern applications, enterprise applications like SAP are often implemented in a distributed fashion and consequently suffer from network degradations resulting in impairments like increased loading delays. While the influence of these impairments on the perceived quality of users is well researched for consumer applications and network services, their impact in a business environment is still unclear. To address this gap we develop a non-intrusive software tool for continuously collecting subjective ratings on the performance of an enterprise application from a large number of employees. Based on the feedback from two field studies in a company we briefly discuss challenges of QoE monitoring in the context of enterprises. As a first step towards building QoE models, we combine the subjective ratings with technical monitoring data and observer a negative correlation between the user satisfaction and the overall load of the server infrastructure. Kathrin Borchert, Matthias Hirth, Thomas Zinner, Anja S. Göritz |
QoMEX | 2 |
| 2016 | ERWIN - enabling the reproducible investigation of waiting times for arbitrary workflowsabstractDelay effects can impact the Quality of Experience of interactive systems, which motivates research assessing delay impairments, mostly for web based systems. Current studies follow individual methodologies and typically assesses individual and custom-made web pages, whose construction requires expert knowledge in web technologies. Further, a range of native, non-web applications cannot be easily modified for delay studies. Thus, a generalized methodology for assessing delay impacts for a broad range of applications that is accessible to researchers without (web) development expertise is still missing. This paper contributes to this open problem by i) presenting a new methodology for reproducible delay assessments in a broad class of systems and ii) presenting an open-source implementation to be used by the community. This methodology particularly aims at making delay assessment available to a broad range of researchers by avoiding programming skills and thus by lowering the barrier for setting-up delay assessments. Thomas Zinner, Matthias Hirth, Valentin Fischer, Oliver Hohlfeld |
QoMEX | 2 |
| 2015 | Text Categorization for Deriving the Application Quality in Enterprises Using Ticketing Systems
Thomas Zinner, Florian Lemmerich, Susanna Schwarzmann, Matthias Hirth, Peter Karg, Andreas Hotho |
DaWaK | 4 |
| 2015 | Crowdsourced network measurements: Benefits and best practices
Matthias Hirth, Tobias Hoßfeld, Marco Mellia, Christian Schwartz, Frank Lehrieder |
Comput. Networks | 1 |
| 2014 | Survey of web-based crowdsourcing frameworks for subjective quality assessmentabstractThe popularity of the crowdsourcing for performing various tasks online increased significantly in the past few years. The low cost and flexibility of crowdsourcing, in particular, attracted researchers in the field of subjective multimedia evaluations and Quality of Experience (QoE). Since online assessment of multimedia content is challenging, several dedicated frameworks were created to aid in the designing of the tests, including the support of the testing methodologies like ACR, DCR, and PC, setting up the tasks, training sessions, screening of the subjects, and storage of the resulted data. In this paper, we focus on the web-based frameworks for multimedia quality assessments that support commonly used crowdsourcing platforms such as Amazon Mechanical Turk and Microworkers. We provide a detailed overview of the crowdsourcing frameworks and evaluate them to aid researchers in the field of QoE assessment in the selection of frameworks and crowdsourcing platforms that are adequate for their experiments. Tobias Hoßfeld, Matthias Hirth, Pavel Korshunov, Philippe Hanhart, Bruno Gardlo, Christian Keimel, Christian Timmerer |
MMSP | 2 |
| 2014 | Best Practices for QoE Crowdtesting: QoE Assessment With CrowdsourcingabstractQuality of Experience (QoE) in multimedia applications is closely linked to the end users' perception and therefore its assessment requires subjective user studies in order to evaluate the degree of delight or annoyance as experienced by the users. QoE crowdtesting refers to QoE assessment using crowdsourcing, where anonymous test subjects conduct subjective tests remotely in their preferred environment. The advantages of QoE crowdtesting lie not only in the reduced time and costs for the tests, but also in a large and diverse panel of international, geographically distributed users in realistic user settings. However, conceptual and technical challenges emerge due to the remote test settings. Key issues arising from QoE crowdtesting include the reliability of user ratings, the influence of incentives, payment schemes and the unknown environmental context of the tests on the results. In order to counter these issues, strategies and methods need to be developed, included in the test design, and also implemented in the actual test campaign, while statistical methods are required to identify reliable user ratings and to ensure high data quality. This contribution therefore provides a collection of best practices addressing these issues based on our experience gained in a large set of conducted QoE crowdtesting studies. The focus of this article is in particular on the issue of reliability and we use video quality assessment as an example for the proposed best practices, showing that our recommended two-stage QoE crowdtesting design leads to more reliable results. Tobias Hoßfeld, Christian Keimel, Matthias Hirth, Bruno Gardlo, Julian Habigt, Klaus Diepold, Phuoc Tran-Gia |
IEEE Trans. Multim. | 3 |
| 2011 | Quantification of YouTube QoE via CrowdsourcingabstractThis paper addresses the challenge of assessing and modeling Quality of Experience (QoE) for online video services that are based on TCP-streaming. We present a dedicated QoE model for You Tube that takes into account the key influence factors (such as stalling events caused by network bottlenecks) that shape quality perception of this service. As second contribution, we propose a generic subjective QoE assessment methodology for multimedia applications (like online video) that is based on crowd sourcing - a highly cost-efficient, fast and flexible way of conducting user experiments. We demonstrate how our approach successfully leverages the inherent strengths of crowd sourcing while addressing critical aspects such as the reliability of the experimental data obtained. Our results suggest that, crowd sourcing is a highly effective QoE assessment method not only for online video, but also for a wide range of other current and future Internet applications. Tobias Hoßfeld, Michael Seufert, Matthias Hirth, Thomas Zinner, Phuoc Tran-Gia, Raimund Schatz |
ISM | 3 |