Maximilian Kapsecker

dblp:327/7768 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2026
0000-0002-3907-0749ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Face, body and person analysis · 67% 3D vision · 33%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Face, body and person analysis › human pose estimation
3d pose estimation
1.012026
A Comparative Assessment of Accuracy in Video-Based Monocular Human Pose Estimation Frameworks · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Computer vision › Face, body and person analysis
human pose estimation
1.012026
A Comparative Assessment of Accuracy in Video-Based Monocular Human Pose Estimation Frameworks · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Computer vision › 3D vision › pose estimation
monocular pose estimation
1.012026
A Comparative Assessment of Accuracy in Video-Based Monocular Human Pose Estimation Frameworks · IEEE Trans. Pattern Anal. Mach. Intell. 2026

Methods — techniques the papers use, named apart from their topics

weighted mean absolute error · 1.0intra-class correlation coefficient · 1.0
YearPublicationVenuePosition
2026 A Comparative Assessment of Accuracy in Video-Based Monocular Human Pose Estimation Frameworks
abstract
In human pose estimation, a comprehensive evaluation of state-of-the-art frameworks is necessary to advance both research and practical applications. This paper presents a thorough review of state-of-the-art 2D and 3D human pose estimation frameworks, analyzing 118 papers and four GitHub repositories, with a focus on frameworks made since 2019. The following frameworks are chosen based on predefined inclusion criteria: AlphaPose, Detectron2, MediaPipe, MeTRAbs, MHFormer, MMPose, MoveNet, OpenPifPaf, OpenPifPaf-vita, OpenPose, PoseFormerV2, rtmlib, StridedTransformer-Pose3D, ultralytics (YOLOv8), ViTPose, and YOLOv7. This paper evaluates these 16 frameworks on an existing, unpublished dataset consisting of exercise videos recorded with a monocular RGB camera and synchronized gold-standard motion capture data. The dataset includes videos of nine individuals performing eight exercises, recorded from two camera views with different planar angles. The analysis evaluates joint angle performance of the frameworks using weighted mean absolute error and weighted intraclass correlation coefficient as quantitative metrics. MeTRAbs emerged as the best overall framework, while AlphaPose, rtmlib, and YOLOv7 were the top 2D performers.
Fabian Kahl, Philipp Wegner, Maximilian Kapsecker, Leon Nissen, Jennifer Faber, Stephan M. Jonas, Lara Marie Reimer
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Facial Landmark Analysis for Detecting Visual Impairment in Mobile LogMAR Test
abstract
Visual impairment is a widespread global health issue that affects millions of people across all ages and backgrounds. Timely intervention is essential for the effective management of eye diseases. Smartphones offer the possibility of continuously recording facial gestures during interaction with the device, whereby changes such as squinting of the eyes could indicate progressive vision loss. In this context, a mobile health application was developed to conduct a digital logMAR test while simultaneously capturing real-time facial features. A total of 37 participants took part in a controlled mobile eye test study. The facial landmarks recorded during the test were analyzed to identify patterns that can distinguish between sequences of letters that were read correctly, partially, or not at all. Specifically, explorative data analysis and receiver operating characteristic curves were employed to determine facial landmarks with high discriminative power in relation to reading ability. The predominant facial regions that showed the most significant change under reduced performance during the vision test were the nose, mouth, and cheeks. Notably, the characteristic maximum squinting of the cheeks stood out with an area under the curve of 0.82. The analysis showed the potential of tracking specific facial features for continuous and unobtrusive vision assessment. It motivates to integrate facial feature analysis into an everyday application such as a web browser and to conduct a study in a non-standardized environment on a larger scale.
Maximilian Kapsecker, Elena Mille, Florian Schweizer, Jens Klinker, Joe Yu, Alexander Leube, Stephan M. Jonas
IEEE J. Biomed. Health Informatics1