VLDB 2026 Research / reviewers in the wild / expert
Hannes Fassold
dblp:63/8105
· DBLP profile ↗
17ranked-venue papers
5as first author
8since 2021 · last 2024
0000-0001-7113-0038ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Data-Efficient Domain Transfer for Instance Segmentation for AR ScenesabstractAugmented Reality (AR) applications rely heavily on the accurate detection and segmentation of objects in order to seamlessly integrate virtual content into real world environments. However, achieving robust object detection and segmentation in AR scenes remains challenging, in particular when specific classes need to be covered. Due to the lack of sufficient real annotated data, training can rely on synthetic data, which has the issue of the domain gap between synthetic and real-world images. This paper addresses the challenge of generalisation in object detection/instance segmentation models trained with synthetic data. It presents techniques to improve the generalisation capability of object detection models trained on synthetic data for real-world applications. Our method is based on the model soup approach, where models trained with different data subsets and hyperparameters are combined to achieve better generalisation performance. The experimental results on the validation dataset of the ADAPT sim2real Object Detection Challenge 2023 demonstrate the effectiveness of our approach in achieving good generalisation performance. This marks a significant step towards more reliable and adaptable object detection and segmentation approaches for AR applications. Stefanie Onsori-Wechtitsch, Hermann Fürntratt, Hannes Fassold, Werner Bailer |
CBMI | 3 |
| 2024 | LiveSkeleton: High-Quality Real-Time Human Tracking and Pose EstimationabstractWe introduce LiveSkeleton, a high-performance human pose estimation algorithm designed for real-time multimedia applications. Our approach integrates robust person detection using Scaled-YoloV4, efficient optical-flow based tracking and the high-quality RTMPose algorithm for skeleton extraction. Operating at 25 frames per second for up to five individuals simultaneously, the system demonstrates robust performance across content from diverse scenarios, including VR/XR environments, surveillance contexts, and healthcare settings. Experimental results confirm that the algorithms performs well in challenging conditions, such as crowded scenes with substantial occlusion, thereby offering significant potential for enhancing human-computer interaction and automated behavior analysis. Hannes Fassold |
ISM | 1 |
| 2024 | Few-Shot Object Detection as a Service: Facilitating Training and Deployment for Domain Experts
Werner Bailer, Mihai Dogariu, Bogdan Ionescu, Hannes Fassold |
MMM (4) | 4 |
| 2024 | Scalable bio-inspired training of Deep Neural Networks with FastHebbabstractRecent work on sample efficient training of Deep Neural Networks (DNNs) proposed a semi-supervised methodology based on biologically inspired Hebbian learning, combined with traditional backprop-based training. Promising results were achieved on various computer vision benchmarks, in scenarios of scarce labeled data availability. However, current Hebbian learning solutions can hardly address large-scale scenarios due to their demanding computational cost. In order to tackle this limitation, in this contribution, we investigate a novel solution, named FastHebb (FH), based on the reformulation of Hebbian learning rules in terms of matrix multiplications, which can be executed more efficiently on GPU. Starting from Soft-Winner-Takes-All (SWTA) and Hebbian Principal Component Analysis (HPCA) learning rules, we formulate their improved FH versions: SWTA-FH and HPCA-FH. We experimentally show that the proposed approach accelerates training speed up to 70 times, allowing us to gracefully scale Hebbian learning experiments on large datasets and network architectures such as ImageNet and VGG. Gabriele Lagani, Fabrizio Falchi, Claudio Gennaro, Hannes Fassold, Giuseppe Amato 0001 |
Neurocomputing | 4 |
| 2023 | People@Places and ToDY: Two Datasets for Scene Classification in Media Production and Archiving
Werner Bailer, Hannes Fassold |
MMM (1) | 2 |
| 2022 | Few-shot Object Detection as a Semi-supervised Learning ProblemabstractThis paper addresses the issue of dealing with few-shot learning settings in which different classes are annotated on different datasets. Each part of the data has exhaustive annotations for only one or a small set of classes, but not for others used in training. It is likely, that unannotated samples of a class exist, potentially impacting the gradient as negative samples. Because of this fact, we argue that few-shot learning is essentially a semi-supervised learning problem. We analyze how approaches from semi-supervised learning can be applied. In particular, the use of soft-sampling to weight the gradient based on overlap of detections and ground truth, and creating missing annotations using a preliminary detector are studied. The use of soft-sampling provides small but consistent improvements, at much lower computational effort than predicting additional annotations. Werner Bailer, Hannes Fassold |
CBMI | 2 |
| 2022 | FastHebb: Scaling Hebbian Training of Deep Neural Networks to ImageNet Level
Gabriele Lagani, Claudio Gennaro, Hannes Fassold, Giuseppe Amato 0001 |
SISAP | 3 |
| 2021 | Detecting speaking persons in videoabstractWe present a novel method for detecting speaking persons in video, by extracting facial landmarks with a neural network and analysing these landmarks statistically over time. Hannes Fassold |
MMSP | 1 |
| 2019 | Resource-Efficient Object Detection by Sharing Backbone CNNsabstractThe detection of objects in image and video has made huge progress in recent years due to the use of deep convolutional neural networks (DNNs), with some network architectures becoming de-facto standards. This paper addresses the problem of sharing a backbone CNN for different tasks, for example, to enable detection of additional classes when an already trained network is available. When using multiple such neural networks, sharing a backbone can save inference time and memory consumption. We study sharing a common backbone between neural networks trained for different tasks (logoness and text block detection) based on Yolo v3. We provide results on the impact of different lengths of the shared backbone on performance and resource efficiency. Werner Bailer, Hannes Fassold |
ISM | 2 |
| 2019 | OmniTrack: Real-Time Detection and Tracking of Objects, Text and Logos in VideoabstractThe automatic detection and tracking of general objects (like persons, animals or cars), text and logos in a video is crucial for many video understanding tasks, and usually real-time processing as required. We propose OmniTrack, an efficient and robust algorithm which is able to automatically detect and track objects, text as well as brand logos in real-time. It combines a powerful deep learning based object detector (YoloV3) with high-quality optical flow methods. Based on the reference YoloV3 C++ implementation, we did some important performance optimizations which will be described. Hannes Fassold, Ridouane Ghermi |
ISM | 1 |
| 2019 | Adapting Computer Vision Algorithms for Omnidirectional VideoabstractOmnidirectional (360°) video has got quite popular because it provides a highly immersive viewing experience. For computer vision algorithms, it poses several challenges, like the special (equirectangular) projection commonly employed and the huge image size. In this work, we give a high-level overview of these challenges and outline strategies how to adapt computer vision algorithm for the specifics of omnidirectional video. Hannes Fassold |
ACM Multimedia | 1 |
| 2016 | Quality Analysis on Mobile Devices for Real-Time Feedback
Stefanie Onsori-Wechtitsch, Hannes Fassold, Marcus Thaler, Krzysztof Kozlowski, Werner Bailer |
MMM (1) | 2 |
| 2013 | Improving Preservation and Access Processes of Audiovisual Media by Content-Based Quality Assessment
Peter Schallauer, Hannes Fassold, Albert Hofmann, Werner Bailer, Stefanie Onsori-Wechtitsch |
MMM (2) | 2 |
| 2013 | A Perceptual Image Sharpness Metric Based on Local Edge Gradient AnalysisabstractIn this letter, a no-reference perceptual sharpness metric based on a statistical analysis of local edge gradients is presented. The method takes properties of the human visual system into account. Based on perceptual properties, a relationship between the extracted statistical features and the metric score is established to form a Perceptual Sharpness Index (PSI). A comparison with state-of-the-art metrics shows that the proposed method correlates highly with human perception and exhibits low computational complexity. In contrast to existing metrics, the PSI performs well for a wide range of blurriness and shows a high degree of invariance for different image contents. Christoph Feichtenhofer, Hannes Fassold, Peter Schallauer |
IEEE Signal Process. Lett. | 2 |
| 2012 | Robust detection of single-frame defects in archived film
Stefanie Onsori-Wechtitsch, Hannes Fassold, Peter Schallauer |
ICPR | 2 |
| 2012 | Automated Visual Quality Analysis for Media ProductionabstractAutomatic quality control for audiovisual media is an important tool in the media production process. In this paper we present tools for assessing the quality of audiovisual content in order to decide about the reusability of archive content. We first discuss automatic detectors for the common impairments noise and grain, video breakups, sharpness, image dynamics and blocking. For the efficient viewing and verification of the automatic results by an operator, three approaches for user interfaces are presented. Finally, we discuss the integration of the tools into a service oriented architecture, focusing on the recent standardization efforts by EBU and AMWA's Joint Task Force on a Framework for Interoperability of Media Services in TV Production (FIMS). Hannes Fassold, Stefanie Onsori-Wechtitsch, Albert Hofmann, Werner Bailer, Peter Schallauer, Roberto Borgotallo, Alberto Messina, Mohan Liu, Patrick Ndjiki-Nya, Peter Altendorf |
ISM | 1 |
| 2009 | Automatic freeze frame detection for video preservationabstractA significant amount of work in film and video preservation is dedicated to quality assessment of the content to be archived or reused out of the archive. This paper proposes automatic content analysis algorithms which reduce manual inspection time in software based preservation environments. We list the requirements for such algorithms and tools and show exemplarily on a freeze frame impairment detector how analysis algorithms need to be designed for meeting the requirements. The evaluation has shown that robustness against other impairments (noise and flickering) is an essential part of the algorithm. Successful detection of freeze frame impairments with a minimum length of three frames is achieved. The consideration of human perception is important to achieve low false detection rates. Analysis results are represented in a MPEG-7 standard compliant way. The proposed defect summary visualization tool enables efficient human exploration of visually impaired content. Peter Schallauer, Hannes Fassold, Werner Bailer |
ICIP | 2 |