EDBT 2026 Demo / reviewers in the wild / expert
Rainer Lienhart
dblp:l/RainerLienhart
· DBLP profile ↗
88ranked-venue papers
19as first author
15since 2021 · last 2026
0000-0003-4007-6889ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 82 · 19 first-author · 13 since 2021Artificial intelligence and machine learning · 14 · 6 since 2021Databases, data management, data science and information retrieval · 8Systems, architecture and hardware · 1Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Uplifting Table Tennis: A Robust, Real-World Application for 3D Trajectory and Spin EstimationabstractObtaining the precise 3D motion of a table tennis ball from standard monocular videos is a challenging problem, as existing methods trained on synthetic data struggle to generalize to the noisy, imperfect ball and table detections of the real world. This is primarily due to the inherent lack of 3D ground truth trajectories and spin annotations for real-world video. To overcome this, we propose a novel two-stage pipeline that divides the problem into a front-end perception task and a back-end 2D-to-3D uplifting task. This separation allows us to train the front-end components with abundant 2D supervision from our newly created TTHQ dataset, while the back-end uplifting network is trained exclusively on physically-correct synthetic data. We specifically re-engineer the uplifting model to be robust to common real-world artifacts, such as missing detections and varying frame rates. By integrating a ball detector and a table keypoint detector, our approach transforms a proof-of-concept uplifting method into a practical, robust, and high-performing end-to-end application for 3D table tennis trajectory and spin analysis. Daniel Kienzle, Katja Ludwig, Julian Lorenz, Shin'ichi Satoh 0001, Rainer Lienhart |
WACV | 5 |
| 2025 | Harnessing Event Sensory Data for Error Pattern Prediction in Vehicles: A Language Model ApproachabstractIn this paper, we draw an analogy between processing natural languages and processing multivariate event streams from vehicles in order to predict when and what error pattern is most likely to occur in the future for a given car. Our approach leverages the temporal dynamics and contextual relationships of our event data from a fleet of cars. Event data is composed of discrete values of error codes as well as continuous values such as time and mileage. Modelled by two causal Transformers, we can anticipate vehicle failures and malfunctions before they happen. Thus, we introduce CarFormer, a Transformer model trained via a new self-supervised learning strategy, and EPredictor, an autoregressive Transformer decoder model capable of predicting when and what error pattern will most likely occur after some error code apparition. Despite the challenges of high cardinality of event types, their unbalanced frequency of appearance and limited labelled data, our experimental results demonstrate the excellent predictive ability of our novel model. Specifically, with sequences of 160 error codes on average, our model is able with only half of the error codes to achieve 80% F1 score for predicting what error pattern will occur and achieves an average absolute error of 58.4 ± 13.2h when forecasting the time of occurrence, thus enabling confident predictive maintenance and enhancing vehicle safety. Hugo Math, Rainer Lienhart, Robin Schön |
AAAI | 2 |
| 2025 | MMMS: Multi-Modal Multi-Surface Interactive SegmentationabstractIn this paper, we present a method to interactively create segmentation masks on the basis of user clicks. We pay particular attention to the segmentation of multiple surfaces that are simultaneously present in the same image. Since these surfaces may be heavily entangled and adjacent, we also present a novel extended evaluation metric that accounts for the challenges of this scenario. Additionally, the presented method is able to use multi-modal inputs to facilitate the segmentation task. At the center of this method is a network architecture which takes as input an RGB image, a number of non-RGB modalities, an erroneous mask, and encoded clicks. Based on this input, the network predicts an improved segmentation mask. We design our architecture such that it adheres to two conditions: (1) The RGB backbone is only available as a black-box. (2) To reduce the response time, we want our model to integrate the interaction-specific information after the image feature extraction and the multi-modal fusion. We refer to the overall task as multi-modal multi-surface interactive segmentation (MMMS). We are able to show the effectiveness of our multi-modal fusion strategy. Using additional modalities, our system reduces the NoC@90 by up to 1.28 clicks per surface on average on DeLiVER and up to 1.19 on MFNet. On top of this, we are able to show that our RGB-only baseline achieves competitive, and in some cases even superior performance when tested in a classical, single-mask interactive segmentation scenario. Robin Schön, Julian Lorenz, Katja Ludwig, Daniel Kienle, Rainer Lienhart |
CBMI | 5 |
| 2025 | 8th ACM International Workshop on Multimedia Content Analysis in Sports (ACM MMSports'25)abstractThe 8th ACM International Workshop on Multimedia Content Analysis in Sports is held in Dublin, Ireland on October 28th, 2025. It is co-located with ACM Multimedia 2025. The goal of this workshop is to bring together researchers and practitioners from academia and industry to address challenges and report progress in mining, analyzing, understanding, and visualizing the multimodal data in sports. The combination of sports and modern technology offers a novel and intriguing field of research with promising approaches for visual broadcast augmentation as well as understanding, statistical analysis, and evaluation in amateur and professional sports. There is a lack of research communities focusing on the fusion of multiple modalities. Thus, this workshop series on multimedia content analysis in sports aims to contribute to the closure of this research gap by bringing together the breadth and depth of these diverse approaches to stimulate each other with new ideas and foster research progress. Rainer Lienhart, Thomas B. Moeslund, Hideo Saito 0001 |
ACM Multimedia | 1 |
| 2024 | Towards Learning Monocular 3D Object Localization from 2D Labels Using the Physical Laws of MotionabstractWe present a novel method for precise 3D object localization in single images from a single calibrated camera using only 2D labels. No expensive 3D labels are needed. Thus, instead of using 3D labels, our model is trained with easy-to-annotate 2D labels along with the physical knowledge of the object’s motion. Given this information, the model can infer the latent third dimension, even though it has never seen this information during training. Our method is evaluated on both synthetic and real-world datasets, and we are able to achieve a mean distance error of just 6 cm in our experiments on real data. The results indicate the method’s potential as a step towards learning 3D object location estimation, where collecting 3D data for training is not feasible. Daniel Kienzle, Katja Ludwig, Julian Lorenz, Rainer Lienhart |
3DV | 4 |
| 2024 | Wseseg: Introducing a Dataset for the Segmentation of Winter Sports Equipment With a Baseline for Interactive SegmentationabstractIn this paper we introduce a new dataset containing instance segmentation masks for ten different categories of winter sports equipment, called WSESeg (Winter Sports Equipment Segmentation)11Available at https://github.com/Schorob/wseseg.. Furthermore, we carry out interactive segmentation experiments on said dataset to explore possibilities for efficient further labeling. The SAM and HQ-SAM models are conceptualized as foundation models for performing user guided segmentation. In order to measure their claimed generalization capability we evaluate them on WSESeg. Since interactive segmentation offers the benefit of creating easily exploitable ground truth data during test-time, we are going to test various online adaptation methods for the purpose of exploring potentials for improvements without having to fine-tune the models explicitly. Our experiments show that our adaptation methods drastically reduce the Failure Rate (FR) and Number of Clicks (NoC) metrics, which generally leads faster to better interactive segmentation results. Robin Schön, Daniel Kienzle, Rainer Lienhart |
CBMI | 3 |
| 2024 | A Fair Ranking and New Model for Panoptic Scene Graph Generation
Julian Lorenz, Alexander Pest, Daniel Kienzle, Katja Ludwig, Rainer Lienhart |
ECCV (61) | 5 |
| 2024 | The STOIC2021 COVID-19 AI challenge: Applying reusable training methodologies to private dataabstractChallenges drive the state-of-the-art of automated medical image analysis. The quantity of public training data that they provide can limit the performance of their solutions. Public access to the training methodology for these solutions remains absent. This study implements the Type Three (T3) challenge format, which allows for training solutions on private data and guarantees reusable training methodologies. With T3, challenge organizers train a codebase provided by the participants on sequestered training data. T3 was implemented in the STOIC2021 challenge, with the goal of predicting from a computed tomography (CT) scan whether subjects had a severe COVID-19 infection, defined as intubation or death within one month. STOIC2021 consisted of a Qualification phase, where participants developed challenge solutions using 2000 publicly available CT scans, and a Final phase, where participants submitted their training methodologies with which solutions were trained on CT scans of 9724 subjects. The organizers successfully trained six of the eight Final phase submissions. The submitted codebases for training and running inference were released publicly. The winning solution obtained an area under the receiver operating characteristic curve for discerning between severe and non-severe COVID-19 of 0.815. The Final phase solutions of all finalists improved upon their Qualification phase solutions. Luuk H. Boulogne, Julian Lorenz, Daniel Kienzle, Robin Schön, Katja Ludwig, Rainer Lienhart, Simon Jégou, Derik Shi, Mayug Maniparambil, Dominik Müller, Silvan Mertes, Niklas Schröter, Fabio Hellmann, Miriam Elia, Ine Dirks, Matías N. Bossa, Abel Díaz Berenguer, Tanmoy Mukherjee, Jef Vandemeulebroucke, Hichem Sahli, Nikos Deligiannis, Panagiotis Gonidakis, Ngoc Dung Huynh, Muhammad Imran Razzak, Mohamed Reda Bouadjenek, Mario Verdicchio, Pasquale Borrelli, Marco Aiello 0003, James A. Meakin, Alexander Lemm, Christoph Russ, Razvan Ionasec, Nikos Paragios, Bram van Ginneken, Marie-Pierre Revel |
Medical Image Anal. | 6 |
| 2023 | MMSports '23: 6th International Workshop on Multimedia Content Analysis in SportsabstractThe sixth ACM International Workshop on Multimedia Content Analysis in Sports (ACM MMSports'23) is part of the ACM International Conference on Multimedia 2023 (ACM Multimedia 2023). The goal of this workshop is to bring together researchers and practitioners from academia and industry to address challenges and report progress in mining, analyzing, understanding, and visualizing multimedia/multimodal data in sports, sports broadcasts, sports games and sports medicine. The combination of sports and modern technology offers a novel and intriguing field of research with promising approaches for visual broadcast augmentation and understanding, for statistical analysis and evaluation, and for sensor fusion during workouts as well as competitions. There is a lack of research communities focusing on the fusion of multiple modalities. We are helping to close this research gap with this workshop series on multimedia content analysis in sports. Hideo Saito 0001, Thomas B. Moeslund, Rainer Lienhart |
ACM Multimedia | 3 |
| 2023 | Uplift and Upsample: Efficient 3D Human Pose Estimation with Uplifting TransformersabstractThe state-of-the-art for monocular 3D human pose estimation in videos is dominated by the paradigm of 2D-to-3D pose uplifting. While the uplifting methods themselves are rather efficient, the true computational complexity depends on the per-frame 2D pose estimation. In this paper, we present a Transformer-based pose uplifting scheme that can operate on temporally sparse 2D pose sequences but still produce temporally dense 3D pose estimates. We show how masked token modeling can be utilized for temporal upsampling within Transformer blocks. This allows to decouple the sampling rate of input 2D poses and the target frame rate of the video and drastically decreases the total computational complexity. Additionally, we explore the option of pre-training on large motion capture archives, which has been largely neglected so far We evaluate our method on two popular benchmark datasets: Human3.6M and MPI-INF-3DHP. With an MPJPE of 45.0 mm and 46.9 mm, respectively, our proposed method can compete with the state-of-the-art while reducing inference time by a factor of 12. This enables real-time throughput with variable consumer hardware in stationary and mobile applications. We release our code and models at https://github.com/goldbricklemon/uplift-upsample-3dhpe Moritz Einfalt, Katja Ludwig, Rainer Lienhart |
WACV | 3 |
| 2022 | Pseudo-Label Noise Suppression Techniques for Semi-Supervised Semantic Segmentation
Sebastian A. Scherer, Robin Schön, Rainer Lienhart |
BMVC | 3 |
| 2022 | Semantically Consistent Image-to-Image Translation for Unsupervised Domain AdaptationabstractUnsupervised Domain Adaptation (UDA) aims to adapt models trained on a source domain to a new target domain where no labelled data is available. In this work, we investigate the problem of UDA from a synthetic computer-generated domain to a similar but real-world domain for learning semantic segmentation. We propose a semantically consistent image-to-image translation method in combination with a consistency regularisation method for UDA. We overcome previous limitations on transferring synthetic images to real looking images. We leverage pseudo-labels in order to learn a generative image-to-image translation model that receives additional feedback from semantic labels on both domains. Our method outperforms state-of-the-art methods that combine image-to-image translation and semi-supervised learning on relevant domain adaptation benchmarks, i.e., on GTA5 to Cityscapes and SYNTHIA to Cityscapes. Stephan Brehm, Sebastian A. Scherer, Rainer Lienhart |
ICAART (2) | 3 |
| 2022 | Synchronized Audio-Visual Frames with Fractional Positional Encoding for Transformers in Video-to-Text TranslationabstractVideo-to-text (VTT) is the task of automatically generating descriptions for short audio-visual video clips. It can help visually impaired people to understand scenes shown in a YouTube video, for example. Transformer architectures have shown great performance in both machine translation and image captioning. In this work, we transfer promising approaches from image captioning and video processing to VTT and develop a straightforward Transformer architecture. Then, we expand this Transformer by a novel way of synchronizing audio and video features in Transformers which we call Fractional Positional Encoding (FPE). We run multiple experiments on the VATEX dataset and improve the CIDEr and BLEU-4 scores by 21.72 and 8.38 points compared to a vanilla Transformer network and achieve state-of-the art results on the MSR-VTT and MSVD datasets. Also, our novel FPE helps increase the CIDEr score by relative 8.6 %. Philipp Harzig, Moritz Einfalt, Rainer Lienhart |
ICIP | 3 |
| 2022 | MMSports'22: 5th International ACM Workshop on Multimedia Content Analysis in SportsabstractThe fifth ACM International Workshop on Multimedia Content Analysis in Sports (ACM MMSports'22) is part of the ACM International Conference on Multimedia 2022 (ACM Multimedia 2022). After two years of pure virtual MMSports workshops due to COVID-19, MMSports'22 is held on-site again. The goal of this workshop is to bring together researchers and practitioners from academia and industry to address challenges and report progress in mining, analyzing, understanding, and visualizing multimedia/multimodal data in sports, sports broadcasts, sports games and sports medicine. The combination of sports and modern technology offers a novel and intriguing field of research with promising approaches for visual broadcast augmentation and understanding, for statistical analysis and evaluation, and for sensor fusion during workouts as well as competitions. There is a lack of research communities focusing on the fusion of multiple modalities. We are helping to close this research gap with this workshop series on multimedia content analysis in sports. Related Workshop Proceedings are available in the ACM DL at: https://dl.acm.org/doi/proceedings/10.1145/3552437. Hideo Saito 0001, Thomas B. Moeslund, Rainer Lienhart |
ACM Multimedia | 3 |
| 2021 | MMSports'21: 4th International Workshop on Multimedia Content Analysis in Sports
Rainer Lienhart, Thomas B. Moeslund, Hideo Saito 0001 |
ACM Multimedia | 1 |
| 2020 | Error Bounds of Projection Models in Weakly Supervised 3D Human Pose EstimationabstractThe current state-of-the-art in monocular 3D human pose estimation is heavily influenced by weakly supervised methods. These allow 2D labels to be used to learn effective 3D human pose recovery either directly from images or via 2D-to-3D pose uplifting. In this paper we present a detailed analysis of the most commonly used simplified projection models, which relate the estimated 3D pose representation to 2D labels: normalized perspective and weak perspective projections. Specifically, we derive theoretical lower bound errors for those projection models under the commonly used mean per-joint position error (MPJPE). Additionally, we show how the normalized perspective projection can be replaced to avoid this guaranteed minimal error. We evaluate the derived lower bounds on the most commonly used 3D human pose estimation benchmark datasets. Our results show that both projection models lead to an inherent minimal error between 19.3mm and 54.7mm, even after alignment in position and scale. This is a considerable share when comparing with recent state-of-the-art results. Our paper thus establishes a theoretical baseline that shows the importance of suitable projection models in weakly supervised 3D human pose estimation. Nikolas Klug, Moritz Einfalt, Stephan Brehm, Rainer Lienhart |
3DV | 4 |
| 2020 | MMSports'20: 3rd International Workshop on Multimedia Content Analysis in SportsabstractThe third ACM International Workshop on Multimedia Content Analysis in Sports (ACM MMSports'20) is part of the ACM International Conference on Multimedia 2020 (ACM Multimedia 2020). Exceptionally, due to the corona pandemic, the workshop is held virtually. The goal of this workshop is to bring together researchers and practitioners from academia and industry to address challenges and report progress in mining, analyzing, understanding and visualizing the multimedia/multimodal data in sports. The combination of sports and modern technology offers a novel and intriguing field of research with promising approaches for visual broadcast augmentation, understanding, statistical analysis and evaluation, and sensor fusion. There is a lack of research communities focusing on the fusion of multiple modalities. We are helping to close this research gap with this workshop series on multimedia content analysis in sports. Rainer Lienhart, Thomas B. Moeslund, Hideo Saito 0001 |
ACM Multimedia | 1 |
| 2020 | Guest Editorial Multimedia Computing With Interpretable Machine LearningabstractThe papers in this special section is to broadly engage the machine learning and multimedia communities on the emerging yet challenging interpretable machine learning. Multimedia is increasingly becoming the “biggest big data,” among the most important and valuable source for insight and information. Many powerful machine learning algorithms, especially deep learning models such as convolutional neural networks (CNNs), have recently achieved outstanding predictive performance in a wide range of multimedia applications, including visual object classification, scene understanding, speech recognition, and activity prediction. Nevertheless, most deep learning algorithms are generally conceived as blackbox methods, and it is difficult to intuitively and quantitatively understand the results of their prediction and inference. Since this lack of interpretability is a major bottleneck in designing more successful predictive models and exploring wider-range useful applications, there has been an explosion of interest in interpreting the representations learned by these models, with profound implications for research into interpretable machine learning in the multimedia community. Yonghong Tian 0001, Cees Snoek, Jingdong Wang 0001, Zhu Liu 0001, Rainer Lienhart, Susanne Boll |
IEEE Trans. Multim. | 5 |
| 2019 | Addressing Data Bias Problems for Chest X-ray Image Report Generation
Philipp Harzig, Yan-Ying Chen, Francine Chen 0001, Rainer Lienhart |
BMVC | 4 |
| 2019 | Automatic Disease Detection and Report Generation for Gastrointestinal Tract ExaminationabstractIn this paper, we present a method to automatically identify diseases from videos of gastrointestinal (GI) tract examinations using a Deep Convolutional Neural Network (DCNN) that processes images from digital endoscopes. Our goal is to aid domain experts by automatically detecting abnormalities and generating a report that summarizes the main findings. We have implemented a model that uses two different DCNN architectures to generate our predictions, which are also capable of running on a mobile device. Using this architecture, we are able to predict findings on individual images. Combined with class activations maps (CAM), we can also automatically generate a textual report describing a video in detail while giving hints about the spatial location of findings and anatomical landmarks. Our work shows one way to use a multi-disease detection pipeline to also generate video reports that summarize key findings. Philipp Harzig, Moritz Einfalt, Rainer Lienhart |
ACM Multimedia | 3 |
| 2019 | MMSports'19: 2nd ACM International Workshop on Multimedia Content Analysis in SportsabstractThe second ACM International Workshop on Multimedia Content Analysis in Sports (ACM MMSports'19) is held in Nice, France on October 25th, 2019 co-located with the ACM International Conference on Multimedia 2019 (ACM Multimedia 2019). The goal of this workshop is to bring together researchers and practitioners from academia and industry to address challenges and report progress in mining, analyzing, understanding and visualizing the multimedia/multimodal data in sports. The combination of sports and modern technology offers a novel and intriguing field of research with promising approaches for visual broadcast augmentation, understanding, statistical analysis and evaluation, and sensor fusion. There is a lack of research communities focusing on the fusion of multiple modalities. We are helping to close this research gap with this workshop series on multimedia content analysis in sports. Rainer Lienhart, Thomas B. Moeslund, Hideo Saito 0001 |
ACM Multimedia | 1 |
| 2018 | Visual Question Answering With a Hybrid Convolution Recurrent ModelabstractVisual Question Answering (VQA) is a relatively new task, which tries to infer answer sentences for an input image coupled with a corresponding question. Instead of dynamically generating answers, they are usually inferred by finding the most probable answer from a fixed set of possible answers. Previous work did not address the problem of finding all possible answers, but only modeled the answering part of VQA as a classification task. To tackle this problem, we infer answer sentences by using a Long Short-Term Memory (LSTM) network that allows us to dynamically generate answers for (image, question) pairs. In a series of experiments, we discover an end-to-end Deep Neural Network structure, which allows us to dynamically answer questions referring to a given input image by using an LSTM decoder network. With this approach, we are able to generate both less common answers, which are not considered by classification models, and more complex answers with the appearance of datasets containing answers that consist of more than three words. Philipp Harzig, Christian Eggert, Rainer Lienhart |
ICMR | 3 |
| 2018 | Session details: Best Paper Session
Rainer Lienhart, Tao Mei 0001 |
ACM Multimedia | 1 |
| 2018 | 1st ACM International Workshop on Multimedia Content Analysis in SportsabstractThe first ACM International Workshop on Multimedia Content Analysis in Sports (ACM MMSports'18) is held in Seoul, South Korea on October 26th, 2018 and is co-located with the ACM International Conference on Multimedia 2018 (ACM Multimedia 2018). The goal of this workshop is to bring together researchers and practitioners from academia and industry to address challenges and report progress in mining and content analysis of multimedia/multimodal data in sports. The combination of sports and modern technology offers a novel and intriguing field of research with promising approaches for visual broadcast augmentation, understanding, statistical analysis and evaluation, and sensor fusion. There is a lack of research communities focusing on the fusion of multiple modalities. We are helping to close this research gap with this first workshop of a serious workshops on multimedia content analysis in sports. Rainer Lienhart, Thomas B. Moeslund, Hideo Saito 0001 |
ACM Multimedia | 1 |
| 2018 | Deep Learning for Multimedia: Science or Technology?abstractDeep learning has been successfully explored in addressing different multimedia topics recent years, ranging from object detection, semantic classification, entity annotation, to multimedia captioning, multimedia question answering and storytelling. Open source libraries and platforms such as Tensorflow, Caffe, MXnet significantly help promote the wide deployment of deep learning in solving real-world applications. On one hand, deep learning practitioners, while not necessary to understand the involved math behind, are able to set up and make use of a complex deep network. One recent deep learning tool based on Keras even provides the graphical interface to enable straightforward 'drag and drop' operation for deep learning programming. On the other hand, however, some general theoretical problems of learning such as the interpretation and generalization, have only achieved limited progress. Most deep learning papers published these days follow the pipeline of designing/modifying network structures - tuning parameters - reporting performance improvement in specific applications. We have even seen many deep learning application papers without one single equation. Theoretical interpretation and the science behind the study are largely ignored. While excited about the successful application of deep learning in classical and novel problems, we multimedia researchers are responsible to think and solve the fundamental topics in deep learning science. Prof. Guanrong Chen recently wrote an editorial note titled 'Science and Technology, not SciTech' [1]. This panel falls into similar discussion and aims to invite prestigious multimedia researchers and active deep learning practitioners to discuss the positioning of deep learning research now and in the future. Specifically, each panelist is asked to present their opinions on the following five questions: 1)How do you think the current phenomenon that deep learning applications are explosively growing, while the general theoretical problems remain slow progress? 2)Do you agree that deployment of deep learning techniques is getting easy (with a low barrier), while deep learning research is difficult (with a high barrier) 3)What do you think are the core problems for deep learning techniques? 4)What do you think are the core problems for deep learning science? 5)What's your suggestion on the multimedia research in the post-deep learning era? Jun Yu 0002, Ramesh Jain 0001, Rainer Lienhart, Peng Cui 0001, Jiashi Feng |
ACM Multimedia | 4 |
| 2018 | Activity-Conditioned Continuous Human Pose Estimation for Performance Analysis of Athletes Using the Example of SwimmingabstractIn this paper we consider the problem of human pose estimation in real-world videos of swimmers. Swimming channels allow filming swimmers simultaneously above and below the water surface with a single stationary camera. These recordings can be used to quantitatively assess the athletes' performance. The quantitative evaluation, so far, requires manual annotations of body parts in each video frame. We therefore apply the concept of CNNs in order to automatically infer the required pose information. Starting with an off-the-shelf architecture, we develop extensions to leverage activity information - in our case the swimming style of an athlete - and the continuous nature of the video recordings. Our main contributions are threefold: (a) We apply and evaluate a fine-tuned Convolutional Pose Machine architecture as a baseline in our very challenging aquatic environment and discuss its error modes, (b) we propose an extension to input swimming style information into the fully convolutional architecture and (c) modify the architecture for continuous pose estimation in videos. With these additions we achieve reliable pose estimates with up to +16% more correct body joint detections compared to the baseline architecture. Moritz Einfalt, Dan Zecha, Rainer Lienhart |
WACV | 3 |
| 2017 | A closer look: Small object detection in faster R-CNNabstractFaster R-CNN is a well-known approach for object detection which combines the generation of region proposals and their classification into a single pipeline. In this paper we apply Faster R-CNN to the task of company logo detection. Motivated by the weak performance of Faster R-CNN on small object instances, we perform a detailed examination of both the proposal and the classification stage, examining their behavior for a wide range of object sizes. Additionally, we look at the influence of feature map resolution on the performance of those stages. We introduce an improved scheme for generating anchor proposals and propose a modification to Faster R-CNN which leverages higher-resolution feature maps for small objects. We evaluate our approach on the Flicker data set improving the detection performance on small object instances. Christian Eggert, Stephan Brehm, Anton Winschel, Dan Zecha, Rainer Lienhart |
ICME | 5 |
| 2017 | Improving Small Object Proposals for Company Logo DetectionabstractMany modern approaches for object detection are two-staged pipelines. The first stage identifies regions of interest which are then classified in the second stage. Faster R-CNN is such an approach for object detection which combines both stages into a single pipeline. In this paper we apply Faster R-CNN to the task of company logo detection. Motivated by its weak performance on small object instances, we examine in detail both the proposal and the classification stage with respect to a wide range of object sizes. We investigate the influence of feature map resolution on the performance of those stages. Christian Eggert, Dan Zecha, Stephan Brehm, Rainer Lienhart |
ICMR | 4 |
| 2016 | Saliency-guided selective magnification for company logo detectionabstractFast R-CNN is a well-known approach to object detection which is generally reported to be robust to scale changes. In this paper we examine the influence of scale within the detection pipeline in the case of company logo detection. We demonstrate that Fast R-CNN encounters problems when handling objects which are significantly smaller than the receptive field of the utilized network. In order to overcome these difficulties, we propose a saliency-guided multiscale approach that does not rely on building a full image pyramid. We use the feature representation computed by Fast R-CNN to directly classify large objects while at the same time predicting salient regions which contain small objects with high probability. Only selected regions are magnified and a new feature representation for these enlarged regions is calculated. Feature representations from both scales are used for classification, improving the detection quality of small objects while keeping the computational overhead low. Compared to a naive magnification strategy we are able to retain 79% of the performance gain while only spending 36% of the computation time. Christian Eggert, Anton Winschel, Dan Zecha, Rainer Lienhart |
ICPR | 4 |
| 2016 | Towards automatic bounding box annotations from weakly labeled images
Christian X. Ries, Fabian Richter 0001, Rainer Lienhart |
Multim. Tools Appl. | 3 |
| 2016 | Erratum to: Towards automatic bounding box annotations from weakly labeled images
Christian X. Ries, Fabian Richter 0001, Rainer Lienhart |
Multim. Tools Appl. | 3 |
| 2015 | Fisher vector encoding of micro color features for (real world) jigsaw puzzlesabstractIn this work we propose a computationally efficient yet effective color encoding scheme for the purpose of jigsaw puzzle solving. Our contribution is twofold: (i) First of all, we show that a Fisher vector encoding gives superior performance over the commonly used compatibility descriptor of stacked color values. (ii) Furthermore, we show experimentally on synthetic and real-world data that our proposed color encoding is more robust in the presence of noise, as compared to three widely used features. Due to the robustness of our proposed descriptor we also anticipate its use to yield performance improvements in other applications, e.g., for the virtual reconstruction of hand-torn documents and archaeological findings. Fabian Richter 0001, Christian Eggert, Rainer Lienhart |
ICDAR | 3 |
| 2015 | On the Benefit of Synthetic Data for Company Logo DetectionabstractIn this paper we explore the benefits of synthetically generated data for the task of company logo detection with deep-learned features in the absence of a large training set. We use pre-trained deep convolutional neural networks for feature extraction and use a set of support vector machines for classifying those features. In order to generate sufficient training examples we synthesize artificial training images. Using a bootstrapping process, we iteratively add new synthesized examples from an unlabeled dataset to the training set. Using this setup we are able to obtain a performance which is close to the performance of the full training set. Christian Eggert, Anton Winschel, Rainer Lienhart |
ACM Multimedia | 3 |
| 2015 | The fertilized forests Decision Forest LibraryabstractSince the introduction of Random Forests in the 80's they have been a frequently used statistical tool for a variety of machine learning tasks. Many different training algorithms and model adaptions demonstrate the versatility of the forests. This variety resulted in a fragmentation of research and code, since each adaption requires its own algorithms and representations. In 2011, Criminisi and Shotton developed a unifying Decision Forest model for many tasks. By identifying the reusable parts and specifying clear interfaces, we extend this approach to an object oriented representation and implementation. This has the great advantage that research on specific parts of the Decision Forest model can be done 'locally' by reusing well-tested and high-performance components. Christoph Lassner, Rainer Lienhart |
ACM Multimedia | 2 |
| 2015 | Norm-Induced Entropies for Decision ForestsabstractThe entropy measurement function is a central element of decision forest induction. The Shannon entropy and other generalized entropies such as the Renyi and Tsallis entropy are designed to fulfill the Khinchin-Shannon axioms. Whereas these axioms are appropriate for physical systems, they do not necessarily model well the artificial system of decision forest induction. In this paper, we show that when omitting two of the four axioms, every norm induces an entropy function. The remaining two axioms are sufficient to describe the requirements for an entropy function in the decision forest context. Furthermore, we introduce and analyze the p-norm-induced entropy, show relations to existing entropies and the relation to various heuristics that are commonly used for decision forest training. In experiments with classification, regression and the recently introduced Hough forests, we show how the discrete and differential form of the new entropy can be used for forest induction and how the functions can simply be fine tuned. The experiments indicate that the impact of the entropy function is limited, however can be a simple and useful post-processing step for optimizing decision forests for high performance applications. Christoph Lassner, Rainer Lienhart |
WACV | 2 |
| 2015 | Key-Pose Prediction in Cyclic Human MotionabstractIn this paper we study the problem of estimating inner cyclic time intervals within repetitive motion sequences of top-class swimmers in a swimming channel. Interval limits are given by temporal occurrences of key-poses, i.e. distinctive postures of the body. A key-pose is defined by means of only one or two specific features of the complete posture. It is often difficult to detect such subtle features directly. We therefore propose the following method: Given that we observe the swimmer from the side, we build a pictorial structure of pose lets to robustly identify random support poses within the regular motion of a swimmer. We formulate a maximum likelihood model which predicts a key-pose given the occurrences of multiple support poses within one stroke. The maximum likelihood can be extended with prior knowledge about the temporal location of a key-pose in order to improve the prediction recall. We experimentally show that our models reliably and robustly detect key-poses with a high precision and that their performance can be improved by extending the framework with additional camera views. Dan Zecha, Rainer Lienhart |
WACV | 2 |
| 2015 | Guest editorial: selected papers from ICIMCS 2012
Zhengjun Zha, Yan Liu 0004, Shin'ichi Satoh 0001, Xinguo Yu, Rainer Lienhart |
Multim. Syst. | 5 |
| 2014 | Evaluation of Discriminative Models for the Reconstruction of Hand-Torn Documents
Fabian Richter 0001, Christian X. Ries, Rainer Lienhart |
ACCV (3) | 3 |
| 2014 | Improving VLAD: Hierarchical coding and a refined local coordinate systemabstractThe enormous growth of image databases calls for new techniques for fast and effective image search that scales with millions of images. Most importantly, the setting requires a compact but also descriptive image signature. Recently, the vector of aggregated local descriptors (VLAD) [1] has received much attention in large-scale image retrieval. In this paper we present two modifications for VLAD which improve the retrieval performance of the signature. Christian Eggert, Stefan Romberg, Rainer Lienhart |
ICIP | 3 |
| 2014 | Partial contour matching for document pieces with content-based priorabstractIn this paper we present a method for aligning shredded document pieces based on outer contours and content-based prior information. Our approach relies on domain-specific knowledge that document pieces must complement each other when aligned correctly. Building on this intuition we propose a variant of MSAC (M-estimator SAmple Consensus) to estimate an hypothesis that recovers the spatial relationship between pairs of pieces. To do so we first approximate their boundaries by polygons from which we define consensus sets between fragments. Each consensus set provides multiple hypotheses for aligning one piece onto the other. An optimal hypothesis is identified by applying a two-stage procedure in which we discard locally inconsistent hypotheses before verifying the remainder for global consistency. Fabian Richter 0001, Christian X. Ries, Stefan Romberg, Rainer Lienhart |
ICME | 4 |
| 2014 | A survey on visual adult image recognition
Christian X. Ries, Rainer Lienhart |
Multim. Tools Appl. | 2 |
| 2013 | Methods and Applications for Distance Based ANN TrainingabstractFeature learning has the aim to take away the hassle of hand-designing features for machine learning tasks. Since the feature design process is tedious and requires a lot of experience, an automated solution is of great interest. However, an important problem in this field is that usually no objective values are available to fit a feature learning function to. Artificial Neural Networks are a sufficiently flexible tool for function approximation to be able to avoid this problem. We show how the error function of an ANN can be modified such that it works solely with objective distances instead of objective values. We derive the adjusted rules for back propagation through networks with arbitrary depths and include practical considerations that must be taken into account to apply difference based learning successfully. On all three benchmark datasets we use, linear SVMs trained on automatically learned ANN features outperform RBF kernel SVMs trained on the raw data. This can be achieved in a feature space with up to only a tenth of dimensions of the number of original data dimensions. We conclude our work with two experiments on distance based ANN training in two further fields: data visualization and outlier detection. Christoph Lassner, Rainer Lienhart |
ICMLA (2) | 2 |
| 2013 | Towards automatic object annotations from global image labelsabstractWe present an approach for automatically devising object annotations in images. Thus, given a set of images which are known to contain a common object, our goal is to find a bounding box for each image which tightly encloses the object. In contrast to regular object detection, we do not assume any previous manual annotations except for binary global image labels. We first use a discriminative color model for initializing our algorithm by very coarse bounding box estimations. We then narrow down these boxes using visual words computed from HOG features. Finally, we apply an iterative algorithm which trains a SVM model based on bag-of-visual-words histograms. During each iteration, the model is used to find better bounding boxes which can be done efficiently by branch and bound. The new bounding boxes are then used to retrain the model. We evaluate our approach for several different classes of publicly available datasets and show that we obtain promising results. Christian X. Ries, Fabian Richter 0001, Rainer Lienhart |
ICMR | 3 |
| 2013 | Bundle min-hashing for logo recognitionabstractWe present a scalable logo recognition technique based on feature bundling. Individual local features are aggregated with features from their spatial neighborhood into bundles. These bundles carry more information about the image content than single visual words. The recognition of logos in novel images is then performed by querying a database of reference images. Stefan Romberg, Rainer Lienhart |
ICMR | 2 |
| 2013 | Learning to Reassemble Shredded DocumentsabstractIn this paper, we address the problem of automatically assembling shredded documents. We propose a two-step algorithmic framework. First, we digitize each fragment of a given document and extract shape- and content-based local features. Based on these multimodal features, we identify pairs of corresponding points on all pairs of fragments using an SVM classifier. Each pair is considered a point of attachment for aligning the respective fragments. In order to restore the layout of the document, we create a document graph in which nodes represent fragments and edges correspond to alignments. We assign weights to the edges by evaluating the alignments using a set of inter-fragment constraints which take into account shape- and content-based information. Finally, we use an iterative algorithm that chooses the edge having the highest weight during each iteration. However, since selecting edges corresponds to combining groups of fragments and thus provides new evidence, we reevaluate the edge weights after each iteration. We quantitatively evaluate the effectiveness of our approach by conducting experiments on a novel dataset. It comprises a total of 120 pages taken from two magazines which have been shredded and annotated manually. We thus provide the means for a quantitative evaluation of assembly algorithms which, to the best of our knowledge, has not been done before. Fabian Richter 0001, Christian X. Ries, Nicolas Cebron, Rainer Lienhart |
IEEE Trans. Multim. | 4 |
| 2013 | Introduction to the special section on the 20th anniversary of the ACM international conference on multimediaabstractintroduction Introduction to the special section on the 20th anniversary of the ACM international conference on multimedia Authors: Klara Nahrstedt View Profile , Rainer Lienhart View Profile , Malcolm Slaney View Profile Authors Info & Claims ACM Transactions on Multimedia Computing, Communications, and ApplicationsVolume 9Issue 1sOctober 2013 Article No.: 32pp 1–3https://doi.org/10.1145/2523001.2523003Published:17 October 2013Publication History 0citation112DownloadsMetricsTotal Citations0Total Downloads112Last 12 Months0Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Klara Nahrstedt, Rainer Lienhart, Malcolm Slaney |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2012 | Learning an object class representation on a continuous viewsphereabstractWe propose an approach to multi-view object class detection and approximate 3D pose estimation. It relies on CAD models as positive training examples and discriminatively learns photometric object parts such that an optimal coverage of intra-class and viewpoint variation is guaranteed. In contrast to previous work, the approach shows a significantly reduced training set dependency while avoiding any manual training supervision or annotation, since it is capable of deriving all relevant information exclusively from the provided set of 3D CAD models and an arbitrary set of 2D negative images. In entirely circumventing semantic or view-based representations, part symmetries and co-occurrences between viewpoints can be efficiently exploited. This, in turn, leads to a significantly lower complexity while still achieving state-of-the-art performance on two current benchmark data sets for two different object classes. Johannes Schels, Joerg Liebelt, Rainer Lienhart |
CVPR | 3 |
| 2012 | Decision Tree Induction from Counterexamples
Nicolas Cebron, Fabian Richter 0001, Rainer Lienhart |
ICPRAM (2) | 3 |
| 2012 | Deriving a discriminative color model for a given object class from weakly labeled training dataabstractThis paper presents a method for creating a discriminative color model for a given object class based on color occurrence statistics. A discriminative color model can be used to classify individual pixels of images with regards to whether they may belong to the wanted object. However, in contrast to existing approaches, we do not exploit pixel-wise object annotations but only global negative and positive image labels. Therefore our approach requires significantly less manual effort. We quantitatively evaluate the performance of our approach on two publicly available datasets and compare it to a baseline approach, which utilizes pixel annotations. The experimental results show that our approach is on par with pixel-wise approaches although requiring only a single global image label. Christian X. Ries, Rainer Lienhart |
ICMR | 2 |
| 2012 | Leveraging community metadata for multimodal image ranking
Fabian Richter 0001, Stefan Romberg, Eva Hörster, Rainer Lienhart |
Multim. Tools Appl. | 4 |
| 2011 | Monocular 3D human pose estimation by classificationabstractWe present a novel approach to 2D and 3D human pose estimation in monocular images by building on and improving recent advances in this field. We take the full body pose as a combination of a 3D pose and a viewpoint and in this way define classes that are then learned by a classifier. Compared to part based approaches, our approach does not suffer from self-occluded body parts since such occlusions are characteristic for certain classes and thus are captured during class definition. Moreover, we significantly relax the requirements posed on training data by the fact that we do neither require labeled viewpoints nor background subtracted images, and the carried out action does not need to be cyclic. By combining an efficient classifier with efficient image features, we present a generic and fast way to estimate human poses in images and achieve comparable results to state-of-the art approaches which we demonstrate on a public benchmark. Thomas Greif, Rainer Lienhart, Debabrata Sengupta |
ICME | 2 |
| 2011 | A graph algorithmic framework for the assembly of shredded documentsabstractIn this paper we propose a framework to address the reassembly of shredded documents. Inspired by the way humans approach this problem we introduce a novel algorithm that iteratively determines groups of fragments that fit together well. We identify such groups by evaluating a set of constraints that takes into account shape- and content-based information of each fragment. Accordingly, we choose the best matching groups of fragments during each iteration and implicitly determine a maximum spanning tree of a graph that represents alignments between the individual fragments. After each iteration we update the graph with respect to additional contextual knowledge. We evaluate the effectiveness of our approach on a dataset of 16 fragmented pages with strongly varying content. The robustness of the proposed algorithm is finally shown in situations in which material is lost. Fabian Richter 0001, Christian X. Ries, Rainer Lienhart |
ICME | 3 |
| 2011 | Building a semantic part-based object class detector from synthetic 3D modelsabstractThis paper presents a new approach for multi-view object class detection based on part models. While most existing approaches have in common that they use real images for training, our approach requires only a database of synthetic 3D models to represent both the appearance and the geometry of an object class. We use semantically equivalent object points on 3D models to build part models and encode the local appearance of the parts by a discriminative learning method that applies AdaBoost to histograms of gradients. The geometric configuration of the parts is represented by spatial distributions which are also directly derived from the 3D models. For recognizing an object in an image, our model provides object hypotheses which are re-ranked with global appearance models. The 2D localization is evaluated on the PASCAL 2006 data set for cars and bicycles, showing that its performance can compete with state-of-the-art detection results. Johannes Schels, Joerg Liebelt, Klaus Schertler, Rainer Lienhart |
ICME | 4 |
| 2011 | Scalable logo recognition in real-world imagesabstractIn this paper we propose a highly effective and scalable framework for recognizing logos in images. At the core of our approach lays a method for encoding and indexing the relative spatial layout of local features detected in the logo images. Based on the analysis of the local features and the composition of basic spatial structures, such as edges and triangles, we can derive a quantized representation of the regions in the logos and minimize the false positive detections. Furthermore, we propose a cascaded index for scalable multi-class recognition of logos.For the evaluation of our system, we have constructed and released a logo recognition benchmark which consists of manually labeled logo images, complemented with non-logo images, all posted on Flickr. The dataset consists of a training, validation, and test set with 32 logo-classes. We thoroughly evaluate our system with this benchmark and show that our approach effectively recognizes different logo classes with high precision. Stefan Romberg, Lluís Garcia Pueyo, Rainer Lienhart, Roelof van Zwol |
ICMR | 3 |
| 2011 | Synthetically trained multi-view object class and viewpoint detection for advanced image retrievalabstractThis paper proposes a novel approach to multi-view object class and viewpoint detection for the retrieval of images showing one or several objects from a given viewpoint, a viewpoint range or any viewpoint in image databases. All detectors are trained exclusively on a few synthetic 3D models without any manual bounding-box, viewpoint or part annotation, making object class and viewpoint detection a scalable learning task. Previous work on this topic relies on the detection of object parts for each individual viewpoint, ignoring the responses of part detectors specific to other viewpoints. Instead, we explicitly exploit appearance ambiguities caused by spurious detections of parts under more than one viewpoint by combining all detector responses in a joint spatial pyramid encoding. We achieve state-of-the-art results in multi-view object class detection and viewpoint determination on current benchmarking data sets and demonstrate increased robustness to partial occlusion. Johannes Schels, Joerg Liebelt, Klaus Schertler, Rainer Lienhart |
ICMR | 4 |
| 2011 | Towards synergy between the open source and the research multimedia communitiesabstractThis panel extends current efforts from the ACM Multimedia 2011 Organization Committee in taking an important step towards open source projects. The panelists include speakers who are among the leading figures from the open source community. The goal is to provide a shared space for discussion and interaction among consolidated and new open source projects and multimedia researchers. Pablo César, Wei Tsang Ooi, Ben Moskowitz, Zohar Babin, Dick C. A. Bulterman, Rainer Lienhart, Robert Richter |
ACM Multimedia | 6 |
| 2010 | Towards universal visual vocabulariesabstractMany content-based image mining systems extract local features from images to obtain an image description based on discrete feature occurrences. Such applications require a visual vocabulary also known as visual codebook or visual dictionary to discretize the extracted high-dimensional features to visual words in an efficient yet accurate way. Once such an application operates on images of a very specific domain the question arises if a vocabulary built from those domain-specific images needs to be used or if a ”universal” visual vocabulary can be used instead. A universal visual vocabulary may be computed from images of a different domain once and then be re-used for various applications and other domains. We therefore evaluate several visual vocabularies from different image domains by determining their performance at pLSA-based image classification on several datasets. We empirically conclude that vocabularies suit our classification tasks equally well disregarding the image domain they were derived from. Christian X. Ries, Stefan Romberg, Rainer Lienhart |
ICME | 3 |
| 2009 | Note onset detection for the transcription of polyphonic piano musicabstractTranscription of music is the process of generating a symbolic representation such as a score sheet or a MIDI file from an audio recording of a piece of music. A statistical machine learning approach for detecting note onsets in polyphonic piano music is presented. An area from the spectrogram of the sound is concatenated into one feature vector. A cascade of boosted classifiers is used for dimensionality reduction and classification in an one-versus-all manner. The presented system achieves an accuracy of 87.4% in onset detection outperforming the best comparison system by 25.1 %. C. Gregor v. d. Boogaart, Rainer Lienhart |
ICME | 2 |
| 2009 | Filtering adult image content with topic modelsabstractProtecting children from exposure to adult content has become a serious problem in the real world. Current statistics show that, for instance, the average age of first Internet exposure to pornography is 11 years, that the largest consumer group of Internet pornography is the age group of 12-to-17- year-olds and that 90% of the 8-to-16-year-olds have viewed porn online. To protect our children, effective algorithms for detecting adult images are needed. In this research we evaluate the use of probabilistic Latent Semantic Analysis (pLSA) for this task. We will show that topic models based on pLSA can detect adult content with a correct positive rate of 92.7%, while only showing off a false positive rate of 1.9%. Even when using grayscale images only, a correct positive rate of 90.8% at a false positive rate of 2% can be achieved. Rainer Lienhart, Rudolf Hauke |
ICME | 1 |
| 2009 | Multimodal pLSA on visual features and tagsabstractThis work studies a new approach for image retrieval on largescale community databases. Our proposed system explores two different modalities: visual features and community-generated metadata, such as tags. We use topic models to derive a high-level representation appropriate for retrieval for each of our images in the database. We evaluate the proposed approach experimentally in a query-by-example retrieval task and compare our results to systems relying solely on visual features or tag features. It is shown that the proposed multimodal system outperforms the unimodal systems by approximately 36%. Stefan Romberg, Eva Hörster, Rainer Lienhart |
ICME | 3 |
| 2008 | Deep networks for image retrieval on large-scale databasesabstractCurrently there are hundreds of millions (high-quality) images in online image repositories such as Flickr. This makes is necessary to develop new algorithms that allow for searching and browsing in those large-scale databases. In this work we explore deep networks for deriving a low-dimensional image representation appropriate for image retrieval. A deep network consisting of multiple layers of features aims to capture higher order correlations between basic image features. We will evaluate our approach on a real world large-scale image database and compare it to image representations based on topic models. Our results show the suitability of the approach for very large databases. Eva Hörster, Rainer Lienhart |
ACM Multimedia | 2 |
| 2007 | Fusing Local Image Descriptors for Large-Scale Image RetrievalabstractOnline image repositories such as Flickr contain hundreds of millions of images and are growing quickly. Along with that the needs for supporting indexing, searching and browsing is becoming more and more pressing. Here we will employ the image content as a source of information to retrieve images and study the representation of images by topic models for content-based image retrieval. We focus on incorporating different types of visual descriptors into the topic modeling context. Three different fusion approaches are explored. The image representations for each fusion approach are learned in an unsupervised fashion, and each image is modeled as a mixture of topics/object parts depicted in the image. However, not all object classes will benefit from all visual descriptors. Therefore, we also investigate which visual descriptor (set) is most appropriate for each of the twelve classes under consideration. We evaluate the presented models on a real world image database consisting of more than 246,000 images. Eva Hörster, Rainer Lienhart |
CVPR | 2 |
| 2007 | PLSA on Large Scale Image DatabasesabstractThe Web and image repositories such as Fickrtrade are the largest image databases in the world. There are billions of images on the web, and hundreds of million high-quality images in image repositories. Currently, these images are indexed based on manually-entered tags and individual and group usage patterns. In this work we a exploring a third information dimension: image features. We are exploring probabilistic latent semantic analysis in order to infer which visual patterns describe each object. We wish to build models that connect words and image features, and use content features and tags to better find similar images. Rainer Lienhart, Malcolm Slaney |
ICASSP (4) | 1 |
| 2006 | Fast Gabor Transformation For Processing High Quality AudioabstractThe Gabor transformation with a Gaussian window has several advantages over classical short time transformations such as the windowed FFT and Wavelets. It allows for perfect localization in time and frequency according to the absolute bound expressed by the Heisenberg uncertainty principle. Furthermore the time-frequency resolution can be chosen as desired. This is bought dearly by the necessity of oversampling and very large windows resulting in high computational and storage costs. To overcome this disadvantages the FFT can be used as an underlying technique to speed up the computation for some rare dedicated time-frequency resolutions. In this paper the use of the FFT is extended to allow choosing the timefrequency resolution arbitrarily, by introducing only a small computational overhead. With the same approach we show, how to use the FFT to compute the Gabor transformation on non-separable lattices. The speed-up factors over an optimized DFT approach range from 2.5 to 100. C. Gregor v. d. Boogaart, Rainer Lienhart |
ICASSP (3) | 2 |
| 2006 | Approximating Optimal Visual Sensor PlacementabstractMany novel multimedia applications use visual sensor arrays. In this paper we address the problem of optimally placing multiple visual sensors in a given space. Our linear programming approach determines the minimum number of cameras needed to cover the space completely at a given sampling frequency. Simultaneously it determines the optimal positions and poses of the visual sensors. We also show how to account for visual sensors with different properties and costs if more than one kind is available, and report performance results. Eva Hörster, Rainer Lienhart |
ICME | 2 |
| 2006 | Calibrating and optimizing poses of visual sensors in distributed platforms
Eva Hörster, Rainer Lienhart |
Multim. Syst. | 2 |
| 2005 | Precise Visibility Determination of Displays in Camera ImagesabstractIn many novel application scenarios such as smart rooms or sensing rooms visual sensors (such as cameras) need to know which visual actuators (such as displays) are visible to them. Often only parts of a display are visible from a camera. Therefore, a novel algorithm for precise visibility determination is presented. The algorithm makes the assumption that the displays are active, i.e., they can be controlled by the application. Under these conditions the algorithm determines precisely where which parts of a display are imaged by a camera Eva Hörster, Rainer Lienhart, Walter Kellermann, Jean-Yves Bouguet |
ICME | 2 |
| 2005 | Position calibration of microphones and loudspeakers in distributed computing platformsabstractWe present a novel algorithm to automatically determine the relative three-dimensional (3-D) positions of audio sensors and actuators in an ad-hoc distributed network of heterogeneous general purpose computing platforms such as laptops, PDAs, and tablets. A closed form approximate solution is derived, which is further refined by minimizing a nonlinear error function. Our formulation and solution accounts for the lack of temporal synchronization among different platforms. We compare two different estimators, one based on the time of flight and the other based on time difference of flight. We also derive an approximate expression for the mean and covariance of the implicitly defined estimator using the implicit function theorem and approximate Taylors' series expansion. The theoretical performance limits for estimating the sensor 3-D positions are derived via the Crame/spl acute/r-Rao bound (CRB) and analyzed, with respect to the number of sensors and actuators, as well as their geometry. We report extensive simulation results and discuss the practical details of implementing our algorithms in a real-life system. Vikas C. Raykar, Igor Kozintsev, Rainer Lienhart |
IEEE Trans. Speech Audio Process. | 3 |
| 2004 | Providing common I/O clock for wireless distributed platformsabstractWe propose a novel synchronization scheme for distributed audio-video input and output on heterogeneous general purpose computers (GPC) such as laptops, tablets, PDA, smart telephones, audio recorders, and camcorders. These devices typically possess sensors such as microphones and possibly cameras, and actuators such as loudspeakers and displays. In order to combine them wirelessly into a distributed array signal processing system, it is necessary to provide relative time synchronization to sensors and actuators. In this work we propose a setup and an algorithm to synchronize input and output for a network of distributed multichannel audio sensors and actuators connected to GPC. An IEEE 802.11 wireless network is used to deliver the global clock to distributed GPC, while the interrupt mechanism is employed to distribute the clock between I/O devices. Experimental results demonstrate a precision in A/D D/A synchronization precision better than 50 /spl mu/s (a couple of samples at 48 kHz). Dmitry Budnikov, Igor Chikalov, Sergey Egorychev, Igor Kozintsev, Rainer Lienhart |
ICASSP (3) | 5 |
| 2004 | Self-aware distributed AV sensor and actuator networks for improved media adaptationabstractMost of the existing research work in the area of media adaptation is concentrated on content adaptation, transcoding and delivery mechanisms without addressing the actual input and output of multimedia data. However, it is the I/O stage of media processing that humans are concerned about. Up until now, most multimedia applications have relied on standalone I/O devices (microphone, headphones, monitor, camera) to capture or render multimedia data. This situation is about to change. Nowadays we are surrounded by a vast number of audio/video (AV) sensors and actuators. They are built into our cellular phones, PDAs, tablets, laptops, and surveillance systems. A natural idea that comes out of this fact is to combine multiple I/O devices into a distributed array of sensors and actuators. The paper shows the feasibility of this idea and shifts media adaptation research away from a single device/stream paradigm towards array multimedia processing. We demonstrate how to transform a network of off-the-shelf devices into a distributed I/O array by providing common time (with tens of microseconds precision) and 3D space coordinates (with a few centimetres precision). We also discuss the implications and potentials of self-calibrating distributed AV-sensor/actuator networks for improved media adaptation. Rainer Lienhart, Igor Kozintsev |
ICME | 1 |
| 2004 | Between context-aware media capture and multimedia content analysis: where do we find the promised land?abstractNo abstract available. Susanne Boll, Dick C. A. Bulterman, Ramesh Jain 0001, Tat-Seng Chua, Rainer Lienhart, Lynn Wilcox, Marc Davis, Svetha Venkatesh |
ACM Multimedia | 5 |
| 2003 | On the importance of exact synchronization for distributed audio signal processingabstractWe propose a new paradigm for implementations of audio array processing algorithms on a network of distributed general-purpose computers. In contrast to currently existing DSP processor-based solutions, our approach offers new possibilities for advanced array signal processing by enabling the usage of general-purpose computing platforms with their superior computational and storage resources. We demonstrate that synchronization of sensors is essential for acoustic blind source separation (BSS) algorithms, and we propose a synchronization scheme that enables BSS on distributed, wirelessly networked computers and can easily be implemented on existing hardware. Rainer Lienhart, Igor Kozintsev, Stefan Wehr, Minerva M. Yeung |
ICASSP (4) | 1 |
| 2003 | A detector tree of boosted classifiers for real-time object detection and trackingabstractThis paper presents a novel tree classifier for complex object detection tasks together with a general framework for real-time object tracking in videos using the novel tree classifier. A boosted training algorithm with a clustering-and-splitting step is employed to construct branches in the nodes recursively, if and only if it improves the discriminative power compared to a single monolithic node classifier and has a lower computational complexity. A mouth tracking system that integrates the tree classifier under the proposed framework is built and tested on XM2FDB database. Experimental results show that the detection accuracy is equal or better than a single or multiple cascade classifier, while being computational less demanding. Rainer Lienhart, Luhong Liang, Alexander Kuranov |
ICME | 1 |
| 2003 | Universal synchronization scheme for distributed audio-video capture on heterogeneous computing platformsabstractWe propose a universal synchronization scheme for distributed audio-video capture on heterogeneous computing devices such as laptops, tablets, PDAs, cellular phones, audio recorders, and camcorders. These devices typically possess sensors such as microphones and possibly cameras. In order to combine them wirelessly into a distributed sensing and computing system, it is necessary to provide relative time synchronization among the distributed sensors. In this work we propose a setup and an algorithm that provide synchronization between sampling times for a network of distributed multi-channel audio sensors connected to general purpose computing (GPC) platforms. Extensive experimental results on distributed acoustic Blind Source Separation (BSS) algorithms validate the performance of our synchronization scheme. Rainer Lienhart, Igor Kozintsev, Stefan Wehr |
ACM Multimedia | 1 |
| 2003 | Position calibration of audio sensors and actuators in a distributed computing platformabstractIn this paper, we present a novel approach to automatically determine the positions of sensors and actuators in an ad-hoc distributed network of general purpose computing platforms. The formulation and solution accounts for the limited precision in temporal synchronization among multiple platforms. The theoretical performance limit for the sensor positions is derived via the Cramer-Rao bound. We analyze the sensitivity of localization accuracy with respect to the number of sensors and actuators as well as their geometry. Extensive Monte Carlo simulation results are reported together with a discussion of the real-time system. In a test platform consisting of 4 speakers and 4 microphones, the sensors' and actuators' three dimensional locations could be estimated with an average bias of 0.08 cm and average standard deviation of 3.8 cm. Vikas C. Raykar, Igor Kozintsev, Rainer Lienhart |
ACM Multimedia | 3 |
| 2002 | An extended set of Haar-like features for rapid object detectionabstractRecently Viola et al. [2001] have introduced a rapid object detection. scheme based on a boosted cascade of simple feature classifiers. In this paper we introduce a novel set of rotated Haar-like features. These novel features significantly enrich the simple features of Viola et al. and can also be calculated efficiently. With these new rotated features our sample face detector shows off on average a 10% lower false alarm rate at a given hit rate. We also present a novel post optimization procedure for a given boosted cascade improving on average the false alarm rate further by 12.5%. Rainer Lienhart, Jochen Maydt |
ICIP (1) | 1 |
| 2002 | A fast method for training support vector machines with a very large set of linear featuresabstractCurrent systems for object detection often use support vector machines (SVM) as the basic classification algorithm. A rather common case is to compute a small set of linear features and then train the classifier on these features. We present a fast method to train and evaluate SVM with many linear features and show results for face detection using a set of 210400 features. The resulting classifier is both more accurate and faster than a classifier trained on raw pixel features, which total up only to 576 features in our case. Jochen Maydt, Rainer Lienhart |
ICME (1) | 2 |
| 2002 | Evaluating and Improving Performance of Multimedia Applications on Simultaneous Multi-ThreadingabstractThis paper presents the study and results of running several core multimedia applications on a simultaneous multithreading (SMT) architecture, including some detailed analysis ranging from memory-bounded kernels to computational-bounded functions. A performance metric to evaluate effective SMT performance gain is introduced, and compared to similar metrics on symmetric multiprocessor (SMP) systems. In addition, we analyze and compare SMT versus SMP systems, and highlight the advantages in the studied applications. The results indicate that sharing the cache in SMT processors can provide better cache locality and thus better performance although sharing the cache can introduce cache conflicts and reduce the actual cache size available for each logical processor. We also propose "mutual prefetching" -a technique to schedule threads so that they prefetch data for each other in order to reduce cache miss penalty. Yen-Kuang Chen, Eric Debes, Rainer Lienhart, Matthew J. Holliman, Minerva M. Yeung |
ICPADS | 3 |
| 2002 | Localizing and segmenting text in images and videosabstractMany images, especially those used for page design on Web pages, as well as videos contain visible text. If these text occurrences could be detected, segmented, and recognized automatically, they would be a valuable source of high-level semantics for indexing and retrieval. We propose a novel method for localizing and segmenting text in complex images and videos. Text lines are identified by using a complex-valued multilayer feed-forward network trained to detect text at a fixed scale and position. The network's output at all scales and positions is integrated into a single text-saliency map, serving as a starting point for candidate text lines. In the case of video, these candidate text lines are refined by exploiting the temporal redundancy of text in video. Localized text lines are then scaled to a fixed height of 100 pixels and segmented into a binary image with black characters on white background. For videos, temporal redundancy is exploited to improve segmentation performance. Input images and videos can be of any size due to a true multiresolution approach. Moreover, the system is not only able to locate and segment text occurrences into large binary images, but is also able to track each text line with sub-pixel accuracy over the entire occurrence in a video, so that one text bitmap is created for all instances of that text line. Therefore, our text segmentation results can also be used for object-based video encoding such as that enabled by MPEG-4. Rainer Lienhart, Axel Wernicke |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2001 | A system for reliable dissolve detection in videosabstractAutomatic shot boundary detection has been an active research area for nearly a decade and has led to high performance detection algorithms for hard cuts, fades and wipes. Reliable dissolve detection, however, is still an unsolved problem. We present the first robust and reliable dissolve detection system. A detection rate of 75% was achieved while reducing the false alarm rate to an acceptable level of 16% on a test video set for which so far the best reported detection and false alarm rate had been 66% and 59%, respectively. In addition, a dissolve's temporal extent is estimated, too. The core ideas of our novel approach are firstly the creation of a dissolve synthesizer capable of creating in principle an infinite number of dissolve examples of any duration from a video database of raw video footage allowing us to use an advanced machine learning algorithm such as neural networks and support vector machines which require large training sets, secondly, two simple features capturing the characteristics of dissolves, thirdly, a fully temporal multi-resolution search based on a fixed position and fixed-scale transition/special effect detector enabling us to determine also the true duration of detected dissolves, and finally, a post processing step which uses global motion estimation to further reduce the number of falsely detected dissolves. Rainer Lienhart, André Zaccarin |
ICIP (3) | 1 |
| 2001 | Scene Determination Based on Video and Audio Features
Silvia Pfeiffer, Rainer Lienhart, Wolfgang Effelsberg |
Multim. Tools Appl. | 2 |
| 2000 | Automatic Text Segmentation and Text Recognition for Video Indexing
Rainer Lienhart, Wolfgang Effelsberg |
Multim. Syst. | 1 |
| 1999 | Abstracting home video automaticallyabstractIn this paper, we present new algorithms for generating amusing, visually appealing and variable video abstracts of home video material automatically.The algorithms make use of a new, empirically motivated approach to cluster time-stamped shots hierachically into meaningful units.Moreover, our algorithms are not restricted to home video but can also be applied to raw video footage in general.1.1 Rainer Lienhart |
ACM Multimedia (2) | 1 |
| 1999 | VisualGREP: A Systematic Method to Compare and Retrieve Video Sequences
Rainer Lienhart, Wolfgang Effelsberg, Ramesh Jain 0001 |
Multim. Tools Appl. | 1 |
| 1996 | Automatic Text Recognition for Video IndexingabstractArticle Automatic text recognition for video indexing Share on Author: Rainer Lienhart University of Mannheim, Praktische Informatik IV, 68131 Mannheim, Germany University of Mannheim, Praktische Informatik IV, 68131 Mannheim, GermanyView Profile Authors Info & Claims MULTIMEDIA '96: Proceedings of the fourth ACM international conference on MultimediaFebruary 1997 Pages 11–20https://doi.org/10.1145/244130.244137Online:01 February 1997Publication History 76citation1,608DownloadsMetricsTotal Citations76Total Downloads1,608Last 12 Months12Last 6 weeks4 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Rainer Lienhart |
ACM Multimedia | 1 |
| 1996 | Abstracting Digital Movies Automatically
Silvia Pfeiffer, Rainer Lienhart, Stephan Fischer 0001, Wolfgang Effelsberg |
J. Vis. Commun. Image Represent. | 2 |
| 1995 | Automatic Recognition of Film GenresabstractNo abstract available. Stephan Fischer 0001, Rainer Lienhart, Wolfgang Effelsberg |
ACM Multimedia | 2 |
| 1995 | Automatic Recognition of Film Genres (Demonstration)abstractNo abstract available. Stephan Fischer 0001, Rainer Lienhart, Wolfgang Effelsberg |
ACM Multimedia | 2 |