VLDB 2026 Research / reviewers in the wild / expert
Klaus Diepold
dblp:27/5714 · also Klaus J. Diepold, Klaus-Jürgen Diepold
· DBLP profile ↗
74ranked-venue papers
6as first author
14since 2021 · last 2025
0000-0003-0439-7511ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 36 · 3 first-author · 2 since 2021Artificial intelligence and machine learning · 25 · 1 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 1 first-author · 5 since 2021Systems, architecture and hardware · 8 · 1 first-author · 3 since 2021Security and privacy · 5Human-computer interaction and ubiquitous computing · 5 · 1 first-authorDatabases, data management, data science and information retrieval · 3Computer networks · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Classifying Music-Induced Emotion Using Multi-Modal Ensembles of EEG and Audio Feature ModelsabstractIn this paper, we present our submission to the EEG-Music Emotion Recognition Challenge at ICASSP 2025. Our work focused on Task 2, where the objective was to classify the emotional state of subjects in a discrete valence-arousal space while they listen to music. Our proposed solution adopts an ensemble approach, integrating electroencephalography (EEG) signals, raw audio signals, and song features. We incorporated a diverse set of models, including Audio Spectrogram Transformer (AST) [1] and a dedicated EEG model. By combining insights from these modalities, we aimed to capture the interplay between music and emotional responses. We achieved a balanced accuracy of 41.34% on held-out data, improving the baseline by more than 11% and thus 2ndplace in the competition. Philipp Paukner, Marisa Ripoll, Dilvan Sabir, Deniz Onat Erdogan, Luca Sacchetto, Klaus Diepold |
ICASSP | 6 |
| 2024 | AI-Enhanced Detection of Cellular Aggregate Biomarkers For Point-of-Care Using Digital Holographic MicroscopyabstractIn the acute clinical setting, stratification of patients presenting with fever is a challenge that often results in high diagnostic costs and inpatient monitoring, leading to substantial resource utilization. Cellular aggregation patterns promise to reveal information about the underlying immunology mechanisms of fever and provide new biomarkers to facilitate triage and diagnosis. However, clinical utilization requires rapid imaging and measurement of these patterns, currently these tools are unavailable. This study introduces the YOLO-Graph model for the fast analysis of the aggregates from high-throughput cellular imaging data, to enhance early diagnostic capabilities in clinical practice. The core of our methodology is built on integrating digital holographic microscopy with the YOLOv8x-p2 neural network and graph model, enabling precise morphological detection and quantitative assessment of cellular aggregates. Our detector achieved 91.14% accuracy for aggregate identification, marking a 13.92% improvement over the benchmark pipeline.Using this rapid analysis of clinical samples, our YOLOGraph model identified specific cellular aggregate patterns during immune reactions. We showed that the mean value of platelet-platelet and leukocyte-platelet aggregates from healthy controls is lower than that of fever and pneumonia patients. Identifying specific patterns associated with clinically important diagnoses and outcomes will allow the development of biomarkers for early triage and diagnosis. Matthew Edward Cove, Kerem Delikoyun, Klaus Diepold, Oliver Hayden, Win Sen Kuan, John Tshon Yit Soong |
BIBM | 4 |
| 2024 | Feature Extraction by Image Transformations for Cell Image ClassificationabstractWhen it comes to decision-making in the medical context, opaque black box models, as accurate as they may be, have still not caught on in clinical practice. To present an alternative to the ever-growing convolutional neural networks, in this paper, we focus on classical image transformations, which are used in established compression methods. Using the example of modern high-throughput cytology by means of label-free quantitative phase microscopy, we want to evaluate the transformed cell features using comprehensible classifiers. We can show that the resulting transparent pipelines are roughly comparable to previous work using classical or black box models. The optimized methods generalize somewhat worse on unknown cells and outliers but represent a good trade-off between transparency and performance. Stefan Röhrl, Franziska Steinmetz, Philipp Paukner, Manuel Lengl, Simon Schumann, David Elias Fresacher, Christian Klenk, Dominik Heim, Martin Knopp, Katja Peschke, Maximilian Reichert, Oliver Hayden, Klaus Diepold |
BIBM | 13 |
| 2024 | Modeling the Emergence of Letter Shapes
Alice Hein, Klaus Diepold |
CogSci | 2 |
| 2023 | Explainable Artificial Intelligence for Cytological Image Analysis
Stefan Röhrl, Hendrik Maier, Manuel Lengl, Christian Klenk, Dominik Heim, Martin Knopp, Simon Schumann, Oliver Hayden, Klaus Diepold |
AIME | 9 |
| 2023 | Comparing Intuitions about Agents' Goals, Preferences and Actions in Human Infants and Video Transformers
Alice Hein, Klaus Diepold |
CogSci | 2 |
| 2023 | Comparing Quadrotor Control Policies for Zero-Shot Reinforcement Learning under Uncertainty and Partial ObservabilityabstractTo alleviate the sample complexity of reinforcement learning algorithms, simulations are a common approach to train control policies before deploying the policy on a real-world robot. However, a gap between simulation and reality generally persists, which endorses the aim to train robust policies already in simulation such that those can be transferred to a real robot at a high success rate. In this paper, we investigate history-dependent policies for drone control in the context of zero-shot transfer learning, where the training is conducted exclusively in simulation. We compare policies represented by feed-forward neural networks with recurrent neural networks and assess both performance and robustness on a real-world quadrotor. Furthermore, we study if an end-to-end learned representation can control a quadrotor based on raw onboard-sensor information only, rendering accurate state estimation from a Kalman filter obsolete. Our results show that recurrent control policies achieve similar performance and robustness as feed-forward policies when acting on state estimates. With raw sensory data, however, recurrent networks offer higher success rates for sim-to-real transfer than feed-forward networks. We also find that recurrent architectures are advantageous when system parameters such as latency are uncertain. Sven Gronauer, Daniel Stümke, Klaus Diepold |
IROS | 3 |
| 2023 | Offline Reinforcement Learning for Quadrotor Control: Overcoming the Ground EffectabstractApplying Reinforcement Learning to solve real-world optimization problems presents significant challenges because of the large amount of data normally required. A popular solution is to train the algorithms in a simulation and transfer the weights to the real system. However, sim-to-real approaches are prone to fail when the Reality Gap is too big, e.g. in robotic systems with complex and non-linear dynamics. In this work, we propose the use of Offline Reinforcement Learning as a viable alternative to sim-to-real policy transfer to address such instances. On the example of a small quadrotor, we show that the ground effect causes problems in an otherwise functioning zero-shot sim-to-real framework. Our sim-to-real experiments show that, even with the explicit modelling of the ground effect and the employing of popular transfer techniques, the trained policies fail to capture the physical nuances necessary to perform a real-world take-off maneuver. Contrariwise, we show that state-of-the-art Offline Reinforcement Learning algorithms represent a feasible, reliable and sample efficient alternative in this use case. Luca Sacchetto, Mathias Korte, Sven Gronauer, Matthias Kissel, Klaus Diepold |
IROS | 5 |
| 2022 | Deep Convolutional Neural Networks with Sequentially Semiseparable Weight MatricesabstractModern Convolutional Neural Networks (CNNs) comprise millions of parameters.Therefore, the use of these networks requires high computing and memory resources.We propose to reduce these resource requirements by using structured matrices.For that, we replace weight matrices of the fully connected classifier part of several pre-trained CNNs by Sequentially Semiseparable (SSS) Matrices.By that, the number of parameters in these layers can be reduced drastically, as well as the number of operations required for evaluating the layer.We show that the combination of approximating the original weight matrices with SSS matrices followed by gradient-descent based training leads to the best prediction results (compared to just approximating or training from scratch). Matthias Kissel, Klaus Diepold |
ESANN | 2 |
| 2022 | A Fuzzy-Cognitive-Maps Approach to Decision-Making in Medical EthicsabstractAlthough machine intelligence is increasingly employed in healthcare, the realm of decision-making in medical ethics remains largely unexplored from a technical perspective. We propose an approach based on fuzzy cognitive maps (FCMs), which builds on Beauchamp and Childress’ prima-facie principles. The FCM’s weights are optimized using a genetic algorithm to provide recommendations regarding the initiation, continuation, or withdrawal of medical treatment. The resulting model approximates the answers provided by our team of medical ethicists fairly well and offers a high degree of interpretability. Possible applications of such a system include informal guidance on medical ethics dilemmas as well as educational purposes. Alice Hein, Lukas J. Meier, Alena Buyx, Klaus Diepold |
FUZZ-IEEE | 4 |
| 2022 | A Comparison of Uncertainty Quantification Methods for Active Learning in Image ClassificationabstractThe adoption of machine learning (ML) technology in real-world settings like medical imaging is currently hampered by a lack of trust in ML models and a lack of labeled data. These two issues are currently addressed in parallel by two subdisciplines of ML: uncertainty quantification (UQ) is concerned with obtaining reliable estimates of a model's confidence in its outputs, and active learning (AL) deals with efficiently training models in low-data regimes. To date, the usefulness of the new methods emerging from the field of UQ for AL remains under-explored. We here take a step to address this by comparing seven different UQ methods on three image classification data sets. Our experiments confirm previous indications in the AL literature that the ranking of sampling strategies can vary greatly across models and data sets. We find that Concrete Dropout, Least Confidence, Smallest Margin, and Entropy sampling consistently outperform Random sampling across data sets, whereas Ensembles, Monte-Carlo Dropout, and Bayes-by-Backprop do not. We also observe that AL training stability is sensitive to data quality. Alice Hein, Stefan Röhrl, Thea Grobel, Manuel Lengl, Nawal Hafez, Martin Knopp, Christian Klenk, Dominik Heim, Oliver Hayden, Klaus Diepold |
IJCNN | 10 |
| 2022 | Convolutional Neural Networks with analytically determined FiltersabstractIn this paper, we propose a new training algorithm for Convolutional Neural Networks (CNNs) based on well-known training methods for neural networks with random weights. Our algorithm analytically determines the filters of the convolutional layers by solving a least squares problem using the Moore-Penrose generalized inverse. The resulting algorithm does not suffer from convergence issues and the training time is drastically reduced compared to traditional CNN training using gradient descent. We validate our algorithm with several standard datasets (MNIST, FashionMNIST and CIFAR10) and show that CNNs trained with our method outperform previous approaches with random or unsupervisedly learned filters in terms of test pre-diction accuracy. Moreover, our approach is up to 25 times faster than training CNNs with equivalent architecture using a gradient-descent based algorithm. Matthias Kissel, Klaus Diepold |
IJCNN | 2 |
| 2022 | Using Simulation Optimization to Improve Zero-shot Policy Transfer of QuadrotorsabstractIn this work, we propose a data-driven approach to optimize the parameters of a simulation such that control policies can be directly transferred from simulation to a real-world quadrotor. Our neural network-based policies take only onboard sensor data as input and run entirely on the embed-ded hardware. In real-world experiments, we compare low-level Pulse-Width Modulated control with higher-level control structures such as Attitude Rate and Attitude, which utilize Proportional-Integral-Derivative controllers to output motor commands. Our experiments show that low-level controllers trained with Reinforcement Learning require a more accurate simulation than higher-level control policies at the expense of being less robust towards parameter uncertainties. Sven Gronauer, Matthias Kissel, Luca Sacchetto, Mathias Korte, Klaus Diepold |
IROS | 5 |
| 2021 | The Successful Ingredients of Policy Gradient AlgorithmsabstractDespite the sublime success in recent years, the underlying mechanisms powering the advances of reinforcement learning are yet poorly understood. In this paper, we identify these mechanisms - which we call ingredients - in on-policy policy gradient methods and empirically determine their impact on the learning. To allow an equitable assessment, we conduct our experiments based on a unified and modular implementation. Our results underline the significance of recent algorithmic advances and demonstrate that reaching state-of-the-art performance may not need sophisticated algorithms but can also be accomplished by the combination of a few simple ingredients. Sven Gronauer, Martin Gottwald, Klaus Diepold |
IJCAI | 3 |
| 2020 | Neural Network Training with Safe Regularization in the Null Space of Batch Activations
Matthias Kissel, Martin Gottwald, Klaus Diepold |
ICANN (2) | 3 |
| 2020 | Dual-Mode Training with Style Control and Quality Enhancement for Road Image Domain AdaptationabstractDealing properly with different viewing conditions remains a key challenge for computer vision in autonomous driving. Domain adaptation has opened new possibilities for data augmentation, translating arbitrary road scene images into different environmental conditions. Although multimodal concepts have demonstrated the capability to separate content and style, we find that existing methods fail to reproduce scenes in the exact appearance given by a reference image. In this paper, we address the aforementioned problem by introducing a style alignment loss between output and reference image. We integrate this concept into a multimodal unsupervised image-to-image translation model with a novel dual-mode training process and additional adversarial losses. Focusing on road scene images, we evaluate our model in various aspects including visual quality and feature matching. Our experiments reveal that we are able to significantly improve both style alignment and image quality in different viewing conditions. Adapting concepts from neural style transfer, our new training approach allows to control the output of multimodal domain adaptation, making it possible to generate arbitrary scenes and viewing conditions for data augmentation. Moritz Venator, Fengyi Shen, Selcuk Aklanoglu, Erich Bruns, Klaus Diepold, Andreas K. Maier |
WACV | 5 |
| 2019 | Decision Process of Autonomous Drones for Environmental MonitoringabstractEnvironmental monitoring has a key role to reduce the human effect on nature and wildlife. Due to intense tracking and observation tasks, environmental monitoring is an expensive solution. Autonomous observation is an open discussion to raise the efficiency of environmental monitoring and reduce the cost of operation. Unmanned aerial vehicles (UAVs) are possible candidates with their proven success in monitoring and tracking. Thus, we are offering a decision process for the autonomy of drones in monitoring tasks. Our system is capable to fly autonomously, e.g., covering a given area, and able to perform certain tasks, e.g., identifying bottles. Our simulation results prove that autonomous drones can be used for a large variety of environmental monitoring tasks. Ömür Yildirim, Klaus Diepold, Revna Acar Vural |
INISTA | 2 |
| 2019 | Sobolev Training with Approximated Derivatives for Black-Box Function Regression with Neural Networks
Matthias Kissel, Klaus Diepold |
ECML/PKDD (2) | 2 |
| 2018 | Deep Reinforcement Learning for Formation ControlabstractContinuing our work on using reinforcement learning for formation control, we present an end-to-end deep learning system which uses only camera images to learn to control the individual system's correct position within the formation. Mnih et al. created AIs playing video games utilizing the same visual input as a human player by employing convolutional neural networks for automatic feature extraction on images. This published work inspired us to employ a similar approach for processing the camera images and controlling the robot. We repeat the same experiment with two completely different camera positions. The results for both positions are very similar and such demonstrate the flexibility of the presented approach. Can Aykin, Martin Knopp, Klaus Diepold |
RO-MAN | 3 |
| 2015 | Crowdsourcing vs. laboratory experiments - QoE evaluation of binaural playback in a teleconference scenario
Thomas Volk, Christian Keimel, Michael Moosmeier, Klaus Diepold |
Comput. Networks | 4 |
| 2014 | Path-finding using reinforcement learning and affective statesabstractDuring decision making and acting in the environment humans appraise decisions and observations with feelings and emotions. In this paper we propose a framework to incorporate an emotional model into the decision making process of a machine learning agent. We use a hierarchical structure to combine reinforcement learning with a dimensional emotional model. The dimensional model calculates two dimensions representing the actual affective state of the autonomous agent. For the evaluation of this combination, we use a reinforcement learning experiment (called Dyna Maze) in which, the agent has to find an optimal path through a maze. Our first results show that the agent is able to appraise the situation in terms of emotions and react according to them. Johannes Feldmaier, Klaus Diepold |
RO-MAN | 2 |
| 2014 | Best Practices for QoE Crowdtesting: QoE Assessment With CrowdsourcingabstractQuality of Experience (QoE) in multimedia applications is closely linked to the end users' perception and therefore its assessment requires subjective user studies in order to evaluate the degree of delight or annoyance as experienced by the users. QoE crowdtesting refers to QoE assessment using crowdsourcing, where anonymous test subjects conduct subjective tests remotely in their preferred environment. The advantages of QoE crowdtesting lie not only in the reduced time and costs for the tests, but also in a large and diverse panel of international, geographically distributed users in realistic user settings. However, conceptual and technical challenges emerge due to the remote test settings. Key issues arising from QoE crowdtesting include the reliability of user ratings, the influence of incentives, payment schemes and the unknown environmental context of the tests on the results. In order to counter these issues, strategies and methods need to be developed, included in the test design, and also implemented in the actual test campaign, while statistical methods are required to identify reliable user ratings and to ensure high data quality. This contribution therefore provides a collection of best practices addressing these issues based on our experience gained in a large set of conducted QoE crowdtesting studies. The focus of this article is in particular on the issue of reliability and we use video quality assessment as an example for the proposed best practices, showing that our recommended two-stage QoE crowdtesting design leads to more reliable results. Tobias Hoßfeld, Christian Keimel, Matthias Hirth, Bruno Gardlo, Julian Habigt, Klaus Diepold, Phuoc Tran-Gia |
IEEE Trans. Multim. | 6 |
| 2013 | Image completion for view synthesis using Markov random fields and efficient belief propagationabstractView synthesis is a process for generating novel views from a scene which has been recorded with a 3-D camera setup. It has important applications in 3-D post-production and 2-D to 3-D conversion. However, a central problem in the generation of novel views lies in the handling of disocclusions. Background content, which was occluded in the original view, may become unveiled in the synthesized view. This leads to missing information in the generated view which has to be filled in a visually plausible manner. We present an inpainting algorithm for disocclusion filling in synthesized views based on Markov random fields and efficient belief propagation. We compare the result to two state-of-the-art algorithms and demonstrate a significant improvement in image quality. Julian Habigt, Klaus Diepold |
ICIP | 2 |
| 2013 | Length-independent refinement of video quality metrics based on multiway data analysisabstractIn previous publications it has been shown that no-reference video quality metrics based on a data analysis approach rather than on modeling the human visual system lead to very promising results and outperform many well-known full-reference metrics. Furthermore, the results improve when taking the temporal structure of the video sequence into account by using multiway analysis methods. This contribution shows a way of refining these multiway quality metrics in order to make them more suitable for real-life applications and maintaining the performance at the same time. Additionally, our results confirm the validity of H.264/AVC bitstream no-reference quality metrics using multiway PLSR by evaluating this concept on an additional dataset. Clemens Horch, Christian Keimel, Julian Habigt, Klaus Diepold |
ICIP | 4 |
| 2013 | Emotional evaluation of bandit problemsabstractIn this paper, we discuss an approach to evaluate decisions made during a multi-armed bandit learning experiment. Usually, the results of machine learning algorithms applied on multi-armed bandit scenarios are rated in terms of earned reward and optimal decisions taken. These criteria are valuable for objective comparison in finite experiments. But learning algorithms used in real scenarios, for example in robotics, need to have instantaneous criteria to evaluate their actual decisions taken. To overcome this problem, in our approach each decision updates the Zürich model which emulates the human sense of feeling secure and aroused. Combining these two feelings results in an emotional evaluation of decision policies and could be used to model the emotional state of an intelligent agent. Johannes Feldmaier, Klaus Diepold |
RO-MAN | 2 |
| 2013 | Analysis Operator Learning and its Application to Image ReconstructionabstractExploiting a priori known structural information lies at the core of many image reconstruction methods that can be stated as inverse problems. The synthesis model, which assumes that images can be decomposed into a linear combination of very few atoms of some dictionary, is now a well established tool for the design of image reconstruction algorithms. An interesting alternative is the analysis model, where the signal is multiplied by an analysis operator and the outcome is assumed to be sparse. This approach has only recently gained increasing interest. The quality of reconstruction methods based on an analysis model severely depends on the right choice of the suitable operator. In this paper, we present an algorithm for learning an analysis operator from training images. Our method is based on l(p)-norm minimization on the set of full rank matrices with normalized columns. We carefully introduce the employed conjugate gradient method on manifolds, and explain the underlying geometry of the constraints. Moreover, we compare our approach to state-of-the-art methods for image denoising, inpainting, and single image super-resolution. Our numerical results show competitive performance of our general approach in all presented applications compared to the specialized state-of-the-art techniques. Simon Hawe, Martin Kleinsteuber, Klaus Diepold |
IEEE Trans. Image Process. | 3 |
| 2012 | Recurrent Takagi-Sugeno fuzzy interpolation for switched linear systems and hybrid automataabstractA novel perspective of fuzzy control is introduced, by combining a continuous Takagi-Sugeno (TSFS) with a discrete-time recurrent fuzzy system (RFS). The developed hybrid dynamic recurrent Takagi-Sugeno fuzzy approach enables a recurrent rule base, which leads to a dynamical interpolation law between linear subsystems. The formalism is applicable to switched systems and hybrid automaton models. The stability concerning switched systems is ensured by a common quadratic Lyapunov function. Additionally, two multiple Lyapunov function-based stability relaxation conditions are shown. For the hybrid automaton case, practical stability is analyzed. The performance of the approach is validated by simulating a car distance control system (hybrid automaton) and an experimental application to an inverted pendulum (switched control). Klaus Diepold, Sebastian J. Pieczona |
FUZZ-IEEE | 1 |
| 2012 | Cartoon-like image reconstruction via constrained ℓp-minimizationabstractThis paper considers the problem of reconstructing images from only a few measurements. A method is proposed that is based on the theory of Compressive Sensing. We introduce a new prior that combines an ℓp-pseudo-norm approximation of the image gradient and the bounded range of the original signal. Ultimately, this leads to a reconstruction algorithm that works particularly well for Cartoon-like images that commonly occur in medical imagery. The arising optimization task is solved by a Conjugate Gradient method that is capable of dealing with large scale problems and easily adapts to extensions of the prior. To overcome the none differentiability of the ℓp-pseudo-norm we employ a Huber-loss term like approximation together with a continuation of the smoothing parameter. Numerical results and a comparison with the state-of-the-art methods show the effectiveness of the proposed algorithm. Simon Hawe, Martin Kleinsteuber, Klaus Diepold |
ICASSP | 3 |
| 2012 | Beyond classical teleoperation: Assistance, cooperation, data reduction, and spatial audioabstractIn this video we present a teleoperation system which is capable of solving complex tasks in human-sized wide area environments. The system consists of two mobile teleoperators controlled by two operators, and offers haptic, visual, and auditory feedback. The task examined here, consists of repairing a robot by removing a computer and replacing a defective hard-drive. To cope with the complexity of such a task, we go beyond classical teleoperation by integrating several advanced software algorithms into the system. Thomas Schauss, Carolina Passenberg, Nikolay Stefanov, Daniela Feth, Iason Vittorias, Angelika Peer, Sandra Hirche, Martin Buss, Martin Rothbucher, Klaus Diepold, Julius Kammerl, Eckehard G. Steinbach |
ICRA | 10 |
| 2012 | QualityCrowd - A framework for crowd-based quality evaluationabstractVideo quality assessment with subjective testing is both time consuming and expensive. An interesting new approach to traditional testing is the so-called crowdsourcing, moving the testing effort into the internet. We therefore propose in this contribution the QualityCrowd framework to effortlessly perform subjective quality assessment with crowdsourcing. QualityCrowd allows codec independent quality assessment with a simple web interface, usable with common web browsers. We compared the results from an online subjective test using this framework with the results from a test in a standardized environment. This comparison shows that QualityCrowd delivers equivalent results within the acceptable inter-lab correlation. While we only consider video quality in this contribution, QualityCrowd can also be used for multimodal quality assessment. Christian Keimel, Julian Habigt, Clemens Horch, Klaus Diepold |
PCS | 4 |
| 2012 | HRTF-based localization and separation of multiple sound sourcesabstractThe human auditory system excels at pinpointing and distinguishing multiple sound sources in noisy and reverberant environments. Mobile robotic platforms implement such capabilities with varying success, classically solving localization and separation independently. This paper presents an algorithm utilizing Head-Related Transfer Function (HRTF) based localization to aid the task of separation. HRTFs for robotic binaural hearing represent the digital emulation of a human's innate direction-dependent filtering for solving the localization problem in a compact and robust manner. The overall result of the presented algorithm for robotic binaural hearing is an HRTF-based localization and separation system, capable of dynamically and intelligently processing simultaneously active sound sources. Martin Rothbucher, Marko Durkovic, Tim Habigt, Hao Shen 0002, Klaus Diepold |
RO-MAN | 5 |
| 2012 | Camera-Pose Estimation via Projective Newton Optimization on the ManifoldabstractDetermining the pose of a moving camera is an important task in computer vision. In this paper, we derive a projective Newton algorithm on the manifold to refine the pose estimate of a camera. The main idea is to benefit from the fact that the 3-D rigid motion is described by the special Euclidean group, which is a Riemannian manifold. The latter is equipped with a tangent space defined by the corresponding Lie algebra. This enables us to compute the optimization direction, i.e., the gradient and the Hessian, at each iteration of the projective Newton scheme on the tangent space of the manifold. Then, the motion is updated by projecting back the variables on the manifold itself. We also derive another version of the algorithm that employs homeomorphic parameterization to the special Euclidean group. We test the algorithm on several simulated and real image data sets. Compared with the standard Newton minimization scheme, we are now able to obtain the full numerical formula of the Hessian with a 60% decrease in computational complexity. Compared with Levenberg-Marquardt, the results obtained are more accurate while having a rather similar complexity. Michel Sarkis, Klaus Diepold |
IEEE Trans. Image Process. | 2 |
| 2011 | Dense disparity maps from sparse disparity measurementsabstractIn this work we propose a method for estimating disparity maps from very few measurements. Based on the theory of Compressive Sensing, our algorithm accurately reconstructs disparity maps only using about 5% of the entire map. We propose a conjugate subgradient method for the arising optimization problem that is applicable to large scale systems and recovers the disparity map efficiently. Experiments are provided that show the effectiveness of the proposed approach and robust behavior under noisy conditions. Simon Hawe, Martin Kleinsteuber, Klaus Diepold |
ICCV | 3 |
| 2011 | No-reference video quality metric for HDTV based on H.264/AVC bitstream featuresabstractNo-reference video quality metrics are becoming ever more popular, as they are more useful in real-life applications compared to full-reference metrics. Many proposed metrics extract features related to human perception from the individual video frames. Hence the video sequences have to be decoded first, before the metrics can be applied. In order to avoid decoding just for quality estimation, we therefore present in this contribution a no-reference metric for HDTV that uses features directly extracted from the H.264/AVC bitstream. We combine these features with the results from subjective tests using a data analysis approach with partial least squares regression to gain a prediction model for the visual quality. For verification, we performed a cross validation. Our results show that the proposed no-reference metric outperforms other metrics and delivers a correlation between the quality prediction and the actual quality of 0.93. Christian Keimel, Manuel Klimpke, Julian Habigt, Klaus Diepold |
ICIP | 4 |
| 2011 | Artificial Cognition in Production SystemsabstractToday's manufacturing and assembly systems have to be flexible to adapt quickly to an increasing number and variety of products, and changing market volumes. To manage these dynamics, several production concepts (e.g., flexible, reconfigurable, changeable or autonomous manufacturing and assembly systems) were proposed and partly realized in the past years. This paper presents the general principles of autonomy and the proposed concepts, methods and technologies to realize cognitive planning, cognitive control and cognitive operation of production systems. Starting with an introduction on the historical context of different paradigms of production (e.g., evolution of production and planning systems), different approaches for the design, planning, and operation of production systems are lined out and future trends towards fully autonomous components of an production system as well as autonomous parts and products are discussed. In flexible production systems with manual and automatic assembly tasks, human-robot cooperation is an opportunity for an ergonomic and economic manufacturing system especially for low lot sizes. The state-of-the-art and a cognitive approach in this area are outlined. Furthermore, introducing self-optimizing and self-learning control systems is a crucial factor for cognitive systems. This principles are demonstrated by a quality assurance and process control in laser welding that is used to perform improved quality monitoring. Finally, as the integration of human workers into the workflow of a production system is of the highest priority for an efficient production, worker guidance systems for manual assembly with environmentally and situationally dependent triggered paths on state-based graphs are described in this paper. Alexander Bannat, Thibault Bautze, Michael Beetz, Jürgen Blume, Klaus Diepold, Christoph Ertelt, Florian Geiger, Thomas Gmeiner, Tobias Gyger, Alois C. Knoll, Christian Lau, Claus Lenz, Martin Ostgathe, Gunther Reinhart, Wolfgang Rösel, Thomas Rühr, Anna Schubö, Kristina Shea, Ingo Stork, Sonja Stork, William Tekouo, Frank Wallhoff, Mathey Wiesbeck, Michael F. Zäh |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2010 | Improving the prediction accuracy of video quality metricsabstractTo improve the prediction accuracy of visual quality metrics for video we propose two simple steps: temporal pooling in order to gain a set of parameters from one measured feature and a correction step using videos of known visual quality. We demonstrate this approach on the well known PSNR. Firstly, we achieve a more accurate quality prediction by replacing the mean luma PSNR by alternative PSNR-based parameters. Secondly, we exploit the almost linear relationship between the output of a quality metric and the subjectively perceived visual quality for individual video sequences. We do this by estimating the parameters of this linear relationship with the help of additionally generated videos of known visual quality. Moreover, we show that this is also true for very different coding technologies. Also we used cross validation to verify our results. Combining these two steps, we achieve for a set of four different high definition videos an increase of the Pearson correlation coefficient from 0.69 to 0.88 for PSNR, outperforming other, more sophisticated full-reference video quality metrics. Christian Keimel, Tobias Oelbaum, Klaus Diepold |
ICASSP | 3 |
| 2010 | High-fidelity telepresence and teleactionabstractThe collaborative research center SFB453 (www.sfb453.de) aims to realize high-fidelity telepresence and teleaction systems. Telepresence and teleaction systems extend the human workspace to remote locations in order to overcome barriers like distance, scaling, danger or the human skin. Using a human-system interface the human operator controls a remotely located teleoperator. Multi-modal feedback in form of visual, auditory, and haptic data is used to increase the feeling of telepresence. Different application areas including minimally invasive surgery, on-orbit servicing, microassembly as well as tele-manufacturing and tele-maintenance are targeted. Robert Bauernschmitt, Martin Buss, Barbara Deml, Klaus Diepold, Berthold Färber, Georg Färber, Ulrich Hagn, Gerd Hirzinger, Sandra Hirche, Alois C. Knoll, Hermann J. Müller, Tobias Ortmaier, Angelika Peer, Michael Popp, Carsten Preusche, Gunther Reinhart, Zhuanghua Shi, Eckehard G. Steinbach, Heinz Ulbrich, Ulrich Walter, Michael F. Zäh |
ICRA | 4 |
| 2010 | Visual quality of current coding technologies at high definition IPTV bitratesabstractHigh definition video over IP based networks (IPTV) has become a mainstay in today's consumer environment. In most applications, encoders conforming to the H.264/AVC standard are used. But even within one standard, often a wide range of coding tools are available that can deliver a vastly different visual quality. Therefore we evaluate in this contribution different coding technologies, using different encoder settings of H.264/AVC, but also a completely different encoder like Dirac. We cover a wide range of different bitrates from ADSL to VDSL and different content, with low and high demand on the encoders. As PSNR is not well suited to describe the perceived visual quality, we conducted extensive subject tests to determine the visual quality. Our results show that for currently common bitrates, the visual quality can be more than doubled, if the same coding technology, but different coding tools are used. Christian Keimel, Julian Habigt, Tim Habigt, Martin Rothbucher, Klaus Diepold |
MMSP | 5 |
| 2010 | Integrating a HRTF-based sound synthesis system into MumbleabstractThis paper describes an integration of a Head Related Transfer Function (HRTF)-based 3D sound convolution engine into the open-source VoIP conferencing software Mumble. Our system allows to virtually place audio contributions of conference participants to different positions around a listener, which helps to overcome the problem of identifying active speakers in an audio conference. Furthermore, using HRTFs to generate 3D sound in virtual 3D space, the listener is able to make use of the cocktail party effect in order to differentiate between several simultaneously active speakers. As a result intelligibility of communication is increased. Martin Rothbucher, Tim Habigt, Johannes Feldmaier, Klaus Diepold |
MMSP | 4 |
| 2010 | Improving the visual quality of AVC/H.264 by combining it with content adaptive depth map compressionabstractThe future of video coding for 3DTV lies in the combination of depth maps and corresponding textures. Most current video coding standards, however, are only optimized for visual quality and are not able to efficiently compress depth maps. We present in this work a content adaptive depth map meshing with tritree and entropy encoding for 3D videos. We show that this approach outperforms the intra frame prediction of AVC/H.264 for the coding of depth maps of still images. We also demonstrate by combining AVC/H.264 with our algorithm that we are able to increase the visual quality of the encoded texture on average by 6 dB. This work is currently limited to still images but an extension to intra coding of 3D video is straightforward. Christian Keimel, Klaus Diepold, Michel Sarkis |
PCS | 2 |
| 2010 | Transient probabilistic recurrent fuzzy systemsabstractA probabilistic-based extension of recurrent fuzzy systems is presented and exemplarily applied to modeling and control of systems in different domains. The system's core-dynamic is described by a recurrent fuzzy system, while further known influencing features are summarized via probability theory using a stochastic automaton. The appropriate conditional probabilities are used to adapt the dynamics of the recurrent fuzzy system depending on its state variables. By allowing transient conditional probabilities, a time-variance is simultaneously achieved. Thus, the developed transient probabilistic recurrent fuzzy system (TP-RFS) is able to handle two kinds of uncertain information (vague and stochastic) and allows slight as well as drastic adaptations of the original recurrent fuzzy system's dynamics. Successful applications of the proposed TP-RFS for modeling different dynamics of an ecological system and for controlling the speed signaling on a highway are shown. Klaus Diepold, Boris Lohmann |
SMC | 1 |
| 2009 | Fast Depth Map Compression and Meshing with Compressed Tritree
Michel Sarkis, Waqar Zia, Klaus Diepold |
ACCV (2) | 3 |
| 2009 | MutanT: A modular and generic tool for multi-sensor data processing
Simon Hawe, Ulrich Kirchmaier, Klaus Diepold |
FUSION | 3 |
| 2009 | No-reference video quality evaluation for high-definition videoabstractA no-reference video quality metric for High-Definition video is introduced. This metric evaluates a set of simple features such as blocking or blurring, and combines those features into one parameter representing visual quality. While only comparably few base feature measurements are used, additional parameters are gained by evaluating changes for these measurements over time and using additional temporal pooling methods. To take into account the different characteristics of different video sequences, the gained quality value is corrected using a low quality version of the received video. The metric is verified using data from accurate subjective tests, and special care was taken to separate data used for calibration and verification. The proposed no-reference quality metric delivers a prediction accuracy of 0.86 when compared to subjective tests, and significantly outperforms PSNR as a quality predictor. Christian Keimel, Tobias Oelbaum, Klaus Diepold |
ICASSP | 3 |
| 2009 | Cryptanalysis of Substitution Cipher Chaining Mode (SCC)abstractIn this paper, we cryptanalyze the substitution cipher chaining mode (SCC-128), which uses three keys. The first key is the encryption key, which we were able to recover with about 240cipher executions and 5 times 28chosen plaintexts. The second key is responsible of generation two layers of masks, we recovered the first layer with 213chosen plaintext and 221cipher executions and our attack to recover the second layer costs only one known sector plaintext and 64 cipher executions. The third key is used to generate the encrypted sector ID, we were able to recover the encrypted sector ID of a sector with I known plaintext and 2 cipher executions for each sector. Mohamed Abo El-Fotouh, Klaus Diepold |
ICC | 2 |
| 2009 | The impact of nonlinear filtering and confidence information on optical flow estimation in a Lucas & Kanade frameworkabstractDetermining optical flow has been a wide field of research for more than 20 years now that has not been solved satisfactorily yet. In this work, we study the influence of a nonlinear smoothing process based on bilateral filtering on a Lucas & Kanade framework for the estimation of optical flow between two image frames. Different confidence measures are used to improve the computation process and detect occlusion and innovation phenomena, explicitly handling discontinuous flow fields. By means of simulations we report that the accuracy can be increased significantly, making this approach interesting for further investigations. Michael Heindlmaier, Lang Yu, Klaus Diepold |
ICIP | 3 |
| 2009 | Depth map compression via compressed sensingabstractWe propose in this paper a new scheme based on compressed sensing to compress a depth map. We first subsample the entity in the frequency domain to take advantage of its compressibility. We then derive a reconstruction scheme to recover the original map from the subsamples using a non-linear conjugate gradient minimization scheme. We preserve the discontinuities of the depth map at the edges and ensure its smoothness elsewhere by incorporating the Total Variation constraint in the minimization. The results we obtained on various test depth maps show that the proposed method leads to lower error rate at high compression ratio when compared to standard image compression techniques like JPEG and JPEG 2000. Michel Sarkis, Klaus Diepold |
ICIP | 2 |
| 2009 | A cognitive system for autonomous robotic weldingabstractCurrently, there is a high demand for autonomous industrial production systems. This paper outlines the development of a cognitive system for autonomous robotic welding. This system is based on dimensionality reduction techniques and Support Vector Machines, allowing the system to learn to separate between acceptable and unacceptable welding results within one batch, and to transfer this ability to a batch with different workpiece properties. It does not aim at a complete and general relationship between all process variables and result quantities, since it has been demonstrated that this is not necessary to reduce significantly the costs of calibrating the welding system. The main objective is to examine a cognitive system that stabilizes robotic welding processes by learning how to improve at least one process steering variable. In order to evaluate and improve the cognitive system, an extensive experimental setup is realized and described. The ability to learn and autonomously adapt to changes in workpiece properties allows the system to reduce the time an expert needs, and relaxes the requirements with respect to workpiece tolerances. Georg Schroth, Ingo Stork, Klaus Diepold |
IROS | 3 |
| 2009 | Optimization of video coding for telepresence applicationsabstractIn several telepresence applications the transmitting end is resource constrained, hence multi-view video has to be source coded and transmitted to a location where stereo matching is performed. The effect of video source coding on stereo matching techniques is presently not well understood. In this paper, we present an analysis which quantitatively demonstrates the impact of standard video coding techniques on the quality of well-known stereo matching algorithms. We propose a quantitative evaluation methodology and framework. Evaluation is also presented for the state-of-the-art Annex H of H.264/AVC (multi-view coding). The results shed a light on the tradeoff between video coding parameters and the quality of dense stereo matching. This enables us to select suitable operating regions for both stereo matching and video coding, for a diverse range of applications. In addition, various system configurations with varying performance-complexity tradeoff are also compared, giving insight on selecting the configuration suitable for a variety of telepresence applications. Waqar Zia, Klaus Diepold, Michel Sarkis |
WACV | 2 |
| 2009 | Calibrating an Automatic Zoom Camera With Moving Least SquaresabstractThe application of zoom camera lenses in machine vision has gained a lot of attention lately. The main difficulty in their employment lies in the accurate estimation of their intrinsic parameters. In this paper, we propose novel approaches to determine these parameters by estimating continuous models of their variations as the focus and the zoom change. The first method is based on the moving least squares (MLS) multiple regression scheme which determines from a predefined number of samples, the complete model of the intrinsic parameters. MLS fits a polynomial function at each focus and zoom setting by using the measured neighboring points. In order to reduce the computational complexity of MLS, we propose another algorithm in which the MLS generated curves are clustered. Then, each cluster is modeled with a single polynomial function. This decreases the complexity of computations for the applications where delay is critical, e.g., telepresence, to the evaluation of simple polynomials. Compared to previous techniques, the proposed algorithms lead to a noticeable increase in the estimation accuracy of the intrinsic parameters. In addition, they are able to generate accurate models of these parameters with only a few measurement points. Michel Sarkis, Christian T. Senft, Klaus Diepold |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2009 | Content Adaptive Mesh Representation of Images Using Binary Space PartitionsabstractThe interest in content adaptive mesh generation of images has been arising lately due to its wide area of applications in image processing. The major issue is to represent an image with a low number of pixels while preserving its content. These pixels or the nonuniform samples are then used to generate a mesh that approximates the corresponding image. This work presents a novel method based on Binary Space Partitions in combination with three clustering schemes to approximate an image with a mesh. The algorithm has the ability to simultaneously reduce the number of pixels and generate the mesh approximation. The idea is to assume each triangle of the mesh as a plane. Consequently, it will be possible to reconstruct the inlying pixels with planar equations defined from the three nodes of each triangle. If a triangle's equation does not have the ability to reconstruct the pixels lying within up to a predefined error, it is split into two new triangles. Tested on several real images, the proposed method leads to reduced size meshes in a fast manner while retaining the visual quality of the reconstructed images. In addition, it is parallelizable due to the property of Binary Space Partitions which facilitates its application in real-time scenarios. Michel Sarkis, Klaus Diepold |
IEEE Trans. Image Process. | 2 |
| 2008 | Dynamic Substitution ModelabstractIn this paper, we present the Dynamic Substitution Model (DSM) and its variant the Static Substitution Model (SSM). In DSM and SSM, the secret encryption key is divided into a primary key and a secondary key. DSM is a model that allows any block cipher to accept a variable length secondary key, this is achieved by substituting some bits of the cipher's expanded key with the secondary key. SSM is a variant of DSM, where the secondary key length and the positions of the replaced bits of the subkeys are determined in the design time. We used the Advanced Encryption Standard (AES) to demonstrate the usage of DSM and SSM models. Mohamed Abo El-Fotouh, Klaus Diepold |
IAS | 2 |
| 2008 | A New Narrow Block Mode of Operations for Disk EncryptionabstractIn this paper, we present a new narrow block mode of operation, the masked code book (MCB), that can be efficiently deployed in disk encryption applications. MCB is characterized by its high-speed in comparison to current state of the art narrow block modes of operation. It is about 80% faster than XTS (when AES with 128-bits key is the underlying cipher). Mohamed Abo El-Fotouh, Klaus Diepold |
IAS | 2 |
| 2008 | The Analysis of Windows Vista Disk Encryption Algorithm
Mohamed Abo El-Fotouh, Klaus Diepold |
DBSec | 2 |
| 2008 | Three dimensional object tracking based on audiovisual fusion using Particle Swarm Optimization
Fakheredine Keyrouz, Ulrich Kirchmaier, Klaus Diepold |
FUSION | 3 |
| 2008 | Efficient content adaptive mesh representation of an image using binary space partitions and singular value decompositionabstractContent adaptive mesh generation is an important research area with many applications in image processing and computer vision. The main issue is to represent an image with the pixels that preserve most of the amount of its information. The obtained pixels are then used to generate a mesh that approximates the original image. This work presents a novel iterative method that simultaneously reduces the number of the pixels and generates the mesh approximation of an image. The main idea is to incorporate binary space partitions along with singular value decomposition to cluster the pixels into planes and thus the nodes of the mesh are nothing but the pixels that define each plane. Compared to previous techniques, the proposed method leads to a 30% reduction in the size of the approximating mesh. In addition, the method minimizes the artifacts obtained from the reconstruction of the original image from the approximating mesh. Michel Sarkis, Oliver Lorscheider, Klaus Diepold |
ICASSP | 3 |
| 2008 | Sparse stereo matching using belief propagationabstractDepth from stereo is an important research field in computer vision due to the wide range of its applications. In this work, we present a stereo matching algorithm based on belief propagation (BP). The algorithm is designed to work on sparse images originating from image content adaptive mesh representation techniques. There, an image is approximated with a mesh. The nodes of the mesh are the non-uniform samples which are the ones that form the sparse image. The key issue in the proposed method is to formulate BP such that it matches a sparse left stereo image with a dense right image to obtain a sparse depth map. Moreover, we propose a simple method that recovers the dense disparity map of the scene from the sparse one using the approximating mesh of the image. The results obtained show that the proposed method leads to an average of 40% improvement in the quality of depth maps when compared to existing sparse stereo matching techniques. Michel Sarkis, Klaus Diepold |
ICIP | 2 |
| 2008 | Multi-modal multi-user telepresence and teleaction systemabstractThe video shows a rich multi-modal multi-user telepresence system, which was developed within the SFB453 funded by the German Research Foundation (www.sfb453.de). As a complex application scenario, the remote repairing of a broken pipe is presented in this paper. The system basically consists of two operator-teleoperator- pairs. While one of the operators interacts with a stationary human- system-interface, the other operator uses a mobile one. Both systems provide visual, auditory, and haptic feedback and enable to control the motion of head, arms, and grippers as well as the locomotion of the corresponding teleoperator. Martin Buss, Angelika Peer, Thomas Schauss, Nikolay Stefanov, Ulrich Unterhinninghofen, Stephan Behrendt, Georg Färber, Jan Leupold, Klaus Diepold, Fakheredine Keyrouz, Michel Sarkis, Peter Hinterseer, Eckehard G. Steinbach, Berthold Färber, Helena Pongrac |
IROS | 9 |
| 2008 | A Fast Encryption Scheme for Networks Applications
Mohamed Abo El-Fotouh, Klaus Diepold |
SECRYPT | 2 |
| 2008 | The Substitution Cipher Chaining Mode
Mohamed Abo El-Fotouh, Klaus Diepold |
SECRYPT | 2 |
| 2008 | A novel biologically inspired neural network solution for robotic 3D sound source sensing
Fakheredine Keyrouz, Klaus Diepold |
Soft Comput. | 2 |
| 2007 | Fast Adaptive Graph-Cuts Based Stereo Matching
Michel Sarkis, Nikolas Dörfler, Klaus Diepold |
ACIVS | 3 |
| 2007 | High performance 3D sound localization for surveillance applicationsabstractOne of the key features of the human auditory system, is its nearly constant omni-directional sensitivity, e.g., the system reacts to alerting signals coming from a direction away from the sight of focused visual attention. In many surveillance situations where visual attention completely fails since the robot cameras have no direct line of sight with the sound sources, the ability to estimate the direction of the sources of danger relying on sound becomes extremely important. We present in this paper a novel method for sound localization in azimuth and elevation based on a humanoid head. The method was tested in simulations as well as in a real reverberant environment. Compared to state-of-the-art localization techniques the method is able to localize with high accuracy 3D sound sources even in the presence of reflections and high distortion. Fakheredine Keyrouz, Klaus Diepold, Shady Keyrouz |
AVSS | 2 |
| 2007 | A Fast and Robust Solution to the Five-Pint Relative Pose Problem using Gauss-Newton Optimization on a ManifoldabstractExtracting the motion parameters of a moving camera is an important issue in computer vision. This is due to the need of numerous emerging applications like telepresence and robot navigation. The key issue is to determine a robust estimate of the (3×3) essential matrix with its five degrees of freedom. In this work, a robust technique to compute the essential matrix is suggested under the assumption that the images are calibrated. The algorithm is a combination of the five-point relative pose problem using an optimization technique on a manifold, with the random sample consensus. The results show that the proposed method delivers faster and more accurate results than the standard techniques. Michel Sarkis, Klaus Diepold, Knut Hüper |
ICASSP (1) | 2 |
| 2007 | A Novel Technique to Model the Variation of the Intrinsic Parameters of an Automatic Zoom Camera using Adaptive Delaunay Meshes Over Moving Least-Squares SurfacesabstractThe accuracy of computer vision systems is highly dependent on the correct estimates of the camera intrinsic parameters. This accuracy is important in numerous applications like telepresence and robot navigation. In this work, a novel technique is proposed to model the variation of the camera's intrinsic parameters as a function of the focus and the zoom. The proposed method computes the complete surfaces of the intrinsic parameters from a predefined number of focus/zoom measurements using a moving least-squares (MLS) regression technique. Then, it approximates the generated MLS surfaces by employing adaptive Delaunay meshes. Compared to a previous technique using bivariate polynomial functions, the new method results in a 94% enhancement of the mean estimation error. In addition, the new method leads to the same accuracy of the results as compared to a previous version of the MLS technique while requiring a less amount of computations. Michel Sarkis, Christian T. Senft, Klaus Diepold |
ICIP (5) | 3 |
| 2007 | Complexity Constrained Robust Video Transmission for Hand-Held DevicesabstractRobust video conversational applications for hand-held devices come with numerous challenges, e.g. real-time processing, complexity constrained devices and small end-to-end delays, etc. Transmission losses of compressed video data result in spatio-temporal error propagation in the decoded video sequence. To ensure some QoS, the video codec has to be well tuned to combat the degradation resulting from losses. Several feedback based error mitigation technique are assessed in this work. The proposed error robustness technique based on reference picture selection (RPS) and error tracking enhances the overall performance of the target system by more than 4 dB for moderate radio link control (RLC) PDU loss rates of 1.5%. This enhancement is achieved without any additional computational complexity. Waqar Zia, Klaus Diepold, Thomas Stockhammer |
ICIP (4) | 2 |
| 2006 | A New Method for Binaural 3-D Localization Based on HrtfsabstractA modern technique for robotic sound source detection using a dataset Head-Related Transfer Functions (HRTFs) is presented. To ensure fast detection, the HRTFs are reduced using three different techniques, namely; Diffuse-Field Equalization, Balanced Model Truncation, and Principle Component Analysis. A new criterion introduced is to be satisfied by a set of output signals from the microphones of a dummy robot head. This criterion is then used to find the sound source location in accordance with the reduced HRTF datasets. The suggested method is verified through simulated examples and further tested in a household environment. This novel technique provides estimates of azimuth and elevation angles in free space by using only two microphones. It also uses a simple algorithm compared to the more complicated algorithms used in similar localization processes. Fakheredine Keyrouz, Youssef Naous, Klaus Diepold |
ICASSP (5) | 3 |
| 2006 | Automatic Model-Order Selection for PCAabstractDetermining the model-order of a given data set is an important task in signal analysis. Principal component analysis (PCA) can be used for this purpose if there is a criterion upon which the correct order can be chosen. In this work, we propose a new and simple technique to determine automatically the rank of a PCA model. Tested with simulated data, the algorithm is able to determine the correct model order efficiently. Applied to video sequences, this method is able to estimate the necessary subspaces that capture the motion and illuminance changes within the different frames. This helps in reducing the storage need/requirements of video sequences and improves the efficiency of context based search and retrieval techniques. Michel Sarkis, Zaher Dawy, Florian Obermeier, Klaus Diepold |
ICIP | 4 |
| 2006 | Performance of Optical Flow Techniques on Graphics HardwareabstractSince graphics cards have become programmable the recent years, numerous computationally intensive algorithms have been implemented on the now called general purpose graphics processing units (GPGPUs). While the results show that GPGPUs regularly outperform CPU based implementations, the question arose how optical flow algorithms can be ported to graphics hardware. To answer the question, the optimal algorithm structure to maximize the performance gained by using graphics cards has to be found. In this paper we compare the performance of two algorithms that are highly different in structure, implemented on both CPU and graphics hardware. Analyzing the results of the CPU and GPGPU implementation, we explore the mapping of the algorithms to the graphics hardware and thereof extract information about a preferred structure of optical flow algorithms for GPGPU based implementation Marko Durkovic, Michael Zwick, Florian Obermeier, Klaus Diepold |
ICME | 4 |
| 2004 | Authentication of MPEG-4-based surveillance videoabstractThe industry is currently starting to use MPEG-4 compressed digital video for surveillance applications. The transition from analog to digital video raises difficulties for using surveillance video in court. Since it is fairly easy to make hard-to-detect modifications to the video stream captured by the cameras, e.g. mask out a specific event or person, a system for proving authenticity and integrity of video streams is needed. This paper presents such a system based on digital signatures embedded in the video stream. Our concept provides a means for proving authenticity and integrity of MPEG-4 digital video streams, while leaving compatibility with standard media players untouched. Michael Pramateftakis, Tobias Oelbaum, Klaus Diepold |
ICIP | 3 |
| 1997 | Actions of noncompact groups and algorithm design: a case studyabstractNumerical matrix computations involving actions of noncompact transformation groups are known to produce numerical problems since the elements of the pertaining matrix representations are inherently unbounded. In this case study we analyse numerical problems occurring in a class of algorithms that is based on actions of the pseudo-orthogonal group O/sub n,m/-a group that is noncompact (hyperbolic geometry) and well established in signal processing (Schur methods). As a major result, it is shown how to exploit the additional degrees of freedom in defining coordinate frames in a Grassmannian setting in order to impose an a priori bound on the norm of the transformation matrices. This way, numerically disastrous situations can be circumvented systematically. Hence, it becomes possible to develop modified algorithms which exhibit superior numerical performance for a large class of problems based on, for example, hyperbolic transformations. Klaus Diepold, Rainer Pauli |
ICASSP | 1 |
| 1993 | Embedding of non-contractive systems in lossless realizations
Klaus Diepold, Rainer Pauli |
ISCAS | 1 |
| 1992 | Schur parametrization of symmetric matrices with any rank profileabstractThe conceptual solution to the parametrization problem for symmetric indefinite matrices P is addressed. Beyond the fact that the symmetric matrix to be parametrized may have positive, negative and vanishing eigenvalues, it may as well comprise singular leading submatrices. For the parametrization, the lossless inverse scattering (LIS) framework is employed, which amounts to the mapping of a given symmetric matrix P onto a lossless and cascaded model structure. This leads to a recursive algorithm for the identification of the model parameters, the so-called Schur parameters, which turn out to form a set of vector-valued quantities to determine the individual lossless layers in the LIS model.> Klaus Diepold, Rainer Pauli |
ICASSP | 1 |
| 1991 | Schur parametrization of symmetric indefinite matricesabstractIt is shown that the generalized Schur algorithm for triangular factorization of symmetric positive definite matrices has a natural extension to the factorization of symmetric indefinite matrices with nonsingular principal submatrices. The proof is constructive and provides for an explicit formulation of the J-orthogonal and triangular matrices involved in the procedure. The (group-theoretic) significance of degenerate transformation steps involving unbounded reflection coefficients is precisely identified. It is found how to assign them an interpretation as Schur parameters and how to get benefit from this knowledge for performing a suitable change of equivalence class during execution, instead of a breakdown of the algorithm.> Klaus Diepold, Rainer Pauli |
ICASSP | 1 |