Eckehard G. Steinbach

dblp:s/EckehardGSteinbach · also Eckehard Steinbach · DBLP profile ↗
← Back
290ranked-venue papers
12as first author
84since 2021 · last 2026
0000-0001-8853-2703ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 196 · 10 first-author · 49 since 2021Artificial intelligence and machine learning · 52 · 2 first-author · 26 since 2021Systems, architecture and hardware · 32 · 15 since 2021Computer networks · 20 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 12 · 3 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 REVNET: Rotation-Equivariant Point Cloud Completion via Vector Neuron Anchor Transformer
Zhifan Ni, Eckehard G. Steinbach
ICPR (2)2
2026 CNN-Based Multipath-Aware Fingerprinting for Indoor Localization in 6G mm-Wave Networks
Majdi Abdmoulah, Nour Neji, Eckehard G. Steinbach
IWCMC3
2026 SMART-LEA: Scalable Multipath-Aware Localization Enhancement for IoT and 6G Sensor Networks
Majdi Abdmoulah, Eckehard G. Steinbach
IWCMC2
2026 MPUrge-MAP: A White-Box Multipath-Assisted Positioning for Indoor Localization in 6G Networks
Majdi Abdmoulah, Eckehard G. Steinbach
IWCMC2
2026 Human-in-the-loop RGB semantic segmentation for intuitive teaching of novel grasping tasks via few-shot adaptation of lightweight neural networks
abstract
Assistive robots are starting to be a part of everyday human life, yet the cognitive abilities of the robots still need to be improved to successfully operate in human environments. Vision-based grasping systems require determining the correct semantic region in the image that is suitable for task-oriented grasping of target objects. The complexity of unstructured environments like households and healthcare facilities necessitates the ability to learn new segmentation tasks from human guidance and to dynamically extend the perceptual skills on a daily basis to improve autonomy. We propose to use a browser-based deep interactive RGB segmentation interface for receiving the human guidance intuitively, allowing non-experts to guide the robots from remote. We further leverage meta learning to quickly adapt lightweight neural networks for custom semantic segmentation tasks with a few human demonstrations, which facilitates a quick, cost-efficient onboard learning on the assistive robots. For training and evaluation of the few-shot learning module with grasp area affordances, we extended a largescale generic grasping dataset with affordance segmentation labels and created a new dataset with 82,944 cluttered scene images and the corresponding segmentation masks. Comparative experiments on the datasets indicate the effectiveness of our few-shot learning modules by reaching high segmentation accuracies at a low computational cost. In addition, the robotic experiments with a 7-DOF manipulator show that the proposed methods outperform the baseline method by a large margin (22 to 25%) in task-oriented grasp precision and success rates. Lastly, the conducted user studies demonstrate the ease of use of the proposed demonstration interface. Consequently, our end-to-end pipeline facilitates real-world deployment of assistive robots in human environments by intuitively teaching new semantic segmentation tasks on a daily basis. TThe generated affordance dataset can be found at https://mediatum.ub.tum.de/1839309 .
Furkan Kaynar, Sudarshan Rajagopalan, Mahmoud Krichene, Eckehard G. Steinbach
Comput. Vis. Image Underst.4
2026 BIR-Adapter: A parameter-efficient diffusion adapter for blind image restoration
abstract
We introduce the BIR-Adapter, a parameter-efficient diffusion adapter for blind image restoration. Diffusion-based restoration methods have demonstrated promising performance in addressing this fundamental problem in computer vision, typically relying on auxiliary feature extractors or extensive fine-tuning of pre-trained models. Building on the observation that large-scale pretrained diffusion models can retain informative representations under image degradations, BIR-Adapter introduces a parameter-efficient, plug-and-play attention mechanism that substantially reduces the number of trained parameters. To further improve reliability, we adapt a sampling guidance mechanism that mitigates hallucinations during restoration. Experiments on synthetic and real-world degradations demonstrate that BIR-Adapter achieves competitive, and in several settings superior, performance compared to state-of-the-art methods while requiring up to 36 × fewer trained parameters. Moreover, the adapter-based design enables integration into existing models. We validate this generality by extending a super-resolution–only diffusion model to handle additional unknown degradations, highlighting the adaptability of our approach for broader image restoration tasks.
Cem Eteke, Alexander Griessel, Wolfgang Kellerer, Eckehard G. Steinbach
Pattern Recognit.4
2025 SMCNet: Supervised Surface Material Classification Using mmWave Radar IQ Signals and Complex-valued CNNs
abstract
Understanding surface material properties is crucial for enhancing indoor robot perception and indoor digital twinning. However, not all sensor modalities typically employed for this task are capable of reliably capturing detailed surface material characteristics. By analyzing the reflected RF signal from a mmWave radar sensor, it is possible to extract information about the reflective material and its composition from a certain surface. We introduce a mmWave MIMO FMCW radar-based surface material classifier SMCNet, employing a complex-valued Convolutional Neural Network (CNN) and complex radar IQ signal input for classifying indoor surface materials. While current radar-based material estimation approaches rely on a fixed sensing distance and constrained setups, our approach incorporates a setup with multiple sensing distances. We trained SMCNet using data from three distinct distances and subsequently tested it on these distances, as well as on two more unseen distances. We reached an overall accuracy of 99.12-99.53% on our test set. Notably, range FFT pre-processing improved accuracy on unknown distances from 25.25% to 58.81% without re-training.
Stefan Hägele, Fabián Seguel, Driton Salihu, Adam Misik, Eckehard G. Steinbach
ICASSP5
2025 FARE: A Deep Learning-Based Framework for Radar-Based Face Recognition and Out-of-Distribution Detection
abstract
In this work, we propose a novel pipeline for face recognition and out-of-distribution (OOD) detection using shortrange FMCW radar. The proposed system utilizes RangeDoppler and micro Range-Doppler Images. The architecture features a primary path (PP) responsible for the classification of in-distribution (ID) faces, complemented by intermediate paths (IPs) dedicated to OOD detection. The network is trained in two stages: first, the PP is trained using triplet loss to optimize ID face classification. In the second stage, the PP is frozen, and the IPs—comprising simple linear autoen-coder networks—are trained specifically for OOD detection. Using our dataset generated with a 60 GHz FMCW radar, our method achieves an ID classification accuracy of 99.30% and an OOD detection AUROC of 96.91%.
Sabri Mustafa Kahya, Boran Hamdi Sivrikaya, Muhammet Sami Yavuz, Eckehard G. Steinbach
ICASSP4
2025 HypCAD: Geometry-Enhanced Hyperbolic Contrastive Learning for CAD Model Retrieval
abstract
Retrieving CAD models for real-world object scans enhances object-level mapping, providing a nuanced spatial understanding crucial for precise interactions in robotics or mixed reality. Commonly, CAD model retrieval is performed by matching features learned in Euclidean space. However, learning discriminative features in Euclidean space faces significant challenges, primarily due to its flat nature and the wide variety of CAD models with different levels of detail. To address the limitations of Euclidean space and improve CAD model retrieval, this paper introduces HypCAD, a contrastive learning framework in hyperbolic space. We present a novel geometry-enhanced hyperbolic distance and utilize a three-component contrastive learning loss to learn hyperbolic feature representations for the CAD model retrieval task. We demonstrate HypCAD’s superior retrieval accuracy through comparisons with baseline contrastive learning methods on both the synthetic ShapeNet dataset and the real-world Scan2CAD dataset.
Adam Misik, Driton Salihu, Heike Brock, Eckehard G. Steinbach
ICASSP5
2025 Deep Learning-Based Perceptual Vibrotactile Codec with Rate Scalability
abstract
We present a vibrotactile codec based on a convolutional neural network that is fully rate-scalable and perceptually optimized. Rate-scalability so far was a rarely addressed problem with deep learning-based codecs. This is achieved through a bit allocation algorithm that perceptually optimizes the signal quantization, utilizing bitrate estimation based on a gaussian entropy model. We compare different bit allocation approaches regarding signal quality and found that our algorithm is effective at reaching the desired bitrate while achieving higher performance than our previously designed classical codec.
Lars Nockenberg, Wenxuan Wei, Mariam Navai, Eckehard G. Steinbach
ICASSP4
2025 MistSense: Versatile Online Detection of Procedural and Execution Mistakes
Constantin Patsch, Marsil Zakour, Driton Salihu, Eckehard G. Steinbach
ICCV5
2025 Real-Time Semantic Video Communication with Temporally Consistent And Controllable Diffusion Models
abstract
This paper introduces CVSC, a real-time-enabled and temporally consistent semantic video communication approach. We minimize the denoising steps of the diffusion model by using the most recent frame and motion information, enabling real-time-capable semantic video synthesis. To ensure temporal consistency, we employ windowed temporal cross-attention. While our objective evaluation highlights the limitations of existing metrics for generative models in semantic video communication, subjective evaluations demonstrate the superiority of our approach in terms of human preference at extremely low bit rates. (< 0.006 bpp).
Cem Eteke, Alexander Griessel, Wolfgang Kellerer, Eckehard G. Steinbach
ICIP4
2025 EQUR: Equivariant Uncertainty Quantification and Refinement for Point Cloud Registration
abstract
Point cloud registration is a crucial task for robotics and mixed reality applications, serving as a foundational component for problems such as 3D reconstruction and localization. Arbitrary poses and real-world artifacts, including noise and occlusions, increase registration uncertainty and limit the performance of current point cloud registration algorithms. This paper proposes a novel, sampling-free uncertainty quantification and refinement method for point cloud registration, termed EQUR. To consistently predict uncertainty with high robustness, we extract equivariant point features, from which we regress an uncertainty score, enabling robust quantification of registration uncertainty. Subsequently, we leverage the estimated registration uncertainty as an auxiliary input to enhance the prediction of transformation refinement terms. We employ an introspective learning strategy to train EQUR based on the errors of a baseline registration model. Through quantitative and qualitative analyses on synthetic ShapeNet and real-world ScanObjectNN datasets, we showcase the effectiveness of EQUR, demonstrating both high accuracies in uncertainty quantification and uncertainty-aided refinement of point cloud registration.
Adam Misik, Driton Salihu, Xiaoang Zhang, Heike Brock, Eckehard G. Steinbach
ICIP5
2025 Model-Mediated Teleoperation with 3D Dynamic Environment Tracking (MMT-DET): A Comparative Study of Task Performance with Time-Domain Passivity Control
abstract
Teleoperation with haptic feedback allows users to interact with remote environments while retaining a sense of touch. However, the stability and transparency of these systems are compromised under communication network delay. This paper presents an augmented Model-Mediated Teleoperation with 3D object and dynamic environment tracking (MMT-DET) by a vision-based algorithm, enabling users to receive haptic feedback in structured dynamic environments while maintaining robustness against network delays. A user study comparing the proposed method with teleoperation using the Time Domain Passivity Approach (TDPA) was conducted. The results demonstrate that our MMT-DET exhibits robustness to varying delays in task performance and outperforms TDPA at higher delay levels.
Diego Fernandez Prado, Jean Elsner, Hamid Sadeghian, Nader Rajaei, Abdeldjallil Naceri, Sami Haddadin, Eckehard G. Steinbach
IROS8
2025 On the Suitability of Perceptual Quality Metrics for Learning-Based Screen Content Compression
abstract
Learned image compression methods tailored for screen content demand perceptual quality metrics that are both accurate and computationally efficient. Traditional metrics such as PSNR and SSIM often fail to capture perceptual distortions specific to screen content, while many advanced learned perceptual metrics are too complex or non-differentiable for practical use with implicit neural codecs. These codecs require loss functions to be evaluated tens of thousands of times during training, imposing a strict complexity limit of a few hundred multiply-accumulate operations per pixel. In this paper, we evaluate several differentiable, low-complexity perceptual metrics on three screen content quality datasets, including a newly collected compression-focused crowdsourced dataset, PerceptualSCC. Our results identify VIF and NLPD as the best-performing metrics; however, only NLPD demonstrates stable behavior in gradient-based optimization experiments. These findings suggest NLPD as a strong candidate for guiding learned screen content compression, balancing perceptual fidelity, computational cost, and optimization stability.
H. Burak Dogaroglu, Hongjie You, Atanas Boev, Elena Alshina, Eckehard G. Steinbach
ISM5
2025 High-Fidelity Semantic Video Communication with Controllable Image-To-Video Diffusion Models
abstract
This work addresses the fidelity problem in low-bitrate real-time semantic video communication, which is crucial for enhancing the user experience. We present I2V-SC, extending baseline diffusion-based Image2Video (I2V) synthesis with ControlNet and distillation tailored to semantic video coding, enabling real-time performance with high fidelity. Evaluations using perceptual, pixel-level, motion, and semantic metrics demonstrate that I2V-SC outperforms baseline I2V and a baseline semantic video communication approach, namely CVSC, in ultra-low-bitrate ($<0.006$bpp) and real-time-enabled settings. Subjective evaluations confirm that I2V-SC further improves the user QoE in terms of overall preferability.
Cem Eteke, Alexander Griessel, Wolfgang Kellerer, Eckehard G. Steinbach
ISM4
2025 A Forearm-Worn Haptic Device for Integrated Tactile and Kinaesthetic Feedback via Skin Stretch
abstract
We propose a forearm-worn haptic device that integrates tactile skin-stretch feedback with kinaesthetic force feedback from a grounded haptic device. The wearable module generates lateral skin deformation corresponding to the tangential components of the applied kinaesthetic force, moving the skin in the opposite direction to produce coherent tactile cues that effectively convey sensations of surface friction and tangential shear, reinforcing the perception of contact with virtual surfaces. This integration aims to enhance both perceptual realism and transparency during haptic interaction. A series of preliminary experiments were conducted in a Chai3D-based virtual environment, where users interacted with virtual walls and corners while receiving combined feedback. The results demonstrate that the proposed device can render directionally accurate and temporally synchronized skin-stretch cues consistent with the kinaesthetic forces, validating its potential for multimodal haptic interaction.
Selin Nur Özsert, Daniel Rodriguez-Guevara, Leonardo Franco, Wenxuan Wei, Eckehard G. Steinbach, Domenico Prattichizzo
ISM5
2025 MPUrge: Merge, Purge, and Urge Multipath Fingerprints into Virtual Transmitters
abstract
With the advent of 6G networks, accurate and reliable location information is increasingly critical for location-based services and joint communication and sensing. In indoor localization (IL), virtual transmitters (VTs) have shown significant potential, particularly in mmWave-dominated 6G environments where specular reflections prevail. Multipath Delay Profile (MDP)-based VT construction algorithms match multipath components (MPCs) delays across multiple reference points to derive anchor-distance pairs, which serve as input for VT trilateration. While relying solely on delay information is advantageous, existing MDP-based approaches suffer from inaccuracies, noise sensitivity, high computational complexity, and extensive data requirements. This paper presents MPUrge, a novel tri-phase iterative MPC matching algorithm that employs a greedy strategy with advanced filtering mechanisms, balancing match quantity and quality while ensuring accurate VT localization based solely on delay data. Experimental results show a 4-to 10-fold improvement in VT construction accuracy and approximately 17% reduction in execution time. MPUrge outperforms state-of-the-art methods in accuracy, robustness, and efficiency while reducing computational requirements and data dependency. It is geometry-agnostic, scalable, and adaptable to any indoor environment. Additionally, MPUrge assigns a confidence score to each VT, enhancing localization system reliability and interpretability.
Majdi Abdmoulah, Eckehard G. Steinbach
IWCMC2
2025 Touch-Augmented Gaussian Splatting for Enhanced 3D Scene Reconstruction
Yue Gao 0001, Xiao Xu 0001, Eckehard G. Steinbach, Daniel Enrique Lucani, Qi Zhang 0013
MMSP3
2025 Multiresolution Contexts for Implicit Neural Codecs
H. Burak Dogaroglu, Cari E. Wiedemann, Eckehard G. Steinbach
PCS3
2025 Kinaesthetic Traffic Characterization and Performance Trade-Offs in Haptic Teleoperation Systems over 5G Standalone Private Networks
abstract
This paper presents a comprehensive characterization of kinaesthetic data traffic and associated performance trade-offs in multimodal bilateral haptic teleoperation over 5G networks. Utilizing two widely adopted haptic devices interacting with virtual environments, precise telemetry on motion and force was obtained across various interaction types and virtual object properties. The study investigates the effects of network resource allocation within a 5G standalone private network by employing three predefined network slicing profiles prioritizing haptic, audio, or video traffic. Performance evaluations were conducted using Quality of Service (QoS) metrics, from which potential implications for Quality of Experience (QoE) were subsequently inferred. The results show the interplay between traffic patterns and resource allocation, identifying configurations that ensure high-fidelity haptic feedback while maintaining acceptable audiovisual quality. These insights lay the groundwork for future research on adaptive network management strategies for haptic-enabled teleoperation systems.
Fernando Hernandez-Gobertti, Daniel Rodriguez-Guevara, Wenxuan Wei, Xiao Xu 0001, Eckehard G. Steinbach, David Gomez-Barquero
PIMRC5
2025 Proactive robot task sequencing through real-time hand motion prediction in human-robot collaboration
abstract
Human–robot collaboration (HRC) is essential for improving productivity and safety across various industries. While reactive motion re-planning strategies are useful, there is a growing demand for proactive methods that predict human intentions to enable more efficient collaboration. This study addresses this need by introducing a framework that combines deep learning-based human hand trajectory forecasting with heuristic optimization for robotic task sequencing. The deep learning model advances real-time hand position forecasting using a multi-task learning loss to account for both hand positions and contact delay regression, achieving state-of-the-art performance on the Ego4D Future Hand Prediction benchmark. By integrating hand trajectory predictions into task planning, the framework offers a cohesive solution for HRC. To optimize task sequencing, the framework incorporates a Dynamic Variable Neighborhood Search (DynamicVNS) heuristic algorithm, which allows robots to pre-plan task sequences and avoid potential collisions with human hand positions. DynamicVNS provides significant computational advantages over the generalized VNS method. The framework was validated on a UR10e robot performing a visual inspection task in a HRC scenario, where the robot effectively anticipated and responded to human hand movements in a shared workspace. Experimental results highlight the system’s effectiveness and potential to enhance HRC in industrial settings by combining predictive accuracy and task planning efficiency. • We enhance hand position forecasting with a novel loss, achieving state-of-the-art results. • Our method unifies forecasting and task planning using the Dynamic TSP with Time Windows. • We enable real-time hand motion prediction and seamless integration into robot planning.
Shyngyskhan Abilkassov, Michael Gentner, Almas Shintemirov, Eckehard G. Steinbach, Mirela Popa
Image Vis. Comput.4
2024 Importance-Driven Semantic Resilience for Challenging Future 6G Channels
abstract
Semantic Communication has recently emerged as a novel communication strategy that prioritizes transmitting meaning over conventional bit-based transmission. By significantly reducing resource requirements in communication tasks such as video conferencing, natural language, and audio transmission, Semantic Communication promises better utilization of challenging, near radio link failure (NRLF) channels. This paper introduces a novel semantic communication framework designed to further enhance the resilience of the transmission of semantics over NRLF channels. Unlike classical, Shannon-based communication that prioritizes the perfect reception of bits, our approach focuses on ensuring the successful synthesis of the message semantics. Our framework leverages the significance of discrete semantics, and a cross-layer strategy to ensure message integrity and comprehension, even under significant loss. Key to our framework is the novel, code block based Proactive Redundancy Transmission (PRT) mechanism prioritizing critical semantics, coupled with a novel error concealment step enabling meaningful reconstruction of non-critical semantics. We establish the resulting importance-driven resilience optimization problem, and introduce and validate preliminary heuristics as an initial attempt to optimize it. We formalize, implement, and evaluate our framework, demonstrating significant improvements in the resilience of semantic communication in NRLF environments. Our evaluations, leveraging a First Order Motion Model (FOMM) for video conferencing synthesis, underscore the resilience of our semantic communication framework against traditional H.265 compression under challenging Channel Block Error Rates (CBLERs). Unlike H.265, which fails to decode under significant CBLERs, our method exhibits remarkable resilience, maintaining perceptual quality even with CBLERs surpassing 75%.
Alexander Griessel, Cem Eteke, Eckehard G. Steinbach, Wolfgang Kellerer
GLOBECOM3
2024 HAROOD: Human Activity Classification and Out-Of-Distribution Detection with Short-Range FMCW Radar
abstract
We propose HAROOD as a short-range FMCW radar-based human activity classifier and out-of-distribution (OOD) detector. It aims to classify human sitting, standing, and walking activities and to detect any other moving or stationary object as OOD. We introduce a two-stage network. The first stage is trained with a novel loss function that includes intermediate reconstruction loss, intermediate contrastive loss, and triplet loss. The second stage uses the first stage’s output as its input and is trained with cross-entropy loss. It creates a simple classifier that performs the activity classification. On our dataset collected by 60 GHz short-range FMCW radar, we achieve an average classification accuracy of 96.51%. Also, we achieve an average AUROC of 95.04% as an OOD detector. Additionally, our extensive evaluations demonstrate the superiority of HAROOD over the state-of-the-art OOD detection methods in terms of standard OOD detection metrics.
Sabri Mustafa Kahya, Muhammet Sami Yavuz, Eckehard G. Steinbach
ICASSP3
2024 Long-Term Action Anticipation Based on Contextual Alignment
abstract
In action anticipation, the model predicts the next future action after a certain observation period. In long-term action anticipation, this idea is further extended to predicting multiple actions and their respective duration. Thus, in this problem setting the model should not only capture relationships between past actions but also predict several future actions that fit into a certain context. Compared to autoregressive models, our model employs an encoder decoder structure to determine future actions and durations in parallel, which prevents the accumulation of prediction errors and reduces the inference time. Furthermore, it is ensured that the predicted actions are aligned with respect to a context representation, which resembles the way humans approach this task as the feasible action set is restricted by the respective context. We evaluate our model on the long-term anticipation benchmark datasets, Breakfast, and 50Salads, where we achieve state-of-the-art results.
Constantin Patsch, Jinghan Zhang 0009, Marsil Zakour, Driton Salihu, Eckehard G. Steinbach
ICASSP6
2024 NPRF: Neural Painted Radiosity Fields for Neural Implicit Rendering and Surface Reconstruction
abstract
In recency, neural signed distance fields have become more popular for reconstructing 3D indoor environments. While great improvements have been made due to missing incident radiance and materials in the surface estimation, current methods cannot reconstruct high-quality surfaces. To address this issue, we propose Neural Painted Radiosity Fields (NPRF), consisting of Neural Radiosity Fields for volumetric surface representation and Neural Painted Scenes for novel view synthesis. Neural Radiosity Fields combine the radiative transfer equation with neural radiosity to estimate 3D surfaces, thus leveraging raytracing to improve the volumetric representation. Neural Painted Scenes employs sparsification and projection of 3D points into 2D images in conjunction with a generative, context-aware inpainting network to produce high-quality novel views. We show that NPRF leads to overall improvements in F-score on the popular ScanNet dataset. Finally, we show that NPRF improves novel view synthesis by a significant margin, giving improvements of up to 25% on PSNR, 53% on LPIPS, and 3% on SSIM.
Driton Salihu, Adam Misik, Constantin Patsch, Eckehard G. Steinbach
ICASSP5
2024 OCTOPUS: Optimized Cross-border TeleOperated Medicine Pouring Using NextGen Seamless Communication Networks
abstract
Teleoperated robotic systems have become instrumental in advancing remote healthcare services, especially in tasks that require precision and expert oversight. The advent of cutting-edge telecommunication infrastructures, such as 5G, has amplified interest in these systems, although their full potential remains untapped. This study delves into the effectiveness of teleoperated robotic systems for medicine dispensing, comparing the performance of Wi-Fi and 5G networks in a transnational setup between two cities - Prague and Munich. We focus on the robot's ability to accurately dispense a predefined volume of a syrup-like substance, simulating a delicate healthcare operation, under the guidance of a distant operator. Our research examines the system's holistic performance in real-world implementation across diverse scenarios, encompassing varying network states and feedback methods. Two primary feedback scenarios are considered: one incorporating real-time video streaming and another offering explicit quantitative data on the dispensed volume. Using a blend of quantitative and qualitative methods, we aim to determine the influence of network type and feedback on task efficacy and user satisfaction. This study provides insights into the potential and hurdles of deploying teleoperated robotic systems in crucial healthcare contexts, guiding future advancements in this domain, especially in scenarios, where precision and dependability are crucial.
Edwin Babaians, Praveen Gorla, Serkut Ayvasik, Jan Plachy, Zdenek Becvar, Wolfgang Kellerer, Eckehard G. Steinbach
ICC7
2024 Adapting Learned Image Codecs To Screen Content Via Adjustable Transformations
abstract
As learned image codecs (LICs) become more prevalent, their low coding efficiency for out-of-distribution data becomes a bottleneck for some applications. To improve the performance of LICs for screen content (SC) images without breaking backwards compatibility, we propose to introduce parameterized and invertible linear transformations into the coding pipeline without changing the underlying baseline codec’s operation flow. We design two neural networks to act as prefilters and postfilters in our setup to increase the coding efficiency and help with the recovery from coding artifacts. Our end-to-end trained solution achieves up to $10 \%$ bitrate savings on SC compression compared to the baseline LICs while introducing only $1 \%$ extra parameters.
H. Burak Dogaroglu, Ahmet Burakhan Koyuncu, Atanas Boev, Elena Alshina, Eckehard G. Steinbach
ICIP5
2024 Real-Time Semantic Video Communication of General Scenes
abstract
This paper presents a real-time semantic video communication method for general scenes, combining lossy semantic map coding with motion compensation to achieve reduced bit rates while maintaining perceptual and semantic quality. Our findings show that semantic image synthesis effectively adapts to minute errors resulting from motion estimation, eliminating the need to transmit the residuals. We recommend the Group of Pictures approach as a more efficient alternative. Comparative assessments against HEVC and VVC confirm the method’s effectiveness. This research paves the way for efficient real-time semantic video communication, addressing the demands of data-intensive visual applications.
Cem Eteke, Alexander Griessel, Wolfgang Kellerer, Eckehard G. Steinbach
ICIP4
2024 Food: Facial Authentication And Out-Of-Distribution Detection With Short-Range FMCW Radar
abstract
This paper proposes a short-range FMCW radar-based facial authentication and out-of-distribution (OOD) detection framework. Our pipeline jointly estimates the correct classes for the in-distribution (ID) samples and detects the OOD samples to prevent their inaccurate prediction. Our reconstruction-based architecture consists of a main convolutional block with one encoder and multi-decoder configuration, and intermediate linear encoder-decoder parts. Together, these elements form an accurate human face classifier and a robust OOD detector. For our dataset, gathered using a 60 GHz short-range FMCW radar, our network achieves an average classification accuracy of 98.07% in identifying in-distribution human faces. As an OOD detector, it achieves an average Area Under the Receiver Operating Characteristic (AUROC) curve of 98.50% and an average False Positive Rate at 95% True Positive Rate (FPR95) of 6.20%. Also, our extensive experiments show that the proposed approach outperforms previous OOD detectors in terms of common OOD detection metrics.
Sabri Mustafa Kahya, Boran Hamdi Sivrikaya, Muhammet Sami Yavuz, Eckehard G. Steinbach
ICIP4
2024 DeepSPF: Spherical SO(3)-Equivariant Patches for Scan-to-CAD Estimation
abstract
Recently, SO(3)-equivariant methods have been explored for 3D reconstruction via Scan-to-CAD. Despite significant advancements attributed to the unique characteristics of 3D data, existing SO(3)-equivariant approaches often fall short in seamlessly integrating local and global contextual information in a widely generalizable manner. Our contributions in this paper are threefold. First, we introduce Spherical Patch Fields, a representation technique designed for patch-wise, SO(3)-equivariant 3D point clouds, anchored theoretically on the principles of Spherical Gaussians. Second, we present the Patch Gaussian Layer, designed for the adaptive extraction of local and global contextual information from resizable point cloud patches. Culminating our contributions, we present Learnable Spherical Patch Fields (DeepSPF) – a versatile and easily integrable backbone suitable for instance-based point networks. Through rigorous evaluations, we demonstrate significant enhancements in Scan-to-CAD performance for point cloud registration, retrieval, and completion: a significant reduction in the rotation error of existing registration methods, an improvement of up to 17\% in the Top-1 error for retrieval tasks, and a notable reduction of up to 30\% in the Chamfer Distance for completion models, all attributable to the incorporation of DeepSPF.
Driton Salihu, Adam Misik, Constantin Patsch, Fabián Seguel, Eckehard G. Steinbach
ICLR6
2024 RFOOD: Real-time Facial Authentication and Out -of-distribution Detection with Short-range FMCW Radar
abstract
Out-of-distribution (OOD) detection is critical for the safe deployment of modern neural network architectures, as it aims to identify samples outside the training domain. In this paper, we introduce RFOOD, a novel OOD detection framework designed for real-time, privacy-preserving facial authentication using low-cost frequency-modulated continuous-wave (FMCW) radar. RFOOD employs both range-Doppler and micro range- Doppler images to enhance the detection accuracy. The architecture consists of a multi-encoder multi-decoder Body Part (BP) and Intermediate Linear Encoder-Decoder (ILED) components. This design allows the system to accurately classify a single individual's face as in-distribution (ID) while identifying all other faces as OOD. On our dataset collected with 60 GHz short-range FMCW radar, RFOOD achieves an Area Under the Receiver Operating Characteristic (AUROC) curve of 94.13 % and a False Positive Rate of 18.12% at a True Positive Rate of 95 % (FPR95). Additionally, RFOOD outperforms state-of-the-art OOD detection methods in common OOD detection metrics and operates in real-time.
Sabri Mustafa Kahya, Muhammet Sami Yavuz, Boran Hamdi Sivrikaya, Eckehard G. Steinbach
ICMLA4
2024 HEGN: Hierarchical Equivariant Graph Neural Network for 9DoF Point Cloud Registration
abstract
Given its wide application in robotics, point cloud registration is a widely researched topic. Conventional methods aim to find a rotation and translation that align two point clouds in 6 degrees of freedom (DoF). However, certain tasks in robotics, such as category-level pose estimation, involve non-uniformly scaled point clouds, requiring a 9DoF transform for accurate alignment. We propose HEGN, a novel equivariant graph neural network for 9DoF point cloud registration. HEGN utilizes equivariance to rotation, translation, and scaling to estimate the transformation without relying on point correspondences. Based on graph representations for both point clouds, we extract equivariant node features aggregated in their local, cross-, and global context. In addition, we introduce a novel node pooling mechanism that leverages the cross-context importance of nodes to pool the graph representation. By repeating the feature extraction and node pooling, we obtain a graph hierarchy. Finally, we determine rotation and translation by aligning equivariant features aggregated over the graph hierarchy. To estimate scaling, we leverage scale information in the vector norm of the equivariant features. We evaluate the effectiveness of HEGN through experiments with the synthetic ModelNet40 dataset and the real-world ScanObjectNN dataset. The results show the superior performance of HEGN in 9DoF point cloud registration and its competitive performance in conventional 6DoF point cloud registration.
Adam Misik, Driton Salihu, Heike Brock, Eckehard G. Steinbach
ICRA5
2024 HPF-SLAM: An Efficient Visual SLAM System Leveraging Hybrid Point Features
abstract
Visual SLAM is an essential tool in diverse applications such as robot perception and extended reality, where feature-based methods are prevalent due to their accuracy and robustness. However, existing methods employ either hand-crafted or solely learnable point features and are thus limited by the feature attributes. In this paper, we propose incorporating hybrid point features efficiently into a single system. By integrating hand-crafted and learnable features, we seek to capitalize on their complementary attributes in both key-point identification and descriptor expressiveness. To this purpose, we design a pre-processing module, which includes extraction, inter-class processing, and post-processing of hybrid point features. We present an efficient matching approach to exclusively perform the data association within the same class of features. Moreover, we design a Hybrid Bag-of-Words (H-BoW) model to deal with hybrid point features in matching and loop-closure-detection. By integrating the proposed framework into a modern feature-based system, we introduce HPF-SLAM. We evaluate the system on EuRoC-MAV and TUM-RGBD benchmarks. The experimental results show that our method consistently surpasses the baseline at comparable speed.
Sebastian Eger, Adam Misik, Rastin Pries, Eckehard G. Steinbach
ICRA6
2024 Sim-to-Real Domain Shift in Online Action Detection
abstract
Human reasoning comprises the ability to understand and reason about the current action solely based on past information. To provide effective assistance in an eldercare or household environment an assistive robot or intelligent assistive system has to assess human actions correctly. Based on this presumption, the task of online action detection determines the current action solely based on the past without access to future information. During inference, the performance of the model is largely impacted by the attributes of the underlying training dataset. However, as high costs and ethical concerns are associated with the real-world data collection process, synthetically created data provides a way to mitigate these problems while providing additional data for the training process of the underlying action detection model to improve performanceDue to the inherent domain shift between the synthetic and real data, we introduce a new egocentric dataset called Human Kitchen Interactions (HKI) to investigate the sim-to-real gap. Our dataset contains in total 100 synthetic and real videos in which 21 different actions are executed in a kitchen environment. The synthetic data is acquired in an egocentric virtual reality (VR) setup while capturing the virtual environment in a game engine. We evaluate state-of-the-art online action detection models on our dataset and provide insights into sim-to-real domain shift. Upon acceptance, we will release our dataset and the corresponding features at https://c-patsch.github.io/HKI/.
Constantin Patsch, Wael Torjmene, Marsil Zakour, Driton Salihu, Eckehard G. Steinbach
IROS6
2024 Enhanced Robotic Assistance for Human Activities through Human-Object Interaction Segment Prediction
abstract
Robotic assistance is a current research topic with high application value and multiple challenges. Assistive robots are used in various scenarios, such as production lines, operating tables, and elderly care. While providing effective assistance, most of the assistance tasks that current robots can perform are limited to predefined tasks. This limitation arises from the insufficiency of the current robot perception system to forecast future human activities. To address this issue, we propose a novel 2-stage robotic assistant for human activities through future human-object interaction (HOI) segment prediction. Unlike previous work focusing on predefined or short-term tasks, our robotic assistant can make predictions for future assistance according to human habits. In the first stage, we propose a visual-based human-object interaction segment prediction method to predict human activities, which enables the robotic system to infer human intention. Moreover, we define the robotic executable tasks as an interactive tuple to keep the robotic assistance normatively consistent with human activity. Meanwhile, a graph convolutional network with geometric features that can predict human-object interaction segments is proposed to provide target manipulation and target object for the assistive robot. In the second stage, we present a mobile task completion process including visual navigation, object localization and grasping. The perception stage is evaluated on the MPHOI dataset and custom-collected SPHOI dataset. Finally, we evaluate our comprehensive framework through real-time experimentation.
Rayene Messaoud, Arne-Christoph Hildebrandt, Marco Baldini, Driton Salihu, Constantin Patsch, Eckehard G. Steinbach
IROS7
2024 Rethinking 3D Geometric Object Features for Enhancing Skeleton-based Action Recognition
abstract
Human action recognition is crucial for intelligent robots, especially in the realm of human-robot collaboration research. Recent advancements in human pose estimation algorithms have shifted the focus of action recognition towards skeleton-based models, which exhibit robustness to changes in background and illumination. However, many state-of-the-art action recognition models rely on 2D skeleton data, neglecting object features. This limitation becomes obvious in complex scenarios where human interactions with objects are crucial, potentially compromising the reliability of assistive robots in understanding human behavior in their environment. To address this issue, we propose a method that effectively integrates 3D geometric object features into skeleton data using graph convolutional neural networks (GCNs). In addition to analyzing the effectiveness of information from different dimensions such as object center position, category, translation, and rotation, we explore various adjacency matrix designs for graph networks. Our model performance is evaluated on two challenging datasets: IKEA ASM and Bimanual Actions. The results demonstrate a significant improvement in action recognition by integrating object features into skeleton-based models. Specifically, on the IKEA-ASM dataset, our approach achieves a frame-wise Top-1 score improvement of 10.8% and an average F1@k improvement of 13.3%, while on the Bimanual Actions dataset, it achieves a frame-wise Top-1 score improvement of 11.4% and an average F1@k improvement of 5.3%, with negligible increases in model complexity.
Driton Salihu, Constantin Patsch, Marsil Zakour, Eckehard G. Steinbach
IROS6
2024 Lossy Coding for Spatially Adaptive Conditioning in Semantic Image Communication
abstract
The increasing demand for high-quality, real-time visual communication and the growing user expectations, coupled with limited network resources, necessitate novel approaches to semantic image communication. This paper presents a method to enhance semantic image communication that combines a novel lossy semantic encoding approach with spatially adaptive semantic image synthesis models. By developing a model-agnostic training augmentation strategy, our approach substantially reduces susceptibility to distortion introduced during encoding, effectively eliminating the need for lossless semantic encoding. Comprehensive evaluation across two spatially adaptive conditioning methods and three popular datasets indicates that this approach enhances semantic image communication at very low bit rate regimes.
Cem Eteke, Alexander Griessel, Wolfgang Kellerer, Eckehard G. Steinbach
VCIP4
2024 Judder Modelling Framework with Perceptual Quality Score Prediction for HDR Videos
abstract
Judder is a motion artifact describing a perceptual mismatch between human visual system and discrete movements on a display. Judder is correlated with video frame rates, motion speeds and brightness. It significantly degrades the perceived video quality. We present a framework for modeling judder, where the input frames are processed to generate optical flow and a sensitivity map. Concurrently, a specifically designed attention map is integrated into the process. These components collectively contribute to the creation of a feature map, which is utilized to compute judder scores. The proposed approach explicitly takes into account the video frame rate, resulting in an accurate judder score to predict the subjective mean opinion scores ultimately. Our assessment of the proposed framework on an HDR video dataset shows that judderness is highly influenced by the frame rate and, to some extent, by motion speed and brightness. The prediction of perceived quality scores shows improvement compared to the baseline framework.
Hongjie You, Nicola Giuliani, Atanas Boev, Elena Alshina, Eckehard G. Steinbach
VCIP6
2024 Efficient Contextformer: Spatio-Channel Window Attention for Fast Context Modeling in Learned Image Compression
abstract
Entropy estimation is essential for the performance of learned image compression. It has been demonstrated that a transformer-based entropy model is of critical importance for achieving a high compression ratio, however, at the expense of a significant computational effort. In this work, we introduce the Efficient Contextformer (eContextformer) – a computationally efficient transformer-based autoregressive context model for learned image compression. The eContextformer efficiently fuses the patch-wise, checkered, and channel-wise grouping techniques for parallel context modeling, and introduces a shifted window spatio-channel attention mechanism. We explore better training strategies and architectural designs and introduce additional complexity optimizations. During decoding, the proposed optimization techniques dynamically scale the attention span and cache the previous attention computations, drastically reducing the model and runtime complexity. Compared to the non-parallel approach, our proposal has ~145x lower model complexity and ~210x faster decoding speed, and achieves higher average bit savings on Kodak, CLIC2020, and Tecnick datasets. Additionally, the low complexity of our context model enables online rate-distortion algorithms, which further improve the compression performance. We achieve up to 17% bitrate savings over the intra coding of Versatile Video Coding (VVC) Test Model (VTM) 16.2 and surpass various learning-based compression models.
Ahmet Burakhan Koyuncu, Panqi Jia, Atanas Boev, Elena Alshina, Eckehard G. Steinbach
IEEE Trans. Circuits Syst. Video Technol.5
2023 Mcrood: Multi-Class Radar Out-Of-Distribution Detection
abstract
Out-of-distribution (OOD) detection has recently received special attention due to its critical role in safely deploying modern deep learning (DL) architectures. This work proposes a reconstruction-based multi-class OOD detector that operates on radar range doppler images (RDIs). The detector aims to classify any moving object other than a person sitting, standing, or walking as OOD. We also provide a simple yet effective pre-processing technique to detect minor human body movements like breathing. The simple idea is called respiration detector (RESPD) and eases the OOD detection, especially for human sitting and standing classes. On our dataset collected by 60GHz short-range FMCW Radar, we achieve AUROCs of 97.45%, 92.13%, and 96.58% for sitting, standing, and walking classes, respectively. We perform extensive experiments and show that our method outperforms state-of-the-art (SOTA) OOD detection methods. Also, our pipeline performs 24 times faster than the second-best method and is very suitable for real-time processing.
Sabri Mustafa Kahya, Muhammet Sami Yavuz, Eckehard G. Steinbach
ICASSP3
2023 Self-Attention Based Action Segmentation Using Intra-And Inter-Segment Representations
abstract
Segmenting activities in untrimmed videos remains a critical challenge to fully understand complex human activity sequences. A correct representation of temporal action relations is key for improving incorrect segmentations. We propose a self-attention-based model that refines initial segmentations by separately considering intra-as well as inter-segment relations between predicted action segments. Furthermore, in order to enhance the training process, we use a similarity-guided regularization technique that ensures intra-segment similarity and the validity of action transitions between adjacent segments. In an extensive evaluation on three public datasets -Georgia Tech Egocentric Activities, 50Salads, and Breakfast -our proposed architecture enhances the backbone model by 6.1% on GTEA, 3.8% on 50Salads, and 3.9% on Breakfast with regard to the F 1@50 metric.
Constantin Patsch, Eckehard G. Steinbach
ICASSP2
2023 COCCA: Point Cloud Completion through Cad Cross-Attention
abstract
3D scene- and object-level scans typically result in sparse and incomplete point clouds. Since dense point clouds of high quality are essential for the 3D reconstruction process, a promising approach is to improve the scan quality by point cloud completion. In this paper, we present COCCA, an extension of point cloud completion networks for scan-to-CAD use cases. The proposed extension is based on cross-attention of features extracted from a scan with rotation-, translation-, and scale-invariant features extracted from a sampled CAD point cloud. With the proposed cross-attention operation, we improve the learning of scan features and the subsequent decoding to a complete shape. We demonstrate the effectiveness of COCCA on the ShapeNet dataset in quantitative and qualitative experiments. COCCA improves the overall completion performance of point cloud completion networks by up to 11.8% for Chamfer Distance and up to 2.2% for F-Score. Our qualitative experiments visualize how COCCA completes point clouds with higher geometric detail. In addition, we demonstrate how completion by COCCA improves the point cloud registration task required for scan-to-CAD alignment.
Adam Misik, Driton Salihu, Heike Brock, Eckehard G. Steinbach
ICIP4
2023 SRI-Graph: A Novel Scene-Robot Interaction Graph for Robust Scene Understanding
abstract
We propose a novel scene-robot interaction graph (SRI-Graph) that exploits the known position of a mobile manipulator for robust and accurate scene understanding. Compared to the state-of-the-art scene graph approaches, the proposed SRI-Graph captures not only the relationships between the objects, but also the relationships between the robot manipulator and objects with which it interacts. To improve the detection accuracy of spatial relationships, we leverage the 3D position of the mobile manipulator in addition to RGB images. The manipulator's ego information is crucial for a successful scene understanding when the relationships are visually uncertain. The proposed model is validated for a real-world 3D robot-assisted feeding task. We release a new dataset named 3DRF-Pos for training and validation. We also develop a tool, named LabelImg-Rel, as an extension of the open-sourced image annotation tool LabelImg for a convenient annotation in robot-environment interaction scenarios*. Our experimental results using the Movo platform show that SRI-Graph outperforms the state-of-the-art approach and improves detection accuracy by up to 9.83%.
Xiao Xu 0001, Mengchen Xiong, Edwin Babaians, Eckehard G. Steinbach
ICRA5
2023 Dynamic Multi-Query Motion Planning with Differential Constraints and Moving Goals
abstract
Planning robot motions in complex environments is a fundamental research challenge and central to the autonomy, efficiency, and ultimately adoption of robots. While often the environment is assumed to be static, real-world settings, such as assembly lines, contain complex shaped, moving obstacles and changing target states. Therein robots must perform safe and efficient motions to achieve their tasks. In repetitive environments and multi-goal settings, reusable roadmaps can substantially reduce the overall query time. Most dynamic roadmap-based planners operate in state-time-space, which is computationally demanding. Interval-based methods store availabilities as node attributes and thereby circumvent the dimensionality increase. However, current approaches do not consider higher-order constraints, which can ultimately lead to collisions during execution. Furthermore, current approaches must replan when the goal changes. To this end, we propose a novel roadmap-based planner for systems with third-order differential constraints operating in dynamic environments with moving goals. We construct a roadmap with availabilities as node attributes. During the query phase, we use a Double-Integrator Minimum Time (DIMT) solver to recursively build feasible trajectories and accurately estimate arrival times. An exit node set in combination with a moving goal heuristic is used to efficiently find the fastest path through the roadmap to the moving goal. We evaluate our method with a simulated UAV operating in dynamic 2D environments and show that it also transfers to a 6-DoF manipulator. We show higher success rates than other state-of-the-art methods both in collision avoidance and reaching a moving goal.
Michael Gentner, Fabian Zillenbiller, André Kraft, Eckehard G. Steinbach
IROS4
2023 Haptic Dataset Augmentation with Subjective QoE Labels using Conditional Generative Adversarial Network
abstract
This paper proposes a novel Generative Adversarial Network (GAN)-based strategy to augment subjective haptic Quality of Experience (QoE) datasets for bilateral teleoperation with haptic feedback without conducting time-consuming subjective experiments. In our previous work, we proposed a multi-assessment fusion approach to predict subjective haptic quality using a collection of objective metrics. This method requires a sufficiently large haptic dataset with QoE labels. The proposed generative approach automatically expands the existing haptic quality dataset by combining a modified conditional GAN (CGAN) and Style GAN (StyleGAN) architecture. The most important feature of our method is that it learns from the labeled training data and focuses on synthesizing signals with artifacts according to new input labels containing the QoE score, time delay, control method, and data reduction information. Extensive experiments are conducted to validate the suitability of the expanded dataset. The results show that our approach is able to generate new data, which match the label and signal distribution of the original data with categorical rank and linear correlation of over 0.85.
Zican Wang, Xiao Xu 0001, Zhenyu Wang 0010, Sarah Shtaierman, Eckehard G. Steinbach
IROS6
2023 Care3D: An Active 3D Object Detection Dataset of Real Robotic-Care Environments
abstract
As labor shortage increases in the health sector, the demand for assistive robotics grows. However, the needed test data to develop those robots is scarce, especially for the application of active 3D object detection, where no real data exists at all. This short paper counters this by introducing such an annotated dataset of real environments. The captured environments represent areas which are already in use in the field of robotic health care research. We further provide ground truth data within one room, for assessing SLAM algorithms running directly on a health care robot.
Michael G. Adam, Sebastian Eger, Martin Piccolrovazzi, Maged Iskandar, Jörn Vogel, Alexander Dietrich, Seongjin Bien, Jon Skerlj, Abdeldjallil Naceri, Eckehard G. Steinbach, Alin Albu-Schäffer, Sami Haddadin, Wolfram Burgard
ISM10
2023 CALC-VFS: Content-adaptive low-complexity Video Frame Synthesis
abstract
We present a content-adaptive, low-complexity video frame synthesis algorithm. Our approach applies the dynamic convolutions content adaptation approach to the widely used frame synthesis algorithm IFRNet. By introducing dynamic convolutions into both the pyramid encoder and the coarse-to-fine decoders of IFRNet, we enforce sparsity, thereby limiting the computationally expensive operations to only the necessary pixels. Training for specific sparsity targets allows us to achieve overall less computational complexity compared to IFRNet while having similar performance. We demonstrate the performance and content adaptivity in two test scenarios and show the savings in computational budget (approximately 20-40%) compared to the baseline IFRNet.
Nicola Giuliani, Hongjie You, Ahmet Burakhan Koyuncu, Atanas Boev, Elena Alshina, Eckehard G. Steinbach
ISM6
2023 Quality of Task Perception based Performance Optimization of Time-delayed Teleoperation
abstract
This paper proposes a Quality-of-Task-Perception (QoTP) based performance optimization approach for bilateral haptic teleoperation. For time-delayed teleoperation, stabilizing control schemes are combined with communication and data reduction algorithms to ensure stability, transparency, and Quality of Experience (QoE). An adaptive control scheme switching strategy to improve the QoE of teleoperation considering network quality of service (QoS) and quality of control (QoC) is proposed in our previous work. In this paper, we introduce a novel concept named quality of task perception (QoTP) to optimize teleoperation from another dimension in addition to QoS and QoC. QoTP represents the pre-cognition of the task and the accuracy of the environment restoration. The proposed optimization approach is applied to a haptic teleoperation system with switchable control schemes (prediction-based or passivity-based). An environment restoration model is set on the leader side using the least squares method (LSM) to fit different environment models and provide force feedback without the influence of round-trip delay. We also evaluate the system performance with different delays, control schemes, and model complexities both objectively and subjectively. Our experiments validate the proposed approach and show that the QoE performance increases when selecting the more accurate environment restoration model in the QoTP dimension considering the system’s computing power.
Xiao Xu 0001, Zican Wang, Zhi Jin 0002, Eckehard G. Steinbach
RO-MAN6
2023 ISSC: Interactive Semantic Shared Control for Haptic Teleoperation
abstract
We propose a novel interactive semantic shared control framework that exploits an active high-level communication loop between the human operator and the robot for time-efficient teleoperation. In shared control approaches, accurate prediction of the operator’s intention is crucial to enable the robot to provide meaningful assistance. Incorrect intention prediction (e.g., target objects to be interacted with) increases the task duration due to conflicts between human behaviors and robot guidance. Unlike existing methods, our approach not only passively observes and predicts the human operator’s input in the haptic control loop, but also actively communicates with the human operator in an additional semantic loop in the form of a speech user interface to optimize the effectiveness of assistance. We evaluate our ISSC framework for a pegin-hole teleoperation task. The experimental results show that the proposed framework significantly outperforms teleoperation without assistance and conventional shared control paradigms regarding task execution efficiency and user control quality, and reduces task completion time by up to 26.68% and 39.00%, respectively.
Xiao Xu 0001, Mengchen Xiong, Edwin Babaians, Zican Wang, Fanle Meng, Eckehard G. Steinbach
RO-MAN7
2023 Subjective video quality assessment of immersive HDR content on head-mounted displays
abstract
High dynamic range (HDR) videos are known to provide better visual quality on HDR TV displays. Head-mounted displays (HMDs) are an integral part of immersive visual experiences. However, typical HMDs are equipped with standard dynamic range (SDR) displays, failing to show details in bright and dark areas of HDR content. Therefore, we aim to evaluate the perceptual quality improvement, in terms of mean opinion score (MOS), when observers view HDR instead of SDR content on immersive displays. We developed a pipeline to render 2D HDR and tone-mapped SDR immersive videos for HMDs. We conducted two single-stimulus subjective evaluation experiments to evaluate (1) the perceived visual difference when comparing one HDR scene with three tone-mapped SDR versions of it, and (2) how frame rates impact the perceptual quality of HDR immersive videos. Our results in MOS show that (1) there is a significant improvement of perceptual quality in HDR compared to tone-mapped SDR content, and (2) HDR immersive videos benefit much more from higher frame rates than SDR videos.
Hongjie You, Nicola Giuliani, Atanas Boev, Elena Alshina, Eckehard G. Steinbach
VCIP6
2023 SGPCR: Spherical Gaussian Point Cloud Representation and its Application to Object Registration and Retrieval
abstract
Retrieving and aligning CAD models from databases with scanned real-world point clouds remains an important topic for 3D reconstruction. Due to zero point-to-point correspondences between the sampled CAD model and the scanned real-world object, an information-rich representation of point clouds is needed. We propose SGPCR, a novel method for representing 3D point clouds by Spherical Gaussians for efficient, stable, and rotation-equivariant representation. We also propose a rotation-invariant convolution to improve the representation quality through a trainable optimization process. In addition, we demonstrate the strengths of SGPCR-based point cloud representation using the fundamental challenge of shape retrieval and point cloud registration on point clouds with zero point-to-point correspondences. Under these conditions, our approach improves registration quality by reducing chamfer distance by up to 90% and rotation root mean square error by up to 86% compared to the state of the art. Furthermore, the proposed SGCPR is used for one-shot shape retrieval and registration and improves retrieval precision by up to 58% over comparable methods.
Driton Salihu, Eckehard G. Steinbach
WACV2
2023 Demo: Remote Robot Control with Haptic Feedback over the Munich 5G Research Hub Testbed
Serkut Ayvasik, Edwin Babaians, Arled Papa, Yash Deshpande, Alba Jano, Wolfgang Kellerer, Eckehard G. Steinbach
WoWMoM7
2022 Contextformer: A Transformer with Spatio-Channel Attention for Context Modeling in Learned Image Compression
Ahmet Burakhan Koyuncu, Han Gao 0001, Atanas Boev, Georgii Gaikov, Elena Alshina, Eckehard G. Steinbach
ECCV (19)6
2022 Evaluation of Video Coding for Machines without Ground Truth
abstract
In the emerging field of video coding for machines, video datasets with pristine video quality and high-quality annotations are required for a comprehensive evaluation. However, existing video datasets with detailed annotations are severely limited in size and video quality. Thus, current methods have to either evaluate their codecs on still images or on already compressed data. To mitigate this problem, we propose an evaluation method based on pseudo ground-truth data from the field of semantic segmentation to the evaluation of video coding for machines. Through extensive evaluation, this paper shows that the proposed ground-truth-agnostic evaluation method results in an acceptable absolute measurement error below 0.7 percentage points on the Bjøntegaard Delta Rate compared to using the true ground truth for mid-range bitrates. We evaluate on the three tasks of semantic segmentation, instance segmentation, and object detection. Lastly, we utilize the ground-truth-agnostic method to measure the coding performances of the VVC compared against HEVC on the Cityscapes sequences. This reveals that the coding position has a significant influence on the task performance.
Kristian Fischer 0001, Markus Hofbauer, Christopher B. Kuhn, Eckehard G. Steinbach, André Kaup
ICASSP4
2022 Bounding Box Disparity: 3D Metrics for Object Detection with Full Degree of Freedom
abstract
The most popular evaluation metric for object detection in 2D images is Intersection over Union (IoU). Existing implementations of the IoU metric for 3D object detection usually neglect one or more degrees of freedom. In this paper, we first derive the analytic solution for three dimensional bounding boxes. As a second contribution, a closed-form solution of the volume-to-volume distance is derived. Finally, the Bounding Box Disparity is proposed as a combined positive continuous metric. We provide open source implementations of the three metrics as standalone python functions, as well as extensions to the Open3D library and as ROS nodes.
Michael G. Adam, Martin Piccolrovazzi, Sebastian Eger, Eckehard G. Steinbach
ICIP4
2022 Reverse Error Modeling for Improved Semantic Segmentation
abstract
We propose the concept of error-reversing autoencoders (ERA) for correcting pixel-wise errors made by an arbitrary semantic segmentation model. For this, we reframe the segmentation model as an error function applied to the ground truth labels. Then, we train an autoencoder to reverse this error function. During testing, the autoencoder reverses the approximated error function to correct the classification errors. We consider two sources of errors. First, we target the errors made by a model despite having being trained with clean, accurately labeled images. In this case, our proposed approach achieves an improvement of around 1% on the Cityscapes data set with the state-of-the-art DeepLabV3+ model. Second, we target errors introduced by compromised images. With JPEG-compressed images as input, our approach improves the segmentation performance by over 70% for high levels of compression. The proposed architecture is simple to implement, fast to train and can be applied to any semantic segmentation model as a post-processing step.
Christopher B. Kuhn, Markus Hofbauer, Goran Petrovic, Eckehard G. Steinbach
ICIP4
2022 CNN-Based Local Tone Mapping in the Perceptual Quantization Domain
abstract
A typical way of local tone mapping (TM) is based on multi-layer decomposition of the source image. For this, the source image is decomposed into a base layer and a detail layer to compress from high dynamic range (HDR) to low dynamic range (LDR). Perceptual quantization (PQ) is a standardized non-linear transfer function for HDR content. It mimics the non-linearity of human vision by compressing more strongly in bright regions and less in dark areas, when converting luminance values to electrical signals. We propose a CNN-based pipeline for local TM, which operates on the low-frequency base layer of the luminance signal, while keeping the detail layer unchanged. The proposed method works entirely in the PQ domain and is adaptable to display peak luminance. The tone-mapped LDR images obtained with our learning-based approach show significant improvements in PSNR, while the network size is reduced compared to previous work. Our experiments on the HDR datasets from Fairchild and Funt show PSNR improvements of 8 dB compared to the state-of-the-art approaches.
Hongjie You, Kai Cui 0003, Eckehard G. Steinbach
ICIP5
2022 PourNet: Robust Robotic Pouring Through Curriculum and Curiosity-based Reinforcement Learning
abstract
Pouring liquids accurately into containers is one of the most challenging tasks for robots as they are unaware of the complex fluid dynamics and the behavior of liquids when pouring. Therefore, it is not possible to formulate a generic pouring policy for real-time applications. In this paper, we propose PourNet, as a generalized solution to pouring different liquids into containers. PourNet is a hybrid planner that uses deep reinforcement learning, for end-effector planning, and Nonlinear Model Predictive Control, for joint planning. In this work, we introduce a novel simulation environment using Unity3D and NVIDIA-Flex to train our agents. By effective choice of the state space, action space and the reward functions, we allow for a direct sim-to-real transfer of the learned skills without additional training. In the simulation, PourNet outperforms state-of-the-art by an average of 4.9g deviation for water-like, and 9.2g deviation for honey-like liquids. In the real-world scenario using Kinova Movo Platform, PourNet achieves an average pouring deviation of 2.3g for dish soap when using a novel pouring container. The average pouring deviation measured for water was 5.5g. All comprehensive experiments and the simulation environment is available at: http://cxdcxd.github.io/RRS/.
Edwin Babaians, Tapan Sharma, Mojtaba Karimi, Sahand Sharifzadeh, Eckehard G. Steinbach
IROS5
2022 Skill-CPD: Real-time Skill Refinement for Shared Autonomy in Manipulator Teleoperation
abstract
Advanced wireless communication networks provide lower latency and a higher transmission rate. Although this is an enabler for many new teleoperation applications, the risk of network instability or packet drop is still unavoidable. Real-time manipulator teleoperation requires data transmission with no discontinuity. Shared autonomy (SA) is a standard method to mitigate this issue. In this way, if the data from the remote side is unavailable, the controller can continue based on the previously observed models. However, due to the spatial gap between human and robot trajectories, indisputable fluctuations occur, which cause issues in teleoperation applications. This motivates us to propose a new skill refinement strategy to modify the previously trained skill and mitigate the sudden unwanted motions within the control takeover phase. To this end, our approach comprises applying the Hidden Semi-Markov Model (HSMM) and Linear Quadratic Tracker (LQT) in combination to learn and predict the user's intentions and then exploiting Coherent Point Drift (CPD) to refine the executable trajectory. We test our method both in simulation and in the real world for 2D English letter drawing and 3D robot-assisted feeding scenarios. Our experimental results using the Kinova® Movo platform show that the proposed refinement approach generates a stable trajectory and mitigates the control switching inconsistency. All comprehensive experiments and source code is available at: http://cxdcxd.github.io/SkillCPD.
Edwin Babaians, Mojtaba Karimi, Xiao Xu 0001, Serkut Ayvasik, Eckehard G. Steinbach
IROS6
2022 Block-based Novel Haptic Data Reduction for Time-delayed Teleoperation
abstract
This work proposes a novel haptic data reduction scheme for time-delayed teleoperation by coding information as blocks. State-of-the-art (SOTA) haptic data reduction approaches are mainly sampled-based schemes. They encode haptic signals sample by sample in order to minimize the introduced coding delay. In contrast, our proposed block-based coding approach transmits a sample block as a single unit (haptic packet). Although it introduces additional algorithmic delays that are proportional to the block length, block coding has benefits since the packet rate is easy to control, the coding approach can be lossless, and the intra-block information can be employed to improve the force feedback quality. We further develop an energy adjustment approach that uses the information in a block to mitigate force oscillations caused by the Time Domain Passivity Approach. Simulation experiments and subjective tests demonstrate that our method reduces network load and significantly increases force feedback quality compared with the SOTA sample-based coding schemes, particularly for mid- to high-latency networks and low packet rates.
Ming Gui, Xiao Xu 0001, Eckehard G. Steinbach
IROS3
2022 To Sparsify or not to Sparsify: Simplifying Visual Feature Maps for Mobile Agents
abstract
Real-time pose estimation is crucial for autonomous agents for motion control and navigation, and it is essential for extended reality applications as well. As the current global positioning systems are not reliable indoors or are not precise enough, visual simultaneous localization and mapping is becoming prevalent in autonomous agents and mobile devices. The visual feature maps can be shared between agents and might be merged into a common map on a server and redistributed to all clients to enable re-localization and co-localization. However, merged maps can grow continuously in density and size that constrained agents are not able to handle anymore.In this paper, we investigate how the map density effects localization and visual odometry performance on low-cost mobile clients. We show that a sparse representation of the original VFM is necessary for real-time visual odometry but at the cost of a degraded re-localization performance.
Sebastian Eger, Rastin Pries, Gábor Sörös, Michael G. Adam, Martin Piccolrovazzi, Eckehard G. Steinbach
ISM6
2022 Measuring the Influence of Image Preprocessing on the Rate-Distortion Performance of Video Encoding
abstract
In this paper, we conduct an extensive analysis of the rate-distortion (RD) performance achieved by using different preprocessing steps before encoding the video. We propose a novel evaluation method called the Mean Saving-Cost Ratio (MSCR) to compare the RD performance for different preprocessing algorithms. We define MSCR as the logarithmic mean ratio of maximum bitrate savings over maximum quality cost for all parameters of a preprocessing algorithm.
Markus Hofbauer, Christopher B. Kuhn, Goran Petrovic, Eckehard G. Steinbach
ISM4
2022 Interactive RGB Image Segmentation via Depth-modified Click Encoding and Estimated Depth
Furkan Kaynar, Adrian Michl, Eckehard G. Steinbach
ISM3
2022 Improving Multimodal Object Detection with Individual Sensor Monitoring
abstract
Multimodal object detection fuses different sensors such as camera or LIDAR to improve the detection performance. However, individual sensor inputs can also be detrimental to a system, for example when sun glare hits a camera. In this work, we propose to monitor each sensor individually to predict when an input would lead to incorrect detections. We first train one detection network for each sensor separately, using only that sensor as input. Then, we record the performance for each single-sensor network and train an introspective performance prediction network for each sensor. Finally, we train a multimodal fusion network where we weight the impact of each sensor with its predicted performance. This allows us to dynamically adapt the fusion to reduce the influence of harmful sensor readings based only on the current data. We apply the proposed concept to the state-of-the-art AVOD architecture and evaluate on the KITTI data set. The proposed sensor monitoring system improves the mean intersection-over-union performance by 4.6%. For inputs with a low predicted performance, the proposed approach outperforms the state of the art by over 10%, demonstrating the potential of using individual sensor monitoring to react to problematic input. The proposed approach can be applied to any fusion network with two or more sensors and could also be used for classification or segmentation tasks.
Christopher B. Kuhn, Markus Hofbauer, Goran Petrovic, Eckehard G. Steinbach
ISM5
2022 Self-Supervised Object Recognition Based on Repeated Re-Capturing of Dynamic Indoor Environments
abstract
Capturing a digital replica of an environment using hand held devices or mobile mapping systems has become increasingly easy in recent years. However, leveraging large amounts of data for various semantic applications is usually impaired by a costly data annotation process. In this paper, we investigate the topic of object recognition in indoor environments without supervision. We approach the problem from a remapping perspective, where we capture RGB images from the same environment at different times and use the naturally occurring changes to identify single objects. In the first step, we create pairs of images from different recordings and generate object candidates using optical flow and an off-the-shelf region proposal algorithm. Then, we use a self-supervised representation learning framework and cluster the extracted objects. We evaluate the performance of several existing clustering methods in an over-clustering setting, since the number of object classes is unknown in an unsupervised setup. Our experimental validation on a real-world dataset shows that the proposed system can successfully recognize objects and pre-annotate a dataset by exploiting a recapturing process.
Martin Piccolrovazzi, Michael G. Adam, Sebastian Eger, Marsil Zakour, Eckehard G. Steinbach
ISM5
2022 S2CMAF: Multi-Method Assessment Fusion for Scan-to-CAD Methods
abstract
Scan-to-CAD-based 3D reconstruction of indoor environments has become increasingly more popular in recent years. The inherent structure of Scan-to-CAD consists of object detection, model retrieval, and alignment. Therefore, a variety of metrics are required to assess these three aspects. This can lead to ambiguous evaluation results and incorrect quality assumptions. To impede the problem of incorrect evaluation, we introduce S2CMAF, a multi-method assessment fusion approach for Scan-to-CAD pipelines. S2CMAF merges several metrics used in evaluating these pipelines into one unique quality score. We show that S2CMAF significantly improves the correlation between Scan-to-CAD results and the ground truth, compared to the conventionally used Scan2CAD benchmark. Additionally, we train S2CMAF using different optimization techniques and demonstrate the advantages of our approach on real-world data.
Driton Salihu, Adam Misik, Markus Hofbauer, Eckehard G. Steinbach
ISM4
2022 Robust Depth Estimation in Foggy Environments Combining RGB Images and mmWave Radar
abstract
In this paper, we propose a robust depth estimation strategy that uses RGB images and mmWave radar data to deal with limited visibility in foggy environments. While the state-of-the-art RGB or LiDAR-based depth estimation works well in scenarios with good visibility, their performance dramatically degrades in the presence of fog. In contrast, mmWave radar sensors are not affected by fog and hence are a promising complement. To leverage this property of mmWave radar, we combine RGB image-based depth estimation with radar information. The proposed combination is an extension of the Sparse-to-Dense (S2D) model. Moreover, a weight-based sensor fusion strategy is presented to improve system performance. Our experiments show that a fog density of meteorological optical range (MOR) less than 50m leads to strongly degraded performance for RGB image-based and LiDAR-based depth estimation. For a MOR of 30m in our dataset, the experiments show an improvement of 26% in mean square error for our proposed approach compared to the combination of RGB images and LiDAR data.
Mengchen Xiong, Xiao Xu 0001, Eckehard G. Steinbach
ISM4
2022 Towards Subjective Experience Prediction for Time-Delayed Teleoperation with Haptic Data Reduction
abstract
This paper presents a novel quality assessment approach for the prediction of the subjective haptic experience in time-delayed teleoperation. With the rapid development of haptic technology in remote robot control and virtual reality, new control schemes and hardware systems are developed to provide high quality human-in-the-loop teleoperation service. Our subjective experiments indicate that the existing objective quality assessment metrics do not sufficiently correlate with the subjective haptic experience of the users. This gap requires expensive and time-consuming subjective experiments to be conducted. To avoid time-consuming experiments and provide a fast and accurate subjective experience prediction, we make an attempt to analyze and explain the mismatch between the subjective and objective haptic signal quality metrics. To this end, extensive subjective experiments and case studies have been conducted for teleoperation with time delay and haptic data reduction. Based on our experimental results, we propose a quality assessment approach that predicts the subjective quality of experience using multiple objective metrics. For the one-dimensional spring model, the Spearman’s rank-order, Kendall’s rank-order and Pearson’s Linearity correlation coefficient (SROCC, KLOCC and PLCC) between the predictions of our model and the results of subjective experiment show remarkable improvement on the correlation between subjective and objective quality assessment.
Zican Wang, Fei Mei, Xiao Xu 0001, Eckehard G. Steinbach
RO-MAN4
2022 Traffic-Aware Multi-View Video Stream Adaptation for Teleoperated Driving
abstract
Remote control of an autonomous vehicle by a human operator requires low delay video transmission to resolve complex situations and ensure safety. The remote operator perceives the current traffic scenario via video streams from multiple cameras. To provide the operator with the best possible scene understanding while matching the available network resources, the video streams need to be automatically adapted. In this paper, we propose a traffic-aware multi-view video stream adaptation scheme. We estimate the importance of each camera view based on the vehicle’s real-time movement in traffic. The resulting prioritization together with the total available transmission rate determines a specific bit-budget for each camera view. We optimize the video quality of each individual video stream for the given bit-budget using a quality-of-experience-driven multi-dimensional adaptation scheme. Additionally, we apply a region-of-interest mask to the rear-facing camera views. The mask removes less important areas from the image which reduces the required bitrate. All modules are implemented to extend the existing TELECARLA framework. We evaluate the proposed traffic-aware adaptation scheme in a user study. We observe a high correlation between the proposed view prioritization module and the subjective ratings obtained in the user study. The region-of-interest masking achieves Bjøntegaard Delta Rate savings of at least 19.8% compared to streaming the full camera view. The overall system improves the VMAF score by 1.86 per camera when considering the importance of the individual camera views as rated by the users. This demonstrates the potential of an individual adaptation for each camera view optimized for the current traffic situation.
Markus Hofbauer, Christopher B. Kuhn, Mariem Khlifi, Goran Petrovic, Eckehard G. Steinbach
VTC Spring5
2022 Preprocessor Rate Control for Adaptive Multi-View Live Video Streaming Using a Single Encoder
abstract
Currently, an increasing number of technical systems are equipped with multiple cameras. Limited by cost and size, they are often restricted to a single hardware encoder. The combination of all views into a single superframe allows for streaming all camera views at the same time, but it prevents individual rate/quality adaptations on those camera views. We propose a preprocessing filter concept that allows for individual rate/quality adaptation while using a single encoder. Additionally, we create a preprocessor model that estimates the required preprocessing filter parameters from the specified encoding parameters. This means our approach can be used with any existing multi-view adaptation scheme designed for controlling multiple encoders. We design both an analytical and a Machine Learning-based bitrate model. Because both models perform equally well, we suggest using either one as the core part of our preprocessor model. Both models are specifically designed for estimating the influence of the quantization parameter, frame rate, frame size, group of pictures length, and a Gaussian low-pass filter on the video bitrate. Furthermore, the rate models outperform state-of-the-art bitrate models by at least 22% regarding the overall root mean square error. Our bitrate models are the first of their kind to consider the influence of a Gaussian low-pass filter. We evaluate the preprocessing approach by streaming six camera views in a teledriving scenario with a single encoder and compare it to using six individual encoders. The experimental results demonstrate that the preprocessing approach achieves bitrates similar to the individual encoders for all views. While achieving a comparable rate and quality for the most important views, our approach requires a total bitrate that is 50% smaller than when using a single encoder approach without preprocessing.
Markus Hofbauer, Christopher B. Kuhn, Goran Petrovic, Eckehard G. Steinbach
IEEE Trans. Circuits Syst. Video Technol.4
2022 Introspective Failure Prediction for Autonomous Driving Using Late Fusion of State and Camera Information
abstract
We present an introspective failure prediction approach for autonomous vehicles. In autonomous driving, complex or unknown scenarios can cause a disengagement of the self-driving system. Disengagements can be triggered either by automatic safety measures or by human intervention. We propose to use recorded disengagement sequences from test drives as training data to learn to predict future failures. The system then learns introspectively from its own previous mistakes. In order to predict failures as early as possible, we propose a machine learning approach where sequences of sensor data are classified as either failure or success. The car itself is treated as a black box. Our method combines two sensor modalities that contain different types of information. An image-based model learns to detect generally challenging situations such as crowded intersections accurately multiple seconds in advance. A state data based model allows to detect fast changes immediately before a failure, such as sudden braking or swerving. The outcome of the individual models is fused by averaging the individual failure probabilities. We evaluate our approach on a data set provided by the BMW Group containing 14 hours of autonomous driving. The proposed late fusion approach allows for predicting failures at an accuracy of more than 85% seven seconds in advance, at a false positive rate of 20%. The proposed method outperforms state-of-the-art failure prediction by more than 15% while being a flexible framework that allows for straightforward addition of further sensor modalities.
Christopher B. Kuhn, Markus Hofbauer, Goran Petrovic, Eckehard G. Steinbach
IEEE Trans. Intell. Transp. Syst.4
2021 Pixel-Wise Failure Prediction For Semantic Video Segmentation
abstract
We propose a pixel-accurate failure prediction approach for semantic video segmentation. The proposed scheme improves previously proposed failure prediction methods which so far disregarded the temporal information in videos. Our approach consists of two main steps: First, we train an LSTM-based model to detect spatio-temporal patterns that indicate pixel-wise misclassifications in the current video frame. Second, we use sequences of failure predictions to train a denoising autoencoder that both refines the current failure prediction and predicts future misclassifications. Since public data sets for this scenario are limited, we introduce the large-scale densely annotated video driving (DAVID) data set generated using the CARLA simulator. We evaluate our approach on the real-world Cityscapes data set and the simulator-based DAVID data set. Our experimental results show that spatiotemporal failure prediction outperforms single-image failure prediction by up to 8.8%. Refining the prediction using a sequence of previous failure predictions further improves the performance by a significant 15.2% and allows to accurately predict misclassifications for future frames. While we focus our study on driving videos, the proposed approach is general and can be easily used in other scenarios as well.
Christopher B. Kuhn, Markus Hofbauer, Ziqin Xu, Goran Petrovic, Eckehard G. Steinbach
ICIP5
2021 NMPC-MP: Real-time Nonlinear Model Predictive Control for Safe Motion Planning in Manipulator Teleoperation
abstract
Motion control and planning for the manipulator are critical components in manipulator teleoperation. Online (real-time) motion control is challenging for active obstacle avoidance and often results in fluctuating and unsafe motion. Offline motion planning, on the other hand, generates precise and secure trajectories for complex manipulation. In this paper, a real-time nonlinear model predictive control based motion planner (NMPC-MP) is designed for teleoperated manipulation. In contrast to traditional NMPC-based approaches, our model considers a complex environment with dynamic obstacles. Our multi-threaded NMPC-MP allows for real-time planning, including dynamic objects. We evaluate our approach both in a simulated environment and with real-world experiments using the Kinova®Movo platform. The comparison to state-of-the-art approaches (e.g., RRT-Connect, CHOMP, and STOMP) shows a significant improvement in real-time motion planning using NMPC-MP. In real-world tests, the proposed planner was applied on a human-shaped dual manipulator setup. Our results show that the NMPC-MP runs in real-time and generates smooth and reliable trajectories. The experiments validate that the planner is able to precisely track active goals from the teleoperator while avoiding self-collision and obstacles.
Siqi Hu, Edwin Babaians, Mojtaba Karimi, Eckehard G. Steinbach
IROS4
2021 QoE-driven Delay-adaptive Control Scheme Switching for Time-delayed Bilateral Teleoperation with Haptic Data Reduction
abstract
Teleoperation systems with haptic feedback allow a human user to remotely interact with a dangerous or inac-cessible environment, perform various tasks, and perceive the haptic feedback. To ensure system stability while maintaining the best possible quality of experience (QoE), different teleoperation control schemes and haptic communication strategies need to be selected to adapt to varying network conditions and teleoperation tasks. In this paper, we propose a QoE-driven control scheme switching approach, which adaptively selects the control scheme that provides the best possible QoE for varying communication delay. A transition period is designed to moderate the artifacts during the switching phase. Haptic data reduction approaches are developed for the switching strategy to match the characteristics of each control scheme. Our experiments verify the feasibility of the proposed scheme. Subjective tests confirm that the proposed adaptive switching scheme is able to achieve a superior user QoE in contrast to a fixed control scheme in the presence of varying communication delay up to 200 ms.
Xiao Xu 0001, Qian Liu 0001, Eckehard G. Steinbach
IROS4
2021 Quality-Blind Compressed Color Image Enhancement with Convolutional Neural Networks
abstract
Lossy compressed images and videos suffer from visible compression artifacts, especially when the bit-rate is low. To improve the quality of the compressed image while keeping the same bit-rate, decoder-side compression artifacts reduction (CAR) becomes important. Recently, convolutional neural networks are adopted for CAR tasks and achieve the state-of-the-art performance. However, most CAR algorithms only focus on the reconstruction of the luminance channel. Also, a separate model usually needs to be trained for each quality factor (QF), which makes these approaches not practical in existing codecs. In this paper, we analyze a quality-blind training strategy and compare it with training separate models for each QF. The testing results with three representative CAR algorithms show the superiority of the quality-blind training compared to separate training. The results for pseudo and real quality-blind CAR tests further prove the generalizability of the quality-blind training for practical CAR tasks.
Kai Cui 0003, Ahmet Burakhan Koyuncu, Atanas Boev, Elena Alshina, Eckehard G. Steinbach
ISCAS5
2021 Trajectory-Based Failure Prediction for Autonomous Driving
abstract
In autonomous driving, complex traffic scenarios can cause situations that require human supervision to resolve safely. Instead of only reacting to such events, it is desirable to predict them early in advance. While predicting the future is challenging, there is a source of information about the future readily available in autonomous driving: the planned trajectory the car intends to drive. In this paper, we propose to analyze the trajectories planned by the vehicle to predict failures early on. We consider sequences of trajectories and use machine learning to detect patterns that indicate impending failures. Since no public data of disengagements of autonomous vehicles is available, we use data provided by development vehicles of the BMW Group. From over six months of test drives, we obtain more than 2600 disengagements of the automated system. We train a Long Short-Term Memory classifier with sequences of planned trajectories that either resulted in successful driving or disengagements. The proposed approach outperforms existing state-of-the-art failure prediction with low-dimensional data by more than 3 % in a Receiver Operating Characteristic analysis. Since our approach makes no assumptions on the underlying system, it can be applied to predict failures in other safety-critical areas of robotics as well.
Christopher B. Kuhn, Markus Hofbauer, Goran Petrovic, Eckehard G. Steinbach
IV4
2021 Convolutional neural network-based post-filtering for compressed YUV420 images and video
abstract
Images and videos compressed with lossy compression algorithms usually suffer from visible distortions, especially when the bitrate is low. To improve the quality without spending extra bitrate, many image and video codecs have built-in filters to mitigate these artifacts. However, most of them are only applied on the luminance channel, while the chrominance channels remain unmodified. While this is partly justified by the observation that the luminance channel usually contains more details and has higher-resolution than the chrominance channels. We observe that the luminance and chrominance channels still have latent correlations. Therefore, the post-filtering of the chrominance channels is also beneficial and can be driven by the information from the luminance channel. In this paper, we propose a 3-stage YUV post-filtering network for compressed YUV420 images and video. The proposed 3-stage structure not only improves the quality of the luminance channel, but also exploits the luma-chroma correlations to improve the quality of the chrominance channels. Our experimental results show that the proposed approach achieves 3.60%/12.75%/14.93% Bj⊘ntegaard Delta bitrate improvement for the Y, U and V channels over the VVC 10.0 codec for All-Intra configuration.
Kai Cui 0003, Ahmet Burakhan Koyuncu, Atanas Boev, Elena Alshina, Eckehard G. Steinbach
PCS5
2021 Parallelized Context Modeling for Faster Image Coding
abstract
Learning-based image compression has reached the performance of classical methods such as BPG. One common approach is to use an autoencoder network to map the pixel information to a latent space and then approximate the symbol probabilities in that space with a context model. During inference, the learned context model provides symbol probabilities, which are used by the entropy encoder to obtain the bitstream. Currently, the most effective context models use autoregression, but autoregression results in a very high decoding complexity due to the serialized data processing. In this work, we propose a method to parallelize the autoregressive process used for image compression. In our experiments, we achieve a decoding speed that is over 8 times faster than the standard autoregressive context model almost without compression performance reduction.
Ahmet Burakhan Koyuncu, Kai Cui 0003, Atanas Boev, Eckehard G. Steinbach
VCIP4
2021 Decoder-Side Motion Vector Refinement in VVC: Algorithm and Hardware Implementation Considerations
abstract
This paper presents an overview of the decoder-side motion vector refinement (DMVR) algorithm in the Versatile Video Coding (VVC) standard. The proposed DMVR algorithm aims to increase the prediction accuracy of the blocks coded in merge mode using the bilateral matching-based refinement method. Compared with previous decoder-side motion vector derivation approaches, the proposed method significantly increases the coding efficiency without signaling additional side information. Furthermore, the hardware implementation considerations of the DMVR design are particularly focused in this study. This paper details and analyzes the novel features of DMVR contributing to the increase in coding efficiency and the reduction in computational complexity and implementation difficulty. Experimental results based on the VVC test model version 8.0 demonstrate that average Bjøntegaard Delta rate savings of 0.80 % and 2.81 % are achieved for the “tool-off” and “tool-on” test configurations, respectively. Moreover, 4 % additional decoding time and negligible additional external memory bandwidth requirements of DMVR based on the common test conditions for VVC are reported.
Han Gao 0001, Semih Esenlik, Jianle Chen, Eckehard G. Steinbach
IEEE Trans. Circuits Syst. Video Technol.5
2021 Geometric Partitioning Mode in Versatile Video Coding: Algorithm Review and Analysis
abstract
This paper presents an overview of the geometric partitioning mode (GPM) algorithm that is a part of the most recent Versatile Video Coding (VVC) standard. The GPM algorithm aims to increase the partitioning precision of moving objects using non-rectangular and asymmetric rectangular partitions on top of the conventional rectangular block partitioning structure of VVC. Novel features of GPM contributing to the increase in coding efficiency and the reduction in encoder and decoder complexity are detailed and analyzed in this paper. Evaluated with VVC test model version 8.0 under the joint video experts team common test conditions, experimental results show that the presented GPM algorithm provides luma Bjøntegaard Delta rate reduction of 0.70% for random access and of 1.55% for low delay with B slices configurations, with roughly 3% to 5% additional encoding time and negligible decoder runtime change. Furthermore, as GPM provides more precise partitions for the boundaries of the moving objects, an improvement of visual quality is seen in GPM coded sequences.
Han Gao 0001, Semih Esenlik, Elena Alshina, Eckehard G. Steinbach
IEEE Trans. Circuits Syst. Video Technol.4
2021 Dual-Stream Multi-Path Recursive Residual Network for JPEG Image Compression Artifacts Reduction
abstract
JPEG is the most widely used lossy image compression standard. When using JPEG with high compression ratios, visual artifacts cannot be avoided. These artifacts not only degrade the user experience but also negatively affect many low-level image processing tasks. Recently, convolutional neural network (CNN)-based compression artifact removal approaches have achieved significant success, however, at the cost of high computational complexity due to an enormous number of parameters. To address this issue, we propose a dual-stream recursive residual network (STRRN) which consists of structure and texture streams for separately reducing the specific artifacts related to high-frequency or low-frequency image components. The outputs of these streams are combined and fed into an aggregation network to further enhance the restored images. By using parameter sharing, the proposed network reduces the total number of training parameters significantly. Moreover, experiments conducted on five commonly used datasets confirm that the proposed STRRN can efficiently reduce the compression artifacts, while using up to 4.6 times less training parameters and 5 times less running time compared to the state-of-the-art approaches.
Zhi Jin 0002, Wenbin Zou, Xia Li 0006, Eckehard G. Steinbach
IEEE Trans. Circuits Syst. Video Technol.5
2021 PVC-SLP: Perceptual Vibrotactile-Signal Compression Based-on Sparse Linear Prediction
abstract
Developing a signal compression technique that is able to achieve a low bit rate while maintaining high perceptual signal quality is a classical signal processing problem vigorously studied for audio, speech, image, and video type of signals. Yet, until recently, there has been limited effort directed toward the compression of vibrotactile signals, which represent a crucial element of rich touch (haptic) information. A vibrotactile signal; produced when stroking a textured surface with a tool-tip or bare-finger; like other signals contains a great deal of redundant and imperceptible information that can be exploited for efficient compression. This paper presents PVC-SLP, a vibrotactile perceptual coding approach. PVC-SLP employs a model of tactile sensitivity; called ASF (Acceleration Sensitivity Function); for perceptual coding. The ASF is inspired by the four channels model that mediate the perception of vibrotactile stimuli in the glabrous skin. The compression algorithm introduces sparsity constraints in a linear prediction scheme both on the residual and the predictor coefficients. The perceptual quantization of the residual is developed through the use of ASF. The quantization parameters of the residual and the predictor coefficients were jointly optimized; by means of both squared error and perceptual quality measures; to find the sweet spot of the rate-distortion curve. PVC-SLP coding performance is evaluated using two publicly available databases that collectively comprise 1281 vibrotactile signals covering 193 material classes. Furthermore, we compare PVC-SLP with a recent vibrotactile compression method and show that PVC-SLP perceptually outperforms existing method by a sizable margin. Most recently, PVC-SLP has been selected to become part of the haptic codec standard currently under preparation by IEEE P1918.1.1, aka Haptic Codecs for the Tactile Internet.
Rania Hassen, Basak Güleçyüz, Eckehard G. Steinbach
IEEE Trans. Multim.3
2021 6DLS: Modeling Nonplanar Frictional Surface Contacts for Grasping Using 6-D Limit Surfaces
abstract
Robot grasping with deformable gripper jaws results in nonplanar surface contacts if the jaws deform to the nonplanar local geometry of an object. The frictional force and torque that can be transmitted through a nonplanar surface contact are both 3-D, resulting in a 6-D frictional wrench (6DFW). Applying traditional planar contact models to such contacts leads to overconservative results as the models do not consider the nonplanar surface geometry and only compute a 3-D subset of the 6DFW. To address this issue, we derive the 6DFW for nonplanar surfaces by combining concepts of differential geometry and Coulomb friction. We also propose two 6-D limit surface (6DLS) models, generalized from well-known 3-D LS (3DLS) models, which describe the friction-motion constraints for a contact. We evaluate the 6DLS models by fitting them to the 6DFW samples obtained from six parametric surfaces and 2932 meshed contacts from finite element method simulations of 24 rigid objects. We further present an algorithm to predict multicontact grasp success by building a grasp wrench space with the 6DLS model of each contact. To evaluate the algorithm, we collected 1035 physical grasps of ten 3-D-printed objects with a KUKA robot and a deformable parallel-jaw gripper. In our experiments, the algorithm achieves 66.8% precision, a metric inversely related to false positive predictions, and 76.9% recall, a metric inversely related to false negative predictions. The 6DLS models increase recall by up to 26.1% over 3DLS models with similar precision.1
Tamay Aykut, Daolin Ma, Eckehard G. Steinbach
IEEE Trans. Robotics4
2020 On the Quality-of-Learning for Haptic Teleoperation-based Skill Transfer over the Tactile Internet
abstract
Transfer of skills and teaching tasks to robots face new challenges when the demonstrations are provided remotely via tele-operation. Not only having a remote operator, but also the communication between the tele-operator and the operator affects the quality of demonstrations. Artifacts introduced by lossy haptic data compression and communication delay deteriorate the system transparency; however, the impact of these on the quality of learning has not been studied yet. In this paper, we construct the bridge between the learning quality and the reduced transparency caused by lossy haptic data compression during teleoperation with haptic feedback. The considered haptic data compression scheme is the previously proposed perceptual dead band-based kinesthetic data reduction approach. The learning quality is assessed both with the mean squared error (MSE) metric on the trajectory level and with the rate of success defined on the task requirement. Our experiments show that the learning quality is reduced significantly for a dead band parameter larger than 20% and 30% for a cube following and peg-in-hole tasks, respectively.
Basak Güleçyüz, Xiao Xu 0001, Andreas Noll, Eckehard G. Steinbach
GLOBECOM4
2020 Minimal Work: A Grasp Quality Metric for Deformable Hollow Objects
abstract
Robot grasping of deformable hollow objects such as plastic bottles and cups is challenging, as the grasp should resist disturbances while minimally deforming the object so as not to damage it or dislodge liquids. We propose minimal work as a novel grasp quality metric that combines wrench resistance and object deformation. We introduce an efficient algorithm to compute the work required to resist an external wrench for a manipulation task by solving a linear program. The algorithm first computes the minimum required grasp force and an estimation of the gripper jaw displacements based on the object's empirical stiffness at different locations. The work done by the jaws is the product of the grasp force and the displacements. Grasps requiring minimal work are considered to be of high quality. We collect 460 physical grasps with a UR5 robot and a Robotiq gripper. We consider a grasp to be successful if it completes the task without damaging the object or dislodging the content. Physical experiments suggest that the minimal work quality metric reaches 74.2% balanced accuracy, a metric that is the raw accuracy normalized by the number of successful and failed real-world grasps, and is up to 24.2% higher than classical wrench-based quality metrics.
Michael Danielczuk, Jeffrey Ichnowski, Jeffrey Mahler, Eckehard G. Steinbach, Kenneth Y. Goldberg
ICRA5
2020 6DFC: Efficiently Planning Soft Non-Planar Area Contact Grasps using 6D Friction Cones
abstract
Analytic grasp planning algorithms typically approximate compliant contacts with soft point contact models to compute grasp quality, but these models are overly conservative and do not capture the full range of grasps available. While area contact models can reduce the number of false negatives predicted by point contact models, they have been restricted to a 3D analysis of the wrench applied at the contact and so are still overly conservative. We extend traditional 3D friction cones and present an efficient algorithm for calculating the 6D friction cone (6DFC) for a non-planar area contact between a compliant gripper and a rigid object. We introduce a novel sampling algorithm to find the 6D friction limit surface for a non-planar area contact and a linearization method for these ellipsoids that reduces the computation of 6DFC constraints to a quadratic program. We show that constraining the wrench applied at the contact in this way increases recall, a metric inversely related to the number of false negative predictions, by 17% and precision, a metric inversely related to the number of false positive predictions, by 2% over soft point contact models on results from 1500 physical grasps on 12 3D printed nonplanar objects with an ABB YuMi robot. The 6DFC algorithm also achieves 6% higher recall with similar precision and 85x faster runtime than a previously proposed area contact model.
Michael Danielczuk, Eckehard G. Steinbach, Kenneth Y. Goldberg
ICRA3
2020 Measuring Driver Situation Awareness Using Region-of-Interest Prediction and Eye Tracking
abstract
With increasing progress in autonomous driving, the human does not have to be in control of the vehicle for the entire drive. A human driver obtains the control of the vehicle in case of an autonomous system failure or when the vehicle encounters an unknown traffic situation it cannot handle on its own. A critical part of this transition to human control is to ensure a sufficient driver situation awareness. Currently, no direct method to explicitly estimate driver awareness exists. In this paper, we propose a novel system to explicitly measure the situation awareness of the driver. Our approach is inspired by methods used in aviation. However, in contrast to aviation, the situation awareness in driving is determined by the detection and understanding of dynamically changing and previously unknown situation elements. Our approach uses machine learning to define the best possible situation awareness. We also propose to measure the actual situation awareness of the driver using eye tracking. Comparing the actual awareness to the target awareness allows us to accurately assess the awareness the driver has of the current traffic situation. To test our approach, we conducted a user study. We measured the situation awareness score of our model for 8 unique traffic scenarios. The results experimentally validate the accuracy of the proposed driver awareness model.
Markus Hofbauer, Christopher B. Kuhn, Lukas Püttner, Goran Petrovic, Eckehard G. Steinbach
ISM5
2020 Adaptive Multi-View Live Video Streaming for Teledriving Using a Single Hardware Encoder
abstract
Teleoperated driving (TOD) is a possible solution to cope with failures of autonomous vehicles. In TOD, the human operator perceives the traffic situation via video streams of multiple cameras from a remote location. Adaptation mechanisms are needed in order to match the available transmission resources and provide the operator with the best possible situation awareness. This includes the adjustment of individual camera video streams according to the current traffic situation. The limited video encoding hardware in vehicles requires the combination of individual camera frames into a larger superframe video. While this enables the encoding of multiple camera views with a single encoder, it does not allow for rate/quality adaptation of the individual views. To this end, we propose a novel concept that uses preprocessing filters to enable individual rate/quality adaptations in the superframe video. The proposed preprocessing filters allow for the usage of existing multidimensional adaptation models in the same way as for individual video streams using multiple encoders. Our experiments confirm that the proposed concept is able to control the spatial, temporal and quality resolution of individual segments in the superframe video. Additionally, we demonstrate the usability of the proposed method by applying it in a multi-view teledriving scenario. We compare our approach to individually encoded video streams and a multiplexing solution without preprocessing. The results show that the proposed approach produces bitrates for the individual video streams which are comparable to the bitrates achieved with separate encoders. While achieving a similar bitrate for the most important views, our approach requires a total bitrate that is 40% smaller compared to the multiplexing approach without preprocessing.
Markus Hofbauer, Christopher B. Kuhn, Goran Petrovic, Eckehard G. Steinbach
ISM4
2020 Better Look Twice - Improving Visual Scene Perception Using a Two-Stage Approach
abstract
Accurate visual scene perception plays an important role in fields such as medical imaging or autonomous driving. Recent advances in computer vision allow for accurate image classification, object detection and even pixel-wise semantic segmentation. Human vision has repeatedly been used as an inspiration for developing new machine vision approaches. In this work, we propose to adapt the “zoom lens model” from psychology for semantic scene segmentation. According to this model, humans first distribute their attention evenly across the entire field of view at low processing power. Then, they follow visual cues to look at a few smaller areas with increased attention. By looking twice, it is possible to refine the initial scene understanding without requiring additional input. We propose to perform semantic segmentation the same way. To obtain visual cues for deciding where to look twice, we use a failure region prediction approach based on a state-of-the-art failure prediction method. Then, the second, focused look is performed by a dedicated classifier that reclassifies the most challenging patches. Finally, pixels predicted to be errors are updated in the original semantic prediction. While focusing only on areas with the highest predicted failure probability, we achieve a classification accuracy of over 63% for the predicted failure regions. After updating the initial semantic prediction of 4000 test images from a large-scale driving data set, we reduce the absolute pixel-wise error of 232 road participants by 10% or more.
Christopher B. Kuhn, Markus Hofbauer, Goran Petrovic, Eckehard G. Steinbach
ISM4
2020 TELECARLA: An Open Source Extension of the CARLA Simulator for Teleoperated Driving Research Using Off-the-Shelf Components
abstract
Teledriving is a possible fallback mode to cope with failures of fully autonomous vehicles. One important requirement for teleoperated vehicles is a reliable low delay data transmission solution, which adapts to the current network conditions to provide the operator with the best possible situation awareness. Currently, there is no easily accessible solution for the evaluation of such systems and algorithms in a fully controllable environment available. To this end we propose an open source framework for teleoperated driving research using low-cost off-the-shelf components. The proposed system is an extension of the open source simulator CARLA, which is responsible for rendering the driving environment and providing reproducible scenario evaluation. As a proof of concept, we evaluated our teledriving solution against CARLA in remote and local driving scenarios. The proposed teledriving system leads to almost identical performance measurements for local and remote driving. In contrast, remote driving using CARLA's client server communication results in drastically reduced operator performance. Further, the framework provides an interface for the adaptation of the temporal resolution and target bitrate of the compressed video streams. The proposed framework reduces the required setup effort for teleoperated driving research in academia and industry.
Markus Hofbauer, Christopher B. Kuhn, Goran Petrovic, Eckehard G. Steinbach
IV4
2020 Introspective Black Box Failure Prediction for Autonomous Driving
abstract
Failures in autonomous driving caused by complex traffic situations or model inaccuracies remain inevitable in the near future. While much research is focused on how to prevent such failures, comparatively little research has been done on predicting them. An early failure prediction would allow for more time to take actions to resolve challenging situations. In this work, we propose an introspective approach to predict future disengagements of the car by learning from previous disengagement sequences. Our method is designed to detect failures as early as possible by using sensor data from up to ten seconds before each disengagement. The car itself is treated as a black box, with only its state data and the number of detected objects being required. Since no model-specific knowledge is needed, our method is applicable to any self-driving system. Currently, no public data of real-life disengagements is available. To test our approach, we therefore use autonomous driving data provided by BMW that was collected with BMW research vehicles over multiple months. We show that an LSTM classifier trained with sequences of state data can predict failures up to seven seconds in advance with an accuracy of more than 80%. This is two seconds earlier than comparable approaches from the literature.
Christopher B. Kuhn, Markus Hofbauer, Goran Petrovic, Eckehard G. Steinbach
IV4
2020 Evaluation of Different Task Distributions for Edge Cloud-based Collaborative Visual SLAM
abstract
In recent years, a variety of visual SLAM (Simultaneous Localization and Mapping) systems have been proposed. These systems allow camera-equipped agents to create a map of the environment and determine their position within this map, even without an available GNSS signal. Visual SLAM algorithms differ mainly in the way the image information is processed and whether the resulting map is represented as a dense point cloud or with sparse feature points. However, most systems have in common that a high computational effort is necessary to create an accurate, correct and up-to-date pose and map. This is a challenge for smaller mobile agents with limited power and computing resources. In this paper, we investigate how the processing steps of a state-of-the-art feature-based visual SLAM system can be distributed among a mobile agent and an edge-cloud server. Depending on the specification of the agent, it can run the complete system locally, offload only the tracking and optimization part, or run nearly all processing steps on the server. For this purpose, the individual processing steps and their resulting data formats are examined and methods are presented how the data can be efficiently transmitted to the server. Our experimental evaluation shows that the CPU load can be reduced for all task distributions which offload part of the pipeline to the server. For agents with low computing power, the processing time for the pose estimation can even be reduced. In addition, the higher computing power of the server allows to increase the frame rate and accuracy for pose estimation.
Sebastian Eger, Rastin Pries, Eckehard G. Steinbach
MMSP3
2020 Fuzzy logic and histogram of normal orientation-based 3D keypoint detection for point clouds
Dmytro Bobkov, Eckehard G. Steinbach
Pattern Recognit. Lett.3
2020 A Flexible Deep CNN Framework for Image Restoration
abstract
Image restoration is a long-standing problem in image processing and low-level computer vision. Recently, discriminative convolutional neural network (CNN)-based approaches have attracted considerable attention due to their superior performance. However, most of these frameworks are designed for one specific image restoration task; hence, they seldom show high performance on other image restoration tasks. To address this issue, we propose a flexible deep CNN framework that exploits the frequency characteristics of different types of artifacts. Hence, the same approach can be employed for a variety of image restoration tasks by adjusting the architecture. For reducing the artifacts with similar frequency characteristics, a quality enhancement network that adopts residual and recursive learning is proposed. Residual learning is utilized to speed up the training process and boost the performance; recursive learning is adopted to significantly reduce the number of training parameters as well as boost the performance. Moreover, lateral connections transmit the extracted features between different frequency streams via multiple paths. One aggregation network combines the outputs of these streams to further enhance the restored images. We demonstrate the capabilities of the proposed framework with three representative applications: image compression artifacts reduction (CAR), image denoising, and single image super-resolution (SISR). Extensive experiments confirm that the proposed framework outperforms the state-of-the-art approaches on benchmark datasets for these applications.
Zhi Jin 0002, Dmytro Bobkov, Wenbin Zou, Xia Li 0006, Eckehard G. Steinbach
IEEE Trans. Multim.6
2019 Efficient Panorama Database Indexing for Indoor Localization
abstract
We consider the task of indoor localization in large-scale environments using visual search on a database of geo-tagged panoramas. In this work we propose an efficient way to represent the database so as to maximize the search accuracy while minimizing the amount of computation required per query. The success of our method is due to a combination of (i) a hierarchical indexing method based on panorama image region information, and (ii) image descriptors aggregated from multiple views sampled finely over the panorama using generalized max pooling (GMP). Experiments on a large indoor dataset show that the complexity is reduced compared to common state-of-the-art retrieval methods such as FLANN (Fast Library for Approximate Nearest Neighbors): our scheme is more than twice as fast as an index based on FLANN while maintaining a similar retrieval performance.
Jean-Baptiste Boin, Dmytro Bobkov, Eckehard G. Steinbach, Bernd Girod
CBMI3
2019 3D Reconstruction of Indoor Geometry using Electromagnetic Multipath Fingerprints
abstract
Knowing the location and the dimensions of walls and objects in an indoor environment is extremely important for both human and robot navigation. Precise blueprints are often not available, and constructing a map of the indoor geometry by hand may not be feasible. Fingerprinting-based localization systems rely on the data received by the user from the access points (APs). Before a system can be deployed, certain properties of the received electromagnetic signal have to be measured at a number of locations throughout the indoor environment. The approach presented in this paper uses the resulting fingerprint map to reconstruct the 3D indoor geometry. The novel approach first calculates the positions of the virtual transmitters (VTs), the reflections of the APs, and then uses the VT positions to reconstruct the indoor geometry. The scheme is designed for multipath-based fingerprints. Several metrics for evaluating the performance of an indoor reconstruction algorithm are also derived. The performance of the proposed approach is validated through simulation. The proposed algorithm is shown to correctly reconstruct up to 66% of the 3D indoor geometry and up to 81% of its 2D perimeter.
Alexandra Zayets, Mohamed Bourguiba, Eckehard G. Steinbach
ICC3
2019 Adaptive Fusion-Based 3D Keypoint Detection for RGB Point Clouds
abstract
We propose a novel keypoint detector for 3D RGB Point Clouds (PCs). The proposed keypoint detector exploits both the 3D structure and the RGB information of the PC data. Keypoint candidates are generated by computing the eigenvalues of the covariance matrix of the PC structure information. Additionally, from the RGB information, we estimate the salient points by an efficient adaptive difference of Gaussian-based operator. Finally, we fuse the resulting two sets of salient points to improve the repeatability of the 3D keypoint detector. The proposed algorithm is compared against the state-of-the-art algorithms on two benchmark datasets. The experimental results show that the proposed scheme outperforms the best existing method by 5.35% and 60.98 points on the SHOT-Kinect dataset and by 5.45% and 145.54 points on the SHOT-SpaceTime dataset in terms of relative and absolute repeatability, respectively.
Dmytro Bobkov, Eckehard G. Steinbach
ICIP3
2019 Edge Cloud-based Augmented Reality
abstract
A convincing augmented reality (AR) experience requires vast computational resources, in particular for three-dimensional mapping of the environment, pose estimation and high-quality rendering of virtual objects. Even today's most powerful mobile devices can not provide such computational resources and consequently limit the achievable quality of the augmentation. To tackle this issue, all computations necessary for AR can be offloaded to the Edge Cloud, such that the mobile device merely acts as a camera and display. This approach introduces additional processing steps, namely video communication, which we carefully evaluate with respect to their influence on the quality of experience and energy consumption. In the evaluation of our prototype, we show that with a Glass-to-Glass delay of about 85 ms, our implementation is competitive against state-of-the-art solutions which run completely locally on a mobile device. Most notably, the additional steps required for offloading contribute little delay, which is often overcompensated by the faster computations in the Edge Cloud. A further benefit is that compared to performing all AR processing locally, offloading reduces the energy consumption in smartphones on average by 50 %. Moreover, the computational resources available for the AR application increase by a factor 10 to 100 through offloading. Finally, offloading enables high-quality AR applications even in low-end mobile devices.
Christoph Bachhuber, Alvaro Sanchez Martinez, Rastin Pries, Sebastian Eger, Eckehard G. Steinbach
MMSP5
2019 Decoder Side Motion Vector Refinement for Versatile Video Coding
abstract
Inter picture prediction is an essential component of today's hybrid video codecs. In order to reduce the bitrate required for motion vector (MV) signaling, the High Efficiency Video Coding (HEVC) standard and the latest Versatile Video Coding (VVC) draft utilize a merge mode to signal the MV. While the merge mode saves the bits for MV indication, it generates inaccurate MVs, which lead to imprecise prediction. To improve the coding performance, novel decoder side motion vector refinement (DMVR) schemes are currently being proposed. The DMVR approach refines the initial MV from the merge mode by searching the block with the smallest matching cost in the previous decoded reference pictures. We present two variants of block matching-based DMVR, namely template matching and bilateral matching. Experimental results obtained with the VTM-2.0 reference software, after integrating our approaches, demonstrate that our proposed methods provide an average luma BD-rate reduction of 4.71% for the template matching-based DMVR and 4.92% for the bilateral matching-based DMVR when using the random access configuration.
Han Gao 0001, Semih Esenlik, Zhijie Zhao, Eckehard G. Steinbach, Jianle Chen
MMSP4
2019 Low-Complexity Geometric Inter-Prediction for Versatile Video Coding
abstract
Non-rectangular block partitioning is a well-known method for improved inter-picture prediction in video coding, enabling better spatial adaptation to the signal properties. This contribution presents the most recent proposal of geometric inter-prediction (GIP) made to the Versatile Video Coding (VVC) standardization activity led by the Joint Video Experts Team (JVET). Implemented in the latest test model VTM-5.0 and evaluated according to the JVET Common Test Conditions, the proposed low-complexity GIP scheme provides objective luma BD-rate reductions of 0.22 % for random access and 0.44 % for low-delay test cases at 7% encoder runtime increase and negligible decoder runtime increase. The coding gain is provided by non-triangular partitioned blocks and in the presence of multiple other VVC coding tools. Furthermore, BD-rate reductions of 2.58 % and 2.78 % can be achieved specifically for pure screen content by employing an adaptive blending filter.
Max Bläser, Han Gao 0001, Semih Esenlik, Elena Alshina, Zhijie Zhao, Christian Rohlfing, Eckehard G. Steinbach
PCS7
2019 Low Complexity Decoder Side Motion Vector Refinement for VVC
abstract
Inter picture prediction is an essential component of today's hybrid video codecs. In order to reduce the motion vector signaling overhead, a merge mode with subsequent decoder side motion vector refinement (DMVR) is current under investigation for the first working draft of the standardization activity on Versatile Video Coding (VVC). While the DMVR method searches the refined MVs at the decoder side, it heavily increases the decoding complexity and the memory bandwidth requirements. To address these issues, a novel low complexity DMVR scheme is proposed in this paper. The low complexity DMVR approach refines the initial MV from the merge mode by searching the block with the smallest matching cost in the previous decoded reference pictures. The proposed low complexity improvements are added to a previously proposed bilateral matching-based DMVR approach. Experimental results obtained with the VTM 2.0 reference software, after integrating our approaches, show that the previously proposed DMVR provides an average luma BD-rate reduction of 4.59% with 32% additional decoding time and the proposed low complexity DMVR provides an average luma BD-rate reduction of 1.67% with only 6% additional decoding time when using the random access configuration.
Han Gao 0001, Semih Esenlik, Zhijie Zhao, Eckehard G. Steinbach, Jianle Chen
PCS4
2019 Joint Pricing and Cache Placement for Video Caching: A Game Theoretic Approach
abstract
Caching can effectively smooth the temporal traffic variability and decrease the redundant data transmission in mobile video delivery. In this paper, we consider a video caching system consisting of a video provider (VP), a mobile network operator (MNO) with a set of cache-enabled base stations (BSs), and multiple mobile users. The VP leases some popular videos to the MNO, while the MNO places these rented videos in local caches of its BSs to save expensive backhaul transmission cost. However, in such a two-sided market, these two entities are competing with each other for their own profit due to their opposite expectation on the video pricing. To address this, we model the competition between the two entities using the framework of Stackelberg games and propose a joint video pricing and cache placement strategy by considering the heterogeneity of video file sizes and exploiting the classic law of demand from the field of economics. The proposed optimization problem is able to jointly maximize the profit of the VP and the MNO by the optimal selection of the video pricing and the cache placement strategy given that price, for both noncooperative BS caching and cooperative BS caching cases. We then develop iterative algorithms based on dynamic programming and gradient ascent, respectively, for these two cases to find the Stackelberg equilibrium (SE). The simulation results further show that the proposed joint optimization formulation follows the law of demand in economics, and the proposed algorithms for both cases can efficiently converge to the SE point that jointly maximizes the profit for both the VP and the MNO.
Junni Zou, Congcong Zhai, Hongkai Xiong, Eckehard G. Steinbach
IEEE J. Sel. Areas Commun.5
2019 The IEEE 1918.1 "Tactile Internet" Standards Working Group and its Standards
abstract
The IEEE “Tactile Internet” (TI) Standards working group (WG), designated the numbering IEEE 1918.1, undertakes pioneering work on the development of standards for the TI. This paper describes the WG, its intentions, and its developing baseline standard and the associated reasoning behind that and touches on a further standard already initiated under its scope: IEEE 1918.1.1 on “Haptic Codecs for the TI.” IEEE 1918.1 and its baseline standard aim to set the framework and act as the foundations for the TI, thereby also serving as a basis for further standards developed on TI within the WG. This paper discusses the aspects of the framework such as its created TI architecture, including the elements, functions, interfaces, and other considerations therein, as well as the novel aspects and differentiating factors compared with, e.g., 5G Ultra-Reliable Low-Latency Communication, where it is noted that the TI will likely operate as an overlay on other networks or combinations of networks. Key foundations of the WG and its baseline standard are also highlighted, including the intended use cases and associated requirements that the standard must serve, and the TI's fundamental definition and assumptions as understood by the WG, among other aspects.
Oliver Holland, Eckehard G. Steinbach, R. Venkatesha Prasad, Qian Liu 0001, Zaher Dawy, Adnan Aijaz, Nikolaos Pappas 0001, Kishor Chandra Joshi, Vijay S. Rao, Sharief Oteafy, Mohamad A. Eid, Mark A. Luden, Amit Bhardwaj, Joachim Sachs, José Araújo
Proc. IEEE2
2019 Adaptive 5G Low-Latency Communication for Tactile InternEt Services
abstract
The tactile internet will enable a new range of capabilities to enable immersive remote operations and interactions with a physical world. Tactile internet use cases span over many fields, from remote operation of industrial applications in, e.g., hazardous environments, via remote-controlled driving in a fully automated intelligent transport system, to remote surgery where the unique expert skills can be delivered to different locations in the world. Fifth-generation (5G) communication will play a fundamental part in this tactile internet vision, as it will provide necessary capabilities for the demanding communication needs in terms of reliability and low latency, for operators or teleoperated systems that are connected wirelessly. This paper provides an overview of tactile internet services and haptic interactions and communication. The 5G functionality for ultrareliable and low-latency services is described in depth and it is shown how 5G new radio (NR) and the evolved long-term evolution (LTE) radio interface can achieve guaranteed low-latency wireless transmission. The costs for providing reliable and low-latency wireless transmission in terms of reduced spectral efficiency and coverage are discussed. The 5G system architecture with a software-based network design based on a distributed cloud platform is presented. It is shown how the 5G network is configured for tactile internet services via multidomain orchestration.
Joachim Sachs, Lars A. A. Andersson, José Araújo, Calin Curescu, Johan Lundsjö, Göran Rune, Eckehard G. Steinbach, Gustav Wikström
Proc. IEEE7
2019 Haptic Codecs for the Tactile Internet
abstract
The Tactile Internet will enable users to physically explore remote environments and to make their skills available across distances. An important technological aspect in this context is the acquisition, compression, transmission, and display of haptic information. In this paper, we present the fundamentals and state of the art in haptic codec design for the Tactile Internet. The discussion covers both kinesthetic data reduction and tactile signal compression approaches. We put a special focus on how limitations of the human haptic perception system can be exploited for efficient perceptual coding of kinesthetic and tactile information. Further aspects addressed in this paper are the multiplexing of audio and video with haptic information and the quality evaluation of haptic communication solutions. Finally, we describe the current status of the ongoing IEEE standardization activity P1918.1.1 which has the ambition to standardize the first set of codecs for kinesthetic and tactile information exchange across communication networks.
Eckehard G. Steinbach, Matti Strese, Mohamad A. Eid, Amit Bhardwaj, Qian Liu 0001, Mohammad Al Ja'afreh, Toktam Mahmoodi, Rania Hassen, Abdulmotaleb El Saddik, Oliver Holland
Proc. IEEE1
2019 On the Minimum Perceptual Temporal Video Sampling Rate and Its Application to Adaptive Frame Skipping
abstract
Media technology, in particular video recording and playback, keeps improving to provide users with high-quality real and virtual visual content. In recent years, increasing the temporal sampling rate of videos and the refresh rate of displays has become one focus of technical innovation. This raises the question, how high the sampling and refresh rates should be? To answer this question, we determine the minimum temporal sampling rate at which a video should be presented to make temporal sampling imperceptible to viewers. Through a psychophysical study, we find that this minimum sampling rate depends on both the speed of the objects in the image plane and the exposure time of the recording camera. We propose a model to compute the required minimum sampling rate based on these two parameters. In addition, state-of-the-art video codecs employ motion vectors from which the local object movement speed can be inferred. Therefore, we present a procedure to compute the minimum sampling rate given an encoded video and camera exposure time. Since the object motion speed in a video may vary, the corresponding minimum frame rate is also varying. This is why the results of this paper are particularly applicable when used together with adaptive frame rate computer generated graphics or novel video communication solutions that drop insignificant frames. In our experiments, we show that videos played back at the minimum adaptive frame rate achieve an average bit rate reduction of 26% compared to constant frame rate playback, while perceptually no difference can be observed.
Christoph Bachhuber, Amit Bhardwaj, Rastin Pries, Eckehard G. Steinbach
IEEE Trans. Circuits Syst. Video Technol.4
2019 Real-Time Global Registration for Globally Consistent RGB-D SLAM
abstract
Real-time globally consistent camera localization is critical for visual simultaneous localization and mapping (SLAM) applications. Regardless the popularity of high efficient pose graph optimization as a backend in SLAM, its deficiency in accuracy can hardly benefit the reconstruction application. An alternative solution for the sake of high accuracy would be global registration, which minimizes the alignment error of all the corresponding observations, yet suffers from high complexity due to the tremendous observations that need to be considered. In this paper, we start by analyzing the complexity bottleneck of global point cloud registration problem, i.e., each observation (three-dimensional point feature) has to be linearized based on its local coordinate (camera poses), which however is nonlinear and dynamically changing, resulting in extensive computation during optimization. We further prove that such nonlinearity can be decoupled into linear component (feature position) and nonlinear components (camera poses), where the former linear one can be effectively represented by its compact second-order statistics, while the latter nonlinear one merely requires six degrees of freedom for each camera pose. Benefiting from the decoupled representation, the complexity can be significantly reduced without sacrifice in accuracy. Experiments show that the proposed algorithm achieves globally consistent pose estimation in real-time via CPU computing, and owns comparable accuracy as state-of-the-art that use GPU computing, enabling the practical usage of globally consistent RGB-D SLAM on highly computationally constrained devices.
Lan Xu 0003, Dmytro Bobkov, Eckehard G. Steinbach, Lu Fang 0001
IEEE Trans. Robotics4
2018 Low-Complexity Fingerprint Matching for Real-Time Indoor Localization Systems
abstract
Due to their robust performance in complex non- line-of-sight environments, fingerprinting-based approaches have long been favored for indoor localization. However, many state-of-the-art schemes rely on a very dense fingerprint map. A larger fingerprint database means a higher achievable localization accuracy, at the same time it means a larger number of required fingerprint comparisons and longer computation delays in the system. This paper presents a novel low-effort approach, that uses a restructured fingerprint database to precisely calculate the location of a user, while comparing the data measured by the user to only a subset of database entries. Several novel approximations of the algorithm are also presented, that trade-off computation complexity and localization accuracy. For comparison, the complexities of a number of state-of-the-art fingerprinting schemes are derived. The effectiveness of the proposed approach is demonstrated through simulation. The presented results show that if the allowed number of fingerprint comparisons is set, the proposed approach and its approximations produce up to 34% lower localization errors than the traditional approach applied to a reduced database.
Alexandra Zayets, Eckehard G. Steinbach
GLOBECOM2
2018 Interpolation and Extrapolation of Multipath Fingerprints Using Virtual Transmitter Placement
abstract
Fingerprinting-based localization approaches generally show the highest localization precision in complex indoor non-line-of-site environments. However, they require the construction of a dense fingerprint map, which can lead to high system deployment and maintenance costs. In this paper, we focus on extending a measured fingerprint map though interpolation and extrapolation. In this way, the accuracy of an indoor localization system (ILS) can be improved without increasing its deployment cost. We focus on channel state information (CSI) based fingerprinting approaches as they have shown superior localization accuracy and robustness. This paper presents a novel algorithm for interpolating and extrapolating CSI and multipath fingerprints by calculating the positions of the so called virtual transmitters (VTs). The performance of the interpolation scheme is validated through simulation. The localization error distribution of an ILS using an interpolated and extrapolated fingerprint map is evaluated and compared to the case when the original map is used. The obtained results demonstrate a 20% decrease in the average localization error through the use of the proposed fingerprint interpolation scheme. Furthermore, the presented results show that when an extrapolated map is used, the probability that a user is precisely localized is less dependent on his position in the indoor environment.
Alexandra Zayets, Eckehard G. Steinbach
ICC2
2018 Color Image Demosaicking Using a 3-Stage Convolutional Neural Network Structure
abstract
Color demosaicking (CDM) is a critical first step for the acquisition of high-quality RGB images with single chip cameras. Conventional CDM approaches are mostly based on interpolation schemes and hand-crafted image priors, which result in unpleasant visual artifacts in some cases. Motivated by the special characteristics of inter-channel correlations (higher correlations for R/G and G/B channels than that for R/B), in this paper, a 3-stage convolutional neural network (CNN) structure for CDM is proposed. In the first stage, the G channel is reconstructed independently. Then, by using the reconstructed G channel as guidance, the R and B channels are recovered in the second stage. Finally, high-quality RGB color images are reconstructed in the third stage. The objective and visual quality evaluation results show that the proposed structure achieves noticeable quality improvements in comparison to the state-of-the-art approaches.
Kai Cui 0003, Zhi Jin 0002, Eckehard G. Steinbach
ICIP3
2018 Robust Map Alignment for Cooperative Visual SLAM
abstract
Generating a map of an unknown environment using visual SLAM with sparse image features is an important task in robotics and other computer vision applications. As the number of available recordings increases, merging maps from multiple sources into an aggregated description of the environment becomes necessary. After identifying similar locations, an important step is the estimation of the transformation that aligns the maps. The aim of this work is to evaluate different methods for computing this transformation and to provide a novel way of estimating the scale difference between maps using a histogram-based scale matching scheme. The proposed approach proves to be more robust than the currently widely used scale estimation methods for loop closure or map merging.
Adrian Garcea, Jiazhen Zhu, Dominik Van Opdenbosch, Eckehard G. Steinbach
ICIP4
2018 Flexible Rate Allocation for Local Binary Feature Compression
abstract
Numerous real-time applications in computer vision rely on finding correspondences between local binary features. In many mobile scenarios, the visual information captured at a sensor node needs to be transmitted to a processing server, which is capable of storing the visual information or executing a complex analysis task. However, not necessarily all the visual information need to be transmitted. In this paper, we present a rate allocation scheme that is capable of categorizing features into classes according to their usefulness and select the amount of data spent on each class to maximize the overall performance of a computer vision task. We demonstrate the approach using ORB, BRISK, and FREAK features and show the improvements on a homography estimation task.
Dominik Van Opdenbosch, Eckehard G. Steinbach
ICIP2
2018 A Delay Compensation Approach for Pan-Tilt-Unit-based Stereoscopic 360 Degree Telepresence Systems Using Head Motion Prediction
abstract
The acceptance of teleoperation applications like tele-driving, tele-surgery, tele-maintenance, etc., is challenged by the quality-reducing effect of end-to-end latency. Particularly, when users wear Head-Mounted Displays to enhance the immersive experience, the lag between head motion and display response leads to unbearable motion sickness, indisposition, and, in the worst case, abortion of the teleoperation session. In this paper, we propose a delay compensation approach with head motion prediction that can be applied to pan-tilt-unit-based stereoscopic telepresence systems. We provide the user with the impression of a 3D 360° video that represents the remote scene without noticing the present delay, even when rotating the head. To this end, we propose a novel prediction paradigm for head motion estimation to substantially mitigate the negative impact of the latency on the quality of experience. We re-implemented state-of-the-art head movement predictors and compare them to our proposed approach by means of qualitative measures. In our experiments, we used two real and independent head motion datasets for validation and tested communication delays between 100-1000ms. Our results show that mean compensation rates of more than 99% are able with our approach.
Tamay Aykut, Chenxi Zou, Dominik Van Opdenbosch, Eckehard G. Steinbach
ICRA5
2018 Selection and Compression of Local Binary Features for Remote Visual SLAM
abstract
In the field of autonomous robotics, Simultaneous Localization and Mapping (SLAM) is still a challenging problem. With cheap visual sensors attracting more and more attention, various solutions to the SLAM problem using visual cues have been proposed. However, current visual SLAM systems are still computationally demanding, especially on embedded devices. In addition, collaborative SLAM approaches emerge using visual information acquired from multiple robots simultaneously to build a joint map. In order to address both challenges, we present an approach for remote visual SLAM where local binary features are extracted at the robot, compressed and sent over a network to a centralized powerful processing node running the visual SLAM algorithm. To this end, we propose a new feature coding scheme including a feature selection stage which ensures that only relevant information is transmitted. We demonstrate the effectiveness of our approach on well-known datasets. With the proposed approach, it is possible to build an accurate map while limiting the data rate to 75 kbits/frame.
Dominik Van Opdenbosch, Martin Oelsch, Adrian Garcea, Tamay Aykut, Eckehard G. Steinbach
ICRA5
2018 Learning-Based Modular Task-Oriented Grasp Stability Assessment
abstract
Assessing grasp stability is essential to prevent the failure of robotic manipulation tasks due to sensory data and object uncertainties. Learning-based approaches are widely deployed to infer the success of a grasp. Typically, the underlying model used to estimate the grasp stability is trained for a specific task, such as lifting, hand-over, or pouring. Since every task has individual stability demands, it is important to adapt the trained model to new manipulation actions. If the same trained model is directly applied to a new task, unnecessary grasp adaptations might be triggered, or in the worst case, the manipulation might fail. To address this issue, we divide the manipulation task used for training into seven sub-tasks, defined as modular tasks. We deploy a learning-based approach and assess the stability for each modular task separately. We further propose analytical features to reduce the dimensionality and the redundancy of the tactile sensor readings. A main task can thereby be represented as a sequence of relevant modular tasks. The stability prediction of the main task is computed based on the inferred success labels of the modular tasks. Our experimental evaluation shows that the proposed feature set lowers the prediction error up to 5.69% compared to other sets used in state-of-the-art methods. Robotic experiments demonstrate that our modular task-oriented stability assessment avoids unnecessary grasp force adaptations and regrasps for various manipulation tasks.
Amit Bhardwaj, Tamay Aykut, Nicolas Alt, Mojtaba Karimi, Eckehard G. Steinbach
IROS7
2018 MID: A Novel Contrast Metric for the MSER Detector
abstract
This paper presents a novel contrast measure for MSER region selection, termed Mean Intensity Difference (MID). The proposed metric is computed between the pixels of an MSER region and its surrounding pixel set. In this work we consider the complementary pixels within the bounding box of a region and alternatively the pixels of the first contour layer as surroundings. To evaluate the proposed contrast metric, a location retrieval task is performed. To this end, SURF descriptors are computed and the Bag-of-Words representation is used as global signature for each image. For the evaluation we use the Devon Island dataset, which is said to have one of the most Mars-like environments on Earth and which comes with GPS ground-truth data. We further integrate the contrast-based methods with the approach of Grid Adaptation. The experimental results show that our contrast metric outperforms state-of-the-art metrics, such as Perceptual Divergence, and yields better performance compared to random region selection. In this work we also evaluate the computational complexity of the methods.
Martin Oelsch, Basak Güleçyüz, Eckehard G. Steinbach
ISM3
2018 HSSIM: An Objective Haptic Quality Assessment Measure for Force-Feedback Signals
abstract
Recent advances in haptic communication cast light onto the promise of full immersion into remote real or virtual environments. The quality of compressed haptic signals is crucial to fulfill this promise. Traditionally, the quality of haptic signals is evaluated through a series of subjective experiments. So far, only very limited attention was directed toward developing objective quality measures for haptic communication. As touch applications continue to grow, the need for efficient haptic quality assessment (HQA) measures is becoming more pronounced. In this work, we attempt to pragmatically approach this problem inspired by recent advances in the visual domain. More specifically, we design a HQA measure based on our knowledge of human haptic perception represented by Stevens power law and a similarity comparison between a reference and a compressed haptic force-feedback signal represented by the SSIM index as a full-reference quality measure. The proposed HSSIM measure provides clear advantages over existing approaches. We validate the performance of the proposed HQA measure with subjective experiments which show high correlation with human quality assessment results.
Rania Hassen, Eckehard G. Steinbach
QoMEX2
2018 Improving Picture Boundary Handling for Video Coding Beyond HEVC
abstract
Block partitioning is an essential component of today's hybrid video codecs. In order to support video sequences with arbitrary spatial resolution, the HEVC standard and the JEM reference software utilize a forced quadtree (QT) method to process CTUs/CUs located at the picture boundaries. While the forced QT method is simple to implement, it is unsatisfactory from a coding efficiency perspective. To improve the compression performance, novel block partitioning methods are currently being proposed. Our approach for picture boundary handling uses a binary tree (BT) based splitting scheme to deal with the boundary located CTUs/CUs. We present two variants with the first one combining forced BT with forced QT partitioning (forced QTBT) and the second one adaptively choosing between BT and QT (adaptive QTBT). Experimental results obtained with the JEM-6.0 reference software after integrating our approaches demonstrate that our proposed methods provide an average luma BD-rate reduction of 0.85% for the forced QTBT scheme and 0.96% for the adaptive QTBT scheme when using the random access configuration.
Han Gao 0001, Zhijie Zhao, Eckehard G. Steinbach, Jianle Chen
VCIP3
2018 Delay Compensation for Actuated Stereoscopic 360 Degree Telepresence Systems with Probabilistic Head Motion Prediction
abstract
Communication delay is a major challenge for the acceptance of telepresence applications. It is particularly critical when the user experiences the remote environment via a Head-Mounted-Display. The lag between head motion and display response results in motion sickness, indisposition, and, at worst, abortion of the telepresence session. In this paper, we propose a delay compensation approach for 3D 360° telepresence systems realized with a mechanically actuated stereoscopic vision system. We further introduce a novel metric to evaluate the achievable level of delay compensation. We investigate state-of-the-art head motion predictors and propose a novel probabilistic prediction paradigm, which can half the mean prediction error and improve the level of delay compensation by up to 26%. The general validity of our approach is shown by means of two independent real head motion datasets. The experimental results verify that average compensation rates of more than 99% can be achieved for communication delays between 100-500ms.
Tamay Aykut, Christoph Burgmair, Mojtaba Karimi, Eckehard G. Steinbach
WACV5
2018 Efficient Map Compression for Collaborative Visual SLAM
abstract
Swarm robotics is receiving increasing interest, because the collaborative completion of tasks, such as the exploration of unknown environments, leads to improved performance and reduced effort. The ability to exchange map information is an essential requirement for collaborative exploration. When moving to large-scale environments, where the communication data rate between the swarm participants is typically limited, efficient compression algorithms and an approach for discarding less informative parts of the map are key for a successful long-term operation. In this paper, we present a novel compression approach for environment maps obtained from a visual SLAM system. We apply feature coding to the visual information to compress the map efficiently. We make use of a minimum spanning tree to connect all features that serve as observations of a single map point. Thereby, we can exploit inter-feature dependencies and obtain an optimal coding order. Additionally, we add a map sparsification step to keep only useful map points by solving a linear integer programming problem, which preserves the map points that exhibit both good compression properties and high observability. We evaluate the proposed method on a standard dataset and show that our approach outperforms state-of-the-art techniques.
Dominik Van Opdenbosch, Tamay Aykut, Nicolas Alt, Eckehard G. Steinbach
WACV4
2018 Efficient Multi-Rate Video Encoding for HEVC-Based Adaptive HTTP Streaming
abstract
Adaptive HTTP streaming requires a video to be encoded at multiple representations, that is, different qualities. Encoding these multiple representations is a computationally complex process, especially when using the recent High Efficiency Video Coding (HEVC) standard. In this paper, we consider a multi-rate HEVC encoder and identify four types of encoding information that can be reused from a high-quality reference encoding to speed up lower quality-dependent encodings. We show that the encoding decisions from the reference cannot be directly reused, as this would harm the overall rate-distortion (RD) performance. Thus, we propose methods to use the encoding information to constrain the RD optimization of the dependent encodings so that the encoding complexity is reduced while the RD performance is kept high. We additionally show that the proposed methods can be combined, leading to an efficient multi-rate encoder that exhibits high RD performance and substantial complexity reduction. Results show that the encoding time for 12 representations at different spatial resolutions and signal qualities can be reduced on average by 38%, while the average bitrate increases by less than 1%.
Damien Schroeder, Adithyan Ilangovan, Martin Reisslein, Eckehard G. Steinbach
IEEE Trans. Circuits Syst. Video Technol.4
2018 On the Minimization of Glass-to-Glass and Glass-to-Algorithm Delay in Video Communication
abstract
Video cameras are increasingly used to provide real-time feedback in automatic control systems, such as autonomous driving and robotics systems. For such highly dynamic applications, the glass-to-glass (G2G) and glass-to-algorithm (G2A) latencies are critical. In this paper, we analyze the latencies in a point-to-point video transmission system and propose novel frame skipping and preemption approaches to reduce the G2G and G2A delays. We implement the proposed approaches in a prototype that shows significantly reduced G2G and G2A latencies as well as reduced transmission bitrate requirements compared with traditional video transmission schemes. In our low-delay video communication prototype, a VGA resolution video is transmitted with average G2G and G2A delays of 21.2 and 11.5 ms, respectively, with off-the-shelf hardware.
Christoph Bachhuber, Eckehard G. Steinbach, Martin Freundl, Martin Reisslein
IEEE Trans. Multim.2
2018 Context-Aware Task Migration for HART-Centric Collaboration over FiWi Based Tactile Internet Infrastructures
abstract
Low task execution time and low energy consumption of collaborating mobile human users and robots are important requirements of emerging human-agent-robot teamwork (HART)-centric Tactile Internet applications. In particular, task migration among mobile HART members has emerged as an important research topic, taking different task types, task deadlines, collaborative node capabilities, and mobility patterns into account. We propose a context-aware task migration scheme for efficiently orchestrating the real-time collaboration among human mobile users, central and decentralized computational agents (cloud/cloudlets), and collaborative robots (cobots) across converged fiber-wireless (FiWi) communications infrastructures. We investigate the problem of whether and, if so, when and where a HART-centric task should be best migrated to. For resource-efficient task execution, the migration decision is made according to given task processing capabilities of cloud/cloudlet agents and cobots, task execution deadline, energy consumption of involved cobots and mobile devices, and task migration latency. We evaluate the performance of our context-aware HART-centric task migration scheme and compare it to conventional task execution without migration. Towards this end, we develop an analytical framework for quantifying its performance in terms of a variety of task migration key performance metrics, including task migration gain-overhead ratio, deadline-miss ratio, task response time, and energy consumption efficiency.
Mahfuzulhoq Chowdhury, Eckehard G. Steinbach, Wolfgang Kellerer, Martin Maier 0001
IEEE Trans. Parallel Distributed Syst.2
2017 Room segmentation in 3D point clouds using anisotropic potential fields
abstract
Emerging applications, such as indoor navigation or facility management, present new requirements of automatic and robust partitioning of indoor 3D point clouds into rooms. Previous research is either based on the Manhattan-world assumption or relies on the availability of the scanner pose information. We address these limitations by following the architectural definition of a room, where the room is an inner free space separated from other spaces through openings or partitions. For this we formulate an anisotropic potential field for 3D environments and illustrate how it can be used for room segmentation in the proposed segmentation pipeline. The experimental results confirm that our method outperforms state-of-the-art methods on a number of datasets including those that violate the Manhattan-world assumption.
Dmytro Bobkov, Martin Kiechle, Sebastian Hilsenbeck, Eckehard G. Steinbach
ICME4
2017 Grasping posture estimation for a two-finger parallel gripper with soft material jaws using a curved contact area friction model
abstract
We present a friction model for the curved contact area between a deformable object and soft parallel gripper jaws for grasping posture estimation. We show that the assumption of a planar contact area leads to an overestimation of the frictional force and torque, which might cause the object to slip. We simulate the contact with the Finite Element Method, then compute the friction wrenches, which are fitted with two limit surface models: an ellipsoid and a convex 4th-order polynomial. Despite a slightly higher fitting error, the ellipsoid limit surface is chosen to compute the grasp quality because of its simplicity. We compare the limit surfaces of our friction model with the planar contact model and show the improved accuracy obtainable with our model. We then apply the presented model for grasping posture estimation by simulating the contact for all grasp candidates. We show a grasp quality map (quality of all grasp candidates) and the best possible grasp location for several deformable objects.
Nicolas Alt, Zhongyao Zhang, Eckehard G. Steinbach
ICRA4
2017 A Testbed for Vision-Based Networked Control Systems
Christoph Bachhuber, Simon Conrady, Michael Schütz, Eckehard G. Steinbach
ICVS4
2017 Robust WiFi-based indoor localization using multipath component analysis
abstract
The number of applications that rely on robust indoor localization is constantly growing. Conventional outdoor localization technologies are generally not suited for indoor use. WiFi-based indoor localization systems are a widely studied alternative since the necessary infrastructure is already available in most buildings. Existing WiFi-based localization algorithms, however, still face challenges such as high sensitivity to changes in the environment, temporal instability, a time consuming calibration process and low accuracy in non-line-of-sight (NLOS) multipath environments. In this paper, we propose a novel fingerprinting-based WiFi indoor localization scheme which operates by extracting and analyzing individual multipath propagation delays. We demonstrate through simulation the robustness of the proposed algorithm to changes in the multipath environment.
Alexandra Zayets, Eckehard G. Steinbach
IPIN2
2017 Survey of Visual Feature Extraction Algorithms in a Mars-like Environment
abstract
This paper presents a performance comparison of several state-of-the-art visual feature extraction algorithms when applied in a poorly-structured environment as found on the planet Mars. So far, no systematic evaluation of feature extraction algorithms in extraterrestrial environments is available. The algorithms in this paper are evaluated using the Devon Island dataset which is said to have one of the most Mars-like environments on Earth. The ground truth for the performance comparison is based on a location retrieval task using the GPS data provided by the dataset. A range of common feature detection and description algorithms is covered including floating point and binary descriptors. With the aim of efficient image retrieval the descriptor vectors are quantized in a bag-of-words model using the k-means algorithm. The results show that the well-known SURF feature provides superior performance over other state-of-the-art feature extraction algorithms in a Mars-like environment.
Martin Oelsch, Dominik Van Opdenbosch, Eckehard G. Steinbach
ISM3
2017 A joint compression scheme for local binary feature descriptors and their corresponding bag-of-words representation
abstract
For real-time computer vision tasks, binary feature descriptors are an efficient alternative to their real-valued counterparts. While providing comparable results for many applications, the computational complexity of extracting and processing binary descriptors is significantly lower. In many application scenarios, the local features are transmitted over a channel with limited capacity and processed at a more powerful central processing unit, which requires efficient compression and transmission approaches. In this paper, we present a compression scheme for local binary features, which jointly encodes the descriptors and their respective Bag-of-Words representation using a shared vocabulary between client and server. By sending the visual word index and the entropy-coded residual vector containing the differences between the visual word and the descriptor, we are able to reduce ORB features to 60.62 % of their uncompressed size.
Dominik Van Opdenbosch, Martin Oelsch, Adrian Garcea, Eckehard G. Steinbach
VCIP4
2017 6DOF decoupled roto-translation alignment of large-scale indoor point clouds
Anas Al-Nuaimi, Sebastian Hilsenbeck, Adrian Garcea, Eckehard G. Steinbach
Comput. Vis. Image Underst.4
2017 A Multiplexing Scheme for Multimodal Teleoperation
abstract
This article proposes an application-layer multiplexing scheme for teleoperation systems with multimodal feedback (video, audio, and haptics). The available transmission resources are carefully allocated to avoid delay-jitter for the haptic signal potentially caused by the size and arrival time of the video and audio data. The multiplexing scheme gives high priority to the haptic signal and applies a preemptive-resume scheduling strategy to stream the audio and video data. The proposed approach estimates the available transmission rate in real time and adapts the video bitrate, data throughput, and force buffer size accordingly. Furthermore, the proposed scheme detects sudden transmission rate drops and applies congestion control to avoid abrupt delay increases and converge promptly to the altered transmission rate. The performance of the proposed scheme is measured objectively in terms of end-to-end signal latencies, packet rates, and peak signal-to-noise ratio (PSNR) for visual quality. Moreover, peak-delay and convergence time measurements are carried out to investigate the performance of the congestion control mode of the system.
Burak Cizmeci, Xiao Xu 0001, Rahul Gopal Chaudhari, Christoph Bachhuber, Nicolas Alt, Eckehard G. Steinbach
ACM Trans. Multim. Comput. Commun. Appl.6
2016 A system for high precision glass-to-glass delay measurements in video communication
abstract
Ultra low delay video transmission is becoming increasingly important. Video-based applications with ultra low delay requirements range from teleoperation scenarios such as controlling drones or telesurgery to autonomous control of dynamic processes using computer vision algorithms applied on real-time video. To evaluate the performance of the video transmission chain in such systems, it is important to be able to precisely measure the glass-to-glass (G2G) delay of the transmitted video. In this paper, we present a low-complexity system that takes a series of pairwise independent measurements of G2G delay and derives performance metrics such as mean delay or minimum delay etc. from the data. The precision is in the sub-millisecond range, mainly limited by the sampling rate of the measurement system. In our implementation, we achieve a G2G measurement precision of 0.5 milliseconds with a sampling rate of 2kHz.
Christoph Bachhuber, Eckehard G. Steinbach
ICIP2
2016 Analyzing LiDAR scan skewing and its impact on scan matching
abstract
This paper presents a systematic study of the scan skewing problem. Scan skewing is the non-rigid deformation of point clouds acquired by LiDAR's and is the result of their sequential scanning nature. We theoretically analyze the impact of skewing on scan matching and subsequently quantify the impact using synthetic LiDAR data with controlled skew distortions. We also show how the Geometric-Algebra LMS, an iterative point set registration filter, can be tuned to incorporate skewing aware weights to reduce the impact of skewing on scan matching. Results with real 3D LiDAR data are also presented.
Anas Al-Nuaimi, Wilder Bezerra Lopes, Paul Zeller, Adrian Garcea, Cássio Guimarães Lopes, Eckehard G. Steinbach
IPIN6
2016 AVLAD: Optimizing the VLAD Image Signature for Specific Feature Descriptors
abstract
Recent works on content-based image retrieval have successfully used the Vector of Locally Aggregated Descriptors (VLAD) as a compact image signature. In this paper, we improve the VLAD signature by tailoring the VLAD representation to the specific properties of the visual feature descriptors. We combine this improvement with the recently proposed hierarchical VLAD approach and demonstrate the effectiveness of this extension for the two well-known feature descriptors SIFT and SURF. Furthermore, we investigate how to efficiently reduce the dimensionality of the resulting representation using different unsupervised dimensionality reduction techniques.
Dominik Van Opdenbosch, Eckehard G. Steinbach
ISM2
2016 6DOF point cloud alignment using geometric algebra-based adaptive filtering
abstract
In this paper we show that a Geometric Algebra-based least-mean-squares adaptive filter (GA-LMS) can be used to recover the 6-degree-of-freedom alignment of two point clouds related by a set of point correspondences. We present a series of techniques that endow the GA-LMS with outlier (false correspondence) resilience to outperform standard least squares (LS) methods that are based on Singular Value Decomposition (SVD). We furthermore show how to derive and compute the step size of the GA-LMS.
Anas Al-Nuaimi, Eckehard G. Steinbach, Wilder Bezerra Lopes, Cássio Guimarães Lopes
WACV2
2016 Low-Complexity and Context-Aware Estimation of Spatial and Temporal Activity Parameters for Automotive Camera Rate Control
abstract
Rate control in video compression adjusts the encoding parameters to reach a certain target bitrate for the encoded video. State-of-the-art rate controllers for hybrid video coding typically employ content-dependent video bitrate models and video quality metrics (VQMs). To capture the content characteristics, temporal and spatial video activity measures are determined from the raw video using computationally complex algorithms that require access to the uncompressed source video. In automotive deployments, however, full access to the uncompressed source video and the internal functions of video encoders is typically not possible. As a remedy, in this paper, we present a low-complexity approach to estimate the temporal activity (TA) and spatial activity (SA) measures for videos that are captured by a front-facing camera of a vehicle, based on the context information of the vehicle. To this end, we exploit information about the dynamics of the vehicle and other vehicles in the field-of-view of the front-facing camera. We apply the estimated TA and SA values to a video bitrate model and an objective VQM and use these models to solve the rate control problem to determine the optimal encoding settings for given bitrate constraints. The proposed low-complexity solution offers a similar accuracy in achieving rate constraints and similar perceptual quality characteristics as a solution that uses the computed TA and SA values, with the advantage that no access to the uncompressed source video stream or the internal functions of the video encoder is required.
Christian Lottermann, Damien Schroeder, Eckehard G. Steinbach
IEEE Trans. Circuits Syst. Video Technol.3
2016 Keypoint Encoding for Improved Feature Extraction From Compressed Video at Low Bitrates
abstract
In many mobile visual analysis applications, compressed video is transmitted over a communication network and analyzed by a server. Typical processing steps performed at the server include keypoint detection, descriptor calculation, and feature matching. Video compression has been shown to have an adverse effect on feature-matching performance. The negative impact of compression can be reduced by using the keypoints extracted from the uncompressed video to calculate descriptors from the compressed video. Based on this observation, we propose to provide these keypoints to the server as side information and to extract only the descriptors from the compressed video. First, we introduce four different frame types for keypoint encoding to address different types of changes in video content. These frame types represent a new scene, the same scene, a slowly changing scene, or a rapidly moving scene, and are determined by comparing features between successive video frames. Then, we propose Intra, Skip, and Inter modes of encoding the keypoints for different frame types. For example, keypoints for new scenes are encoded using the Intra mode, and keypoints for unchanged scenes are skipped. As a result, the bitrate of the side information related to keypoint encoding is significantly reduced. Finally, we present pairwise matching and image retrieval experiments conducted to evaluate the performance of the proposed approach using the Stanford mobile augmented reality dataset and 720p format videos. The results show that the proposed approach offers significantly improved feature matching and image retrieval performance at a given bitrate.
Jianshu Chao, Eckehard G. Steinbach
IEEE Trans. Multim.2
2016 Deep Learning for Surface Material Classification Using Haptic and Visual Information
abstract
When a user scratches a hand-held rigid tool across an object surface, an acceleration signal can be captured, which carries relevant information about the surface material properties. More importantly, such haptic acceleration signals can be used together with surface images to jointly recognize the surface material. In this paper, we present a novel deep learning method dealing with the surface material classification problem based on a fully convolutional network, which takes the aforementioned acceleration signal and a corresponding image of the surface texture as inputs. Compared to the existing surface material classification solutions which rely on a careful design of hand-crafted features, our method automatically extracts discriminative features utilizing advanced deep learning methodologies. Experiments performed on the TUM surface material database demonstrate that our method achieves state-of-the-art classification accuracy robustly and efficiently.
Haitian Zheng, Lu Fang 0001, Mengqi Ji, Matti Strese, Yigitcan Özer, Eckehard G. Steinbach
IEEE Trans. Multim.6
2015 Transparency analysis of client-server-based multi-rate haptic interaction with deformable objects
abstract
In this paper we describe a client-server architecture for haptic interaction with simulated deformable objects. The computationally expensive object deformation is computed on the server at a low temporal update rate and transmitted to the clients. There, an intermediate representation of the deformable object is used to locally render haptic force feedback displayed to the user at the required rate of 1 kHz. Based on a one-dimensional deformable object, we analyze the transparency of this multi-rate architecture for a single client interaction. The delay introduced by the deformation simulation and the client-server communication leads to increased rendered forces at the clients compared to a reference scenario without delay. We propose a method that adaptively adjusts the stiffness used in the local force rendering at the client to compensate for this. The evaluation shows that the proposed method successfully compensates the effect of delay in the tested delay range of up to 100 ms.
Clemens Schuwerk, Xiao Xu 0001, Wolfgang Freund, Eckehard G. Steinbach
World Haptics4
2015 Surface classification using acceleration signals recorded during human freehand movement
abstract
When a tool is used to tap onto an object or it is dragged over the object surface, vibrations are induced in the tool that can be captured using acceleration sensors. Based on these signals, this paper presents an approach for tool-mediated surface classification which is robust against varying scan-time parameters. We examine freehand recordings of 69 textures and propose a classification system that uses perception-related features such as hardness, roughness and friction as well as selected features adapted from speech recognition such as modified cepstral coefficients. We focus on mitigating the effect of varying contact force and hand speed conditions on these features as a prerequisite for a robust machine-learning-based approach for surface classification. Our system works without explicit scan force and velocity measurements. Experimental results show that our proposed approach allows for successful classification of surface textures under varying freehand movement conditions. The proposed features lead to a classification accuracy of 95% when combined with a Naive Bayes Classifier.
Matti Strese, Clemens Schuwerk, Eckehard G. Steinbach
World Haptics3
2015 Haptic data reduction for time-delayed teleoperation using the time domain passivity approach
abstract
We propose a perceptual haptic data reduction approach for teleoperation systems which use the time domain passivity approach (TDPA) as their control architecture for dealing with time-varying communication delay. Our goal is to reduce the packet rate over the communication network while preserving system stability in the presence of time-varying and unknown delays. Compared to the existing wave variable-based (WV-based) haptic data reduction approaches, our proposed scheme leads to smaller distortion in the force signals and robustly deals with time-varying delays. Experiments show that our proposed approach can reduce the average packet rate by up to 80%, without introducing significant distortion. In addition, the proposed approach outperforms the existing WV-based approaches in both packet rate reduction and subjective preference for the tested communication delays.
Xiao Xu 0001, Burak Cizmeci, Clemens Schuwerk, Eckehard G. Steinbach
World Haptics4
2015 Objective quality prediction for haptic texture signal compression
abstract
Perceptual quality for media compression algorithms is traditionally evaluated through user studies. Such studies are time consuming, laborious and expensive, slowing down the development of new signal processing algorithms. To address this problem, a number of algorithmic quality prediction methodologies have been developed in the audio and video fields, something that is currently lacking in haptics research. In this paper, we present a novel method for predicting the perceptual quality degradation of compressed haptic texture signals. For this purpose, abstract perceptual features like Roughness, Brightness, etc. that capture the subjective experience of textures are exploited, in addition to low-level psychophysical models from the literature. As compared to the state-of-the-art, the presented prediction methodology shows an approximately 30% improvement in explaining the variance in the perceptual data.
Rahul Gopal Chaudhari, Yongjae Yoo, Clemens Schuwerk, Seungmoon Choi, Eckehard G. Steinbach
ICASSP5
2015 Network-aware video level encoding for uplink adaptive HTTP streaming
abstract
We study the uplink delivery of live video using adaptive HTTP streaming (AHS). In AHS, the process of simultaneously creating video levels at different rates is computationally demanding and quickly exceeds the computational capacity of mobile devices. As a remedy, we propose a network-aware video level selection approach which reduces the number of levels that need to be encoded. To this end, we develop an algorithm which selects a reduced set of video levels from a static pre-defined set based on TCP uplink throughput information. More specifically, during session start-up and after inter-RAN handovers, TCP uplink throughput information from a remote database is used while otherwise, actual TCP uplink throughput measurements are performed. We test the proposed approach in an automotive scenario to upstream the video of a vehicle's front-facing camera to a remote video portal. Our results show that our proposed network-aware video level selection approach leads to a significant reduction of the number of video levels that need to be encoded. At the same time, a similar quality of experience is achieved in terms of mean subjective quality of the delivered video segments, interrupted playback duration due to stalling events, and number of quality switches when compared to an implementation which considers the static full set of video levels.
Christian Lottermann, Serhan Gul, Damien Schroeder, Eckehard G. Steinbach
ICC4
2015 Low-complexity block size decision for HEVC intra coding using binary image feature descriptors
abstract
We present a novel low-complexity algorithm for feature-based block size decision in HEVC intra coding. Our approach evaluates a set of point pairs within a coding unit (CU) in order to determine whether a CU should be further split into smaller sub CUs for coding. We apply a modified version of the Binary Robust Independent Elementary Feature (BRIEF) descriptor, which in its original version is usually applied to describe local image properties in the context of image/video analysis. While the elements of the original BRIEF descriptor describe which pixel of a pixel pair has the higher value, the modified descriptor evaluates whether the difference between the pixel values is above a pre-defined threshold, which is determined by training. If a certain number of pixel pair differences exceed their corresponding threshold the CU is split into its four sub CUs. Furthermore, we restrict the feature point pairs to be located within different potential sub CUs. For an adaptive training approach, we achieved in our experiments an average encoding time reduction of 58% compared to the HEVC reference software HM12.0 with an average rate increase of 3.8% at equal quality and an average encoding time reduction of 65% with an average rate increase of 5.98% at equal quality for an offline training approach.
Walther Geuder, Peter Amon, Eckehard G. Steinbach
ICIP3
2015 Block structure reuse for multi-rate high efficiency video coding
abstract
Adaptive HTTP streaming requires a video to be encoded at multiple independently decodable bitrates. Encoding of multiple bitrates is a complex and time consuming process, especially with the new video coding standard HEVC. In this paper, we analyze the relation of HEVC block structures across encodings of a video at different bitrates. We propose a multi-rate encoding method for HEVC to decrease the overall encoding complexity while keeping the rate-distortion performance as high as possible. Experimental results show an average encoding time decrease of 27% compared to the reference HEVC encoder without degrading the rate-distortion performance.
Damien Schroeder, Patrick Rehm, Eckehard G. Steinbach
ICIP3
2015 Keypoint encoding and transmission for improved feature extraction from compressed images
abstract
In many mobile visual analysis scenarios, compressed images are transmitted over a communication network for analysis at a server. Often, the processing at the server includes some form of feature extraction and matching. Image compression has been shown to have an adverse effect on feature matching performance. To address this issue, we propose to signal the feature keypoints as side information to the server, and extract only the feature descriptors from the compressed images. To this end, we propose an approach to efficiently encode the locations, scales, and orientations of keypoints extracted from the original image. Furthermore, we propose a new approach for selecting relevant yet fragile keypoints as side information for the image, thus further reducing the data volume. We evaluate the performance of our approach using the Stanford mobile augmented reality dataset. Results show that the feature matching performance is significantly improved for images at low bitrate.
Jianshu Chao, Eckehard G. Steinbach, Lexing Xie
ICME2
2015 Activity recognition on handheld devices for pedestrian indoor navigation
abstract
We propose an inertial sensor-based approach to activity recognition for pedestrian indoor navigation. In the considered scenario a mobile device is held in a hand in front of the user. The recognized activities are the ones relevant to positioning in multi-floor buildings: walking and going up or down the stairs. To model the time dependency between consecutive activities we employ a Hidden Markov Model (HMM). For efficient quantization of continuous features, we apply a random forest classifier. For verification of the proposed algorithm, we conducted experiments with 12 participants and 4 different mobile devices. In our comparison to state-of-the-art approaches, we implement and evaluate major classification algorithms, such as nearest neighbour, decision tree and dynamic Bayesian Network. In the experiments we show the trade-off between computational complexity and classification performance. Furthermore, we demonstrate that the complexity of the HMM can be significantly reduced by replacing it with a dynamic Bayesian network with negligible impact on classification performance. The best of our proposed classifier achieves a classification accuracy of 91% for new users, which offers a 30% improvement compared to state-of-the-art approaches.
Dmytro Bobkov, Ferdinand Grimm, Eckehard G. Steinbach, Sebastian Hilsenbeck, Georg Schroth
IPIN3
2015 Multi-rate encoding for HEVC-based adaptive HTTP streaming with multiple resolutions
abstract
Adaptive HTTP streaming requires a video to be encoded at different rates and qualities called representations. The encoding of multiple representations with the new video coding standard HEVC is computationally complex. In this paper, we propose a multi-rate encoding method which reduces the complexity of encoding a video at multiple spatial resolutions. We first examine block structure similarities at different resolutions and propose a method to derive the block structure for a low resolution representation from a reference high resolution encoding. The derived block structure is used to speed up the encoding of low resolution representations. We further consider the content of the videos in order to achieve a rate-distortion (RD) performance similar to independent HEVC encoding. Experimental results show that the encoding time can be reduced by 50% on average for a low resolution video without degrading the RD performance.
Damien Schroeder, Adithyan Ilangovan, Eckehard G. Steinbach
MMSP3
2015 Macroblock level rate control for low delay H.264/AVC based video communication
abstract
In this paper, we propose a macro-block (MB) level rate control algorithm for low delay H.264/AVC video communication based on the ρ domain rate model. In the proposed algorithm, an exponential model is used to characterize the relation between ρ and the quantization step (Qstep) at the MB level, with which the quantization parameter (QP) for a MB can be obtained. Furthermore, a switched QP calculation scheme is introduced to obtain the QP for each MB to avoid large deviation of the actual frame size from the target bit budget. Compared with the original ρ domain rate control, the proposed method can achieve better video quality and improved bit-rate accuracy. Meanwhile, the computational complexity is also significantly reduced.
Min Gao 0002, Burak Cizmeci, Michael Eiler, Eckehard G. Steinbach, Debin Zhao, Wen Gao 0001
PCS4
2015 Compensating the Effect of Communication Delay in Client-Server-Based Shared Haptic Virtual Environments
abstract
Shared haptic virtual environments can be realized using a client-server architecture. In this architecture, each client maintains a local copy of the virtual environment (VE). A centralized physics simulation running on a server calculates the object states based on haptic device position information received from the clients. The object states are sent back to the clients to update the local copies of the VE, which are used to render interaction forces displayed to the user through a haptic device. Communication delay leads to delayed object state updates and increased force feedback rendered at the clients. In this article, we analyze the effect of communication delay on the magnitude of the rendered forces at the clients for cooperative multi-user interactions with rigid objects. The analysis reveals guidelines on the tolerable communication delay. If this delay is exceeded, the increased force magnitude becomes haptically perceivable. We propose an adaptive force rendering scheme to compensate for this effect, which dynamically changes the stiffness used in the force rendering at the clients. Our experimental results, including a subjective user study, verify the applicability of the analysis and the proposed scheme to compensate the effect of time-varying communication delay in a multi-user SHVE.
Clemens Schuwerk, Xiao Xu 0001, Rahul Gopal Chaudhari, Eckehard G. Steinbach
ACM Trans. Appl. Percept.4
2015 A Novel Rate Control Framework for SIFT/SURF Feature Preservation in H.264/AVC Video Compression
abstract
This paper presents a novel rate control framework for H.264/Advanced Video Coding-based video coding that improves the preservation of gradient-based features like scale-invariant feature transform or speeded up robust feature compared with the default rate control algorithm in the JM reference software. First, a criterion (matching score) for feature preservation on the basis of the bag-of-features concept is proposed. Then, the matching scores are collected as a function of the quantization parameters and analyzed for different feature types. With this analysis, macroblocks are categorized into different groups before encoding. Our rate control algorithm assigns different quantization parameters to each group according to the importance of the group for feature extraction. The experimental results show that our rate control algorithm achieves the desired target bit rate, and more features are preserved compared with videos encoded using the default rate control. The proposed approach not only improves feature preservation, but also leads to a noticeable performance improvement in a real image retrieval system. The rate control framework proposed in this paper is fully standard compatible.
Jianshu Chao, Robert Huitl, Eckehard G. Steinbach, Damien Schroeder
IEEE Trans. Circuits Syst. Video Technol.3
2015 QoE-Based Traffic and Resource Management for Adaptive HTTP Video Delivery in LTE
abstract
There is a growing interest in over-the-top (OTT) dynamic adaptive streaming over Hypertext Transfer Protocol (HTTP) (DASH) services. In mobile DASH, a client controls the streaming rate and the base station in the mobile network decides on the resource allocation. Different from the majority of previous works that focus on client-based rate adaptation mechanisms, this paper investigates the mobile network potential for enhancing the user quality-of-experience (QoE) in multiuser OTT DASH. Specifically, we first present proactive and reactive QoE optimization approaches for adapting the adaptive HTTP video delivery in an long-term evolution network. We then show, using subjective experiments, that by taking a proactive role in determining the transmission and streaming rates, the network operator can provide a better video quality and a fairer QoE across the streaming users. Furthermore, we consider the playout buffer time of the clients and propose a novel playout buffer-dependent approach that determines for each client the streaming rate for future video segments according to its buffer time and the achievable QoE under current radio conditions. In addition, we show that by jointly solving for the streaming and transmission rates, the wireless network resources are more efficiently allocated among the users and substantial gains in the user perceived video quality can be achieved.
Ali El Essaili, Damien Schroeder, Eckehard G. Steinbach, Dirk Staehle, Mohammed Shehada
IEEE Trans. Circuits Syst. Video Technol.3
2015 QoE-Based Cross-Layer Optimization for Uplink Video Transmission
abstract
We study the problem of resource-efficient uplink distribution of user-generated video content over fourth-generation mobile networks. This is challenged by (1) the capacity-limited and time-variant uplink channel, (2) the resource-hungry upstreamed videos and their dynamically changing complexity, and (3) the different playout times of the video consumers. To address these issues, we propose a systematic approach for quality-of-experience (QoE)-based resource optimization and uplink transmission of multiuser generated video content. More specifically, we present an analytical model for distributed scalable video transmission at the mobile producers which considers these constraints. This is complemented by a multiuser cross-layer optimizer in the mobile network which determines the transmission capacity for each mobile terminal under current cell load and radio conditions. Both optimal and low-complexity solutions are presented. Simulation results for LTE uplink transmission show that significant gains in perceived video quality can be achieved by our cross-layer resource optimization scheme. In addition, the distributed optimization at the mobile producers can further improve the user experience across the different types of video consumers.
Ali El Essaili, Zibin Wang, Eckehard G. Steinbach
ACM Trans. Multim. Comput. Commun. Appl.3
2015 QoE-Based SVC Layer Dropping in LTE Networks Using Content-Aware Layer Priorities
abstract
The increasing popularity of mobile video streaming applications has led to a high volume of video traffic in mobile networks. As the base station, for instance, the eNB in LTE networks, has limited physical resources, it can be overloaded by this traffic. This problem can be addressed by using Scalable Video Coding (SVC), which allows the eNB to drop layers of the video streams to dynamically adapt the bitrate. The impact of bitrate adaptation on the Quality of Experience (QoE) for the users depends on the content characteristics of videos. As the current mobile network architectures do not support the eNB in obtaining video content information, QoE optimization schemes with explicit signaling of content information have been proposed. These schemes, however, require the eNB or a specific optimization module to process the video content on the fly in order to extract the required information. This increases the computation and signaling overhead significantly, raising the OPEX for mobile operators. To address this issue, in this article, a content-aware (CA) priority marking and layer dropping scheme is proposed. The CA priority indicates a transmission order for the layers of all transmitted videos across all users, resulting from a comparison of their utility versus rate characteristics. The CA priority values can be determined at the P-GW on the fly, allowing mobile operators to control the priority marking process. Alternatively, they can be determined offline at the video servers, avoiding real-time computation in the core network. The eNB can perform content-aware SVC layer dropping using only the priority values. No additional content processing is required. The proposed scheme is lightweight both in terms of architecture and computation. The improvement in QoE is substantial and very close to the performance obtained with the computation and signaling-intensive QoE optimization schemes.
Dirk Staehle, Gerald Kunzmann, Eckehard G. Steinbach, Wolfgang Kellerer
ACM Trans. Multim. Comput. Commun. Appl.4
2014 Graph-based data fusion of pedometer and WiFi measurements for mobile indoor positioning
abstract
We propose a graph-based, low-complexity sensor fusion approach for ubiquitous pedestrian indoor positioning using mobile devices. We employ our fusion technique to combine relative motion information based on step detection with WiFi signal strength measurements. The method is based on the well-known particle filter methodology. In contrast to previous work, we provide a probabilistic model for location estimation that is formulated directly on a fully discretized, graph-based representation of the indoor environment. We generate this graph by adaptive quantization of the indoor space, removing irrelevant degrees of freedom from the estimation problem. We evaluate the proposed method in two realistic indoor environments using real data collected from smartphones. In total, our dataset spans about 20 kilometers in distance walked and includes 13 users and four different mobile device types. Our results demonstrate that the filter requires an order of magnitude less particles than state-of-the-art approaches while maintaining an accuracy of a few meters. The proposed low-complexity solution not only enables indoor positioning on less powerful mobile devices, but also saves much-needed resources for location-based end-user applications which run on top of a localization service.
Sebastian Hilsenbeck, Dmytro Bobkov, Georg Schroth, Robert Huitl, Eckehard G. Steinbach
UbiComp5
2014 Bit rate estimation for H.264/AVC video encoding based on temporal and spatial activities
abstract
We present a novel bit rate model for H.264/AVC video encoding which is based on the quantization parameter, the frame rate as well as temporal and spatial activity measures. With the proposed model, it is possible to trade-off the frame rate versus the quantization parameter to achieve a target bit rate. Our model depends on video activity measures that can be easily calculated from the uncompressed video. In our experiments, the model achieves a Pearson correlation of 0.99 and a root-mean-square error of less than 5% with the measured bit rate values, as verified by statistical analysis.
Christian Lottermann, Alexander Machado, Damien Schroeder, Eckehard G. Steinbach
ICIP5
2014 Camera-based indoor positioning using scalable streaming of compressed binary image signatures
abstract
Recent progress in the field of content-based image retrieval has enabled camera-based indoor positioning. The matching of smart-phone recordings with a database of geo-referenced images allows for meter accurate infrastructure-free localization. In mobile scenarios, however, three major constraints have to be considered: limited computational resources of mobile devices, limited network capacity and the need for scalability in large buildings. To address these issues, we modify the state-of-the-art Vector of Locally Aggregated Descriptors (VLAD) image signature to work with recently emerging binary feature descriptors. We show that this results in a substantial reduction in the overall computational complexity, which enables the matching of image signatures directly on the mobile device. The specific properties of this signature form the basis of our proposed scalable streaming approach that preemptively loads image signatures of reference images in the vicinity of the user onto the mobile device to mitigate the effect of network latency. In order to provide efficient streaming, we compress the signatures by exploiting the similarities of spatially neighboring reference images. In combination, the contributions of this paper lead to an indoor localization system, which allows instantaneous camera-based indoor positioning with very low requirements on the available network connection.
Dominik Van Opdenbosch, Georg Schroth, Robert Huitl, Sebastian Hilsenbeck, Adrian Garcea, Eckehard G. Steinbach
ICIP6
2014 Camera context based estimation of spatial and temporal activity parameters for video quality metrics in automotive applications
abstract
We present a low complexity approach for the estimation of the temporal and spatial activity parameters of videos which are captured by a front-facing camera of a vehicle based on context information of the vehicle. The estimated parameters are integrated into an objective video quality metric, which can be used to determine the perceptual quality of a compressed video stream. Our proposed video quality metric has very low computational complexity, which makes it suitable for live video streaming applications. It shows a high Pearson correlation of 0.98 with an average root-mean-square error of 6%, as verified by statistical analysis with data from subjective tests.
Christian Lottermann, Alexander Machado, Damien Schroeder, Wolfgang Hintermaier, Eckehard G. Steinbach
ICME5
2014 Energy-efficient and QoE-driven adaptive HTTP streaming over LTE
abstract
LTE networks offer broadband wireless access to mobile users who can benefit from high data rate applications such as video streaming. In order to improve the user satisfaction, Quality-of-Experience (QoE) based resource allocation for multiple streaming users in an LTE cell has been studied. However, the high energy consumption of these applications has not been considered. In this paper, we propose adaptive Discontinuous Reception (DRX) parameters for LTE that reduce the energy consumption of mobile devices without degrading the video quality of adaptive HTTP streaming users. Furthermore, we extend the QoE-optimized resource allocation by additionally considering the power consumption of the mobile devices. Simulation results show the benefits of using the proposed adaptive DRX parameters and that further energy saving gains can be achieved by including the power consumption in the optimization problem.
Damien Schroeder, Ali El Essaili, Eckehard G. Steinbach
WCNC4
2014 MDVQM: A novel multidimensional no-reference video quality metric for video transcoding
Fan Zhang 0026, Eckehard G. Steinbach
J. Vis. Commun. Image Represent.2
2013 Dynamic model displacement for model-mediated teleoperation
abstract
In this paper, we study and extend the concept of model-mediated teleoperation (MMT) for teleaction systems which provide live video feedback from the remote side with a time delay. In MMT, the haptic feedback is rendered locally on the operator side using a simple object surface model in order to keep the haptic control loop stable in the presence of communication delays. Because the live video from the remote side is received with delay, this results in a visual-haptic asynchrony for the displayed interaction events. In addition, sudden model parameter updates can lead to “model-jump” effects for the displayed haptic feedback. Both effects degrade the user experience and system performance. To address these issues, we propose an extension of MMT which we call model-displaced teleoperation (MDT) in this paper. In MDT, we adaptively shift the position of the local surface model to delay the haptic contact with the environment, thus compensating the visual-haptic asynchrony and avoiding the model-jump effect. As the haptic feedback is still rendered locally, the advantages of the MMT approach are retained and instabilities in the haptic interaction are avoided. In our experiments, we determine the optimal displacement compromise between visual-haptic asynchrony, the model-jump effect and perceived distance errors. Moreover, the subjective experience and objective task performance of the proposed MDT and the original MMT for a teleoperation setup with soft objects are evaluated. Our results show that the users prefer the MDT method compared to MMT once the communication delay between the teleoperator and the operator exceeds 50ms. In addition, the task error rate is reduced by about 50% and the subjects are better able to control their contact force for system delays larger than 50ms if the MDT method is employed.
Xiao Xu 0001, Giulia Paggetti, Eckehard G. Steinbach
World Haptics3
2013 Quality-of-experience driven adaptive HTTP media delivery
abstract
This paper presents a Quality of Experience (QoE) driven approach for multi-user resource optimization in Dynamic Adaptive Streaming over HTTP (DASH) over next generation wireless networks. Our objective is to enhance the user experience in adaptive HTTP streaming by jointly considering the characteristics of the media content and the available wireless resources in the operator network. Specifically, we propose a proactive QoE-based approach for rewriting the client HTTP requests at a proxy in the mobile network. The advantage of the proposed approach is its applicability for over-the-top (OTT) streaming as it requires no adaptation of the media content. We compare our proposed scheme to both reactive QoE-optimized and to standard-DASH HTTP streaming. Our contributions are: 1) We first show that standard OTT DASH leads to unsatisfactory performance since the content agnostic resource allocation by the LTE scheduler is far from optimal, and we can achieve a clear QoE improvement when considering the content characteristics. 2) We additionally show that proactively rewriting the client requests gives control of the video content adaptation to the network operator which has better information than the client on the load and radio conditions in the cell. This results in additional gains in user perceived video quality. 3) A standard unmodified DASH client remains unaware of the proposed rewriting of the HTTP requests and can decode and play the redirected media segments.
Ali El Essaili, Damien Schroeder, Dirk Staehle, Mohammed Shehada, Wolfgang Kellerer, Eckehard G. Steinbach
ICC6
2013 Reconstruction of transparent objects in unstructured scenes with a depth camera
abstract
The visual 3D reconstruction of transparent objects in unstructured scenes is challenging due to the complex image formation principles underlying their visual appearance. Most state-of-the-art reconstruction methods ignore this problem and assume Lambertian reflection. Yet, transparent objects are relevant scene information for applications in intelligent robotics (such as grasping) or virtual reality. In this work, we present an approach to detect non-planar transparent objects, like bottles or glasses, by specifically searching for geometry inconsistencies caused by refraction or reflection. Depth information is acquired using a Kinect sensor, which is moved within the scene in order to acquire multiple views. The individual measurements are combined into a 3D volume, yielding the objects' location and a rough shape estimate. Results are presented using various household objects made of glass or plastic.
Nicolas Alt, Patrick Rives, Eckehard G. Steinbach
ICIP3
2013 On the design of a novel JPEG quantization table for improved feature detection performance
abstract
Keypoint or interest point detection is the first step in many computer vision algorithms. The detection performance of the state-of-the-art detectors is, however, strongly influenced by compression artifacts, especially at low bit rates. In this paper, we design a novel quantization table for the widely-used JPEG compression standard which leads to improved feature detection performance. After analyzing several popular scale-space based detectors, we propose a novel quantization table which is based on the observed impact of scale-space processing on the DCT basis functions. Experimental results show that the novel quantization table outperforms the JPEG default quantization table in terms of feature repeatability, number of correspondences, matching score, and number of correct matches.
Jianshu Chao, Eckehard G. Steinbach
ICIP3
2013 Speeded-up SURF: Design of an efficient multiscale feature detector
abstract
We present a fast and highly performant multiscale feature detector which is based on the established SURF algorithm. The additional speed-up is achieved by linearizing the SURF detector in a way that its detection characteristics are preserved. Our evaluations show that the proposed suSURF detector is roughly 30% faster without significant sacrifices in feature quality. The key points detected by suSURF are compatible with standard SURF features and those found by other determinant of Hessian (DoH) based detectors.
Florian Schweiger, Georg Schroth, Robert Huitl, Yasir Latif, Eckehard G. Steinbach
ICIP5
2013 Fast relocalization for visual odometry using binary features
abstract
State-of-the-art visual odometry algorithms achieve remarkable efficiency and accuracy. Under realistic conditions, however, tracking failures are inevitable and to continue tracking, a recovery strategy is required. In this paper, we propose a relocalization system that enables realtime, 6D pose recovery for wide baselines. Our approach targets specifically resource-constrained hardware such as mobile phones. By exploiting the properties of low-complexity binary feature descriptors, nearest-neighbor search is performed efficiently using Locality Sensitive Hashing. Our method does not require time-consuming offline training of hash tables and it can be applied to any visual odometry system. We provide a thorough evaluation of effectiveness, robustness and runtime on an indoor test sequence with available ground truth poses. We investigate the system parameterization and compare the relocalization performance for the three binary descriptors BRIEF, unscaled BRIEF and ORB. In contrast to previous work on mobile visual odometry, we are able to quickly recover from tracking failures within maps with thousands of 3D feature points.
Jeremy Straub, Sebastian Hilsenbeck, Georg Schroth, Robert Huitl, Andreas Möller, Eckehard G. Steinbach
ICIP6
2013 Towards the design of an intuitive multi-view video navigation interface based on spatial information
abstract
A Multi-View Video (MVV) is a set of related videos that capture an interesting scene from different perspectives at overlapping times. The work at hand is concerned with the design of innovative user interfaces (UI) for viewing MVVs. As a first contribution four different MVV UIs are designed. While different in design, their common aim is to allow a pleasant viewing and perspective switching experience by reducing the cognitive effort associated with constructing a mental map of the scene. This is achieved by incorporating the spatial relationships of the available views in the UI elements. As a second contribution a quality model is developed and a methodical evaluation process is designed. This is used to evaluate and compare the UIs. In a third contribution we use principal component analysis (PCA) to reveal information about the perceptual quality space which helps validating our proposed quality model. Based on the findings, a series of conclusions for best design practices are provided.
Silviu Apostu, Anas Al-Nuaimi, Eckehard G. Steinbach, Michael Fahrmair, Xiaohang Song, Andreas Möller
Mobile HCI3
2013 Performance comparison of various feature detector-descriptor combinations for content-based image retrieval with JPEG-encoded query images
abstract
We study the impact of JPEG compression on the performance of an image retrieval system for different feature detector-descriptor combinations. The VLBenchmarks retrieval framework is used to compare a total of 60 detector-descriptor combinations for a dataset with JPEG-encoded query images. Our results show that among all tested detectors, the Hessian-Affine detector leads to the most robust performance in the presence of strong JPEG compression. Additionally, we compare the retrieval gains of the different detector-descriptor pairs after processing the JPEG-encoded query images with different deblocking filters. The results illustrate that for the MSER, MFD and WαSH detectors, the retrieval results benefit from two of the deblocking approaches at low bit rate irrespective of what descriptor the detectors are combined with. The same two deblocking filters are found to increase the retrieval performance for the MROGH descriptor when combined with most of the tested detectors.
Jianshu Chao, Anas Al-Nuaimi, Georg Schroth, Eckehard G. Steinbach
MMSP4
2013 Fully Automatic and Frame-Accurate Video Synchronization Using Bitrate Sequences
abstract
Video synchronization is an essential processing step in many multimedia applications, and various methods have been proposed in the literature each of which addresses the problem from a different point of vantage. In this article, we present an information theoretic approach to video synchronization, based on the state-of-the-art in hybrid video coding. Time series derived from the videos' instantaneous bitrate demand are correlated in a robust manner employing the recently published ConCor algorithm. We enhance ConCor with integrated normalization capabilities in order to improve its shape-oriented matching performance. Furthermore, we present a mathematical framework to derive the most suitable ConCor parameters given a specific class of input videos. In an extensive experimental analysis, we give an insight into the representation of synchronization-relevant scene changes with bitrate data, and examine the influence of encoding parameters on the synchronization performance. Experiments on diverse video input substantiate the reliable performance of our easy to implement, yet effective video synchronization algorithm which distinguishes itself in that it operates largely without manual intervention.
Florian Schweiger, Georg Schroth, Michael Eichhorn, Anas Al-Nuaimi, Burak Cizmeci, Michael Fahrmair, Eckehard G. Steinbach
IEEE Trans. Multim.7
2012 Exploiting prior knowledge in mobile visual location recognition
abstract
Mobile visual location recognition needs to be performed in real-time for location based services to be perceived as useful. We describe and validate an approach that eliminates the network delay by preloading partial visual vocabularies to the mobile device. Retrieval performance is significantly increased by composing partial vocabularies based on the uncertainty about the location of the client. This way, prior knowledge is efficiently integrated into the matching process. Based on compressed feature sets, infrequently uploaded from the mobile device, the server estimates the client location and its uncertainty by fusing consecutive query results using a particle filter.
Georg Schroth, Robert Huitl, Mohammad Abu-Alqumsan, Florian Schweiger, Eckehard G. Steinbach
ICASSP5
2012 SIFT feature-preserving bit allocation for H.264/AVC video compression
abstract
Compression artifacts in low-quality videos strongly influence the performance of feature matching algorithms. In order to achieve reasonable feature matching performance even for low bit rate video, we propose to allocate the bit budget during compression such that the important features are preserved. Specifically, we present two bit allocation approaches to preserve the strongest SIFT features for H.264 encoded videos. For both approaches, we first categorize the Macroblocks in a Group of Pictures into several groups according to the scale specific characteristics of SIFT features. In our first approach a novel R-D model based on the matching score is applied to allocate the bit budget to these groups. In our second approach, in order to reduce the computational complexity, we analyze the detector characteristics of correctly matched pairs and propose a R-D optimization method based on the repeatability metric. Our experiments show that both approaches achieve better feature preservation when compared to standard video encoding which is optimized for maximum picture quality. The proposed approaches are fully standard compatible and the encoded videos can be decoded by any H.264 decoder.
Jianshu Chao, Eckehard G. Steinbach
ICIP2
2012 TUMindoor: An extensive image and point cloud dataset for visual indoor localization and mapping
abstract
Recent advances in the field of content-based image retrieval (CBIR) have made it possible to quickly search large image databases using photographs or video sequences as a query. With appropriately tagged images of places, this technique can be applied to the problem of visual location recognition. While this task has attracted large interest in the community, most existing approaches focus on outdoor environments only. This is mainly due to the fact that the generation of an indoor dataset is elaborate and complex. In order to allow researchers to advance their approaches towards the challenging field of CBIR-based indoor localization and to facilitate an objective comparison of different algorithms, we provide an extensive, high resolution indoor dataset. The free for use dataset includes realistic query sequences with ground truth as well as point cloud data, enabling a localization system to perform 6-DOF pose estimation.
Robert Huitl, Georg Schroth, Sebastian Hilsenbeck, Florian Schweiger, Eckehard G. Steinbach
ICIP5
2012 A Quality-of-Experience driven bidding game for uplink video transmission in next generation mobile networks
abstract
Centralized approaches to solve resource allocation problems for wireless real-time multimedia communications have been intensively studied but require the availability of meta information about the multimedia content and channel information of all users. In this paper, we propose a Quality of Experience (QoE) driven bidding game for de-centralized uplink resource allocation among multiple mobile video producers. Different from previous works, the price per resource unit is defined on a Mean Opinion Score (MOS) scale and users bid for the resources that maximize their own utility function. Simulations in an LTE environment show the benefits of our distributed approach in terms of convergence time and QoE performance, compared to a centralized greedy scheme and a state-of-the-art game-theoretic approach.
Damien Schroeder, Ali El Essaili, Eckehard G. Steinbach, Zoran Despotovic, Wolfgang Kellerer
ICIP3
2012 Real-time compression of point cloud streams
abstract
We present a novel lossy compression approach for point cloud streams which exploits spatial and temporal redundancy within the point data. Our proposed compression framework can handle general point cloud streams of arbitrary and varying size, point order and point density. Furthermore, it allows for controlling coding complexity and coding precision. To compress the point clouds, we perform a spatial decomposition based on octree data structures. Additionally, we present a technique for comparing the octree data structures of consecutive point clouds. By encoding their structural differences, we can successively extend the point clouds at the decoder. In this way, we are able to detect and remove temporal redundancy from the point cloud data stream. Our experimental results show a strong compression performance of a ratio of 14 at 1 mm coordinate precision and up to 40 at a coordinate precision of 9 mm.
Julius Kammerl, Nico Blodow, Radu Bogdan Rusu, Suat Gedikli, Michael Beetz, Eckehard G. Steinbach
ICRA6
2012 Beyond classical teleoperation: Assistance, cooperation, data reduction, and spatial audio
abstract
In this video we present a teleoperation system which is capable of solving complex tasks in human-sized wide area environments. The system consists of two mobile teleoperators controlled by two operators, and offers haptic, visual, and auditory feedback. The task examined here, consists of repairing a robot by removing a computer and replacing a defective hard-drive. To cope with the complexity of such a task, we go beyond classical teleoperation by integrating several advanced software algorithms into the system.
Thomas Schauss, Carolina Passenberg, Nikolay Stefanov, Daniela Feth, Iason Vittorias, Angelika Peer, Sandra Hirche, Martin Buss, Martin Rothbucher, Klaus Diepold, Julius Kammerl, Eckehard G. Steinbach
ICRA12
2012 Scale-preserving long-term visual odometry for indoor navigation
abstract
We present a visual odometry system for indoor navigation with a focus on long-term robustness and consistency. As our work is targeting mobile phones, we employ monocular SLAM to jointly estimate a local map and the device's trajectory. We specifically address the problem of estimating the scale factor of both, the map and the trajectory. State-of-the-art solutions approach this problem with an Extended Kalman Filter (EKF), which estimates the scale by fusing inertial and visual data, but strongly relies on good initialization and takes time to converge. Each visual tracking failure introduces a new arbitrary scale factor, forcing the filter to re-converge. We propose a fast and robust method for scale initialization that exploits basic geometric properties of the learned local map. Using random projections, we efficiently compute geometric properties from the feature point cloud produced by the visual SLAM system. From these properties (e.g., corridor width or height) we estimate scale changes caused by tracking failures and update the EKF accordingly. As a result, previously achieved convergence is preserved despite re-initializations of the map. To minimize the time required to continue tracking after failure, we perform recovery and re-initialization in parallel. This increases the time available for recovery and hence the likelihood for success, thus allowing almost seamless tracking. Moreover, fewer re-initializations are necessary. We evaluate our approach using extensive and diverse indoor datasets. Results demonstrate that errors and convergence times for scale estimation are considerably reduced, thus ensuring consistent and accurate scale estimation. This enables long-term odometry despite of tracking failures which are inevitable in realistic scenarios.
Sebastian Hilsenbeck, Andreas Möller, Robert Huitl, Georg Schroth, Matthias Kranz, Eckehard G. Steinbach
IPIN6
2012 Low bitrate source-filter model based compression of vibrotactile texture signals in haptic teleoperation
abstract
Vibrotactile signals convey the touch-based characteristics of object surfaces felt through a tool. They particularly enhance the quality of human-machine interactions by providing realistic haptic perception of textures. In this paper, inspired by the similarities observed between vibrotactile texture signals and speech signals, we present a novel vibrotactile texture codec for bilateral teleoperation, based on well-known speech coding techniques. The proposed low bitrate, high quality codec preserves not only the spectral signature vital to the general feel of the texture, but also important temporal features of the texture signal. We report a compression ratio of 8:1 (12.5 %) with a constant output bitrate of 4 kbps, and we validate the perceptual transparency of the codec via rigorous subjective tests and analyses.
Rahul Gopal Chaudhari, Burak Cizmeci, Katherine J. Kuchenbecker, Seungmoon Choi, Eckehard G. Steinbach
ACM Multimedia5
2012 Virtual reference view generation for CBIR-based visual pose estimation
abstract
Determining the pose of a mobile device based on visual information is a promising approach to solve the indoor localization problem. We present an approach that transforms localized images along a mapping trajectory into virtual viewpoints that cover a set of densely sampled camera positions and orientations in a confined environment. The viewpoints are represented by their respective bag-of-features vectors and image retrieval techniques are applied to determine the most likely pose of query images at very low computational complexity. As virtual image locations and orientations are decoupled from actual image locations, the system is able to work with sparse reference imagery and copes well with perspective distortion. Experiments confirm that pose retrieval performance is significantly improved.
Robert Huitl, Georg Schroth, Sebastian Hilsenbeck, Florian Schweiger, Eckehard G. Steinbach
ACM Multimedia5
2012 ConCor+: Robust and confident video synchronization using consensus-based Cross-correlation
abstract
Consensus-based Cross-correlation (ConCor) is a recently presented algorithm for robust synchronization of noisy and corrupted signals. ConCor has a number of interdependent parameters that need to be set correctly to guarantee good performance. In this paper we analyse the effects of the individual parameters on ConCor's behaviour and performance. As a second contribution, we show that a parameter sweep with subsequent majority voting can be used to boost ConCor's performance and produce a trustworthy confidence measure. As a final contribution we show how the proposed extension also allows performing multi-modal (joint audio-video) synchronization of casual multi-perspective video recordings enabling superior matching performance.
Anas Al-Nuaimi, Burak Cizmeci, Florian Schweiger, Roman Katz, Sinan Taifour, Eckehard G. Steinbach, Michael Fahrmair
MMSP6
2012 Haptic Communications
abstract
Audiovisual communications is at the core of multimedia systems that allow users to interact across distances. It is common understanding that both audio and video are required for high-quality interaction. While audiovisual information provides a user with a satisfactory impression of being present in a remote environment, physical interaction and manipulation is not supported. True immersion into a distant environment and efficient distributed collaboration require the ability to physically interact with remote objects and to literally get in touch with other people. Touching and manipulating objects remotely becomes possible if we augment traditional audiovisual communications by the haptic modality. Haptic communications is a relatively young field of research that has the potential to substantially improve human-human and human-machine interaction. In this paper, we discuss the state-of-the-art in haptic communications both from psychophysical and technical points of view. From a human perception point of view, we mainly focus on the multimodal integration of video and haptics and the improved performance that can be achieved when combining them. We also discuss how the human adapts to discrepancies and synchronization errors between different modalities, a research area which is typically referred to as perceptual learning. From a technical perspective, we address perceptual coding of haptic information and the transmission of haptic data streams over resource-constrained and potentially lossy networks in the presence of unpredictable and time-varying communication delays. In this context, we also discuss the need for objective quality metrics for haptic communication. Throughout the paper, we stress the fact that haptic communications is not meant as a replacement of traditional audiovisual communications but rather as an additional dimension for telepresence that will allow us to advance in our quest for truly immersive communication.
Eckehard G. Steinbach, Sandra Hirche, Marc O. Ernst, Fernanda Brandi, Rahul Gopal Chaudhari, Julius Kammerl, Iason Vittorias
Proc. IEEE1
2011 Towards an objective quality evaluation framework for haptic data reduction
abstract
High packet rates in telepresence and teleaction systems pose grave challenges to teleoperation over existing communication infrastructure like the Internet. To counter these issues, efficient perceptually motivated packet-rate reduction schemes have been developed. These schemes are conventionally evaluated for perceived quality via subjective user tests. Such tests are time-consuming, expensive and require precise control of experimental conditions. Computer modeling of telepresence sessions can, on the other hand, bring repeatability, ease of observation, definite control over system parameters and task description and fairness of comparison. In this paper, we present first steps towards a methodology and a framework to model and simulate a networked haptic interaction and evaluate it objectively for the quality of experience. Towards this purpose, we model the human control action and haptic perception process in teleoperation. Our results show that simulations of these models for a range of data reduction scheme parameters produce quality estimates whose trend is comparable to carefully performed subjective user tests.
Rahul Gopal Chaudhari, Eckehard G. Steinbach, Sandra Hirche
World Haptics2
2011 Joint source-channel rate control for pixel-domain distributed video coding
abstract
We study the scenario of pixel-domain distributed video coding for noisy transmission environments and propose a method to allocate the available rate between source coding and channel coding to generate a robust video stream. Having observed in experiments the uncertainty of the source and the channel coding rate, we model them as random variables via offline training, estimate the decoding failure probability and calculate the mean end-to-end distortion. Adaptive quantization is performed for each slice to minimize its mean end-to-end distortion. With this joint source-channel rate allocation, we compare the robustness of two coding prototypes, namely distributed video coding and distributed video coding with forward error correction. According to our experimental results, under same total bit budget, the distributed video coding only scheme proves more robust than the latter one and the gain is up to 1 dB in PSNR.
Eckehard G. Steinbach, Chang Wen Chen
ICASSP2
2011 Rapid image retrieval for mobile location recognition
abstract
Recognizing the location and orientation of a mobile device from captured images is a promising application of image retrieval algorithms. Matching the query images to an existing georeferenced database like Google Street View enables mobile search for location related media, products, and services. Due to the rapidly changing field of view of the mobile device caused by constantly changing user attention, very low retrieval times are essential. These can be significantly reduced by performing the feature quantization on the handheld and transferring compressed Bag-of-Feature vectors to the server. To cope with the limited processing capabilities of handhelds, the quantization of high dimensional feature descriptors has to be performed at very low complexity. To this end, we introduce in this paper the novel Multiple Hypothesis Vocabulary Tree (MHVT) as a step towards real-time mobile location recognition. The MHVT increases the probability of assigning matching feature descriptors to the same visual word by introducing an overlapping buffer around the separating hyperplanes to allow for a soft quantization and an adaptive clustering approach. Further, a novel framework is introduced that allows us to integrate the probability of correct quantization in the distance calculation using an inverted file scheme. Our experiments demonstrate that our approach achieves query times reduced by up to a factor of 10 when compared to the state-of-the-art.
Georg Schroth, Anas Al-Nuaimi, Robert Huitl, Florian Schweiger, Eckehard G. Steinbach
ICASSP5
2011 Improved ρ-domain rate control with accurate header size estimation
abstract
p-domain rate control has been shown to be a simple yet effective rate control approach for DCT-based hybrid video codecs. When it is applied to H.264/AVC, improvements need to be made because of the large amount of header information and the QP-dependent Rate Distortion Optimization (RDO). In this paper, we propose a novel rate model to estimate the size of header information accurately. We present a two-stage rate control algorithm which combines the proposed header rate model and the p-domain source model. In comparison to the original p-domain rate control algorithm without header size estimation, our scheme not only improves the PSNR of the decoded video, but also achieves the target bit rates more accurately.
Fan Zhang 0026, Eckehard G. Steinbach
ICASSP2
2011 QoE-Based Cross-Layer Optimization of Wireless Video with Unperceivable Temporal Video Quality Fluctuation
abstract
This paper proposes a novel approach for Quality of Experience (QoE) driven cross-layer optimization for wireless video transmission. We formulate the cross-layer optimization problem with a constraint on the temporal fluctuation of the video quality. Our objective is to minimize the temporal change of the video quality as perceivable quality fluctuations negatively affect the overall quality of experience. The proposed QoE scheme jointly optimizes the application layer and the lower layers of a wireless protocol stack. It allocates network resources and performs rate adaptation such that the fluctuations lie within the range of unperceivable changes. We determine corresponding perception thresholds via extensive subjective tests and evaluate the proposed scheme using an OPNET High Speed Downlink Packet Access (HSDPA) emulator. Our simulation results show that the proposed approach leads to a noticeable improvement of overall user satisfaction for the provided video delivery service when compared to state-of-the-art approaches.
Srisakul Thakolsri, Wolfgang Kellerer, Eckehard G. Steinbach
ICC3
2011 Preserving SIFT features in JPEG-encoded images
abstract
For image compression applications where the information sink is not a person but a computer algorithm, the image encoder should control the encoding process in such a way that the important and relevant features of the image are preserved after compression. In this paper, our goal is to preserve the strongest SIFT features for JPEG-encoded images. We analyze the relevant characteristics of SIFT features and categorize the image Macroblocks into several groups. Then we propose a novel rate-distortion model which is based on the SIFT feature matching score. The dependency between the quantization table in the JPEG file and the common Lagrange multiplier is obtained from a training image database. Then for a given image quality we exploit this relationship to perform R-D optimization for each group. Our results show that the proposed algorithm achieves better feature preservation when compared to standard JPEG encoding. The proposed approach is fully standard compatible.
Jianshu Chao, Eckehard G. Steinbach
ICIP2
2011 A comparison of the error resiliency of bit-plane based and symbol based pixel-domain distributed video coding
abstract
This work studies the error resilience of pixel-domain distributed video coding in noisy wireless transmission environment. Turbo codes are used to implement the DVC coder and the AWGN model is assumed for the transmission channel. The goal is to find out whether symbol based coding or bit-plane based coding is more robust against channel noise. We compare the two in the context of joint source-channel coding to ensure a high error resilience. First, we propose a framework to estimate the end-to-end distortion for these two schemes. Next, we allocate the rate between source coding and channel coding, aiming at a minimum end-to-end distortion. Then, we simulate the two schemes, setting the same bit budget for them, and compare their error resilience. Experimental results show that symbol based coding outperforms bit-plane based coding by up to 0.7 dB in PSNR of the decoded video.
Eckehard G. Steinbach, Chang Wen Chen
ICIP2
2011 QoE-driven resource optimization for user generated video content in next generation mobile networks
abstract
The increasing popularity of user-generated content and the high quality upstreaming capabilities of mobile phones indicate a prevalence of video traffic in the uplink of next generation mobile net works. Need arises for optimizing the network resource allocation while preserving the user satisfaction. In this paper, we propose a service-centric approach for uplink distribution of real-time user generated content based on the Quality of Experience (QoE) and popularity of the video content. In case of limited network resources, the proposed approach assigns more resources for popular contents while maintaining a minimum guaranteed QoE for the less popular ones. We compare our service-centric approach with a QoE-driven one that does not consider video popularity and evaluate both approaches for the uplink of an LTE system. The simulation results show that a significant gain in terms of average user satisfaction can be achieved.
Ali El Essaili, Eckehard G. Steinbach, Daniele Munaretto, Srisakul Thakolsri, Wolfgang Kellerer
ICIP2
2011 Image-based object detection under varying illumination in environments with specular surfaces
abstract
Image-based environment representations capture the appearance of the surroundings of a mobile robot and are useful for the detection of novelty. However, image-based novelty detection can be impaired by illumination effects. In this paper we present an approach for the image-based detection of novel objects in a scene under varying lighting conditions and in the presence of objects with specular surfaces. The computation of an illumination-invariant image-based environment representation allows for the extraction of the shading of the environment from camera images. Using statistical models infered from the luminance and the saturation component of the shading images, specularities and shadows are detected and suppressed in the process of novelty detection. Experimental results show that the proposed method outperforms two recently presented reference approaches for illumination-invariant change detection in images.
Werner Maier 0001, Michael Eschey, Eckehard G. Steinbach
ICIP3
2011 A novel full-reference video quality metric and its application to wireless video transmission
abstract
In this paper, we present a novel objective video quality metric that captures the trade-off between the picture quality and the temporal resolution of a compressed video. The proposed metric is based on PSNR, frame rate as well as spatial and temporal activity measures that are obtained from the video. The content-independency of the metric makes it useful for the dynamic optimization of wireless video transmission. With the proposed metric, it is possible to adjust the trade-off between spatial and temporal qualities such that the user satisfaction is maximized. Our metric is very accurate, as verified by statistical analysis with data from subjective tests. We integrate the metric into a real-time wireless video transmission system and show in our experiments that the system, with our metric's ability to predict perceptual quality, can deliver significantly improved perceptual quality for arbitrary videos over a wide range of channel conditions.
Eckehard G. Steinbach
ICIP2
2011 Synchronization of presentation slides and lecture videos using bit rate sequences
abstract
The temporal synchronization of presentation slides and lecture videos enables us to enhance the user experience of online lecture viewing systems. In this work, we present a novel approach to robustly and reliably detect and recognize slide changes, which is based on bit rate sequences. By exploiting readily available information, the computational complexity of the approach is very low. In contrast to prior work, no requirements on the amount or type of texture, motion of foreground objects, or text size on the slide are imposed. Experimental results show the ability to detect even minor slide changes and demonstrate the robustness against occlusions, foreground motion, and camera motion.
Georg Schroth, Ngai-Man Cheung, Eckehard G. Steinbach, Bernd Girod
ICIP3
2011 HTTP-based scalable video streaming over mobile networks
abstract
This paper proposes an adaptive HTTP-based video streaming framework for mobile networks using the Scalable Video Coding (SVC) extension to the H.264 standard. We present a method to statistically estimate the channel at the mobile client and use it in our work to adapt the bit rate of the video. The adaptation takes into account both the quality contributions and the probability of successful timely decoding for different video segments. The simulation results show significant improvements in terms of a reduction of playback interruptions and improved perceived quality of service.
Ktawut Tappayuthpijarn, Thomas Stockhammer, Eckehard G. Steinbach
ICIP3
2011 Surprise-driven acquisition of visual object representations for cognitive mobile robots
abstract
Robots in a household environment have to execute a variety of tasks including carrying objects. In order to grasp the correct objects for a desired action it is indispensable that the robot is able to recognize the objects in its environment. From time to time, the robot will encounter new unknown objects which it has never seen before. In order to recognize them in a later task the robot has to acquire an internal representation of them. In this paper, we present an approach for the autonomous acquisition of visual object representations in a cluttered environment. Guided by surprise, the robot detects novel objects in a familiar environment, selects local image features which represent their appearance and stores them in a database. Experimental results show that our method for the detection of surprising events reliably directs the robot's attention to the novel objects and that the recognition behavior based on our acquired object representations outperforms a state-of-the-art approach.
Werner Maier 0001, Eckehard G. Steinbach
ICRA2
2011 Exploiting Text-Related Features for Content-based Image Retrieval
abstract
Distinctive visual cues are of central importance for image retrieval applications, in particular, in the context of visual location recognition. While in indoor environments typically only few distinctive features can be found, outdoors dynamic objects and clutter significantly impair the retrieval performance. We present an approach which exploits text, a major source of information for humans during orientation and navigation, without the need for error-prone optical character recognition. To this end, characters are detected and described using robust feature descriptors like SURF. By quantizing them into several hundred visual words we consider the distinctive appearance of the characters rather than reducing the set of possible features to an alphabet. Writings in images are transformed to strings of visual words termed visual phrases, which provide significantly improved distinctiveness when compared to individual features. An approximate string matching is performed using N-grams, which can be efficiently combined with an inverted file structure to cope with large datasets. An experimental evaluation on three different datasets shows significant improvement of the retrieval performance while reducing the size of the database by two orders of magnitude compared to state-of-the-art. Its low computational complexity makes the approach particularly suited for mobile image retrieval applications.
Georg Schroth, Sebastian Hilsenbeck, Robert Huitl, Florian Schweiger, Eckehard G. Steinbach
ISM5
2011 A novel real-time video data scheduling approach for driver assistance services
abstract
Traffic shaping and data scheduling are important components of an IP-camera based in-vehicle sensor network for driver assistance services. To this end, we present an IP-camera specific traffic shaping scheme as well as a service specific data scheduling approach. In order to determine a suitable configuration of the service specific scheduling approach, relevant driver assistance services are described and analyzed. In our work we assume that the video data is carried by an in-car switched IP/Ethernet network. As an implementation study, an IP-camera sensor network has been developed in order to investigate the proposed traffic shaping as well as data scheduling mechanisms. Our experimental results show that the IEEE 1588 v2 standard is suitable for the proposed data scheduling approach and that as a result reduced network resource overprovisioning is required.
Wolfgang Hintermaier, Eckehard G. Steinbach
Intelligent Vehicles Symposium2
2011 Perceptual coding of recorded telemanipulation sessions
abstract
Efficient recording and replay of telemanipulation sessions require the compression of the occurring visual-haptic signals. To this end, we propose two distinct approaches which take into consideration known limitations of human visual and haptic perception. We compare both schemes to the state-of-the-art approaches and show that a substantial additional data reduction can be achieved while maintaining good overall haptic and visual performances.
Fernanda Brandi, Eckehard G. Steinbach
ACM Multimedia2
2011 Consensus-based cross-correlation
abstract
Cross-correlation is a classical similarity measure with broad applications in multimedia signal processing. While it is robust against uncorrelated noise in the input signals, it is severely affected by systematic disturbances which lead to biased results. To overcome this limitation, we propose in this paper consensus-based cross-correlation (ConCor) to deal with heavily corrupted signal parts that derail regular cross-correlation. ConCor builds upon the widely adopted RANSAC algorithm to reliably identify and eliminate corrupt signal parts at limited additional complexity. Our approach is universal in that it can be combined with existing cross-correlation variants. We apply ConCor in two example applications, namely video synchronization and template matching. Our experimental results demonstrate the improved robustness and accuracy when compared to classical cross-correlation.
Florian Schweiger, Georg Schroth, Michael Eichhorn, Eckehard G. Steinbach, Michael Fahrmair
ACM Multimedia4
2011 QoE-driven live and on-demand LTE uplink video transmission
abstract
We consider the joint upstreaming of live and on-demand user-generated video content over LTE using a Quality-of-Experience driven approach. We contribute to the state-of-the-art work on multimedia scheduling in three aspects: 1) we jointly optimize the transmission of live and time-shifted video under scarce uplink resources by transmitting a basic quality in realtime and uploading a refined quality for on-demand consumption. 2) We propose a producer-consumer deadline-aware scheduling algorithm that incorporates both the physical state of the mobile producer (e.g., cache fullness) and the scheduled playout time at the end-user. 3) We show that the scheduling decisions in 1) and 2) can be determined locally for each mobile producer. We additionally present an analytical framework for de-centralized scalable video transmission and prove that there exists an optimal solution to our problem. Simulation results for LTE uplink further demonstrate the significance of our proposed optimization on the overall user experience.
Ali El Essaili, Damien Schroeder, Eckehard G. Steinbach, Wolfgang Kellerer
MMSP4
2011 Influence of Image/Video Compression on Night Vision Based Pedestrian Detection in an Automotive Application
abstract
With the introduction of IP/Ethernet networks in automobiles, video compression becomes mandatory for efficient video transfer. Lossy image and video compression standards, such as JPEG and H.264/AVC, achieve high compression efficiency, but also introduce irreversible modification to the video data. On the one hand, this can lead to perceivable image degradation, and on the other hand, affect the performance of image processing algorithms. The pedestrian detection application based on night vision technology, is one example of a driver assistance service using a video stream as input. In this paper, we examine the impact of image/video compression on pedestrian detection, which employs a far infrared (FIR) sensor. In order to measure application quality, metrics used for driver assistance system validation, are investigated through hardware in the loop (HIL) simulation.
Tankred Hase, Wolfgang Hintermaier, Andreas Frey, Tobias Strobel, Uwe Baumgarten, Eckehard G. Steinbach
VTC Spring6
2010 Error-Resilient Video Transmission for Short-Range Point-to-Point Wireless Communication
abstract
In this paper, we present a novel error resiliency approach for short-range point-to-point wireless video transmission systems that have particularly stringent delay constraints. We propose to utilize per-packet feedback information that can be generated and transmitted instantaneously from the receiver to the transmitter in such systems to help control the transmission errors. A framework for error resilient video transmission under low-latency constraints is presented, which incorporates the instantaneous feedback into the video encoder to stop error propagation, as well as into a channel adaptive retransmission scheme at the transmitter to reduce the residual packet loss rate. The channel adaptive retransmission scheme is integrated into the system without introducing any additional delay, which is achieved by dynamically adjusting the resource allocation between video source coding and retransmission, where different resource allocation strategies are applied for different channel conditions. The performance of the proposed framework is evaluated in a real-time video transmission system, showing significant video quality improvements for a wide range of channel conditions.
Fan Zhang 0026, Eckehard G. Steinbach
ICCCN3
2010 High-fidelity recording, compression, and replay of visual-haptic telepresence sessions
abstract
In this paper, we propose a novel framework for the efficient recording and playback of visual-haptic telepresence sessions. During recording, all involved modalities are synchronously encoded and stored. During replay, interaction with the tele-environment is no longer possible and the displayed force feedback does not relate to the user's interaction anymore. Furthermore, position guidance by applying force-feedback prevents the presentation of the recorded haptic impressions. In our experiments, we have discovered that the human operator can easily associate the previously recorded movements and actions by following the endeffector position of the teleoperator in the displayed visual recording and relate the presented force feedback to its observation. We describe and investigate this observation which allows the user to immerse into the previously recorded telemanipulation session. Conducted psychophysical experiments investigate the performance of our framework in conjunction with perceptual offline coding for haptic signals. Our user studies reveal that our approach can provide high-fidelity visual-haptic experiences for learning, entertainment and/or performance analysis purposes.
Julius Kammerl, Eckehard G. Steinbach
ICIP2
2010 Video synchronization using bit rate profiles
abstract
We present a novel approach for the temporal synchronization of multiple videos which is based on cross-correlating bit rate profiles. The proposed scheme determines the temporal offset without major restrictions on viewing angles, camera properties and camera motion. We propose two extensions of the basic algorithm which reduce the influence of camera motion and distracting background objects. Additionally, we describe how to optimally combine different bit rate components in order to further improve the reliability of the synchronization scheme. The proposed approach, when combined with the three extensions, leads to a reliable, robust, and frame accurate temporal alignment of videos at remarkably low complexity.
Georg Schroth, Florian Schweiger, Michael Eichhorn, Eckehard G. Steinbach, Michael Fahrmair, Wolfgang Kellerer
ICIP4
2010 A novel coordinated adaptive video streaming framework for Scalable Video over mobile networks
abstract
This paper presents a novel adaptive video streaming framework designed for OFDM-based mobile networks. The proposed adaptation algorithm exploits the benefits of the Scalable Video Coding (SVC) extension of the H.264/AVC standard together with cross-layer information from the mobile terminals in performing adaptation that maximizes streaming qualities amongst all participating users. Simulations for adaptive video streaming along with the well-known TFRC-based and no adaptation scenarios were conducted in a realistic mobile network simulator. Our results show significant improvements in all aspects for the introduced framework over both the TFRC-based and no adaptation cases.
Ktawut Tappayuthpijarn, Thomas Stockhammer, Eckehard G. Steinbach
ICIP3
2010 High-fidelity telepresence and teleaction
abstract
The collaborative research center SFB453 (www.sfb453.de) aims to realize high-fidelity telepresence and teleaction systems. Telepresence and teleaction systems extend the human workspace to remote locations in order to overcome barriers like distance, scaling, danger or the human skin. Using a human-system interface the human operator controls a remotely located teleoperator. Multi-modal feedback in form of visual, auditory, and haptic data is used to increase the feeling of telepresence. Different application areas including minimally invasive surgery, on-orbit servicing, microassembly as well as tele-manufacturing and tele-maintenance are targeted.
Robert Bauernschmitt, Martin Buss, Barbara Deml, Klaus Diepold, Berthold Färber, Georg Färber, Ulrich Hagn, Gerd Hirzinger, Sandra Hirche, Alois C. Knoll, Hermann J. Müller, Tobias Ortmaier, Angelika Peer, Michael Popp, Carsten Preusche, Gunther Reinhart, Zhuanghua Shi, Eckehard G. Steinbach, Heinz Ulbrich, Ulrich Walter, Michael F. Zäh
ICRA18
2010 Illumination-invariant image-based novelty detection in a cognitive mobile robot's environment
abstract
Image-based scene representations enable a mobile robot to make a realistic prediction of its environment. Hence, it is able to rapidly detect changes in its surroundings by comparing a virtual image generated from previously acquired reference images and its current observation. This facilitates attentional control to novel events. However, illumination effects can impair attentional control if the robot does not take them into account. To address this issue, we present in this paper an approach for the acquisition of illumination-invariant scene representations. Using multiple spatial image sequences which are captured under varying illumination conditions the robot computes an illumination-invariant image-based environment model. With this representation and statistical models about the illumination behavior, the robot is able to robustly detect texture changes in its environment under different lighting. Experimental results show high-quality images which are free of illumination effects as well as more robust novelty detection compared to state-of-the-art methods.
Werner Maier 0001, Fengqing Bao, Elmar Mair, Eckehard G. Steinbach, Darius Burschka
ICRA4
2010 A SOA-based middleware concept for in-vehicle service discovery and device integration
abstract
We present a novel middleware approach for in-vehicle service discovery and device integration that is based on the concept of Service-oriented Architectures (SOA). In order to be able to identify a suitable SOA system, we define a series of criteria the SOA platform should fulfill. These criteria take into account the specific requirements of the automotive domain. Based on these criteria we compare nine different SOA standards and show why the Device Profile for Web Services (DPWS) standard is most suitable as the middleware of an in-car infotainment and communication system. Furthermore, we present a prototypical implementation of an in-vehicle network based on the DPWS middleware. Hence, we enable the integration of nomadic devices by providing their functionalities within the network and consequently including their graphical output into the in-vehicle Human Machine Interface (HMI). Due to the system design it is also possible to use nomadic devices as an additional display to present the HMI. Finally, an in-vehicle controller is included into our SOA and therefore available to all other devices within our system.
Michael Eichhorn, Martin Pfannenstein, Daniel Muhra, Eckehard G. Steinbach
Intelligent Vehicles Symposium4
2010 A system architecture for IP-camera based driver assistance applications
abstract
We present a novel system architecture for automotive IP-camera based driver assistance applications. We assume that the video data is carried by an in-car switched IP/Ethernet network. The proposed architecture borrows concepts from the CE industry and hence fully takes advantage of the corresponding price and performance advantages. As an implementation study, a top view service is investigated. The suitability of the proposed approach is evaluated by introducing performance measures reflecting the special requirements of the driver assistance domain. To meet the real-time requirements of the driver assistance services we investigate and compare different architectural approaches for IP network cameras. Based on the results obtained with the prototypical implementations, an appropriate solution for future IP-camera based driver assistance services is proposed.
Wolfgang Hintermaier, Eckehard G. Steinbach
Intelligent Vehicles Symposium2
2010 Error-resilient perceptual coding for networked haptic interaction
abstract
The performance of haptic interaction across communication networks critically depends on the successful reconstruction of the bidirectionally transmitted haptic signals, and hence on the quality of the communication channel. We propose a novel error-resilient data reduction scheme for haptic communication which exploits known limits of human haptic perception. Particularly, we show that missing haptic information due to packet loss may strongly impair the user's experience during haptic interaction. We present and compare methods that eliminate the disturbing artifacts resulting out of packet loss. Our approach keeps the estimated impact of packet losses below human perception thresholds. A tree of possible cases (packets received or not received) and their respective occurrence probabilities is maintained at the sender side, and the system predicts unacceptable error cases to decide whether extra packets should be sent. We introduce different criteria that can be employed to trigger additional packets. In our experiments, we evaluate both the objective data reduction performance and the subjective system transparency by performing extensive tests using packet loss probability and round trip time as parameters. The proposed scheme shows excellent performances in terms of data reduction while sustaining good subjective ratings for a wide range of packet loss values and round trip times.
Fernanda Brandi, Julius Kammerl, Eckehard G. Steinbach
ACM Multimedia3
2010 Qoe-based rate adaptation scheme selection for resource-constrained wireless video transmission
abstract
This paper proposes a Quality of Experience (QoE) based rate adaptation scheme selection approach for multi-user wireless video delivery. Transcoding and packet dropping are used as examples of rate adaptation schemes, and we investigate their impact on user perceived video quality. In the presence of constrained computation resources, the most suitable rate adaptation scheme is determined for each video stream such that the overall quality degradation is minimized. The proposed scheme selection approach is integrated with QoE-based resource allocation in presence of constrained transmission resources. Simulation results obtained from an emulated High Speed Downlink Packet Access (HSDPA) show that the QoE-based approach leads to significant improvements of user perceived quality compared to other approaches including a non-optimized HSDPA systems
Srisakul Thakolsri, Wolfgang Kellerer, Eckehard G. Steinbach
ACM Multimedia3
2010 Addressing the uncertainty in critical rate estimation for pixel-domain Wyner-Ziv video coding
abstract
Distributed video coding is typically treated as a channel coding problem among others. The encoder generates parity bits (or syndrome bits) for the source and transmits part of them to the decoder for a certain target quality. The decoder tries to reconstruct the source using the received parity bits along with the side information available. In this paper, we aim to estimate the critical number of parity bits to transmit. Having observed the uncertainty of the critical rate, we model it as a random variable, use its distribution to calculate the decoding failure probability and formulate the expected distortion. We allocate a certain bit budget among different bit-planes such that the expected distortion is minimized. Moreover, we introduce fast decoding at the encoder, which helps us to estimate the critical rate far more accurately. Eventually, we achieve up to 1.5 dB gain in rate-distortion performance at high bit rates.
Abdul Rehman 0001, Eckehard G. Steinbach
VCIP3
2010 Guest editorial: Wireless video transmission
abstract
The 20 papers in this special issue on wireless video transmissions are divided into four categories: wireless video streaming, optimization and scheduling; wireless video broadcast/multicast; retransmission techniques; and video quality assessment.
Jiangzhou Wang, Mingxi Fan, Xiaohu You 0001, Xi Zhang 0005, Eckehard G. Steinbach, Laurence B. Milstein
IEEE J. Sel. Areas Commun.6
2010 Introduction to the best papers of ACM multimedia 2009
abstract
No abstract available.
Changsheng Xu, Eckehard G. Steinbach, Abdulmotaleb El Saddik, Michelle X. Zhou
ACM Trans. Multim. Comput. Commun. Appl.2
2009 Estimation of Location Uncertainty for Scale Invariant Features Points
abstract
Image feature points are the basis for numerous computer vision tasks, such as pose estimation or object detection. State of the art algorithms detect features that are invariant to scale and orientation changes. While feature detectors and descriptors have been widely studied in terms of stability and repeatability, their localisation error has often been assumed to be uniform and insignificant. We argue that this assumption does not hold for scale-invariant feature detectors and demonstrate that the detection of features at different image scales actually has an influence on the localisation accuracy. A general framework to determine the uncertainty of multi-scale image features is introduced. This uncertainty is represented via anisotropic covariances with varying orientation and magnitude. We apply our framework to the well-known SIFT and SURF algorithms, detail its implementation and make it available. Finally the usefulness of such covariance estimates for bundle adjustment and homography computation is illustrated.
Bernhard Zeisl, Pierre Fite Georgel, Florian Schweiger, Eckehard G. Steinbach, Nassir Navab
BMVC4
2009 CAMP: A framework for Cooperation Among Mobile Prosumers
abstract
We propose a novel framework for cooperation among mobile prosumers (CAMP) that aims at establishing semantic relations between multimedia content acquired by mobile users, focussing especially on audio-visual media. We discuss the basic properties of such a system and describe a number of possible applications. We also address technical requirements and possible solutions, and show a preliminary implementation of a cooperative video sharing platform.
Florian Schweiger, Eckehard G. Steinbach, Michael Fahrmair, Wolfgang Kellerer
ICME2
2009 Visual homing and surprise detection for cognitive mobile robots using image-based environment representations
abstract
One important feature of a cognitive system is to perceive and understand its environment and to adapt its actions to changes and unforeseen situations. In this paper, we propose a scheme for visual surprise detection in cognitive mobile robots. With the robot's observation and a set of reference images previously acquired near its current viewpoint, a pixel-wise surprise trigger is computed using Bayesian probabilistic inference techniques. With appropriate mathematical approximations this algorithm can be implemented on modern graphics hardware which nearly allows for real-time surprise detection. In order to refer to prior observations, a mobile robot has to be able to re-localize itself with respect to its environment. Thus, we also present two online image-based homing algorithms which both facilitate the computation of location-independent surprise triggers. Experiments show acceptable results in terms of robust and fast detection of unexpected changes in the environment.
Werner Maier 0001, Elmar Mair, Darius Burschka, Eckehard G. Steinbach
ICRA4
2009 Adaptive video streaming over a mobile network with TCP-friendly rate control
abstract
This paper investigates the performance of TCP-Friendly Rate Control (TFRC) to control the transmission rate of scalable video streams when used in a mobile network. The streams are encoded using the Scalable Video Coding (SVC) extension of the H.264/AVC standard. Adding or removing the layers is decided based on the TFRC during varying channel conditions of the mobile network. We conduct simulations in various realistic use cases, evaluate and compare the performance with and without TFRC-based adaptation. The results show significant improvements in terms of lower loss rate, delay, required buffer size and less playback interruption.
Ktawut Tappayuthpijarn, Günther Liebl, Thomas Stockhammer, Eckehard G. Steinbach
IWCMC4
2009 On the role of multimodal communication in telesurgery systems
abstract
Telesurgery systems integrate multimodal communication and robotic technologies to enable surgical procedures to be performed from remote locations. They allow human surgeons to intuitively control laparoscopic instruments and to navigate within the human body. In this paper, we present selected topics on multimodal interaction in the context of telesurgery applications. These are results from the collaborative research project SFB 453 on ldquoHigh-Fidelity Telepresence and Teleactionrdquo which is funded by the German Research Foundation in the larger Munich area. The focus in this paper is on multimodal information processing and communication including simulation of surgical targets in the human body. Furthermore, we present an overview of our advanced multimodal telesurgery demonstrators that provide a comprehensive platform for our collaborative telepresence research.
Robert Bauernschmitt, Eva U. Braun, Martin Buss, Florian A. Fröhlich, Sandra Hirche, Gerd Hirzinger, Julius Kammerl, Alois C. Knoll, Rainer Konietschke, Bernhard Kübler, Rüdiger Lange, Hermann Georg Mayer, Markus Rank, Gerhard Schillhuber, Christoph Staub, Eckehard G. Steinbach, Andreas Tobergte, Heinz Ulbrich, Iason Vittorias
MMSP16
2009 Models and analysis of streaming video transmission over wireless fading channels
Michel T. Ivrlac, Ruly Lai-U Choi, Eckehard G. Steinbach, Josef A. Nossek
Signal Process. Image Commun.3
2009 Compression of Bayer-Pattern Video Sequences Using Adjusted Chroma Subsampling
abstract
Most consumer digital color cameras are equipped with a single chip. Such cameras capture only one color component per pixel (e.g., Bayer pattern) instead of an RGB triple. Conventionally, missing color components at each pixel are interpolated from its neighboring pixels, so that full color images are constructed. This process is typically referred to as demosaicing. After demosaicing, the full resolution RGB video frames are converted into YUV color space. U and V are then typically subsampled by a factor of four and the resulting video data in the 4:2:0 format become the input for the video encoder. In this letter, we look into the weakness of the conventional scheme and propose a novel solution for compressing Bayer-pattern video data. The novelty of our work lies largely in the chroma subsampling. We properly choose the locations to calculate the chroma pixels U and V according to the positions of B and R pixels in the Bayer pattern and this leads to higher quality of the reconstructed images. In our experiments, we have observed an improvement in composite peak-signal-to-noise ratio performance of up to 1.5 dB at the same encoding rate. Based on this highly efficient approach, we propose also a low-complexity method which saves almost half of the computation at the expense of a small loss in coding efficiency.
Mingzhe Sun, Eckehard G. Steinbach
IEEE Trans. Circuits Syst. Video Technol.3
2009 Proxy-Based Reference Picture Selection for Error Resilient Conversational Video in Mobile Networks
abstract
We propose a frame dependency management strategy for error robust transmission of conversational video in mobile networks. We consider an end-to-end video transmission scenario that involves both a wireless uplink as well as a wireless downlink plus some intermediate wireline network transmission. We also investigate the special cases of an end-to-end scenario where only a wireless uplink or a wireless downlink is present. We cope with packet loss on the downlink by retransmitting lost packets from the base station to the receiver for error recovery. Retransmissions are enabled by using fixed-distance reference picture selection during encoding with a prediction distance that corresponds to the round-trip time of the downlink combined with accelerated decoding. We deal with transmission errors on the uplink by sending acknowledgments and predicting the next frame to encode from those slices that have been correctly received by the base station. We show that these two separate approaches for uplink and downlink efficiently complement one another and the resulting end-to-end scheme is characterized by very low computational complexity. We compare our scheme to several state-of-the-art error resiliency approaches and report significant improvements.
Wei Tu 0001, Eckehard G. Steinbach
IEEE Trans. Circuits Syst. Video Technol.2
2009 Traffic Shaping for Resource-Efficient In-Vehicle Communication
abstract
In-vehicle communication has become complex and costly due to the growing number of automotive network systems applied for different data types. In this work, our previously proposed in-vehicle network architecture that is based on Internet protocol (IP) and full-duplex switched Ethernet (IP/Ethernet) is further investigated for real-time audio and video streaming. Quality-of-service (QoS) and resource usage are analyzed for selected IP/Ethernet-based network topologies. Traffic shaping is used to reduce the required network resources and consequently the cost. A novel traffic shaping algorithm is presented that outperforms other traffic shapers in terms of resource usage when applied to variable bit rate video sources in the proposed double star topology. In addition, a new architecture design is introduced for traffic shaper implementation in switches which operates on a per stream basis. Analytical and simulation results confirm that the proposed network architecture with traffic shaping is well-adapted for in-vehicle communication.
Mehrnoush Rahmani, Ktawut Tappayuthpijarn, Benjamin Krebs, Richard Bogenberger, Eckehard G. Steinbach
IEEE Trans. Ind. Informatics5
2009 Proxy Caching for Video-on-Demand Using Flexible Starting Point Selection
abstract
In this paper, we propose a novel proxy caching scheme for video-on-demand (VoD) services. Our approach is based on the observation that streaming video users searching for some specific content or scene pay most attention to the initial delay, while a small shift of the starting point is acceptable. We present results from subjective VoD tests that relate waiting time and starting point deviation to user satisfaction. Based on this relationship as well as the dynamically changing popularity of video segments, we propose an efficient segment-based caching algorithm, which maximizes the user satisfaction by trading off between the initial delay and the deviation of starting point. Our caching scheme supports interactive video cassette recorder (VCR) functionalities and enables cache replacement with a much finer granularity compared to previously proposed segment-based approaches. Our experimental results show a significantly improved user satisfaction for our scheme compared to conventional caching schemes.
Eckehard G. Steinbach, Muhammad Muhammad
IEEE Trans. Multim.2
2008 A Theoretical Analysis of Data Reduction Using the Weber Quantizer
abstract
We present a theoretical analysis of a perceptual coding approach, the so called Weber quantizer. Extensive studies performed by experimental psychologists and physiologists have unveiled one major conclusion: human perception often follows Weber's law. Ernst Weber was an experimental physiologist who in 1834 first discovered the following implication DeltaI = kl, where DeltaI is the so called difference threshold or the just noticeable difference (JND). It describes the smallest amount of change of an (arbitrary) stimulus I which can be detected just as often as it cannot be detected and defines the Weber bound at [(1 - k)I, (1 + k)I].
Julius Kammerl, Peter Hinterseer, Subhasis Chaudhuri, Eckehard G. Steinbach
DCC4
2008 Progressive rendering from RDTC optimized streams
abstract
Recently, we have introduced a framework for the compression and interactive streaming of image-based scene representations based on hybrid video coding techniques. In that scheme, the storage rate (R), distortion (D), transmission rate (T), and decoding complexity (C) are traded off against each other to minimize the user perceived delay in a remote navigation scenario. In this work, we extend that scheme by introducing a technique for progressive transmission and rendering from such RDTC compressed image-based scene representations. Real-time experiments show that a trade off between the user perceived delay and the rendering quality can be achieved and that the delay can be significantly reduced compared to independent or rate-distortion optimized compression.
Ingo Bauermann, Werner Maier 0001, Eckehard G. Steinbach
ICME3
2008 Flexible distribution of computational complexity between the encoder and the decoder in distributed video coding
abstract
Based on our previous work on distributed video coding with low complexity motion estimation at the encoder, we propose in this paper a more flexible system, providing the possibility that the encoder and the decoder cooperate in motion estimation and share the computational bulk of it. In this system, not only the encoder but also the decoder is capable of improving the coding efficiency by spending more computational power in motion estimation. Experimental results show that the cooperation of the encoder and the decoder can reduce the total computational complexity and at the same time improve the coding efficiency. In addition, we propose a simple scheme to order the intensity of motion activity for different macroblocks. When the encoder has limited computational power, we apply motion estimation only to the macroblocks with relatively large motion. This leads to an efficient use of computational power at the encoder.
Eckehard G. Steinbach
ICME2
2008 Dynamic segment based proxy caching for Video on Demand
abstract
In this paper, we propose a novel proxy caching scheme for Video on Demand (VoD) services. Our approach is based on an observation we have made during subjective VoD performance evaluation tests. We have found that when users are seeking for some specific content, they pay most attention to the initial delay, while a small shift of the starting point is acceptable. Meanwhile, the random access using the slide bar sometimes even makes this small shift unnoticeable. Based on this fact as well as the popularity of video segments, we propose a dynamic segment based caching algorithm. It adjusts the segment size according to its hitting rate and maximizes the user satisfaction by optimally trading off between the initial delay and the deviation of the starting point. Our experimental results show a significantly improved user satisfaction compared to previous caching approaches.
Eckehard G. Steinbach
ICME3
2008 Adaptive multi-source video multicast
abstract
We introduce an adaptive rate control scheme for multi-source video streaming in peer-to-peer (P2P) networks. The rate control scheme is applied in the context of a novel multi-source video multicast framework, where each source distributes its own video sequence while additionally forwarding video data received from other sources. Multiple videos are distributed to all requesting peers exploiting full collaboration between the sources, the requesting peers and helper peers. The rate adaptation is jointly decided for all sources based on their upload capacities and the required video quality. Our goal is to obtain either dasiasame throughputpsila or dasiasame video qualitypsila for all streams. To achieve the same video quality for all streams, we use scalable video coding techniques. Our evaluation made on PlanetLab shows that our solution achieves a better video quality or throughput adaptation compared to independent rate allocation performed at each source.
Francisco de Asís López-Fuentes, Eckehard G. Steinbach
ICME2
2008 Evaluation of segment-based proxy caching for video on demand
abstract
In this paper, we derive an analytical model for the evaluation of the performance of a Video on Demand (VoD) system. The model estimates the mean waiting time achievable by the Popularity-Aware Partial cAching (PAPA) algorithm from our previous work. Two approximation strategies are proposed for low computational complexity. Furthermore, we also consider the influence of a starting point shift on the quality of experience and combine the two factors into a universal user satisfaction metric. In order to find the relation between the two impairments, waiting time and starting point shift, sophisticated subjective tests are performed. With the final score model, a more comprehensive evaluation of the system can be obtained with very low computational complexity.
Muhammad Muhammad, Eckehard G. Steinbach
ICME3
2008 A comparative study of network transport protocols for in-vehicle media streaming
abstract
We analyze and compare various transport protocols in the context of wireless in-vehicle IP-based audio and video communication. We determine the most appropriate transport protocol and discuss its benefits for an application in the car. The analyses are accomplished based on the IEEE 802.11 standard. A testbed is used to measure and compare quality of service values such as throughput, jitter and media quality at the receiver. In the experiments, the traditional protocols TCP and UDP showed the best performance.
Mehrnoush Rahmani, Andrea Pettiti, Ernst W. Biersack, Eckehard G. Steinbach, Joachim Hillebrand
ICME4
2008 Joint calibration of a camera triplet and a laser rangefinder
abstract
This paper studies the calibration of a multi-sensor capture device for the acquisition of image-based scene representations. It is composed of cameras and a laser rangefinder. We use a plane based approach to calibrate the compound system, determining both intrinsic and extrinsic parameters. The novelty of our approach is that the cameras and the scanner are treated jointly. We have tested our method both in simulations and on real data and observe very good calibration results.
Florian Schweiger, Ingo Bauermann, Eckehard G. Steinbach
ICME3
2008 Multi-source video multicast in peer-to-peer networks
abstract
We propose a novel framework for multi-source video streaming in peer-to-peer (P2P) networks. Multiple videos are distributed to all requesting peers exploiting full collaboration between the sources, the requesting peers and helper peers. Each source distributes its own video sequence while additionally forwarding video data received from other sources. A single peer is selected to redistribute a particular video block to the peers which would like to receive the videos. Our goal is to maximize the overall throughput or alternatively the aggregate video quality of multiple concurrent streaming sessions. We also consider the special cases of “same throughput” or “same video quality” for all streams. We formulate the rate allocation and redistribution as an optimization problem and evaluate our framework for three different scenarios. In the first scenario, the rate allocation is jointly decided for all participating peers. In the second scenario, the rate allocation is also decided jointly, but additionally either same rate or same video quality streams are enforced. Our third scenario assumes separate distribution for every source. In this case, the peers divide their upload capacity equally among the different video sequences. Our results show the superior performance of joint rate allocation compared to independent allocation and the effectiveness of our framework.
Francisco de Asís López-Fuentes, Eckehard G. Steinbach
IPDPS2
2008 Multi-modal multi-user telepresence and teleaction system
abstract
The video shows a rich multi-modal multi-user telepresence system, which was developed within the SFB453 funded by the German Research Foundation (www.sfb453.de). As a complex application scenario, the remote repairing of a broken pipe is presented in this paper. The system basically consists of two operator-teleoperator- pairs. While one of the operators interacts with a stationary human- system-interface, the other operator uses a mobile one. Both systems provide visual, auditory, and haptic feedback and enable to control the motion of head, arms, and grippers as well as the locomotion of the corresponding teleoperator.
Martin Buss, Angelika Peer, Thomas Schauss, Nikolay Stefanov, Ulrich Unterhinninghofen, Stephan Behrendt, Georg Färber, Jan Leupold, Klaus Diepold, Fakheredine Keyrouz, Michel Sarkis, Peter Hinterseer, Eckehard G. Steinbach, Berthold Färber, Helena Pongrac
IROS13
2008 Deadband-based offline-coding of haptic media
abstract
In this work, a novel perceptual coding approach for offline compression of haptic media is presented. Our scheme exploits the properties of human haptic perception and hides coding artifacts introduced by lossy compression below the human perception thresholds. We combine the concept of Just Noticeable Differences with predictive coding in order to achieve high coding efficiency without impairing user perception. In our experiments, we apply the proposed lossy compression scheme to haptic data that has been recorded during a telemanipulation session in a virtual environment. Our user studies reveal that our approach leads to strong data reduction up to 100:1 while preserving a high-quality haptic experience during playback of the compressed haptic data streams.
Julius Kammerl, Eckehard G. Steinbach
ACM Multimedia2
2008 RDTC Optimized Compression of Image-Based Scene Representations (Part I): Modeling and Theoretical Analysis
abstract
Rendering of virtual views in interactive streaming of compressed image-based scene representations requires random access to arbitrary parts of the reference image data. The degree of interframe dependencies exploited during encoding has an impact on the transmission and decoding time and, at the same time, delimits the (storage) rate-distortion (RD) tradeoff that can be achieved. In this work, we extend the classical RD optimization approach using hybrid video coding concepts to a tradeoff between the storage rate (R), distortion (D), transmission data rate (T), and decoding complexity (C). We present a theoretical model for this RDTC space with a focus on the decoding complexity and, in addition, the impact of client side caching on the RDTC measures is considered and evaluated. Experimental results qualitatively match those predicted by our theoretical models and show that an adaptation of the encoding process to scenario specific parameters like computational power of the receiver and channel throughput can significantly reduce the user-perceived delay or required storage for RDTC optimized streams compared to RD optimized or independently encoded scene representations.
Ingo Bauermann, Eckehard G. Steinbach
IEEE Trans. Image Process.2
2008 RDTC Optimized Compression of Image-Based Scene Representations (Part II): Practical Coding
abstract
Interactive streaming of compressed image-based scene representations requires random access to the reference image data. The degree of interframe dependencies exploited during encoding has an impact on the transmission and decoding time and, at the same time, delimits the (storage) rate-distortion (RD) tradeoff that can be achieved. The transmission data rate and the decoding complexity at the client have received attention in the literature, but their incorporation into the optimization procedure for compression and streaming is missing. If scenario-specific measures are considered, the traditional RD optimization can be extended to a tradeoff between the (storage) rate (R), distortion (D), transmission data rate (T), and decoding complexity (C). In the first part of this sequel of papers, we have theoretically analyzed the RDTC space for the compression of densely sampled image-based scene representations. In this second part, we consider practical RDTC optimization. We propose a modeling and encoding parameter selection procedure that allows us to adapt the compression to scenario-specific properties. The impact of client side caching is considered and evaluated using an experimental testbed. Our results show a significant reduction of the user perceived delay, memory consumption or required minimum channel and storage bitrate for RDTC optimized streams compared to classical RD optimized or independently encoded scene representations.
Ingo Bauermann, Eckehard G. Steinbach
IEEE Trans. Image Process.2
2008 Multimedia Applications in Mobile/Wireless Context
abstract
The five articles in this special section focus on multimedia applications in mobile/wireless context. Two papers focus on video transmission over 802.11a/b and WiMax, the next discusses a visual sensor network in wireless environments, and the last two propose supporting technologies for two mobile applications that rely on multimedia interactions.
Bo Shen 0003, Wei Tsang Ooi, Giacomo Morabito, Eckehard G. Steinbach
IEEE Trans. Multim.4
2007 Cross-Layer Optimization With Model-Based Parameter Exchange
abstract
Cross-layer optimization (CLO) promises significant gains in comparison to a conventional system design, which does not allow for information exchange across layers. One of the key challenges in CLO is the exchange of parameters between optimizer and layers. In this paper a model-based approach is presented that drastically reduces the amount of parameters that are to be exchanged. The optimizer employs models of the respective layers that emulate the communication system within the optimizer. The layers then only need to pass a small number of model parameters to the optimizer. This general concept is applied to CLO between application (APP) layer and medium access control (MAC) layer of a radio communications system. Our proposed model for the MAC layer is suitable for a transmitter without instantaneous channel state information (CSI). Simulation results demonstrate that the proposed model-based CLO is able to exploit the available diversity to enhance the system capacity. Dependent on the application characteristics, the same perceived quality in terms of mean opinion score (MOS) is maintained, while increasing the number of served users by up to 25%. Compared to known CLO approaches, much fewer parameters need to be exchanged.
Andreas Saul, Shoaib Khan, Gunther Auer, Wolfgang Kellerer, Eckehard G. Steinbach
ICC5
2007 A Novel Signal Reconstruction Algorithm for Perception Based Data Reduction in Haptic Signal Communication
abstract
The performance and immersiveness of telepresence and teleaction systems critically depend on the quality of the communication between the operator and the teleoperator. High packet rate for low delay data exchange in internet-based tele-operation is one of the major challenges in this context. A psychophysical^ motivated data reduction scheme, earlier presented by the authors, solves this problem by the use of signal amplitude dependent deadband quantization. This scheme successfully limits the amount of transmitted packets, but as soon as the size of the deadband exceeds a certain limit, noticeable and disturbing artifacts are introduced into the haptic signals. In this paper, we present a novel signal reconstruction technique, based on a signal adaptive synthesis filter. It removes discontinuities in the displayed signals and enables the use of strong deadband based data reduction without perceptible disturbance in the telepresence system.
Julius Kammerl, Peter Hinterseer, Eckehard G. Steinbach
ICCCN3
2007 Analysis of the Decoding-Complexity of Compressed Image-Based Scene Representations
abstract
Interactive navigation in image-based scenes requires random access to the compressed reference image data. When using state of the art block-based hybrid video coding techniques, the degree of inter and intra block dependencies introduced during compression has an impact on the effort required to access reference image data and therefore delimits the response time for interactive applications. In this work a theoretical model for the decoding-complexity of compressed image-based scene representations is presented and evaluated. Results show the validity of the model. Additionally, results for decoding-complexity constrained rate distortion optimization (RDC) using our model show the benefit of incorporating the computational power of client devices into the compression process.
Ingo Bauermann, Eckehard G. Steinbach
ICIP (6)2
2007 Encoding Parameter Estimation for RDTC Optimized Compression and Streaming of Image-Based Scene Representations
abstract
Remote navigation in image-based scene representations requires random access to parts of the compressed reference image data to compose virtual views. The degree of dependencies introduced during compression has an impact on the effort that is required to access reference image data and at the same time delimits the rate-distortion (RD) tradeoff that can be achieved. If a limited channel bitrate and computational power of client devices are taken into account, encoding can be performed in a RD optimal manner with respect to the expected maximum transmission data rate (T) and decoding complexity (C). In this work we present a practical framework for parameter estimation for RDTC optimal encoding of image-based scene representations.
Ingo Bauermann, Eckehard G. Steinbach
ICIP (6)2
2007 Multiple Description Video Transcoding
abstract
In this paper we introduce the concept of multiple description video transcoding (MDVT). MDVT converts a single description encoded video into two or more descriptions at an intermediate node in the network. The objective of our MDVT approach is to adapt the video transmission to a multi-radio environment where two or more independent transmission paths exist between the intermediate node and the receiver. The sender does not have to be aware of the transcoding process and the multi-path transmission. MDVT can for instance be applied for multi-mode terminals that are simultaneously connected to two wireless access technologies, e.g., UMTS and WLAN. We compare MDVT with multiple description coding at the sender (MDC-S) as well as with MDC at the intermediate node (MDC-I) where the incoming single description video is decoded and re-encoded into multiple descriptions and the transmission is optimized separately for each path. We present a fast greedy method that can be used to perform multiple description video transcoding in real-time at low complexity. Our experimental results show that we can achieve performance similar to MDC-S where the sender has to be aware of the availability of multiple paths. Compared to MDC-I we observe more than 2 dB gain in reconstruction quality.
Ali El Essaili, Shoaib Khan, Wolfgang Kellerer, Eckehard G. Steinbach
ICIP (6)4
2007 Wyner-Ziv Video Coding Based on Turbo Codes Exploiting Perfect Knowledge of Parity Bits
abstract
In a Wyner-Ziv video coding system based on turbo codes, parity bits are generated for a Wyner-Ziv frame and transmitted to the decoder. In this context of distributed source coding, we can assume that the transmission of the parity bits is guaranteed to be error-free by the lower network layers. In our work, we propose a simple method to exploit the guaranteed reliability of the parity bits so that the turbo decoding achieves a better performance. The proposed scheme is to assign the parity bits a higher weight than the information bits in turbo decoding. Our simulation on binary sequences shows that the proposed scheme enables turbo codes to come closer to the Slepian-Wolf bound in distributed lossless source coding. When it is applied to Wyner-Ziv video coding, the bit rate is reduced by 2% -4% for a given video quality. Although the gain is relatively small, it comes for free, because all that we have to do is to multiply the parity bits with a sufficiently large factor before the decoding process starts.
Eckehard G. Steinbach
ICME2
2007 Fast Motion Estimation-Based Reference Frame Generation in Wyner-Ziv Residual Video Coding
abstract
In practical Wyner-Ziv video coding, every frame is encoded independently of others, but decoded based on the side information generated from adjacent frames. In Wyner-Ziv residual coding of video, the residual of a frame with respect to a reference frame is Wyner-Ziv encoded, which leads to a higher coding efficiency than directly Wyner-Ziv encoding the original frame. In previous work, the reference frame is directly copied from the previously reconstructed frame. In this paper, we generate the reference Wyner-Ziv frame at the encoder using low complexity fast motion search. Experimental results show that the proposed scheme provides a significant gain in video coding efficiency with the encoding complexity being only slightly increased. Using our approach, Wyner-Ziv residual video encoding becomes flexible and allows us to trade-off the encoding complexity and the overall rate-distortion performance. Moreover, the reference frame can help us in refining the side information generated at the decoder, which contributes to the improvement of the overall video coding efficiency.
Eckehard G. Steinbach
ICME2
2007 Joint Network and Rate Allocation for Simultaneous Wireless Applications
abstract
We address the problem of rate allocation and network/path selection for multiple users, running simultaneous applications over multiple parallel access networks. Our joint optimization problem consists of finding the appropriate application rate allocation and network parameters for each individual user, such that an overall quality metric is maximized. We compare our solution to other solutions based on throughput optimization strategies through extensive simulations, and we show the superiority of our approach. Furthermore, our solution proves to be more robust in dynamic systems, when clients can join/leave the access networks.
Dan Jurca, Wolfgang Kellerer, Eckehard G. Steinbach, Shoaib Khan, Srisakul Thakolsri, Pascal Frossard
ICME3
2007 RD-Optimized Rate Shaping for Multiple Scalable Video Streams
abstract
Delivering streaming video over wired and wireless networks poses many challenges, primarily due to the throughput variations caused by time-varying network conditions. Scalable video coding gives an elegant way to adapt the video streams to the available transmission resource. In this paper, we consider a multi-user scenario wherein multiple scalable video streams compete for a shared transmission link with limited forwarding capacity. We propose a rate-distortion (RD) optimized rate shaping approach to improve the overall video quality. For this, compact RD side information is sent along with the sequences. Our simulation results show that significant improvements are achieved by the proposed RD-optimized rate shaping approach compared to conventional priority-based rate shaping.
Rajyalakshmi Mahalingam, Wei Tu 0001, Eckehard G. Steinbach
ICME3
2007 A Flexible Starting Point Based Partial Caching Algorithm for Video on Demand
abstract
In this paper, we propose a novel proxy caching scheme for Video on Demand (VoD) services. Our approach is based on an observation we have made during subjective VoD performance evaluation tests. We have found that when users are seeking for some specific content, they pay most attention to the initial delay, while a small shift of the starting point is acceptable. Based on this observation as well as the dynamic popularity of video segments, we propose an efficient segment based caching algorithm. The approach is applied on a proxy and minimizes the average initial playout delay at the clients. Our experimental results show a significant reduction of the average initial waiting time compared to conventional caching approaches.
Lian Shen, Eckehard G. Steinbach
ICME3
2007 Proximity-Aware Collaborative Multicast for Small P2P Communities
abstract
In this paper we describe a novel solution for delay sensitive one-to-many content distribution in P2P networks based on cooperative m-ary trees. Our scheme maximizes the overall throughput while minimizing end-to-end delay by exploiting the full upload capacities of the participating peers and their proximity relationship. Our delivery scheme is based on cooperation between the source, the content-requesting peers and the helper peers. In our solution, the source splits the content into several blocks and feeds them into multiple m-ary trees rooted at the source. Every peer contributes its upload capacity by being a forwarding peer in at least one of the m-ary trees. Our performance evaluation shows that our proposal achieves similar throughput as the best known solution in the literature (mutualcast) while at the same time reducing content delivery delay.
Francisco de Asís López-Fuentes, Eckehard G. Steinbach
IPDPS2
2007 Joint Network and Rate Allocation for Video Streaming over Multiple Wireless Networks
abstract
Abstract — We address the problem of video streaming over multiple parallel networks. In the context of multiple users, accessing different types of applications, we are looking for efficient ways of allocating network resources and selecting network paths for each application, in order to maximize the overall systems performance. Our optimization joint problem consists of finding the appropriate application rate allocation and network parameters for each individual user, such that a universal system quality metric is maximized. A specific mapping between the requirements of each considered application and the overall quality metric is introduced, and our results are compared to other solutions based on throughput optimization strategies. The superiority and robustness of our approach is shown through extensive simulations in constant and dynamic systems, when clients can join/leave the access networks. Furthermore, we introduce heuristic algorithms which can obtain good results and are inexpensive in terms of computation and execution time. I.
Dan Jurca, Wolfgang Kellerer, Eckehard G. Steinbach, Shoaib Khan, Srisakul Thakolsri, Pascal Frossard
ISM3
2007 Hierarchical collaborative multicast
abstract
In this paper, we propose and evaluate a novel solution for delay sensitive one-to-many content distribution in P2P networks based on hierarchical clustering. Our delivery scheme involves cooperation among the participating peers. The source splits the content into variable size blocks and sends the blocks to a subset of peers. All peers redistribute parts of the content to other peers. Our approach adopts a tree structure as the global structure to achieve scalability, while fully connected small clusters based on proximity are built in each hierarchical level of the tree, taking advantage of the higher transmission capacities among neighboring peers. Every peer contributes its upload capacity to the system by being a forwarding peer within a cluster. Our evaluation made on PlanetLab shows that our proposal achieves a good performance, short delivery time and scalability.
Francisco de Asís López-Fuentes, Eckehard G. Steinbach
ACM Multimedia2
2007 Multiuser MIMO: Principle, Performance in Measured Channels and Applicable Service
abstract
The exploitation of multiuser diversity and the application of multiple antennas at transmitter and receiver are considered to be key technologies for future highly bandwidth-efficient wireless systems. We combine both ideas in a downlink multicarrier transmission scheme where multiple users compete for the available resources in time, frequency and space. The instantaneous channel impulse responses for all users are assumed to be perfectly known at the transmitter. Our proposed algorithm allocates each spatial dimension on a subcarrier to the user which has the highest channel tap gain on the respective spatial dimension. The scheduling strategy is optimized for sum capacity maximization. In this paper, we restrict ourselves to a more illustrative description of the idea rather then providing mathematical details. We demonstrate the potential of the proposed scheme by capacity results for measured real world channels in a large office environment. Finally, video streaming is used as a potential application with high data rate and low latency demands. It is shown that the proposed method has the potential to exploit multiuser diversity while still providing stable video streams even though QoS constraints are not explicitly taken into account by the scheduler.
Gerhard Bauch 0001, Pedro Tejera, Christian Guthy, Wolfgang Utschick, Josef A. Nossek, Markus Herdin, Jorgen Nielsen, Jørgen Bach Andersen, Eckehard G. Steinbach, Shoaib Khan
VTC Spring9
2006 Perception-Based Compression of Haptic Data Streams Using Kalman Filters
abstract
In order to realize truly immersive and stable telepresence and teleaction over the Internet it is necessary to keep delay, data rate, and packet rate of haptic data streams as low as possible. In addition, the compression scheme for haptic data, which is necessary to achieve those goals, must be fast enough to work on a sample by sample basis to not add further delay. This paper presents an approach that reduces haptic data traffic in networked telepresence and teleaction systems to a small fraction of the original rate without impairing performance by using fast Kalman filters on the input signals combined with model based prediction of haptic signals. Our approach reduces the number of transmitted packets to 9.8% (velocity) and 6.2% (force) of the original rate without impairing immersiveness
Peter Hinterseer, Eckehard G. Steinbach, Subhasis Chaudhuri
ICASSP (5)2
2006 Comparison of Physical-Layer and Application-Layer Diversity for Video Streaming Over Wireless Rayleigh Fading Channels
abstract
In this paper, we compare the performance of the diversity approach at the application layer and that at the physical layer for H.264/AVC encoded streaming video transmission in a wireless system with two independent channels. The application layer diversity is achieved by using multiple description source coding, while the physical layer diversity is obtained by channel coding over the two channels. Numerical results show that PHY layer diversity always outperforms the application layer diversity.
Ruly Lai-U Choi, Michel T. Ivrlac, Eckehard G. Steinbach, Josef A. Nossek
ICIP3
2006 Telepresence Across Networks: A Combined Deadband and Prediction Approach
abstract
Executing telepresence operations across networks poses a number of challenges. In particular, the transfer of certain data types requires high sampling rates, such as haptic and velocity data. This paper presents a two-step approach which reduces the total amount of data sent across a network, while attempting to reconstruct the original data series. The first step utilises a deadband approach to eliminate the transfer of redundant data. In this sense, redundant refers to the fact that the data will be undetectable or of no consequence at the receiving end. The second step is a prediction implementation to best reconstruct the reduced series to its original form. This combined deadband and prediction approach is applied to velocity data extracted from a telepresence scenario, resulting in an 86% reduction in the amount of data transferred, and good data reconstruction properties
Michael F. Zäh, Stella M. Clarke, Peter Hinterseer, Eckehard G. Steinbach
IV4
2006 Application-driven cross-layer optimization for mobile multimedia communication using a common application layer quality metric
abstract
This paper proposes a cross-layer optimization framework that provides efficient allocation of wireless network resources across multiple types of applications to maximize network capacity and user satisfaction. We define a novel optimization scheme based on the Mean Opinion Score (MOS) as the unifying metric. Our experiments, applied to scenarios where users simultaneously run three types of applications, such as realtime voice, video conferencing and file download, confirm that MOS-based optimization leads to significant improvement in terms of user perceived quality when compared to throughput-based optimization.
Shoaib Khan, Svetoslav Duhovnikov, Eckehard G. Steinbach, Marco Sgroi, Wolfgang Kellerer
IWCMC3
2005 A novel, psychophysically motivated transmission approach for haptic data streams in telepresence and teleaction systems
abstract
One of the key challenges in telepresence and teleaction systems is the fact that a global control loop is closed over a communication network. The transmission delay of haptic information is extremely critical. Therefore, new data samples from the haptic sensors are typically immediately forwarded to the receiver which leads to a large number of packets being generated when using the Internet as the communication infrastructure. We present a novel approach to reduce the number of packets and, therefore, the amount data communicated in a telepresence and teleaction system. Our method uses a passive deadband transmission approach which only delivers data packets over the network when the sampled sensor data changes more than a given threshold value. The threshold value is determined by psychophysical experiments. This approach leads to a considerable reduction (up to 90%) of packet rate and data rate without sacrificing the fidelity and immersiveness of the system.
Peter Hinterseer, Eckehard G. Steinbach, Sandra Hirche, Martin Buss
ICASSP (2)2
2005 Multiple Description Video Coding Using Motion-Compensated Lifted 3D Wavelet Decomposition
abstract
We design a multiple description (MD) video coding scheme based on the motion compensated (MC) lifted wavelet transform. A major advantage of basing MD video coding on motion-compensated lifted 3D wavelet decomposition is that it does not require any mismatch control as in a hybrid codec which was previously achieved by sending drift compensation data. We propose to perform the temporal decomposition of a group of pictures and then create multiple descriptions for each temporally transformed frame. We use polyphase subsampling with an appropriate amount of oversampling to form the descriptions and to control the redundancy among them. We determine the appropriate amount of redundancy as a function of description loss probability. Our experimental results show that the controlled introduction of redundancy leads to significant improvements in lossy transmission environments.
Aditya Mavlankar, Eckehard G. Steinbach
ICASSP (2)2
2005 Analysis of distortion due to packet loss in streaming video transmission over wireless communication links
abstract
In this paper, we provide an accurate and fully analytical model for the distortion due to lost frames in wireless video transmission. Our analysis combines the properties of the video sequence and those of the wireless transmission link. Built on the long-term average properties of the video source, we first develop a model for the average distortion due to the loss of frames, and then derive a mathematical model for analysis of transmission errors occurring during wireless transmission. Experimental results show a surprisingly accurate behaviour of the proposed model.
Ruly Lai-U Choi, Michel T. Ivrlac, Eckehard G. Steinbach, Josef A. Nossek
ICIP (1)3
2005 Sequence-level models for distortion-rate behaviour of compressed video
abstract
In this paper, two empirical models for the sequence-level distortion-rate performance of predictive video source encoding are proposed. They require very limited amount of empirical data, namely three pairs of rate and distortion, in order to set up the model parameters. The advantages of these proposed models are the robustness towards measurement noise, low number of model parameters with a closed form setup solution, and high accuracy. Experimental validation using H.264/AVC encoded video test sequences reports high accuracy of these two proposed models.
Ruly Lai-U Choi, Michel T. Ivrlac, Eckehard G. Steinbach, Josef A. Nossek
ICIP (2)3
2005 Adaptive resource allocation and frame scheduling for wireless multi-user video streaming
abstract
We propose an application-driven multi-user resource allocation and frame scheduling concept for wireless video streaming. Our approach is based on joint optimization of the application layer, the data link layer and the physical layer. For this, key parameters from these three layers are abstracted. The abstracted parameters at the application layer describe the rate-distortion characteristics of the pre-encoded video streams. At the lower layers they describe the current transmission characteristics of all users. The outcome of the joint optimization leads to adaptive resource allocation at the lower layers and an adaptive decision on which frames to send on the application layer. We show that for our scenario the expected video quality at the client side can be described analytically which leads to low complexity joint optimization. The performance of our approach is demonstrated using a real-time testbed implementation.
Shoaib Khan, Eckehard G. Steinbach, Marco Sgroi, Wolfgang Kellerer
ICIP (3)3
2005 Cross-layer optimization for wireless video streaming-performance and cost
abstract
Cross-layer design (CLD) is a new paradigm for network architecture that allows us to make better use of network resources by optimizing across the boundaries of traditional network layers. Previous work has shown that applying CLD to mobile multimedia communication systems may lead to significant performance improvements. In this paper we also consider the other side of the coin, i.e., the additional computation and communication overhead introduced by CLD. We evaluate the performance improvements and the cost of cross-layer optimization using a wireless multi-user video streaming example.
Shoaib Khan, Marco Sgroi, Eckehard G. Steinbach, Wolfgang Kellerer
ICME3
2005 Proxy-based reference picture selection for real-time video transmission over mobile networks
abstract
We propose a framework for error robust real-time video transmission over wireless networks. In our approach, we cope with packet loss on the downlink by retransmitting lost packets from the base station (BS) to the receiver for error recovery. Retransmissions are enabled by using fixed-distance reference picture selection during encoding with a prediction distance that corresponds to the round-trip-time between the BS and the receiver. We deal with transmission errors on the uplink by sending acknowledgements and predicting the next frame from the most recent frame that has been positively acknowledged by the BS. We show that these two separate approaches for uplink and downlink nicely fit together. We compare our approach to state-of-the art error resilience approaches that employ random intra update of macroblocks and FEC across packets for error resilience. At the same bit rate and packet loss rate we observe improvements of up to 4.5 dB for our scheme.
Eckehard G. Steinbach
ICME2
2005 Affine multipicture motion-compensated prediction
abstract
Affine motion compensation is combined with long-term memory motion-compensated prediction. The idea is to determine several affine motion parameter sets on subareas of the image. Then, for each affine motion parameter set, a complete reference picture is warped and inserted into the multipicture buffer. Given the multipicture buffer of decoded pictures and affine warped versions thereof, block-based translational motion-compensated prediction and Lagrangian coder control are utilized. The affine motion parameters are transmitted as side information requiring additional bit rate. Hence, the utility of each reference picture and, with that, each affine motion parameter set is tested for its rate-distortion efficiency. The combination of affine and long-term memory motion-compensated prediction provides a highly efficient video compression scheme in terms of rate-distortion performance. The two incorporated multipicture concepts complement each other well providing almost additive rate-distortion gains. When warping the prior decoded picture, average bit-rate savings of 15% against TMN-10, the test model of ITU-T Recommendation H.263, are reported for the case that 20 warped reference pictures are used. When employing 20 warped reference pictures and 10 decoded reference pictures, average bit-rate savings of 24% can be obtained for a set of eight test sequences. These bit-rate savings correspond to gains in PSNR between 0.8-3 dB. For some cases, the combination of affine and long-term memory motion-compensated prediction provides more than additive gains.
Thomas Wiegand 0001, Eckehard G. Steinbach, Bernd Girod
IEEE Trans. Circuits Syst. Video Technol.2
2004 Cross layer optimization for wireless multi-user video streaming
Ruly Lai-U Choi, Wolfgang Kellerer, Eckehard G. Steinbach
ICIP3
2004 A performance comparison of multiple description video streaming in peer-to-peer and content delivery networks
abstract
We examine the performance of peer-to-peer (P2P) media streaming using multiple description coding (MDC) and compare it with that of a content delivery network (CDN). In both approaches, multiple servers simultaneously serve one requesting client with complementary descriptions. This approach improves reliability and decreases the data rate a server has to provide. We have implemented both approaches in the ns-2 network simulator. The experimental results indicate that the user perceived video quality of MDC-based streaming in a P2P network can be significantly better than in a CDN, despite the high degree of unreliability of the P2P network.
Shoaib Khan, Rüdiger Schollmeier, Eckehard G. Steinbach
ICME3
2004 Proxy-based error tracking for H.264 based real-time video transmission in mobile environments
abstract
Error tracking (ET) is an error resilience technique for real-time video transmission over error-prone communication channels. In this paper we propose proxy-based ET for communication scenarios where the sender is in the wired Internet and the receiver is connected via a wireless link. Our main assumptions in this work are that there is a strong imbalance between the transmission rates available in the wired and the wireless Internet and that the round-trip delay is mainly caused by the wired part of the connection. We also assume that the majority of the packet loss is caused by the wireless link. In order to allow the proxy server to perform error tracking, an additional update stream is sent through the wired network. This additional information is used by the proxy to improve the performance on the wireless link We show that under these assumptions proxy-based error tracking leads to significantly improved performance for H.264 based real-time video communication in comparison to traditional end-to-end error tracking
Eckehard G. Steinbach
ICME2
2004 Disposal of Explosive Ordnances by Use of a Bimanual Haptic Telepresence System
abstract
The paper outlines a novel approach for performing explosive ordnance disposal by use of a bimanual haptic telepresence system. This system enables an operator to perceive multimodal feedback from a remote environment for proper task execution. The developed experimental setup, comprising a bimanual human system interface and the corresponding bimanual teleoperator for use of both hands, is presented in detail. The teleoperation control architecture is discussed as well as a local model-based impedance control algorithm for manipulator control. Human-system performance is improved by means of stereo visualization of the tele-environment together with overlayed Augmented Reality assistance, and an algorithm for avoidance of dangerous manipulator configurations supported by augmented force feedback. Furthermore, a recently developed UDP communication library is presented for system interconnection taking into account compression of haptic data. Thus efficient and low delay of data transfer is ensured. The usability and effectiveness of the developed bimanual telepresence system are demonstrated by focusing a relevant task scenario, such as demining operations in a remote environment.
Alexander Kron, Günther Schmidt 0001, Bernd Petzold, Michael F. Zäh, Peter Hinterseer, Eckehard G. Steinbach
ICRA6
2004 Adaptive media playout for low-delay video streaming over error-prone channels
abstract
When media is streamed over best-effort networks, media data is buffered at the client to protect against playout interruptions due to packet losses and random delays. While the likelihood of an interruption decreases as more data is buffered, the latency that is introduced increases. In this paper we show how adaptive media playout (AMP), the variation of the playout speed of media frames depending on channel conditions, allows the client to buffer less data, thus introducing less delay, for a given buffer underflow probability. We proceed by defining models for the streaming media system and the random, lossy, packet delivery channel. Our streaming system model buffers media at the client, and combats packet losses with deadline-constrained automatic repeat request (ARQ). For the channel, we define a two-state Markov model that features state-dependent packet loss probability. Using the models, we develop a Markov chain analysis to examine the tradeoff between buffer underflow probability and latency for AMP-augmented video streaming. The results of the analysis, verified with simulation experiments, indicate that AMP can greatly improve the tradeoff, allowing reduced latencies for a given buffer underflow probability.
Mark Kalman, Eckehard G. Steinbach, Bernd Girod
IEEE Trans. Circuits Syst. Video Technol.2
2003 A scalable virtual programmable real-time testbed for rapid multimedia service creation and evaluation
abstract
In this paper we describe a new flexible and universally applicable testbed for Internet-based multimedia applications. Our approach combines the technology of programmable networks with emulated network environments and delivers a virtual programmable testbed (VPT). In this work we show that a VPT is well suited for rapid creation and evaluation of distributed multimedia applications. There are three advantages of our approach compared to previously proposed testbed solutions: first, the software for processing IP datagrams, so called atomic modules, only depends on the programmable platform deployed, making it reusable. Second, loading several atomic modules concurrently allows the easy setup of customized test environments, e.g., channel models. Third, the application being developed can be tested under various network conditions (failures, congestion, etc.) at an early stage. Based on our VPT, distributed multimedia services can be created and evaluated both fast and economically in emulated environments without expensive and complex hardware setup.
Christian Bachmeir, Peter Tabery, Serdar Uzumcu, Eckehard G. Steinbach
ICME4
2002 Geometry refinement for light field compression
abstract
In geometry-aided light field compression, a geometry model is used for disparity-compensated prediction of light field images from already encoded light field images. This geometry model, however, may have limited accuracy. We present an algorithm that refines a geometry model to improve the overall light field compression efficiency. This algorithm uses an optical-flow technique to explicitly minimize the disparity-compensated prediction error. Results from experiments performed on both real and synthetic data sets show bit-rate reductions of approximately 10% using the improved geometry model over a silhouette-reconstructed geometry model.
Peter Eisert, Prashant Ramanathan, Eckehard G. Steinbach, Bernd Girod
ICIP (2)3
2002 Rate-distortion optimized video streaming with adaptive playout
abstract
We propose a new scheme for streaming media systems that combines adaptive media playout (AMP) with rate-distortion (R-D)optimized packet transmission. AMP, the client-controlled, adaptive modification of the media playout rate, allows us to flexibly adjust the playout deadlines of individual packets, and can therefore reduce reconstruction distortion. This added flexibility incurs a subjective cost, however. We introduce functions that assess the subjective cost of a schedule of playout rate modifications, and we show how to optimize the schedule with respect to these costs and distortion. Because the optimal playout rate schedule and the R-D optimal transmission schedule are interdependent, we solve the two problems jointly. In simulations that model a receiver-driven scenario, results for a short media clip show a more than 2 dB improvement in mean PSNR for R-D optimized transmission scheduling combined with a moderate amount of AMP, over R-D optimized transmission scheduling alone.
Mark Kalman, Eckehard G. Steinbach, Bernd Girod
ICIP (3)2
2002 R-D optimized media streaming enhanced with adaptive media playout
abstract
Streaming media systems buffer media data at the client to improve reconstruction quality at the cost of latency. Reconstruction quality may be further improved with rate-distortion optimized packet transmission. In this paper we show how the tradeoff between buffering latency and reconstruction quality can be improved with adaptive media playout (AMP) - the client-controlled manipulation of the playout speeds of media frames. We then show how to incorporate AMP into an existing framework for rate-distortion optimized packet transmission. The resulting receiver-driven scheme jointly optimizes packet transmissions and playout speeds with respect to rate, distortion and the subjective cost of playout speed variations. Results for a short media clip show a more than 2 dB improvement in mean PSNR for R-D optimized transmission combined with a moderate amount of AMP, over R-D optimized transmission alone.
Mark Kalman, Eckehard G. Steinbach, Bernd Girod
ICME (1)2
2002 A real-time Internet streaming media testbed
abstract
We describe a real-time LAN-based testbed that allows us to investigate the behavior of streaming media applications under various network conditions. For commercially available streaming media applications, we are interested to see how they perform over next-generation wireline and wireless networks. For future streaming media applications, the testbed is an invaluable tool for the development and verification of new algorithms. Our testbed implementation is based on Linux Divert Sockets and supports a straightforward integration of various packet erasure and delay models. Individual IP-packets are diverted to a user process where they are delayed or deleted according to the desired channel model. The testbed has been used to investigate the flow-control behavior of existing streaming media systems over wireless networks. Our experiments confirm that flow-control algorithms that consider lost packets to be the result of network congestion, as employed today in wireline streaming, are not suited for wireless networks, where loss is mainly due to link impairments.
Wolfgang Kellerer, Eckehard G. Steinbach, Peter Eisert, Bernd Girod
ICME (2)2
2001 Adaptive playout for low latency video streaming
abstract
Network variations in video streaming require sufficient data to be prebuffered at the client prior to playout. This receiver buffer prevents the display process from starvation in case of network congestion. Prebuffering, however, is also responsible for the major part of the delay between requesting a media stream and playing it at the receiver. We show how adaptive media playout can be employed to reduce the delay introduced by the receiver buffer while preserving the same resilience against buffer underflow as in non-adaptive media playout. For adaptive media playout we adjust the playout speed of the media packets depending on the condition of the channel and the current client buffer fullness. We employ a two-state Markov channel model to analyze the buffer underflow-delay trade-off for our adaptive playout strategy and show that for typical parameters the average end-to-end delay can be reduced by 1 to 2 seconds.
Bernd Girod, Nikolaus Färber, Eckehard G. Steinbach
ICIP (1)3
2001 Real-time voice communication over the internet using packet path diversity
abstract
The quality of real-time voice communication over best-effort networks is mainly determined by the delay and loss characteristics observed along the network path. Excessive playout buffering at the receiver is prohibitive and significantly delayed packets have to be discarded and considered as late loss. We propose to improve the tradeoff among delay, late loss rate, and speech quality using multi-stream transmission of real-time voice over the Internet, where multiple redundant descriptions of the voice stream are sent over independent network paths. Scheduling the playout of the received voice packets is based on a novel multi-stream adaptive playout scheduling technique that uses a Lagrangian cost function to trade delay versus loss. Experiments over the Internet suggest largely uncorrelated packet erasure and delay jitter characteristics for different network paths which leads to a noticeable path diversity gain. We observe significant reductions in mean end-to-end latency and loss rates as well as improved speech quality when compared to FEC protected single-path transmission at the same data rate. In addition to our Internet measurements, we analyze the performance of the proposed multi-path voice communication scheme using the ns network simulator for different network topologies, including shared network links.
Yi J. Liang, Eckehard G. Steinbach, Bernd Girod
ACM Multimedia2
2001 Multi-stream voice over IP using packet path diversity
abstract
We propose multi-stream transmission of real-time voice over best-effort packet networks such as today's Internet, where multiple redundant descriptions of the voice stream are sent over independent network paths. At the receiver, multi-stream adaptive playout scheduling is employed to improve the tradeoff among delay, late loss rate, and speech quality. Experiments over the Internet suggest largely uncorrelated statistical characteristics, such as erasure probability and delay jitter, for different network paths, which leads to a noticeable path diversity gain. We have obtained significant reductions in mean end-to-end latency and loss rates compared to FEC protected single-path transmission at the same data rate. The speech quality perceived by the receiver is evaluated using the recently standardized objective quality measure PESQ. In our experiments, we observe gains of more than 0.4 PESQ score for voice transmission with packet path diversity.
Yi J. Liang, Eckehard G. Steinbach, Bernd Girod
MMSP2
2000 3-D Reconstruction of Real-World Objects Using Extended Voxels
abstract
In this paper we present a voxel-based 3-D reconstruction technique that computes a set of non-transparent object surface voxels from a given set of calibrated camera views. We show that the quality of the reconstruction strongly depends on the accuracy of the computed voxel projection in the image plane and discuss different approximations of the exact projection. The most simple and computationally least demanding approximation is obtained when projecting point voxels, i.e., voxels without spatial extent. However, correct occlusion handling is not possible for point voxel volumes leading to reconstruction artifacts. The most accurate projection is obtained by computing the exact outline of the projected voxels. This projection is computationally most demanding but allows correct occlusion handling during reconstruction. Experimental results that compare the reconstruction quality for point and exact voxel projection show that it is worthwhile computing the tract image plane footprint of the projected voxels.
Eckehard G. Steinbach, Bernd Girod, Peter Eisert, Arnulf Betz
ICIP1
2000 An Image-Domain Cost Function for 3-D Rigid Body Motion Estimation
abstract
We derive an image-domain cost function for calibrated 3-D rigid body motion estimation from two views. The cost function is based on the accumulation of minimum mean squared error values along the epipolar lines for a large number of small measurement windows. In comparison to other non-linear, but feature-based formulations, the minimization of this cost function does not require pre-computed point correspondences. Several speedups are derived that keep the computational complexity similar to other non-linear formulations of the 3-D rigid body motion estimation problem. The main advantage of the proposed technique is that no effort has to be devoted to the reliable determination of point correspondences.
Eckehard G. Steinbach, Bernd Girod
ICPR1
2000 3-D Object Reconstruction Using Spatially Extended Voxels and Multi-Hypothesis Voxel Coloring
abstract
We describe a voxel-based 3-D reconstruction technique from multiple calibrated camera views that makes explicit use of the finite size footprint of a voxel when projected into the image plane. We derive a class of computationally efficient axis-aligned volume traversal orders that ensure that a processed voxel cannot occlude previously processed voxels. For each view, one out of 79 different cases of volume traversal is identified depending on the relative position between camera and voxel volume. Views belonging to the same visibility class can be processed simultaneously. Our voxel coloring strategy is based on a color hypothesis test that ensures the consistency of the projected reconstruction with the original images. A surface voxel list is constantly updated during reconstruction ensuring that only a minimum number of voxels has to be processed. Experimental results that compare the reconstruction quality for voxels with and without spatial extent underscore the conclusion that it is worthwhile taking into account the exact footprint of the projected voxels.
Eckehard G. Steinbach, Bernd Girod, Peter Eisert, Arnulf Betz
ICPR1
2000 Automatic reconstruction of stationary 3-D objects from multiple uncalibrated camera views
abstract
A system for the automatic reconstruction of real-world objects from multiple uncalibrated camera views is presented. The camera position and orientation for all views, the 3-D shape of the rigid object, as well as the associated color information, are recovered from the image sequence. The system proceeds in four steps. First, the internal camera parameters describing the imaging geometry are calibrated using a reference object. Second, an initial 3-D description of the object is computed from two views. This model information is then used in a third step to estimate the camera positions for all available views using a novel linear 3-D motion and shape estimation algorithm. The main feature of this third step is the simultaneous estimation of 3-D camera-motion parameters and object shape refinement with respect to the initial 3-D model. The initial 3-D shape model exhibits only a few degrees of freedom and the object shape refinement is defined as flexible deformation of the initial shape model. Our formulation of the shape deformation allows the object texture to slide on the surface, which differs from traditional flexible body modeling. This novel combined shape and motion estimation using sliding texture considerably improves the calibration data of the individual views in comparison to fixed-shape model based camera-motion estimation. Since the shape model used for model based camera-motion estimation is only approximate, a volumetric 3-D reconstruction process is initiated in the fourth step that combines the information from ail views simultaneously. The recovered object consists of a set of voxels with associated color information that describes even fine structures and details of the object. New views of the object can be rendered from the recovered 3-D model, which has potential applications in virtual reality or multimedia systems and the emerging field of video coding using 3-D scene models.
Peter Eisert, Eckehard G. Steinbach, Bernd Girod
IEEE Trans. Circuits Syst. Video Technol.2
1999 Multi-hypothesis, volumetric reconstruction of 3-D objects from multiple calibrated camera views
abstract
In this paper we present a volumetric method for the 3-D reconstruction of real world objects from multiple calibrated camera views. The representation of the objects is fully volume-based and no explicit surface description is needed. The approach is based on multi-hypothesis tests of the voxel model back-projected into the image planes. All camera views are incorporated in the reconstruction process simultaneously and no explicit data fusion is needed. In a first step each voxel of the viewing volume is filled with several color hypotheses originating from different camera views. This leads to an overcomplete representation of the 3-D object and each voxel typically contains multiple hypotheses. In a second step only those hypotheses remain in the voxels which are consistent with all camera views where the voxel is visible. Voxels without a valid hypothesis are considered to be transparent. The methodology of our approach combines the advantages of silhouette-based and image feature-based methods. Experimental results on real and synthetic image data show the excellent visual quality of the voxel-based 3-D reconstruction.
Peter Eisert, Eckehard G. Steinbach, Bernd Girod
ICASSP2
1999 3-D Image Models and Compression: Synthetic Hybrid or Natural Fit?
abstract
This paper highlights recent advances in image compression aided by 3-D geometry information. As two examples, we present a model-aided video coder for efficient compression of head-and-shoulder scenes and a geometry-aided coder for 4-D light fields for image-based rendering. Both examples illustrate that an explicit representation of 3-D geometry is advantageous if many views of the same 3-D object or scene have to be encoded. Waveform-coding and 3-D model-based coding can be combined in a rate-distortion framework, such that the generality of waveform coding and the efficiency of 3-D models are available where needed.
Bernd Girod, Peter Eisert, Marcus A. Magnor, Eckehard G. Steinbach, Thomas Wiegand 0001
ICIP (2)4
1999 Using Multiple Global Motion Models for IMPRO VED Block-Based Video Coding
abstract
A novel motion representation and estimation scheme is presented that leads to improved coding efficiency of block-based video coders like, e.g., H.263. The proposed scheme is based on motion-compensated prediction from more than one reference frame. The reference pictures are warped or globally motion-compensated versions of the previous frame. Affine motion models are used as warping parameters approximating the motion vector field. Motion compensation is performed using standard block matching in the multiple reference frames buffer. The frame reference and the affine motion parameters are transmitted as side information. The approach is incorporated into an H.263-based video codec at minor syntax changes and embedded into a rate-constrained motion estimation and macroblock mode decision frame work. In contrast to conventional global motion compensation, where one motion model is transmitted, we show that multiple global motion models are of benefit in terms of coding efficiency. Significant coding gains in comparison to TMN-10, the test model of H.263+, are achieved that provide bit-rate savings between 20% and 35% for the various sequences tested.
Eckehard G. Steinbach, Thomas Wiegand 0001, Bernd Girod
ICIP (2)1
1999 Long-Term Memory Prediction Using Affine Motion Compensation
abstract
Long-term memory prediction extends motion compensation from the previous frame to several past frames with the result of increased coding efficiency. We demonstrate that combining long-term memory prediction with affine motion compensation leads to further coding gains. For that, various affine motion parameter sets are estimated between frames in the long-term memory buffer and the current frame. Motion compensation is conducted using standard block matching in the multiple reference frame buffer. The picture reference and the affine motion parameters are transmitted as side information. The technique is embedded into a hybrid video coder mainly following the H.263 standard. The coder control employs Lagrangian optimization for the motion estimation and macroblock mode decision. Significant bit-rate savings between 20 and 50% are achieved for the sequences tested over TMN-10 the test model of H.263+. These bit-rate savings correspond to gains in PSNR between 0.8 and 3 dB.
Thomas Wiegand 0001, Eckehard G. Steinbach, Bernd Girod
ICIP (1)2
1998 Motion-based analysis and segmentation of image sequences using 3-D scene models
Eckehard G. Steinbach, Peter Eisert, Bernd Girod
Signal Process.1
1997 Robust estimation of multi-component motion in image sequences using the epipolar constraint
abstract
Given two frames of a dynamic scene with several rigid body objects undergoing different motions in the three-dimensional space, we robustly estimate the motion and structure of each object. The least median of squares (LMedS) estimator is integrated into a robust 3D motion parameter estimation and scene structure recovery framework to deal with the multi-motion problem. Experimental results underline the capability of the approach to deal successfully with multi-component motion. We apply the approach presented in this paper to the problem of automatic insertion of artificial objects in real image sequences.
Eckehard G. Steinbach, Subhasis Chaudhuri, Bernd Girod
ICASSP1
1997 Data-Driven Multi-Frame 3D Motion Estimation
abstract
We investigate how the temporal evolution of image point correspondences over multiple frames can be exploits for 3D motion and structure estimation. In comparison to other multi-frame approaches we do not separate the establishment of feature point correspondences and the 3D motion parameter computation. As a result our approach does not suffer from the limitations of the traditional two step approaches where even small errors in point correspondence can lead to large errors in 3D motion parameters. Experimental results on synthetic and real video sequences show that the estimation error is reduced when processing more than two frames simultaneously.
Eckehard G. Steinbach, Subhasis Chaudhuri, Bernd Girod
ICIP (1)1
1997 Standard compatible extension of H.263 for robust video transmission in mobile environments
abstract
In this paper we address the problem of robust video transmission in error prone environments. The approach is compatible with the ITU-T video coding standard H.263. Fading situations in mobile networks are tolerated and the image quality degradation due to spatio-temporal error propagation is minimized utilizing a feedback channel between transmitter and receiver carrying acknowledgment information. In a first step, corrupted group of blocks (GOB's) are concealed to avoid annoying artifacts caused by decoding of an erroneous bit stream. The GOB and the corresponding frame number are reported to the transmitter via the back channel. The encoder evaluates the negative acknowledgments and reconstructs the spatial and temporal error propagation. A low complexity algorithm for real-time reconstruction of spatio-temporal error propagation is described in detail. Rapid error recovery is achieved by INTRA refreshing image regions (macroblocks) bearing visible distortion. The feedback channel method does not introduce additional delay and is particularly relevant for real-time conversational services in mobile networks. Experimental results with bursty bit error sequences simulating a Digital European Cordless Telephony (DECT) channel are presented with different combinations of forward error correction (FEC), automatic repeat on request (ARQ), and the proposed error compensation technique, Compared to the case where FEC and ARQ are used for error correction, a gain of up to 3 dB peak signal-to-noise ratio (PSNR) is observed if error compensation is employed additionally.
Eckehard G. Steinbach, Nikolaus Färber, Bernd Girod
IEEE Trans. Circuits Syst. Video Technol.1
1996 Estimation of rigid body motion and scene structure from image sequences using a novel epipolar transform
abstract
Model-based video communication systems attempt to extract information about the three-dimensional structure of the scene to be transmitted. The relative motion between camera and objects can be a powerful depth cue, if the objects are rigid or almost rigid. We present a new algorithm for estimation of rigid body motion parameters and scene structure from monocular image sequences. A novel epipolar image transform is utilized to preserve the relevant information from mean squared displaced frame difference (DFD) surfaces and thus overcomes the inherent limitations of feature correspondence methods. Our algorithm conducts a coarse-to-fine search in 5-dimensional parameter space. Relative depth values are computed for each measurement window. Experimental results are presented to demonstrate the performance of the new algorithm.
Eckehard G. Steinbach, Bernd Girod
ICASSP1
1996 3D motion and scene structure estimation with motion dependent distortion of measurement windows
abstract
We present a two stage technique for estimation of rigid body motion and scene structure from monocular image sequences. The robustness of motion recovery is considerably increased by motion dependent distortion of rectangular measurement windows. In the first stage, we estimate a set of candidate motion parameter sets from which the optimum is determined in a second step evaluating motion based image content deformation together with corresponding depth values. Experimental results show that the motion parameter estimation error can be reduced significantly.
Eckehard G. Steinbach, Alan Hanjalic, Bernd Girod
ICIP (1)1