Thomas Sikora

dblp:s/ThomasSikora · DBLP profile ↗
← Back
184ranked-venue papers
14as first author
7since 2021 · last 2025
0000-0001-5695-9614ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 172 · 13 first-author · 6 since 2021Artificial intelligence and machine learning · 8Databases, data management, data science and information retrieval · 6Applied, interdisciplinary, general and emerging computing · 4 · 1 first-authorSecurity and privacy · 3Systems, architecture and hardware · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2025 3D SMoE Splatting for Edge-aware Realtime Radiance Field Rendering
abstract
Steered Mixtures-of-Experts (SMoE) is an existing regression framework that has previously been applied for modeling and compression of 2D images and higher-dimensional imagery, including compression of light fields and light-field video. SMoE models are sparse, edge-aware representations that allow rendering of imagery with few Gaussians with excellent quality. In this paper a novel, edge-aware "3D SMoE Splatting" (3DSMoES) framework for 3D rendering is introduced, adopted to fit into the existing "3D Gaussian Splatting" (3DGS) CUDA optimization pipeline. Here, SMoE regression serves as a "plug-and-play" solution that replaces the established 3DGS regression as a novel workhorse. 3DSMoES achieves significant visual quality gains with drastically fewer Gaussian kernels compared to 3DGS. We observe up to approximately 4dB improvement in PSNR on individual scenes with kernel reductions between 20 to 50 percent. The sparse models are significantly faster to train and allow up to 30-50 percent improved rendering speeds.
Yi-Hsin Li, Thomas Sikora, Sebastian Knorr, Mårten Sjöström
SIGGRAPH Asia2
2025 Fusion Networks for Advanced Block-Overlapping SMoE Image Regression
abstract
Our challenge is to develop efficient block-overlapping regression strategies for clean and noisy images. Block-based approaches allow massive parallel computation to achieve fast rendering - which is important for many applications. In this paper we focus on block-overlapping regression with sparse Steered-Mixture-of-Experts (SMoE) models which have shown to provide excellent results for divers applications. Block-overlapping SMoE regression provides multi-hypotheses for estimating each pixel value from clean or noisy observations - usually simple averaging of hypotheses is performed. We introduce and investigate several linear and nonlinear Filters/Networks for fusing the estimates more efficiently. On the clean Kodak image test set our best Fusion Network can improve over the "average" filter with an impressive 4.6-5.7 dB gain. For noisy observations the gain is around 1.6 dB - with denoising results in the same ball-park as BM3D.The developed strategies and results appear of more general interest beyond SMoE regression and have the potential to improve also on the plethora of other block-regression approaches such as DNN Autoencoders and kernel methods in 2D and 3D Gaussian Splatting
Thomas Sikora
VCIP2
2025 Adaptive Segmentation-Based Initialization for Steered Mixture of Experts Image Regression
abstract
Kernel image regression methods have demonstrated excellent efficiency in various image processing tasks, including image and light-field compression, Gaussian Splatting, denoising and super-resolution. The estimation of parameters for these methods commonly employs gradient descent iterative optimization, which poses a significant computational burden for many applications. In this paper, we introduce a novel adaptive segmentation-based initialization method targeted for optimizing Steered-Mixture-of Experts (SMoE) gating networks and RadialBasis-Function (RBF) networks with steering kernels. The novel initialization method allocates kernels into pre-calculated image segments. The optimal number of kernels, kernel positions, and steering parameters are derived per segment in an iterative optimization and kernel sparsification procedure. The kernel information from local segments is then transferred into a global initialization, ready for use in iterative optimization of SMoE, RBF, and related kernel image regression methods. Results demonstrate significant improvements in both objective and subjective quality compared to regular grid, K-Means, deeplearning-based, and previous segmentation-based initialization methods. The proposed initialization method reduces kernel usage by 70% compared to other initialization methods while maintaining the same reconstruction quality. Furthermore, by generating initial parameters closer to optimized results, convergence time is reduced, achieving overall runtime savings of up to 50% compared to prior methods. Additionally, the method supports parallel computation, with initialization time halved when using four GPUs compared to one
Yi-Hsin Li, Sebastian Knorr, Mårten Sjöström, Thomas Sikora
IEEE Trans. Multim.4
2023 Edge-Aware Autoencoder Design for Real-Time Mixture-of-Experts Image Compression
abstract
Steered-Mixtures-of-Experts (SMoE) models provide sparse, edge-aware representations, applicable to many use-cases in image processing. This includes denoising, super-resolution and compression of 2D- and higher dimensional pixel data. Recent works for image compression indicate that compression of images based on SMoE models can provide competitive performance to the state-of-the-art. Unfortunately, the iterative model-building process at the encoder comes with excessive computational demands. In this paper we introduce a novel edge-aware Autoencoder (AE) strategy designed to avoid the time-consuming iterative optimization of SMoE models. This is done by directly mapping pixel blocks to model parameters for compression, in spirit similar to recent works on “unfolding” of algorithms, while maintaining full compatibility to the established SMoE framework. With our plug-in AE encoder, we achieve a quantum-leap in performance with encoder run-time savings by a factor of 500 to 1000 with even improved image reconstruction quality. For image compression the plug-in AE encoder has real-time properties and improves RD-performance compared to our previous works.
Elvira Fleig, Jonas Geistert, Erik Bochinski, Rolf Jongebloed, Thomas Sikora
ISCAS5
2023 Segmentation-based Initialization for Steered Mixture of Experts
abstract
The Steered-Mixture-of-Experts (SMoE) model is an edge-aware kernel representation that has successfully been explored for the compression of images, video, and higher-dimensional data such as light fields. The present work aims to leverage the potential for enhanced compression gains through efficient kernel reduction. We propose a fast segmentation-based strategy to identify a sufficient number of kernels for representing an image and giving initial kernel parametrization. The strategy implies both reduced memory footprint and reduced computational complexity for the subsequent parameter optimization, resulting in an overall faster processing time. Fewer kernels, when combined with the inherent sparsity of the SMoEs, further enhance the overall compression performance. Empirical evaluations demonstrate a gain of 0.3-1.0 dB in PSNR for a constant number of kernels, and the use of 23 % less kernels and 25 % less time for constant PSNR. The results highlight the feasibility and practicality of the approach, positioning it as a valuable solution for various image-related applications, including image compression.
Yi-Hsin Li, Mårten Sjöström, Sebastian Knorr, Thomas Sikora
VCIP4
2022 Utilizing Crowd Collectiveness to Enhance Bottleneck Detection Based on the Lagrangian Framework
abstract
At large events and other crowded areas narrowed passages create a major safety risk for the involved people. We present a novel approach to automatically detect the formation of bottleneck situations and, thus, aid decision - makers to mitigate potential safety threats. In our work we analyse the dynamics of motion by using the Lagrangian approach, known from the analysis of dynamic systems, to understand movements of groups of people. Characteristic congestion patterns can be recognised in the complex time - dependent movement dynamics with the help of Lagrangian measures. For this purpose, crowd collectiveness is introduced as a new measure, which can recognise random-looking movements of single individuals as group movements. With the new measure, an accuracy gain of 5% can be achieved for the detection of bottleneck situations compared to state-of-the-art approaches.
Maik Simon, Erik Bochinski, Markus Kuchhold, Thomas Sikora
AVSS4
2022 Accelerated Deep Lossless Image Coding with Unified Paralleleized GPU Coding Architecture
abstract
We propose Deep Lossless Image Coding (DLIC), a full resolution learned lossless image compression algorithm. Our algorithm is based on a neural network combined with an entropy encoder. The neural network performs a density estimation on each pixel of the source image. The density estimation is then used to code the target pixel, beating FLIF in terms of compression rate. Similar approaches have been attempted. However, long run times make them unfeasible for real world applications. We introduce a parallelized GPU based implementation, allowing for encoding and decoding of grayscale, 8-bit images in less than one second. Because DLIC uses a neural network to estimate the probabilities used for the entropy coder, DLIC can be trained on domain specific image data. We demonstrate this capability by adapting and training DLIC with Magnet Resonance Imaging (MRI) images.
Benjamin Lukas Cajus Barzen, Fedor Glazov, Jonas Geistert, Thomas Sikora
PCS4
2020 Edge Oriented Hierarchical Motion Estimation For Video Coding
abstract
Efficient video compression relies heavily on mitigating the temporal redundancy that exists between successive video frames. This is achieved through effective motion modelling. In conventional video coding standards, the motion of the current frame is modelled from the neighbouring frames using block-based motion estimation techniques. However, as the motion discontinuities are tied to the moving objects in a video frame, the block-based techniques are unable to model the actual motion of individual objects. In this paper, an object-based hierarchical motion estimation and prediction technique for high-efficiency video coding (HEVC) is proposed. We use an edge position difference (EPD) similarity measure, which has the ability to align the largest object in the frames, to estimate the motion of the object in the current frame from the neighbouring one. In other words, it estimates the largest object's motion instead of the whole frame's global motion. The proposed method gradually models all of the objects' motions and establishes a prediction of the current frame. The predicted frame is then exploited as an additional reference frame in the HEVC compression algorithm. Our experimental results demonstrate that our proposed approach achieves a bit rate savings with a peak signal to noise ratio (PSNR) gain over the HEVC standard.
Md. Asikuzzaman, Ashek Ahmmed, Mark R. Pickering, Thomas Sikora
ICIP4
2020 Steered Mixture-of-Experts for Light Field Images and Video: Representation and Coding
abstract
Research in light field (LF) processing has heavily increased over the last decade. This is largely driven by the desire to achieve the same level of immersion and navigational freedom for camera-captured scenes as it is currently available for CGI content. Standardization organizations such as MPEG and JPEG continue to follow conventional coding paradigms in which viewpoints are discretely represented on 2-D regular grids. These grids are then further decorrelated through hybrid DPCM/transform techniques. However, these 2-D regular grids are less suited for high-dimensional data, such as LFs. We propose a novel coding framework for higher-dimensional image modalities, called Steered Mixture-of-Experts (SMoE). Coherent areas in the higher-dimensional space are represented by single higher-dimensional entities, called kernels. These kernels hold spatially localized information about light rays at any angle arriving at a certain region. The global model consists thus of a set of kernels which define a continuous approximation of the underlying plenoptic function. We introduce the theory of SMoE and illustrate its application for 2-D images, 4-D LF images, and 5-D LF video. We also propose an efficient coding strategy to convert the model parameters into a bitstream. Even without provisions for high-frequency information, the proposed method performs comparable to the state of the art for low-to-mid range bitrates with respect to subjective visual quality of 4-D LF images. In case of 5-D LF video, we observe superior decorrelation and coding performance with coding gains of a factor of 4x in bitrate for the same quality. At least equally important is the fact that our method inherently has desired functionality for LF rendering which is lacking in other state-of-the-art techniques: (1) full zero-delay random access, (2) light-weight pixel-parallel view reconstruction, and (3) intrinsic view interpolation and super-resolution.
Ruben Verhack, Thomas Sikora, Glenn Van Wallendael, Peter Lambert
IEEE Trans. Multim.2
2019 Video-based Bottleneck Detection utilizing Lagrangian Dynamics in Crowded Scenes
abstract
Avoiding bottleneck situations in crowds is critical for the safety and comfort of people at large events or in public transportation. Based on the work of Lagrangian motion analysis we propose a novel video-based bottleneck-detector by identifying characteristic stowage patterns in crowd-movements captured by optical flow fields. The Lagrangian framework allows to assess complex time-dependent crowd-motion dynamics at large temporal scales near the bottleneck by two dimensional Lagrangian fields. In particular we propose long-term temporal filtered Finite Time Lyapunov Exponents (FTLE) fields that provide towards a more global segmentation of the crowd movements and allows to capture its deformations when a crowd is passing a bottleneck. Finally, these deformations are used for an automatic spatio-temporal detection of such situations. The performance of the proposed approach is shown in extensive evaluations on the existing Jülich and AGO-RASET datasets, that we have updated with ground truth data for spatio-temporal bottleneck analysis.
Maik Simon, Markus Kuchhold, Tobias Senst, Erik Bochinski, Thomas Sikora
AVSS5
2019 Quantized and Regularized Optimization for Coding Images Using Steered Mixtures-of-Experts
abstract
Compression algorithms that employ Mixtures-of-Experts depart drastically from standard hybrid block-based transform domain approaches as in JPEG and MPEG coders. In previous works we introduced the concept of Steered Mixtures-of-Experts (SMoEs) to arrive at sparse representations of signals. SMoEs are gating networks trained in a machine learning approach that allow individual experts to explain and harvest directional long-range correlation in the N-dimensional signal space. Previous results showed excellent potential for compression of images and videos but the reconstruction quality was mainly limited to low and medium image quality. In this paper we provide evidence that SMoEs can compete with JPEG2000 at mid-and high-range bit-rates. To this end we introduce a SMoE approach for compression of color images with specialized gates and steering experts. A novel machine learning approach is introduced that optimizes RD-performance of quantized SMoEs towards SSIM using fake quantization. We drastically improve our previous results and outperform JPEG by up to 42%.
Rolf Jongebloed, Erik Bochinski, Lieven Lange, Thomas Sikora
DCC4
2018 Extending IOU Based Multi-Object Tracking by Visual Information
abstract
Today's multi-object tracking approaches benefit greatly from nearly perfect object detections when following the popular tracking-by-detection scheme. This allows for extremely simple but accurate tracking methods which completely rely on the input detections as the high-speed IOU tracker. For real world applications, few missing detections cause a high number of ID switches and fragmentations which degrades the quality of the tracks significantly. We show that this problem can be efficiently overcome if the tracker falls back to visual single-object tracking in cases where no object detection is available. In several experiments we show for different visual trackers that the number of ID switches and fragmentations can be reduced by a large amount while maintaining high tracking speeds and outperforming the state-of-the art for the UA-DETRAC and VisDrone datasets.
Erik Bochinski, Tobias Senst, Thomas Sikora
AVSS3
2018 UA-DETRAC 2018: Report of AVSS2018 & IWT4S Challenge on Advanced Traffic Monitoring
abstract
A desirable smart traffic-monitoring and street-safety system can elicit and support the intervention of law enforcement agencies or medical staff. Recently, there has been a dramatically higher demand for such smart systems. To this end, the International Workshop on Traffic and Street Surveillance for Safety and Security (IWT4S) was organized in conjunction with the 15th IEEE International Conference on Advanced Video and Signal-based Surveillance (AVSS 2018). Our goal is to advance the state-of-the-art detection and tracking algorithms and provide a comprehensive performance evaluation for them. We evaluate 5 submitted detection and 7 submitted tracking methods on the large-scale UA-DETRAC benchmark, and the results are shared publicly on the website http://detrac-db. rit.albany.edu. We expect this challenge to advance the research and development of new detection and tracking methods for transportation applications.
Siwei Lyu, Ming-Ching Chang, Dawei Du, Wenbo Li 0001, Yi Wei 0006, Marco Del Coco, Pierluigi Carcagnì, Arne Schumann, Bharti Munjal, Dinh-Quoc-Trung Dang, Doo-Hyun Choi, Erik Bochinski, Fabio Galasso, Filiz Bunyak, Guna Seetharaman, Jang-Woon Baek, Jong Taek Lee, Kannappan Palaniappan, Kil-Taek Lim, Kiyoung Moon, Kwang-Ju Kim, Lars Wilko Sommer, Meltem Brandlmaier, Minsung Kang, Moongu Jeon, Noor Al-Shakarji, Oliver Acatay, Pyong-Kun Kim, Sikandar Amin, Thomas Sikora, Tien Ba Dinh, Tobias Senst, Vu-Gia-Hy Che, Young-Chul Lim, Yun-Su Chung
AVSS30
2018 Optical Flow Dataset and Benchmark for Visual Crowd Analysis
abstract
The performance of optical flow algorithms greatly depends on the specifics of the content and the application for which it is used. Existing and well established optical flow datasets are limited to rather particular contents from which none is close to crowd behavior analysis; whereas such applications heavily utilize optical flow. We introduce a new optical flow dataset exploiting the possibilities of a recent video engine to generate sequences with ground-truth optical flow for large crowds in different scenarios. We break with the development of the last decade of introducing ever increasing displacements to pose new difficulties. Instead we focus on real-world surveillance scenarios where numerous small, partly independent, non rigidly moving objects observed over a long temporal range pose a challenge. By evaluating different optical flow algorithms, we find that results of established datasets can not be transferred to these new challenges. In exhaustive experiments we are able to provide new insight into optical flow for crowd analysis. Finally, the results have been validated on the real-world UCF crowd tracking benchmark while achieving competitive results compared to more sophisticated state-of-the-art crowd tracking approaches.
Gregory Schröder, Tobias Senst, Erik Bochinski, Thomas Sikora
AVSS4
2018 Regularized Gradient Descent Training of Steered Mixture of Experts for Sparse Image Representation
abstract
The Steered Mixture-of-Experts (SMoE) framework targets a sparse space-continuous representation for images, videos, and light fields enabling processing tasks such as approximation, denoising, and coding. The underlying stochastic processes are represented by a Gaussian Mixture Model, traditionally trained by the Expectation-Maximization (EM) algorithm. We instead propose to use the MSE of the regressed imagery for a Gradient Descent optimization as primary training objective. Further, we extend this approach with regularization terms to enforce desirable properties like the sparsity of the model or noise robustness of the training process. Experimental evaluations show that our approach outperforms the state-of-the-art consistently by 1.5 dB to 6.1 dB PSNR for image representation.
Erik Bochinski, Rolf Jongebloed, Michael Tok, Thomas Sikora
ICIP4
2018 Scale-Adaptive Real-Time Crowd Detection and Counting for Drone Images
abstract
We propose a scale-adaptive crowd detection and counting approach for drone images. Based on local feature points and density estimation considering the image scale, we detect dense crowds over multiple distances and introduce an extremely fast counting strategy with high accuracy for our detected crowd regions. We compare our results with a recent CNN-based state-of-the-art approach and validate both methods for different scaling factors on a novel crowd dataset. The results show that our proposed method outperforms the pre-trained CNN-based approach and receives very precise counting results for different zoom factors, resolutions and crowd sizes. Its low computational complexity makes it highly suitable for real-time analysis or embedded systems.
Markus Kuchhold, Maik Simon, Volker Eiselein, Thomas Sikora
ICIP4
2018 Restricted Boltzmann Machine Image Compression
abstract
We propose a novel lossy block-based image compression approach. Our approach builds on non-linear autoencoders that can, when properly trained, explore non-linear statistical dependencies in the image blocks for redundancy reduction. In contrast the DCT employed in JPEG is inherently restricted to exploration of linear dependencies using a second-order statistics framework. The coder is based on pre-trained class-specific Restricted Boltzmann Machines (RBM). These machines are statistical variants of neural network autoencoders that directly map pixel values in image blocks into coded bits. Decoders can be implemented with low computational complexity in a codebook design. Experimental results show that our RBM-codec outperforms JPEG at high compression rates, both in terms of PSNR, SSIM and subjective results.
Markus Kuchhold, Maik Simon, Thomas Sikora
PCS3
2018 An MSE Approach For Training And Coding Steered Mixtures Of Experts
abstract
Previous research has shown the interesting properties and potential of Steered Mixtures-of-Experts (SMoE) for image representation, approximation, and compression based on EM optimization. In this paper we introduce an MSE optimization method based on Gradient Descent for training SMoEs. This allows improved optimization towards PSNR and SSIM and de-coupling of experts and gates. In consequence we can now generate very high quality SMoE models with significantly reduced model complexity compared to previous work and much improved edge representations. Based on this strategy a block-based image coder was developed using Mixture-of-Experts that uses very simple experts with very few model parameters. Experimental evaluations shows that a significant compression gain can be achieved compared to JPEG for low bit rates.
Michael Tok, Rolf Jongebloed, Lieven Lange, Erik Bochinski, Thomas Sikora
PCS5
2018 Progressive Modeling of Steered Mixture-of-Experts for Light Field Video Approximation
abstract
Steered Mixture-of-Experts (SMoE) is a novel framework for the approximation, coding, and description of image modalities. The future goal is to arrive at a representation for Six Degrees-of-Freedom (6DoF) image data. The goal of this paper is to introduce SMoE for 4D light field videos by including the temporal dimension. However, these videos contain vast amounts of samples due to the large number of views per frame. Previous work on static light field images mitigated the problem by hard subdividing the modeling problem. However, such a hard subdivision introduces visually disturbing block artifacts on moving objects in dynamic image data. We propose a novel modeling method that does not result in block artifacts while minimizing the computational complexity and which allows for a varying spread of kernels in the spatio-temporal domain. Experiments validate that we can progressively model light field videos with increasing objective quality up to 0.97 SSIM.
Ruben Verhack, Glenn Van Wallendael, Martijn Courteaux, Peter Lambert, Thomas Sikora
PCS5
2017 High-Speed tracking-by-detection without using image information
abstract
Tracking-by-detection is a common approach to multi-object tracking. With ever increasing performances of object detectors, the basis for a tracker becomes much more reliable. In combination with commonly higher frame rates, this poses a shift in the challenges for a successful tracker. That shift enables the deployment of much simpler tracking algorithms which can compete with more sophisticated approaches at a fraction of the computational cost. We present such an algorithm and show with thorough experiments its potential using a wide range of object detectors. The proposed method can easily run at 100K fps while outperforming the state-of-the-art on the DETRAC vehicle tracking dataset.
Erik Bochinski, Volker Eiselein, Thomas Sikora
AVSS3
2017 Assessing post-detection filters for a generic pedestrian detector in a tracking-by-detection scheme
abstract
Tracking-by-detection becomes more and more popular for visual pedestrian tracking applications. However, it requires accurate and reliable detections in order to obtain good results. In this work, we propose two different post-detection filters designed to enhance the performance of custom person detectors. Using a popular deformable-parts-based pedestrian detector as a baseline, a detailed comparison over multiple test videos is performed and the gain of both algorithms is proven. Further analysis shows that the improved detection outcomes also lead to improved tracking results. We thus found that the usage of the proposed post-detection filters is recommendable as they do not impose a high computational load and are not limited to a specific detector method.
Volker Eiselein, Erik Bochinski, Thomas Sikora
AVSS3
2017 Sequential sensor fusion combining probability hypothesis density and kernelized correlation filters for multi-object tracking in video data
abstract
This work applies the Gaussian Mixture Probability Hypothesis Density (GMPHD) Filter to multi-object tracking in video data. In order to take advantage of additional visual information, Kernelized Correlation Filters (KCF) are evaluated as a possible extension of the GMPHD tracking-by-detection scheme to enhance its performance. The baseline GMPHD filter and its extension are evaluated on the UA-DETRAC benchmark, showing that combining both methods leads to a higher recall and a better quality of object tracks to the cost of increased computational complexity and increased sensitivity to false-positives.
Tino Kutschbach, Erik Bochinski, Volker Eiselein, Thomas Sikora
AVSS4
2017 UA-DETRAC 2017: Report of AVSS2017 & IWT4S Challenge on Advanced Traffic Monitoring
abstract
The rapid advances of transportation infrastructure have led to a dramatic increase in the demand for smart systems capable of monitoring traffic and street safety. Fundamental to these applications are a community-based evaluation platform and benchmark for object detection and multi-object tracking. To this end, we organize the AVSS2017 Challenge on Advanced Traffic Monitoring, in conjunction with the International Workshop on Traffic and Street Surveillance for Safety and Security (IWT4S), to evaluate the state-of-the-art object detection and multi-object tracking algorithms in the relevance of traffic surveillance. Submitted algorithms are evaluated using the large-scale UA-DETRAC benchmark and evaluation protocol. The benchmark, the evaluation toolkit and the algorithm performance are publicly available from the website http://detrac-db.rit.albany.edu.
Siwei Lyu, Ming-Ching Chang, Dawei Du, Longyin Wen, Honggang Qi, Yuezun Li, Yi Wei 0006, Lipeng Ke, Tao Hu 0011, Marco Del Coco, Pierluigi Carcagnì, Dmitriy Anisimov, Erik Bochinski, Fabio Galasso, Filiz Bunyak, Hao Ye 0005, Hong Wang 0014, Kannappan Palaniappan, Koray Ozcan, Li Wang 0033, Liang Wang 0001, Martin Lauer, Nattachai Watcharapinchai, Nenghui Song, Noor Al-Shakarji, Sikandar Amin, Sitapa Watcharapinchai, Tatiana Khanova, Thomas Sikora, Tino Kutschbach, Volker Eiselein, Wei Tian 0001, Xiangyang Xue 0001, Xiaoyi Yu, Yao Lu 0028, Yingbin Zheng, Yongzhen Huang, Yuqi Zhang 0001
AVSS31
2017 Color prediction in image coding using Steered Mixture-of-Experts
abstract
We propose a novel approach for modeling and coding color in images and video. Luminance is linearly correlated with chrominance locally, as such we can predict color given the luma value. Using the Steered Mixture-of-Experts (SMoE) approach, the image is viewed as a stochastic process over 5 random variables including the 2-D pixel locations, 1 luminance and 2 chrominance values. We model this process as a continuous joint density function by fitting a K-modal 5-D Gaussian Mixture Model (GMM). As such, the chroma values are predicted as the expectation of the conditional density. To validate, the technique was integrated within JPEG showing PSNR gains in the lower bitrate regions. A deeper analysis of the tolerance of the activation function is given through recycling color models in video sequences, yielding a high quality reconstruction over a considerable range of frames.
Ruben Verhack, Simon Van De Keer, Glenn Van Wallendael, Thomas Sikora, Peter Lambert
ICASSP4
2017 Hyper-parameter optimization for convolutional neural network committees based on evolutionary algorithms
abstract
In a broad range of computer vision tasks, convolutional neural networks (CNNs) are one of the most prominent techniques due to their outstanding performance. Yet it is not trivial to find the best performing network structure for a specific application because it is often unclear how the network structure relates to the network accuracy. We propose an evolutionary algorithm-based framework to automatically optimize the CNN structure by means of hyper-parameters. Further, we extend our framework towards a joint optimization of a committee of CNNs to leverage specialization and cooperation among the individual networks. Experimental results show a significant improvement over the state-of-the-art on the well-established MNIST dataset for hand-written digits recognition.
Erik Bochinski, Tobias Senst, Thomas Sikora
ICIP3
2017 A consistent two-level metric for evaluation of automated abandoned object detection methods
abstract
Scientific interest in automated abandoned object detection algorithms using visual information is high and many related systems have been published in recent years. However, most evaluation techniques rely only on statistical evaluation on the object level. Therefore and due to benchmarks with commonly only few abandoned objects and a non-standardized evaluation procedure, an objective performance comparison between different methods is generally hard. We propose a new evaluation metric which is focused on an end-user application case and an evaluation protocol which eliminates uncertainties in previous performance assessments. Using two variants of an abandoned object detection method, we show the features of the novel metric on multiple datasets proving its advantages over previously used measures.
Patrick Krusch, Erik Bochinski, Volker Eiselein, Thomas Sikora
ICIP4
2017 Steered mixture-of-experts for light field coding, depth estimation, and processing
abstract
The proposed framework, called Steered Mixture-of-Experts (SMoE), enables a multitude of processing tasks on light fields using a single unified Bayesian model. The underlying assumption is that light field rays are instantiations of a non-linear or non-stationary random process that can be modeled by piecewise stationary processes in the spatial domain. As such, it is modeled as a space-continuous Gaussian Mixture Model. Consequently, the model takes into account different regions of the scene, their edges, and their development along the spatial and disparity dimensions. Applications presented include light field coding, depth estimation, edge detection, segmentation, and view interpolation. The representation is compact, which allows for very efficient compression yielding state-of-the-art coding results for low bit-rates. Furthermore, due to the statistical representation, a vast amount of information can be queried from the model even without having to analyze the pixel values. This allows for “blind” light field processing and classification.
Ruben Verhack, Thomas Sikora, Lieven Lange, Rolf Jongebloed, Glenn Van Wallendael, Peter Lambert
ICME2
2017 Crowd Violence Detection Using Global Motion-Compensated Lagrangian Features and Scale-Sensitive Video-Level Representation
abstract
Lagrangian theory provides a rich set of tools for analyzing non-local, long-term motion information in computer vision applications. Based on this theory, we present a specialized Lagrangian technique for the automated detection of violent scenes in video footage. We present a novel feature using Lagrangian direction fields that is based on a spatio-temporal model and uses appearance, background motion compensation, and long-term motion information. To ensure appropriate spatial and temporal feature scales, we apply an extended bag-of-words procedure in a late-fusion manner as a classification scheme on a per-video basis. We demonstrate that the temporal scale, captured by the Lagrangian integration time parameter, is crucial for violence detection and show how it correlates to the spatial scale of characteristic events in the scene. The proposed system is validated on multiple public benchmarks and non-public, real-world data from the London Metropolitan Police. Our experiments confirm that the inclusion of Lagrangian measures is a valuable cue for automated violence detection and increases the classification performance considerably compared with the state-of-the-art methods.
Tobias Senst, Volker Eiselein, Alexander Kuhn, Thomas Sikora
IEEE Trans. Inf. Forensics Secur.4
2016 Training a convolutional neural network for multi-class object detection using solely virtual world data
abstract
Convolutional neural networks are a popular choice for current object detection and classification systems. Their performance improves constantly but for effective training, large, hand-labeled datasets are required. We address the problem of obtaining customized, yet large enough datasets for CNN training by synthesizing them in a virtual world, thus eliminating the need for tedious human interaction for ground truth creation. We developed a CNN-based multi-class detection system that was trained solely on virtual world data and achieves competitive results compared to state-of-the-art detection systems.
Erik Bochinski, Volker Eiselein, Thomas Sikora
AVSS3
2016 Robust local optical flow: Long-range motions and varying illuminations
abstract
Sparse motion estimation with local optical flow methods is fundamental for a wide range of computer vision application. Classical approaches like the pyramidal Lucas-Kanade method (PLK) or more sophisticated approaches like the Robust Local Optical Flow (RLOF) fail when it comes to environments with illumination changes and/or long-range motions. In this work we focus on these limitations and propose a novel local optical flow framework taking into account an illumination model to deal with varying illumination and a prediction step based on a perspective global motion model to deal with long-range motions. Experimental results shows tremendous improvements, e.g. 56% smaller error for dense motion fields on the KITTI and an about 76% smaller error for sparse motion fields on the Sintel dataset.
Tobias Senst, Jonas Geistert, Thomas Sikora
ICIP3
2016 A universal image coding approach using sparse steered Mixture-of-Experts regression
abstract
Our challenge is the design of a “universal” bit-efficient image compression approach. The prime goal is to allow reconstruction of images with high quality. In addition, we attempt to design the coder and decoder “universal”, such that MPEG-7-like low-and mid-level descriptors are an integral part of the coded representation. To this end, we introduce a sparse Mixture-of-Experts regression approach for coding images in the pixel domain. The underlying stochastic process of the pixel amplitudes are modelled as a 3-dimensional and multi-modal Mixture-of-Gaussians with K modes. This closed form continuous analytical model is estimated using the Expectation-Maximization algorithm and describes segments of pixels by local 3-D Gaussian steering kernels with global support. As such, each component in the mixture of experts steers along the direction of highest correlation. The conditional density then serves as the regression function. Experiments show that a considerable compression gain is achievable compared to JPEG for low bitrates for a large class of images, while forming attractive low-level descriptors for the image, such as the local segmentation boundaries, direction of intensity flow and the distribution of these parameters over the image.
Ruben Verhack, Thomas Sikora, Lieven Lange, Glenn Van Wallendael, Peter Lambert
ICIP2
2016 Robust local optical flow: Dense motion vector field interpolation
abstract
Optical flow methods integrating sparse point correspondences have made significant contribution in the field of optical flow estimation. Especially for the goal of estimating motion accurately and efficiently, sparse-to-dense interpolation schemes for feature point matches have shown outstanding performances. Concurrently, local optical flow methods have been significantly improved with respect to long-range motion estimation in environments with varying illumination. This motivates us to propose a sparse-to-dense approach based on the Robust Local Optical Flow method. Compared to state-of-the-art methods the proposed approach is significantly faster while retaining competitive accuracy on Middlebury, KITTI 2015 and MPI-Sintel data-set.
Jonas Geistert, Tobias Senst, Thomas Sikora
PCS3
2016 Video representation and coding using a sparse steered mixture-of-experts network
abstract
In this paper, we introduce a novel approach for video compression that explores spatial as well as temporal redundancies over sequences of many frames in a unified framework. Our approach supports “compressed domain vision” capabilities. To this end, we developed a sparse Steered Mixture-of-Experts (SMoE) regression network for coding video in the pixel domain. This approach drastically departs from the established DPCM/Transform coding philosophy. Each kernel in the Mixture-of-Experts network steers along the direction of highest correlation, both in spatial and temporal domain, with local and global support. Our coding and modeling philosophy is embedded in a Bayesian framework and shows strong resemblance to Mixture-of-Experts neural networks. Initial experiments show that at very low bit rates the SMoE approach can provide competitive performance to H.264.
Lieven Lange, Ruben Verhack, Thomas Sikora
PCS3
2015 Image guided phase unwrapping for real-time 3D-scanning
abstract
3D-reconstructions produced by active 3D-scanning systems based on structured light can achieve high accuracy reconstructions of the scene surfaces. Structured light algorithms based on phase measuring triangulation (PMT) utilize phase-shifted sinusoidal patterns projected into the scene for a precise determination of correspondencies. The number of patterns used for that purpose may vary depending on the design of the algorithm. No matter how many patterns are required, all of these algorithms suffer from the acquisition time needed to record all patterns sequentially. In case of a dynamic scene the sequential acquisition of images lead to the capture of dynamic objects in different poses which in turn result in erroneous reconstructions depending on the object's velocity. Our goal is to achieve a more robust result during dynamic scene capture as well as better scene reconstruction rate. Two novel approaches are presented to reduce the amount of required patterns for a high-accuracy 3D-reconstruction. This is achieved by incorporating passive matching techniques in the phase-unwrapping stage of the algorithm, allowing to drop one half of the sinusoidal patterns.
Thilo Borgmann, Michael Tok, Thomas Sikora
PCS3
2015 A novel Kernel PCA/KLT approach for transform coding of waveforms
abstract
A novel Kernel PCA/Kernel KLT transform (S-KPCA) is introduced which incorporates higher order statistics into the design of the transform matrix using a Reproducing Kernel Hilbert Space (RKHS) approach. The goal is to arrive at an orthonormal transform matrix E with column eigenvectors that allow reconstruction of an input vector with few coefficients and superior signal fidelity. In contrast to the well known Kernel PCA the number of the generated transform coefficients is not dependent on the size of the training set and the “pre-image problem” is avoided completely. Results indicate that the derived transform is more compact than the standard PCA/KLT in terms of fidelity measures in RKHS.
Thomas Sikora
PCS1
2015 Motion modeling for motion vector coding in HEVC
abstract
During the standardization of HEVC, new motion information coding and prediction schemes such as temporal motion vector prediction have been investigated to reduce the spatial redundancy of motion vector fields used for motion compensated inter prediction. In this paper a general motion model based vector coding scheme is introduced. This scheme includes estimation, coding and dynamic recombination of parametric motion models to generate vector predictors and merge candidates for all common HEVC inter coding settings. Bit rate reductions of up to 4.9% indicate that higher order motion models can increase the efficiency of motion information coding in modern hybrid video coding standards.
Michael Tok, Volker Eiselein, Thomas Sikora
PCS3
2015 Lossless image compression based on Kernel Least Mean Squares
abstract
This paper introduces a novel approach for coding luminance images using kernel-based adaptive filtering and context-adaptive arithmetic coding. This approach tackles the problem that is present in current image and video coders; these coders depend on assumptions of the image and are constrained by the linearity of their predictors. The efficacy of the predictors determines the compression gain. The goal is to create a generic image coder that learns and adapts to the characteristics of the signals and handles nonlinearity in the prediction. Results show that pixel luminance prediction using the Kernel Least Mean Squares (KLMS) yields a significant gain compared to the standard Least Mean Squares algorithm. By coding the residual using a Context-Adaptive Arithmetic Coder (CAAC), the codec is able to outperform the current industry standards of lossless image coding. An average bitrate reduction of more than 2.5% is found for the used test set.
Ruben Verhack, Lieven Lange, Peter Lambert, Rik Van de Walle, Thomas Sikora
PCS5
2015 Spatio-temporal crowd density model in a human detection and tracking framework
Hajer Fradi, Volker Eiselein, Jean-Luc Dugelay, Ivo Keller, Thomas Sikora
Signal Process. Image Commun.5
2015 Evaluation of the wavelet image two-line coder: A low complexity scheme for image compression
Stephan Rein, Frank H. P. Fitzek, Clemens Gühmann, Thomas Sikora
Signal Process. Image Commun.4
2014 Theoretical Considerations Concerning Pixelwise Temporal Filtering
abstract
Temporal in loop filters present one possible way to reduce noise introduced in compressed video sequences at low bit rates. Some of these filtering approaches make use of the quantized and generally noisy motion information conveyed in the bit stream generated by the encoder. One key feature of such filters is an adaptive filter length depending on the image content and the quality of the motion field. This paper derives mathematical equations to model the behaviour of one such filter in the presence of noisy motion vectors. The predicted optimal filter lengths are demonstrated to have a global optimum. They also show strong correlation with a real-world implementation of the previously introduced Temporal Trajectory Filter based on the HEVC main profile.
Marko Esche, Michael Tok, Thomas Sikora
DCC3
2014 Cross based robust local optical flow
abstract
In many computer vision applications local optical flow methods are still a widely used. Such methods, like the Pyramidal Lucas Kanade and the Robust Local Optical Flow, have to address the trade-off between run time and accuracy. In this work we propose an extension to these methods that improves the accuracy especially at object boundaries. This extension makes use of the cross based variable support region generation proposed in [1] accounting for local intensity discontinuities. In the evaluation using Middlebury data set we prove the ability of the proposed extension to increase the accuracy by a slight increase of run time.
Tobias Senst, Thilo Borgmann, Ivo Keller, Thomas Sikora
ICIP4
2014 Crowd analysis in non-static cameras using feature tracking and multi-person density
abstract
We propose a new methodology for crowd analysis by introducing the concept of Multi-Person Density. Using a state-of-the-art feature tracking algorithm, representative low-level features and their long-term motion information are extracted and combined into a human detection model. In contrast to previously proposed techniques, the proposed method takes small camera motion into account and is not affected by camera shaking. This increases the robustness of separating crowd features from background and thus opens a whole new field for application of these techniques in non-static CCTV cameras. We show the effectiveness of our approach on various test videos and compare it to state-of-the-art people counting methods.
Tobias Senst, Volker Eiselein, Ivo Keller, Thomas Sikora
ICIP4
2014 Lossy image coding in the pixel domain using a sparse steering kernel synthesis approach
abstract
Kernel regression has been proven successful for image de-noising, deblocking and reconstruction. These techniques lay the foundation for new image coding opportunities. In this paper, we introduce a novel compression scheme: Sparse Steering Kernel Synthesis Coding (SSKSC). This pre- and postprocessor for JPEG performs non-uniform sampling based on the smoothness of an image, and reconstructs the missing pixels using adaptive kernel regression. At the same time, the kernel regression reduces the blocking artifacts from the JPEG coding. Crucial to this technique is that non-uniform sampling is performed while maintaining only a small overhead for signalization. Compared to JPEG, SSKSC achieves a compression gain for low bits-per-pixel regions of 50% or more for PSNR and SSIM. A PSNR gain is typically in the 0.0-0.5 bpp range, and an SSIM gain can mostly be achieved in the 0.0-1.0 bpp range.
Ruben Verhack, Andreas Krutz, Peter Lambert, Rik Van de Walle, Thomas Sikora
ICIP5
2014 Real-time generation of multi-view video plus depth content using mixed narrow and wide baseline
Frederik Zilly, Christian Riechert, Marcus Müller 0001, Peter Eisert, Thomas Sikora, Peter Kauff
J. Vis. Commun. Image Represent.5
2014 Adaptively Splitted GMM With Feedback Improvement for the Task of Background Subtraction
abstract
Per pixel adaptive Gaussian mixture models (GMMs) have become a popular choice for the detection of change in the video surveillance domain because of their ability to cope with many challenges characteristic for surveillance systems in real time with low memory requirements. Since their first introduction in the surveillance domain, GMM has been enhanced in many directions. In this paper, we present a study of some relevant GMM approaches and analyze their underlying assumptions and design decisions. Based on this paper, we show how these systems can be further improved by means of a variance controlling scheme and the incorporation of region analysis-based feedback. The proposed system has been thoroughly evaluated using the extensive data set of the IEEE Workshop on Change Detection, showing an outranking performance in comparison with state-of-the-art methods.
Rubén Heras Evangelio, Michael Pätzold, Ivo Keller, Thomas Sikora
IEEE Trans. Inf. Forensics Secur.4
2014 Creating Experts From the Crowd: Techniques for Finding Workers for Difficult Tasks
abstract
Crowdsourcing is currently used for a range of applications, either by exploiting unsolicited user-generated content, such as spontaneously annotated images, or by utilizing explicit crowdsourcing platforms such as Amazon Mechanical Turk to mass-outsource artificial-intelligence-type jobs. However, crowdsourcing is most often seen as the best option for tasks that do not require more of people than their uneducated intuition as a human being. This article describes our methods for identifying workers for crowdsourced tasks that are difficult for both machines and humans. It discusses the challenges we encountered in qualifying annotators and the steps we took to select the individuals most likely to do well at these tasks.
Luke R. Gottlieb, Gerald Friedland, Jaeyoung Choi 0002, Pascal Kelm, Thomas Sikora
IEEE Trans. Multim.5
2013 Enhancing human detection using crowd density measures and an adaptive correction filter
abstract
In this paper we present a method of improving a human detector by means of crowd density information. Human detection is especially challenging in crowded scenes which makes it important to introduce additional knowledge into the detection process. We compute crowd density maps in order to estimate the spatial distribution of people in the scene and show how it is possible to enhance the detection results of a state-of-the-art human detector by this information. The proposed method applies a self-adaptive, dynamic parametrization and as an additional contribution uses scene-adaptive learning of the human aspect ratio in order to reduce false positive detections in crowded areas. We evaluate our method on videos from different datasets and demonstrate how our system achieves better results than the baseline algorithm.
Volker Eiselein, Hajer Fradi, Ivo Keller, Thomas Sikora, Jean-Luc Dugelay
AVSS4
2013 Multiple cue indexing and summarization of surveillance video
abstract
In this paper we propose a system for the summarization of safety and security surveillance video. By combining the information provided by multiple analysis cues, we improve the quality of the information extracted out of the analyzed video sequences with respect to the state-of-the-art approaches, therefore, being able to generate summaries that better align with the content of the original video. The proposed system has been tested using an extensive set of surveillance sequences, showing compression ratios ranging from 11 to 114, depending on the video content and the configuration of the system.
Rubén Heras Evangelio, Ivo Keller, Thomas Sikora
AVSS3
2013 Efficient Quadtree Compression for Temporal Trajectory Filtering
abstract
Summary form only given. Spatial in loop filters are a well established tool to improve the compression performance of today's video codecs. Temporal denoising and deblocking filters have recently also received some attention, because of their ability to stabilize pictures and to reduce flickering artifacts. One such filter, the previously introduced Quad tree-based Temporal Trajectory Filter, can produce good results, provided that the associated quad tree is sufficiently detailed. In this paper a novel, generally applicable scheme to compress such quad tree information is presented. In addition, the performance of the filter within the current HEVC test model HM 8.0 is investigated.
Marko Esche, Michael Tok, Alexander Glantz, Andreas Krutz, Thomas Sikora
DCC5
2013 A Parametric Merge Candidate for High Efficiency Video Coding
abstract
Block based motion compensated prediction still is the main technique used for temporal redundancy reduction in modern hybrid video codecs. However, the resulting motion vector fields are highly redundant as well. So, motion vector prediction and difference coding are used to compress such vector fields. A drawback of common motion vector prediction techniques is their inability to predict complex motion such as rotation and zoom in an efficient way. We present a novel Merge candidate for improving already existing vector prediction techniques based on higher order motion models to overcome this issue. To transmit the needed models, an efficient compression scheme is utilized. The improvement results in bit rate savings of 1.7% in average and up to 4% respectively.
Michael Tok, Marko Esche, Alexander Glantz, Andreas Krutz, Thomas Sikora
DCC5
2013 Adpative dense vector field interpolation for temporal filtering
abstract
Inloop filters are a well known tool to improve the compression performance of hybrid video codecs. These filters generally work in the spatial domain. If temporal filters are used instead, their performance strongly depends on the available motion data. In this paper a new method to retrieve accurate motion information per pixel is introduced and evaluated within the HEVC test model HM 8.0. The interpolated motion vector field is significnatly closer to the real object motion than the block-based motion data. An average bit rate reduction of 0.4% is observed with a maximum reduction of 2.0%.
Marko Esche, Michael Tok, Thomas Sikora
ICIP3
2013 Consensus-based multiview texturing and depth-map completion
abstract
We describe an active depth imaging system based on phase measuring triangulation. Typically depth-maps generated with such 3D scanning systems suffer from occlusions and imperfections, especially in the vicinity of depth discontinuities. Applying multiple color images, captured with a camera array, for view synthesis from the erroneous depth-maps can result in severe texturing artifacts. Our consensus-based approach greatly reduces these artifacts by comparing the similarity of the multiview texture images during the blending process to detect outliers in the form of foreground texture projected on background surfaces and specular ambiguity. Additionally, the approach is applied to dramatically improve the depth-maps by generating multiple depth-map hypotheses and selecting the areas of each that have the highest consensus with the set of multiview texture images. Our approach yields accurate and occlusion-free depth-maps in real-time.
Kai Ide, Ivo Keller, Thomas Sikora
ICIP3
2013 Robust local optical flow estimation using bilinear equations for sparse motion estimation
abstract
This article presents a theoretical framework to decrease the computation effort of the Robust Local Optical Flow method which is based on the Lucas Kanade method. We show mathematically, how to transform the iterative scheme of the feature tracker into a system of bilinear equations and thus estimate the motion vectors directly by analyzing its zeros. Furthermore, we show that it is possible to parallelise our approach efficiently on a GPU, thus, outperforming the current OpenCV-OpenCL implementation of the pyramidal Lucas Kanade method in terms of runtime and accuracy. Finally, an evaluation is given for the Middlebury Optical Flow and the KITTI datasets.
Tobias Senst, Jonas Geistert, Ivo Keller, Thomas Sikora
ICIP4
2013 A dynamic model buffer for parametric motion vector prediction in random-access coding scenarios
abstract
Motion compensated inter prediction is a powerful tool used in modern hybrid video codecs to reduce the temporal redundancy of video sequences. However, the motion information needed for motion compensation is highly redundant as well. Thus, motion vector prediction and difference coding is a common method in modern video codecs. During the standardization of HEVC, new methods for motion prediction such as temporal motion vector prediction have been analyzed. This paper presents a method for motion vector prediction from perspective motion models in random access scenarios with hierarchical group of picture structures. To enable this kind of prediction a dynamic buffer system for generating, compressing and transmitting the underlying motion models is introduced. Bit rate reductions of up to 5% underline the performance of the complete system.
Michael Tok, Marko Esche, Thomas Sikora
ICIP3
2013 Human vs machine: establishing a human baseline for multimodal location estimation
abstract
Over the recent years, the problem of video location estimation (i.e., estimating the longitude/latitude coordinates of a video without GPS information) has been approached with diverse methods and ideas in the research community and significant improvements have been made. So far, however, systems have only been compared against each other and no systematic study on human performance has been conducted. Based on a human-subject study with 11,900 experiments, this article presents a human baseline for location estimation for different combinations of modalities (audio, audio/video, audio/video/text). Furthermore, this article compares state-of-the-art location estimation systems with the human baseline. Although the overall performance of humans' multimodal video location estimation is better than current machine learning approaches, the difference is quite small: For 41% of the test set, the machine's accuracy was superior to the humans. We present case studies and discuss why machines did better for some videos and not for others. Our analysis suggests new directions and priorities for future work on the improvement of location inference algorithms.
Jaeyoung Choi 0002, Howard Lei, Venkatesan N. Ekambaram, Pascal Kelm, Luke R. Gottlieb, Thomas Sikora, Kannan Ramchandran, Gerald Friedland
ACM Multimedia6
2013 DCT-based features for categorisation of social media in compressed domain
abstract
These days the sharing of videos is very popular in social networks. Many of these social media websites such as Flickr, Facebook and YouTube allows the user to manually label their uploaded videos with textual information. However, the manually labelling for a large set of social media is still boring and error-prone. For this reason we present a algorithm for categorisation of videos in social media platforms without decoding them. The paper shows a data-driven approach which makes use of global and local features from the compressed domain and achieves a mean average precision of 0.2498 on the Blip10k dataset. In comparison with existing retrieval approaches at the MediaEval Tagging Task 2012 we will show the effectiveness and high accuracy relative to the state-of-the art solutions.
Sebastian Schmiedeke, Pascal Kelm, Thomas Sikora
MMSP3
2013 Blip10000: a social video dataset containing SPUG content for tagging and retrieval
abstract
The increasing amount of digital multimedia content available is inspiring potential new types of user interaction with video data. Users want to easily find the content by searching and browsing. For this reason, techniques are needed that allow automatic categorisation, searching the content and linking to related information. In this work, we present a dataset that contains comprehensive semi-professional user-generated (SPUG) content, including audiovisual content, user-contributed metadata, automatic speech recognition transcripts, automatic shot boundary files, and social information for multiple 'social levels'. We describe the principal characteristics of this dataset and present results that have been achieved on different tasks.
Sebastian Schmiedeke, Isabelle Ferrané, Maria Eskevich, Christoph Kofler, Martha A. Larson, Yannick Estève, Lori Lamel, Gareth J. F. Jones, Thomas Sikora
MMSys10
2013 Motion-based object segmentation using hysteresis and bidirectional inter-frame change detection in sequences with moving camera
Marina Georgia Arvanitidou, Michael Tok, Alexander Glantz, Andreas Krutz, Thomas Sikora
Signal Process. Image Commun.5
2013 Monte-Carlo-Based Parametric Motion Estimation Using a Hybrid Model Approach
abstract
Parametric motion estimation is an important task for various video processing applications, such as analysis, segmentation, and coding. The process for such an estimation has to satisfy three requirements. It has to be fast, accurate, and robust in the presence of arbitrarily moving foreground objects. We introduce a two-step simplification scheme, suitable for Monte-Carlo-based perspective motion model estimation. For complexity reduction, the Helmholtz tradeoff estimator as well as random sample consensus are enhanced with this scheme and applied on Kanade-Lucas-Tomasi features as well as on video stream macroblock motion vector fields. For the feature-based estimation, good trackable features are detected and tracked on raw video sequences. For the block-based approach, motion vector fields from encoded H.264/AVC video streams are used. Results indicate that the complexity of the whole estimation process can be reduced by a factor of up to 10000 compared to state-of-the-art methods without losing estimation precision.
Michael Tok, Alexander Glantz, Andreas Krutz, Thomas Sikora
IEEE Trans. Circuits Syst. Video Technol.4
2013 Adaptive Image Warping for Hole Prevention in 3D View Synthesis
abstract
Increasing popularity of 3D videos calls for new methods to ease the conversion process of existing monocular video to stereoscopic or multi-view video. A popular way to convert video is given by depth image-based rendering methods, in which a depth map that is associated with an image frame is used to generate a virtual view. Because of the lack of knowledge about the 3D structure of a scene and its corresponding texture, the conversion of 2D video, inevitably, however, leads to holes in the resulting 3D image as a result of newly-exposed areas. The conversion process can be altered such that no holes become visible in the resulting 3D view by superimposing a regular grid over the depth map and deforming it. In this paper, an adaptive image warping approach as an improvement to the regular approach is proposed. The new algorithm exploits the smoothness of a typical depth map to reduce the complexity of the underlying optimization problem that is necessary to find the deformation, which is required to prevent holes. This is achieved by splitting a depth map into blocks of homogeneous depth using quadtrees and running the optimization on the resulting adaptive grid. The results show that this approach leads to a considerable reduction of the computational complexity while maintaining the visual quality of the synthesized views.
Nils Plath, Sebastian Knorr, Lutz Goldmann, Thomas Sikora
IEEE Trans. Image Process.4
2012 Real-Time Multi-human Tracking Using a Probability Hypothesis Density Filter and Multiple Detectors
abstract
The Probability Hypothesis Density (PHD) filter is a multi-object Bayes filter which has recently attracted a lot of interest in the tracking community mainly for its linear complexity and its ability to deal with high clutter especially in radar/sonar scenarios. In the computer vision community however, underlying constraints are different from radar scenarios and have to be taken into account when using the PHD filter. In this article, we propose a new tree-based path extraction algorithm for a Gaussian Mixture PHD filter in Computer Vision applications. We also investigate how an additional benefit can be achieved by using a second human detector and justify an approximation for multiple sensors in low-clutter scenarios.
Volker Eiselein, Daniel Arp, Michael Pätzold, Thomas Sikora
AVSS4
2012 Splitting Gaussians in Mixture Models
abstract
Gaussian mixture models have been extensively used and enhanced in the surveillance domain because of their ability to adaptively describe multimodal distributions in real-time with low memory requirements. Nevertheless, they still often suffer from the problem of converging to poor solutions if the main mode stretches and thus over-dominates weaker distributions. Based on the results of the Split and Merge EM algorithm, in this paper we propose a solution to this problem. Therefore, we define an appropriate splitting operation and the corresponding criterion for the selection of candidate modes, for the case of background subtraction. The proposed method achieves better background models than state-of-the-art approaches and is low demanding in terms of processing time and memory requirements, therefore making it especially appealing in the surveillance domain.
Rubén Heras Evangelio, Michael Pätzold, Thomas Sikora
AVSS3
2012 Boosting Multi-hypothesis Tracking by Means of Instance-Specific Models
abstract
In this paper we present a visual person tracking-by-detection system based on on-line-learned instance-specific information along with the kinematic relation of measurements provided by a generic person-category detector. The proposed system is able to initialize tracks on individual persons and start learning their appearance even in crowded situations and does not require that a person enters the scene separately. For that purpose we integrate the process of learning instance-specific models into a standard MHT-framework. The capability of the system to eliminate detections-to-object association ambiguities occurring from missed detections or false ones is demonstrated by experiments for counting and tracking applications using very long video sequences on challenging outdoor scenarios.
Michael Pätzold, Rubén Heras Evangelio, Thomas Sikora
AVSS3
2012 Clustering Motion for Real-Time Optical Flow Based Tracking
abstract
The selection of regions or sets of points to track is a key task in motion-based video analysis, which has significant performance effects in terms of accuracy and computational efficiency. Computational efficiency is an unavoidable requirement in video surveillance applications. Well established methods, e.g. Good Features to Track, select points to be tracked based on appearance features such as cornerness and therefore neglecting the motion exhibited by the selected points. In this paper, we propose an interest point selection method that takes into account the motion of previously tracked points in order to constrain the number of point trajectories needed. By defining pair-wise temporal affinities between trajectories and representing them in a minimum spanning tree, we achieve a very efficient clustering. The number of trajectories assigned to each motion cluster is adapted by initializing and removing tracked points by means of feed-back. Compared to the KLT tracker, we save up to 65% of the points to track, therefore gaining in efficiency while not scarifying accuracy.
Tobias Senst, Rubén Heras Evangelio, Ivo Keller, Thomas Sikora
AVSS4
2012 Detecting People Carrying Objects Utilizing Lagrangian Dynamics
abstract
The availability of dense motion information in computer vision domain allows for the effective application of Lagrangian techniques that have their origin in fluid flow analysis and dynamical systems theory. A well established technique that has been proven to be useful in image-based crowd analysis are Finite Time Lyapunov Exponents (FTLE). Based on this, we present a method to detect people carrying object and describe a methodology how to apply established flow field methods onto the problem of describing individuals. Further, we reinterpret Lagrangian features in relation to the underlying motion process and show their applicability towards the appearance modeling of pedestrians. This definition allows to increase performance of state-of-the-art methods and is shown to be robust under varying parameter settings and different optical flow extraction approaches.
Tobias Senst, Alexander Kuhn, Holger Theisel, Thomas Sikora
AVSS4
2012 Resampling of multiple camera point cloud data
abstract
Even though the sampling grid of a digital camera is uniform and rectangular, the depth information reconstructed from multi-camera samples is neither uniform nor rectangular due to the rotations and translations of the camera coordinate systems with respect to each other. In order to facilitate further processing, e.g. compression of multi camera depth data, resampling of non-uniform multi-camera depth samples to a uniform rectangular grid is advantageous. In this paper, we introduce and compare two resampling methods using a voxel-based and a triangular-mesh-based approach.
Oliver Belaifa, Robert Skupin, Engin Kurutepe, Thomas Sikora
CCNC4
2012 Quadtree-based temporal trajectory filtering
abstract
In both the HEVC draft and in H.264/AVC, in-loop filters are employed to improve the subjective and the objective quality of compressed video sequences. These filters use spatial information from a single frame only. Temporal Trajectory Filtering (TTF) constitutes an alternative approach which performs filtering in the temporal domain instead. In this work, a combination of the TTF with a quadtree partitioning algorithm for applying different filter parameters to different image regions is proposed and investigated. Experiments were conducted in the environment of the HEVC test model HM 3.0. Bit rate reductions of up to 9% for the low delay high efficiency setting of HEVC are reported.
Marko Esche, Alexander Glantz, Andreas Krutz, Michael Tok, Thomas Sikora
ICIP5
2012 Multi-camera rectification using linearized trifocal tensor
Frederik Zilly, Christian Riechert, Marcus Müller 0001, Wolfgang Waizenegger, Thomas Sikora, Peter Kauff
ICPR5
2012 Performance Evaluation of Feature Detection for Local Optical Flow Tracking
Tobias Senst, Brigitte Unger, Ivo Keller, Thomas Sikora
ICPRAM (2)4
2012 Cross-modal categorisation of user-generated video sequences
abstract
This paper describes the possibilities of cross-modal classification of multimedia documents in social media platforms. Our framework predicts the user-chosen category of consumer-produced video sequences based on their textual and visual features. These text resources---includes metadata and automatic speech recognition transcripts---are represented as bags of words and the video content is represented as a bag of clustered local visual features. The contribution of the different modalities is investigated and how they should be combined if sequences lack certain resources. Therefore, several classification methods are evaluated, varying the resources. The paper shows an approach that achieves a mean average precision of 0.3977 using user-contributed metadata in combination with clustered SURF.
Sebastian Schmiedeke, Pascal Kelm, Thomas Sikora
ICMR3
2012 Human action recognition using Lagrangian descriptors
abstract
Human action recognition requires the description of complex motion patterns in image sequences. In general, these patterns span varying temporal scales. In this context, Lagrangian methods have proven to be valuable for crowd analysis tasks such as crowd segmentation. In this paper, we show that, besides their potential in describing large scale motion patterns, Lagrangian methods are also well suited to model complex individual human activities over variable time intervals. We use Finite Time Lyapunov Exponents and time-normalized arc length measures in a linear SVM classification scheme. We evaluated our method on the Weizmann and KTH datasets. The results demonstrate that our approach is promising and that human action recognition performance is improved by fusing Lagrangian measures.
Esra Acar, Tobias Senst, Alexander Kuhn, Ivo Keller, Holger Theisel, Sahin Albayrak, Thomas Sikora
MMSP7
2012 A Lagrangian framework for video analytics
abstract
The extraction of motion patterns from image sequences based on the optical flow methodology is an important and timely topic among visual multi media applications. In this work we will present a novel framework that combines the optical flow methodology from image processing with methods developed for the Lagrangian analysis of time-dependent vector fields. The Lagrangian approach has been proven to be a valuable and powerful tool to capture the complex dynamic motion behavior within unsteady vector fields. To come up with a compact and applicable framework, this paper will provide concepts on how to compute trajectory-based Lagrangian measures in series of optical flow fields, a set of basic measures to capture the essence of the motion behavior within the image, and a compact hierarchical, feature-based description of the resulting motion features. The resulting framework will bee shown to be suitable for an automated image analysis as well as compact visual analysis of image sequences in its spatio-temporal context. We show its applicability for the task of motion feature description and extraction on different temporal scales, crowd motion analysis, and automated detection of abnormal events within video sequences.
Alexander Kuhn, Tobias Senst, Ivo Keller, Thomas Sikora, Holger Theisel
MMSP4
2012 Weighted temporal long trajectory filtering for video compression
abstract
In the context of the HEVC standardization activity, in-loop filters such as the adaptive loop filter and the deblocking filter are currently under investigation. Both filters work in the spatial domain only, despite the temporal correlation within video sequences. In this work a previously introduced filter, that uses temporal information for deblocking and denoising instead, is integrated into the HEVC test model HM 3.0. It is shown how the filter is to be adapted to work in combination with the adaptive loop filter for the HEVC low-delay profile. In addition, an optimal weighting function for the filtered luma samples based on the qunatization parameter is derived. Bit rate reductions of up to 7.6% are reported for individual sequences.
Marko Esche, Alexander Glantz, Andreas Krutz, Michael Tok, Thomas Sikora
PCS5
2012 Adaptive global motion temporal filtering
abstract
The emerging standardization on high efficiency video coding (HEVC) has brought a huge improvement in terms of coding performance comparing to existing standards and approaches. One tool that provides a significant portion of the coding gain so far is loop filtering. In H.264/AVC, a deblocking filter has been used to reduce blocking artifacts. HEVC added loop filter approaches to this deblocking method that further reduce noise in the decoded video frame. All these techniques work spatially. Besides that, work has been done to further improve the quality of decoded frames by applying a temporal filtering approach. In this work, we propose adaptive global motion temporal filtering (AGMTF) that reduces noise along temporal trajectories. Experimental results show that the coding performance of the current HEVC test model HM 4.0 can be improved by up to 8.8% and 3.7% in average over a large bit rate range using this technique.
Andreas Krutz, Alexander Glantz, Michael Tok, Thomas Sikora
PCS4
2012 Parametric motion vector prediction for hybrid video coding
abstract
Motion compensated prediction still is the main technique for redundancy reduction in modern hybrid video codecs. However, the resulting motion vector fields are highly redundant as well. Thus, motion vector prediction and difference coding are used for compressing. One drawback of all common motion vector prediction techniques is, that they are not able to predict complex motion as rotation and zoom efficiently. We present a novel parametric motion vector predictor (PMVP), based on higher-order motion models to overcome this issue. To transmit the needed motion models, an efficient compression scheme is utilized. This scheme is based on transformation, quantization and difference coding. By incorporating this predictor into the HEVC test model HM 3.2 gains of up to 2.42% are achieved.
Michael Tok, Alexander Glantz, Andreas Krutz, Thomas Sikora
PCS4
2012 Lossy parametric motion model compression for global motion temporal filtering
abstract
It has been shown that techniques using higher-order motion parameters outperform common translational motion compensated prediction for hybrid video coders. A critical issue is the transmission of accurate higher-order motion parameters with as little additional bits as possible to maximize the compression gain of the whole system. For that, we propose a compression scheme for perspective motion models using transformation before quantization and temporal redundancy reduction and integrate this scheme into a video coding environment using adaptive global motion temporal filtering. Experimental results show that using the proposed compression scheme for the perspective motion models, the BD-rate can be improved up to 8.8% in average in the higher bit rate range and up to 7.7% in average in the lower bit rate range compared to the latest version of the HEVC test model HM 4.0.
Michael Tok, Andreas Krutz, Alexander Glantz, Thomas Sikora
PCS4
2012 Adaptive Temporal Trajectory Filtering for Video Compression
abstract
Most in-loop filters currently being employed in video compression algorithms use spatial information from a single frame of the video sequence only. In this paper, a new filter is introduced and investigated that combines both spatial and temporal information to provide subjective and objective quality improvement. The filter only requires a small overhead on slice level while using the temporal information conveyed in the bit stream to reconstruct the individual motion trajectory of every pixel in a frame at both encoder and decoder. This information is then used to perform pixel-wise adaptive motion-compensated temporal filtering. It is shown that the filter performs better than the state-of-the-art codec H.264/AVC over a large range of sequences and bit rates. Additionally, the filter is compared with another, Wiener-based in-loop filtering approach and a complexity analysis of both algorithms is conducted.
Marko Esche, Alexander Glantz, Andreas Krutz, Thomas Sikora
IEEE Trans. Circuits Syst. Video Technol.4
2012 Adaptive Global Motion Temporal Filtering for High Efficiency Video Coding
abstract
Coding artifacts in video codecs can be reduced using several spatial in-loop filters that are part of the emerging video coding standard High Efficiency Video Coding (HEVC). In this paper, we introduce the concept of global motion temporal filtering. A theoretical framework for a concept combining the temporal overlapping of several noisy versions of the same signal is introduced. This includes a model of the motion estimation error. As an important result, it is shown that an optimum number of framesNfor filtering exists. An implementation of the concept based on several versions of the HEVC test model using global motion-compensated temporal filtering shows that significant gains can be achieved.
Andreas Krutz, Alexander Glantz, Michael Tok, Marko Esche, Thomas Sikora
IEEE Trans. Circuits Syst. Video Technol.5
2012 Robust Local Optical Flow for Feature Tracking
abstract
This paper is motivated by the problem of local motion estimation via robust regression with linear models. In order to increase the robustness of the motion estimates, we propose a novel robust local optical flow approach based on a modified Hampel estimator. We show the deficiencies of the least squares estimator used by the standard Kanade-Lucas-Tomasi (KLT) tracker when the assumptions made by Lucas-Kanade are violated. We propose a strategy to adapt the window sizes to cope with the generalized aperture problem. Finally, we evaluate our method on the Middlebury and MIT dataset and show that the algorithm provides excellent feature tracking performance with only slightly increased computational complexity compared to KLT. To facilitate further development, the presented algorithm can be downloaded from http://www.nue.tu-berlin.de/menue/forschung/projekte/rlof.
Tobias Senst, Volker Eiselein, Thomas Sikora
IEEE Trans. Circuits Syst. Video Technol.3
2011 Complementary background models for the detection of static and moving objects in crowded environments
abstract
In this paper we propose the use of complementary background models for the detection of static and moving objects in crowded video sequences. One model is devoted to accurately detect motion, while the other aims to achieve a representation of the empty scene. The differences in foreground detection of the complementary models are used to identify new static regions. A subsequent analysis of the detected regions is used to ascertain if an object was placed in or removed from the scene. Static objects are prevented from being incorporated into the empty scene model. Removed objects are rapidly dropped from both models. In this way, we build a very precise model of the empty scene and improve the foreground segmentation results of a single background model. The system was validated with several public datasets, showing many advantages over state-of-the-art static objects and foreground detectors.
Rubén Heras Evangelio, Thomas Sikora
AVSS2
2011 Real-time person counting by propagating networks flows
abstract
In this paper we present a system that tracks multiple persons by detection in real-time. We introduce a measure for similarity of detections which segments significant information from background clutter by using statistical information obtained during the learning phase of the detector. In order to track multiple persons we map the detections into flow networks utilizing this measure. A continuous real-time processing of video streams is accomplished by analyzing only small chunks of detections consecutively using different networks. By propagating the result of one network into the subsequent one a temporal consistent association is achieved. The system was evaluated using a standard video sequence containing a crowded scene and an own dataset with very long sequences. The results demonstrate that the system performs comparable to other systems while meeting real-time requirements.
Michael Pätzold, Thomas Sikora
AVSS2
2011 On building decentralized wide-area surveillance networks based on ONVIF
abstract
In this paper we present a decentralized surveillance network composed of IP video cameras, analysis devices and a central node which collects information and displays it in a 3D model of the complete area. The exchange of information between all components in the surveillance network takes place according to the ONVIF specification, therefore ensuring interoperability between products complying with the specification and flexibility regarding the integration of new devices and services. The collected information is displayed in a 3D model of the surveilled area, therefore providing a comfortable overview of the activity in large environments and offering the user an intuitive way to eventually interact with network devices.
Tobias Senst, Michael Pätzold, Rubén Heras Evangelio, Volker Eiselein, Ivo Keller, Thomas Sikora
AVSS6
2011 Feature-based global motion estimation using the Helmholtz principle
abstract
Global motion estimation is an important task for various video processing techniques. The estimation itself has to be robust in presence of arbitrarily moving foreground objects. For that task, two different kinds of estimation methods exist. On the one hand, pixel-based approaches deliver more precise results and work more robust on video sequences with foreground objects. On the other hand, when working on encoded video streams, block-based methods can be used for a much faster but often less precise estimation. We propose a two step estimation method based on the determination and tracking of feature points of video frames and robust motion model estimation using the Helmholtz principle. Therefore, good trackable features are detected and tracked in video sequences. Subsequently, a perspective motion model is derived from the resulting correspondencies by removing feature pairs not belonging to global motion.
Michael Tok, Alexander Glantz, Andreas Krutz, Thomas Sikora
ICASSP4
2011 Temporal trajectory filtering for bi-directional predicted frames
abstract
In this work the application of a temporal in-loop filtering approach for B-frames in video compression based on the Temporal Trajectory Filter (TTF) is investigated. The TTF constructs temporal pixel trajectories for individual image points in the P-frames of a video sequence, which can be utilized to improve the quality of the reconstructed frames used for prediction. It is shown, how this concept can be adapted to B-frames despite the fact that these already use temporal motion information to a great extent through the flexible choice of reference frames and prediction modes. The proposed filter has been integrated into the H.264/AVC encoder using the extended profile with hierarchical B-frames and was tested on a wide range of sequences. The filter produces bit rate reductions of up to -4% with an average of -1.6% over all tested sequences while also improving the subjective quality of the decoded video.
Marko Esche, Andreas Krutz, Alexander Glantz, Thomas Sikora
ICIP4
2011 A block-adaptive skip mode for inter prediction basedonparametric motion models
abstract
Motion compensated prediction (MCP) in hybrid video coding estimates a translational motion vector for a given block which is then used for residual computation. However, when complex motion like zoom, rotation, and perspective transformation occur, the translational model assumption does not always hold. This may result in higher residual energy and splitting of blocks, respectively. This paper proposes a skip mode based on higher-order parametric motion models. Often, these models provide a better prediction quality resulting in lower residual energy and larger block sizes. The proposed technique estimates a higher-order motion model between two given pictures. The encoder decides in terms of rate-distortion optimization whether to use the new skip mode for a block and therefore not to transfer any additional information like coefficient data. Experimental evaluation shows that the proposed technique can improve the coding performance of next generation video coding standards significantly.
Alexander Glantz, Michael Tok, Andreas Krutz, Thomas Sikora
ICIP4
2011 Theoretical consideration of global motion temporal filtering
abstract
A widely used technique to reduce the noise variance of a signal is a temporal overlapping of several noisy versions of it. It will be shown that the same idea can be applied for video sequences. Several versions of the current frame can be aligned using motion compensation that adjacent frames represent a noisy version of the current frame. In a first theoretical calculation of this concept combining the temporal overlapping of several noisy versions of the same signal and a rate-distortion equation, it has been shown that a theoretical bit rate reduction of 1/2 log2(N) is possible. In this work, the concept will be advanced to be closer to practice by adding a model for the motion estimation error. It will be shown that the derived theoretical equation confirms the practice and models the behavior of a video encoding environment using parametric motion compensated temporal filtering very well.
Andreas Krutz, Alexander Glantz, Thomas Sikora
ICIP3
2011 Efficient real-time local optical flow estimation by means of integral projections
abstract
In this paper we present an approach for the efficient computation of optical flow fields in real-time and provide implementation details. Proposing a modification of the popular Lucas-Kanade energy functional based on integral projections allows us to speed up the method notably. We show the potential of this method which can compute dense flow fields of 640×480 pixels at a speed of 4 fps in a GPU implementation based on the OpenCL framework. Working on sparse optical flow fields of up to 17,000 points, we reach execution times of 70 fps. Optical flow methods are used in many different areas, the proposed method speeds up current surveillance algorithms used for scene description and crowd analysis or Augmented Reality and robot navigation applications.
Tobias Senst, Volker Eiselein, Michael Pätzold, Thomas Sikora
ICIP4
2011 Short-term motion-based object segmentation
abstract
Motion-based segmentation approaches employ either longterm motion information, which is computationally demanding, or suffer from lack of accuracy when employing short-term information. We present an automatic motion-based object segmentation algorithm for video sequences with moving camera, employing short-term motion information solely. For every frame, two error frames are generated using motion compensation. They are combined and a thresholding segmentation algorithm is applied. Recent advances in the field of global motion estimation enable outlier elimination in the background area, and thus a more precise definition of the foreground is achieved. We propose a simple and effective error frame generation and consider spatial error localization. Thus, we achieve improved performance compared with a previously proposed short-term motion-based method and provide subjective as well as objective evaluation.
Marina Georgia Arvanitidou, Michael Tok, Andreas Krutz, Thomas Sikora
ICME4
2011 Multi-modal, multi-resource methods for placing Flickr videos on the map
abstract
We present three approaches for placing videos in Flickr on the world map. The toponym extraction and geo lookup approach makes use of external resources to identify toponyms in the metadata and associate them with geo-coordinates. The metadata-based region model approach uses a k-nearest-neighbour classifier trained over geographical regions. Videos are represented using their metadata in a text space with reduced dimensionality. The visual region model approach uses a support vector machine also trained over geographical regions. Videos are represented using low-level feature vectors from multiple key frames. Voting methods are used to form a single decision for each video. We compare the approaches experimentally, highlighting the importance of using appropriate metadata features and suitable regions as the basis of the region model. The best performance is achieved by the geo-lookup approach used with fallback to the visual region model when the video metadata contains no toponym.
Pascal Kelm, Sebastian Schmiedeke, Thomas Sikora
ICMR3
2011 Detection of static objects for the task of video surveillance
abstract
Detecting static objects in video sequences has a high relevance in many surveillance scenarios like airports and railwaystations. In this paper we propose a system for the detection of static objects in crowded scenes that, based on the detection of two background models learning at different rates, classifies pixels with the help of a finite-state machine. The background is modelled by two mixtures of Gaussians with identical parameters except for the learning rate. The state machine provides the meaning for the interpretation of the results obtained from background subtraction and can be used to incorporate additional information cues, obtaining thus a flexible system specially suitable for real-life applications. The system was built in our surveillance application and successfully validated with several public datasets.
Rubén Heras Evangelio, Tobias Senst, Thomas Sikora
WACV3
2011 Robust modified L2 local optical flow estimation and feature tracking
abstract
This paper describes a robust method for the local optical flow estimation and the KLT feature tracking performed on the GPU. Therefore we present an estimator based on the L2norm with robust characteristics. In order to increase the robustness at discontinuities we propose a strategy to adapt the used region size. The GPU implementation of our approach achieves real-time (>;25 fps) performance for High Definition (HD) video sequences while tracking several thousands of points. The benefit of the suggested enhancement is illustrated on the Middlebury optical flow benchmark.
Tobias Senst, Volker Eiselein, Rubén Heras Evangelio, Thomas Sikora
WACV4
2011 Detecting people carrying objects based on an optical flow motion model
abstract
Detecting people carrying objects is a commonly formulated problem as a first step to monitor interactions between people and objects. Recent work relies on a precise foreground object segmentation, which is often difficult to achieve in video surveillance sequences due to a bad contrast of the foreground objects with the scene background, abrupt changing light conditions and small camera vibrations. In order to cope with these difficulties we propose an approach based on motion statistics. Therefore we use a Gaussian mixture motion model (GMMM) and, based on that model, we define a novel speed and direction independent motion descriptor in order to detect carried baggage as those regions not fitting in the motion description model of an average walking person. The system was tested with the public dataset PETS2006 and a more challenging dataset including abrupt lighting changes and bad color contrast and compared with existing systems, showing very promising results.
Tobias Senst, Rubén Heras Evangelio, Thomas Sikora
WACV3
2010 Counting People in Crowded Environments by Fusion of Shape and Motion Information
abstract
Knowing the number of people in a crowded scene is of big interest in the surveillance scene. In the past, this problem has been tackled mostly in an indirect, statistical way. This paper presents a direct, counting by detection, method based on fusing spatial information received from an adapted Histogram of Oriented Gradients-algorithm (HOG) with temporal information by exploiting distinctive motion characteristics of different human body parts. For that purpose, this paper defines a measure for uniformity of motion. Furthermore, the system performance is enhanced by validating the resulting human hypotheses by tracking and applying a coherent motion detection. The approach is illustrated with an experimental evaluation.
Michael Pätzold, Rubén Heras Evangelio, Thomas Sikora
AVSS3
2010 Global motion temporal filtering for in-loop deblocking
abstract
One of the most severe problems in hybrid video coding is its block-based approach, which leads to distortions called blocking artifacts. These artifacts affect not only the subjective perception at the receiver but also the motion compensated prediction (MCP) that generates a prediction signal from previously decoded pictures. It is therefore directly connected to the amount of data that has to be transmitted. In this paper, we propose a technique called global motion temporal filtering for blocking artifact reduction. Other than common deblocking techniques, this approach does not reduce the blocking artifacts spatially. Filtering is performed temporally using a set of neighboring pictures from the picture buffer. This approach is incorporated into an H.264/AVC reference software. Experimental evaluation shows that the proposed technique significantly improves the quality in terms of rate-distortion performance.
Alexander Glantz, Andreas Krutz, Thomas Sikora
ICIP3
2010 Robust global motion estimation using motion vectors of variable size blocks and automatic motion model selection
abstract
A new approach for highly robust and precise global motion estimation (GME) using motion vectors (MVs) is presented. We show that this approach obtains precise higher-order short-term motion parameters for global motion using motion vectors solely. The approach is general and works for different mathematical methods including least-squares and Newton-Raphson method. We show that the approach is suitable for fixed block sizes from plain full-search block-matching as well as for arbitrary block sizes from video streams compressed with H.264/AVC reference encoder. The proposed approach is compared against four other known MV-based GME (MV-GME) methods. Our results show that the approach is significantly more robust and obtains higher precision for global motion parameters in terms of background motion compensation, especially if moving objects occur. In addition, the results are as good as results from precise pixel-based GME methods or even better while the presented MV-GME methods have very low computational costs.
Martin Haller, Andreas Krutz, Thomas Sikora
ICIP3
2010 Compressed domain global motion estimation using the Helmholtz Tradeoff Estimator
abstract
Several algorithms for global motion estimation in video sequences using pixel- or block-based approaches have been published. Most known pixel-based methods lack in performance while when using block-based algorithms working on motion vectors, robustness to outliers and accuracy is missing. In this paper we present the fundamentals of a significantly improved, robust block-based method for global motion estimation in compressed domain following the generic Helmholtz principle. To this aim, we use motion vector fields as provided by MPEG data streams. Background PSNR values for four motion compensated test sequences show that our new method delivers results comparable to more complex algorithms.
Michael Tok, Alexander Glantz, Marina Georgia Arvanitidou, Andreas Krutz, Thomas Sikora
ICIP5
2010 Background modeling for video coding: From sprites to Global Motion Temporal filtering
abstract
Techniques for modeling the background of a video sequence can be useful in alternative video coding approaches. Sprite coding has been evolved to provide high quality decoded video frames after transmission by a reduced amount of bits. However, it has also been shown that this works only for a certain kind of video sequences. It also meets quality limits due to the Sprite generation step. To tackle this problem, multiple Sprites have been proposed. Considering the improved quality coming with multiple Sprites, a method has been developed which provides an optimal quality, i.e. local background Sprite generation. This method can be used to conduct Global Motion Temporal Filtering (GMTF) of distorted video frames. In a first application, GMTF is applied as a post-processing deblocking filter in a video coding environment, which is experimentally evaluated.
Andreas Krutz, Alexander Glantz, Thomas Sikora
ISCAS3
2010 Subjective evaluation of scalable video coding for content distribution
abstract
This paper investigates the influence of the combination of the scalability parameters in scalable video coding (SVC) schemes on the subjective visual quality. We aim at providing guidelines for an adaptation strategy of SVC that can select the optimal scalability options for resource-constrained networks. Extensive subjective tests are conducted by using two different scalable video codecs and high definition contents. The results are analyzed with respect to five dimensions, namely, codec, content, spatial resolution, temporal resolution, and frame quality.
Jong-Seok Lee, Francesca De Simone, Naeem Ramzan, Zhijie Zhao, Engin Kurutepe, Thomas Sikora, Jörn Ostermann, Ebroul Izquierdo, Touradj Ebrahimi
ACM Multimedia6
2010 A novel inloop filter for video-compression based on temporal pixel trajectories
abstract
The objective of this work is to investigate the performance of a new inloop filter for video compression, which uses temporal rather than spatial information to improve the quality of reference frames used for prediction. The new filter has been integrated into the H.264/AVC baseline encoder and tested on a wide range of sequences. Experimental results show that the filter achieves a bit rate reduction of up to 12% and more than 4% on average without increasing the complexity of either encoder or decoder significantly.
Marko Esche, Andreas Krutz, Alexander Glantz, Thomas Sikora
PCS4
2010 Adaptive global motion temporal prediction for video coding
abstract
Depending on the content of a video sequence and the settings used for encoding it, the amount of bits spent for the transmission of motion vector information can be enormous and in some cases even take the largest fraction of the bit rate. This is not always necessary since often wide areas, i.e. background or large foreground regions, fit the same global motion. Additionally, a global motion model using sophisticated interpolation techniques can be a better representation of movement in these regions than a motion vector that has only quarter-pel accuracy. This is true especially if scaling, rotation or perspective transformation occur. This paper presents a novel prediction technique that is based on global motion compensation and temporal filtering of previously decoded pictures. The new approach is incorporated into an H.264/AVC reference software. The new encoder outperforms the reference by up to 14%.
Alexander Glantz, Andreas Krutz, Thomas Sikora
PCS3
2010 Recent advances in video coding using static background models
abstract
Sprite coding, as standardized in MPEG-4 Visual, can result in superior performance compared to common hybrid video codecs both objectively and subjectively. However, state-of-the-art video coding standard H.264/AVC clearly outperforms MPEG-4 Visual sprite coding in broad bit rate ranges. Based on the sprite coding idea, this paper proposes a video coding technique that merges the advantages of H.264/AVC and sprite coding. For that, sophisticated algorithms for global motion estimation, sprite generation and object segmentation - all needed for thorough sprite coding - are incorporated into an H.264/AVC coding environment. The proposed approach outperforms H.264/AVC especially in lower bit rate ranges. Savings up to 21% can be achieved.
Andreas Krutz, Alexander Glantz, Thomas Sikora
PCS3
2010 Automatic MPEG-4 sprite coding - Comparison of integrated object segmentation algorithms
Alexander Glantz, Andreas Krutz, Thomas Sikora, Paulo J. L. Nunes, Fernando Pereira 0001
Multim. Tools Appl.3
2010 Dynamic Spectral Envelope Modeling for Timbre Analysis of Musical Instrument Sounds
abstract
We present a computational model of musical instrument sounds that focuses on capturing the dynamic behavior of the spectral envelope. A set of spectro-temporal envelopes belonging to different notes of each instrument are extracted by means of sinusoidal modeling and subsequent frequency interpolation, before being subjected to principal component analysis. The prototypical evolution of the envelopes in the obtained reduced-dimensional space is modeled as a nonstationary Gaussian Process. This results in a compact representation in the form of a set of prototype curves in feature space, or equivalently of prototype spectro-temporal envelopes in the time-frequency domain. Finally, the obtained models are successfully evaluated in the context of two music content analysis tasks: classification of instrument samples and detection of instruments in monaural polyphonic mixtures.
Juan José Burred, Axel Röbel, Thomas Sikora
IEEE Trans. Speech Audio Process.3
2009 Polyphonic musical instrument recognition based on a dynamic model of the spectral envelope
abstract
We propose a new method for detecting the musical instruments that are present in single-channel mixtures. Such a task is of interest for audio and multimedia content analysis and indexing applications. The approach is based on grouping sinusoidal trajectories according to common onsets, and comparing each group's overall amplitude evolution with a set of pre-trained probabilistic templates describing the temporal evolution of the spectral envelopes of a given set of instruments. Classification is based on either an Euclidean or a probabilistic definition of timbral similarity, both of which are compared with respect to detection accuracy.
Juan José Burred, Axel Röbel, Thomas Sikora
ICASSP3
2009 Motion-based object segmentation using local background sprites
abstract
It is well known that video material with a static background allows easier segmentation than that with a moving background. One approach to segmentation of sequences with a moving background is to use preprocessing to create a static background, after which conventional background subtraction techniques can be used for segmenting foreground objects. It has been recently shown that global motion estimation and/or background sprite generation techniques are reliable. We propose a new background modeling technique for object segmentation using local background sprite generation. Experimental results show the excellent performance of this new method compared to recent algorithms proposed.
Andreas Krutz, Alexander Glantz, Thilo Borgmann, Michael R. Frater, Thomas Sikora
ICASSP5
2009 Incorporating prior knowledge on the digital media creation process into audio classifiers
abstract
In the process of music content creation, a wide range of typical audio effects such as reverberation, equalization or dynamic compression are very commonly used. Despite the fact that such effects have a clear impact on the audio features, they are rarely taken into account when building an automatic audio classifier. In this paper, it is shown that the incorporation of prior knowledge of the digital media creation chain can clearly improve the robustness of the audio classifiers, which is demonstrated on a task of musical instrument recognition. The proposed system is based on a robust feature selection strategy, on a novel use of the virtual support vector machines technique and a specific equalization used to normalize the signals to be classified. The robustness of the proposed system is experimentally evidenced using a rather large and varied sound database.
Maxime Lardeur, Slim Essid, Gaël Richard, Martin Haller, Thomas Sikora
ICASSP5
2009 Video coding using global motion temporal filtering
abstract
Recent deblocking techniques are based on spatial filtering. We present a new deblocking technique based on temporal filtering of spatially aligned frames. This approach is used in an H.264/AVC coding environment. The algorithm estimates the ideal amount of frames used for temporal filtering at the encoder side. In that way it is assured that the receiver is presented with the best possible visual quality in terms of structural similarity. Theoretical consideration of the problem proves the concept of the new approach. Experimental evaluation shows that the new temporal deblocking filter significantly improves visual quality and reduces bit rate compared to common H.264/AVC deblocking by up to 18%.
Alexander Glantz, Andreas Krutz, Martin Haller, Thomas Sikora
ICIP4
2009 Rate-distortion optimization for automatic sprite video coding using H.264/AVC
abstract
Sprite-based video coding offers higher compression efficiency than conventional block-based hybrid video coders. In sprite coding a sequence is divided into a model of its background, i.e. a so-called background sprite image, and a foreground object sequence. These are then encoded and transmitted to the receiver. At the decoder the output video sequence is synthesized using content previously transmitted. The background information of the video sequence is, other than the foreground object sequence, encoded indirectly, meaning for a sequence with N frames one single image is encoded and not N background frames. The background for a single frame is reconstructed by coordinate transformation at the decoder. An important issue to address here is an appropriate rate-distortion technique for optimization of the sprite-based coding approach, since foreground and background are encoded independently. In this paper, the Lagrangian cost function is considered for rate-constrained encoder control. Experimental evaluation shows the correct use of the optimization method and the gain in terms of rate-distortion performance compared to H.264/AVC.
Andreas Krutz, Alexander Glantz, Michael R. Frater, Thomas Sikora
ICIP4
2009 Multimodal person search combining information fusion and relevance feedback
abstract
With the increasing amount of multimedia data, efficient tools for search and retrieval are needed. Since people are naturally one of the most interesting objects within these documents, a system for multimodal person search and retrieval has been developed. It combines the audiovisual analysis of persons with the query by example paradigm and relevance feedback to provide an efficient tool for searching multimedia data. For the relevance feedback, one and two class approaches are considered and compared to each other. Multimodal fusion techniques are used to exploit the complementary character of the audio and video information. The experimental results prove that multimodal person search and retrieval is feasible and more efficient than manual exploration.
Lutz Goldmann, Amjad Samour, Touradj Ebrahimi, Thomas Sikora
MMSP4
2009 Automating multi-camera self-calibration
abstract
We demonstrate a convenient and accurate method for fully automatic camera calibration. The method needs at least two cameras and one projector to function, but the cameras need not to be synchronized. By projecting a predefined black and white sequence into the cameras' field of view a large number of individual points are tagged by a binary bit sequence over time. This solves the correspondence problem among the adjacent views and furthermore allows for Forward Error Correction (FEC) yielding a dense error free and subpixel accurate point cloud which is used for internal and external camera calibration. Experimental results and comparison with varying permutations of the projection sequence are given at the end of this paper. Finally, the method gives instant feedback to the user as the resulting calibration point cloud is in fact a 3D scan of the arbitrary calibration scene which can be easily visualized.
Kai Ide, Steffen Siering, Thomas Sikora
WACV3
2008 Towards Fully Automatic Image Segmentation Evaluation
Lutz Goldmann, Tomasz Adamek, Peter Vajda, Mustafa Karaman, Roland Mörzinger, Eric Galmar, Thomas Sikora, Noel E. O'Connor, Thien Ha-Minh, Touradj Ebrahimi, Peter Schallauer, Benoit Huet
ACIVS7
2008 More robust face recognition by considering occlusion information
abstract
This paper addresses one of the main challenges of face recognition (FR): facial occlusions. Currently, the human brain is the most robust known FR approach towards partially occluded faces. Nevertheless, it is still not clear if humans recognize faces using a holistic or a component-based strategy, or even a combination of both. In this paper, three different approaches based on principal component analysis (PCA) are analyzed. The first one, a holistic approach, is the well-known eigenface approach. The second one, a component-based method, is a variation of the eigenfeatures approach, and finally, the third one, a near-holistic method, is an extension of the lophoscopic principal component analysis (LPCA). So the main contributions of this paper are: The three different strategies are compared and analyzed for identifying partially occluded faces and furthermore it explores how a priori knowledge about present occlusions can be used to improve the recognition performance.
Antonio Rama, Francesc Tarres, Lutz Goldmann, Thomas Sikora
FG4
2008 Extending H.264/AVC with a background sprite prediction mode
abstract
The latest standardized hybrid video codec, H.264/AVC, significantly outperforms earlier video coding standards. Despite combining improved and new algorithms within this codec, it is still possible to find methods which lead to a higher coding efficiency. We tackle the prediction problem adding a new prediction mode to the codec. It has been shown that the generation of a background sprite image containing all the background information of a certain sequence is very useful e.g. for object-based video coding. We use a pre-generated background sprite image for creating a new prediction mode in the encoder loop. For the current frame to be compensated, blocks reconstructed from the background sprite are used beside the remaining modes to calculate the residual. The rate-distortion optimization decides which mode is taken. Experimental results show the improvement using the new sprite prediction (SP) mode with the considered test sequences.
Matthias Kunter, Philipp Krey, Andreas Krutz, Thomas Sikora
ICIP4
2008 Multi-view video streaming over P2P networks with low start-up delay
abstract
We propose to stream multi-view video over a multi-tree peer- to-peer (P2P) network using the NUEPMuT protocol. Each view of the multi-view video is streamed over an independent P2P streaming tree and each peer only contributes upload capacity in a single tree, in order to limit the adverse effects of ungraceful peer departures. Additionally, we introduce a quick join procedure to reduce the start-up delay for the first data packet after a join request. Continuity index and decoded video quality performance for simulcast and MVC encoding in a large topology under different settings are reported, in addition to the improvements achieved by the quick join procedure.
Engin Kurutepe, Thomas Sikora
ICIP2
2008 Noise filtering method for color images based on LDA and nonlinear diffusion
abstract
The purpose of noise filtering for images is to preserve features such as edge or corners in images, while reducing noise. Recent noise filtering algorithms based on diffusion equation shows the satisfactory results to some extent, if the noise is additive Gaussian noise. However, if the noise is not additive Gaussian noise, the filtering result is not satisfactory. In this paper, we propose a noise filtering method for color images based on LDA and nonlinear diffusion, which makes use of a common diffusion control. Experimental results with images degraded by additive Gaussian noise, salt and pepper noise, and multiplicative noise are presented.
Woong Hee Kim, Thomas Sikora
ICME2
2008 Recent developments in panoramic image generation and sprite coding
abstract
The composition of panoramic images has recently received considerable attention. While panoramic images were first used mainly as a flexible visualization technique, they also found application in video coding, video enhancement, format conversion, and content analysis. The topic has enlarged and diverged into many specialized research directions, which makes it difficult to stay in touch with recent developments. This paper intends to give an overview of the current state of research, including recent developments. Two of the applications of sprite coding and global-motion estimation are presented in more detail to provide some insights into the system aspects.
Dirk Farin, Martin Haller, Andreas Krutz, Thomas Sikora
MMSP4
2008 On the detection and localization of facial occlusions and its use within different scenarios
abstract
Face analysis is a very active research field, due to its large variety of applications and the different challenges (illumination, pose, expressions or occlusions) the methods need to cope with. Facial occlusions are one of the biggest challenges since they are difficult to model and have a large influence on the performance of subsequent analysis modules. This paper describes a face detection/classification module that allows to detect and localize faces and present occlusions and discusses the use of this additional information within different application scenarios. The approach is evaluated on two databases with realistic occlusions and performs very well for the different detection/classification tasks. It achieves a f-measure of over 97% for face detection and around 86% for component detection. Regarding the occlusion detection, the proposed approach reaches a recognition rate above 91% for both faces and components.
Lutz Goldmann, Antonio Rama, Thomas Sikora, Francesc Tarres
MMSP3
2008 Camera motion-constraint video codec selection
abstract
In recent years advanced video codecs have been developed, such as standardized in MPEG-4. The latest video codec H.264/AVC provides compression performance superior to previous standards, but is based on the same basic motion-compensated-DCT architecture. However, for certain types of video, it has been shown that it is possible to outperform the H.264/AVC using an object-based video codec. Towards a general-purpose object-based video coding system we present an automated approach to separate a video sequences into sub-sequences regarding its camera motion type. Then, the sub-sequences are coded either with an object-based codec or the common H.264/AVC. Applying different video codecs for different kinds of camera motion, we achieve a higher overall coding gain for the video sequence. In first experimental evaluations, we demonstrate the excellence performance of this approach on two test sequences.
Andreas Krutz, Sebastian Knorr, Matthias Kunter, Thomas Sikora
MMSP4
2008 Multimedia Retrieval and Delivery: Essential Metadata Challenges and Standards
abstract
Multimedia information retrieval (MIR) and delivery plays an important role in many application domains due to the increasing need to identify, filter, and manage growing amounts of data, notably multimedia information. To efficiently manage and exchange multimedia information, interoperability between coded data and metadata is required and standardization is central to achieving the necessary level of interoperability. In the context of this paper, the term retrieval refers to the process by which a user, human or machine, identifies the content it needs, and the term delivery refers to the adaptive transport and consumption of the identified content in a particular context or usage environment. Both the retrieval and delivery processes may require content and context metadata. This paper will argue that maximum quality of experience depends not only on the content itself (and thus content metadata) but also on the consumption conditions (thus context metadata). Additionally, the rights and protection conditions have become critically important in recent years, especially with the explosion of electronic music commerce and different ldquoshoppingrdquo conditions. This paper will review existing multimedia standards related to information retrieval and adaptive delivery of multimedia content, emphasizing the need for such standards, and will show how these standards can help the development, dissemination, and valorization of MIR research results. Moreover, it will also discuss limitations of the current standards and anticipate what future standardization activities are relevant and needed. Due to space limitations, the paper will mainly concentrate on MPEG standards although many other relevant standards are also reviewed and discussed.
Fernando Pereira 0001, Anthony Vetro, Thomas Sikora
Proc. IEEE3
2008 Stereoscopic 3D from 2D video with super-resolution capability
Sebastian Knorr, Matthias Kunter, Thomas Sikora
Signal Process. Image Commun.3
2007 Confocal Disparity Estimation and Recovery of Pinhole Image for Real-Aperture Stereo Camera Systems
abstract
A single dense depth estimation using stereo or defocus cannot produce a reliable result due to the ambiguity problem. In this paper, we propose a novel anisotropic disparity estimation embedding a stereo confocal constraint for real-aperture stereo camera systems. If the focal length of a real-aperture stereo camera is just changed, the depth range is localized in a focused object which can be discriminated from defocused blurring. The focal depth plane is estimated by the displacement of tensors which are derived from generalized 2D Gaussian, since the point spread functions (PSF) in defocused blurring can be approximated by a shift-invariant Gaussian function. We localize the isotropic propagation in blurring over invariance by a sparse Laplacian kernel in Poisson solution. The matching of real-aperture stereo images is performed by observing the focal consistency. However, the isotropic propagation cannot exactly hold a non-parallel surface to the lens plane i.e., unequifocal surface. An anisotropic regularization term is employed to suppress the isotropic propagation near the non-parallel surface boundary. Our method achieves an accurate dense disparity map by sampling the disparities in focal points from multiple defocus stereo images. The pels in focal points are utilized to recover the pinhole image (i.e. an ideally focused image for all different depths).
Jangheon Kim, Thomas Sikora
ICIP (5)2
2007 An Image-Based Rendering (IBR) Approach for Realistic Stereo View Synthesis of TV Broadcast Based on Structure from Motion
abstract
In the past years, the 3D display technology has become a booming branch of research with fast technical progress. Hence, the 3D conversion of already existing 2D video material increases more and more in popularity. In this paper, a new approach for realistic stereo view synthesis (RSVS) of existing 2D video material is presented. The intention of our work is not a real-time conversion of existing video material with a deduction in stereo perception, but rather a more realistic off-line conversion with high accuracy. Our approach is based on structure from motion techniques and uses image-based rendering to reconstruct the desired stereo views for each video frame. The algorithm is tested on several TV broadcast videos, as well as on sequences captured with a single handheld camera. Finally, some simulation results will show the remarkable performance of this approach.
Sebastian Knorr, Thomas Sikora
ICIP (6)2
2007 Window-Based Image Registration using Variable Window Sizes
abstract
We present an unsupervised image registration algorithm to estimate the background object motion in a real video sequence. The algorithm is based on a Gaussian minimisation technique. It has been shown earlier that initialization of such an approach is very important to achieve the motion parameters of the background object precisely, and that the use of a windowing technique can give better background object motion estimation results, even with large background occlusions. In some cases, however, the fixed window size initializes the gradient descent algorithm in a sub-optimal way. Here, another window size would bring the desired estimation direction. In this paper, we present a technique where variable window sizes are used to prevent these outliers. Experimental results show that the technique works very well with the considered test sequences.
Andreas Krutz, Michael R. Frater, Thomas Sikora
ICIP (5)3
2007 Object-Based Multiple Sprite Coding of Unsegmented Videos using H.264/AVC
abstract
In spite of recent progress in the development of hybrid block-based video codecs, it has been shown that for low-bitrate scenarios there is still coding gain applying object-based techniques. We present a sprite-based codec, based on latest H.264 features using an inbuilt segmentation approach for scenes recorded by a rotating camera. The segmentation itself is built up on reliable background estimation from the sprite and short-term image registration. Moreover, we generate multiple sprites based on physical camera parameter estimation that overcome three of the main drawbacks of sprite coding techniques. First, the coding cost for the sprite image is minimized. Second, multiple sprites allow temporal background refresh and finally, registration error accumulation is kept very small. Experimental results show that this coding approach significantly outperforms latest H.264 extensions applying hierarchical B pictures.
Matthias Kunter, Andreas Krutz, Michael Drose, Michael R. Frater, Thomas Sikora
ICIP (1)5
2007 Extensible Platform for Multimedia Analysis (XPMA)
abstract
We present a new software-platform with an open and programming-language-independent structure, to improve the reusability of developed multimedia analysis components. A user can design his own, application-oriented multimedia system by combining these components via XML and evaluate the experimental results.
Ronald Glasberg, Pascal Kelm, Thomas Sikora
ICME4
2007 Optimal multiple sprite generation based on physical camera parameter estimation
abstract
We present a robust and computational low complex method to estimate the physical camera parameters, intrinsic and extrinsic, for scene shots captured by cameras applying pan, tilt, rotation, and zoom. These parameters are then used to split a sequence of frames into several subsequences in an optimal way to generate multiple sprites. Hereby, optimal means a minimal usage of memory while keeping or even improving the reconstruction quality of the scene background. Since wide angles between two frames of a scene shot cause geometrical distortions using a perspective mapping it is necessary to part the shot into several subsequences. In our approach it is not mandatory that all frames of a subsequence are adjacent frames in the original scene. Furthermore the angle-based classification allows frame reordering and makes our approach very powerful.
Matthias Kunter, Andreas Krutz, Mrinal Mandal 0001, Thomas Sikora
VCIP4
2007 Towards 3-D scene reconstruction from broadcast video
Evren Imre, Sebastian Knorr, Burak Özkalayci, Ugur Topay, A. Aydin Alatan, Thomas Sikora
Signal Process. Image Commun.6
2007 Components and Their Topology for Robust Face Detection in the Presence of Partial Occlusions
abstract
This paper presents a novel approach for automatic and robust object detection. It utilizes a component-based approach that combines techniques from both statistical and structural pattern recognition domain. While the component detection relies on Haar-like features and an AdaBoost trained classifier cascade, the topology verification is based on graph matching techniques. The system was applied to face detection and the experiments show its outstanding performance in comparison to conventional face detection approaches. Especially in the presence of partial occlusions, uneven illumination, and out-of-plane rotations, it yields higher robustness. Furthermore, this paper provides a comprehensive review of recent approaches for object detection and gives an overview of available databases for face detection.
Lutz Goldmann, Ullrich J. Mönich, Thomas Sikora
IEEE Trans. Inf. Forensics Secur.3
2006 Extending Single-View Scalable Video Coding to Multi-View Based on H.264/AVC
abstract
An extension of single-view scalable video coding to multi-view is presented in this paper. Scalable video coding is recently developed in the Joint Video Team of ISO/IEC MPEG and ITU-T VCEG named Joint Scalable Video Model. The model includes temporal, spatial and quality scalability enhancing a H.264/AVC base layer. To remove redundancy between views a hierarchical decomposition in a similar way to the temporal direction is applied. The codec is based on this technology and supports open-loop as well as closed-loop controlled encoding. The advantage of this approach lies in its compatibility to the state of the art single-view video codec H.264/AVC and its simple decomposition structure. Encoding a base view using H.264/AVC syntax, any standard single-view decoder is able to decode the data. The hierarchical decomposition structure allows efficient access to all views and frames inside a view. This is especially important for video-based-rendering and multi-view displays, which have different requirements. The chosen decomposition structure also supports parallel processing. Gain in objective as well as subjective quality was achieved for some test sequences using a single layer. The results were compared to JSVM 5.1 (simulcast).
Michael Drose, Carsten Clemens, Thomas Sikora
ICIP3
2006 Extracting High Level Semantics by Means of Speech, Audio, and Image Primitives in Surveillance Applications
abstract
Traditional surveillance systems are usually based on visual information only. With the emerging multimedia analysis techniques, interests are changing towards systems that incorporate multiple sensors and different modalities, which leads to new ways of analyzing this multimedia data and more sophisticated applications. This paper shortly reviews the ideas of traditional surveillance systems and explains actual research interests in this domain. Then, it focuses on the typical structure, goals, and applications of multimedia surveillance systems. These issues are supported by short descriptions of selected analysis steps of such a system currently under development. Some experimental results are given to illustrate the extracted semantics and to assess the performance of the individual steps.
Lutz Goldmann, Amjad Samour, Mustafa Karaman, Thomas Sikora
ICIP4
2006 Prioritized Sequential 3D Reconstruction in Video Sequences with Multiple Motions
abstract
In this study, an algorithm is proposed to solve the multi-frame structure from motion (MFSfM) problem for monocular video sequences in dynamic scenes. The algorithm uses the epipolar criterion to segment the features belonging to independently moving objects. Once the features are segmented, corresponding objects are reconstructed individually by using a sequential algorithm, which is also capable of prioritizing the frame pairs with respect to their reliability and information content, thus achieving a fast and accurate reconstruction through efficient processing of the available data. A tracker is utilized to increase the baseline distance between views and to improve the F-matrix estimation, which is beneficial to both the segmentation and the 3D structure estimation processes. The experimental results demonstrate that our approach has the potential to effectively deal with the multi-body MFSfM problem in a generic video sequence.
Evren Imre, Sebastian Knorr, A. Aydin Alatan, Thomas Sikora
ICIP4
2006 Robust Anisotropic Disparity Estimation with Perceptual Maximum Variation Modeling
abstract
We present a robust anisotropic dense disparity estimation algorithm which employs perceptual maximum variation modeling. Edge-preserving dense disparity vectors are estimated using a coarse-to-fine diffusive method on iteratively filtered images, i.e. the scale-space. While an energy-minimization framework adjusts local disparity, the edges are efficiently preserved by anisotropic disparity-field diffusion. However, the localization at weak image edges which have small brightness variations is fundamentally difficult. In this paper, perceptual maximum variation modeling prevents the delocalization flow over edges, e.g. over-diffusion and back-diffusion, computed by evaluating small variations. We perform disparity-field diffusion on a perceptually optimized color space, which combines the small differences in both brightness and chromaticity. Additionally a consistency constraint is employed in the modeling to avoid the influence of global color distributions and to enhance important edges as the human vision system does. The experimental results show the excellent localization performance preserving the disparity discontinuity of each object.
Jangheon Kim, Thomas Sikora
ICIP2
2006 Windowed Image Registration for Robust Mosaicing of Scenes with Large Background Occlusions
abstract
We propose an enhanced window-based approach to local image registration for robust video mosaicing in scenes with arbitrarily moving foreground objects. Unlike other approaches, we estimate accurately the image transformation without any pre-segmentation even if large background regions are occluded. We apply a windowed hierarchical frame-to-frame registration based on image pyramid decomposition. In the lowest resolution level phase correlation for initial parameter estimation is used while in the next levels robust Newton-based energy minimization of the compensated image mean-squared error is conducted. To overcome the degradation error caused by spatial image interpolation due to the warping process, i.e. aliasing effects from under-sampling, final pixel values are assigned in an up-sampled image domain using a Daubechies bi-orthogonal synthesis filter. Experimental results show the excellent performance of the method compared to recently published methods. The image registration is sufficiently accurate to allow open-loop parameter accumulation for long-term motion estimation.
Andreas Krutz, Michael R. Frater, Matthias Kunter, Thomas Sikora
ICIP4
2006 Recognizing Commercials in Real-Time using Three Visual Descriptors and a Decision-Tree
abstract
We present a new approach for classifying mpeg-2 video sequences as `commercial' or `non-commercial' by analyzing specific color, texture and motion features of consecutive frames in real-time. This is part of the well-known video-genre-classification problem, where popular TV-broadcast genres like cartoon, commercial, music, news and sports are studied. Such applications have also been discussed in the context of MPEG-7. In our method the extracted features from three visual descriptors are logically combined using a decision tree to produce a reliable recognition. The results demonstrate a high identification rate based on a large collection of 200 representative video sequences (40 `commercials' and 4*40 `non-commercials') gathered from free digital TV-broadcasting in Germany
Ronald Glasberg, Cengiz Tas, Thomas Sikora
ICME3
2006 Audiovisual Anchorperson Detection for Topic-Oriented Navigation in Broadcast News
abstract
This paper presents a content-based audiovisual video analysis technique for anchorperson detection in broadcast news. For topic-oriented navigation in newscasts, a segmentation of the topic boundaries is needed. As the anchorperson gives a strong indication for such boundaries, the presented technique automatically determines that high-level information for video indexing from MPEG-2 videos and stores the results in an MPEG-7 conform format. The multimodal analysis process is carried out separately in the auditory and visual modality, and the decision fusion forms the final anchorperson segments
Martin Haller, Hyoung-Gook Kim, Thomas Sikora
ICME3
2006 Improved Image Registration using the Up-Sampled Domain
abstract
We consider the warping problem which appears in well-known image registration algorithms that use higher-order motion models. An implementation inspired by recent work and a new image registration algorithm are used for the analysis. Both approaches rely on frame-to-frame estimation. The key technique is the well-established gradient descent approach for the estimation of higher-order motion parameters. We show that using up-sampled input images in the last step of the algorithms improves the accuracy of the estimated motion parameters. It can be seen in the experimental results that the performance of the image registration algorithms increases significantly only by applying the gradient descent on up-sampled images in comparison to recent algorithms developed
Andreas Krutz, Michael R. Frater, Thomas Sikora
MMSP3
2006 Multi-view synthesis: A novel view creation approach for free viewpoint video
Eddie Cooke, Peter Kauff, Thomas Sikora
Signal Process. Image Commun.3
2005 Distortion Estimation for Temporal Layered Video Coding
abstract
We present a recursive block based decoder distortion estimation model for temporal layered video transmission, based on a DPCM structure. Each block in a video frame is modeled as a sample from an AR(1) source. The correlation coefficient of this source depends on the loop filtering effects, whereas the additional noise term on the motion compensated block difference on the quantization distortion of the block. Distortion estimations are compared to simulation results, and the model is shown to accurately capture the video distortion in various lossy streaming scenarios. The low implementation complexity, and high estimation accuracy of the proposed technique makes it particularly attractive for adaptive video communication applications, that try to optimize the streaming policy.
Sila Ekmekci Flierl, Pascal Frossard, Thomas Sikora
ICASSP (2)3
2005 Hybrid Speaker-Based Segmentation System Using Model-Level Clustering
abstract
We present a hybrid speaker-based segmentation, which combines metric-based and model-based techniques. Without a priori information about the number of speakers and speaker identities, the speech stream is segmented in three stages: (1) the most likely speaker changes are detected; (2) to group segments of identical speakers, a two-level clustering algorithm is performed using a Bayesian information criterion (BIC) and HMM model scores - every cluster is assumed to contain only one speaker; (3) the speaker models are reestimated from each cluster by HMM. Finally a resegmentation step performs a more refined segmentation using these speaker models. To measure the performance, we compare the segmentation results of the proposed hybrid method versus metric-based segmentation. Results show that the hybrid approach using two-level clustering significantly outperforms direct metric-based segmentation.
Hyoung-Gook Kim, Daniel Ertelt, Thomas Sikora
ICASSP (1)3
2005 Hybrid recursive energy-based method for robust optical flow on large motion fields
abstract
We present a new reliable hybrid recursive method for optical flow estimation. The method efficiently combines the advantage of discrete motion estimation and optical flow estimation in a recursive block-to-pixel estimation scheme. Integrated local and global approaches using the robust statistic of anisotropic diffusion remove outliers from the estimated motion field. We separately describe the process with two frameworks i.e. an incremental updating framework and a robust energy minimization framework. With robust error norms of Perona and Marik anisotropic diffusion, the formulation usually leads to non-convex optimization problems. Thus, the solution has many local minima, and convergence to the global minima is not guaranteed. Our hybrid recursive energy-based method employs a hierarchical block-to-pixel estimation concept to prevent this problem. The experimental results prove the excellent performance on several large motion fields.
Jangheon Kim, Thomas Sikora
ICIP (1)2
2005 Comparison of different phone-based spoken document retrieval methods with text and spoken queries
abstract
This study compares four phone-based spoken document retrieval (SDR) approaches. In all cases, the indexing and retrieval system uses phonetic information only. The first retrieval method is based on the vector space model, using phone 3-grams as indexing terms. This approach is compared with 2 string-matching methods. A fourth method, combining the VSM approach with the slot detection step of stringmatching techniques is proposed. This method is tested on a collection of short German spoken documents, using three different sets of queries: text queries, clean spoken queries and noisy spoken queries.
Nicolas Moreau, Thomas Sikora
INTERSPEECH3
2005 Special Issue on Advances in Video Coding and Delivery
Wenwu Zhu 0001, Ming-Ting Sun, Liang-Gee Chen, Thomas Sikora
Proc. IEEE4
2005 Trends and Perspectives in Image and Video Coding
abstract
The objective of the paper is to provide an overview on recent trends and future perspectives in image and video coding. Here, I review the rapid development in the field during the past 40 years and outline current state-of-the art strategies for coding images and videos. These and other coding algorithms are discussed in the context of international JPEG, JPEG 2000, MPEG-1/2/4, and H.261/3/4 standards. Novel techniques targeted at achieving higher compression gains, error robustness, and network/device adaptability are described and discussed.
Thomas Sikora
Proc. IEEE1
2005 Announcement
Thomas Sikora
IEEE Trans. Circuits Syst. Video Technol.1
2004 Comparison of MPEG-7 audio spectrum projection features and MFCC applied to speaker recognition, sound classification and audio segmentation
abstract
We evaluate the MPEG-7 audio spectrum projection (ASP) features for general sound recognition performance against the well established MFCC. The recognition tasks of interest are speaker recognition, sound classification, and segmentation of audio using sound/speaker identification. For sound classification we use three approaches: direct approach; hierarchical approach without hints; hierarchical approach with hints. For audio segmentation, the MPEG-7 ASP features and MFCCs are used to train hidden Markov models (HMM) for individual speakers and sounds. The trained sound/speaker models are then used to segment conversational speech involving a given subset of people in panel discussion television programs. Results show that the MFCC approach yields a sound/speaker recognition rate superior to MPEG-7 implementations.
Hyoung-Gook Kim, Thomas Sikora
ICASSP (5)2
2004 Audio content description with wavelets and neural nets
abstract
Precision audio content description is one of the key components of next generation Internet multimedia search machines. We examine the usability of a combination of 39 different wavelets and three different types of neural nets for precision audio content description. More specifically, we develop a novel wavelet dispersion measure that measures obtained ranks of wavelet coefficients. Our dispersion measure in conjunction with a probabilistic radial basis neural network trained by only three independent example sets obtains a success rate of approximately 78% in identifying unknown complex classical music movements.
Stephan Rein, Martin Reisslein, Thomas Sikora
ICASSP (4)3
2004 Recursive decoder distortion estimation based on AF(1) source modeling for video
abstract
We introduce a recursive block based decoder distortion estimation technique for video and present results showing the accordance of the estimation results with simulation results. Each block in a frame, with all its corresponding blocks along the video sequence, are modeled as an AR(1) source where the correlation coefficient of the source depends on the loop filtering effects and the additional noise term on the motion compensated block difference as well as on the quantization distortion of the block. The distortion term for each block of each frame in the video sequence is calculated recursively depending on the packet loss rate of the channel.
Sila Ekmekci Flierl, Thomas Sikora
ICIP2
2004 A gradient based approach for stereoscopic error concealment
abstract
Error concealment is an important field of research in image processing. Many methods have been applied to conceal block losses in monocular images. We present a concealment strategy for block loss in stereoscopic image pairs. Unlike the error concealment techniques used for monocular images, the information of the associated image is utilized, i.e., by means of a projective transformation model, pixel values from the associated stereo image are warped to their corresponding positions in the lost block. The stereoscopic depth perception is much less affected in our approach than using monoscopic error concealment techniques.
Matthias Kunter, Sebastian Knorr, Carsten Clemens, Thomas Sikora
ICIP4
2004 Speech enhancement based on smoothing of spectral noise floor
Hyoung-Gook Kim, Thomas Sikora
INTERSPEECH2
2004 Phonetic confusion based document expansion for spoken document retrieval
abstract
This paper presents a phone-based approach of spoken document retrieval (SDR), developed in the framework of the emerging MPEG-7 standard. We describe an indexing and retrieval system that uses phonetic information only. The retrieval method is based on the vector space IR model, using phone N-grams as indexing terms. We propose a technique to expand the representation of documents by means of phone confusion probabilities in order to improve the retrieval performance. This method is tested on a collection of short German spoken documents, using 10 city names as queries.
Nicolas Moreau, Hyoung-Gook Kim, Thomas Sikora
INTERSPEECH3
2004 Evaluation of distance measures for MPEG-7 melody contours
abstract
In query by humming (QBH) systems, the melody contour is often used as a symbolic description of music. The MelodyContour description scheme (DS) defined by MPEG-7 is a standardized representation of melody contours. For melody comparison in a QBH system, a distance measure is required. This paper evaluates different distance measures for the MFEG-7 MelodyContour DS. The use of each measure is discussed.
Jan-Mark Batke, Gunnar Eisenberg, Gunnar Weishaupt, Thomas Sikora
MMSP4
2004 Temporal layered vs. multistate video coding
abstract
Multiple Description Video Coding (MDC) and Layered Coding (LC) are both error-resilient source coding techniques used for transmission over error-prone channels. Both techniques generate multiple streams. The streams generated by MDC correspond to different descriptions of the same source whereas the streams produced by LC are differentiated as base and enhancement layer streams. Moreover whereas the MDC streams are independently decodable the decoding of the enhancement layer stream is dependent on the decoding of the base layer stream. In this work we concentrate on specific MDC and LC schemes, i.e. Multi-State Video Coding (MSVC) and Temporal Layered Coding (TLC). MSVC was introduced by John Apostolopoulos and it was showed that if each frame is transmitted in a separate packet and if motion information for each lost frame is also lost, MSVC outperforms Single Layer Coding (SC) in recovering from single as well as burst losses. Here we compared MSVC with TLC as an extension of SC based on transmission simulations over lossy channels under the assumption that the motion vectors are always available. Using different coding modes and specific reconstruction methods average reconstructed frame PSNR (peak signal to noise ratio) is measured and compared. Results show that when motion vectors are received TLC performs better than MSVC for every coding option tested. The performance difference is bigger for low motion sequences.
Sila Ekmekci Flierl, Thomas Sikora
VCIP2
2004 Human body posture recognition using MPEG-7 descriptors
abstract
This paper presents a novel approach to human body posture recognition based on the MPEG-7 contour-based shape descriptor and the widely used projection histogram. A combination of them was used to recognize the main posture and the view of a human based on the binary object mask obtained by the segmentation process. The recognition is treated as a typical pattern recognition task and is carried out through a hierarchy of classifiers. Therefore various structures both hierachical and non-hierarchical, in combination with different classifiers, are compared to each other with respect to recognition performance and computational complexity. Based on this an optimal system design with recognition rates of 95.59% for the main posture, 77.84% for the view and 79.77% in combination is achieved.
Lutz Goldmann, Mustafa Karaman, Thomas Sikora
VCIP3
2004 Fast Index Filtering in Vector Approximation File
Wladyslaw Skarbek, Thomas Sikora, Grzegorz Galinski
Fundam. Informaticae2
2004 Audio classification based on MPEG-7 spectral basis representations
abstract
In this paper, we present an MPEG-7-based audio classification and retrieval technique targeted for analysis of film material. The technique consists of low-level descriptors and high-level description schemes. For low-level descriptors, low-dimensional features such as audio spectrum projection based on audio spectrum basis descriptors is produced in order to find a balanced tradeoff between reducing dimensionality and retaining maximum information content. High-level description schemes are used to describe the modeling of reduced-dimension features, the procedure of audio classification, and retrieval. A classifier based on continuous hidden Markov models is applied. The sound model state path, which is selected according to the maximum-likelihood model, is stored in an MPEG-7 sound database and used as an index for query applications. Various experiments are presented where the speaker- and sound-recognition rates are compared for different feature extraction methods. Using independent component analysis, we achieved better results than normalized audio spectrum envelope and principal component analysis in a speaker recognition system. In audio classification experiments, audio sounds are classified into selected sound classes in real time with an accuracy of 96%.
Hyoung-Gook Kim, Nicolas Moreau, Thomas Sikora
IEEE Trans. Circuits Syst. Video Technol.3
2004 Announcement
Thomas Sikora
IEEE Trans. Circuits Syst. Video Technol.1
2003 Speaker recognition using MPEG-7 descriptors
abstract
Our purpose is to evaluate the efficiency of MPEG-7 audio descriptors for speaker recognition. The upcoming MPEG-7 standard provides audio feature descriptors, which are useful for many applications. One example application is a speaker recognition system, in which reduced-dimension log-spectral features based on MPEG-7 descriptors are used to train hidden Markov models for individual speakers. The feature extraction based on MPEG-7 descriptors consists of three main stages: Normalized Audio Spectrum Envelope (NASE), Principal Component Analysis (PCA) and Independent Component Analysis (ICA). An experimental study is presented where the speaker recognition rates are compared for different feature extraction methods. Using ICA, we achieved better results than NASE and PCA in a speaker recognition system.
Hyoung-Gook Kim, Edgar Berdahl, Nicolas Moreau, Thomas Sikora
INTERSPEECH4
2003 Enhancement of noisy speech for noise robust front-end and speech reconstruction at back-end of DSR system
abstract
This paper presents a speech enhancement method for noise robust front-end and speech reconstruction at the back-end of Distributed Speech Recognition (DSR). The speech noise removal algorithm is based on a two stage noise filtering LSAHT by log spectral amplitude speech estimator (LSA) and harmonic tunneling (HT) prior to feature extraction. The noise reduced features are transmitted with some parameters, viz., pitch period, the number of harmonic peaks from the mobile terminal to the server along noise-robust mel-frequency cepstral coefficients. Speech reconstruction at the back end is achieved by sinusoidal speech representation. Finally, the performance of the system is measured by the segmental signal-noise ratio, MOS tests, and the recognition accuracy of an Automatic Speech Recognition (ASR) in comparison to other noise reduction methods.
Hyoung-Gook Kim, Markus Schwab, Nicolas Moreau, Thomas Sikora
INTERSPEECH4
2003 Model-based unbalanced multiple description video transmission using path diversity
Sila Ekmekci Flierl, Thomas Sikora
VCIP2
2002 Hierarchical image database browsing environment with embedded relevance feedback
abstract
We address the user-navigation through large volumes of image data. A tree structured K-means clustering is introduced which will hierarchically group images into similar groups. Providing the nodes of the different levels with representative image samples leads to different "image maps" similar to street maps with various resolutions of details. The user can zoom into various cluster levels to obtain more or less detail if required. Further a new query refinement method is introduced. The retrieval process is controlled by learning from positive examples from the user, often called the relevance feedback of the user. The combination of the relevance feedback and the hierarchical structure together with a three-dimensional visualization of the "image maps" leads to an intuitive browsing environment. The results obtained verify the attractiveness of the approach for navigation and retrieval applications.
Thomas Meiers, Thomas Sikora, Ivo Keller
ICIP (2)2
2002 Media semantics: who needs it and why?
abstract
Article Share on Media semantics: who needs it and why? Authors: Chitra Dorai Organizers OrganizersView Profile , Andreas Mauthe View Profile , Frank Nack Organizers OrganizersView Profile , Lloyd Rutledge View Profile , Thomas Sikora View Profile , Herbert Zettl View Profile Authors Info & Claims MULTIMEDIA '02: Proceedings of the tenth ACM international conference on MultimediaDecember 2002 Pages 580–583https://doi.org/10.1145/641007.641123Online:01 December 2002Publication History 9citation676DownloadsMetricsTotal Citations9Total Downloads676Last 12 Months8Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Chitra Dorai, Andreas Mauthe, Frank Nack, Lloyd Rutledge, Thomas Sikora, Herbert Zettl
ACM Multimedia5
2001 Visualization and navigation in image database applications based on MPEG-7 descriptors
abstract
Summary form only given, as follows. The MPEG-7 standard is currently being developed to specify standardized interfaces that help users or agents to identify, filter, browse and efficiently retrieve audio-visual material. The purpose of the presentation is to provide an overview of the scope and potential of the MPEG-7 standard, with particular emphasis on no-text descriptors for visual content. Based on the MPEG-7 specifications, the implementation of a wealth of applications is possible. Various application scenarios are presented and discussed.
Thomas Sikora
ICIP (3)1
2001 Introduction to the special issue on MPEG-7
abstract
MPEG-7 has ignited interest in industry with broad interests in media, content management, content provision, broadcasting, computers, telecommunications, as well as researchers and academia. Overall, the MPEG-7 standard is currently in a fairly mature stage and is on its way to final standardization later this year. The pace of progress in development of MPEG-7 has been quite frantic. This special issue provides snapshot of the status of MPEG-7 technology as of around April 2001. This special issue contains only invited papers from experts who are involved in MPEG-7 standardization. It is designed to be useful to beginners as well as experts alike as it covers various topics at different depths.
Shih-Fu Chang, Atul Puri, Thomas Sikora, HongJiang Zhang
IEEE Trans. Circuits Syst. Video Technol.3
2001 Overview of the MPEG-7 standard
abstract
MPEG-7, formally known as the Multimedia Content Description Interface, includes standardized tools (descriptors, description schemes, and language) enabling structural, detailed descriptions of audio-visual information at different granularity levels (region, image, video segment, collection) and in different areas (content description, management, organization, navigation, and user interaction). It aims to support and facilitate a wide range of applications, such as media portals, content broadcasting, and ubiquitous multimedia. We present a high-level overview of the MPEG-7 standard. We first discuss the scope, basic terminology, and potential applications. Next, we discuss the constituent components. Then, we compare the relationship with other standards to highlight its capabilities.
Shih-Fu Chang, Thomas Sikora, Atul Puri
IEEE Trans. Circuits Syst. Video Technol.2
2001 The MPEG-7 visual standard for content description-an overview
abstract
S.696-702
Thomas Sikora
IEEE Trans. Circuits Syst. Video Technol.1
2000 Image visualization based on MPEG-7 color descriptors
Thomas Meiers, H. Czernoch-Peters, L. Ihlenburg, Thomas Sikora
VCIP4
2000 Special Issue on Shape Coding for Emerging Multimedia Applications
Minoru Etoh, Jörn Ostermann, Thomas Sikora
Signal Process. Image Commun.3
2000 A family of single-user autostereoscopic displays with head-tracking capabilities
abstract
We present prototypes of autostereoscopic displays which allow single users to experience stereoscopic vision without the need for special eye glasses or helmet-mounted displays. The design of the displays is based on lenticular raster plates and includes a number of novel concepts for tracking of raster plates or projection lenses to account for changes of the viewers position in front of the screen. Applications envisioned include 3-D multimedia desktop visualization for medical and biological imaging, design, and architecture, as well as computer games and 3-D virtual reality in general. Concepts and results for both high-resolution flat liquid-crystal panel monitors for PC desktop applications as well as large screen high resolution displays using rear-projection technology are discussed.
Reinhard Borner, Bernd Duckstein, Oliver Machui, Hans Roder, Thomas Sinnig, Thomas Sikora
IEEE Trans. Circuits Syst. Video Technol.6
1999 Special Issue On Representation And Coding Of Images And Video II
King Ngi Ngan, Sethuraman Panchanathan, Thomas Sikora, Ming-Ting Sun
IEEE Trans. Circuits Syst. Video Technol.3
1999 Real-time estimation of long-term 3-D motion parameters for SNHC face animation and model-based coding applications
abstract
We present two recursive methods for the real-time estimation of long-term three-dimensional (3-D) motion parameters from monocular image sequences suitable for synthetic/natural hybrid coding face animation and model-based coding applications. Based on feature point extractions in energy frame, the 3-D motion parameters of a human face are estimated with a predictive approach. The first method uses a recursive linear least squares approach and the second employs a nonlinear extended Kalman filter, which does not rely on a linearized model of the face motion. Both methods perform a prediction and correction loop at every time step. Compared to other methods described in the literature, the recursive and predictive structure of the proposed estimation process solves the problem of error accumulation in long-term motion estimation. This makes the estimation stable and consistent over long periods. Experimental results are presented for synthetic data and real image sequences, which demonstrate the performance of the estimation methods and compare the two approaches.
Aljoscha Smolic, Bela Makai, Thomas Sikora
IEEE Trans. Circuits Syst. Video Technol.3
1999 Long-term global motion estimation and its application for sprite coding, content description, and segmentation
abstract
We present a new technique for long-term global motion estimation of image objects. The estimated motion parameters describe the continuous and time-consistent motion over the whole sequence relatively to a fixed reference coordinate system. The proposed method is suitable for the estimation of affine motion parameters as well as for higher order motion models like the parabolic model-combining the advantages of feature matching and optical flow techniques. A hierarchical strategy is applied for the estimation, first translation, affine motion, and finally higher order motion parameters, which is robust and computationally efficient. A closed-loop prediction scheme is applied to avoid the problem of error accumulation in long-term motion estimation. The presented results indicate that the proposed technique is a very accurate and robust approach for long-term global motion estimation, which can be used for applications such as MPEG-4 sprite coding or MPEG-7 motion description. We also show that the efficiency of global motion estimation can be significantly increased if a higher order motion model is applied, and we present a new sprite coding scheme for on-line applications. We further demonstrate that the proposed estimator serves as a powerful tool for segmentation of video sequences.
Aljoscha Smolic, Thomas Sikora, Jens-Rainer Ohm
IEEE Trans. Circuits Syst. Video Technol.2
1998 Image sequence analysis for emerging interactive multimedia services-the European COST 211 framework
abstract
Flexibility and efficiency of coding, content extraction, and content-based search are key research topics in the field of interactive multimedia. Ongoing ISO MPEG-4 and MPEG-7 activities are targeting standardization to facilitate such services. European COST Telecommunications activities provide a framework for research collaboration. At present a significant effort of the COST 211/sup ter/ group activities is dedicated toward image and video sequence analysis and segmentation-an important technological aspect for the success of emerging object-based MPEG-4 and MPEG-7 multimedia applications. The current work of COST 211 is centered around the test model, called the analysis model (AM). The essential feature of the AM is its ability to fuse information from different sources to achieve a high-quality object segmentation. The current information sources are the intermediate results from frame-based (still) color segmentation, motion vector based segmentation, and change-detection-based segmentation. Motion vectors, which form the basis for the motion vector based intermediate segmentation, are estimated from consecutive frames. A recursive shortest spanning tree (RSST) algorithm is used to obtain intermediate color and motion vector based segmentation results. A rule-based region processor fuses the intermediate results; a postprocessor further refines the final segmentation output. The results of the current AM are satisfactory.
A. Aydin Alatan, Levent Onural, Michael Wollborn, Roland Mech, Ertem Tuncel, Thomas Sikora
IEEE Trans. Circuits Syst. Video Technol.6
1998 Guest Editorial
King Ngi Ngan, Sethuraman Panchanathan, Thomas Sikora, Ming-Ting Sun
IEEE Trans. Circuits Syst. Video Technol.3
1998 Special Issue On Representation And Coding Of Images And Video I [Guest Editorial]
King Ngi Ngan, Sethuraman Panchanathan, Thomas Sikora, Ming-Ting Sun
IEEE Trans. Circuits Syst. Video Technol.3
1997 Functional coding of video using a shape-adaptive DCT algorithm and an object-based motion prediction toolbox
abstract
This paper presents an object-based layered video coding scheme which achieves very high compression efficiency along with the provision for advanced content-based functionalities, e.g., content-based scalability or content-based access and manipulation of video data. In a first step, a video sequence is segmented into several arbitrarily shaped "object layers." To achieve the desired content-based functionalities, a baseline shape-adaptive discrete cosine transform (DCT) coding algorithm is introduced which can be seen as an extension of conventional block-based DCT coding schemes (e.g., H.261, H.263, MPEG-1, or MPEG-2) toward coding of arbitrarily shaped image content. In order to increase compression efficiency, the baseline object-based layered approach can be extended with an object-based motion prediction toolbox. Using this toolbox, the coding scheme can potentially select specific prediction techniques for every object layer to be coded. To illustrate the concept, an extension of the baseline shape-adaptive DCT algorithm with a technique for global background motion estimation and compensation is described which significantly improves the compression efficiency of suitable video sequences compared to standard MPEG coding schemes.
Peter Kauff, Bela Makai, S. Rauthenberg, Ulrich Gölz, Jan L. P. de Lameillieure, Thomas Sikora
IEEE Trans. Circuits Syst. Video Technol.6
1997 The MPEG-4 video standard verification model
abstract
The MPEG-4 standardization phase has the mandate to develop algorithms for audio-visual coding allowing for interactivity, high compression, and/or universal accessibility and portability of audio and video content. In addition to the conventional "frame"-based functionalities of the MPEG-1 and MPEG-2 standards, the MPEG-4 video coding algorithm will also support access and manipulation of "objects" within video scenes. The January 1996 MPEG Video Group meeting witnessed the definition of the first version of the MPEG-4 video verification model-a milestone in the development of the MPEG-4 standard. The primary intent of the video verification model is to provide a fully defined core video coding algorithm platform for the development of the standard. As such, the structure of the MPEG-4 video verification model already gives some indication about the tools and algorithms that will be provided by the final MPEG-4 standard. The paper describes the scope of the MPEG-4 video standard and outlines the structure of the MPEG-4 video verification model under development.
Thomas Sikora
IEEE Trans. Circuits Syst. Video Technol.1
1997 Optimal Wiener interpolation filters for multiresolution coding of images
abstract
A design approach is presented which allows the optimization of coefficients for symmetric and separable finite impulse response (FIR) interpolation filters for multiresolution coding schemes. The interpolation filters serve as optimal inverse filters in the Wiener sense and are designed to match the characteristics of the specific filters used for decimation as well as for the statistics of "typical" images to be reconstructed. Applied to the coding of test images in a four-level progressive pyramid scheme, the optimal interpolation filters generated substantially improved rate-distortion results compared to conventional filters.
Thomas Sikora
IEEE Trans. Circuits Syst. Video Technol.1
1996 The European COST211ter activities-research towards advanced algorithms for coding of video signals at very low bit rates
abstract
A significant effort of the COST211ter group activities is dedicated towards the contribution for the MPEG-4 video group activities. This paper provides an overview of the COST telecommunication framework and discusses the technical simulation model approach taken by the COST211ter group. Video coding with application to multimedia services is discussed.
Thomas Sikora, Jörn Ostermann
ICIP (3)1
1996 Compression algorithms for software coding of video
Thomas Sikora, Eric Viscito
Signal Process. Image Commun.1
1996 Optimal block-overlapping synthesis transforms for coding images and video at very low bitrates
abstract
In this paper we address the problem of coding images and video at very low bitrates using conventional hybrid differential pulse code modulation/discrete cosine transform (DPCM/DCT) coding schemes with reduced blocking artifacts. To this end, optimal block-overlapping synthesis filterbanks (block-overlapping inverse transform kernels) with minimum least-squares reconstruction error properties are derived which replace the conventional inverse DCT if only a few DCT coefficients are transmitted to the receiver. Results show that the proposed inverse block-overlapping transforms can significantly outperform the conventional inverse DCT both in terms of objective and subjective quality measures.
Thomas Sikora
IEEE Trans. Circuits Syst. Video Technol.1
1995 Digital video coding standards and their role in video communications
abstract
The efficient digital representation of image and video signals has been subject of considerable research over the past 20 years. With the growing availability of digital transmission links, progress in signal processing, VLSI technology and image compression research, visual communications has become more feasible than ever. Digital video coding technology has developed into a mature field and a diversity of products has been developed-targeted for a wide range of emerging applications, such as video on demand, digital TV/HDTV broadcasting, and multimedia image/video database services. With the increased commercial interest in video communications the need for international image and video coding standards arose. Standardization of video coding algorithms holds the promise of large markets for video communication equipment. Interoperability of implementations from different vendors enables the consumer to access video from a wider range of services and VLSI implementations of coding algorithms conforming to international standards can be manufactured at considerably reduced costs. The purpose of this paper is to provide an overview of today's image and video coding standards and their role in video communications. The different coding algorithms developed for each standard are reviewed and the commonalities between the standards are discussed.>
Ralf Schafer, Thomas Sikora
Proc. IEEE2
1995 Low complexity shape-adaptive DCT for coding of arbitrarily shaped image segments
Thomas Sikora
Signal Process. Image Commun.1
1995 Efficiency of shape-adaptive 2-D transforms for coding of arbitrarily shaped image segments
abstract
We introduce a formula to compute an optimum 2-D shape-adaptive Karhunen-Loeve transform (KLT) suitable for coding pels in arbitrarily-shaped image segments. The efficiency of the KLT on a 2-D AR(1) process is used to benchmark two other shape-adaptive transforms described in literature. It is shown that the optimum KLT significantly outperforms the well known shape-adaptive DCT method introduced by Gilge et al. (1989) for coding Segments of arbitrary shape in intraframe coding mode. A statistical transform gain close to the Gilge-method can be achieved with a shape-adaptive DCT algorithm introduced by Sikora and Makai (see Proc. Workshop Image Anal. Image Coding, Berlin, FRG, Nov. 1993) which is implemented with much lower complexity.>
Thomas Sikora, Sven Bauer, Bela Makai
IEEE Trans. Circuits Syst. Video Technol.1
1995 Shape-adaptive DCT for generic coding of video
abstract
A low complexity shape-adaptive DCT algorithm suitable for coding pels in arbitrarily shaped image segments is introduced. In contrast to other techniques described in literature the proposed algorithm is based on predefined orthogonal sets of DCT basis functions and does not require more computations than a normal block DCT. It is shown that the shape-adaptive DCT algorithm can be easily incorporated into existing block-based JPEG, H.261, or MPEG coding schemes. Thus segment or object based coding of images and video can be provided with backward compatibility to existing coding standards. As an important feature, with the proposed technique additional content based functionalities currently discussed in the MPEG-4 standardization phase can be readily achieved.>
Thomas Sikora, Bela Makai
IEEE Trans. Circuits Syst. Video Technol.1