Martin Kleinsteuber

dblp:31/677 · DBLP profile ↗
← Back
36ranked-venue papers
3as first author
6since 2021 · last 2023
0000-0002-4323-9260ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 22 · 3 first-author · 2 since 2021Artificial intelligence and machine learning · 15 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Theory of computation · 2Systems, architecture and hardware · 1Computer networks · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Representation and self-supervised learning · 52% Transfer learning and domain adaptation · 10% Graph learning · 10%
Computer graphics and multimedia
8 papers
Image and video processing · 65% Visualization and visual analytics · 35%
Databases, data mining, and information retrieval
1 paper
Data mining · 100%

Topics — the 30 heaviest of 40, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › sparse coding
dictionary learning
0.942017
Dynamical Textures Modeling via Joint Video Dictionary Learning · IEEE Trans. Image Process. 2017
Trace Quotient Meets Sparsity: A Method for Learning Low Dimensional Image Representations · CVPR 2016
Sample Complexity of Dictionary Learning and Other Matrix Factorizations · IEEE Trans. Inf. Theory 2015
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
sparse coding
0.722020
Trace Quotient with Sparsity Priors for Learning Low Dimensional Image Representations · IEEE Trans. Pattern Anal. Mach. Intell. 2020
Trace Quotient Meets Sparsity: A Method for Learning Low Dimensional Image Representations · CVPR 2016
Natural language and speech › Language models and text generation
compositional generalization
0.612022
On Leveraging Variational Graph Embeddings for Open World Compositional Zero-Shot Learning · ACM Multimedia 2022
Machine learning › Transfer learning and domain adaptation › zero-shot learning
compositional zero-shot learning
0.612022
On Leveraging Variational Graph Embeddings for Open World Compositional Zero-Shot Learning · ACM Multimedia 2022
Machine learning › Graph learning › graph autoencoder
variational graph autoencoder
0.612022
On Leveraging Variational Graph Embeddings for Open World Compositional Zero-Shot Learning · ACM Multimedia 2022
Data mining
clustering
0.612022
Cluster-Aware Heterogeneous Information Network Embedding · WSDM 2022
Data mining › clustering
graph clustering
0.612022
Cluster-Aware Heterogeneous Information Network Embedding · WSDM 2022
Data mining › structured data mining
graph mining
0.612022
Cluster-Aware Heterogeneous Information Network Embedding · WSDM 2022
Data mining › representation learning › graph representation learning
heterogeneous information network embedding
0.612022
Cluster-Aware Heterogeneous Information Network Embedding · WSDM 2022
Machine learning › Representation and self-supervised learning › representation learning › latent representation learning
low-dimensional representation learning
0.412020
Trace Quotient with Sparsity Priors for Learning Low Dimensional Image Representations · IEEE Trans. Pattern Anal. Mach. Intell. 2020
Visualization and visual analytics
dimensionality reduction
0.412020
Glyphboard: Visual Exploration of High-Dimensional Data Combining Glyphs with Dimensionality Reduction · IEEE Trans. Vis. Comput. Graph. 2020
Visualization and visual analytics › visual encoding
glyph-based visualization
0.412020
Glyphboard: Visual Exploration of High-Dimensional Data Combining Glyphs with Dimensionality Reduction · IEEE Trans. Vis. Comput. Graph. 2020
Visualization and visual analytics
high-dimensional data visualization
0.412020
Glyphboard: Visual Exploration of High-Dimensional Data Combining Glyphs with Dimensionality Reduction · IEEE Trans. Vis. Comput. Graph. 2020
Image and video processing › feature extraction
feature learning
0.312018
Model-Based Learning of Local Image Features for Unsupervised Texture Segmentation · IEEE Trans. Image Process. 2018
Image and video processing
image segmentation
0.312018
Model-Based Learning of Local Image Features for Unsupervised Texture Segmentation · IEEE Trans. Image Process. 2018
Image and video processing › image segmentation › variational segmentation
mumford-shah functional
0.312018
Model-Based Learning of Local Image Features for Unsupervised Texture Segmentation · IEEE Trans. Image Process. 2018
Image and video processing › image segmentation
texture segmentation
0.312018
Model-Based Learning of Local Image Features for Unsupervised Texture Segmentation · IEEE Trans. Image Process. 2018
Image and video processing
variational methods
0.312018
Model-Based Learning of Local Image Features for Unsupervised Texture Segmentation · IEEE Trans. Image Process. 2018
Computer vision › Video understanding and tracking › motion analysis
dynamic texture modeling
0.312017
Dynamical Textures Modeling via Joint Video Dictionary Learning · IEEE Trans. Image Process. 2017
Machine learning › Representation and self-supervised learning › representation learning
dimensionality reduction
0.212016
Trace Quotient Meets Sparsity: A Method for Learning Low Dimensional Image Representations · CVPR 2016
Machine learning › Representation and self-supervised learning
matrix factorization
0.212015
Sample Complexity of Dictionary Learning and Other Matrix Factorizations · IEEE Trans. Inf. Theory 2015
Machine learning › Representation and self-supervised learning › matrix factorization
nonnegative matrix factorization
0.212015
Sample Complexity of Dictionary Learning and Other Matrix Factorizations · IEEE Trans. Inf. Theory 2015
Machine learning › Learning theory
sample complexity
0.212015
Sample Complexity of Dictionary Learning and Other Matrix Factorizations · IEEE Trans. Inf. Theory 2015
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › sparse coding › dictionary learning
sparse dictionary learning
0.212015
Sample Complexity of Dictionary Learning and Other Matrix Factorizations · IEEE Trans. Inf. Theory 2015
Image and video processing › image segmentation
interactive segmentation
0.212014
Co-Sparse Textural Similarity for Interactive Segmentation · ECCV (6) 2014
Computer vision › 3D vision
depth estimation
0.212013
A Joint Intensity and Depth Co-sparse Analysis Model for Depth Map Super-resolution · ICCV 2013
Image and video processing › super-resolution › image super-resolution
depth super-resolution
0.212013
A Joint Intensity and Depth Co-sparse Analysis Model for Depth Map Super-resolution · ICCV 2013
Image and video processing
image reconstruction
0.212013
Analysis Operator Learning and its Application to Image Reconstruction · IEEE Trans. Image Process. 2013
Visualization and visual analytics › visual analytics
visual analytics system
0.112020
Glyphboard: Visual Exploration of High-Dimensional Data Combining Glyphs with Dimensionality Reduction · IEEE Trans. Vis. Comput. Graph. 2020
Computer vision › 3D vision › depth estimation › stereo depth estimation
dense disparity estimation
0.112011
Dense disparity maps from sparse disparity measurements · ICCV 2011

Methods — techniques the papers use, named apart from their topics

contrastive learning · 1.1sparse coding · 0.7variational inference · 0.6variational graph autoencoder · 0.6meta-path · 0.6deep metric learning · 0.6trace quotient criterion · 0.4glyph-based visualization · 0.4dimensionality reduction · 0.4conjugate gradient · 0.4unsupervised learning · 0.3convolutional features · 0.3transition matrix · 0.3markov random process · 0.3riemannian optimization · 0.2elastic net · 0.2co-sparse analysis · 0.2bimodal modeling · 0.2
YearPublicationVenuePosition
2023 Barlow Graph Auto-Encoder for Unsupervised Network Embedding
abstract
Network embedding has emerged as a promising research field for network analysis. Recently, an approach, named Barlow Twins, has been proposed for self-supervised learning in computer vision by applying the redundancy-reduction principle to the embedding vectors corresponding to two distorted versions of the image samples. Motivated by this, we propose Barlow Graph Auto-Encoder, a simple yet effective architecture for learning network embedding. It aims to maximize the similarity between the embedding vectors of immediate and larger neighborhoods of a node while minimizing the redundancy between the components of these projections. In addition, we also present the variational counterpart named Barlow Variational Graph Auto-Encoder. We demonstrate the effectiveness of our approach in learning multiple graph-related tasks, i.e., link prediction, clustering, and downstream node classification, by providing extensive comparisons with several well-known techniques on eight benchmark datasets.
Rayyan Ahmad Khan, Martin Kleinsteuber
AISTATS2
2022 On Leveraging Variational Graph Embeddings for Open World Compositional Zero-Shot Learning
abstract
Humans are able to identify and categorize novel compositions of known concepts. The task in Compositional Zero-Shot learning (CZSL) is to learn composition of primitive concepts, i.e. objects and states, in such a way that even their novel compositions can be zero-shot classied. In this work, we do not assume any prior knowledge on the feasibility of novel compositions, i.e. open-world setting, where infeasible compositions dominate the search space. We propose a Compositional Variational Graph Autoencoder (CVGAE) approach for learning the variational embeddings of the primitive concepts (nodes) as well as feasibility of their compositions (via edges). Such modelling makes CVGAE scalable to real-world application scenarios. This is in contrast to SOTA method, CGE, which is computationally very expensive. e.g. for benchmark C-GQA dataset, CGE requires 3.94×10^5 nodes, whereas CVGAE requires only 1323 nodes. We learn a mapping of the graph and image embeddings onto a common embedding space. CVGAE adopts a deep metric learning approach and learns a similarity metric in this space via bi-directional contrastive loss between projected graph and image embeddings. We validate the eectiveness of our approach on three benchmark datasets. We also demonstrate via an image retrieval task that the representations learnt by CVGAE are better suited for compositional generalization.
Muhammad Umer Anwaar, Zhihui Pan, Martin Kleinsteuber
ACM Multimedia3
2022 Cluster-Aware Heterogeneous Information Network Embedding
abstract
Heterogeneous Information Network (HIN) embedding refers to the low-dimensional projections of the HIN nodes that preserve the HIN structure and semantics. HIN embedding has emerged as a promising research field for network analysis as it enables downstream tasks such as clustering and node classification. In this work, we propose VaCA-HINE for joint learning of cluster embeddings as well as cluster-aware HIN embedding. We assume that the connected nodes are highly likely to fall in the same cluster, and adopt a variational approach to preserve the information in the pairwise relations in a cluster-aware manner. In addition, we deploy contrastive modules to simultaneously utilize the information in multiple meta-paths, thereby alleviating the meta-path selection problem - a challenge faced by many of the famous HIN embedding approaches. The HIN embedding, thus learned, not only improves the clustering performance but also preserves pairwise proximity as well as the high-order HIN structure. We show the effectiveness of our approach by comparing it with many competitive baselines on three real-world datasets on clustering and downstream node classification.
Rayyan Ahmad Khan, Martin Kleinsteuber
WSDM2
2021 A Contrastive Learning Approach for Compositional Zero-Shot Learning
abstract
An object can be in several states. For different states (attributes) the object could look dramatically different. Thus, the smart information retrieval systems of the future need to learn good state-object representations. Such a system should not only be able to recognize state-object compositions unseen during training but also be able to retrieve images based on multi-modal (image-text) query. In the literature, these tasks are treated separately. In this work, we propose a unified model, ContraNet, which leverages the rich semantics of the state-object to learn multimodal representation in a contrastive manner. We adopt a deep metric learning approach and learn a multimodal representation by pulling similar images and texts closer to each other and pushing apart different ones. Our autoencoder based model learns the text-aware representation of image which is suitable for both tasks. The reconstruction losses provide additional regularization for learning of the representation. Our approach outperforms the state-of-the-art (SOTA) methods on widely-used benchmarks. Specifically, on the task of state-object composition, ContraNet achieves 8.7% and 8.1% performance gain on UT-Zappos and MIT-States on best HM metric, respectively. For the image retrieval task, ContraNet surpasses the SOTA performance by 4% on MIT-States and 5.3% on Fashion200k.
Muhammad Umer Anwaar, Rayyan Ahmad Khan, Zhihui Pan, Martin Kleinsteuber
ICMI4
2021 Unsupervised Learning of Joint Embeddings for Node Representation and Community Detection
Rayyan Ahmad Khan, Muhammad Umer Anwaar, Omran Kaddah, Zhiwei Han, Martin Kleinsteuber
ECML/PKDD (2)5
2021 Compositional Learning of Image-Text Query for Image Retrieval
abstract
In this paper, we investigate the problem of retrieving images from a database based on a multi-modal (imagetext) query. Specifically, the query text prompts some modification in the query image and the task is to retrieve images with the desired modifications. For instance, a user of an E-Commerce platform is interested in buying a dress, which should look similar to her friend's dress, but the dress should be of white color with a ribbon sash. In this case, we would like the algorithm to retrieve some dresses with desired modifications in the query dress. We propose an autoencoder based model, ComposeAE, to learn the composition of image and text query for retrieving images. We adopt a deep metric learning approach and learn a metric that pushes composition of source image and text query closer to the target images. We also propose a rotational symmetry constraint on the optimization problem. Our approach is able to outperform the state-of-the-art method TIRG [24] on three benchmark datasets, namely: MIT-States, Fashion200k and Fashion IQ. In order to ensure fair comparison, we introduce strong baselines by enhancing TIRG method. To ensure reproducibility of the results, we publish our code here: https://github.com/ecom-research/ComposeAE.
Muhammad Umer Anwaar, Egor Labintcev, Martin Kleinsteuber
WACV3
2020 Epitomic Variational Graph Autoencoder
abstract
Variational autoencoder (VAE) is a widely used generative model for learning latent representations. Burda et al. [3] in their seminal paper showed that learning capacity of VAE is limited by over-pruning. It is a phenomenon where a significant number of latent variables fail to capture any information about the input data and the corresponding hidden units become inactive. This adversely affects learning diverse and interpretable latent representations. As variational graph autoencoder (VGAE) extends VAE for graph-structured data, it inherits the over-pruning problem. In this paper, we adopt a model based approach and propose epitomic VGAE (EVGAE), a generative variational framework for graph datasets which successfully mitigates the over-pruning problem and also boosts the generative ability of VGAE. We consider EVGAE to consist of multiple sparse VGAE models, called epitomes, that are groups of latent variables sharing the latent space. This approach aids in increasing active units as epitomes compete to learn better representation of the graph data. We verify our claims via experiments on three benchmark datasets. Our experiments show that EVGAE has a better generative ability than VGAE. Moreover, EVGAE outperforms VGAE on link prediction task in citation networks.
Rayyan Ahmad Khan, Muhammad Umer Anwaar, Martin Kleinsteuber
ICPR3
2020 Trace Quotient with Sparsity Priors for Learning Low Dimensional Image Representations
abstract
This work studies the problem of learning appropriate low dimensional image representations. We propose a generic algorithmic framework, which leverages two classic representation learning paradigms, i.e., sparse representation and the trace quotient criterion, to disentangle underlying factors of variation in high dimensional images. Specifically, we aim to learn simple representations of low dimensional, discriminant factors by applying the trace quotient criterion to well-engineered sparse representations. We construct a unified cost function, coined as the SPARse LOW dimensional representation (SparLow) function, for jointly learning both a sparsifying dictionary and a dimensionality reduction transformation. The SparLow function is widely applicable for developing various algorithms in three classic machine learning scenarios, namely, unsupervised, supervised, and semi-supervised learning. In order to develop efficient joint learning algorithms for maximizing the SparLow function, we deploy a framework of sparse coding with appropriate convex priors to ensure the sparse representations to be locally differentiable. Moreover, we develop an efficient geometric conjugate gradient algorithm to maximize the SparLow function on its underlying Riemannian manifold. Performance of the proposed SparLow algorithmic framework is investigated on several image processing tasks, such as 3D data visualization, face/digit recognition, and object/scene categorization.
Xian Wei, Hao Shen 0002, Martin Kleinsteuber
IEEE Trans. Pattern Anal. Mach. Intell.3
2020 Glyphboard: Visual Exploration of High-Dimensional Data Combining Glyphs with Dimensionality Reduction
abstract
Rigorous data science is interdisciplinary at its core. In order to make sense of high-dimensional data, data scientists need to enter into a dialogue with domain experts. We present Glyphboard, a visualization tool that aims to support this dialogue. Glyphboard is a zoomable user interface that combines well-known methods such as dimensionality reduction and glyph-based visualizations in a novel, seamless, and integrated tool. While the dimensionality reduction affords a quick overview over the data, glyph-based visualizations are able to show the most relevant dimensions in the data set at one glance. We contribute an open-source prototype of Glyphboard, a general exchange format for high-dimensional data, and a case study with nine data scientists and domain experts from four exemplary domains in order to evaluate how the different visualization and interaction features of Glyphboard are used.
Dietrich Kammer, Mandy Keck, Thomas Gründer, Alexander Maasch, Thomas Thom, Martin Kleinsteuber, Rainer Groh 0001
IEEE Trans. Vis. Comput. Graph.6
2019 Reconstructible Nonlinear Dimensionality Reduction via Joint Dictionary Learning
abstract
This paper presents a parametric low-dimensional (LD) representation learning method that allows to reconstruct high-dimensional (HD) input vectors in an unsupervised manner. Under the assumption that the HD data and its LD representation share the same or similar local sparse structure, the proposed method achieves reconstructible dimensionality reduction via jointly learning dictionaries in both the original HD data space and its LD representation space. By regarding the sparse representation as a smooth function with respect to a specific dictionary, we construct an encoding-decoding block for learning LD representations from sparse coefficients of HD data. It is expected that this learning process preserves the desirable structure of HD data in the LD representation space, and simultaneously allows a reliable reconstruction from the LD space back to the original HD space. In addition, the proposed single layer encoding-decoding block can be easily extended to deep learning structures. Numerical experiments on both synthetic data sets and real images show that the proposed method achieves strongly competitive and robust performance in data DR, reconstruction, and synthesis, even on heavily corrupted data. The proposed method can be used as an alternative approach to compressive sensing (CS); however, it can outperform the traditional CS methods in: 1) task-driven learning problems, such as 2-D/3-D data visualization, and 2) data reconstruction at a lower dimensional space.
Xian Wei, Hao Shen 0002, Martin Kleinsteuber, Yi Lu Murphey
IEEE Trans. Neural Networks Learn. Syst.6
2018 Alignment Distances on Systems of Bags
abstract
Recent research in image and video recognition indicates that many visual processes can be thought of as being generated by a time-varying generative model. A nearby descriptive model for visual processes is thus a statistical distribution that varies over time. Specifically, modeling visual processes as streams of histograms generated by a kernelized linear dynamic system turns out to be efficient. We refer to such a model as a system of bags. In this paper, we investigate systems of bags with special emphasis on dynamic scenes and dynamic textures. Parameters of linear dynamic systems suffer from ambiguities. In order to cope with these ambiguities in the kernelized setting, we develop a kernelized version of the alignment distance. For its computation, we use a Jacobi-type method and prove its convergence to a set of critical points. We employ it as a dissimilarity measure on Systems of Bags. As such, it outperforms other known dissimilarity measures for kernelized linear dynamic systems, in particular the Martin distance and the Maximum singular value distance, in every tested classification setting. A considerable margin can be observed in settings, where classification is performed with respect to an abstract mean of video sets. For this scenario, the presented approach can outperform the state-of-the-art techniques, such as dynamic fractal spectrum or orthogonal tensor dictionary learning.
Alexander Sagel, Martin Kleinsteuber
IEEE Trans. Circuits Syst. Video Technol.2
2018 Model-Based Learning of Local Image Features for Unsupervised Texture Segmentation
abstract
Features that capture well the textural patterns of a certain class of images are crucial for the performance of texture segmentation methods. The manual selection of features or designing new ones can be a tedious task. Therefore, it is desirable to automatically adapt the features to a certain image or class of images. Typically, this requires a large set of training images with similar textures and ground truth segmentation. In this work, we propose a framework to learn features for texture segmentation when no such training data is available. The cost function for our learning process is constructed to match a commonly used segmentation model, the piecewise constant Mumford-Shah model. This means that the features are learned such that they provide an approximately piecewise constant feature image with a small jump set. Based on this idea, we develop a two-stage algorithm which first learns suitable convolutional features and then performs a segmentation. We note that the features can be learned from a small set of images, from a single image, or even from image patches. The proposed method achieves a competitive rank in the Prague texture segmentation benchmark, and it is effective for segmenting histological images.
Martin Kiechle, Martin Storath, Andreas Weinmann, Martin Kleinsteuber
IEEE Trans. Image Process.4
2017 Towards Glyph-based visualizations for big data clustering
abstract
Data Analysts have to deal with an ever-growing amount of data resources. One way to make sense of this data is to extract features and use clustering algorithms to group items according to a similarity measure. Algorithm developers are challenged when evaluating the performance of the algorithm since it is hard to identify features that influence the clustering. Moreover, many algorithms can be trained using a semi-supervised approach, where human users provide ground truth samples by manually grouping single items. Hence, visualization techniques are needed that help data analysts achieve their goal in evaluating Big data clustering algorithms. In this context, Multidimensional Scaling (MDS) has become a prominent visualization tool. In this paper, we propose a combination with glyphs that can provide a detailed view of specific features involved in MDS. In consequence, human users can understand, adjust, and ultimately improve clustering algorithms. We present a thorough glyph design, which is founded in a comprehensive survey of related work and report the results of a controlled experiments, where participants solved data analysis tasks with both glyphs and a traditional textual display of data values.
Mandy Keck, Dietrich Kammer, Thomas Gründer, Thomas Thom, Martin Kleinsteuber, Alexander Maasch, Rainer Groh 0001
VINCI5
2017 Dynamical Textures Modeling via Joint Video Dictionary Learning
abstract
Video representation is an important and challenging task in the computer vision community. In this paper, we consider the problem of modeling and classifying video sequences of dynamic scenes which could be modeled in a dynamic textures (DTs) framework. At first, we assume that image frames of a moving scene can be modeled as a Markov random process. We propose a sparse coding framework, named joint video dictionary learning (JVDL), to model a video adaptively. By treating the sparse coefficients of image frames over a learned dictionary as the underlying "states", we learn an efficient and robust linear transition matrix between two adjacent frames of sparse events in time series. Hence, a dynamic scene sequence is represented by an appropriate transition matrix associated with a dictionary. In order to ensure the stability of JVDL, we impose several constraints on such transition matrix and dictionary. The developed framework is able to capture the dynamics of a moving scene by exploring both the sparse properties and the temporal correlations of consecutive video frames. Moreover, such learned JVDL parameters can be used for various DT applications, such as DT synthesis and recognition. Experimental results demonstrate the strong competitiveness of the proposed JVDL approach in comparison with the state-of-the-art video representation methods. Especially, it performs significantly better in dealing with DT synthesis and recognition on heavily corrupted data.
Xian Wei, Hao Shen 0002, Martin Kleinsteuber, Zhongfeng Wang 0001
IEEE Trans. Image Process.5
2016 Trace Quotient Meets Sparsity: A Method for Learning Low Dimensional Image Representations
abstract
This paper presents an algorithm that allows to learn low dimensional representations of images in an unsupervised manner. The core idea is to combine two criteria that play important roles in unsupervised representation learning, namely sparsity and trace quotient. The former is known to be a convenient tool to identify underlying factors, and the latter is known as a disentanglement of underlying discriminative factors. In this work, we develop a generic cost function for learning jointly a sparsifying dictionary and a dimensionality reduction transformation. It leads to several counterparts of classic low dimensional representation methods, such as Principal Component Analysis, Local Linear Embedding, and Laplacian Eigenmap. Our proposed optimisation algorithm leverages the efficiency of geometric optimisation on Riemannian manifolds and a closed form solution to the elastic net problem.
Xian Wei, Hao Shen 0002, Martin Kleinsteuber
CVPR3
2016 Active classifier selection for RGB-D object categorization using a Markov random field ensemble method
abstract
In this work, a new ensemble method for the task of category recognition in different environments is presented. The focus is on service robotic perception in an open environment, where the robot’s task is to recognize previously unseen objects of predefined categories, based on training on a public dataset. We propose an ensemble learning approach to be able to flexibly combine complementary sources of information (different state-of-the-art descriptors computed on color and depth images), based on a Markov Random Field (MRF). By exploiting its specific characteristics, the MRF ensemble method can also be executed as a Dynamic Classifier Selection (DCS) system. In the experiments, the committee- and topology-dependent performance boost of our ensemble is shown. Despite reduced computational costs and using less information, our strategy performs on the same level as common ensemble approaches. Finally, the impact of large differences between datasets is analyzed.
Maximilian Durner, Zoltan-Csaba Marton, Ulrich Hillenbrand, Martin Kleinsteuber
ICMV5
2016 Joint learning dictionary and discriminative features for high dimensional data
abstract
Recently, sparse representation (SR) over a redundant dictionary has become a popular way of representing the data. It has been verified as an efficient and useful tool to promote the discrimination between signals. This work develops a joint learning approach to find the low dimensional discriminative features for high dimensional data. To avoid the high computational cost of direct sparse coding on large scale input data, we first learn SR in an orthogonal projected space over a task-driven sparsifying dictionary. We then exploit the discriminative projection on SR. The whole learning process is treated as an optimization problem of trace quotient maximization, which involves an orthogonal projection on original data space, a dictionary and a discriminative projection on sparse codes. The related cost function is well defined on a product manifold of the Stiefel manifold, the Oblique manifold and the Grassmann manifold. Finally, we employ a stochastic gradient descent algorithm on the smooth product manifold to maximize the cost function. Our numerical experiments on visual recognition demonstrate the effectiveness of the proposed algorithm, in comparison with the state of the arts.
Xian Wei, Hao Shen 0002, Martin Kleinsteuber, Yi Lu Murphey
ICPR4
2016 Network Volume Anomaly Detection and Identification in Large-Scale Networks Based on Online Time-Structured Traffic Tensor Tracking
abstract
This paper addresses network anomography, that is, the problem of inferring network-level anomalies from indirect link measurements. This problem is cast as a low-rank subspace tracking problem for normal flows under incomplete observations and an outlier detection problem for abnormal flows. Since traffic data is large-scale time-structured data accompanied with noise and outliers under partial observations, an efficient modeling method is essential. To this end, this paper proposes an online subspace tracking of a Hankelized time-structured traffic tensor for normal flows based on the Candecomp/PARAFAC decomposition exploiting the recursive least squares algorithm. We estimate abnormal flows as outlier sparse flows via sparsity maximization in the underlying under-constrained linear-inverse problem. A major advantage is that our algorithm estimates normal flows by low-dimensional matrices with time-directional features as well as the spatial correlation of multiple links without using the past observed measurements and the past model parameters. Extensive numerical evaluations show that the proposed algorithm achieves faster convergence per iteration of model approximation and better volume anomaly detection performance compared to state-of-the-art algorithms.
Hiroyuki Kasai, Wolfgang Kellerer, Martin Kleinsteuber
IEEE Trans. Netw. Serv. Manag.3
2015 A Bimodal Co-sparse Analysis Model for Image Processing
Martin Kiechle, Tim Habigt, Simon Hawe, Martin Kleinsteuber
Int. J. Comput. Vis.4
2015 Sample Complexity of Dictionary Learning and Other Matrix Factorizations
abstract
Many modern tools in machine learning and signal processing, such as sparse dictionary learning, principal component analysis, non-negative matrix factorization, K-means clustering, and so on, rely on the factorization of a matrix obtained by concatenating high-dimensional vectors from a training collection. While the idealized task would be to optimize the expected quality of the factors over the underlying distribution of training vectors, it is achieved in practice by minimizing an empirical average over the considered collection. The focus of this paper is to provide sample complexity estimates to uniformly control how much the empirical average deviates from the expected cost function. Standard arguments imply that the performance of the empirical predictor also exhibit such guarantees. The level of genericity of the approach encompasses several possible constraints on the factors (tensor product structure, shift-invariance, sparsity...), thus providing a unified perspective on the sample complexity of several widely used matrix factorization schemes. The derived generalization bounds behave proportional to (log (n)/n)1/2with respect to the number of samples n for the considered matrix factorization techniques.
Rémi Gribonval, Rodolphe Jenatton, Francis R. Bach, Martin Kleinsteuber, Matthias Seibert
IEEE Trans. Inf. Theory4
2014 Co-Sparse Textural Similarity for Interactive Segmentation
Claudia Nieuwenhuis, Simon Hawe, Martin Kleinsteuber, Daniel Cremers
ECCV (6)3
2014 An adaptive dictionary learning approach for modeling dynamical textures
abstract
Video representation is an important and challenging task in the computer vision community. In this paper, we assume that image frames of a moving scene can be modeled as a Markov random process. We propose a sparse coding framework, named adaptive video dictionary learning (AVDL), to model a video adaptively. The developed framework is able to capture the dynamics of a moving scene by exploring both sparse properties and the temporal correlations of consecutive video frames. The proposed method is compared with state of the art video processing methods on several benchmark data sequences, which exhibit appearance changes and heavy occlusions.
Xian Wei, Hao Shen 0002, Martin Kleinsteuber
ICASSP3
2014 An ontology-based approach for decentralized monitoring and diagnostics
abstract
Modern decentralized industrial applications demand the design of application-independent solutions for monitoring and diagnostics systems (MDSs) that exhibit a high degree of flexibility and re-utilization. To achieve this, we propose an ontology-based approach that adheres to the Meta Object Facility (MOF) paradigm for engineering and maintenance of MDSs. The key of our approach is to built a decentralized system architecture implemented on a semantic technology stack. Our architecture allows for storing plant engineering expert knowledge and the monitoring and diagnosis rules in formalized OWL models. The plant models can then be processed by the rules to compute monitoring states and diagnose causes of faults. This paper specifically focuses on a system implementation in alignment to requirements of the industrial domain. Based on these requirements, alternative knowledge-based tools and techniques are compared to evaluate the effectiveness of our approach.
Lisa Abele, Stephan Grimm, Sonja Zillner, Martin Kleinsteuber
INDIN4
2014 pROST: a smoothed ℓp-norm robust online subspace tracking method for background subtraction in video
Florian Seidel, Clemens Hage, Martin Kleinsteuber
Mach. Vis. Appl.3
2013 Separable Dictionary Learning
abstract
Many techniques in computer vision, machine learning, and statistics rely on the fact that a signal of interest admits a sparse representation over some dictionary. Dictionaries are either available analytically, or can be learned from a suitable training set. While analytic dictionaries permit to capture the global structure of a signal and allow a fast implementation, learned dictionaries often perform better in applications as they are more adapted to the considered class of signals. In imagery, unfortunately, the numerical burden for (i) learning a dictionary and for (ii) employing the dictionary for reconstruction tasks only allows to deal with relatively small image patches that only capture local image information. The approach presented in this paper aims at overcoming these drawbacks by allowing a separable structure on the dictionary throughout the learning process. On the one hand, this permits larger patch-sizes for the learning phase, on the other hand, the dictionary is applied efficiently in reconstruction tasks. The learning procedure is based on optimizing over a product of spheres which updates the dictionary as a whole, thus enforces basic dictionary properties such as mutual coherence explicitly during the learning procedure. In the special case where no separable structure is enforced, our method competes with state-of-the-art dictionary learning methods like K-SVD.
Simon Hawe, Matthias Seibert, Martin Kleinsteuber
CVPR3
2013 A Joint Intensity and Depth Co-sparse Analysis Model for Depth Map Super-resolution
abstract
High-resolution depth maps can be inferred from low-resolution depth measurements and an additional high-resolution intensity image of the same scene. To that end, we introduce a bimodal co-sparse analysis model, which is able to capture the interdependency of registered intensity and depth information. This model is based on the assumption that the co-supports of corresponding bimodal image structures are aligned when computed by a suitable pair of analysis operators. No analytic form of such operators exist and we propose a method for learning them from a set of registered training signals. This learning process is done offline and returns a bimodal analysis operator that is universally applicable to natural scenes. We use this to exploit the bimodal co-sparse analysis model as a prior for solving inverse problems, which leads to an efficient algorithm for depth map super-resolution.
Martin Kiechle, Simon Hawe, Martin Kleinsteuber
ICCV3
2013 Averaging complex subspaces via a Karcher mean approach
Knut Hüper, Martin Kleinsteuber, Hao Shen 0002
Signal Process.2
2013 Analysis Based Blind Compressive Sensing
abstract
In this letter, we address the problem of blindly reconstructing compressively sensed signals by exploiting the co-sparse analysis model. In the analysis model it is assumed that a signal multiplied by an analysis operator results in a sparse vector. We propose an algorithm that learns the operator adaptively during the reconstruction process. The arising optimization problem is tackled via a geometric conjugate gradient approach. Different types of sampling noise are handled by simply exchanging the data fidelity term. Numerical experiments are performed for measurements corrupted with Gaussian as well as impulsive noise to show the effectiveness of our method.
Julian Wörmann, Simon Hawe, Martin Kleinsteuber
IEEE Signal Process. Lett.3
2013 Analysis Operator Learning and its Application to Image Reconstruction
abstract
Exploiting a priori known structural information lies at the core of many image reconstruction methods that can be stated as inverse problems. The synthesis model, which assumes that images can be decomposed into a linear combination of very few atoms of some dictionary, is now a well established tool for the design of image reconstruction algorithms. An interesting alternative is the analysis model, where the signal is multiplied by an analysis operator and the outcome is assumed to be sparse. This approach has only recently gained increasing interest. The quality of reconstruction methods based on an analysis model severely depends on the right choice of the suitable operator. In this paper, we present an algorithm for learning an analysis operator from training images. Our method is based on l(p)-norm minimization on the set of full rank matrices with normalized columns. We carefully introduce the employed conjugate gradient method on manifolds, and explain the underlying geometry of the constraints. Moreover, we compare our approach to state-of-the-art methods for image denoising, inpainting, and single image super-resolution. Our numerical results show competitive performance of our general approach in all presented applications compared to the specialized state-of-the-art techniques.
Simon Hawe, Martin Kleinsteuber, Klaus Diepold
IEEE Trans. Image Process.2
2012 Cartoon-like image reconstruction via constrained ℓp-minimization
abstract
This paper considers the problem of reconstructing images from only a few measurements. A method is proposed that is based on the theory of Compressive Sensing. We introduce a new prior that combines an ℓp-pseudo-norm approximation of the image gradient and the bounded range of the original signal. Ultimately, this leads to a reconstruction algorithm that works particularly well for Cartoon-like images that commonly occur in medical imagery. The arising optimization task is solved by a Conjugate Gradient method that is capable of dealing with large scale problems and easily adapts to extensions of the prior. To overcome the none differentiability of the ℓp-pseudo-norm we employ a Huber-loss term like approximation together with a continuation of the smoothing parameter. Numerical results and a comparison with the state-of-the-art methods show the effectiveness of the proposed algorithm.
Simon Hawe, Martin Kleinsteuber, Klaus Diepold
ICASSP2
2012 Tracking Solutions of Time Varying Linear Inverse Problems
Martin Kleinsteuber, Simon Hawe
ICPRAM (1)1
2012 Blind Source Separation With Compressively Sensed Linear Mixtures
abstract
This work studies the problem of simultaneously separating and reconstructing signals from compressively sensed linear mixtures. We assume that all source signals share a common sparse representation basis. The approach combines classical Compressive Sensing (CS) theory with a linear mixing model. It allows the mixtures to be sampled independently of each other. If samples are acquired in the time domain, this means that the sensors need not be synchronized. Since Blind Source Separation (BSS) from a linear mixture is only possible up to permutation and scaling, factoring out these ambiguities leads to a minimization problem on the so-called oblique manifold. We develop a geometric conjugate subgradient method that scales to large systems for solving the problem. Numerical results demonstrate the promising performance of the proposed algorithm compared to several state of the art methods.
Martin Kleinsteuber, Hao Shen 0002
IEEE Signal Process. Lett.1
2011 Dense disparity maps from sparse disparity measurements
abstract
In this work we propose a method for estimating disparity maps from very few measurements. Based on the theory of Compressive Sensing, our algorithm accurately reconstructs disparity maps only using about 5% of the entire map. We propose a conjugate subgradient method for the arising optimization problem that is applicable to large scale systems and recovers the disparity map efficiently. Experiments are provided that show the effectiveness of the proposed approach and robust behavior under noisy conditions.
Simon Hawe, Martin Kleinsteuber, Klaus Diepold
ICCV2
2008 Local Convergence Analysis of FastICA and Related Algorithms
abstract
The FastICA algorithm is one of the most prominent methods to solve the problem of linear independent component analysis (ICA). Although there have been several attempts to prove local convergence properties of FastICA, rigorous analysis is still missing in the community. The major difficulty of analysis is because of the well-known sign-flipping phenomenon of FastICA, which causes the discontinuity of the corresponding FastICA map on the unit sphere. In this paper, by using the concept of principal fiber bundles, FastICA is proven to be locally quadratically convergent to a correct separation. Higher order local convergence properties of FastICA are also investigated in the framework of a scalar shift strategy. Moreover, as a parallelized version of FastICA, the so-called QR FastICA algorithm, which employs the QR decomposition (Gram-Schmidt orthonormalization process) instead of the polar decomposition, is shown to share similar local convergence properties with the original FastICA.
Hao Shen 0002, Martin Kleinsteuber, Knut Hüper
IEEE Trans. Neural Networks2
2007 An Intrinsic CG Algorithm for Computing Dominant Subspaces
abstract
In this paper, a conjugate gradient method on the complex Grabmann manifold is proposed that computes the k-principal components of a Hermitian (n × n)-matrix. The algorithm is at most of order O(n2k) and yields locally good convergence results.
Martin Kleinsteuber, Knut Hüper
ICASSP (4)1
2006 Spin Dynamics: A Paradigm for Time Optimal Control on Compact Lie Groups
Gunther Dirr, Uwe Helmke, Knut Hüper, Martin Kleinsteuber
J. Glob. Optim.4