Federico Raue

dblp:169/4880 · DBLP profile ↗
← Back
35ranked-venue papers
4as first author
21since 2021 · last 2026
0000-0002-8604-6207ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 28 · 4 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 10 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 A Study in Dataset Distillation for Image Super-Resolution
Tobias Dietz, Brian B. Moser, Tobias Christian Nauen, Federico Raue, Stanislav Frolov, Andreas Dengel 0001
ICPR (1)4
2026 HyperCore: Coreset Selection Under Noise via Hypersphere Models
Brian B. Moser, Arundhati S. Shanbhag, Tobias Nauen, Stanislav Frolov, Federico Raue, Joachim Folz, Andreas Dengel 0001
ICPR (1)5
2026 A Low-Resolution Image is Worth 1 ˟ 1 Words: Enabling Fine Image Super-Resolution with Transformers and TaylorShift
Sanath Budakegowdanadoddi Nagaraju, Brian B. Moser, Tobias Christian Nauen, Stanislav Frolov, Federico Raue, Andreas Dengel 0001
ICPR (1)5
2025 When 512×512 is Not Enough: Local Degradation-Aware Multi-Diffusion for Extreme Image Super-Resolution
abstract
Large-scale, pre-trained Text-to-Image (T2I) diffusion models have gained significant popularity in image synthesis and have shown unexpected potential in image Super-Resolution (SR). However, they are usually trained with a resolution limit of 512×512, making scaling beyond this resolution an unresolved but necessary challenge. To address this limitation, we propose a novel approach that enables them to generate 2K, 4K, and even 8K images without any additional training. Our method leverages MultiDiffusion, which distributes the generation across multiple diffusion paths, and local degradation-aware prompt extraction, which guides the T2I model according to its low-resolution input. As a result, we unlock higher resolutions, allowing T2I diffusion to be applied to image SR tasks without limitation on resolution.
Brian B. Moser, Stanislav Frolov, Tobias Christian Nauen, Federico Raue, Andreas Dengel 0001
ICIP4
2025 Distill the Best, Ignore the Rest: A Study in Latent Dataset Distillation on Core-Sets
abstract
Latent dataset distillation, which exploits pre-trained generative priors, has gained significant interest in recent years because it can be applied agnostic to any distillation algorithm and addresses two significant limitations of classical distillation algorithms: cross-architecture generalization and high-resolution synthesis. However, existing approaches typically distill from the entire dataset, potentially including non-beneficial samples. We introduce a novel "Prune First, Distill After" framework that systematically prunes datasets via loss-based sampling prior to latent distillation. By leveraging pruning before classical distillation techniques and generative priors, we create a representative coreset that leads to enhanced generalization for unseen architectures - a significant challenge of current distillation methods. More specifically, our proposed framework significantly boosts distilled quality, achieving up to a 5.2 percentage points accuracy increase even with substantial dataset pruning, i.e., removing 80% of the original dataset prior to distillation. Overall, our experimental results highlight the advantages of our easy-sample prioritization and cross-architecture robustness, paving the way for more effective and high-quality dataset distillation.
Brian B. Moser, Federico Raue, Tobias Christian Nauen, Stanislav Frolov, Andreas Dengel 0001
IJCNN2
2025 Unlocking Dataset Distillation with Diffusion Models
abstract
Dataset distillation seeks to condense datasets into smaller but highly representative synthetic samples. While diffusion models now lead all generative benchmarks, current distillation methods avoid them and rely instead on GANs or autoencoders, or, at best, sampling from a fixed diffusion prior. This trend arises because naive backpropagation through the long denoising chain leads to vanishing gradients, which prevents effective synthetic sample optimization. To address this limitation, we introduce Latent Dataset Distillation with Diffusion Models (LD3M), the first method to learn gradient-based distilled latents and class embeddings end-to-end through a pre-trained latent diffusion model. A linearly decaying skip connection, injected from the initial noisy state into every reverse step, preserves the gradient signal across dozens of timesteps without requiring diffusion weight fine-tuning. Across multiple ImageNet subsets at $128\times128$ and $256\times256$, LD3M improves downstream accuracy by up to 4.8 percentage points (1 IPC) and 4.2 points (10 IPC) over the prior state-of-the-art. The code for LD3M is provided at https://github.com/Brian-Moser/prune_and_distill.
Brian B. Moser, Federico Raue, Sebastian Palacio, Stanislav Frolov, Andreas Dengel 0001
NeurIPS2
2025 Dynamic Attention-Guided Diffusion for Image Super-Resolution
abstract
Diffusion models in image Super-Resolution (SR) treat all image regions uniformly, which risks compromising the overall image quality by potentially introducing artifacts during denoising of less-complex regions. To address this, we propose “You Only Diffuse Areas” (YODA), a dynamic attention-guided diffusion process for image SR. YODA selectively focuses on spatial regions defined by attention maps derived from the low-resolution images and the current de-noising time step. This time-dependent targeting enables a more efficient conversion to high-resolution outputs by focusing on areas that benefit the most from the iterative refinement process, i.e., detail-rich objects. We empirically validate YODA by extending leading diffusion-based methods SR3, DiffBIR, and SRDiff. Our experiments demonstrate new state-of-the-art performances in face and general SR tasks across PSNR, SSIM, and LPIPS metrics. As a side effect, we find that YODA reduces color shift issues and stabilizes training with small batches.
Brian B. Moser, Stanislav Frolov, Federico Raue, Sebastian Palacio, Andreas Dengel 0001
WACV3
2025 Which Transformer to Favor: A Comparative Analysis of Efficiency in Vision Transformers
abstract
Self-attention in Transformers comes with a high computational cost because of their quadratic computational complexity, but their effectiveness in addressing problems in language and vision has sparked extensive research aimed at enhancing their efficiency. However, diverse experimental conditions, spanning multiple input domains, prevent a fair comparison based solely on reported results, posing challenges for model selection. To address this gap in comparability, we perform a large-scale benchmark of more than 45 models for image classification, evaluating key efficiency aspects, including accuracy, speed, and memory usage. Our benchmark provides a standardized baseline for efficiency-oriented transformers. We analyze the results based on the Pareto front - the boundary of optimal models. Surprisingly, despite claims of other models being more efficient, ViT remains Pareto optimal across multiple metrics. We observe that hybrid attention-CNN models exhibit remarkable inference memory- and parameter-efficiency. Moreover, our benchmark shows that using a larger model in general is more efficient than using higher resolution images. Thanks to our holistic evaluation, we provide a centralized resource for practitioners and researchers, facilitating informed decisions when selecting or developing efficient transformers.11https://github.com/tobna/WhatTransformerToFavor
Tobias Christian Nauen, Sebastian Palacio, Federico Raue, Andreas Dengel 0001
WACV3
2025 Diffusion Models, Image Super-Resolution, and Everything: A Survey
abstract
Diffusion models (DMs) have disrupted the image super-resolution (SR) field and further closed the gap between image quality and human perceptual preferences. They are easy to train and can produce very high-quality samples that exceed the realism of those produced by previous generative methods. Despite their promising results, they also come with new challenges that need further research: high computational demands, comparability, lack of explainability, color shifts, and more. Unfortunately, entry into this field is overwhelming because of the abundance of publications. To address this, we provide a unified recount of the theoretical foundations underlying DMs applied to image SR and offer a detailed analysis that underscores the unique characteristics and methodologies within this domain, distinct from broader existing reviews in the field. This article articulates a cohesive understanding of DM principles and explores current research avenues, including alternative input domains, conditioning techniques, guidance mechanisms, corruption spaces, and zero-shot learning approaches. By offering a detailed examination of the evolution and current trends in image SR through the lens of DMs, this article sheds light on the existing challenges and charts potential future directions, aiming to inspire further innovation in this rapidly advancing area.
Brian B. Moser, Arundhati S. Shanbhag, Federico Raue, Stanislav Frolov, Sebastian Palacio, Andreas Dengel 0001
IEEE Trans. Neural Networks Learn. Syst.3
2024 A Study in Dataset Pruning for Image Super-Resolution
Brian B. Moser, Federico Raue, Andreas Dengel 0001
ICANN (2)2
2024 Waving Goodbye to Low-Res: A Diffusion-Wavelet Approach for Image Super-Resolution
abstract
Image Super-Resolution (SR) remains challenging, particularly in achieving high-quality details without extensive computational cost. Existing methods often struggle to balance the trade-off between image quality, especially in high-frequency details, and computational efficiency. In this paper, we present a novel Diffusion-Wavelet (DiWa) approach for bridging this gap. It leverages the strengths of diffusion models and discrete wavelet transformation. By enabling the diffusion model to operate in the frequency domain, our models effectively hallucinate highfrequency information for SR images on the wavelet spectrum, resulting in high-quality and detailed reconstructions in image space. Quantitatively, our method outperforms other state-ofthe-art diffusion-based SR methods, namely SR3 and SRDiff, regarding PSNR, SSIM, and LPIPS on both face (8x scaling) and general (4x scaling) SR benchmarks. Meanwhile, using the frequency domain allows us to use fewer parameters than the compared models: 92M parameters instead of 550M compared to SR3 and 9.3M instead of 12M compared to SRDiff. Additionally, DiWa outperforms other state-of-the-art generative methods on general SR datasets while saving inference time (ca. 250 %).
Brian B. Moser, Stanislav Frolov, Federico Raue, Sebastian Palacio, Andreas Dengel 0001
IJCNN3
2024 SphereCraft: A Dataset for Spherical Keypoint Detection, Matching and Camera Pose Estimation
abstract
This paper introduces SphereCraft, a dataset specifically designed for spherical keypoint detection, matching, and camera pose estimation. The dataset addresses the limitations of existing datasets by providing extracted keypoints from various detectors, along with their ground truth correspondences. Synthetic scenes with photo-realistic rendering and accurate 3D meshes are included, as well as real-world scenes acquired from different spherical cameras. SphereCraft enables the development and evaluation of algorithms targeting multiple camera viewpoints, advancing the state-of-the-art in computer vision tasks involving spherical images. Our dataset is available at https://dfki.github.io/spherecraftweb/.
Christiano Couto Gava, Yunmin Cho, Federico Raue, Sebastian Palacio, Alain Pagani, Andreas Dengel 0001
WACV3
2023 DWA: Differential Wavelet Amplifier for Image Super-Resolution
Brian B. Moser, Stanislav Frolov, Federico Raue, Sebastian Palacio, Andreas Dengel 0001
ICANN (2)3
2023 Sequential Spatial Transformer Networks for Salient Object Classification
David Dembinsky, Fatemeh Azimi, Federico Raue, Jörn Hees, Sebastian Palacio, Andreas Dengel 0001
ICPRAM3
2023 Cross-Domain Transformation for Outlier Detection on Tabular Datasets
abstract
The overwhelming success of Deep Learning approaches in recent years is often driven by the availability of large public datasets. However, in some domains like finance, creating and sharing realistic datasets is hindered by secrecy or privacy concerns. This can lead to a mismatch, where approaches that have proven to work well on public, research-oriented datasets end up underperforming when applied to real-world (private) datasets. In this work, we focus on the task of Outlier Detection (OD) and bridge the above gap by building an autoencoder based Deep Learning approach that can transform samples between two tabular datasets (e.g., a private and public one). The goal of our approach is that transformed samples become similar to the target dataset, while inliers remain inliers and outliers remain outliers. Among others, after successful transformation, this allows applying of proven methods on public datasets to internal datasets, even if they are of different dimensionality (rows and columns). To evaluate our approach, we introduce metrics to measure dataset similarity and the quality of transformed samples. Our experimental results show that combining public datasets with transformed samples of other datasets leads to higher dataset similarity while sustaining performance w.r.t. common OD algorithms.
Dayananda Herurkar, Timur Sattarov, Jörn Hees, Sebastian Palacio, Federico Raue, Andreas Dengel 0001
IJCNN5
2023 Hitchhiker's Guide to Super-Resolution: Introduction and Recent Advances
abstract
With the advent of Deep Learning (DL), Super-Resolution (SR) has also become a thriving research area. However, despite promising results, the field still faces challenges that require further research, e.g., allowing flexible upsampling, more effective loss functions, and better evaluation metrics. We review the domain of SR in light of recent advances and examine state-of-the-art models such as diffusion (DDPM) and transformer-based SR models. We critically discuss contemporary strategies used in SR and identify promising yet unexplored research directions. We complement previous surveys by incorporating the latest developments in the field, such as uncertainty-driven losses, wavelet networks, neural architecture search, novel normalization methods, and the latest evaluation techniques. We also include several visualizations for the models and methods throughout each chapter to facilitate a global understanding of the trends in the field. This review ultimately aims at helping researchers to push the boundaries of DL applied to SR.
Brian B. Moser, Federico Raue, Stanislav Frolov, Sebastian Palacio, Jörn Hees, Andreas Dengel 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Audioclip: Extending Clip to Image, Text and Audio
abstract
The rapidly evolving field of sound classification has greatly benefited from the methods of other domains. Today, the trend is to fuse domain-specific tasks and approaches together, which provides the community with new outstanding models.We present AudioCLIP – an extension of the CLIP model that handles audio in addition to text and images. Utilizing the AudioSet dataset, our proposed model incorporates the ESResNeXt audio-model into the CLIP framework, thus enabling it to perform multimodal classification and keeping CLIP’s zero-shot capabilities.AudioCLIP achieves new state-of-the-art results in the Environmental Sound Classification (ESC) task and out-performs others by reaching accuracies of 97.15 % on ESC-50 and 90.07 % on UrbanSound8K. Further, it sets new baselines in the zero-shot ESC-task on the same datasets (69.40 % and 68.78 %, respectively).We also asses the influence of different training setups on the final performance of the proposed model. For the sake of reproducibility, our code is published.
Andrey Guzhov, Federico Raue, Jörn Hees, Andreas Dengel 0001
ICASSP2
2022 Self-supervised Test-time Adaptation on Video Data
abstract
In typical computer vision problems revolving around video data, pre-trained models are simply evaluated at test time, without adaptation. This general approach clearly cannot capture the shifts that will likely arise between the distributions from which training and test data have been sampled. Adapting a pre-trained model to a new video en-countered at test time could be essential to avoid the potentially catastrophic effects of such shifts. However, given the inherent impossibility of labeling data only available at test-time, traditional "fine-tuning" techniques cannot be lever-aged in this highly practical scenario. This paper explores whether the recent progress in test-time adaptation in the image domain and self-supervised learning can be lever-aged to adapt a model to previously unseen and unlabelled videos presenting both mild (but arbitrary) and severe covariate shifts. In our experiments, we show that test-time adaptation approaches applied to self-supervised methods are always beneficial, but also that the extent of their effectiveness largely depends on the specific combination of the algorithms used for adaptation and self-supervision, and also on the type of covariate shift taking place.
Fatemeh Azimi, Sebastian Palacio, Federico Raue, Jörn Hees, Luca Bertinetto, Andreas Dengel 0001
WACV3
2021 ESResNe(X)t-fbsp: Learning Robust Time-Frequency Transformation of Audio
abstract
Environmental Sound Classification (ESC) is a rapidly evolving field that recently demonstrated the advantages of application of visual domain techniques to the audio-related tasks. Previous studies indicate that the domain-specific modification of cross-domain approaches show a promise in pushing the whole area of ESC forward. In this paper, we present a new time-frequency transformation layer that is based on complex frequency B-spline (fbsp) wavelets. Being used with a high-performance audio classification model, the proposed fbsp-layer provides an accuracy improvement over the previously used Short-Time Fourier Transform (STFT) on standard datasets. We also investigate the influence of different pre-training strategies, including the joint use of two large-scale datasets for weight initialization: ImageNet and AudioSet. Our proposed model out-performs other approaches by achieving accuracies of 95.20 % on the ESC-50 and 89.14 % on the UrbanSound8K datasets. Additionally, we assess the increase of model robustness against additive white Gaussian noise and reduction of an effective sample rate introduced by the proposed layer and demonstrate that the fbsp-layer improves the model's ability to withstand signal perturbations, in comparison to STFT-based training. For the sake of reproducibility, our code is made available.
Andrey Guzhov, Federico Raue, Jörn Hees, Andreas Dengel 0001
IJCNN2
2021 ANP-W2V: Effects of Composition Methods for Embedding Adjective-Noun Pairs
abstract
Adjective-Noun Pairs (ANPs) are often used for affective computing in textual and visual domains. Due to the training cost of more current models (e.g., transformers) many approaches still rely on Word2Vec to compute ANP embeddings by combining individual word embeddings. However, when combining an adjective and a noun into an ANP, there can potentially be more complex interactions which cannot be accounted for by using Word2Vec alone. To solve these challenges, we propose ANP- W2V, an approach that puts adjectives and nouns in different embedding spaces and hence outperforms the baselines based on Word2Vec. In this paper, we do a comprehensive comparison, where we systematically evaluate the role of six different fusion methods in four different tasks with different embedding sizes. The chosen tasks not only gauge the external and internal relationships in the ANPs but can also check whether ANP embeddings capture more complex interactions between adjectives and nouns.
Dayananda Herurkar, Philipp Blandfort, Federico Raue, Jörn Hees, Andreas Dengel 0001
IJCNN3
2021 Adversarial text-to-image synthesis: A review
abstract
With the advent of generative adversarial networks, synthesizing images from text descriptions has recently become an active research area. It is a flexible and intuitive way for conditional image generation with significant progress in the last years regarding visual realism, diversity, and semantic alignment. However, the field still faces several challenges that require further research efforts such as enabling the generation of high-resolution images with multiple objects, and developing suitable and reliable evaluation metrics that correlate with human judgement. In this review, we contextualize the state of the art of adversarial text-to-image synthesis models, their development since their inception five years ago, and propose a taxonomy based on the level of supervision. We critically examine current strategies to evaluate text-to-image synthesis models, highlight shortcomings, and identify new areas of research, ranging from the development of better datasets and evaluation metrics to possible improvements in architectural design and model training. This review complements previous surveys on generative adversarial networks with a focus on text-to-image synthesis which we believe will help researchers to further advance the field.
Stanislav Frolov, Tobias Hinz, Federico Raue, Jörn Hees, Andreas Dengel 0001
Neural Networks3
2020 DartsReNet: Exploring New RNN Cells in ReNet Architectures
Brian B. Moser, Federico Raue, Jörn Hees, Andreas Dengel 0001
ICANN (1)2
2020 Revisiting Sequence-to-Sequence Video Object Segmentation with Multi-Task Loss and Skip-Memory
abstract
Video Object Segmentation (VOS) is an active research area of the visual domain. One of its fundamental subtasks is semi-supervised / one-shot learning: given only the segmentation mask for the first frame, the task is to provide pixel-accurate masks for the object over the rest of the sequence. Despite much progress in the last years, we noticed that many of the existing approaches lose objects in longer sequences, especially when the object is small or briefly occluded. In this work, we build upon a sequence-to-sequence approach that employs an encoder-decoder architecture together with a memory module for exploiting the sequential data. We further improve this approach by proposing a model that manipulates multiscale spatio-temporal information using memory-equipped skip connections. Furthermore, we incorporate an auxiliary task based on distance classification which greatly enhances the quality of edges in segmentation masks. We compare our approach to the state of the art and show considerable improvement in the contour accuracy metric and the overall segmentation accuracy. Our source code and the pre-trained weights are publicly available11https://github.com/fatemehazimi990/RS2S.
Fatemeh Azimi, Benjamin Bischke, Sebastian Palacio, Federico Raue, Jörn Hees, Andreas Dengel 0001
ICPR4
2020 ESResNet: Environmental Sound Classification Based on Visual Domain Models
abstract
Environmental Sound Classification (ESC) is an active research area in the audio domain and has seen a lot of progress in the past years. However, many of the existing approaches achieve high accuracy by relying on domain-specific features and architectures, making it harder to benefit from advances in other fields (e.g., the image domain). Additionally, some of the past successes have been attributed to a discrepancy of how results are evaluated (i.e., on unofficial splits of the UrbanSound8K (US8K) dataset), distorting the overall progression of the field. The contribution of this paper is twofold. First, we present a model that is inherently compatible with mono and stereo sound inputs. Our model is based on simple log-power Short-Time Fourier Transform (STFT) spectrograms and combines them with several well-known approaches from the image domain (i.e., ResNet, Siamese-like networks and attention). We investigate the influence of cross-domain pre-training, architectural changes, and evaluate our model on standard datasets. We find that our model out-performs all previously known approaches in a fair comparison by achieving accuracies of 97.0 % (ESC-10), 91.5 % (ESC-50) and 84.2 % / 85.4 % (US8K mono / stereo). Second, we provide a comprehensive overview of the actual state of the field, by differentiating several previously reported results on the US8K dataset between official or unofficial splits. For better reproducibility, our code (including any re- implementations) is made available.
Andrey Guzhov, Federico Raue, Jörn Hees, Andreas Dengel 0001
ICPR2
2020 P ≈ NP, at least in Visual Question Answering
abstract
In recent years, progress in the Visual Question Answering (VQA) field has largely been driven by public challenges and large datasets. One of the most widely-used of these is the VQA 2.0 dataset, consisting of polar (“yes/no”) and non-polar questions. Looking at the question distribution over all answers, we find that the answers “yes” and “no” account for 38% of the questions (19% per class), while the remaining 62% are spread over the remaining 3127 answers (0.02% per class). While several sources of biases have been investigated in the field, the effects of such an over-representation of polar questions remain unclear. In this paper, we measure the potential confounding factors when polar and non-polar samples are used jointly to train a baseline VQA classifier, and compare it to an upper bound where the over-representation of polar questions is excluded from the training. Further, we perform cross-over experiments to analyze how well the feature spaces of polar and non-polar samples align. Contrary to expectations, we find no evidence of counterproductive effects in the joint training of unbalanced classes. In fact, by exploring the intermediate feature space of visual-text embeddings, we find that the feature space of polar questions already encodes sufficient structure to answer many non-polar questions. Our results indicate that the polar (P) and the non-polar (NP) feature spaces are strongly aligned, hence the expression P ≈ NP.
Shailza Jolly, Sebastian Palacio, Joachim Folz, Federico Raue, Jörn Hees, Andreas Dengel 0001
ICPR4
2019 A Reinforcement Learning Approach for Sequential Spatial Transformer Networks
Fatemeh Azimi, Federico Raue, Jörn Hees, Andreas Dengel 0001
ICANN (1)2
2019 Conditional GANs for Image Captioning with Sentiments
Tushar Karayil, Asif Irfan, Federico Raue, Jörn Hees, Andreas Dengel 0001
ICANN (4)3
2019 Comparison Between U-Net and U-ReNet Models in OCR Tasks
Brian B. Moser, Federico Raue, Jörn Hees, Andreas Dengel 0001
ICANN (3)2
2019 Fusion Strategies for Learning User Embeddings with Neural Networks
abstract
Growing amounts of online user data motivate the need for automated processing techniques. In case of user ratings, one interesting option is to use neural networks for learning to predict ratings given an item and a user. While training for prediction, such an approach at the same time learns to map each user to a vector, a so-called user embedding. Such embeddings can for example be valuable for estimating user similarity. However, there are various ways how item and user information can be combined in neural networks, and it is unclear how the way of combining affects the resulting embeddings.In this paper, we run an experiment on movie ratings data, where we analyze the effect on embedding quality caused by several fusion strategies in neural networks. For evaluating embedding quality, we propose a novel measure, Pair-Distance Correlation, which quantifies the condition that similar users should have similar embedding vectors. We find that the fusion strategy affects results in terms of both prediction performance and embedding quality. Surprisingly, we find that prediction performance not necessarily reflects embedding quality. This suggests that if embeddings are of interest, the common tendency to select models based on their prediction ability should be reconsidered.
Philipp Blandfort, Tushar Karayil, Federico Raue, Jörn Hees, Andreas Dengel 0001
IJCNN3
2018 What Do Deep Networks Like to See?
abstract
We propose a novel way to measure and understand convolutional neural networks by quantifying the amount of input signal they let in. To do this, an autoencoder (AE) was fine-tuned on gradients from a pre-trained classifier with fixed parameters. We compared the reconstructed samples from AEs that were fine-tuned on a set of image classifiers (AlexNet, VGG16, ResNet-50, and Inception v3) and found substantial differences. The AE learns which aspects of the input space to preserve and which ones to ignore, based on the information encoded in the backpropagated gradients. Measuring the changes in accuracy when the signal of one classifier is used by a second one, a relation of total order emerges. This order depends directly on each classifier's input signal but it does not correlate with classification accuracy or network size. Further evidence of this phenomenon is provided by measuring the normalized mutual information between original images and auto-encoded reconstructions from different fine-tuned AEs. These findings break new ground in the area of neural network understanding, opening a new way to reason, debug, and interpret their results. We present four concrete examples in the literature where observations can now be explained in terms of the input signal that a model uses.
Sebastian Palacio, Joachim Folz, Jörn Hees, Federico Raue, Damian Borth, Andreas Dengel 0001
CVPR4
2018 Symbol Grounding Association in Multimodal Sequences with Missing Elements
abstract
In this paper, we extend a symbolic association framework for being able to handle missing elements in multimodal sequences. The general scope of the work is the symbolic associations of object-word mappings as it happens in language development in infants. In other words, two different representations of the same abstract concepts can associate in both directions. This scenario has been long interested in Artificial Intelligence, Psychology, and Neuroscience. In this work, we extend a recent approach for multimodal sequences (visual and audio) to also cope with missing elements in one or both modalities. Our method uses two parallel Long Short-Term Memories (LSTMs) with a learning rule based on EM-algorithm. It aligns both LSTM outputs via Dynamic Time Warping (DTW). We propose to include an extra step for the combination with the max operation for exploiting the common elements between both sequences. The motivation behind is that the combination acts as a condition selector for choosing the best representation from both LSTMs. We evaluated the proposed extension in the following scenarios: missing elements in one modality (visual or audio) and missing elements in both modalities (visual and sound). The performance of our extension reaches better results than the original model and similar results to individual LSTM trained in each modality.
Federico Raue, Andreas Dengel 0001, Thomas M. Breuel, Marcus Liwicki
J. Artif. Intell. Res.1
2017 Classless Association Using Neural Networks
Federico Raue, Sebastian Palacio, Andreas Dengel 0001, Marcus Liwicki
ICANN (2)1
2016 Symbolic Association Using Parallel Multilayer Perceptron
Federico Raue, Sebastian Palacio, Thomas M. Breuel, Wonmin Byeon, Andreas Dengel 0001, Marcus Liwicki
ICANN (2)1
2015 Scene labeling with LSTM recurrent neural networks
abstract
This paper addresses the problem of pixel-level segmentation and classification of scene images with an entirely learning-based approach using Long Short Term Memory (LSTM) recurrent neural networks, which are commonly used for sequence classification. We investigate two-dimensional (2D) LSTM networks for natural scene images taking into account the complex spatial dependencies of labels. Prior methods generally have required separate classification and image segmentation stages and/or pre- and post-processing. In our approach, classification, segmentation, and context integration are all carried out by 2D LSTM networks, allowing texture and spatial model parameters to be learned within a single model. The networks efficiently capture local and global contextual information over raw RGB values and adapt well for complex scene images. Our approach, which has a much lower computational complexity than prior methods, achieved state-of-the-art performance over the Stanford Background and the SIFT Flow datasets. In fact, if no pre- or post-processing is applied, LSTM networks outperform other state-of-the-art approaches. Hence, only with a single-core Central Processing Unit (CPU), the running time of our approach is equivalent or better than the compared state-of-the-art approaches which use a Graphics Processing Unit (GPU). Finally, our networks' ability to visualize feature maps from each layer supports the hypothesis that LSTM networks are overall suited for image processing tasks.
Wonmin Byeon, Thomas M. Breuel, Federico Raue, Marcus Liwicki
CVPR3
2015 Parallel sequence classification using recurrent neural networks and alignment
abstract
The aim of this work is to investigate Long Short-Term Memory (LSTM) for finding the semantic associations between two parallel text lines of different instances of the same class sequence. In this work, we propose a new model called class-less classifier, which is cognitive motivated by a simplified version of the infants learning. The presented model not only learns the semantic association but also learns the relation between the labels and the classes. In addition, our model uses two parallel class-less LSTM networks and the learning rule is based on the alignment of both networks. For testing purposes, a parallel sequence dataset is generated based on MNIST dataset, which is a standard dataset for handwritten digit recognition. The results of our model were similar to the standard LSTM.
Federico Raue, Wonmin Byeon, Thomas M. Breuel, Marcus Liwicki
ICDAR1