VLDB 2026 Research / reviewers in the wild / expert
Sebastian Palacio
dblp:166/0222
· DBLP profile ↗
22ranked-venue papers
2as first author
15since 2021 · last 2025
0000-0002-8656-9569ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 7 since 2021Systems, architecture and hardware · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DriveAIAgent: A Multi-Agent System for Industrial Drive Commissioning and TroubleshootingabstractCommissioning is a critical phase in the lifecycle of an industrial drive system, significantly affecting overall performance and reliability. It involves configuring drives and motors for specific applications—such as mixing, pumping, or operating conveyors—and often requires managing complex interdependencies. Traditionally, commissioning is performed manually by domain experts through multiple steps, requiring them to navigate various resources ranging from drive datasheets to vendor-specific tools for drive parameter adjustment. This manual process is time-consuming, and errors can cause significant delays, sometimes lasting several days. In this paper, we propose DriveAIAgent, a novel system architecture and methodology that integrates generative AI-powered multi-agent systems to automate and enhance the commissioning and troubleshooting of industrial drives. Our agentic system utilizes expert tools to interact with and respond to external environments, incorporating a human-in-the-loop approach. The evaluation shows that our system achieves 100% accuracy in identifying encoder parameters during the initial commissioning phase, as well as very high solution relevance (94.73%) and correctness (86.29%) for troubleshooting across different drive types. Overall, DriveAIAgent can streamline complex interactions and significantly reduce the time required for commissioning and troubleshooting industrial drive systems while maintaining the same level of quality as a human expert. Virendra Ashiwal, Marcus Ritter, Sebastian Palacio, Nicolai Schoch |
ETFA | 3 |
| 2025 | Automatic Validation of Unstructured LLM-Generated Outputs: An Approach for Q&A Applications in Process AutomationabstractLarge language models (LLMs) have gained significant attention for their success in natural language generation and their widespread adoption across various business sectors. In process automation (PA) engineering, LLMs can be utilized to assist with tasks such as PLC programming, software development, and data processing. Despite their advantages, LLMs face challenges like biased responses, non-deterministic outputs and model hallucinations. Therefore, it is essential to validate the responses generated by LLMs before relying on them. This work presents a novel workflow to automatically validate unstructured outputs generated by LLMs used in question answering (Q&A) applications based on the user-specified documents in the PA domain. The Domain of Validity (DoV) of these LLMs is first estimated based on the user-specified documents. The proposed workflow uses the estimated DoV to evaluate the relevance of the outputs generated by the LLM with respect to the user documents. A regular expression approach is included in the workflow to check whether the mentioned tag names in the LLM output, if any, exist in the user-specified documents. Furthermore, an LLM-as-a-judge approach is used then to cross-check the LLM generated outputs with the ground truth information. The proposed workflow enables users to automatically detect invalid LLM outputs. This allows the user to focus on cases that truly require human judgment and act accordingly by re-prompting the LLM, adding documents, or switching models. A fictitious PA example is employed to demonstrate the versatility and benefits of the developed workflow. Mohamed Elsheikh, Nicolai Schoch, Sebastian Palacio, Nika Strem, Katharina Stark, Mario Hoernicke |
ETFA | 3 |
| 2025 | Unlocking Dataset Distillation with Diffusion ModelsabstractDataset distillation seeks to condense datasets into smaller but highly representative synthetic samples. While diffusion models now lead all generative benchmarks, current distillation methods avoid them and rely instead on GANs or autoencoders, or, at best, sampling from a fixed diffusion prior. This trend arises because naive backpropagation through the long denoising chain leads to vanishing gradients, which prevents effective synthetic sample optimization. To address this limitation, we introduce Latent Dataset Distillation with Diffusion Models (LD3M), the first method to learn gradient-based distilled latents and class embeddings end-to-end through a pre-trained latent diffusion model. A linearly decaying skip connection, injected from the initial noisy state into every reverse step, preserves the gradient signal across dozens of timesteps without requiring diffusion weight fine-tuning. Across multiple ImageNet subsets at $128\times128$ and $256\times256$, LD3M improves downstream accuracy by up to 4.8 percentage points (1 IPC) and 4.2 points (10 IPC) over the prior state-of-the-art. The code for LD3M is provided at https://github.com/Brian-Moser/prune_and_distill. Brian B. Moser, Federico Raue, Sebastian Palacio, Stanislav Frolov, Andreas Dengel 0001 |
NeurIPS | 3 |
| 2025 | Dynamic Attention-Guided Diffusion for Image Super-ResolutionabstractDiffusion models in image Super-Resolution (SR) treat all image regions uniformly, which risks compromising the overall image quality by potentially introducing artifacts during denoising of less-complex regions. To address this, we propose “You Only Diffuse Areas” (YODA), a dynamic attention-guided diffusion process for image SR. YODA selectively focuses on spatial regions defined by attention maps derived from the low-resolution images and the current de-noising time step. This time-dependent targeting enables a more efficient conversion to high-resolution outputs by focusing on areas that benefit the most from the iterative refinement process, i.e., detail-rich objects. We empirically validate YODA by extending leading diffusion-based methods SR3, DiffBIR, and SRDiff. Our experiments demonstrate new state-of-the-art performances in face and general SR tasks across PSNR, SSIM, and LPIPS metrics. As a side effect, we find that YODA reduces color shift issues and stabilizes training with small batches. Brian B. Moser, Stanislav Frolov, Federico Raue, Sebastian Palacio, Andreas Dengel 0001 |
WACV | 4 |
| 2025 | Which Transformer to Favor: A Comparative Analysis of Efficiency in Vision TransformersabstractSelf-attention in Transformers comes with a high computational cost because of their quadratic computational complexity, but their effectiveness in addressing problems in language and vision has sparked extensive research aimed at enhancing their efficiency. However, diverse experimental conditions, spanning multiple input domains, prevent a fair comparison based solely on reported results, posing challenges for model selection. To address this gap in comparability, we perform a large-scale benchmark of more than 45 models for image classification, evaluating key efficiency aspects, including accuracy, speed, and memory usage. Our benchmark provides a standardized baseline for efficiency-oriented transformers. We analyze the results based on the Pareto front - the boundary of optimal models. Surprisingly, despite claims of other models being more efficient, ViT remains Pareto optimal across multiple metrics. We observe that hybrid attention-CNN models exhibit remarkable inference memory- and parameter-efficiency. Moreover, our benchmark shows that using a larger model in general is more efficient than using higher resolution images. Thanks to our holistic evaluation, we provide a centralized resource for practitioners and researchers, facilitating informed decisions when selecting or developing efficient transformers.11https://github.com/tobna/WhatTransformerToFavor Tobias Christian Nauen, Sebastian Palacio, Federico Raue, Andreas Dengel 0001 |
WACV | 2 |
| 2025 | Diffusion Models, Image Super-Resolution, and Everything: A SurveyabstractDiffusion models (DMs) have disrupted the image super-resolution (SR) field and further closed the gap between image quality and human perceptual preferences. They are easy to train and can produce very high-quality samples that exceed the realism of those produced by previous generative methods. Despite their promising results, they also come with new challenges that need further research: high computational demands, comparability, lack of explainability, color shifts, and more. Unfortunately, entry into this field is overwhelming because of the abundance of publications. To address this, we provide a unified recount of the theoretical foundations underlying DMs applied to image SR and offer a detailed analysis that underscores the unique characteristics and methodologies within this domain, distinct from broader existing reviews in the field. This article articulates a cohesive understanding of DM principles and explores current research avenues, including alternative input domains, conditioning techniques, guidance mechanisms, corruption spaces, and zero-shot learning approaches. By offering a detailed examination of the evolution and current trends in image SR through the lens of DMs, this article sheds light on the existing challenges and charts potential future directions, aiming to inspire further innovation in this rapidly advancing area. Brian B. Moser, Arundhati S. Shanbhag, Federico Raue, Stanislav Frolov, Sebastian Palacio, Andreas Dengel 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | TaylorShift: Shifting the Complexity of Self-attention from Squared to Linear (and Back) Using Taylor-Softmax
Tobias Christian Nauen, Sebastian Palacio, Andreas Dengel 0001 |
ICPR (6) | 2 |
| 2024 | Waving Goodbye to Low-Res: A Diffusion-Wavelet Approach for Image Super-ResolutionabstractImage Super-Resolution (SR) remains challenging, particularly in achieving high-quality details without extensive computational cost. Existing methods often struggle to balance the trade-off between image quality, especially in high-frequency details, and computational efficiency. In this paper, we present a novel Diffusion-Wavelet (DiWa) approach for bridging this gap. It leverages the strengths of diffusion models and discrete wavelet transformation. By enabling the diffusion model to operate in the frequency domain, our models effectively hallucinate highfrequency information for SR images on the wavelet spectrum, resulting in high-quality and detailed reconstructions in image space. Quantitatively, our method outperforms other state-ofthe-art diffusion-based SR methods, namely SR3 and SRDiff, regarding PSNR, SSIM, and LPIPS on both face (8x scaling) and general (4x scaling) SR benchmarks. Meanwhile, using the frequency domain allows us to use fewer parameters than the compared models: 92M parameters instead of 550M compared to SR3 and 9.3M instead of 12M compared to SRDiff. Additionally, DiWa outperforms other state-of-the-art generative methods on general SR datasets while saving inference time (ca. 250 %). Brian B. Moser, Stanislav Frolov, Federico Raue, Sebastian Palacio, Andreas Dengel 0001 |
IJCNN | 4 |
| 2024 | ObjBlur: A Curriculum Learning Approach With Progressive Object-Level Blurring for Improved Layout-to-Image Generation
Stanislav Frolov, Brian B. Moser, Sebastian Palacio, Andreas Dengel 0001 |
ACM Multimedia | 3 |
| 2024 | SphereCraft: A Dataset for Spherical Keypoint Detection, Matching and Camera Pose EstimationabstractThis paper introduces SphereCraft, a dataset specifically designed for spherical keypoint detection, matching, and camera pose estimation. The dataset addresses the limitations of existing datasets by providing extracted keypoints from various detectors, along with their ground truth correspondences. Synthetic scenes with photo-realistic rendering and accurate 3D meshes are included, as well as real-world scenes acquired from different spherical cameras. SphereCraft enables the development and evaluation of algorithms targeting multiple camera viewpoints, advancing the state-of-the-art in computer vision tasks involving spherical images. Our dataset is available at https://dfki.github.io/spherecraftweb/. Christiano Couto Gava, Yunmin Cho, Federico Raue, Sebastian Palacio, Alain Pagani, Andreas Dengel 0001 |
WACV | 4 |
| 2023 | DWA: Differential Wavelet Amplifier for Image Super-Resolution
Brian B. Moser, Stanislav Frolov, Federico Raue, Sebastian Palacio, Andreas Dengel 0001 |
ICANN (2) | 4 |
| 2023 | Sequential Spatial Transformer Networks for Salient Object Classification
David Dembinsky, Fatemeh Azimi, Federico Raue, Jörn Hees, Sebastian Palacio, Andreas Dengel 0001 |
ICPRAM | 5 |
| 2023 | Cross-Domain Transformation for Outlier Detection on Tabular DatasetsabstractThe overwhelming success of Deep Learning approaches in recent years is often driven by the availability of large public datasets. However, in some domains like finance, creating and sharing realistic datasets is hindered by secrecy or privacy concerns. This can lead to a mismatch, where approaches that have proven to work well on public, research-oriented datasets end up underperforming when applied to real-world (private) datasets. In this work, we focus on the task of Outlier Detection (OD) and bridge the above gap by building an autoencoder based Deep Learning approach that can transform samples between two tabular datasets (e.g., a private and public one). The goal of our approach is that transformed samples become similar to the target dataset, while inliers remain inliers and outliers remain outliers. Among others, after successful transformation, this allows applying of proven methods on public datasets to internal datasets, even if they are of different dimensionality (rows and columns). To evaluate our approach, we introduce metrics to measure dataset similarity and the quality of transformed samples. Our experimental results show that combining public datasets with transformed samples of other datasets leads to higher dataset similarity while sustaining performance w.r.t. common OD algorithms. Dayananda Herurkar, Timur Sattarov, Jörn Hees, Sebastian Palacio, Federico Raue, Andreas Dengel 0001 |
IJCNN | 4 |
| 2023 | Hitchhiker's Guide to Super-Resolution: Introduction and Recent AdvancesabstractWith the advent of Deep Learning (DL), Super-Resolution (SR) has also become a thriving research area. However, despite promising results, the field still faces challenges that require further research, e.g., allowing flexible upsampling, more effective loss functions, and better evaluation metrics. We review the domain of SR in light of recent advances and examine state-of-the-art models such as diffusion (DDPM) and transformer-based SR models. We critically discuss contemporary strategies used in SR and identify promising yet unexplored research directions. We complement previous surveys by incorporating the latest developments in the field, such as uncertainty-driven losses, wavelet networks, neural architecture search, novel normalization methods, and the latest evaluation techniques. We also include several visualizations for the models and methods throughout each chapter to facilitate a global understanding of the trends in the field. This review ultimately aims at helping researchers to push the boundaries of DL applied to SR. Brian B. Moser, Federico Raue, Stanislav Frolov, Sebastian Palacio, Jörn Hees, Andreas Dengel 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Self-supervised Test-time Adaptation on Video DataabstractIn typical computer vision problems revolving around video data, pre-trained models are simply evaluated at test time, without adaptation. This general approach clearly cannot capture the shifts that will likely arise between the distributions from which training and test data have been sampled. Adapting a pre-trained model to a new video en-countered at test time could be essential to avoid the potentially catastrophic effects of such shifts. However, given the inherent impossibility of labeling data only available at test-time, traditional "fine-tuning" techniques cannot be lever-aged in this highly practical scenario. This paper explores whether the recent progress in test-time adaptation in the image domain and self-supervised learning can be lever-aged to adapt a model to previously unseen and unlabelled videos presenting both mild (but arbitrary) and severe covariate shifts. In our experiments, we show that test-time adaptation approaches applied to self-supervised methods are always beneficial, but also that the extent of their effectiveness largely depends on the specific combination of the algorithms used for adaptation and self-supervision, and also on the type of covariate shift taking place. Fatemeh Azimi, Sebastian Palacio, Federico Raue, Jörn Hees, Luca Bertinetto, Andreas Dengel 0001 |
WACV | 2 |
| 2020 | Revisiting Sequence-to-Sequence Video Object Segmentation with Multi-Task Loss and Skip-MemoryabstractVideo Object Segmentation (VOS) is an active research area of the visual domain. One of its fundamental subtasks is semi-supervised / one-shot learning: given only the segmentation mask for the first frame, the task is to provide pixel-accurate masks for the object over the rest of the sequence. Despite much progress in the last years, we noticed that many of the existing approaches lose objects in longer sequences, especially when the object is small or briefly occluded. In this work, we build upon a sequence-to-sequence approach that employs an encoder-decoder architecture together with a memory module for exploiting the sequential data. We further improve this approach by proposing a model that manipulates multiscale spatio-temporal information using memory-equipped skip connections. Furthermore, we incorporate an auxiliary task based on distance classification which greatly enhances the quality of edges in segmentation masks. We compare our approach to the state of the art and show considerable improvement in the contour accuracy metric and the overall segmentation accuracy. Our source code and the pre-trained weights are publicly available11https://github.com/fatemehazimi990/RS2S. Fatemeh Azimi, Benjamin Bischke, Sebastian Palacio, Federico Raue, Jörn Hees, Andreas Dengel 0001 |
ICPR | 3 |
| 2020 | P ≈ NP, at least in Visual Question AnsweringabstractIn recent years, progress in the Visual Question Answering (VQA) field has largely been driven by public challenges and large datasets. One of the most widely-used of these is the VQA 2.0 dataset, consisting of polar (“yes/no”) and non-polar questions. Looking at the question distribution over all answers, we find that the answers “yes” and “no” account for 38% of the questions (19% per class), while the remaining 62% are spread over the remaining 3127 answers (0.02% per class). While several sources of biases have been investigated in the field, the effects of such an over-representation of polar questions remain unclear. In this paper, we measure the potential confounding factors when polar and non-polar samples are used jointly to train a baseline VQA classifier, and compare it to an upper bound where the over-representation of polar questions is excluded from the training. Further, we perform cross-over experiments to analyze how well the feature spaces of polar and non-polar samples align. Contrary to expectations, we find no evidence of counterproductive effects in the joint training of unbalanced classes. In fact, by exploring the intermediate feature space of visual-text embeddings, we find that the feature space of polar questions already encodes sufficient structure to answer many non-polar questions. Our results indicate that the polar (P) and the non-polar (NP) feature spaces are strongly aligned, hence the expression P ≈ NP. Shailza Jolly, Sebastian Palacio, Joachim Folz, Federico Raue, Jörn Hees, Andreas Dengel 0001 |
ICPR | 2 |
| 2020 | Contextual Classification Using Self-Supervised Auxiliary Models for Deep Neural NetworksabstractThe following topics are dealt with: learning (artificial intelligence); feature extraction; neural nets; image classification; object detection; image segmentation; image representation; video signal processing; pattern classification; computer vision. Sebastian Palacio, Philipp Engler, Jörn Hees, Andreas Dengel 0001 |
ICPR | 1 |
| 2020 | Adversarial Defense based on Structure-to-Signal AutoencodersabstractAdversarial attacks have exposed the intricacies of the complex loss surfaces approximated by neural networks. In this paper, we present a defense strategy against gradient-based attacks, on the premise that input gradients need to expose information about the semantic manifold for attacks to be successful. We propose an architecture based on compressive autoencoders (AEs) with a two-stage training scheme, creating not only an architectural bottleneck but also a representational bottleneck. We show that the proposed mechanism yields robust results against a collection of gradient-based attacks under challenging white-box conditions. This defense is attack-agnostic and can, therefore, be used for arbitrary pre-trained models, while not compromising the original performance. These claims are supported by experiments conducted with state-of-the-art image classifiers (ResNet50 and Inception v3), on the full ImageNet validation set. Experiments, including counterfactual analysis, empirically show that the robustness stems from a shift in the distribution of input gradients, which mitigates the effect of tested adversarial attack methods. Gradients propagated through the proposed AEs represent less semantic information and instead point to low-level structural features. Joachim Folz, Sebastian Palacio, Jörn Hees, Andreas Dengel 0001 |
WACV | 2 |
| 2018 | What Do Deep Networks Like to See?abstractWe propose a novel way to measure and understand convolutional neural networks by quantifying the amount of input signal they let in. To do this, an autoencoder (AE) was fine-tuned on gradients from a pre-trained classifier with fixed parameters. We compared the reconstructed samples from AEs that were fine-tuned on a set of image classifiers (AlexNet, VGG16, ResNet-50, and Inception v3) and found substantial differences. The AE learns which aspects of the input space to preserve and which ones to ignore, based on the information encoded in the backpropagated gradients. Measuring the changes in accuracy when the signal of one classifier is used by a second one, a relation of total order emerges. This order depends directly on each classifier's input signal but it does not correlate with classification accuracy or network size. Further evidence of this phenomenon is provided by measuring the normalized mutual information between original images and auto-encoded reconstructions from different fine-tuned AEs. These findings break new ground in the area of neural network understanding, opening a new way to reason, debug, and interpret their results. We present four concrete examples in the literature where observations can now be explained in terms of the input signal that a model uses. Sebastian Palacio, Joachim Folz, Jörn Hees, Federico Raue, Damian Borth, Andreas Dengel 0001 |
CVPR | 1 |
| 2017 | Classless Association Using Neural Networks
Federico Raue, Sebastian Palacio, Andreas Dengel 0001, Marcus Liwicki |
ICANN (2) | 2 |
| 2016 | Symbolic Association Using Parallel Multilayer Perceptron
Federico Raue, Sebastian Palacio, Thomas M. Breuel, Wonmin Byeon, Andreas Dengel 0001, Marcus Liwicki |
ICANN (2) | 2 |