EDBT 2026 Demo / reviewers in the wild / expert
Vasileios Belagiannis
dblp:75/7627
· DBLP profile ↗
51ranked-venue papers
6as first author
35since 2021 · last 2026
0000-0003-0960-8453ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 35 · 6 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 4 first-author · 16 since 2021Systems, architecture and hardware · 9 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LSA: Localized Semantic Alignment for Enhancing Temporal Consistency in Traffic Video Generation
Mirlan Karimov, Teodora Spasojevic, Markus Braun 0003, Julian Wiederer, Vasileios Belagiannis, Marc Pollefeys |
IV | 5 |
| 2026 | GroupEnsemble: Efficient Uncertainty Estimation for DETR-based Object Detection
Yutong Yang, Katarina Popovic, Julian Wiederer, Markus Braun 0003, Vasileios Belagiannis, Bin Yang 0009 |
IV | 5 |
| 2025 | AGENTS-LLM: Augmentative GENeration of Challenging Traffic Scenarios with an Agentic LLM FrameworkabstractRare, yet critical, scenarios pose a significant challenge in testing and evaluating autonomous driving planners. Relying solely on real-world driving scenes requires collecting massive datasets to capture these scenarios. While automatic generation of traffic scenarios appears promising, data-driven models require extensive training data and often lack fine-grained control over the output. Moreover, generating novel scenarios from scratch can introduce a distributional shift from the original training scenes which undermines the validity of evaluations especially for learning-based planners. To sidestep this, recent work proposes to generate challenging scenarios by augmenting original scenarios from the test set. However, this involves the manual augmentation of scenarios by domain experts. An approach that is unable to meet the demands for scale in the evaluation of self-driving systems. Therefore, this paper introduces a novel LLM-agent based framework for augmenting real-world traffic scenarios using natural language descriptions, addressing the limitations of existing methods. A key innovation is the use of an agentic design, enabling fine-grained control over the output and maintaining high performance even with smaller, cost-effective LLMs. Extensive human expert evaluation demonstrates our framework’s ability to accurately adhere to user intent, generating high quality augmented scenarios comparable to those created manually. Salil Bhatnagar, Markus Mazzola, Vasileios Belagiannis, Igor Gilitschenski, Luigi Palmieri, Simon Razniewski, Marcel Hallgarten |
IROS | 4 |
| 2025 | PSumSim: A Simulator for Partial-Sum Quantization in Analog Matrix-Vector MultipliersabstractAs AI and its applications evolve, efficient hardware is required to run the novel algorithms. Compute platforms with a high degree of parallelism, such as matrix-vector multipliers, meet the need to process large homogeneous loads of operations. However, most of the matrix-vector multiplications required by the AI algorithms are larger than what the actual hardware supports. The operations must therefore be tiled into blocks that fit on the given hardware. Finally, the partial sums generated by the hardware for each tile must be accumulated or concatenated into the complete result. Especially with mixed-signal compute-in-memory architectures, this can lead to quantization on two levels. First, an ADC quantizes the partial sums generated for each tile. Then, the algorithm performs another quantization of the final result to limit bitwidth and resource consumption in adjacent computations. While quantizing only the partial sums or only the final results has been studied extensively, the combination of the two has yet to be investigated. This work introduces a simulator to understand the effects of quantization caused by multiple quantization steps on different levels. It is based on a generic, stochastic representation of value probabilities using histograms. Common operations such as scaling, rounding, and accumulation are implemented in this representation, allowing the effect of quantization to be studied in a matrix-vector-multiplier application. It is shown, that the selection of tilesize, ADC bitwidth and clipping technique form a complex trade-off, which can be solved using PSumSim. PSumSim is available under https://github.com/Joschua-Conrad/PSumSim. Joschua Conrad, Simon Wilhelmstätter, Holger Mandry, Paul Kässer, Ahmed Abdelaal, Rohan Asthana, Vasileios Belagiannis, Maurits Ortmanns |
ISCAS | 7 |
| 2025 | Diffusion Model Guided Sampling with Pixel-Wise Aleatoric Uncertainty EstimationabstractDespite the remarkable progress in generative modelling, current diffusion models lack a quantitative approach to assess image quality. To address this limitation, we propose to estimate the pixel-wise aleatoric uncertainty during the sampling phase of diffusion models and utilise the uncertainty to improve the sample generation quality. The uncertainty is computed as the variance of the denoising scores with a perturbation scheme that is specifically designed for diffusion models. We then show that the aleatoric uncertainty estimates are related to the second-order derivative of the diffusion noise distribution. We evaluate our uncertainty estimation algorithm and the uncertainty-guided sampling on the ImageNet and CIFAR-10 datasets. In our comparisons with the related work, we demonstrate promising results in filtering out low quality samples. Furthermore, we show that our guided approach leads to better sample generation in terms of FID scores. Michele De Vita, Vasileios Belagiannis |
WACV | 2 |
| 2025 | Revisiting Gradient-Based Uncertainty for Monocular Depth EstimationabstractMonocular depth estimation, similar to other image-based tasks, is prone to erroneous predictions due to ambiguities in the image, for example, caused by dynamic objects or shadows. For this reason, pixel-wise uncertainty assessment is required for safety-critical applications to highlight the areas where the prediction is unreliable. We address this in a post hoc manner and introduce gradient-based uncertainty estimation for already trained depth estimation models. To extract gradients without depending on the ground truth depth, we introduce an auxiliary loss function based on the consistency of the predicted depth and a reference depth. The reference depth, which acts as pseudo ground truth, is in fact generated using a simple image or feature augmentation, making our approach simple and effective. To obtain the final uncertainty score, the derivatives w.r.t. the feature maps from single or multiple layers are calculated using back-propagation. We demonstrate that our gradient-based approach is effective in determining the uncertainty without re-training using the two standard depth estimation benchmarks KITTI and NYU. In particular, for models trained with monocular sequences and therefore most prone to uncertainty, our method outperforms related approaches. Julia Hornauer, Amir El-Ghoussani, Vasileios Belagiannis |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Differentiable Cost Model for Neural-Network Accelerator Regarding Memory HierarchyabstractDedicated neural-network inference-processors improve latency and power of the computing devices. They use custom memory hierarchies that take into account the flow of operators present in neural networks and convolutional layers. For efficient implementation, such network topologies can greatly benefit from hardware-cost optimization using automated network-architecture search. Thereby, cost functions predict the suitability of a network topology for a given type of inference hardware. A differentiable neural-architecture search that optimizes both weights and topology in a single training requires cost models to be differentiable in the dimensions of weight and activation matrices. State-of-the-art differentiable cost models require time-consuming system-level measurements or simulation results, or do not encounter the hardware structure at all. This work presents a simple yet effective procedure for deriving a differentiable neural-network-accelerator cost-model that is suitable for any type of accelerator. It is based on hardware-independent parameterization and a novel differentiable divide-ceil function, as well as hardware-specific modeling. The resulting differentiable model can be reconfigured to the actual hardware size and memory structure to predict the inference energy for an exact network topology. The modeling and prediction are demonstrated for a state-of-the-art SRAM-based inference-accelerator and for the Eyeriss accelerator, inferring different state-of-the-art neural networks, resulting in excellent agreement with measured hardware. Joschua Conrad, Simon Wilhelmstätter, Rohan Asthana, Vasileios Belagiannis, Maurits Ortmanns |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2025 | Translation Consistent Semi-Supervised Segmentation for 3D Medical Imagesabstract3D medical image segmentation methods have been successful, but their dependence on large amounts of voxel-level annotated data is a disadvantage that needs to be addressed given the high cost to obtain such annotation. Semi-supervised learning (SSL) solves this issue by training models with a large unlabelled and a small labelled dataset. The most successful SSL approaches are based on consistency learning that minimises the distance between model responses obtained from perturbed views of the unlabelled data. These perturbations usually keep the spatial input context between views fairly consistent, which may cause the model to learn segmentation patterns from the spatial input contexts instead of the foreground objects. In this paper, we introduce the Translation Consistent Co-training (TraCoCo) which is a consistency learning SSL method that perturbs the input data views by varying their spatial input context, allowing the model to learn segmentation patterns from foreground objects. Furthermore, we propose a new Confident Regional Cross entropy (CRC) loss, which improves training convergence and keeps the robustness to co-training pseudo-labelling mistakes. Our method yields state-of-the-art (SOTA) results for several 3D data benchmarks, such as the Left Atrium (LA), Pancreas-CT (Pancreas), and Brain Tumor Segmentation (BraTS19). Our method also attains best results on a 2D-slice benchmark, namely the Automated Cardiac Diagnosis Challenge (ACDC), further demonstrating its effectiveness. Our code, training logs and checkpoints are available at https://github.com/yyliu01/ TraCoCo. Yuyuan Liu, Yu Tian 0001, Chong Wang 0012, Yuanhong Chen, Fengbei Liu, Vasileios Belagiannis, Gustavo Carneiro 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2024 | ItTakesTwo: Leveraging Peer Representations for Semi-supervised LiDAR Semantic Segmentation
Yuyuan Liu, Yuanhong Chen, Hu Wang 0005, Vasileios Belagiannis, Ian D. Reid 0001, Gustavo Carneiro 0001 |
ECCV (1) | 4 |
| 2024 | Infrastructure-based Perception with Cameras and Radars for Cooperative Driving ScenariosabstractRoadside infrastructure has enjoyed widespread adoption for various tasks such as traffic surveillance, traffic monitoring, control of traffic flow, and prioritization of public transit and emergency vehicles. As automated driving functions and vehicle communications continue to be researched, cooperative and connected driving scenarios can now be realized. Cooperative driving, however, imposes stringent environmental perception and model requirements. In particular, road users, including pedestrians and cyclists, must be reliably detected and accurately localized. Furthermore, the perception framework must have low latency to provide up-to-date information. In this work, we present a refined, camera-based reference point detector design that does not rely on annotated infrastructure datasets and incorporates fusion with cost-effective radar sensor data to increase system reliability, if available. The reference point detector design is realized with box and instance segmentation object detector models to extract object ground points. In parallel, objects are extracted from radar target data through a clustering pipeline and fused with camera object detections. To demonstrate the real-world applicability of our approaches for cooperative driving scenarios, we provide an extensive evaluation of data from a real test site. Alexander Tsaregorodtsev, Michael Buchholz, Vasileios Belagiannis |
IV | 3 |
| 2023 | Out-of-Distribution Detection for Monocular Depth EstimationabstractIn monocular depth estimation, uncertainty estimation approaches mainly target the data uncertainty introduced by image noise. In contrast to prior work, we address the uncertainty due to lack of knowledge, which is relevant for the detection of data not represented by the training distribution, the so-called out-of-distribution (OOD) data. Motivated by anomaly detection, we propose to detect OOD images from an encoder-decoder depth estimation model based on the reconstruction error. Given the features extracted with the fixed depth encoder, we train an image decoder for image reconstruction using only in-distribution data. Consequently, OOD images result in a high reconstruction error, which we use to distinguish between in- and out-of-distribution samples. We built our experiments on the standard NYU Depth V2 and KITTI benchmarks as in-distribution data. Our post hoc method performs astonishingly well on different models and outperforms existing uncertainty estimation approaches without modifying the trained encoder-decoder depth estimation model. Julia Hornauer, Adrian Holzbock, Vasileios Belagiannis |
ICCV | 3 |
| 2023 | Residual Pattern Learning for Pixel-wise Out-of-Distribution Detection in Semantic SegmentationabstractSemantic segmentation models classify pixels into a set of known ("in-distribution") visual classes. When deployed in an open world, the reliability of these models depends on their ability to not only classify in-distribution pixels but also to detect out-of-distribution (OoD) pixels. Historically, the poor OoD detection performance of these models has motivated the design of methods based on model re-training using synthetic training images that include OoD visual objects. Although successful, these re-trained methods have two issues: 1) their in-distribution segmentation accuracy may drop during re-training, and 2) their OoD detection accuracy does not generalise well to new contexts outside the training set (e.g., from city to country context). In this paper, we mitigate these issues with: (i) a new residual pattern learning (RPL) module that assists the segmentation model to detect OoD pixels with minimal deterioration to inlier segmentation accuracy; and (ii) a novel context-robust contrastive learning (CoroCL) that enforces RPL to robustly detect OoD pixels in various contexts. Our approach improves by around 10% FPR and 7% AuPRC previous state-of-the-art in Fishyscapes, Segment-Me-If-You-Can, and RoadAnomaly datasets. Yuyuan Liu, Choubo Ding, Yu Tian 0001, Guansong Pang, Vasileios Belagiannis, Ian D. Reid 0001, Gustavo Carneiro 0001 |
ICCV | 5 |
| 2023 | Joint Out-of-Distribution Detection and Uncertainty Estimation for Trajectory PredictionabstractDespite the significant research efforts on trajectory prediction for automated driving, limited work exists on assessing the prediction reliability. To address this limitation we propose an approach that covers two sources of error, namely novel situations with out-of-distribution (OOD) detection and the complexity in in-distribution (ID) situations with uncertainty estimation. We introduce two modules next to an encoder-decoder network for trajectory prediction. Firstly, a Gaussian mixture model learns the probability density function of the ID encoder features during training, and then it is used to detect the OOD samples in regions of the feature space with low likelihood. Secondly, an error regression network is applied to the encoder, which learns to estimate the trajectory prediction error in supervised training. During inference, the estimated prediction error is used as the uncertainty. In our experiments, the combination of both modules outperforms the prior work in OOD detection and uncertainty estimation, on the Shifts robust trajectory prediction dataset by 2.8 % and 10.1%, respectively. The code is publicly available44project page: https://github.com/againerju/joodu. Julian Wiederer, Julian Schmidt, Ulrich Kressel, Klaus Dietmayer, Vasileios Belagiannis |
IROS | 5 |
| 2023 | RESET: Revisiting Trajectory Sets for Conditional Behavior PredictionabstractIt is desirable to predict the behavior of traffic participants conditioned on different planned trajectories of the autonomous vehicle. This allows the downstream planner to estimate the impact of its decisions. Recent approaches for conditional behavior prediction rely on a regression decoder, meaning that coordinates or polynomial coefficients are regressed. In this work we revisit set-based trajectory prediction, where the probability of each trajectory in a predefined trajectory set is determined by a classification model, and first-time employ it to the task of conditional behavior prediction. We propose RESET, which combines a new metric-driven algorithm for trajectory set generation with a graph-based encoder. For unconditional prediction, RESET achieves comparable performance to a regression-based approach. Due to the nature of set-based approaches, it has the advantageous property of being able to predict a flexible number of trajectories without influencing runtime or complexity. For conditional prediction, RESET achieves reasonable results with late fusion of the planned trajectory, which was not observed for regression-based approaches before. This means that RESET is computationally lightweight to combine with a planner that proposes multiple future plans of the autonomous vehicle, as large parts of the forward pass can be reused. Julian Schmidt, Pascal Huissel, Julian Wiederer, Julian Jordan, Vasileios Belagiannis, Klaus Dietmayer |
IV | 5 |
| 2023 | Automated Static Camera Calibration with Intelligent VehiclesabstractConnected and cooperative driving requires precise calibration of the roadside infrastructure for having a reliable perception system. To solve this requirement in an automated manner, we present a robust extrinsic calibration method for automated geo-referenced camera calibration. Our method requires a calibration vehicle equipped with a combined GNSS/RTK receiver and an inertial measurement unit (IMU) for self-localization. In order to remove any requirements for the target’s appearance and the local traffic conditions, we propose a novel approach using hypothesis filtering. Our method does not require any human interaction with the information recorded by both the infrastructure and the vehicle. Furthermore, we do not limit road access for other road users during calibration. We demonstrate the feasibility and accuracy of our approach by evaluating our approach on synthetic datasets as well as a real-world connected intersection, and deploying the calibration on real infrastructure. Our source code is publicly available1. Alexander Tsaregorodtsev, Adrian Holzbock, Jan Strohbeck, Michael Buchholz, Vasileios Belagiannis |
IV | 5 |
| 2023 | Knowing What to Label for Few Shot Microscopy Image Cell SegmentationabstractIn microscopy image cell segmentation, it is common to train a deep neural network on source data, containing different types of microscopy images, and then fine-tune it using a support set comprising a few randomly selected and annotated training target images. In this paper, we argue that the random selection of unlabelled training target images to be annotated and included in the support set may not enable an effective fine-tuning process, so we propose a new approach to optimise this image selection process. Our approach involves a new scoring function to find informative unlabelled target images. In particular, we propose to measure the consistency in the model predictions on target images against specific data augmentations. However, we observe that the model trained with source datasets does not reliably evaluate consistency on target images. To alleviate this problem, we propose novel self-supervised pretext tasks to compute the scores of unlabelled target images. Finally, the top few images with the least consistency scores are added to the support set for oracle (i.e., expert) annotation and later used to fine-tune the model to the target images. In our evaluations that involve the segmentation of five different types of cell images, we demonstrate promising results on several target test sets compared to the random selection approach as well as other selection approaches, such as Shannon's entropy and Monte-Carlo dropout. Youssef Dawoud, Arij Bouazizi, Katharina Ernst, Gustavo Carneiro 0001, Vasileios Belagiannis |
WACV | 5 |
| 2023 | Heatmap-based Out-of-Distribution DetectionabstractOur work investigates out-of-distribution (OOD) detection as a neural network output explanation problem. We learn a heatmap representation for detecting OOD images while visualizing in- and out-of-distribution image regions at the same time. Given a trained and fixed classifier, we train a decoder neural network to produce heatmaps with zero response for in-distribution samples and high response heatmaps for OOD samples, based on the classifier features and the class prediction. Our main innovation lies in the heatmap definition for an OOD sample, as the normalized difference from the closest in-distribution sample. The heatmap serves as a margin to distinguish between in- and out-of-distribution samples. Our approach generates the heatmaps not only for OOD detection, but also to indicates in- and out-of-distribution regions of the input image. In our evaluations, our approach mostly outperforms the prior work on fixed classifiers, trained on CIFAR-10, CIFAR-100 and Tiny ImageNet. The code is publicly available at: https://github.com/jhornauer/heatmap_ood. Julia Hornauer, Vasileios Belagiannis |
WACV | 2 |
| 2023 | ParticleAugment: Sampling-based data augmentationabstractWe present an automated data augmentation approach for image classification . The problem is formulated as a Monte Carlo sampling problem where the goal is to approximate the optimal augmentation policies using a policy mixture distribution. We propose a particle filter scheme for the policy search where the probability of applying a set of augmentation operations forms the state of the filter. The policy performance is measured based on the loss function difference between a reference model and the actual model. This performance measure is then used to re-weight the particles and finally update the policy distribution. In our experiments, we show that our formulation for automated augmentation reaches promising results on CIFAR-10, CIFAR-100, and ImageNet datasets using the standard network architectures for this problem. By comparing with the related work, our method reaches a balance between the computational cost of policy search and the model performance. The source code of our approach is publicly available. Alexander Tsaregorodtsev, Vasileios Belagiannis |
Comput. Vis. Image Underst. | 2 |
| 2023 | LongReMix: Robust learning with high confidence samples in a noisy label environment
Filipe R. Cordeiro, Ragav Sachdeva, Vasileios Belagiannis, Ian D. Reid 0001, Gustavo Carneiro 0001 |
Pattern Recognit. | 3 |
| 2023 | ScanMix: Learning from Severe Label Noise via Semantic Clustering and Semi-Supervised LearningabstractWe propose a new training algorithm, ScanMix, that explores semantic clustering and semi-supervised learning (SSL) to allow superior robustness to severe label noise and competitive robustness to non-severe label noise problems, in comparison to the state of the art (SOTA) methods. ScanMix is based on the expectation maximisation framework, where the E-step estimates the latent variable to cluster the training images based on their appearance and classification results, and the M-step optimises the SSL classification and learns effective feature representations via semantic clustering. We present a theoretical result that shows the correctness and convergence of ScanMix, and an empirical result that shows that ScanMix has SOTA results on CIFAR-10/-100 (with symmetric, asymmetric and semantic label noise), Red Mini-ImageNet (from the Controlled Noisy Web Labels), Clothing1M and WebVision. In all benchmarks with severe label noise, our results are competitive to the current SOTA. Ragav Sachdeva, Filipe R. Cordeiro, Vasileios Belagiannis, Ian D. Reid 0001, Gustavo Carneiro 0001 |
Pattern Recognit. | 3 |
| 2022 | Multi-Task Edge Prediction in Temporally-Dynamic Video Graphs
Osman Ülger, Julian Wiederer, Mohsen Ghafoorian, Vasileios Belagiannis, Pascal Mettes |
BMVC | 4 |
| 2022 | Perturbed and Strict Mean Teachers for Semi-supervised Semantic SegmentationabstractConsistency learning using input image, feature, or network perturbations has shown remarkable results in semi-supervised semantic segmentation, but this approach can be seriously affected by inaccurate predictions of unlabelled training images. There are two consequences of these inaccurate predictions: 1) the training based on the “strict” cross-entropy (CE) loss can easily overfit prediction mistakes, leading to confirmation bias; and 2) the perturbations applied to these inaccurate predictions will use potentially erroneous predictions as training signals, degrading consistency learning. In this paper, we address the prediction accuracy problem of consistency learning methods with novel extensions of the mean-teacher (MT) model, which include a new auxiliary teacher, and the replacement of MT's mean square error (MSE) by a stricter confidence-weighted cross-entropy (Conf-CE) loss. The accurate prediction by this model allows us to use a challenging combination of network, input data and feature perturbations to improve the consistency learning generalisation, where the feature perturbations consist of a new adversarial perturbation. Results on public benchmarks show that our approach achieves remarkable improvements over the previous SOTA methods in the field.11Supported by Australian Research Council through grants DP180103232 and FT190100525. Our code is available at https://github.com/yyliu01/PS-MT. Yuyuan Liu, Yu Tian 0001, Yuanhong Chen, Fengbei Liu, Vasileios Belagiannis, Gustavo Carneiro 0001 |
CVPR | 5 |
| 2022 | ACPL: Anti-curriculum Pseudo-labelling for Semi-supervised Medical Image ClassificationabstractEffective semi-supervised learning (SSL) in medical image analysis (MIA) must address two challenges: 1) work effectively on both multi-class (e.g., lesion classification) and multi-label (e.g., multiple-disease diagnosis) problems, and 2) handle imbalanced learning (because of the high variance in disease prevalence). One strategy to explore in SSL MIA is based on the pseudo labelling strategy, but it has a few shortcomings. Pseudo-labelling has in general lower accuracy than consistency learning, it is not specifically design for both multi-class and multi-label problems, and it can be challenged by imbalanced learning. In this paper, unlike traditional methods that select confident pseudo label by threshold, we propose a new SSL algorithm, called anti-curriculum pseudo-labelling (ACPL), which introduces novel techniques to select informative unlabelled samples, improving training balance and allowing the model to work for both multi-label and multi-class problems, and to estimate pseudo labels by an accurate ensemble of classifiers (improving pseudo label accuracy). We run extensive experiments to evaluate ACPL on two public medical image classification benchmarks: Chest X-Ray 14 for thorax disease multi-label classification and ISIC2018 for skin lesion multi-class classification. Our method outperforms previous SOTA SSL methods on both datasets11Supported by Australian Research Council through grants DP180103232 and FT190100525.22Code is available at https://github.com/FBLADL/ACPL. Fengbei Liu, Yu Tian 0001, Yuanhong Chen, Yuyuan Liu, Vasileios Belagiannis, Gustavo Carneiro 0001 |
CVPR | 5 |
| 2022 | Gradient-Based Uncertainty for Monocular Depth Estimation
Julia Hornauer, Vasileios Belagiannis |
ECCV (20) | 2 |
| 2022 | Lightweight Monocular Depth Estimation through Guided DecodingabstractWe present a lightweight encoder-decoder architecture for monocular depth estimation, specifically designed for embedded platforms. Our main contribution is the Guided Upsampling Block (GUB) for building the decoder of our model. Motivated by the concept of guided image filtering, GUB relies on the image to guide the decoder on upsampling the feature representation and the depth map reconstruction, achieving high resolution results with fine-grained details. Based on multiple GUBs, our model outperforms the related methods on the NYU Depth V2 dataset in terms of accuracy while delivering up to 35.1 fps on the NVIDIA Jetson Nano and up to 144.5 fps on the NVIDIA Xavier NX. Similarly, on the KITTI dataset, inference is possible with up to 23.7 fps on the Jetson Nano and 102.9 fps on the Xavier NX. Our code and models are made publicly available14https://github.com/mic-rud/GuidedDecoding. Michael Rudolph 0006, Youssef Dawoud, Ronja Güldenring, Lazaros Nalpantidis, Vasileios Belagiannis |
ICRA | 5 |
| 2022 | MotionMixer: MLP-based 3D Human Body Pose ForecastingabstractIn this work, we present MotionMixer, an efficient 3D human body pose forecasting model based solely on multi-layer perceptrons (MLPs). MotionMixer learns the spatial-temporal 3D body pose dependencies by sequentially mixing both modalities. Given a stacked sequence of 3D body poses, a spatial-MLP extracts fine-grained spatial dependencies of the body joints. The interaction of the body joints over time is then modelled by a temporal MLP. The spatial-temporal mixed features are finally aggregated and decoded to obtain the future motion. To calibrate the influence of each time step in the pose sequence, we make use of squeeze-and-excitation (SE) blocks. We evaluate our approach on Human3.6M, AMASS, and 3DPW datasets using the standard evaluation protocols. For all evaluations, we demonstrate state-of-the-art performance, while having a model with a smaller number of parameters. Our code is available at: https://github.com/MotionMLP/MotionMixer. Arij Bouazizi, Adrian Holzbock, Ulrich Kressel, Klaus Dietmayer, Vasileios Belagiannis |
IJCAI | 5 |
| 2022 | A Spatio-Temporal Multilayer Perceptron for Gesture RecognitionabstractGesture recognition is essential for the interaction of autonomous vehicles with humans. While the current approaches focus on combining several modalities like image features, keypoints and bone vectors, we present neural network architecture that delivers state-of-the-art results only with body skeleton input data. We propose the spatio-temporal multilayer perceptron for gesture recognition in the context of autonomous vehicles. Given 3D body poses over time, we define temporal and spatial mixing operations to extract features in both domains. Additionally, the importance of each time step is re-weighted with Squeeze-and-Excitation layers. An extensive evaluation of the TCG and Drive& Act datasets is provided to showcase the promising performance of our approach. Furthermore, we deploy our model to our autonomous vehicle to show its real-time capability and stable execution. Adrian Holzbock, Alexander Tsaregorodtsev, Youssef Dawoud, Klaus Dietmayer, Vasileios Belagiannis |
IV | 5 |
| 2022 | A Multi-Task Recurrent Neural Network for End-to-End Dynamic Occupancy Grid MappingabstractA common approach for modeling the environment of an autonomous vehicle are dynamic occupancy grid maps, in which the surrounding is divided into cells, each containing the occupancy and velocity state of its location. Despite the advantage of modeling arbitrary shaped objects, the used algorithms rely on hand-designed inverse sensor models and semantic information is missing. Therefore, we introduce a multi-task recurrent neural network to predict grid maps providing occupancies, velocity estimates, semantic information and the driveable area. During training, our network architecture, which is a combination of convolutional and recurrent layers, processes sequences of raw lidar data, that is represented as bird’s eye view images with several height channels. The multi-task network is trained in an end-to-end fashion to predict occupancy grid maps without the usual preprocessing steps consisting of removing ground points and applying an inverse sensor model. In our evaluations, we show that our learned inverse sensor model is able to overcome some limitations of a geometric inverse sensor model in terms of representing object shapes and modeling freespace. Moreover, we report a better runtime performance and more accurate semantic predictions for our end-to-end approach, compared to our network relying on measurement grid maps as input data. Marcel Schreiber, Vasileios Belagiannis, Claudius Gläser, Klaus Dietmayer |
IV | 2 |
| 2022 | NVUM: Non-volatile Unbiased Memory for Robust Medical Image Classification
Fengbei Liu, Yuanhong Chen, Yu Tian 0001, Yuyuan Liu, Chong Wang 0012, Vasileios Belagiannis, Gustavo Carneiro 0001 |
MICCAI (3) | 6 |
| 2021 | Learning Temporal 3D Human Pose Estimation with Pseudo-LabelsabstractWe present a simple, yet effective, approach for self-supervised 3D human pose estimation. Unlike the prior work, we explore the temporal information next to the multi-view self-supervision. During training, we rely on triangulating 2D body pose estimates of a multiple-view camera system. A temporal convolutional neural network is trained with the generated 3D ground-truth and the geometric multi-view consistency loss, imposing geometrical constraints on the predicted 3D body skeleton. During inference, our model receives a sequence of 2D body pose estimates from a single-view to predict the 3D body pose for each of them. An extensive evaluation shows that our method achieves state-of-the-art performance in the Human3.6M and MPI-INF-3DHP benchmarks. Our code and models are publicly available at https://github.com/vru2020/TM_HPE/. Arij Bouazizi, Ulrich Kressel, Vasileios Belagiannis |
AVSS | 3 |
| 2021 | PropMix: Hard Sample Filtering and Proportional MixUp for Learning with Noisy Labels
Filipe R. Cordeiro, Vasileios Belagiannis, Ian D. Reid 0001, Gustavo Carneiro 0001 |
BMVC | 2 |
| 2021 | Self-Supervised 3D Human Pose Estimation with Multiple-View GeometryabstractWe present a self-supervised learning algorithm for 3D human pose estimation of a single person based on a multiple-view camera system and 2D body pose estimates for each view. To train our model, represented by a deep neural network, we propose a four-loss function learning algorithm, which does not require any 2D or 3D body pose ground-truth. The proposed loss functions make use of the multiple-view geometry to reconstruct 3D body pose estimates and impose body pose constraints across the camera views. Our approach utilizes all available camera views during training, while the inference is single-view. In our evaluations, we show promising performance on Human3.6M and HumanEva benchmarks, while we also present a generalization study on MPI-INF-3DHP dataset, as well as several ablation results. Overall, we outperform all self-supervised learning methods and reach comparable results to supervised and weakly-supervised learning approaches. Our code and models are publicly available11Source Code: https://gi-thub.com/vru2020jpose_3D/. Arij Bouazizi, Julian Wiederer, Ulrich Kressel, Vasileios Belagiannis |
FG | 4 |
| 2021 | Dynamic Occupancy Grid Mapping with Recurrent Neural NetworksabstractModeling and understanding the environment is an essential task for autonomous driving. In addition to the detection of objects, in complex traffic scenarios the motion of other road participants is of special interest. Therefore, we propose to use a recurrent neural network to predict a dynamic occupancy grid map, which divides the vehicle surrounding in cells, each containing the occupancy probability and a velocity estimate. During training, our network is fed with sequences of measurement grid maps, which encode the lidar measurements of a single time step. Due to the combination of convolutional and recurrent layers, our approach is capable to use spatial and temporal information for the robust detection of static and dynamic environment. In order to apply our approach with measurements from a moving ego-vehicle, we propose a method for ego-motion compensation that is applicable in neural network architectures with recurrent layers working on different resolutions. In our evaluations, we compare our approach with a state-of-the-art particle-based algorithm on a large publicly available dataset to demonstrate the improved accuracy of velocity estimates and the more robust separation of the environment in static and dynamic area. Additionally, we show that our proposed method for ego-motion compensation leads to comparable results in scenarios with stationary and with moving ego-vehicle. Marcel Schreiber, Vasileios Belagiannis, Claudius Gläser, Klaus Dietmayer |
ICRA | 2 |
| 2021 | EvidentialMix: Learning with Combined Open-set and Closed-set Noisy LabelsabstractThe efficacy of deep learning depends on large-scale data sets that have been carefully curated with reliable data acquisition and annotation processes. However, acquiring such large-scale data sets with precise annotations is very expensive and time-consuming, and the cheap alternatives often yield data sets that have noisy labels. The field has addressed this problem by focusing on training models under two types of label noise: 1) closed-set noise, where some training samples are incorrectly annotated to a training label other than their known true class; and 2) open-set noise, where the training set includes samples that possess a true class that is (strictly) not contained in the set of known training labels. In this work, we study a new variant of the noisy label problem that combines the open-set and closed-set noisy labels, and introduce a benchmark evaluation to assess the performance of training algorithms under this setup. We argue that such problem is more general and better reflects the noisy label scenarios in practice. Furthermore, we propose a novel algorithm, called EvidentialMix, that addresses this problem and compare its performance with the state-of-the-art methods for both closed-set and open-set noise on the proposed benchmark. Our results show that our method produces superior classification results and better feature representations than previous state-of-the-art methods. The code is available at https:/github.com/ragavsachdeva/EvidentialMix. Ragav Sachdeva, Filipe R. Cordeiro, Vasileios Belagiannis, Ian D. Reid 0001, Gustavo Carneiro 0001 |
WACV | 3 |
| 2021 | Forecasting People Trajectories and Head Poses by Jointly Reasoning on Tracklets and VisletsabstractIn this article, we explore the correlation between people trajectories and their head orientations. We argue that people trajectory and head pose forecasting can be modelled as a joint problem. Recent approaches on trajectory forecasting leverage short-term trajectories (aka tracklets) of pedestrians to predict their future paths. In addition, sociological cues, such as expected destination or pedestrian interaction, are often combined with tracklets. In this article, we propose MiXing-LSTM (MX-LSTM) to capture the interplay between positions and head orientations (vislets) thanks to a joint unconstrained optimization of full covariance matrices during the LSTM backpropagation. We additionally exploit the head orientations as a proxy for the visual attention, when modeling social interactions. MX-LSTM predicts future pedestrians location and head pose, increasing the standard capabilities of the current approaches on long-term trajectory forecasting. Compared to the state-of-the-art, our approach shows better performances on an extensive set of public benchmarks. MX-LSTM is particularly effective when people move slowly, i.e., the most challenging scenario for all other models. The proposed approach also allows for accurate predictions on a longer time horizon. Irtiza Hasan, Francesco Setti, Theodore Tsesmelis, Vasileios Belagiannis, Sikandar Amin, Alessio Del Bue, Marco Cristani, Fabio Galasso |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2020 | Deep Learning versus High-order Recurrent Neural Network based Decoding for Convolutional CodesabstractIn the last decade, deep neural networks (DNNs) have shown impressive results in various fields such as image classification, speech recognition, or playing the abstract strategy board game Go. Recently, also an increased interest in the application of DNNs to physical layer problems in digital communications can be observed. We use a DNN for one-shot decoding of convolutional (self-orthogonal) codes. An advantage of this use case is the unlimited amount of labeled data for training. A disadvantage is, that the number of code words to be learned increases exponentially with the dimension of the code. We compare the performance of the DNN-based decoding with iterative threshold decoding (ITD). Here, a discrete-time high-order recurrent neural network (HORNN) is used as a computational model for ITD. Unfolding the HORNN in time, we arrive at a DNN with a special structure as defined by the HORNN. With a training procedure we can optimize the performance of this unfolded HORNN (uHORNN). The advantage of this approach is that the structure of the uHORNN is determined by the structure of iterative threshold decoding. Only a few weights, which are shared within and between the layers, must be adapted to optimize the network. In this way we combine the advantages of both approaches, the structured approach of the HORNN as a computational model, on the one hand, and the training based optimization as given by a DNN on the other hand. Werner G. Teich, Vasileios Belagiannis |
GLOBECOM | 3 |
| 2020 | Motion Estimation in Occupancy Grid Maps in Stationary Settings Using Recurrent Neural NetworksabstractIn this work, we tackle the problem of modeling the vehicle environment as dynamic occupancy grid map in complex urban scenarios using recurrent neural networks. Dynamic occupancy grid maps represent the scene in a bird's eye view, where each grid cell contains the occupancy probability and the two dimensional velocity. As input data, our approach relies on measurement grid maps, which contain occupancy probabilities, generated with lidar measurements. Given this configuration, we propose a recurrent neural network architecture to predict a dynamic occupancy grid map, i.e. filtered occupancy and velocity of each cell, by using a sequence of measurement grid maps. Our network architecture contains convolutional long-short term memories in order to sequentially process the input, makes use of spatial context, and captures motion. In the evaluation, we quantify improvements in estimating the velocity of braking and turning vehicles compared to the state-of-the-art. Additionally, we demonstrate that our approach provides more consistent velocity estimates for dynamic objects, as well as, less erroneous velocity estimates in static area. Marcel Schreiber, Vasileios Belagiannis, Claudius Gläser, Klaus Dietmayer |
ICRA | 2 |
| 2020 | Multiple Trajectory Prediction with Deep Temporal and Spatial Convolutional Neural NetworksabstractAutomated vehicles need to not only perceive their environment, but also predict the possible future behavior of all detected traffic participants in order to safely navigate in complex scenarios and avoid critical situations, ranging from merging on highways to crossing urban intersections. Due to the availability of datasets with large numbers of recorded trajectories of traffic participants, deep learning based approaches can be used to model the behavior of road users. This paper proposes a convolutional network that operates on rasterized actor-centric images which encode the static and dynamic actor-environment. We predict multiple possible future trajectories for each traffic actor, which include position, velocity, acceleration, orientation, yaw rate and position uncertainty estimates. To make better use of the past movement of the actor, we propose to employ temporal convolutional networks (TCNs) and rely on uncertainties estimated from the previous object tracking stage. We evaluate our approach on the public "Argoverse Motion Forecasting" dataset, on which it won the first prize at the Argoverse Motion Forecasting Challenge, as presented on the NeurIPS 2019 workshop on "Machine Learning for Autonomous Driving". Jan Strohbeck, Vasileios Belagiannis, Johannes Müller 0003, Marcel Schreiber, Martin Herrmann, Daniel Wolf, Michael Buchholz |
IROS | 2 |
| 2020 | Traffic Control Gesture Recognition for Autonomous VehiclesabstractA car driver knows how to react on the gestures of the traffic officers. Clearly, this is not the case for the autonomous vehicle, unless it has road traffic control gesture recognition functionalities. In this work, we address the limitation of the existing autonomous driving datasets to provide learning data for traffic control gesture recognition. We introduce a dataset that is based on 3D body skeleton input to perform traffic control gesture classification on every time step. Our dataset consists of 250 sequences from several actors, ranging from 16 to 90 seconds per sequence. To evaluate our dataset, we propose eight sequential processing models based on deep neural networks such as recurrent networks, attention mechanism, temporal convolutional networks and graph convolutional networks. We present an extensive evaluation and analysis of all approaches for our dataset, as well as real-world quantitative evaluation. The code and dataset is publicly available4. Julian Wiederer, Arij Bouazizi, Ulrich Kressel, Vasileios Belagiannis |
IROS | 4 |
| 2017 | Recurrent Human Pose EstimationabstractWe propose a ConvNet model for predicting 2D human body poses in an image. The model regresses a heatmap representation for each body keypoint, and is able to learn and represent both the part appearances and the context of the part configuration. We make the following three contributions: (i) an architecture combining a feed forward module with a recurrent module, where the recurrent module can be run iteratively to improve the performance; (ii) the model can be trained end-to-end and from scratch, with auxiliary losses incorporated to improve performance; (iii) we investigate whether keypoint visibility can also be predicted. The model is evaluated on two benchmark datasets. The result is a simple architecture that achieves performance on par with the state of the art, but without the complexity of a graphical model stage (or layers). Vasileios Belagiannis, Andrew Zisserman |
FG | 1 |
| 2016 | Deeper Depth Prediction with Fully Convolutional Residual NetworksabstractThis paper addresses the problem of estimating the depth map of a scene given a single RGB image. We propose a fully convolutional architecture, encompassing residual learning, to model the ambiguous mapping between monocular images and depth maps. In order to improve the output resolution, we present a novel way to efficiently learn feature map up-sampling within the network. For optimization, we introduce the reverse Huber loss that is particularly suited for the task at hand and driven by the value distributions commonly present in depth maps. Our model is composed of a single architecture that is trained end-to-end and does not rely on post-processing techniques, such as CRFs or other additional refinement steps. As a result, it runs in real-time on images or videos. In the evaluation, we show that the proposed model contains fewer parameters and requires fewer training data than the current state of the art, while outperforming all approaches on depth estimation. Code and models are publicly available. Iro Laina, Christian Rupprecht 0001, Vasileios Belagiannis, Federico Tombari, Nassir Navab |
3DV | 3 |
| 2016 | Real-time localization of articulated surgical instruments in retinal microsurgery
Nicola Rieke, David Joseph Tan, Chiara Amat di San Filippo, Federico Tombari, Mohamed Alsheakhali, Vasileios Belagiannis, Abouzar Eslami, Nassir Navab |
Medical Image Anal. | 6 |
| 2016 | Parsing human skeletons in an operating room
Vasileios Belagiannis, Xinchao Wang, Horesh Ben Shitrit, Kiyoshi Hashimoto, Ralf Stauder, Yoshimitsu Aoki, Michael Kranzfelder, Armin Schneider, Pascal Fua, Slobodan Ilic, Hubertus Feußner, Nassir Navab |
Mach. Vis. Appl. | 1 |
| 2016 | 3D Pictorial Structures Revisited: Multiple Human Pose EstimationabstractWe address the problem of 3D pose estimation of multiple humans from multiple views. The transition from single to multiple human pose estimation and from the 2D to 3D space is challenging due to a much larger state space, occlusions and across-view ambiguities when not knowing the identity of the humans in advance. To address these problems, we first create a reduced state space by triangulation of corresponding pairs of body parts obtained by part detectors for each camera view. In order to resolve ambiguities of wrong and mixed parts of multiple humans after triangulation and also those coming from false positive detections, we introduce a 3D pictorial structures (3DPS) model. Our model builds on multi-view unary potentials, while a prior model is integrated into pairwise and ternary potential functions. To balance the potentials' influence, the model parameters are learnt using a Structured SVM (SSVM). The model is generic and applicable to both single and multiple human pose estimation. To evaluate our model on single and multiple human pose estimation, we rely on four different datasets. We first analyse the contribution of the potentials and then compare our results with related work where we demonstrate superior performance. Vasileios Belagiannis, Sikandar Amin, Mykhaylo Andriluka, Bernt Schiele, Nassir Navab, Slobodan Ilic |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2016 | AggNet: Deep Learning From Crowds for Mitosis Detection in Breast Cancer Histology ImagesabstractThe lack of publicly available ground-truth data has been identified as the major challenge for transferring recent developments in deep learning to the biomedical imaging domain. Though crowdsourcing has enabled annotation of large scale databases for real world images, its application for biomedical purposes requires a deeper understanding and hence, more precise definition of the actual annotation task. The fact that expert tasks are being outsourced to non-expert users may lead to noisy annotations introducing disagreement between users. Despite being a valuable resource for learning annotation models from crowdsourcing, conventional machine-learning methods may have difficulties dealing with noisy annotations during training. In this manuscript, we present a new concept for learning from crowds that handle data aggregation directly as part of the learning process of the convolutional neural network (CNN) via additional crowdsourcing layer (AggNet). Besides, we present an experimental study on learning from crowds designed to answer the following questions. 1) Can deep CNN be trained with data collected from crowdsourcing? 2) How to adapt the CNN to train on multiple types of annotation datasets (ground truth and crowd-based)? 3) How does the choice of annotation and aggregation affect the accuracy? Our experimental setup involved Annot8, a self-implemented web-platform based on Crowdflower API realizing image annotation tasks for a publicly available biomedical image database. Our results give valuable insights into the functionality of deep CNN learning from crowd annotations and prove the necessity of data aggregation integration. Shadi Albarqouni, Christoph Baur, Felix Achilles, Vasileios Belagiannis, Stefanie Demirci, Nassir Navab |
IEEE Trans. Medical Imaging | 4 |
| 2015 | Robust Optimization for Deep RegressionabstractConvolutional Neural Networks (ConvNets) have successfully contributed to improve the accuracy of regression-based methods for computer vision tasks such as human pose estimation, landmark localization, and object detection. The network optimization has been usually performed with L2 loss and without considering the impact of outliers on the training process, where an outlier in this context is defined by a sample estimation that lies at an abnormal distance from the other training sample estimations in the objective space. In this work, we propose a regression model with ConvNets that achieves robustness to such outliers by minimizing Tukey's biweight function, an M-estimator robust to outliers, as the loss function for the ConvNet. In addition to the robust loss, we introduce a coarse-to-fine model, which processes input images of progressively higher resolutions for improving the accuracy of the regressed values. In our experiments, we demonstrate faster convergence and better generalization of our robust loss function for the tasks of human pose estimation and age estimation from face images. We also show that the combination of the robust loss function with the coarse-to-fine model produces comparable or better results than current state-of-the-art approaches in four publicly available human pose estimation datasets. Vasileios Belagiannis, Christian Rupprecht 0001, Gustavo Carneiro 0001, Nassir Navab |
ICCV | 1 |
| 2015 | Revisiting Robust Visual Tracking Using Pixel-Wise Posteriors
Falk Schubert, Daniele Casaburo, Dirk Dickmanns, Vasileios Belagiannis |
ICVS | 4 |
| 2015 | Surgical Tool Tracking and Pose Estimation in Retinal Microsurgery
Nicola Rieke, David Joseph Tan, Mohamed Alsheakhali, Federico Tombari, Chiara Amat di San Filippo, Vasileios Belagiannis, Abouzar Eslami, Nassir Navab |
MICCAI (1) | 6 |
| 2014 | 3D Pictorial Structures for Multiple Human Pose EstimationabstractIn this work, we address the problem of 3D pose estimation of multiple humans from multiple views. This is a more challenging problem than single human 3D pose estimation due to the much larger state space, partial occlusions as well as across view ambiguities when not knowing the identity of the humans in advance. To address these problems, we first create a reduced state space by triangulation of corresponding body joints obtained from part detectors in pairs of camera views. In order to resolve the ambiguities of wrong and mixed body parts of multiple humans after triangulation and also those coming from false positive body part detections, we introduce a novel 3D pictorial structures (3DPS) model. Our model infers 3D human body configurations from our reduced state space. The 3DPS model is generic and applicable to both single and multiple human pose estimation. In order to compare to the state-of-the art, we first evaluate our method on single human 3D pose estimation on HumanEva-I [22] and KTH Multiview Football Dataset II [8] datasets. Then, we introduce and evaluate our method on two datasets for multiple human 3D pose estimation. In order to compare to the state-of-the art, we first evaluate our method on single human 3D pose estimation on HumanEva-I [22] and KTH Multiview Football Dataset II [8] datasets. Then, we introduce and evaluate our method on two datasets for multiple human 3D pose estimation. Vasileios Belagiannis, Sikandar Amin, Mykhaylo Andriluka, Bernt Schiele, Nassir Navab, Slobodan Ilic |
CVPR | 1 |
| 2014 | Fully Automatic Catheter Localization in C-Arm Images Using ℓ1-Sparse Coding
Fausto Milletari, Vasileios Belagiannis, Nassir Navab, Pascal Fallavollita |
MICCAI (2) | 2 |
| 2012 | Segmentation Based Particle Filtering for Real-Time 2D Object Tracking
Vasileios Belagiannis, Falk Schubert, Nassir Navab, Slobodan Ilic |
ECCV (4) | 1 |