EDBT 2026 Demo / reviewers in the wild / expert
Driton Salihu
dblp:307/2975
· DBLP profile ↗
19ranked-venue papers
4as first author
19since 2021 · last 2025
0000-0001-9152-0770ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 9 since 2021Systems, architecture and hardware · 7 · 7 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SMCNet: Supervised Surface Material Classification Using mmWave Radar IQ Signals and Complex-valued CNNsabstractUnderstanding surface material properties is crucial for enhancing indoor robot perception and indoor digital twinning. However, not all sensor modalities typically employed for this task are capable of reliably capturing detailed surface material characteristics. By analyzing the reflected RF signal from a mmWave radar sensor, it is possible to extract information about the reflective material and its composition from a certain surface. We introduce a mmWave MIMO FMCW radar-based surface material classifier SMCNet, employing a complex-valued Convolutional Neural Network (CNN) and complex radar IQ signal input for classifying indoor surface materials. While current radar-based material estimation approaches rely on a fixed sensing distance and constrained setups, our approach incorporates a setup with multiple sensing distances. We trained SMCNet using data from three distinct distances and subsequently tested it on these distances, as well as on two more unseen distances. We reached an overall accuracy of 99.12-99.53% on our test set. Notably, range FFT pre-processing improved accuracy on unknown distances from 25.25% to 58.81% without re-training. Stefan Hägele, Fabián Seguel, Driton Salihu, Adam Misik, Eckehard G. Steinbach |
ICASSP | 3 |
| 2025 | HypCAD: Geometry-Enhanced Hyperbolic Contrastive Learning for CAD Model RetrievalabstractRetrieving CAD models for real-world object scans enhances object-level mapping, providing a nuanced spatial understanding crucial for precise interactions in robotics or mixed reality. Commonly, CAD model retrieval is performed by matching features learned in Euclidean space. However, learning discriminative features in Euclidean space faces significant challenges, primarily due to its flat nature and the wide variety of CAD models with different levels of detail. To address the limitations of Euclidean space and improve CAD model retrieval, this paper introduces HypCAD, a contrastive learning framework in hyperbolic space. We present a novel geometry-enhanced hyperbolic distance and utilize a three-component contrastive learning loss to learn hyperbolic feature representations for the CAD model retrieval task. We demonstrate HypCAD’s superior retrieval accuracy through comparisons with baseline contrastive learning methods on both the synthetic ShapeNet dataset and the real-world Scan2CAD dataset. Adam Misik, Driton Salihu, Heike Brock, Eckehard G. Steinbach |
ICASSP | 2 |
| 2025 | MistSense: Versatile Online Detection of Procedural and Execution Mistakes
Constantin Patsch, Marsil Zakour, Driton Salihu, Eckehard G. Steinbach |
ICCV | 4 |
| 2025 | EQUR: Equivariant Uncertainty Quantification and Refinement for Point Cloud RegistrationabstractPoint cloud registration is a crucial task for robotics and mixed reality applications, serving as a foundational component for problems such as 3D reconstruction and localization. Arbitrary poses and real-world artifacts, including noise and occlusions, increase registration uncertainty and limit the performance of current point cloud registration algorithms. This paper proposes a novel, sampling-free uncertainty quantification and refinement method for point cloud registration, termed EQUR. To consistently predict uncertainty with high robustness, we extract equivariant point features, from which we regress an uncertainty score, enabling robust quantification of registration uncertainty. Subsequently, we leverage the estimated registration uncertainty as an auxiliary input to enhance the prediction of transformation refinement terms. We employ an introspective learning strategy to train EQUR based on the errors of a baseline registration model. Through quantitative and qualitative analyses on synthetic ShapeNet and real-world ScanObjectNN datasets, we showcase the effectiveness of EQUR, demonstrating both high accuracies in uncertainty quantification and uncertainty-aided refinement of point cloud registration. Adam Misik, Driton Salihu, Xiaoang Zhang, Heike Brock, Eckehard G. Steinbach |
ICIP | 2 |
| 2024 | Long-Term Action Anticipation Based on Contextual AlignmentabstractIn action anticipation, the model predicts the next future action after a certain observation period. In long-term action anticipation, this idea is further extended to predicting multiple actions and their respective duration. Thus, in this problem setting the model should not only capture relationships between past actions but also predict several future actions that fit into a certain context. Compared to autoregressive models, our model employs an encoder decoder structure to determine future actions and durations in parallel, which prevents the accumulation of prediction errors and reduces the inference time. Furthermore, it is ensured that the predicted actions are aligned with respect to a context representation, which resembles the way humans approach this task as the feasible action set is restricted by the respective context. We evaluate our model on the long-term anticipation benchmark datasets, Breakfast, and 50Salads, where we achieve state-of-the-art results. Constantin Patsch, Jinghan Zhang 0009, Marsil Zakour, Driton Salihu, Eckehard G. Steinbach |
ICASSP | 5 |
| 2024 | NPRF: Neural Painted Radiosity Fields for Neural Implicit Rendering and Surface ReconstructionabstractIn recency, neural signed distance fields have become more popular for reconstructing 3D indoor environments. While great improvements have been made due to missing incident radiance and materials in the surface estimation, current methods cannot reconstruct high-quality surfaces. To address this issue, we propose Neural Painted Radiosity Fields (NPRF), consisting of Neural Radiosity Fields for volumetric surface representation and Neural Painted Scenes for novel view synthesis. Neural Radiosity Fields combine the radiative transfer equation with neural radiosity to estimate 3D surfaces, thus leveraging raytracing to improve the volumetric representation. Neural Painted Scenes employs sparsification and projection of 3D points into 2D images in conjunction with a generative, context-aware inpainting network to produce high-quality novel views. We show that NPRF leads to overall improvements in F-score on the popular ScanNet dataset. Finally, we show that NPRF improves novel view synthesis by a significant margin, giving improvements of up to 25% on PSNR, 53% on LPIPS, and 3% on SSIM. Driton Salihu, Adam Misik, Constantin Patsch, Eckehard G. Steinbach |
ICASSP | 1 |
| 2024 | DeepSPF: Spherical SO(3)-Equivariant Patches for Scan-to-CAD EstimationabstractRecently, SO(3)-equivariant methods have been explored for 3D reconstruction via Scan-to-CAD.
Despite significant advancements attributed to the unique characteristics of 3D data, existing SO(3)-equivariant approaches often fall short in seamlessly integrating local and global contextual information in a widely generalizable manner.
Our contributions in this paper are threefold.
First, we introduce Spherical Patch Fields, a representation technique designed for patch-wise, SO(3)-equivariant 3D point clouds, anchored theoretically on the principles of Spherical Gaussians.
Second, we present the Patch Gaussian Layer, designed for the adaptive extraction of local and global contextual information from resizable point cloud patches.
Culminating our contributions, we present Learnable Spherical Patch Fields (DeepSPF) – a versatile and easily integrable backbone suitable for instance-based point networks.
Through rigorous evaluations, we demonstrate significant enhancements in Scan-to-CAD performance for point cloud registration, retrieval, and completion: a significant reduction in the rotation error of existing registration methods, an improvement of up to 17\% in the Top-1 error for retrieval tasks, and a notable reduction of up to 30\% in the Chamfer Distance for completion models, all attributable to the incorporation of DeepSPF. Driton Salihu, Adam Misik, Constantin Patsch, Fabián Seguel, Eckehard G. Steinbach |
ICLR | 1 |
| 2024 | HEGN: Hierarchical Equivariant Graph Neural Network for 9DoF Point Cloud RegistrationabstractGiven its wide application in robotics, point cloud registration is a widely researched topic. Conventional methods aim to find a rotation and translation that align two point clouds in 6 degrees of freedom (DoF). However, certain tasks in robotics, such as category-level pose estimation, involve non-uniformly scaled point clouds, requiring a 9DoF transform for accurate alignment. We propose HEGN, a novel equivariant graph neural network for 9DoF point cloud registration. HEGN utilizes equivariance to rotation, translation, and scaling to estimate the transformation without relying on point correspondences. Based on graph representations for both point clouds, we extract equivariant node features aggregated in their local, cross-, and global context. In addition, we introduce a novel node pooling mechanism that leverages the cross-context importance of nodes to pool the graph representation. By repeating the feature extraction and node pooling, we obtain a graph hierarchy. Finally, we determine rotation and translation by aligning equivariant features aggregated over the graph hierarchy. To estimate scaling, we leverage scale information in the vector norm of the equivariant features. We evaluate the effectiveness of HEGN through experiments with the synthetic ModelNet40 dataset and the real-world ScanObjectNN dataset. The results show the superior performance of HEGN in 9DoF point cloud registration and its competitive performance in conventional 6DoF point cloud registration. Adam Misik, Driton Salihu, Heike Brock, Eckehard G. Steinbach |
ICRA | 2 |
| 2024 | Sim-to-Real Domain Shift in Online Action DetectionabstractHuman reasoning comprises the ability to understand and reason about the current action solely based on past information. To provide effective assistance in an eldercare or household environment an assistive robot or intelligent assistive system has to assess human actions correctly. Based on this presumption, the task of online action detection determines the current action solely based on the past without access to future information. During inference, the performance of the model is largely impacted by the attributes of the underlying training dataset. However, as high costs and ethical concerns are associated with the real-world data collection process, synthetically created data provides a way to mitigate these problems while providing additional data for the training process of the underlying action detection model to improve performanceDue to the inherent domain shift between the synthetic and real data, we introduce a new egocentric dataset called Human Kitchen Interactions (HKI) to investigate the sim-to-real gap. Our dataset contains in total 100 synthetic and real videos in which 21 different actions are executed in a kitchen environment. The synthetic data is acquired in an egocentric virtual reality (VR) setup while capturing the virtual environment in a game engine. We evaluate state-of-the-art online action detection models on our dataset and provide insights into sim-to-real domain shift. Upon acceptance, we will release our dataset and the corresponding features at https://c-patsch.github.io/HKI/. Constantin Patsch, Wael Torjmene, Marsil Zakour, Driton Salihu, Eckehard G. Steinbach |
IROS | 5 |
| 2024 | Enhanced Robotic Assistance for Human Activities through Human-Object Interaction Segment PredictionabstractRobotic assistance is a current research topic with high application value and multiple challenges. Assistive robots are used in various scenarios, such as production lines, operating tables, and elderly care. While providing effective assistance, most of the assistance tasks that current robots can perform are limited to predefined tasks. This limitation arises from the insufficiency of the current robot perception system to forecast future human activities. To address this issue, we propose a novel 2-stage robotic assistant for human activities through future human-object interaction (HOI) segment prediction. Unlike previous work focusing on predefined or short-term tasks, our robotic assistant can make predictions for future assistance according to human habits. In the first stage, we propose a visual-based human-object interaction segment prediction method to predict human activities, which enables the robotic system to infer human intention. Moreover, we define the robotic executable tasks as an interactive tuple to keep the robotic assistance normatively consistent with human activity. Meanwhile, a graph convolutional network with geometric features that can predict human-object interaction segments is proposed to provide target manipulation and target object for the assistive robot. In the second stage, we present a mobile task completion process including visual navigation, object localization and grasping. The perception stage is evaluated on the MPHOI dataset and custom-collected SPHOI dataset. Finally, we evaluate our comprehensive framework through real-time experimentation. Rayene Messaoud, Arne-Christoph Hildebrandt, Marco Baldini, Driton Salihu, Constantin Patsch, Eckehard G. Steinbach |
IROS | 5 |
| 2024 | Rethinking 3D Geometric Object Features for Enhancing Skeleton-based Action RecognitionabstractHuman action recognition is crucial for intelligent robots, especially in the realm of human-robot collaboration research. Recent advancements in human pose estimation algorithms have shifted the focus of action recognition towards skeleton-based models, which exhibit robustness to changes in background and illumination. However, many state-of-the-art action recognition models rely on 2D skeleton data, neglecting object features. This limitation becomes obvious in complex scenarios where human interactions with objects are crucial, potentially compromising the reliability of assistive robots in understanding human behavior in their environment. To address this issue, we propose a method that effectively integrates 3D geometric object features into skeleton data using graph convolutional neural networks (GCNs). In addition to analyzing the effectiveness of information from different dimensions such as object center position, category, translation, and rotation, we explore various adjacency matrix designs for graph networks. Our model performance is evaluated on two challenging datasets: IKEA ASM and Bimanual Actions. The results demonstrate a significant improvement in action recognition by integrating object features into skeleton-based models. Specifically, on the IKEA-ASM dataset, our approach achieves a frame-wise Top-1 score improvement of 10.8% and an average F1@k improvement of 13.3%, while on the Bimanual Actions dataset, it achieves a frame-wise Top-1 score improvement of 11.4% and an average F1@k improvement of 5.3%, with negligible increases in model complexity. Driton Salihu, Constantin Patsch, Marsil Zakour, Eckehard G. Steinbach |
IROS | 3 |
| 2023 | COCCA: Point Cloud Completion through Cad Cross-Attentionabstract3D scene- and object-level scans typically result in sparse and incomplete point clouds. Since dense point clouds of high quality are essential for the 3D reconstruction process, a promising approach is to improve the scan quality by point cloud completion. In this paper, we present COCCA, an extension of point cloud completion networks for scan-to-CAD use cases. The proposed extension is based on cross-attention of features extracted from a scan with rotation-, translation-, and scale-invariant features extracted from a sampled CAD point cloud. With the proposed cross-attention operation, we improve the learning of scan features and the subsequent decoding to a complete shape. We demonstrate the effectiveness of COCCA on the ShapeNet dataset in quantitative and qualitative experiments. COCCA improves the overall completion performance of point cloud completion networks by up to 11.8% for Chamfer Distance and up to 2.2% for F-Score. Our qualitative experiments visualize how COCCA completes point clouds with higher geometric detail. In addition, we demonstrate how completion by COCCA improves the point cloud registration task required for scan-to-CAD alignment. Adam Misik, Driton Salihu, Heike Brock, Eckehard G. Steinbach |
ICIP | 2 |
| 2023 | Modeling Action Spatiotemporal Relationships Using Graph-Based Class-Level Attention Network for Long-Term Action DetectionabstractIn recent years, Action Detection has become an active research topic in various fields such as human-robot interaction and assistive robots. Most of the previous methods in this field focus on temporally processing the action representation, without considering the dependencies among the action classes. However, actions that occur in a video are constantly related, and this correlation could offer effective clues for detection tasks. In this work, we propose to exploit the information of related action classes with the help of a graph neural network in conjunction with temporal modeling. We introduce the attention-based temporal class module (ATC), which models the inherent action dependencies on the graph and learns action-specific features among temporal dimensions with a dual-branch attention mechanism. Further, we present the Graph-based Class-level Attention Network (GCAN), which is built upon ATC modules with increasing temporal receptive fields to handle actions instances in complex untrimmed videos. Our network is evaluated on two challenging benchmark datasets with dense annotations: Charades and MultiTHUMOS. Experimental results show that our approach demonstrates highly competitive results with a significantly reduced model complexity. Driton Salihu, Marsil Zakour, Constantin Patsch |
IROS | 3 |
| 2023 | SGPCR: Spherical Gaussian Point Cloud Representation and its Application to Object Registration and RetrievalabstractRetrieving and aligning CAD models from databases with scanned real-world point clouds remains an important topic for 3D reconstruction. Due to zero point-to-point correspondences between the sampled CAD model and the scanned real-world object, an information-rich representation of point clouds is needed. We propose SGPCR, a novel method for representing 3D point clouds by Spherical Gaussians for efficient, stable, and rotation-equivariant representation. We also propose a rotation-invariant convolution to improve the representation quality through a trainable optimization process. In addition, we demonstrate the strengths of SGPCR-based point cloud representation using the fundamental challenge of shape retrieval and point cloud registration on point clouds with zero point-to-point correspondences. Under these conditions, our approach improves registration quality by reducing chamfer distance by up to 90% and rotation root mean square error by up to 86% compared to the state of the art. Furthermore, the proposed SGCPR is used for one-shot shape retrieval and registration and improves retrieval precision by up to 58% over comparable methods. Driton Salihu, Eckehard G. Steinbach |
WACV | 1 |
| 2023 | Reconstruction with robustness: A semantic prior guided face super-resolution framework for multiple degradations
Hongjun Wu 0003, Huanrong Zhang, Zhi Jin 0002, Driton Salihu, Jianfang Hu |
Image Vis. Comput. | 5 |
| 2022 | AnaCoNGA: Analytical HW-CNN Co-Design Using Nested Genetic AlgorithmsabstractWe present AnaCoNGA, an analytical co-design methodology, which enables two genetic algorithms to evaluate the fitness of design decisions on layer-wise quantization of a neural network and hardware (HW) resource allocation. We embed a hardware architecture search (HAS) algorithm into a quantization strategy search (QSS) algorithm to evaluate the hardware design Pareto-front of each considered quantization strategy. We harness the speed and flexibility of analytical HW-modeling to enable parallel HW-CNN co-design. With this approach, the QSS is focused on seeking high-accuracy quantization strategies which are guaranteed to have efficient hardware designs at the end of the search. Through AnaCoNGA, we improve the accuracy by 2.88 p.p. with respect to a uniform 2-bit ResNet20 on CIFAR-10, and achieve a 35% and 37% improvement in latency and DRAM accesses, while reducing LUT and BRAM resources by 9% and 59% respectively, when compared to a standard edge variant of the accelerator. The nested genetic algorithm formulation also reduces the search time by 51% compared to an equivalent, sequential co-design formulation. Nael Fasfous, Manoj Rohit Vemparala, Alexander Frickenstein, Emanuele Valpreda, Driton Salihu, Julian Höfer, Anmol Singh, Naveen Shankar Nagaraja, Hans-Jörg Vögel, Nguyen Anh Vu Doan, Maurizio Martina, Jürgen Becker 0001, Walter Stechele |
DATE | 5 |
| 2022 | S2CMAF: Multi-Method Assessment Fusion for Scan-to-CAD MethodsabstractScan-to-CAD-based 3D reconstruction of indoor environments has become increasingly more popular in recent years. The inherent structure of Scan-to-CAD consists of object detection, model retrieval, and alignment. Therefore, a variety of metrics are required to assess these three aspects. This can lead to ambiguous evaluation results and incorrect quality assumptions. To impede the problem of incorrect evaluation, we introduce S2CMAF, a multi-method assessment fusion approach for Scan-to-CAD pipelines. S2CMAF merges several metrics used in evaluating these pipelines into one unique quality score. We show that S2CMAF significantly improves the correlation between Scan-to-CAD results and the ground truth, compared to the conventionally used Scan2CAD benchmark. Additionally, we train S2CMAF using different optimization techniques and demonstrate the advantages of our approach on real-world data. Driton Salihu, Adam Misik, Markus Hofbauer, Eckehard G. Steinbach |
ISM | 1 |
| 2021 | Hardware-Aware Mixed-Precision Neural Networks using In-Train Quantization
Manoj Rohit Vemparala, Nael Fasfous, Lukas Frickenstein, Alexander Frickenstein, Anmol Singh, Driton Salihu, Christian Unger, Naveen Shankar Nagaraja, Walter Stechele |
BMVC | 6 |
| 2021 | HW-FlowQ: A Multi-Abstraction Level HW-CNN Co-design Quantization MethodologyabstractModel compression through quantization is commonly applied to convolutional neural networks (CNNs) deployed on compute and memory-constrained embedded platforms. Different layers of the CNN can have varying degrees of numerical precision for both weights and activations, resulting in a large search space. Together with the hardware (HW) design space, the challenge of finding the globally optimal HW-CNN combination for a given application becomes daunting. To this end, we propose HW-FlowQ, a systematic approach that enables the co-design of the target hardware platform and the compressed CNN model through quantization. The search space is viewed at three levels of abstraction, allowing for an iterative approach for narrowing down the solution space before reaching a high-fidelity CNN hardware modeling tool, capable of capturing the effects of mixed-precision quantization strategies on different hardware architectures (processing unit counts, memory levels, cost models, dataflows) and two types of computation engines (bit-parallel vectorized, bit-serial). To combine both worlds, a multi-objective non-dominated sorting genetic algorithm (NSGA-II) is leveraged to establish a Pareto-optimal set of quantization strategies for the target HW-metrics at each abstraction level. HW-FlowQ detects optima in a discrete search space and maximizes the task-related accuracy of the underlying CNN while minimizing hardware-related costs. The Pareto-front approach keeps the design space open to a range of non-dominated solutions before refining the design to a more detailed level of abstraction. With equivalent prediction accuracy, we improve the energy and latency by 20% and 45% respectively for ResNet56 compared to existing mixed-precision search methods. Nael Fasfous, Manoj Rohit Vemparala, Alexander Frickenstein, Emanuele Valpreda, Driton Salihu, Nguyen Anh Vu Doan, Christian Unger, Naveen Shankar Nagaraja, Maurizio Martina, Walter Stechele |
ACM Trans. Embed. Comput. Syst. | 5 |