Constantin Patsch

dblp:315/6962 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
11since 2021 · last 2025
0000-0002-7546-6732ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2025 MistSense: Versatile Online Detection of Procedural and Execution Mistakes
Constantin Patsch, Marsil Zakour, Driton Salihu, Eckehard G. Steinbach
ICCV1
2024 Long-Term Action Anticipation Based on Contextual Alignment
abstract
In action anticipation, the model predicts the next future action after a certain observation period. In long-term action anticipation, this idea is further extended to predicting multiple actions and their respective duration. Thus, in this problem setting the model should not only capture relationships between past actions but also predict several future actions that fit into a certain context. Compared to autoregressive models, our model employs an encoder decoder structure to determine future actions and durations in parallel, which prevents the accumulation of prediction errors and reduces the inference time. Furthermore, it is ensured that the predicted actions are aligned with respect to a context representation, which resembles the way humans approach this task as the feasible action set is restricted by the respective context. We evaluate our model on the long-term anticipation benchmark datasets, Breakfast, and 50Salads, where we achieve state-of-the-art results.
Constantin Patsch, Jinghan Zhang 0009, Marsil Zakour, Driton Salihu, Eckehard G. Steinbach
ICASSP1
2024 NPRF: Neural Painted Radiosity Fields for Neural Implicit Rendering and Surface Reconstruction
abstract
In recency, neural signed distance fields have become more popular for reconstructing 3D indoor environments. While great improvements have been made due to missing incident radiance and materials in the surface estimation, current methods cannot reconstruct high-quality surfaces. To address this issue, we propose Neural Painted Radiosity Fields (NPRF), consisting of Neural Radiosity Fields for volumetric surface representation and Neural Painted Scenes for novel view synthesis. Neural Radiosity Fields combine the radiative transfer equation with neural radiosity to estimate 3D surfaces, thus leveraging raytracing to improve the volumetric representation. Neural Painted Scenes employs sparsification and projection of 3D points into 2D images in conjunction with a generative, context-aware inpainting network to produce high-quality novel views. We show that NPRF leads to overall improvements in F-score on the popular ScanNet dataset. Finally, we show that NPRF improves novel view synthesis by a significant margin, giving improvements of up to 25% on PSNR, 53% on LPIPS, and 3% on SSIM.
Driton Salihu, Adam Misik, Constantin Patsch, Eckehard G. Steinbach
ICASSP4
2024 DeepSPF: Spherical SO(3)-Equivariant Patches for Scan-to-CAD Estimation
abstract
Recently, SO(3)-equivariant methods have been explored for 3D reconstruction via Scan-to-CAD. Despite significant advancements attributed to the unique characteristics of 3D data, existing SO(3)-equivariant approaches often fall short in seamlessly integrating local and global contextual information in a widely generalizable manner. Our contributions in this paper are threefold. First, we introduce Spherical Patch Fields, a representation technique designed for patch-wise, SO(3)-equivariant 3D point clouds, anchored theoretically on the principles of Spherical Gaussians. Second, we present the Patch Gaussian Layer, designed for the adaptive extraction of local and global contextual information from resizable point cloud patches. Culminating our contributions, we present Learnable Spherical Patch Fields (DeepSPF) – a versatile and easily integrable backbone suitable for instance-based point networks. Through rigorous evaluations, we demonstrate significant enhancements in Scan-to-CAD performance for point cloud registration, retrieval, and completion: a significant reduction in the rotation error of existing registration methods, an improvement of up to 17\% in the Top-1 error for retrieval tasks, and a notable reduction of up to 30\% in the Chamfer Distance for completion models, all attributable to the incorporation of DeepSPF.
Driton Salihu, Adam Misik, Constantin Patsch, Fabián Seguel, Eckehard G. Steinbach
ICLR4
2024 Sim-to-Real Domain Shift in Online Action Detection
abstract
Human reasoning comprises the ability to understand and reason about the current action solely based on past information. To provide effective assistance in an eldercare or household environment an assistive robot or intelligent assistive system has to assess human actions correctly. Based on this presumption, the task of online action detection determines the current action solely based on the past without access to future information. During inference, the performance of the model is largely impacted by the attributes of the underlying training dataset. However, as high costs and ethical concerns are associated with the real-world data collection process, synthetically created data provides a way to mitigate these problems while providing additional data for the training process of the underlying action detection model to improve performanceDue to the inherent domain shift between the synthetic and real data, we introduce a new egocentric dataset called Human Kitchen Interactions (HKI) to investigate the sim-to-real gap. Our dataset contains in total 100 synthetic and real videos in which 21 different actions are executed in a kitchen environment. The synthetic data is acquired in an egocentric virtual reality (VR) setup while capturing the virtual environment in a game engine. We evaluate state-of-the-art online action detection models on our dataset and provide insights into sim-to-real domain shift. Upon acceptance, we will release our dataset and the corresponding features at https://c-patsch.github.io/HKI/.
Constantin Patsch, Wael Torjmene, Marsil Zakour, Driton Salihu, Eckehard G. Steinbach
IROS1
2024 Enhanced Robotic Assistance for Human Activities through Human-Object Interaction Segment Prediction
abstract
Robotic assistance is a current research topic with high application value and multiple challenges. Assistive robots are used in various scenarios, such as production lines, operating tables, and elderly care. While providing effective assistance, most of the assistance tasks that current robots can perform are limited to predefined tasks. This limitation arises from the insufficiency of the current robot perception system to forecast future human activities. To address this issue, we propose a novel 2-stage robotic assistant for human activities through future human-object interaction (HOI) segment prediction. Unlike previous work focusing on predefined or short-term tasks, our robotic assistant can make predictions for future assistance according to human habits. In the first stage, we propose a visual-based human-object interaction segment prediction method to predict human activities, which enables the robotic system to infer human intention. Moreover, we define the robotic executable tasks as an interactive tuple to keep the robotic assistance normatively consistent with human activity. Meanwhile, a graph convolutional network with geometric features that can predict human-object interaction segments is proposed to provide target manipulation and target object for the assistive robot. In the second stage, we present a mobile task completion process including visual navigation, object localization and grasping. The perception stage is evaluated on the MPHOI dataset and custom-collected SPHOI dataset. Finally, we evaluate our comprehensive framework through real-time experimentation.
Rayene Messaoud, Arne-Christoph Hildebrandt, Marco Baldini, Driton Salihu, Constantin Patsch, Eckehard G. Steinbach
IROS6
2024 Rethinking 3D Geometric Object Features for Enhancing Skeleton-based Action Recognition
abstract
Human action recognition is crucial for intelligent robots, especially in the realm of human-robot collaboration research. Recent advancements in human pose estimation algorithms have shifted the focus of action recognition towards skeleton-based models, which exhibit robustness to changes in background and illumination. However, many state-of-the-art action recognition models rely on 2D skeleton data, neglecting object features. This limitation becomes obvious in complex scenarios where human interactions with objects are crucial, potentially compromising the reliability of assistive robots in understanding human behavior in their environment. To address this issue, we propose a method that effectively integrates 3D geometric object features into skeleton data using graph convolutional neural networks (GCNs). In addition to analyzing the effectiveness of information from different dimensions such as object center position, category, translation, and rotation, we explore various adjacency matrix designs for graph networks. Our model performance is evaluated on two challenging datasets: IKEA ASM and Bimanual Actions. The results demonstrate a significant improvement in action recognition by integrating object features into skeleton-based models. Specifically, on the IKEA-ASM dataset, our approach achieves a frame-wise Top-1 score improvement of 10.8% and an average F1@k improvement of 13.3%, while on the Bimanual Actions dataset, it achieves a frame-wise Top-1 score improvement of 11.4% and an average F1@k improvement of 5.3%, with negligible increases in model complexity.
Driton Salihu, Constantin Patsch, Marsil Zakour, Eckehard G. Steinbach
IROS4
2024 CSPN: A Category-Specific Processing Network for Low-Light Image Enhancement
abstract
Images captured in low-light conditions usually suffer from degradation problems. Recently, numerous deep learning-based methods are proposed for low-light image enhancement. They either focus on performance improvement with negligence of computational complicity, or are extremely computationally efficient networks with poor performance. In this work, we intend to figure out a solution, which strikes a balance between computational cost and performance. Moreover, we observe that different regions of an image contain different amounts of information, where the region with less information is easier to restore than that with more information. Hence, we propose to crop a low-light image into patches and classify these patches into “simple”, “medium” and “hard” categories based on their involved information. Then, we enhance different patch categories with different network complexities, therefore, a Category-specific Processing Network (CSPN) is proposed to achieve the computational complexity and performance balance. The patch classification is implemented by the proposed Grey-Level Co-occurrence Matrix (GLCM) entropy-based algorithm, which measures the content complexity of an image by analyzing the statistics of the difference between pixels. As the frequency domain contains exclusive feature information that is beneficial for improving image quality, the wavelet transform is introduced during the enhancement. Extensive experimental results demonstrate the superiority of our proposed CSPN over other state-of-the-art methods in various datasets with the least amount of computational cost.
Hongjun Wu 0003, Luwei Tu, Constantin Patsch, Zhi Jin 0002
IEEE Trans. Circuits Syst. Video Technol.4
2023 Self-Attention Based Action Segmentation Using Intra-And Inter-Segment Representations
abstract
Segmenting activities in untrimmed videos remains a critical challenge to fully understand complex human activity sequences. A correct representation of temporal action relations is key for improving incorrect segmentations. We propose a self-attention-based model that refines initial segmentations by separately considering intra-as well as inter-segment relations between predicted action segments. Furthermore, in order to enhance the training process, we use a similarity-guided regularization technique that ensures intra-segment similarity and the validity of action transitions between adjacent segments. In an extensive evaluation on three public datasets -Georgia Tech Egocentric Activities, 50Salads, and Breakfast -our proposed architecture enhances the backbone model by 6.1% on GTEA, 3.8% on 50Salads, and 3.9% on Breakfast with regard to the F 1@50 metric.
Constantin Patsch, Eckehard G. Steinbach
ICASSP1
2023 Modeling Action Spatiotemporal Relationships Using Graph-Based Class-Level Attention Network for Long-Term Action Detection
abstract
In recent years, Action Detection has become an active research topic in various fields such as human-robot interaction and assistive robots. Most of the previous methods in this field focus on temporally processing the action representation, without considering the dependencies among the action classes. However, actions that occur in a video are constantly related, and this correlation could offer effective clues for detection tasks. In this work, we propose to exploit the information of related action classes with the help of a graph neural network in conjunction with temporal modeling. We introduce the attention-based temporal class module (ATC), which models the inherent action dependencies on the graph and learns action-specific features among temporal dimensions with a dual-branch attention mechanism. Further, we present the Graph-based Class-level Attention Network (GCAN), which is built upon ATC modules with increasing temporal receptive fields to handle actions instances in complex untrimmed videos. Our network is evaluated on two challenging benchmark datasets with dense annotations: Charades and MultiTHUMOS. Experimental results show that our approach demonstrates highly competitive results with a significantly reduced model complexity.
Driton Salihu, Marsil Zakour, Constantin Patsch
IROS6
2022 Automatic Recognition of Human Activities Combining Model-based AI and Machine Learning
Constantin Patsch, Marsil Zakour, Rahul Gopal Chaudhari
ICAART (3)1