VLDB 2026 Research / reviewers in the wild / expert
Marc Comino
dblp:198/8982 · also Marc Comino-Trinidad
· DBLP profile ↗
16ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0001-5621-7565ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 3 first-author · 12 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PHYSPLAT: A Framework for Photorealistic Hybrid Simulation of Real and Synthetic Elements using 3D Gaussian Splatting
Mario Alfonso-Arsuaga, Henar Dominguez-Elvira, Jorge Casas-Guerrero, Andrea Castiella-Aguirrezabala, Lorenzo Costabile-Dominguez, Jorge García-González 0001, Maria Naranjo-Almeida, Marc Comino, Jorge Lopez-Moreno |
WACV | 8 |
| 2026 | Bridging BIM and reality: A hardware-optimized registration pipeline for Mixed Reality in indoor construction environmentsabstractBuilding Information Modeling (BIM) has transformed the Architecture, Engineering, and Construction (AEC) industry by digitizing project data, yet its full potential remains unrealized due to persistent gaps between virtual models and physical sites. These gaps contribute to inefficiencies, with studies reporting substantial waste in labor and coordination. Extended Reality (XR) technologies offer a promising solution by enabling immersive, real-scale visualization of BIM models on-site. This article introduces a Mixed Reality (MR) application for Microsoft HoloLens 2 that superimposes BIM representations onto construction environments at a 1:1 scale, supporting real-time detection of differences between the as-designed BIM model and the as-built construction on site. We present a robust registration pipeline that integrates commercial XR hardware with advanced algorithms to achieve precise alignment under challenging conditions. To validate the system, we conducted a controlled user study comparing three registration paradigms (manual gesture-based, QR-assisted, and fully automatic) and analyzing their impact on alignment accuracy and user experience (UX) in AEC-related tasks. Results show that our automatic approach provides advantages over state-of-the-art alternatives and significantly improves registration precision and usability ratings over the baseline methods. Furthermore, the study demonstrates that alignment errors strongly influence spatial perception and decision-making, highlighting the necessity of high-fidelity registration for effective MR integration in construction workflows. Marcos Arroyo-Ruiz, Gonzalo Gomez-Nogales, José Antonio Gómez-Hernández, Carlos Andújar, Marc Comino |
Comput. Graph. | 5 |
| 2026 | GNOCHI: Generative Neural mOdel for Close Human-Human InteractionsabstractAbstract Creating realistic 3D human‐human interactions in virtual environments is challenging due to the high degrees of freedom in human body and the need for physically accurate poses that do not collide with each other. Traditional methods for humanhuman interaction are based on motion tracking or 3D body reconstruction, but lack generative capabilities. Recent generative methods enable the synthesis of individual or interacting motions via text or image input, but generally fall short in modeling close interactions. This paper introduces a novel generative model for close 3D human‐human interactions using a conditional variational autoencoder (cVAE), which generates poses for one human conditioned on the pose of another, allowing for controlled and diverse interaction synthesis. To train our model, we address two underlying long‐standing challenges in the field of human‐human interaction: data scarcity, for which we propose an automated supervised data augmentation strategy that generates synthetic yet realistic interaction poses; and collision awareness in generative approaches, for which we propose a self‐supervised loss based on a collision resolution technique using volumetric proxies to ensure physically correct interactions. We extensively evaluate the capabilities of our model, and demonstrate a wide variety of plausible and physically correct interactions, not possible to generate with current state‐of‐the‐art methods. Gonzalo Gomez-Nogales, Marc Comino, Andrés Casado-Elvira, Dan Casas |
Comput. Graph. Forum | 2 |
| 2025 | 3DGStrands: Personalized 3D Gaussian splatting for realistic hair representation and animationabstractWe introduce a novel method for generating a personalized 3D Gaussian Splatting (3DGS) hair representation from an unorganized set of photographs. Our approach begins by leveraging an out-of-the-shelf method to estimate a strand-organized point cloud representation of the hair. This point cloud serves as the foundation for constructing a 3DGS model that accurately preserves the hair’s geometric structure while visually fitting the appearance in the photographs. Our model seamlessly integrates with the standard 3DGS rendering pipeline, enabling efficient volumetric rendering of complex hairstyles. Furthermore, we demonstrate the versatility of our approach by applying the Material Point Method (MPM) to simulate realistic hair physics directly on the 3DGS model, achieving lifelike hair animation. To the best of our knowledge, this is the first method to simulate hair dynamics within a 3DGS model . This work paves the way for future research that can leverage the flexible nature of 3DGS to fit more complex hair material models or enable physics properties estimation through dynamic tracking. Henar Dominguez-Elvira, Mario Alfonso-Arsuaga, Ana Barrueco-Garcia, Marc Comino |
Comput. Graph. | 4 |
| 2025 | Detecting anomalies in dense 3D crowdsabstractEstimating the behavior of dense 3D crowds is crucial for applications in security, surveillance, and planning. Detecting events in such crowds from a single video, the most common scenario, is challenging due to ambiguities, occlusions, and complex human behavior. To address this, we propose a method that overlays pixel-based labels on video data to highlight anomalies in dense 3D crowds movement. Our key contribution is a data-driven, image-based model trained on features derived from 3D virtual crowd animations of articulated characters that mimic real crowds at a micro-level. By using training data based on captured dense crowd trajectories and realistic 3D motions, we can analyze and detect anomalies in complex real-world scenarios. Additionally, while acquiring ground-truth data from diverse viewpoints is difficult in real-world settings, our virtual simulator allows rendering scenes from multiple perspectives, enabling the training of models robust to viewpoint variations. We demonstrate qualitatively and quantitatively that our method can detect anomalies in much denser crowds than existing methods. Melania Prieto-Martín, Marc Comino, Dan Casas |
Comput. Graph. | 2 |
| 2025 | Accurate hand contact detection from RGB images via image-to-image translationabstractHand tracking is a growing research field that can potentially provide a natural interface to interact with virtual environments. However, despite the impressive recent advances, the 3D tracking of two interacting hands from RGB video remains an open problem. While current methods are able to infer the 3D pose of two hands in interaction reasonably, residual errors in depth, shape, and pose estimation prevent the accurate detection of hand-to-hand contact. To mitigate these errors, in this paper, we propose an image-based data-driven method to estimate the contact in hand-to-hand interactions. Our method is built on top of 3D hand trackers that predict the articulated pose of two hands, enriching them with camera-space probability maps of contact points. To train our method, we first feed motion capture data of interacting hands into a physics-based hand simulator, and compute dense 3D contact points. We then render such contact maps from various viewpoints and create a dataset of pairs of pixel-to-surface hand images and their corresponding contact labels. Finally, we train an image-to-image network that learns to translate pixel-to-surface correspondences to contact maps. At inference time, we estimate pixel-to-surface correspondences using state-of-the-art hand tracking and then use our network to predict accurate hand-to-hand contact. We qualitatively and quantitatively validate our method in real-world data and demonstrate that our contact predictions are more accurate than state-of-the-art hand-tracking methods. Suzanne Sorli, Marc Comino, Dan Casas |
Comput. Graph. | 2 |
| 2024 | Resolving Collisions in Dense 3D Crowd AnimationsabstractWe propose a novel contact-aware method to synthesize highly-dense 3D crowds of animated characters. Existing methods animate crowds by, first, computing the 2D global motion approximating subjects as 2D particles and, then, introducing individual character motions without considering their surroundings. This creates the illusion of a 3D crowd, but, with density, characters frequently intersect each other since character-to-character contact is not modeled. We tackle this issue and propose a general method that considers any crowd animation and resolves existing residual collisions. To this end, we take a physics-based approach to model contacts between articulated characters. This enables the real-time synthesis of 3D high-density crowds with dozens of individuals that do not intersect each other, producing an unprecedented level of physical correctness in animations. Under the hood, we model each individual using a parametric human body incorporating a set of 3D proxies to approximate their volume. We then build a large system of articulated rigid bodies, and use an efficient physics-based approach to solve for individual body poses that do not collide with each other while maintaining the overall motion of the crowd. We first validate our approach objectively and quantitatively. We then explore relations between physical correctness and perceived realism based on an extensive user study that evaluates the relevance of solving contacts in dense crowds. Results demonstrate that our approach outperforms existing methods for crowd animation in terms of geometric accuracy and overall realism. Gonzalo Gomez-Nogales, Melania Prieto-Martín, Cristian Romero, Marc Comino, Pablo Ramon-Prieto, Anne-Hélène Olivier, Ludovic Hoyet, Miguel A. Otaduy, Julien Pettré, Dan Casas |
ACM Trans. Graph. | 4 |
| 2023 | SMPLitex: A Generative Model and Dataset for 3D Human Texture Estimation from Single Image
Dan Casas, Marc Comino |
BMVC | 2 |
| 2023 | Foreword to the Special Section on CEIG 2023
Ana Serrano, Marc Comino, Jesús Gimeno |
Comput. Graph. | 2 |
| 2022 | Sweep Encoding: Serializing Space Subdivision Schemes for Optimal SlicingabstractSlicing a model (computing thin slices of a geometric or volumetric model with a sweeping plane) is necessary for several applications ranging from 3D printing to medical imaging. This paper introduces a technique designed to compute these slices efficiently, even for huge and complex models. We voxelize the volume of the model at a required resolution and show how to encode this voxelization in an out-of-core octree using a novel Sweep Encoding linearization. This approach allows for efficient slicing with bounded cost per slice. We discuss specific applications, including 3D printing, and compare these octrees’ performance against the standard representations in the literature. Marc Comino, Àlvar Vinacua, A. Carruesco, Antoni Chica, Pere Brunet |
Comput. Aided Des. | 1 |
| 2022 | Gain compensation across LIDAR scansabstractHigh-end Terrestrial Lidar Scanners are often equipped with RGB cameras that are used to colorize the point samples. Some of these scanners produce panoramic HDR images by encompassing the information of multiple pictures with different exposures. Unfortunately, exported RGB color values are not in an absolute color space, and thus point samples with similar reflectivity values might exhibit strong color differences depending on the scan the sample comes from. These color differences produce severe visual artifacts if, as usual, multiple point clouds colorized independently are combined into a single point cloud. In this paper we propose an automatic algorithm to minimize color differences among a collection of registered scans. The basic idea is to find correspondences between pairs of scans, i.e. surface patches that have been captured by both scans. If the patches meet certain requirements, their colors should match in both scans. We build a graph from such pair-wise correspondences, and solve for the gain compensation factors that better uniformize color across scans. The resulting panoramas can be used to colorize the point clouds consistently. We discuss the characterization of good candidate matches, and how to find such correspondences directly on the panorama images instead of in 3D space. We have tested this approach to uniformize color across scans acquired with a Leica RTC360 scanner, with very good results. Imanol Muñoz-Pandiella, Marc Comino, Carlos Andújar, Oscar Argudo, Carles Bosch, Antoni Chica, Beatriz Martínez 0003 |
Comput. Graph. | 2 |
| 2022 | PERGAMO: Personalized 3D Garments from Monocular VideoabstractAbstract Clothing plays a fundamental role in digital humans. Current approaches to animate 3D garments are mostly based on realistic physics simulation, however, they typically suffer from two main issues: high computational run‐time cost, which hinders their deployment; and simulation‐to‐real gap, which impedes the synthesis of specific real‐world cloth samples. To circumvent both issues we propose PERGAMO, a data‐driven approach to learn a deformable model for 3D garments from monocular images. To this end, we first introduce a novel method to reconstruct the 3D geometry of garments from a single image, and use it to build a dataset of clothing from monocular videos. We use these 3D reconstructions to train a regression model that accurately predicts how the garment deforms as a function of the underlying body pose. We show that our method is capable of producing garment animations that match the real‐world behavior, and generalizes to unseen body motions extracted from motion capture dataset. Andrés Casado-Elvira, Marc Comino, Dan Casas |
Comput. Graph. Forum | 2 |
| 2019 | Multi-View Image FusionabstractWe present an end-to-end learned system for fusing multiple misaligned photographs of the same scene into a chosen target view. We demonstrate three use cases: 1) color transfer for inferring color for a monochrome view, 2) HDR fusion for merging misaligned bracketed exposures, and 3) detail transfer for reprojecting a high definition image to the point of view of an affordable VR180-camera. While the system can be trained end-to-end, it consists of three distinct steps: feature extraction, image warping and fusion. We present a novel cascaded feature extraction method that enables us to synergetically learn optical flow at different resolution levels. We show that this significantly improves the network's ability to learn large disparities. Finally, we demonstrate that our alignment architecture outperforms a state-of-the art optical flow network on the image warping task when both systems are trained in an identical manner. Marc Comino, Ricardo Martin-Brualla, Florian Kainz, Janne Kontkanen |
ICCV | 1 |
| 2018 | Segmentation of aerial images for plausible detail synthesis
Oscar Argudo, Marc Comino, Antoni Chica, Carlos Andújar, Felipe Lumbreras |
Comput. Graph. | 2 |
| 2018 | Sensor-aware Normal Estimation for Point Clouds from 3D Range ScansabstractAbstract Normal vectors are essential for many point cloud operations, including segmentation, reconstruction and rendering. The robust estimation of normal vectors from 3D range scans is a challenging task due to undersampling and noise, specially when combining points sampled from multiple sensor locations. Our error model assumes a Gaussian distribution of the range error with spatially‐varying variances that depend on sensor distance and reflected intensity, mimicking the features of Lidar equipment. In this paper we study the impact of measurement errors on the covariance matrices of point neighborhoods. We show that covariance matrices of the true surface points can be estimated from those of the acquired points plus sensor‐dependent directional terms. We derive a lower bound on the neighbourhood size to guarantee that estimated matrix coefficients will be within a predefined error with a prescribed probability. This bound is key for achieving an optimal trade‐off between smoothness and fine detail preservation. We also propose and compare different strategies for handling neighborhoods with samples coming from multiple materials and sensors. We show analytically that our method provides better normal estimates than competing approaches in noise conditions similar to those found in Lidar equipment. Marc Comino, Carlos Andújar, Antoni Chica, Pere Brunet |
Comput. Graph. Forum | 1 |
| 2017 | Error-aware construction and rendering of multi-scan panoramas from massive point clouds
Marc Comino, Carlos Andújar, Antoni Chica, Pere Brunet |
Comput. Vis. Image Underst. | 1 |