Marco Agus

dblp:13/6215 · DBLP profile ↗
← Back
38ranked-venue papers
8as first author
19since 2021 · last 2026
0000-0003-2752-3525ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 36 · 8 first-author · 17 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 NeMoCo: Self-supervised contrastive learning for ultrastructural 3D neuroscience morphologies
abstract
Volume electron microscopy (EM) now enables nanometric-scale 3D reconstructions of neural tissue, opening the door to quantitative, morphology-driven neuroscience beyond connectivity alone. While previous studies relied on handcrafted descriptors and classical machine learning for morphology analysis, recent progress in deep learning for 3D shape understanding offers new opportunities to learn robust, task-specific representations directly from geometric data. In this paper we present NeMoCo , a geometry learning framework that targets the key practical bottleneck in connectomics and ultrastructural analysis: the scarcity and cost of dense expert annotations for the long tail of neurite and organelle phenotypes. NeMoCo formulates representation learning for EM-derived neurite meshes in a self-supervised Momentum Contrast (MoCo) style. We use DiffusionNet (Sharp et al., 2022) as a mesh encoder with intrinsic spectral descriptors (HKS) and train with a momentum-updated teacher encoder and a large memory bank of negatives. To learn invariances that are essential in practice, we generate paired geometric views via controlled affine transformations and resolution changes (including mesh decimation), encouraging embeddings to be stable under nuisance variability while remaining discriminative. We provide an extensive study of augmentation strength and temperature, and evaluate learned representations through frozen retrieval and non-parametric classification (frozen kNN), as well as downstream supervised fine-tuning under limited labels. NeMoCo demonstrates that MoCo-style self-supervision yields robust neurite morphology embeddings on EM meshes, improving label-efficiency and offering a scalable foundation for retrieval, clustering, and phenotype discovery in ultrastructural neuroscience. All the data and the code used for NeMoCo are available at https://github.com/Uzshah/NeMoCo .
Humaira Shaffique, Uzair Shah, Mahmood Alzubaidi, Jens Schneider 0002, Pierre J. Magistretti, Corrado Calì, Mowafa Househ, Marco Agus
Graph. Model.8
2025 Towards Culturally Aware AI: Quantified Evaluation of Relevance and Similarity in AI-Generated Images
Wala Elsharif, Mahmood Alzubaidi, James She, Marco Agus
CGI (3)4
2025 PanoFloor: Reconstruction and Immersive Exploration of Large Multi-Room Scenes from a Minimal Set of Registered Panoramic Images Using Denoised Density Maps
abstract
We introduce a deep learning approach to automatically generate 3D floor plans and immersive multi-room virtual visit experiences from a small set of co-registered 360° panoramas - down to just one per room. We integrate novel neural networks that leverage panoramic image broad context and large annotated room datasets to build a geometric and visual graph. Nodes represent stereo-viewable multiple-center-of-projection (MCOP) 360° images at the capture locations, while arcs connect them with paths through doors, avoiding clutter and minimizing disocclusions to maximize visual quality. The process starts with depth prediction and floor-plan projection to create a comprehensive but noisy global density map, which is refined via a latent diffusion model. A segmentation network then extracts room layouts, openings, and clutter. This structured representation is lifted to a visual one by creating a 360° stereo-explorable MCOP representation at each node, produced using a view-synthesis network from the original image and its predicted depth map. Arc paths are then computed using an optimization process that considers structural constraints, including openings and obstacles, while minimizing visual discontinuities, occlusions, and disocclusions. Finally, 360° video transitions are synthesized using a specialized view-synthesis network to obtain a fully precomputed WebXR-ready explorable representation that can be efficiently experienced on Head-Mounted-Displays with limited graphics capabilities. The extracted floor plan not only aids in documenting the captured building but can also enhance immersive experiences by serving as a live map of the building. Our experiments show that the method achieves state-of-the-art reconstruction from sparse inputs and supports compelling immersive visits.
Giovanni Pintore, Sara Jashari, Marco Agus, Enrico Gobbetti
ISMAR3
2025 Deep learning for brain electron microscopy segmentation: Advances, challenges, and future directions in connectomics and ultrastructure analysis
abstract
This systematic review and meta-analysis comprehensively analyzes deep learning approaches for brain electron microscopy (EM) segmentation, addressing the critical challenge of extracting neuroanatomical information at nanometer resolution. Following PRISMA guidelines, we identified 60 studies through structured database searches, with quantitative meta-analysis of 27 studies (46 experiments) across 10 datasets providing the first unified benchmark comparison in this domain. Our analysis reveals a field transitioning from traditional CNN approaches toward foundation models and hybrid architectures. The meta-analysis demonstrates that foundation models outperform traditional CNNs by 13%–35% across key metrics, with the 3D Transformer + U-Net achieving the highest composite score (0.954) across five datasets. Meta-analysis confirms significant advantages for foundation models in instance-based metrics (Cohen’s d = − 6 . 44 ), while only 26% of experiments validate across multiple datasets. Four key evolutionary trends emerge: (1) transition from 2D to 3D architectures optimized for ultrastructural complexity; (2) development of topology-preserving loss functions and evaluation metrics (clDice, ERL) that prioritize neural connectivity over pixel-wise accuracy; (3) emergence of self-supervised and foundation model adaptation techniques reducing annotation dependency; and (4) evolution toward specialized architectures capturing long-range dependencies critical for neural structures. Performance analysis reveals that mitochondria segmentation achieves highest accuracy (Jaccard scores 87.2–90.5%), while computational requirements vary from single-GPU implementations to distributed systems with 48 GPUs for teravoxel-scale volumes. Despite progress, reproducibility challenges persist with only 54% of studies providing public code repositories. These advances drive innovation in 3D computer vision, establish new benchmarks for volumetric instance segmentation, and address fundamental challenges in processing massive biological datasets. Our unified benchmarks and comprehensive analysis provide a foundation for systematic progress tracking and evidence-based method selection, positioning brain EM segmentation to enable large-scale connectomics studies and detailed neuroanatomical mapping across scales.
Uzair Shah, Mahmood Alzubaidi, Marco Agus, Corrado Calì, Pierre J. Magistretti, Mowafa Househ
Comput. Graph.3
2025 AI-guided immersive exploration of brain ultrastructure for collaborative analysis and education
abstract
We introduce NeuroVerse, a framework for exploring 3D nanometric-scale reconstructions of neural and glial cellular processes in the central nervous system. Using image stacks from volume electron microscopy, NeuroVerse generates 3D mesh models through a SAM2-based segmentation pipeline and integrates absorption signals for deployment in a Metaverse environment. The framework includes a SAM2 adapter optimized for biological microscopy imaging, adapted with feature enhancement blocks and dual decoders to improve the segmentation of complex cellular structures. An interactive virtual AI agent, powered by Heygen and OpenAI models with domain-specific knowledge, provides semi-real-time assistance. NeuroVerse supports education and collaborative analysis for neuroanatomy and neuroscience. It includes a pipeline for the creation of 3D models, automated segmentation, mesh reconstruction, and heatmap computation, optimized for the Spatial.io ecosystem. Contributions include a virtual anatomy lab for neuroanatomy education and collaborative sessions on spatial morphology correlation and neuroenergetic absorption models. Evaluations show that the SAM2 adapter preserves fine cellular details and manages irregular boundaries. Preliminary sessions indicate potential to enhance neuroscience education, improve remote collaboration among scientists, and provide access to advanced neuroscientific data and tools. Evaluation of the virtual AI agent confirms its ability to provide context-aware support, interpret complex cellular structures, and facilitate understanding through semi-real-time assistance for students analyzing neural and glial reconstructions. NeuroVerse combines imaging, segmentation, and AI technologies within an immersive Metaverse platform for neuroscience education and research. • Development of digital twin of the Human Anatomy Institute of the University of Turin. • AI-Based pipeline for EM image segmentation and Data interpretation. • Case study for multiple usage for Data Analysis and Education in the Metaverse.
Uzair Shah, Marco Agus, Daniya Boges, Hamad Aldous, Vanessa Chiappini, Mahmood Alzubaidi, Markus Hadwiger, Pierre J. Magistretti, Mowafa Househ, Corrado Calì
Comput. Graph.2
2025 DDD++: Exploiting Density map consistency for Deep Depth estimation in indoor environments
abstract
We introduce a novel deep neural network designed for fast and structurally consistent monocular 360° depth estimation in indoor settings. Our model generates a spherical depth map from a single gravity-aligned or gravity-rectified equirectangular image, ensuring the predicted depth aligns with the typical depth distribution and structural features of cluttered indoor spaces, which are generally enclosed by walls, floors, and ceilings. By leveraging the distinctive vertical and horizontal patterns found in man-made indoor environments, we propose a streamlined network architecture that incorporates gravity-aligned feature flattening and specialized vision transformers. Through flattening, these transformers fully exploit the omnidirectional nature of the input without requiring patch segmentation or positional encoding. To further enhance structural consistency, we introduce a novel loss function that assesses density map consistency by projecting points from the predicted depth map onto a horizontal plane and a cylindrical proxy. This lightweight architecture requires fewer tunable parameters and computational resources than competing methods. Our comparative evaluation shows that our approach improves depth estimation accuracy while ensuring greater structural consistency compared to existing methods. For these reasons, it promises to be suitable for incorporation in real-time solutions, as well as a building block in more complex structural analysis and segmentation methods.
Giovanni Pintore, Marco Agus, Alberto Signoroni, Enrico Gobbetti
Graph. Model.2
2025 Autocleandeepfood: auto-cleaning and data balancing transfer learning for regional gastronomy food computing
abstract
Abstract Food computing has emerged as a promising research field, employing artificial intelligence, deep learning, and data science methodologies to enhance various stages of food production pipelines. To this end, the food computing community has compiled a variety of data sets and developed various deep-learning architectures to perform automatic classification. However, automated food classification presents a significant challenge, particularly when it comes to local and regional cuisines, which are often underrepresented in available public-domain data sets. Nevertheless, obtaining high-quality, well-labeled, and well-balanced real-world labeled images is challenging since manual data curation requires significant human effort and is time-consuming. In contrast, the web has a potentially unlimited source of food data but tapping into this resource has a good chance of corrupted and wrongly labeled images. In addition, the uneven distribution among food categories may lead to data imbalance problems. All these issues make it challenging to create clean data sets for food from web data. To address this issue, we present AutoCleanDeepFood, a novel end-to-end food computing framework for regional gastronomy that contains the following components: (i) a fully automated pre-processing pipeline for custom data sets creation related to specific regional gastronomy, (ii) a transfer learning-based training paradigm to filter out noisy labels through loss ranking, incorporating a Russian Roulette probabilistic approach to mitigate data imbalance problems, and (iii) a method for deploying the resulting model on smartphones for real-time inferences. We assess the performance of our framework on a real-world noisy public domain data set, ETH Food-101, and two novel web-collected datasets, MENA-150 and Pizza-Styles. We demonstrate the filtering capabilities of our proposed method through embedding visualization of the feature space using the t-SNE dimension reduction scheme. Our filtering scheme is efficient and effectively improves accuracy in all cases, boosting performance by 0.96, 0.71, and 1.29% on MENA-150, ETH Food-101, and Pizza-Styles, respectively.
Nauman Ullah Gilal, Marwa Qaraqe, Jens Schneider 0002, Marco Agus
Vis. Comput.4
2024 Deep synthesis and exploration of omnidirectional stereoscopic environments from a single surround-view panoramic image
Giovanni Pintore, Alberto Jaspe-Villanueva, Markus Hadwiger, Jens Schneider 0002, Marco Agus, Fabio Marton, Fabio Bettio, Enrico Gobbetti
Comput. Graph.5
2024 Evaluating machine learning technologies for food computing from a data set perspective
abstract
Abstract Food plays an important role in our lives that goes beyond mere sustenance. Food affects behavior, mood, and social life. It has recently become an important focus of multimedia and social media applications. The rapid increase of available image data and the fast evolution of artificial intelligence, paired with a raised awareness of people’s nutritional habits, have recently led to an emerging field attracting significant attention, called food computing, aimed at performing automatic food analysis. Food computing benefits from technologies based on modern machine learning techniques, including deep learning, deep convolutional neural networks, and transfer learning. These technologies are broadly used to address emerging problems and challenges in food-related topics, such as food recognition, classification, detection, estimation of calories and food quality, dietary assessment, food recommendation, etc. However, the specific characteristics of food image data, like visual heterogeneity, make the food classification task particularly challenging. To give an overview of the state of the art in the field, we surveyed the most recent machine learning and deep learning technologies used for food classification with a particular focus on data aspects. We collected and reviewed more than 100 papers related to the usage of machine learning and deep learning for food computing tasks. We analyze their performance on publicly available state-of-art food data sets and their potential for usage in multimedia food-related applications for various needs (communication, leisure, tourism, blogging, reverse engineering, etc.). In this paper, we perform an extensive review and categorization of available data sets: to this end, we developed and released an open web resource in which the most recent existing food data sets are collected and mapped to the corresponding geographical regions. Although artificial intelligence methods can be considered mature enough to be used in basic food classification tasks, our analysis of the state-of-the-art reveals that challenges related to the application of this technology need to be addressed. These challenges include, among others: poor representation of regional gastronomy, incorporation of adaptive learning schemes, and reverse engineering for automatic food creation and replication.
Nauman Ullah Gilal, Khaled Al-Thelaya, Jumana Khalid Al-Saeed, Mohamed M. Abdallah 0001, Jens Schneider 0002, James She, Jawad Hussain Awan, Marco Agus
Multim. Tools Appl.8
2023 Noise2Seg: Automatic Few-Shot Selection from Noisy Web Data for Underwater Tropical Fishes Segmentation
abstract
We present a novel automatic few-shot selection (FSS) approach with a noisy web-collected dataset for underwater tropical fishes. Underwater image segmentation is of utmost importance in marine biology and underwater robotics. However, obtaining high-quality annotated underwater images is a challenging task due to the scarcity of underwater imaging systems and the high cost of data acquisition. Alternatively, web data is a readily available source of images, but it frequently contains web-corrupted noisy labels, making it challenging to curate a clean dataset for underwater tropical fish segmentation. Generally, the manual selection of high-quality images from web data requires human supervision and expert proofreading, which is both expensive and time-consuming. To address this issue, we propose Noise2Seg, an automatic processing framework composed by: (i) an automated web scrapping tool, (ii) a robust and automatic FSS process using a loss ranking scheme; (iii) a manual annotation component using “Roboflow”, (iv) the latest You Only Look Once version 8 (YOLOv8) for underwater fish detection and segmentation. Additionally, we curated and annotated a novel dataset for the segmentation of tropical fishes from Qatar marine ecosystem (Qatar Tropical Fishes-10 QTF-10). We compared the performance of manual and automatic FSS using mean average precision (mAP) score; manual FSS achieved all mAP (91.5%, 93.4%), and minimum mAP (19.5%, 54.8%), while automatic FSS achieved all mAP (92%, 99.5%), and minimum mAP (24.9%, 99.5%) with 5 and 10 images per class, respectively. The code and dataset related to this paper can be found on GitHub11https://github.com/GilalNauman/Automatic-Few-shot-Selction-ISNCC/tree/main, accessed on 1st of March 2023.
Nauman Ullah Gilal, Fahad Majeed, Khaled Al-Thelaya, Mehak Khan, Jens Schneider 0002, Marco Agus
ISNCC6
2023 Ultraliser: a framework for creating multiscale, high-fidelity and geometrically realistic 3D models for in silico neuroscience
abstract
Ultraliser is a neuroscience-specific software framework capable of creating accurate and biologically realistic 3D models of complex neuroscientific structures at intracellular (e.g. mitochondria and endoplasmic reticula), cellular (e.g. neurons and glia) and even multicellular scales of resolution (e.g. cerebral vasculature and minicolumns). Resulting models are exported as triangulated surface meshes and annotated volumes for multiple applications in in silico neuroscience, allowing scalable supercomputer simulations that can unravel intricate cellular structure-function relationships. Ultraliser implements a high-performance and unconditionally robust voxelization engine adapted to create optimized watertight surface meshes and annotated voxel grids from arbitrary non-watertight triangular soups, digitized morphological skeletons or binary volumetric masks. The framework represents a major leap forward in simulation-based neuroscience, making it possible to employ high-resolution 3D structural models for quantification of surface areas and volumes, which are of the utmost importance for cellular and system simulations. The power of Ultraliser is demonstrated with several use cases in which hundreds of models are created for potential application in diverse types of simulations. Ultraliser is publicly released under the GNU GPL3 license on GitHub (BlueBrain/Ultraliser). SIGNIFICANCE: There is crystal clear evidence on the impact of cell shape on its signaling mechanisms. Structural models can therefore be insightful to realize the function; the more realistic the structure can be, the further we get insights into the function. Creating realistic structural models from existing ones is challenging, particularly when needed for detailed subcellular simulations. We present Ultraliser, a neuroscience-dedicated framework capable of building these structural models with realistic and detailed cellular geometries that can be used for simulations.
Marwan Abdellah, Juan Jose Garcia-Cantero, Nadir Román Guerrero, Alessandro Foni, Jay S. Coggan, Corrado Calì, Marco Agus, Eleftherios Zisis, Daniel X. Keller, Markus Hadwiger, Pierre J. Magistretti, Henry Markram, Felix Schürmann
Briefings Bioinform.7
2023 SPIDER: A framework for processing, editing and presenting immersive high-resolution spherical indoor scenes
abstract
Today’s Extended Reality (XR) applications that call for specific Diminished Reality (DR) strategies to hide specific classes of objects are increasingly using 360° cameras, which can capture entire areas in a single picture. In this work, we present an interactive-based image processing, editing and rendering system named SPIDER, that takes a spherical 360° indoor scene as input. The system is composed of a novel integrated deep learning architecture for extracting geometric and semantic information of full and empty rooms, based on gated and dilated convolutions, followed by a super-resolution module for improving the resolution of the color and depth signals. The obtained high resolution representations allow users to perform interactive exploration and basic editing operations on the reconstructed indoor scene, namely: (i) rendering of the scene in various modalities (point cloud, polygonal, wireframe) (ii) refurnishing (transferring portions of rooms) (iii) deferred shading through the usage of precomputed normal maps. These kinds of scene editing and manipulations can be used for assessing the inference from deep learning models and enable several Mixed Reality applications in areas such as furniture retails, interior designs, and real estates. Moreover, it can also be useful in data augmentation, arts, designs, and paintings. We report on the performance improvement of the various processing components on public domain spherical image indoor datasets.
Muhammad Tukur, Giovanni Pintore, Enrico Gobbetti, Jens Schneider 0002, Marco Agus
Graph. Model.5
2023 Deep Scene Synthesis of Atlanta-World Interiors from a Single Omnidirectional Image
abstract
We present a new data-driven approach for extracting geometric and structural information from a single spherical panorama of an interior scene, and for using this information to render the scene from novel points of view, enhancing 3D immersion in VR applications. The approach copes with the inherent ambiguities of single-image geometry estimation and novel view synthesis by focusing on the very common case of Atlanta-world interiors, bounded by horizontal floors and ceilings and vertical walls. Based on this prior, we introduce a novel end-to-end deep learning approach to jointly estimate the depth and the underlying room structure of the scene. The prior guides the design of the network and of novel domain-specific loss functions, shifting the major computational load on a training phase that exploits available large-scale synthetic panoramic imagery. An extremely lightweight network uses geometric and structural information to infer novel panoramic views from translated positions at interactive rates, from which perspective views matching head rotations are produced and upsampled to the display size. As a result, our method automatically produces new poses around the original camera at interactive rates, within a working area suitable for producing depth cues for VR applications, especially when using head-mounted displays connected to graphics servers. The extracted floor plan and 3D wall structure can also be used to support room exploration. The experimental results demonstrate that our method provides low-latency performance and improves over current state-of-the-art solutions in prediction accuracy on available commonly used indoor panoramic benchmarks.
Giovanni Pintore, Fabio Bettio, Marco Agus, Enrico Gobbetti
IEEE Trans. Vis. Comput. Graph.3
2022 Probabilistic Occlusion Culling using Confidence Maps for High-Quality Rendering of Large Particle Data
abstract
Achieving high rendering quality in the visualization of large particle data, for example from large-scale molecular dynamics simulations, requires a significant amount of sub-pixel super-sampling, due to very high numbers of particles per pixel. Although it is impossible to super-sample all particles of large-scale data at interactive rates, efficient occlusion culling can decouple the overall data size from a high effective sampling rate of visible particles. However, while the latter is essential for domain scientists to be able to see important data features, performing occlusion culling by sampling or sorting the data is usually slow or error-prone due to visibility estimates of insufficient quality. We present a novel probabilistic culling architecture for super-sampled high-quality rendering of large particle data. Occlusion is dynamically determined at the sub-pixel level, without explicit visibility sorting or data simplification. We introduce confidence maps to probabilistically estimate confidence in the visibility data gathered so far. This enables progressive, confidence-based culling, helping to avoid wrong visibility decisions. In this way, we determine particle visibility with high accuracy, although only a small part of the data set is sampled. This enables extensive super-sampling of (partially) visible particles for high rendering quality, at a fraction of the cost of sampling all particles. For real-time performance with millions of particles, we exploit novel features of recent GPU architectures to group particles into two hierarchy levels, combining fine-grained culling with high frame rates.
Mohamed Ibrahim 0006, Peter Rautek, Guido Reina, Marco Agus, Markus Hadwiger
IEEE Trans. Vis. Comput. Graph.4
2022 Instant Automatic Emptying of Panoramic Indoor Scenes
abstract
Nowadays 360° cameras, capable to capture full environments in a single shot, are increasingly being used in a variety of Extended Reality (XR) applications that require specific Diminished Reality (DR) techniques to conceal selected classes of objects. In this work, we present a new data-driven approach that, from an input 360° image of a furnished indoor space automatically returns, with very low latency, an omnidirectional photorealistic view and architecturally plausible depth of the same scene emptied of all clutter. Contrary to recent data-driven inpainting methods that remove single user-defined objects based on their semantics, our approach is holistically applied to the entire scene, and is capable to separate the clutter from the architectural structure in a single step. By exploiting peculiar geometric features of the indoor environment, we shift the major computational load on the training phase and having an extremely lightweight network at prediction time. Our end-to-end approach starts by calculating an attention mask of the clutter in the image based on the geometric difference between full and empty scene. This mask is then propagated through gated convolutions that drive the generation of the output image and its depth. Returning the depth of the resulting structure allows us to exploit, during supervised training, geometric losses of different orders, including robust pixel-wise geometric losses and high-order 3D constraints typical of indoor structures. The experimental results demonstrate that our method provides interactive performance and outperforms current state-of-the-art solutions in prediction accuracy on available commonly used indoor panoramic benchmarks. In addition, our method presents consistent quality results even for scenes captured in the wild and for data for which there is no ground truth to support supervised training.
Giovanni Pintore, Marco Agus, Eva Almansa, Enrico Gobbetti
IEEE Trans. Vis. Comput. Graph.2
2021 SliceNet: Deep Dense Depth Estimation From a Single Indoor Panorama Using a Slice-Based Representation
abstract
We introduce a novel deep neural network to estimate a depth map from a single monocular indoor panorama. The network directly works on the equirectangular projection, exploiting the properties of indoor 360° images. Starting from the fact that gravity plays an important role in the design and construction of man-made indoor scenes, we propose a compact representation of the scene into vertical slices of the sphere, and we exploit long- and short-term relationships among slices to recover the equirectangular depth map. Our design makes it possible to maintain high-resolution information in the extracted features even with a deep network. The experimental results demonstrate that our method outperforms current state-of-the-art solutions in prediction accuracy, particularly for real-world data.
Giovanni Pintore, Marco Agus, Eva Almansa, Jens Schneider 0002, Enrico Gobbetti
CVPR2
2021 InShaDe: Invariant Shape Descriptors for visual 2D and 3D cellular and nuclear shape analysis and classification
abstract
We present a shape processing framework for visual exploration of cellular nuclear envelopes extracted from microscopic images arising in histology and neuroscience. The framework is based on a novel shape descriptor of closed contours in 2D and 3D. In 2D, it relies on a geodesically uniform resampling of discrete curves to compute unsigned curvatures at vertices and edges based on discrete differential geometry. Our descriptor is, by design, invariant under translation, rotation, and parameterization. We achieve the latter invariance under parameterization shifts by using elliptic Fourier analysis on the resulting curvature vectors. Uniform scale-invariance is optional and is a result of scaling curvature features to z-scores. We further augment the proposed descriptor with feature coefficients obtained through sparse coding of the extracted cellular structures using K-sparse autoencoders. For the analysis of 3D shapes, we compute mean curvatures based on the Laplace-Beltrami operator on triangular meshes, followed by computing a spherical parameterization through mean curvature flow. Finally, we compute the Spherical Harmonics decomposition to obtain invariant energy coefficients. Our invariant descriptors provide an embedding into a fixed-dimensional feature space that can be used for various applications, e.g., as input features for deep and shallow learning techniques or as input for dimension reduction schemes to provide a visual reference for clustering shape collections. We demonstrate the capabilities of our framework in the context of visual analysis and unsupervised classification of 2D histology images and 3D nuclear envelopes extracted from serial section electron microscopy stacks.
Khaled Al-Thelaya, Marco Agus, Nauman Ullah Gilal, Yin Yang 0001, Giovanni Pintore, Enrico Gobbetti, Corrado Calì, Pierre J. Magistretti, William Mifsud, Jens Schneider 0002
Comput. Graph.2
2021 Deep3DLayout: 3D reconstruction of an indoor layout from a spherical panoramic image
abstract
Recovering the 3D shape of the bounding permanent surfaces of a room from a single image is a key component of indoor reconstruction pipelines. In this article, we introduce a novel deep learning technique capable to produce, at interactive rates, a tessellated bounding 3D surface from a single 360° image. Differently from prior solutions, we fully address the problem in 3D, significantly expanding the reconstruction space of prior solutions. A graph convolutional network directly infers the room structure as a 3D mesh by progressively deforming a graph-encoded tessellated sphere mapped to the spherical panorama, leveraging perceptual features extracted from the input image. Important 3D properties of indoor environments are exploited in our design. In particular, gravity-aligned features are actively incorporated in the graph in a projection layer that exploits the recent concept of multi head self-attention, and specialized losses guide towards plausible solutions even in presence of massive clutter and occlusions. Extensive experiments demonstrate that our approach outperforms current state of the art methods in terms of accuracy and capability to reconstruct more complex environments.
Giovanni Pintore, Eva Almansa, Marco Agus, Enrico Gobbetti
ACM Trans. Graph.3
2021 The Mixture Graph-A Data Structure for Compressing, Rendering, and Querying Segmentation Histograms
abstract
In this paper, we present a novel data structure, called the Mixture Graph. This data structure allows us to compress, render, and query segmentation histograms. Such histograms arise when building a mipmap of a volume containing segmentation IDs. Each voxel in the histogram mipmap contains a convex combination (mixture) of segmentation IDs. Each mixture represents the distribution of IDs in the respective voxel's children. Our method factorizes these mixtures into a series of linear interpolations between exactly two segmentation IDs. The result is represented as a directed acyclic graph (DAG) whose nodes are topologically ordered. Pruning replicate nodes in the tree followed by compression allows us to store the resulting data structure efficiently. During rendering, transfer functions are propagated from sources (leafs) through the DAG to allow for efficient, pre-filtered rendering at interactive frame rates. Assembly of histogram contributions across the footprint of a given volume allows us to efficiently query partial histograms, achieving up to 178 x speed-up over naive parallelized range queries. Additionally, we apply the Mixture Graph to compute correctly pre-filtered volume lighting and to interactively explore segments based on shape, geometry, and orientation using multi-dimensional transfer functions.
Khaled Al-Thelaya, Marco Agus, Jens Schneider 0002
IEEE Trans. Vis. Comput. Graph.2
2020 AtlantaNet: Inferring the 3D Indoor Layout from a Single $360^\circ $ Image Beyond the Manhattan World Assumption
Giovanni Pintore, Marco Agus, Enrico Gobbetti
ECCV (8)2
2020 Foreword to the special section on smart tool and applications for graphics (STAG 2019)
Marco Agus, Massimiliano Corsini, Ruggero Pintus
Comput. Graph.1
2020 Virtual reality framework for editing and exploring medial axis representations of nanometric scale neural structures
abstract
We present a novel virtual reality (VR) based framework for the exploratory analysis of nanoscale 3D reconstructions of cellular structures acquired from rodent brain samples through serial electron microscopy. The system is specifically targeted on medial axis representations (skeletons) of branched and tubular structures of cellular shapes, and it is designed for providing to domain scientists: i) effective and fast semi-automatic interfaces for tracing skeletons directly on surface-based representations of cells and structures, ii) fast tools for proofreading, i.e., correcting and editing of semi-automatically constructed skeleton representations, and iii) natural methods for interactive exploration, i.e., measuring, comparing, and analyzing geometric features related to cellular structures based on medial axis representations. Neuroscientists currently use the system for performing morphology studies on sparse reconstructions of glial cells and neurons extracted from a sample of the somatosensory cortex of a juvenile rat. The framework runs in a standard PC and has been tested on two different display and interaction setups: PC-tethered stereoscopic head-mounted display (HMD) with 3D controllers and tracking sensors, and a large display wall with a standard gamepad controller. We report on a user study that we carried out for analyzing user performance on different tasks using these two setups.
Daniya Boges, Marco Agus, Ronell Sicat, Pierre J. Magistretti, Markus Hadwiger, Corrado Calì
Comput. Graph.2
2019 Virtual environment for processing medial axis representations of 3D nanoscale reconstructions of brain cellular structures
abstract
We present a novel immersive environment for the interactive analysis of nanoscale cellular reconstructions of rodent brain samples acquired through electron microscopy. The system is focused on medial axis representations (skeletons) of branched and tubular structures of brain cells, and it is specifically designed for: i) effective semi-automatic creation of skeletons from surface-based representations of cells and structures ii) fast proofreading, i.e., correcting and editing of semi-automatically constructed skeleton representations, and iii) useful exploration, i.e., measuring, comparing, and analyzing geometric features related to cellular structures based on medial axis representations. The application runs in a standard PC-tethered virtual reality (VR) setup with a head mounted display (HMD), controllers, and tracking sensors. The system is currently used by neuroscientists for performing morphology studies on sparse reconstructions of glial cells and neurons extracted from a sample of the somatosensory cortex of a juvenile rat.
Daniya Boges, Corrado Calì, Pierre J. Magistretti, Markus Hadwiger, Ronell Sicat, Marco Agus
VRST6
2019 Interactive Volumetric Visual Analysis of Glycogen-derived Energy Absorption in Nanometric Brain Structures
abstract
Abstract Digital acquisition and processing techniques are changing the way neuroscience investigation is carried out. Emerging applications range from statistical analysis on image stacks to complex connectomics visual analysis tools targeted to develop and test hypotheses of brain development and activity. In this work, we focus on neuroenergetics, a field where neuroscientists analyze nanoscale brain morphology and relate energy consumption to glucose storage in form of glycogen granules. In order to facilitate the understanding of neuroenergetic mechanisms, we propose a novel customized pipeline for the visual analysis of nanometric‐level reconstructions based on electron microscopy image data. Our framework supports analysis tasks by combining i) a scalable volume visualization architecture able to selectively render image stacks and corresponding labelled data, ii) a method for highlighting distance‐based energy absorption probabilities in form of glow maps, and iii) a hybrid connectivitybased and absorption‐based interactive layout representation able to support queries for selective analysis of areas of interest and potential activity within the segmented datasets. This working pipeline is currently used in a variety of studies in the neuroenergetics domain. Here, we discuss a test case in which the framework was successfully used by domain scientists for the analysis of aging effects on glycogen metabolism, extracting knowledge from a series of nanoscale brain stacks of rodents somatosensory cortex.
Marco Agus, Corrado Calì, Ali K. Al-Awami, Enrico Gobbetti, Pierre J. Magistretti, Markus Hadwiger
Comput. Graph. Forum1
2019 A framework for GPU-accelerated exploration of massive time-varying rectilinear scalar volumes
abstract
Abstract We introduce a novel flexible approach to spatiotemporal exploration of rectilinear scalar volumes. Our out‐of‐core representation, based on per‐frame levels of hierarchically tiled non‐redundant 3D grids, efficiently supports spatiotemporal random access and streaming to the GPU in compressed formats. A novel low‐bitrate codec able to store into fixed‐size pages a variable‐rate approximation based on sparse coding with learned dictionaries is exploited to meet stringent bandwidth constraint during time‐critical operations, while a near‐lossless representation is employed to support high‐quality static frame rendering. A flexible high‐speed GPU decoder and raycasting framework mixes and matches GPU kernels performing parallel object‐space and image‐space operations for seamless support, on fat and thin clients, of different exploration use cases, including animation and temporal browsing, dynamic exploration of single frames, and high‐quality snapshots generated from near‐lossless data. The quality and performance of our approach are demonstrated on large data sets with thousands of multi‐billion‐voxel frames.
Fabio Marton, Marco Agus, Enrico Gobbetti
Comput. Graph. Forum2
2019 Culling for Extreme-Scale Segmentation Volumes: A Hybrid Deterministic and Probabilistic Approach
abstract
With the rapid increase in raw volume data sizes, such as terabyte-sized microscopy volumes, the corresponding segmentation label volumes have become extremely large as well. We focus on integer label data, whose efficient representation in memory, as well as fast random data access, pose an even greater challenge than the raw image data. Often, it is crucial to be able to rapidly identify which segments are located where, whether for empty space skipping for fast rendering, or for spatial proximity queries. We refer to this process as culling. In order to enable efficient culling of millions of labeled segments, we present a novel hybrid approach that combines deterministic and probabilistic representations of label data in a data-adaptive hierarchical data structure that we call the label list tree. In each node, we adaptively encode label data using either a probabilistic constant-time access representation for fast conservative culling, or a deterministic logarithmic-time access representation for exact queries. We choose the best data structures for representing the labels of each spatial region while building the label list tree. At run time, we further employ a novel query-adaptive culling strategy. While filtering a query down the tree, we prune it successively, and in each node adaptively select the representation that is best suited for evaluating the pruned query, depending on its size. We show an analysis of the efficiency of our approach with several large data sets from connectomics, including a brain scan with more than 13 million labeled segments, and compare our method to conventional culling approaches. Our approach achieves significant reductions in storage size as well as faster query times.
Johanna Beyer, Haneen Mohammed, Marco Agus, Ali K. Al-Awami, Hanspeter Pfister, Markus Hadwiger
IEEE Trans. Vis. Comput. Graph.3
2018 GLAM: Glycogen-derived Lactate Absorption Map for visual analysis of dense and sparse surface reconstructions of rodent brain structures on desktop systems and virtual environments
Marco Agus, Daniya Boges, Nicolas Gagnon, Pierre J. Magistretti, Markus Hadwiger, Corrado Calì
Comput. Graph.1
2018 SparseLeap: Efficient Empty Space Skipping for Large-Scale Volume Rendering
abstract
Recent advances in data acquisition produce volume data of very high resolution and large size, such as terabyte-sized microscopy volumes. These data often contain many fine and intricate structures, which pose huge challenges for volume rendering, and make it particularly important to efficiently skip empty space. This paper addresses two major challenges: (1) The complexity of large volumes containing fine structures often leads to highly fragmented space subdivisions that make empty regions hard to skip efficiently. (2) The classification of space into empty and non-empty regions changes frequently, because the user or the evaluation of an interactive query activate a different set of objects, which makes it unfeasible to pre-compute a well-adapted space subdivision. We describe the novel SparseLeap method for efficient empty space skipping in very large volumes, even around fine structures. The main performance characteristic of SparseLeap is that it moves the major cost of empty space skipping out of the ray-casting stage. We achieve this via a hybrid strategy that balances the computational load between determining empty ray segments in a rasterization (object-order) stage, and sampling non-empty volume data in the ray-casting (image-order) stage. Before ray-casting, we exploit the fast hardware rasterization of GPUs to create a ray segment list for each pixel, which identifies non-empty regions along the ray. The ray-casting stage then leaps over empty space without hierarchy traversal. Ray segment lists are created by rasterizing a set of fine-grained, view-independent bounding boxes. Frame coherence is exploited by re-using the same bounding boxes unless the set of active objects changes. We show that SparseLeap scales better to large, sparse data than standard octree empty space skipping.
Markus Hadwiger, Ali K. Al-Awami, Johanna Beyer, Marco Agus, Hanspeter Pfister
IEEE Trans. Vis. Comput. Graph.4
2017 Mobile graphics
abstract
course Share on Mobile graphics Authors: Marco Agus KAUST & CRS4 KAUST & CRS4View Profile , Enrico Gobbetti CRS4 CRS4View Profile , Fabio Marton CRS4 CRS4View Profile , Giovanni Pintore CRS4 CRS4View Profile , Pere-Pau Vázquez UPC UPCView Profile Authors Info & Claims SA '17: SIGGRAPH Asia 2017 CoursesNovember 2017 Article No.: 12Pages 1–259https://doi.org/10.1145/3134472.3134483Published:27 November 2017Publication History 1citation239DownloadsMetricsTotal Citations1Total Downloads239Last 12 Months19Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Marco Agus, Enrico Gobbetti, Fabio Marton, Giovanni Pintore, Pere-Pau Vázquez
SIGGRAPH ASIA (Courses)1
2016 Omnidirectional image capture on mobile devices for fast automatic generation of 2.5D indoor maps
abstract
We introduce a light-weight automatic method to quickly capture and recover 2.5D multi-room indoor environments scaled to real-world metric dimensions. To minimize the user effort required, we capture and analyze a single omni-directional image per room using widely available mobile devices. Through a simple tracking of the user movements between rooms, we iterate the process to map and reconstruct entire floor plans. In order to infer 3D clues with a minimal processing and without relying on the presence of texture or detail, we define a specialized spatial transform based on catadioptric theory to highlight the room's structure in a virtual projection. From this information, we define a parametric model of each room to formalize our problem as a global optimization solved by Levenberg-Marquardt iterations. The effectiveness of the method is demonstrated on several challenging real-world multi-room indoor scenes.
Giovanni Pintore, Valeria Garro, Fabio Ganovelli, Enrico Gobbetti, Marco Agus
WACV5
2015 Adaptive Recommendations for Enhanced Non-linear Exploration of Annotated 3D Objects
abstract
Abstract We introduce a novel approach for letting casual viewers explore detailed 3D models integrated with structured spatially associated descriptive information organized in a graph. Each node associates a subset of the 3D surface seen from a particular viewpoint to the related descriptive annotation, together with its author‐defined importance. Graph edges describe, instead, the strength of the dependency relation between information nodes, allowing content authors to describe the preferred order of presentation of information. At run‐time, users navigate inside the 3D scene using a camera controller, while adaptively receiving unobtrusive guidance towards interesting viewpoints and history‐ and location‐dependent suggestions on important information, which is adaptively presented using 2D overlays displayed over the 3D scene. The capabilities of our approach are demonstrated in a real‐world cultural heritage application involving the public presentation of sculptural complex on a large projection‐based display. A user study has been performed in order to validate our approach.
Marcos Balsa, Marco Agus, Fabio Marton, Enrico Gobbetti
Comput. Graph. Forum2
2012 Natural exploration of 3D massive models on large-scale light field displays using the FOX proximal navigation technique
Fabio Marton, Marco Agus, Enrico Gobbetti, Giovanni Pintore, Marcos Balsa
Comput. Graph.2
2011 Interactive Multiscale Tensor Reconstruction for Multiresolution Volume Visualization
abstract
Large scale and structurally complex volume datasets from high-resolution 3D imaging devices or computational simulations pose a number of technical challenges for interactive visual analysis. In this paper, we present the first integration of a multiscale volume representation based on tensor approximation within a GPU-accelerated out-of-core multiresolution rendering framework. Specific contributions include (a) a hierarchical brick-tensor decomposition approach for pre-processing large volume data, (b) a GPU accelerated tensor reconstruction implementation exploiting CUDA capabilities, and (c) an effective tensor-specific quantization strategy for reducing data transfer bandwidth and out-of-core memory footprint. Our multiscale representation allows for the extraction, analysis and display of structural features at variable spatial scales, while adaptive level-of-detail rendering methods make it possible to interactively explore large datasets within a constrained memory footprint. The quality and performance of our prototype system is evaluated on large structurally complex datasets, including gigabyte-sized micro-tomographic volumes.
Susanne K. Suter, José Antonio Iglesias Guitián, Fabio Marton, Marco Agus, Andreas Elsener, Christoph P. E. Zollikofer, Meenakshisundaram Gopi, Enrico Gobbetti, Renato Pajarola
IEEE Trans. Vis. Comput. Graph.4
2009 An interactive 3D medical visualization system based on a light field display
Marco Agus, Fabio Bettio, Andrea Giachetti 0001, Enrico Gobbetti, José Antonio Iglesias Guitián, Fabio Marton, Giovanni Pintore
Vis. Comput.1
2008 GPU Accelerated Direct Volume Rendering on an Interactive Light Field Display
abstract
Abstract We present a GPU accelerated volume ray casting system interactively driving a multi‐user light field display. The display, driven by a single programmable GPU, is based on a specially arranged array of projectors and a holographic screen and provides full horizontal parallax. The characteristics of the display are exploited to develop a specialized volume rendering technique able to provide multiple freely moving naked‐eye viewers the illusion of seeing and manipulating virtual volumetric objects floating in the display workspace. In our approach, a GPU ray‐caster follows rays generated by a multiple‐center‐of‐projection technique while sampling pre‐filtered versions of the dataset at resolutions that match the varying spatial accuracy of the display. The method achieves interactive performance and provides rapid visual understanding of complex volumetric data sets even when using depth oblivious compositing techniques.
Marco Agus, Enrico Gobbetti, José Antonio Iglesias Guitián, Fabio Marton, Giovanni Pintore
Comput. Graph. Forum1
2003 Adaptive techniques for real-time haptic and visual simulation of bone dissection
abstract
Bone dissection is an important component of many surgical procedures. In this paper we discuss adaptive techniques for providing real-time haptic and visual feedback during a virtual bone dissection simulation. The simulator is being developed as a component of a training system for temporal bone surgery. We harness the difference in complexity and frequency requirements of the visual and haptic simulations by modeling the system as a collection of loosely coupled concurrent components. The haptic component exploits a multi-resolution representation of the first two moments of the bone characteristic function to rapidly compute contact forces and determine bone erosion. The visual component uses a time-critical particle system evolution method to simulate secondary visual effects, such as bone debris accumulation, blooding, irrigation, and suction.
Marco Agus, Andrea Giachetti 0001, Enrico Gobbetti, Gianluigi Zanetti, Antonio Zorcolo
VR1
2003 Hierarchical Higher Order sFace Cluster Radiosity for Global Illumination Walkthroughs of Complex Non-Diffuse Environments
abstract
Abstract We present an algorithm for simulating global illumination in scenes composed of highly tessellated objects withdiffuse or moderately glossy reflectance. The solution method is a higher order extension of the face cluster radiositytechnique. It combines face clustering, multiresolution visibility, vector radiosity, and higher order baseswith a modified progressive shooting iteration to rapidly produce visually continuous solutions with limited memoryrequirements. The output of the method is a vector irradiance map that partitions input models into areaswhere global illumination is well approximated using the selected basis. The programming capabilities of moderncommodity graphics architectures are exploited to render illuminated models directly from the vector irradiancemap, exploiting hardware acceleration for approximating view dependent illumination during interactive walkthroughs.Using this algorithm, visually compelling global illumination solutions for scenes of over one millioninput polygons can be computed in minutes and examined interactively on common graphics personal computers. Categories and Subject Descriptors (according to ACM CCS): I.3.3 [Computer Graphics]: Picture and Image Generation; I.3.7 [Computer Graphics]: Three‐Dimensional Graphics and Realism.
Enrico Gobbetti, Leonardo Spanò, Marco Agus
Comput. Graph. Forum3
2002 Real-Time Haptic and Visual Simulation of Bone Dissection
abstract
Bone dissection is an important component of many surgical procedures. In this paper, we discuss a haptic and visual implementation of a bone-cutting burr that is being developed as a component of a training system for temporal bone surgery. We use a physically motivated model to describe the burr-bone interaction, which includes haptic force evaluation, the bone erosion process and the resulting debris. The current implementation, directly operating on a voxel discretization of patient-specific 3D CT and MRI data, is efficient enough to provide real-time feedback on a low-end multiprocessing PC platform.
Marco Agus, Andrea Giachetti 0001, Enrico Gobbetti, Gianluigi Zanetti, Antonio Zorcolo
VR1