Marco Cipriano

dblp:267/9230 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
3since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Representation and self-supervised learning · 28% Segmentation and scene understanding · 25% Vision and language · 20%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%

Topics — the 10 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Segmentation and scene understanding
scene understanding
1.012026
MINGLE: VLMs for Semantically Complex Region Detection in Urban Scenes · AAAI 2026
Computer vision › Vision and language › multimodal reasoning
vision-language model reasoning
1.012026
MINGLE: VLMs for Semantically Complex Region Detection in Urban Scenes · AAAI 2026
Machine learning › Generative modeling
autoregressive model
0.912025
Vector Grimoire: Codebook-based Shape Generation under Raster Image Supervision · ICML 2025
Machine learning › Representation and self-supervised learning › vector quantization
codebook learning
0.912025
Vector Grimoire: Codebook-based Shape Generation under Raster Image Supervision · ICML 2025
Machine learning › Representation and self-supervised learning
vector quantization
0.912025
Vector Grimoire: Codebook-based Shape Generation under Raster Image Supervision · ICML 2025
Visual content generation and editing
vector graphics generation
0.912025
Vector Grimoire: Codebook-based Shape Generation under Raster Image Supervision · ICML 2025
Machine learning › Learning paradigms › semi-supervised learning › graph-based semi-supervised learning
label propagation
0.612022
Improving Segmentation of the Inferior Alveolar Nerve through Deep Label Propagation · CVPR 2022
Computer vision › Segmentation and scene understanding
medical image segmentation
0.612022
Improving Segmentation of the Inferior Alveolar Nerve through Deep Label Propagation · CVPR 2022
Computer vision › Image recognition and object detection › object detection › category-specific object detection
person detection
0.312026
MINGLE: VLMs for Semantically Complex Region Detection in Urban Scenes · AAAI 2026
Computer vision › Vision and language › vision-language generation
text-guided generation
0.312025
Vector Grimoire: Codebook-based Shape Generation under Raster Image Supervision · ICML 2025

Methods — techniques the papers use, named apart from their topics

vector quantization · 1.7autoregressive transformer · 1.7deep label propagation · 1.1convolutional neural network · 1.1vision-language model · 1.0spatial aggregation · 1.0depth estimation · 1.0
YearPublicationVenuePosition
2026 MINGLE: VLMs for Semantically Complex Region Detection in Urban Scenes
abstract
Understanding group-level social interactions in public spaces is crucial for urban planning, informing the design of socially vibrant and inclusive environments. Detecting such interactions from images involves interpreting subtle visual cues such as relations, proximity and co-movement – semantically complex signals that go beyond traditional object detection. To address this challenge, we introduce a social group region detection task, which requires inferring and spatially grounding visual regions defined by abstract interpersonal relations. We propose MINGLE (Modeling INterpersonal Group-Level Engagement), a modular three-stage pipeline that integrates: (1) off-the-shelf human detection and depth estimation, (2) VLM-based reasoning to classify pairwise social affiliation, and (3) a lightweight spatial aggregation algorithm to localize socially connected groups. To support this task and encourage future research, we present a new dataset of 100K urban street-view images annotated with bounding boxes and labels for both individuals and socially interacting groups. The annotations combine human-created labels and outputs from the MINGLE pipeline, ensuring semantic richness and broad coverage of real world scenarios.
Liu Liu 0018, Alexandra Schild, Marco Cipriano, Fatimeh Al Ghannam, Freya Tan, Gerard de Melo, Andres Sevtsuk
AAAI3
2025 Vector Grimoire: Codebook-based Shape Generation under Raster Image Supervision
abstract
Scalable Vector Graphics (SVG) is a popular format on the web and in the design industry. However, despite the great strides made in generative modeling, SVG has remained underexplored due to the discrete and complex nature of such data. We introduce GRIMOIRE, a text-guided SVG generative model that is comprised of two modules: A Visual Shape Quantizer (VSQ) learns to map raster images onto a discrete codebook by reconstructing them as vector shapes, and an Auto-Regressive Transformer (ART) models the joint probability distribution over shape tokens, positions and textual descriptions, allowing us to generate vector graphics from natural language. Unlike existing models that require direct supervision from SVG data, GRIMOIRE learns shape image patches using only raster image supervision which opens up vector generative modeling to significantly more data. We demonstrate the effectiveness of our method by fitting GRIMOIRE for closed filled shapes on the MNIST and Emoji, and for outline strokes on icon and font data, surpassing previous image-supervised methods in generative quality and vector-supervised approach in flexibility.
Marco Cipriano, Moritz Feuerpfeil, Gerard de Melo
ICML1
2022 Improving Segmentation of the Inferior Alveolar Nerve through Deep Label Propagation
abstract
Many recent works in dentistry and maxillofacial imagery focused on the Inferior Alveolar Nerve (IAN) canal detection. Unfortunately, the small extent of available 3D maxillofacial datasets has strongly limited the performance of deep learning-based techniques. On the other hand, a huge amount of sparsely annotated data is produced every day from the regular procedures in the maxillofacial practice. Despite the amount of sparsely labeled images being significant, the adoption of those data still raises an open problem. Indeed, the deep learning approach frames the presence of dense annotations as a crucial factor. Recent efforts in literature have hence focused on developing label propagation techniques to expand sparse annotations into dense labels. However, the proposed methods proved only marginally effective for the purpose of segmenting the alveolar nerve in CBCT scans. This paper exploits and publicly releases a new 3D densely annotated dataset, through which we are able to train a deep label propagation model which obtains better results than those available in literature. By combining a segmentation model trained on the 3D annotated data and label propagation, we significantly improve the state of the art in the Inferior Alveolar Nerve segmentation.
Marco Cipriano, Stefano Allegretti, Federico Bolelli, Federico Pollastri, Costantino Grana
CVPR1
2020 Alzheimer's Garden: Understanding Social Behaviors of Patients with Dementia to Improve Their Quality of Life
abstract
Abstract This paper aims at understanding the social behavior of people with dementia through the use of technology, specifically by analyzing localization data of patients of an Alzheimer’s assisted care home in Italy. The analysis will allow to promote social relations by enhancing the facility’s spaces and activities, with the ultimate objective of improving residents’ quality of life. To assess social wellness and evaluate the effectiveness of the village areas and activities, this work introduces measures of sociability for both residents and places. Our data analysis is based on classical statistical methods and innovative machine learning techniques. First, we analyze the correlation between relational indicators and factors such as the outdoor temperature and the patients’ movements inside the facility. Then, we use statistical and accessibility analyses to determine the spaces residents appreciate the most and those in need of enhancements. We observe that patients’ sociability is strongly related to the considered factors. From our analysis, outdoor areas result less frequented and need spatial redesign to promote accessibility and attendance among patients. The data awareness obtained from our analysis will also be of great help to caregivers, doctors, and psychologists to enhance assisted care home social activities, adjust patient-specific treatments, and deepen the comprehension of the disease.
Gloria Bellini, Marco Cipriano, Nicola De Angeli, Jacopo Pio Gargano, Matteo Gianella, Gianluca Goi, Gabriele Rossi, Andrea Masciadri, Sara Comai
ICCHP (2)2
2020 The color out of space: learning self-supervised representations for Earth Observation imagery
abstract
The recent growth in the number of satellite images fosters the development of effective deep-learning techniques for Remote Sensing (RS). However, their full potential is untapped due to the lack of large annotated datasets. Such a problem is usually countered by fine-tuning a feature extractor that is previously trained on the ImageNet dataset. Unfortunately, the domain of natural images differs from the RS one, which hinders the final performance. In this work, we propose to learn meaningful representations from satellite imagery, leveraging its high-dimensionality spectral bands to reconstruct the visible colors. We conduct experiments on land cover classification (BigEarthNet) and West Nile Virus detection, showing that colorization is a solid pretext task for training a feature extractor. Furthermore, we qualitatively observe that guesses based on natural images and colorization rely on different parts of the input. This paves the way to an ensemble model that eventually outperforms both the above-mentioned techniques.
Stefano Vincenzi, Angelo Porrello, Pietro Buzzega, Marco Cipriano, Pietro Fronte, Roberto Cuccu, Carla Ippoliti, Annamaria Conte, Simone Calderara
ICPR4