Bruno Lepri

dblp:99/6489 · DBLP profile ↗
← Back
79ranked-venue papers
6as first author
29since 2021 · last 2026
0000-0003-1275-2333ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 26 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 25 · 1 first-author · 16 since 2021Human-computer interaction and ubiquitous computing · 21 · 3 first-author · 2 since 2021Databases, data management, data science and information retrieval · 10 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 3 since 2021Theory of computation · 2Systems, architecture and hardware · 1Computer networks · 1
YearPublicationVenuePosition
2026 Don't Stop the Multi-Party! On Generating Synthetic Written Multi-Party Conversations with Constraints
abstract
Written Multi-Party Conversations (WMPCs) are widely studied across disciplines, with social media as a primary data source due to their accessibility. However, these datasets raise privacy concerns and often reflect platform-specific properties. For example, interactions between speakers may be limited due to rigid platform structures (e.g., threads, tree-like discussions), which yield overly simplistic interaction patterns (e.g., one-to-one ``reply-to'' links). This work explores the feasibility of generating synthetic WMPCs with instruction-tuned Large Language Models (LLMs) by providing deterministic constraints such as dialogue structure and participants’ stance. We investigate two complementary strategies of leveraging LLMs in this context: (i.) LLMs as WMPC generators, where we task the LLM to generate a whole WMPC at once and (ii.) LLMs as WMPC parties, where the LLM generates one turn of the conversation at a time (made of speaker, addressee and message), provided the conversation history. We next introduce an analytical framework to evaluate compliance with the constraints, content quality, and interaction complexity for both strategies. Finally, we assess the level of obtained WMPCs via human and LLM-as-a-judge evaluations. We find stark differences among LLMs, with only some being able to generate high-quality WMPCs. We also find that turn-by-turn generation yields better conformance to constraints and higher linguistic variability than generating WMPCs in one pass. Nonetheless, our structural and qualitative evaluation indicates that both generation strategies can yield high-quality WMPCs.
Nicolò Penzo, Marco Guerini, Bruno Lepri, Goran Glavas, Sara Tonelli
AAAI3
2026 ESPLoRA: Enhanced Spatial Precision with Low-Rank Adaption in Text-to-Image Diffusion Models for High-Definition Synthesis
Andrea Rigo, Luca Stornaiuolo, Mauro Martino, Bruno Lepri, Nicu Sebe
ICPR (6)4
2026 Insight in Sight: Complaint Detection and Aspect-Based Reasoning Through Visually-Grounded Reviews With VLLMs
Apoorva Singh, Soumitra Ghosh, Karanjot Singh, Bruno Lepri
IEEE Trans. Comput. Soc. Syst.4
2025 Fully-Geometric Cross-Attention for Point Cloud Registration
abstract
Point cloud registration approaches often fail when the overlap between point clouds is low due to noisy point correspondences. This work introduces a novel cross-attention mechanism tailored for Transformer-based architectures that tackles this problem, by fusing information from coordinates and features at the super-point level between point clouds. This formulation has remained unexplored primarily because it must guarantee rotation and translation invariance since point clouds reside in different and independent reference frames. We integrate the Gromov-Wasserstein distance into the cross-attention formulation to jointly compute distances between points across different point clouds and account for their geometric structure. By doing so, points from two distinct point clouds can attend to each other under arbitrary rigid transformations. At the point level, we also devise a self-attention mechanism that aggregates the local geometric structure information into point features for fine matching. Our formulation boosts the number of inlier correspondences, thereby yielding more precise registration results compared to state-of-the-art approaches. We have conducted an extensive evaluation on 3DMatch, 3DLoMatch, KITTI, and 3DCSR datasets. Project page: https://github.com/twowwj/FLAT.
Weijie Wang 0002, Guofeng Mei, Jian Zhang 0002, Nicu Sebe, Bruno Lepri, Fabio Poiesi
3DV5
2025 SMoSE: Sparse Mixture of Shallow Experts for Interpretable Reinforcement Learning in Continuous Control Tasks
abstract
Continuous control tasks often involve high-dimensional, dynamic, and non-linear environments. State-of-the-art performance in these tasks is achieved through complex closed-box policies that are effective, but suffer from an inherent opacity. Interpretable policies, while generally underperforming compared to their closed-box counterparts, advantageously facilitate transparent decision-making within automated systems. Hence, their usage is often essential for diagnosing and mitigating errors, supporting ethical and legal accountability, and fostering trust among stakeholders. In this paper, we propose SMoSE, a novel method to train sparsely activated interpretable controllers, based on a top-1 Mixture-of-Experts architecture. SMoSE combines a set of interpretable decision-makers, trained to be experts in different basic skills, and an interpretable router that assigns tasks among the experts. The training is carried out via state-of-the-art Reinforcement Learning algorithms, exploiting load-balancing techniques to ensure fair expert usage. We then distill decision trees from the weights of the router, significantly improving the ease of interpretation. We evaluate SMoSE on six benchmark environments from MuJoCo: our method outperforms recent interpretable baselines and narrows the gap with non-interpretable state-of-the-art algorithms.
Mátyás Vincze, Laura Ferrarotti, Leonardo Lucio Custode, Bruno Lepri, Giovanni Iacca
AAAI4
2025 Separation Power of Equivariant Neural Networks
abstract
The separation power of a machine learning model refers to its ability to distinguish between different inputs and is often used as a proxy for its expressivity. Indeed, knowing the separation power of a family of models is a necessary condition to obtain fine-grained universality results. In this paper, we analyze the separation power of equivariant neural networks, such as convolutional and permutation-invariant networks. We first present a complete characterization of inputs indistinguishable by models derived by a given architecture. From this results, we derive how separability is influenced by hyperparameters and architectural choices—such as activation functions, depth, hidden layer width, and representation types. Notably, all non-polynomial activations, including ReLU and sigmoid, are equivalent in expressivity and reach maximum separation power. Depth improves separation power up to a threshold, after which further increases have no effect. Adding invariant features to hidden representations does not impact separation power. Finally, block decomposition of hidden representations affects separability, with minimal components forming a hierarchy in separation power that provides a straightforward method for comparing the separation power of models.
Marco Pacini, Xiaowen Dong 0001, Bruno Lepri, Gabriele Santin
ICLR3
2025 FreeInsert: Disentangled Text-Guided Object Insertion in 3D Gaussian Scene without Spatial Priors
Weijie Wang 0002, Qiang Li 0048, Nicu Sebe, Bruno Lepri, Weizhi Nie
ACM Multimedia5
2025 Bridging Theory and Practice in Link Representation with Graph Neural Networks
abstract
Graph Neural Networks (GNNs) are widely used to compute representations of node pairs for downstream tasks such as link prediction. Yet, theoretical understanding of their expressive power has focused almost entirely on graph-level representations. In this work, we shift the focus to links and provide the first comprehensive study of GNN expressiveness in link representation. We introduce a unifying framework, the $k_\phi$-$k_\rho$-$m$ framework, that subsumes existing message-passing link models and enables formal expressiveness comparisons. Using this framework, we derive a hierarchy of state-of-the-art methods and offer theoretical tools to analyze future architectures. To complement our analysis, we propose a synthetic evaluation protocol comprising the first benchmark specifically designed to assess link-level expressiveness. Finally, we ask: does expressiveness matter in practice? We use a graph symmetry metric that quantifies the difficulty of distinguishing links and show that while expressive models may underperform on standard benchmarks, they significantly outperform simpler ones as symmetry increases, highlighting the need for dataset-aware model selection.
Veronica Lachi, Francesco Ferrini, Antonio Longa, Bruno Lepri, Andrea Passerini, Manfred Jaeger
NeurIPS4
2025 On Universality Classes of Equivariant Networks
abstract
Equivariant neural networks provide a principled framework for incorporating symmetry into learning architectures and have been extensively analyzed through the lens of their *separation power*, that is, the ability to distinguish inputs modulo symmetry. This notion plays a central role in settings such as graph learning, where it is often formalized via the Weisfeilern–Leman hierarchy. In contrast, the *universality* of equivariant models—their capacity to approximate target functions—remains comparatively underexplored. In this work, we investigate the approximation power of equivariant neural networks beyond separation constraints. We show that separation power does not fully capture expressivity: models with identical separation power may differ in their approximation ability. To demonstrate this, we characterize the universality classes of shallow invariant networks, providing a general framework for understanding which functions these architectures can approximate. Since equivariant models reduce to invariant ones under projection, this analysis yields sufficient conditions under which shallow equivariant networks fail to be universal. Conversely, we identify settings where shallow models do achieve separation-constrained universality. These positive results, however, depend critically on structural properties of the symmetry group, such as the existence of adequate normal subgroups, which may not hold in important cases like permutation symmetry.
Marco Pacini, Gabriele Santin, Bruno Lepri, Shubhendu Trivedi
NeurIPS3
2025 You Don't Bring Me Flowers: Mitigating Unwanted Recommendations Through Conformal Risk Control
abstract
Recommenders are significantly shaping online information consumption.While effective at personalizing content, these systems increasingly face criticism for propagating irrelevant, unwanted, and even harmful recommendations.Such content degrades user satisfaction and contributes to significant societal issues, including misinformation, radicalization, and erosion of user trust.Although platforms offer mechanisms to mitigate exposure to undesired content, these mechanisms are often insufficiently effective and slow to adapt to users' feedback.This paper introduces an intuitive, modelagnostic, and distribution-free method that uses conformal risk control to provably bound unwanted content in personalized recommendations by leveraging simple binary feedback on items.We also address a limitation of traditional conformal risk control approaches, i.e., the fact that the recommender can provide a smaller set of recommended items, by leveraging implicit feedback on consumed items to expand the recommendation set while ensuring robust risk mitigation.Our experimental evaluation on data coming from a popular online video-sharing platform demonstrates that our approach ensures an effective and controllable reduction of unwanted recommendations with minimal effort.The source code is available here: https://github.com/geektoni/mitigating-harm-recsys.
Giovanni De Toni, Erasmo Purificato, Emilia Gómez, Andrea Passerini, Bruno Lepri, Cristian Consonni
RecSys5
2025 T2TD: Text-3D Generation Model Based on Prior Knowledge Guidance
abstract
In recent years, 3D models have been utilized in many applications, such as auto-drivers, 3D reconstruction, VR, and AR. However, the scarcity of 3D model data does not meet its practical demands. Thus, generating high-quality 3D models efficiently from textual descriptions is a promising but challenging way to solve this problem. In this paper, inspired by the creative mechanisms of human imagination, which concretely supplement the target model from ambiguous descriptions built upon human experiential knowledge, we propose a novel text-3D generation model (T2TD). T2TD aims to generate the target model based on the textual description with the aid of experiential knowledge. Its target creation process simulates the imaginative mechanisms of human beings. In this process, we first introduce the text-3D knowledge graph to preserve the relationship between 3D models and textual semantic information, which provides related shapes like humans' experiential information. Second, we propose an effective causal inference model to select useful feature information from these related shapes, which can remove the unrelated structure information and only retain solely the feature information strongly related to the textual description. Third, we adopt a novel multi-layer transformer structure to progressively fuse this strongly related structure information and textual information, compensating for the lack of structural information, and enhancing the final performance of the 3D generation model. The final experimental results demonstrate that our approach significantly improves 3D model generation quality and outperforms the SOTA methods on the text2shape datasets.
Weizhi Nie, Rui-dong Chen, Weijie Wang 0002, Bruno Lepri, Nicu Sebe
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 Putting Context in Context: the Impact of Discussion Structure on Text Classification
abstract
Nicolò Penzo, Antonio Longa, Bruno Lepri, Sara Tonelli, Marco Guerini. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Nicolò Penzo, Antonio Longa, Bruno Lepri, Sara Tonelli, Marco Guerini
EACL (1)3
2024 Do LLMs suffer from Multi-Party Hangover? A Diagnostic Approach to Addressee Recognition and Response Selection in Conversations
abstract
Assessing the performance of systems to classify Multi-Party Conversations (MPC) is challenging due to the interconnection between linguistic and structural characteristics of conversations.Conventional evaluation methods often overlook variances in model behavior across different levels of structural complexity on interaction graphs.In this work, we propose a methodological pipeline to investigate model performance across specific structural attributes of conversations.As a proof of concept we focus on Response Selection and Addressee Recognition tasks, to diagnose model weaknesses.To this end, we extract representative diagnostic subdatasets with a fixed number of users and a good structural variety from a large and open corpus of online MPCs.We further frame our work in terms of data minimization, avoiding the use of original usernames to preserve privacy, and propose alternatives to using original text messages.Results show that response selection relies more on the textual content of conversations, while addressee recognition requires capturing their structural dimension.Using an LLM in a zero-shot setting, we further highlight how sensitivity to prompt variations is task-dependent.
Nicolò Penzo, Maryam Sajedinia, Bruno Lepri, Sara Tonelli, Marco Guerini
EMNLP3
2024 A Characterization Theorem for Equivariant Networks with Point-wise Activations
abstract
Equivariant neural networks have shown improved performance, expressiveness and sample complexity on symmetrical domains. But for some specific symmetries, representations, and choice of coordinates, the most common point-wise activations, such as ReLU, are not equivariant, hence they cannot be employed in the design of equivariant neural networks. The theorem we present in this paper describes all possibile combinations of representations, choice of coordinates and point-wise activations to obtain an equivariant layer, generalizing and strengthening existing characterizations. Notable cases of practical relevance are discussed as corollaries. Indeed, we prove that rotation-equivariant networks can only be invariant, as it happens for any network which is equivariant with respect to connected compact groups. Then, we discuss implications of our findings when applied to important instances of equivariant networks. First, we completely characterize permutation equivariant networks such as Invariant Graph Networks with point-wise nonlinearities and their geometric counterparts, highlighting a plethora of models whose expressive power and performance are still unknown. Second, we show that feature spaces of disentangled steerable convolutional neural networks are trivial representations.
Marco Pacini, Xiaowen Dong 0001, Bruno Lepri, Gabriele Santin
ICLR3
2024 UVMap-ID: A Controllable and Personalized UV Map Generative Model
abstract
Recently, diffusion models have made significant strides in synthesizing realistic 2D human images based on provided text prompts. Building upon this, researchers have extended 2D text-to-image diffusion models into the 3D domain for generating human textures (UV Maps). However, some important problems about UV Map Generative models are still not solved, i.e., how to generate personalized texture maps for any given face image, and how to define and evaluate the quality of these generated texture maps. To solve the above problems, we introduce a novel method, UVMap-ID, which is a controllable and personalized UV Map generative model. Unlike traditional large-scale training methods in 2D, we propose to fine-tune a pre-trained text-to-image diffusion model which is integrated with a face fusion module for achieving ID-driven customized generation. To support the finetuning strategy, we introduce a small-scale attribute-balanced training dataset, including high-quality textures with labeled text and Face ID. Additionally, we introduce some metrics to evaluate the multiple aspects of the textures. Finally, both quantitative and qualitative analyses demonstrate the effectiveness of our method in controllable and personalized UV Map generation.
Weijie Wang 0002, Jichao Zhang, Chang Liu 0030, Xia Li 0005, Xingqian Xu, Humphrey Shi, Nicu Sebe, Bruno Lepri
ACM Multimedia8
2024 Preference Elicitation in Interactive and User-centered Algorithmic Recourse: an Initial Exploration
abstract
Algorithmic Recourse aims to provide actionable explanations, or recourse plans, to overturn potentially unfavourable decisions taken by automated machine learning models. In this paper, we propose an interaction paradigm based on a guided interaction pattern aimed at both eliciting the users’ preferences and heading them toward effective recourse interventions. In a fictional task of money lending, we compare this approach with an exploratory interaction pattern based on a combination of alternative plans and the possibility of freely changing the configurations by the users themselves. Our results suggest that users may recognize that the guided interaction paradigm improves efficiency. However, they also feel less freedom to experiment with “what-if” scenarios. Nevertheless, the time spent on the purely exploratory interface tends to be perceived as a lack of efficiency, which reduces attractiveness, perspicuity, and dependability. Conversely, for the guided interface, more time on the interface seems to increase its attractiveness, perspicuity, and dependability while not impacting the perceived efficiency. That might suggest that this type of interfaces should combine these two approaches by trying to support exploratory behavior while gently pushing toward a guided effective solution.
Seyedehdelaram Esfahani, Giovanni De Toni, Bruno Lepri, Andrea Passerini, Katya Tentori, Massimo Zancanaro
UMAP3
2024 Spatial entropy as an inductive bias for vision transformers
abstract
Abstract Recent work on Vision Transformers (VTs) showed that introducing a local inductive bias in the VT architecture helps reducing the number of samples necessary for training. However, the architecture modifications lead to a loss of generality of the Transformer backbone, partially contradicting the push towards the development of uniform architectures, shared, e.g., by both the Computer Vision and the Natural Language Processing areas. In this work, we propose a different and complementary direction, in which a local bias is introduced using an auxiliary self-supervised task, performed jointly with standard supervised training. Specifically, we exploit the observation that the attention maps of VTs, when trained with self-supervision, can contain a semantic segmentation structure which does not spontaneously emerge when training is supervised. Thus, we explicitly encourage the emergence of this spatial clustering as a form of training regularization. In more detail, we exploit the assumption that, in a given image, objects usually correspond to few connected regions, and we propose a spatial formulation of the information entropy to quantify this object-based inductive bias. By minimizing the proposed spatial entropy, we include an additional self-supervised signal during training. Using extensive experiments, we show that the proposed regularization leads to equivalent or better results than other VT proposals which include a local bias by changing the basic Transformer architecture, and it can drastically boost the VT final accuracy when using small-medium training sets. The code is available at https://github.com/helia95/SAR .
Elia Peruzzo, Enver Sangineto, Marco De Nadai, Wei Bi, Bruno Lepri, Nicu Sebe
Mach. Learn.6
2023 Trajectory test-train overlap in next-location prediction datasets
Massimiliano Luca, Luca Pappalardo, Bruno Lepri, Gianni Barlacchi
Mach. Learn.3
2023 Synthesizing explainable counterfactual policies for algorithmic recourse with program synthesis
abstract
Abstract Being able to provide counterfactual interventions—sequences of actions we would have had to take for a desirable outcome to happen—is essential to explain how to change an unfavourable decision by a black-box machine learning model (e.g., being denied a loan request). Existing solutions have mainly focused on generating feasible interventions without providing explanations of their rationale. Moreover, they need to solve a separate optimization problem for each user. In this paper, we take a different approach and learn a program that outputs a sequence of explainable counterfactual actions given a user description and a causal graph. We leverage program synthesis techniques, reinforcement learning coupled with Monte Carlo Tree Search for efficient exploration, and rule learning to extract explanations for each recommended action. An experimental evaluation on synthetic and real-world datasets shows how our approach, FARE (eFficient counterfActual REcourse), generates effective interventions by making orders of magnitude fewer queries to the black-box classifier with respect to existing solutions, with the additional benefit of complementing them with interpretable explanations.
Giovanni De Toni, Bruno Lepri, Andrea Passerini
Mach. Learn.2
2023 Play&Go Corporate: An End-to-End Solution for Facilitating Urban Cyclability
abstract
Mobility plays a fundamental role in modern cities. How citizens experience the urban environment, access city core services, and participate in city life, strongly depends on its mobility organization and efficiency. The challenges that municipalities face are very ambitious: on the one hand, administrators must guarantee their citizens the right to mobility and to easily access local services; on the other hand, they need to minimize the economic, social, and environmental costs of the mobility system. Municipalities are increasingly facing problems of traffic congestion, road safety, energy dependency and air pollution, and therefore encouraging a shift towards sustainable mobility habits based on active mobility is of central importance. Active modes, such as cycling, should be particularly encouraged, especially for local recurrent journeys (e.g., home–to–school, home–to–work). In this context, addressing and mitigating commuter-generated traffic requires engaging public and private stakeholders through innovative and collaborative approaches that focus not only on supply (e.g., roads and vehicles) but also on transportation demand management. In this paper, we present and end-to-end solution, called Play&Go Corporate, for enabling urban cyclability and its concrete exploitation in the realization of a home-to-work sustainable mobility campaign (i.e., BIKE2 WORK) targeting employees of public and private companies. To evaluate the effectiveness of the proposed solution we developed two analyses: the first to carefully analyze the user experience and any behaviour change related to the BIKE2 WORK mobility campaign, and the second to demonstrate how exploiting the collected data we can potentially inform and guide the involved municipality (i.e., Ferrara, a city in Northern Italy) in improving urban cyclability.
Antonio Bucchiarone, Simone Bassanelli, Massimiliano Luca, Simone Centellegher, Piergiorgio Cipriano, Luca Giovannini, Bruno Lepri, Annapaola Marconi
IEEE Trans. Intell. Transp. Syst.7
2023 ISF-GAN: An Implicit Style Function for High-Resolution Image-to-Image Translation
abstract
Recently, there has been an increasing interest in image editing methods that employ pre-trained unconditional image generators (e.g., StyleGAN). However, applying these methods to translate images to multiple visual domains remains challenging. Existing works do not often preserve the domain-invariant part of the image (e.g., the identity in human face translations), or they do not usually handle multiple domains or allow for multi-modal translations. This work proposes an implicit style function (ISF) to straightforwardly achieve multi-modal and multi-domain image-to-image translation from pre-trained unconditional generators. The ISF manipulates the semantics of a latent code to ensure that the image generated from the manipulated code lies in the desired visual domain. Our human faces and animal image manipulations show significantly improved results over the baselines. Our model enables cost-effective multi-modal unsupervised image-to-image translations at high resolution using pre-trained unconditional GANs. The code and data are available at:https://github.com/yhlleo/stylegan-mmuit.
Linchao Bao, Nicu Sebe, Bruno Lepri, Marco De Nadai
IEEE Trans. Multim.5
2022 Understanding and Rewiring Cities
Bruno Lepri, Simone Centellegher, Marco De Nadai
ADBIS1
2022 An efficient procedure for mining egocentric temporal motifs
abstract
Abstract Temporal graphs are structures which model relational data between entities that change over time. Due to the complex structure of data, mining statistically significant temporal subgraphs, also known as temporal motifs, is a challenging task. In this work, we present an efficient technique for extracting temporal motifs in temporal networks. Our method is based on the novel notion of egocentric temporal neighborhoods, namely multi-layer structures centered on an ego node. Each temporal layer of the structure consists of the first-order neighborhood of the ego node, and corresponding nodes in sequential layers are connected by an edge. The strength of this approach lies in the possibility of encoding these structures into a unique bit vector, thus bypassing the problem of graph isomorphism in searching for temporal motifs. This allows our algorithm to mine substantially larger motifs with respect to alternative approaches. Furthermore, by bringing the focus on the temporal dynamics of the interactions of a specific node, our model allows to mine temporal motifs which are visibly interpretable. Experiments on a number of complex networks of social interactions confirm the advantage of the proposed approach over alternative non-egocentric solutions. The egocentric procedure is indeed more efficient in revealing similarities and discrepancies among different social environments, independently of the different technologies used to collect data, which instead affect standard non-egocentric measures.
Antonio Longa, Giulia Cencetti, Bruno Lepri, Andrea Passerini
Data Min. Knowl. Discov.3
2022 Optimizing City-Scale Traffic Through Modeling Observations of Vehicle Movements
abstract
The capability of traffic-information systems to sense the movement of millions of users and offer trip plans through mobile phones has enabled a new way of optimizing city traffic dynamics, turning transportation big data into insights and actions in a closed-loop and evaluating this approach in the real world. Existing research has applied dynamic Bayesian networks and deep neural networks to make traffic predictions from floating car data, utilized dynamic programming and simulation approaches to identify how people normally travel with dynamic traffic assignment for policy research, and introduced Markov decision processes and reinforcement learning to optimally control traffic signals. However, none of these works utilized floating car data to suggest departure times and route choices in order to optimize city traffic dynamics. In this paper, we present a study showing that floating car data can lead to lower average trip time, higher on-time arrival ratio, and higher Charypar-Nagel score compared with how people normally travel. The study is based on optimizing a partially observable discrete-time decision process and is evaluated in one synthesized scenario, one partly synthesized scenario, and three real-world scenarios. This study points to the potential of a “living lab” approach where we learn, predict, and optimize behaviors in the real world.
Fan Yang 0057, Alina Vereshchaka, Bruno Lepri, Wen Dong 0001
IEEE Trans. Intell. Transp. Syst.3
2021 Smoothing the Disentangled Latent Style Space for Unsupervised Image-to-Image Translation
abstract
Image-to-Image (I2I) multi-domain translation models are usually evaluated also using the quality of their semantic interpolation results. However, state-of-the-art models frequently show abrupt changes in the image appearance during interpolation, and usually perform poorly in interpolations across domains. In this paper, we propose a new training protocol based on three specific losses which help a translation network to learn a smooth and disentangled latent style space in which: 1) Both intra- and inter-domain interpolations correspond to gradual changes in the generated images and 2) The content of the source image is better preserved during the translation. Moreover, we propose a novel evaluation metric to properly measure the smoothness of latent style space of I2I translation models. The proposed method can be plugged in existing translation approaches, and our extensive experiments on different datasets show that it can significantly boost the quality of the generated images and the graduality of the interpolations.
Enver Sangineto, Linchao Bao, Haoxian Zhang, Nicu Sebe, Bruno Lepri, Wei Wang 0108, Marco De Nadai
CVPR7
2021 Click to Move: Controlling Video Generation with Sparse Motion
abstract
This paper introduces Click to Move (C2M), a novel framework for video generation where the user can control the motion of the synthesized video through mouse clicks specifying simple object trajectories of the key objects in the scene. Our model receives as input an initial frame, its corresponding segmentation map and the sparse motion vectors encoding the input provided by the user. It outputs a plausible video sequence starting from the given frame and with a motion that is consistent with user input. Notably, our proposed deep architecture incorporates a Graph Convolution Network (GCN) modelling the movements of all the objects in the scene in a holistic manner and effectively combining the sparse user motion information and image features. Experimental results show that C2M outperforms existing methods on two publicly available datasets, thus demonstrating the effectiveness of our GCN framework at modelling object interactions. The source code is publicly available at https://github.com/PierfrancescoArdino/C2M.
Pierfrancesco Ardino, Marco De Nadai, Bruno Lepri, Elisa Ricci 0001, Stéphane Lathuilière
ICCV3
2021 Efficient Training of Visual Transformers with Small Datasets
abstract
Visual Transformers (VTs) are emerging as an architectural paradigm alternative to Convolutional networks (CNNs). Differently from CNNs, VTs can capture global relations between image elements and they potentially have a larger representation capacity. However, the lack of the typical convolutional inductive bias makes these models more data hungry than common CNNs. In fact, some local properties of the visual domain which are embedded in the CNN architectural design, in VTs should be learned from samples. In this paper, we empirically analyse different VTs, comparing their robustness in a small training set regime, and we show that, despite having a comparable accuracy when trained on ImageNet, their performance on smaller datasets can be largely different. Moreover, we propose an auxiliary self-supervised task which can extract additional information from images with only a negligible computational overhead. This task encourages the VTs to learn spatial relations within an image and makes the VT training much more robust when training data is scarce. Our task is used jointly with the standard (supervised) training and it does not depend on specific architectural choices, thus it can be easily plugged in the existing VTs. Using an extensive evaluation with different VTs and datasets, we show that our method can improve (sometimes dramatically) the final accuracy of the VTs. Our code is available at: https://github.com/yhlleo/VTs-Drloc.
Enver Sangineto, Wei Bi, Nicu Sebe, Bruno Lepri, Marco De Nadai
NeurIPS5
2021 Between privacy and security: the factors that drive intentions to use cyber-security applications
abstract
Installing security applications is a common way to protect against malicious apps, phishing emails, and other threats in mobile operating systems. While these applications can provide essential security protections, they also tend to access large amounts of people's sensitive information. Therefore, individuals need to evaluate the trade-off between the security features and the privacy invasion when deciding on which protection mechanisms to use. In this paper, we examine factors affecting the willingness to install mobile security applications by taking into account the invasion levels and security features of cyber-security applications. To this end, we propose a visual language that depicts the coverage of different security features as well as privacy intrusiveness levels. Our user study (n=300) shows that users assessing security applications find their trade-off balance in highly secure apps with a medium level of privacy invasion. The results indicate that a low privacy invasion might signal that the security application provides less security. We discuss these findings in the context of understanding the trade-off between privacy and security.
Hadas Chassidim, Christos Perentis, Eran Toch, Bruno Lepri
Behav. Inf. Technol.4
2021 Land Use Classification With Point of Interests and Structural Patterns
abstract
In this paper, we present a framework for performing automatic analysis of Land Use Zones based on Location-Based Social Networks (LBSNs). We model city areas using a hierarchical structure of POIs extracted from foursquare. We encode such structures in kernel machines, e.g., Support Vector Machines, using a new Tree Kernel, i.e., the Hierarchical POI Kernel (HPK), which can take the importance of the individual POIs into account during the substructure matching. This way, HPK projects structures in the space of all their possible substructures such that each dimension corresponds to a semantic structural feature, weighted according to the discriminative power of POIs . We generated four different datasets for the following cities: Barcelona, Lisbon, Amsterdam and Milan, where we trained and tested our models. The results show that our approach largely outperforms previous work and standard baseline built on simple features, such as counts of different POIs. Finally, we apply a mining algorithm to extract the most relevant features (tree fragments) from the implicit TK space according to the weights the kernel machine assigned to them. Our approach can produce an explicit set of representative features that can be used to classify and characterize urban areas.
Gianni Barlacchi, Bruno Lepri, Alessandro Moschitti
IEEE Trans. Knowl. Data Eng.2
2020 Understanding Individual Behaviour: From Virtual to Physical Patterns
abstract
As "Big Data" has become pervasive, an increasing amount of research has connected the dots between human behaviour in the offline and online worlds. Consequently, researchers have exploited these new findings to create models that better predict different aspects of human life and recommend future behaviour. To date, however, we do not yet fully understand the similarities and differences of human behaviour in these virtual and physical worlds. Here, we analyse and discuss the mobility and application usage of 400,000 individuals over eight months. We find an astonishing similarity between people's mobility in the physical space and how they move from app to app in smartphones. Our data shows that individuals use and visit a finite number of apps and places, but they keep exploring over time. In particular, two distinct profiles of individuals emerge: those that keep changing places and services, and those that are stable over time, named as "explorers" and "keepers". We see these findings as crucial to enrich a discussion for the potentials and the challenges of building human-centric AI systems, which might leverage recent results in Computational Social Science.
Marco De Nadai, Bruno Lepri, Nuria Oliver
ECAI2
2020 Semantic-Guided Inpainting Network for Complex Urban Scenes Manipulation
abstract
Manipulating images of complex scenes to reconstruct, insert and/or remove specific object instances is a challenging task. Complex scenes contain multiple semantics and objects, which are frequently cluttered or ambiguous, thus hampering the performance of inpainting models. Conventional techniques often rely on structural information such as object contours in multi-stage approaches that generate unreliable results and boundaries. In this work, we propose a novel deep learning model to alter a complex urban scene by removing a user-specified portion of the image and coherently inserting a new object (e.g. car or pedestrian) in that scene. Inspired by recent works on image inpainting, our proposed method leverages the semantic segmentation to model the content and structure of the image, and learn the best shape and location of the object to insert. To generate reliable results, we design a new decoder block that combines the semantic segmentation and generation task to guide better the generation of new objects and scenes, which have to be semantically consistent with the image. Our experiments, conducted on two large-scale datasets of urban scenes (Cityscapes and Indian Driving), show that our proposed approach successfully address the problem of semantically-guided inpainting of complex urban scene.
Pierfrancesco Ardino, Elisa Ricci 0001, Bruno Lepri, Marco De Nadai
ICPR4
2020 Retrieval Guided Unsupervised Multi-domain Image to Image Translation
abstract
Image to image translation aims to learn a mapping that transforms an image from one visual domain to another. Recent works assume that images descriptors can be disentangled into a domain-invariant content representation and a domain-specific style representation. Thus, translation models seek to preserve the content of source images while changing the style to a target visual domain. However, synthesizing new images is extremely challenging especially in multi-domain translations, as the network has to compose content and style to generate reliable and diverse images in multiple domains. In this paper we propose the use of an image retrieval system to assist the image-to-image translation task. First, we train an image-to-image translation model to map images to multiple domains. Then, we train an image retrieval model using real and generated images to find images similar to a query one in content but in a different domain. Finally, we exploit the image retrieval system to fine-tune the image-to-image translation model and generate higher quality images. Our experiments show the effectiveness of the proposed solution and highlight the contribution of the retrieval network, which can benefit from additional unlabeled data and help image-to-image translation models in the presence of scarce data.
Raul Gomez, Marco De Nadai, Dimosthenis Karatzas, Bruno Lepri, Nicu Sebe
ACM Multimedia5
2020 Describe What to Change: A Text-guided Unsupervised Image-to-image Translation Approach
abstract
Manipulating visual attributes of images through human-written text is a very challenging task. On the one hand, models have to learn the manipulation without the ground truth of the desired output. On the other hand, models have to deal with the inherent ambiguity of natural language. Previous research usually requires either the user to describe all the characteristics of the desired image or to use richly-annotated image captioning datasets. In this work, we propose a novel unsupervised approach, based on image-to-image translation, that alters the attributes of a given image through a command-like sentence such as "change the hair color to black". Contrarily to state-of-the-art approaches, our model does not require a human-annotated dataset nor a textual description of all the attributes of the desired image, but only those that have to be modified. Our proposed model disentangles the image content from the visual attributes, and it learns to modify the latter using the textual description, before generating a new image from the content and the modified attribute representation. Because text might be inherently ambiguous (blond hair may refer to different shadows of blond, e.g. golden, icy, sandy), our method generates multiple stochastic versions of the same translation. Experiments show that the proposed model achieves promising performances on two large-scale public datasets: CelebA and CUB. We believe our approach will pave the way to new avenues of research combining textual and speech commands with visual attributes.
Marco De Nadai, Deng Cai 0002, Xavier Alameda-Pineda, Nicu Sebe, Bruno Lepri
ACM Multimedia7
2020 Learning Mobility Flows from Urban Features with Spatial Interaction Models and Neural Networks**To appear in the Proceedings of 2020 IEEE International Conference on Smart Computing (SMARTCOMP 2020)
abstract
A fundamental problem of interest to policy makers, urban planners, and other stakeholders involved in urban development is assessing the impact of planning and construction activities on mobility flows. This is a challenging task due to the different spatial, temporal, social, and economic factors influencing urban mobility flows. These flows, along with the influencing factors, can be modelled as attributed graphs with both node and edge features characterising locations in a city and the various types of relationships between them. In this paper, we address the problem of assessing origin-destination (OD) car flows between a location of interest and every other location in a city, given their features and the structural characteristics of the graph. We propose three neural network architectures, including graph neural networks (GNN), and conduct a systematic comparison between the proposed methods and state-of-the-art spatial interaction models, their modifications, and machine learning approaches. The objective of the paper is to address the practical problem of estimating potential flow between an urban project location and other locations in the city, where the features of the project location are known in advance. We evaluate the performance of the models on a regression task using a custom data set of attributed car OD flows in London. We also visualise the model performance by showing the spatial distribution of flow residuals across London.
Gevorg Yeghikyan, Felix L. Opolka, Mirco Nanni, Bruno Lepri, Pietro Liò
SMARTCOMP4
2020 Modelling Taxi Drivers' Behaviour for the Next Destination Prediction
abstract
In this paper, we study how to model taxi drivers' behavior and geographical information for an interesting and challenging task: the next destination prediction in a taxi journey. Predicting the next location is a well-studied problem in human mobility, which finds several applications in real-world scenarios, from optimizing the efficiency of electronic dispatching systems to predicting and reducing the traffic jam. This task is normally modeled as a multiclass classification problem, where the goal is to select, among a set of already known locations, the next taxi destination. We present a Recurrent Neural Network (RNN) approach that models the taxi drivers' behavior and encodes the semantics of visited locations by using geographical information from Location-Based Social Networks (LBSNs). In particular, the RNNs are trained to predict the exact coordinates of the next destination, overcoming the problem of producing, in output, a limited set of locations, seen during the training phase. The proposed approach was tested on the ECML/PKDD Discovery Challenge 2015 dataset-based on the city of Porto-, obtaining better results with respect to the competition winner, whilst using less information, and on Manhattan and San Francisco datasets.
Alberto Rossi, Gianni Barlacchi, Monica Bianchini, Bruno Lepri
IEEE Trans. Intell. Transp. Syst.4
2019 Detecting Permanent and Intermittent Purchase Hotspots via Computational Stigmergy
abstract
The analysis of credit card transactions allows gaining new insights into the spending occurrences and mobility behavior of large numbers of individuals at an unprecedented scale. However, unfolding such spatiotemporal patterns at a community level implies a non-trivial system modeling and parametrization, as well as, a proper representation of the temporal dynamic. In this work we address both those issues by means of a novel computational technique, i.e. computational stigmergy. By using computational stigmergy each sample position is associated with a digital pheromone deposit, which aggregates with other deposits according to their spatiotemporal proximity. By processing transactions data with computational stigmergy, it is possible to identify high-density areas (hotspots) occurring in different time and days, as well as, analyze their consistency over time. Indeed, a hotspot can be permanent, i.e. present throughout the period of observation, or intermittent, i.e. present only in certain time and days due to community level occurrences (e.g. nightlife). Such difference is not only spatial (where the hotspot occurs) and temporal (when the hotspot occurs) but affects also which people visit the hotspot. The proposed approach is tested on a real-world dataset containing the credit card transaction of 60k users between 2014 and 2015.
Antonio L. Alfeo, Mario G. C. A. Cimino, Bruno Lepri, Alex Pentland, Gigliola Vaglini
ICPRAM3
2019 Urban Swarms: A new approach for autonomous waste management
abstract
Modern cities are growing ecosystems that face new challenges due to the increasing population demands. One of the many problems they face nowadays is waste management, which has become a pressing issue requiring new solutions. Swarm robotics systems have been attracting an increasing amount of attention in the past years and they are expected to become one of the main driving factors for innovation in the field of robotics. The research presented in this paper explores the feasibility of a swarm robotics system in an urban environment. By using bio-inspired foraging methods such as multi-place foraging and stigmergy-based navigation, a swarm of robots is able to improve the efficiency and autonomy of the urban waste management system in a realistic scenario. To achieve this, a diverse set of simulation experiments was conducted using real-world GIS data and implementing different garbage collection scenarios driven by robot swarms. Results presented in this research show that the proposed system outperforms current approaches. Moreover, results not only show the efficiency of our solution, but also give insights about how to design and customize these systems.
Antonio L. Alfeo, Eduardo Castelló Ferrer, Yago Lizarribar 0001, Arnaud Grignard, Luis Alonso Pastor, Dylan T. Sleeper, Mario G. C. A. Cimino, Bruno Lepri, Gigliola Vaglini, Kent Larson, Marco Dorigo, Alex Pentland
ICRA8
2019 Following Wrong Suggestions: Self-blame in Human and Computer Scenarios
Andrea Beretta, Massimo Zancanaro, Bruno Lepri
INTERACT (3)3
2019 Gesture-to-Gesture Translation in the Wild via Category-Independent Conditional Maps
abstract
Recent works have shown Generative Adversarial Networks (GANs) to be particularly effective in image-to-image translations. However, in tasks such as body pose and hand gesture translation, existing methods usually require precise annotations, e.g. key-points or skeletons, which are time-consuming to draw. In this work, we propose a novel GAN architecture that decouples the required annotations into a category label - that specifies the gesture type - and a simple-to-draw category-independent conditional map - that expresses the location, rotation and size of the hand gesture. Our architecture synthesizes the target gesture while preserving the background context, thus effectively dealing with gesture translation in the wild. To this aim, we use an attention module and a rolling guidance approach, which loops the generated images back into the network and produces higher quality images compared to competing works. Thus, our GAN learns to generate new images from simple annotations without requiring key-points or skeleton labels. Results on two public datasets show that our method outperforms state of the art approaches both quantitatively and qualitatively. To the best of our knowledge, no work so far has addressed the gesture-to-gesture translation in the wild by requiring user-friendly annotations.
Marco De Nadai, Gloria Zen, Nicu Sebe, Bruno Lepri
ACM Multimedia5
2019 Assessing Refugees' Integration via Spatio-Temporal Similarities of Mobility and Calling Behaviors
abstract
In Turkey, the increasing tension, due to the presence of 3.4 million Syrian refugees, demands the formulation of effective integration policies. Moreover, their design requires tools aimed at understanding the integration of refugees despite the complexity of this phenomenon. In this work, we propose a set of metrics aimed at providing insights and assessing the integration of Syrian refugees, by analyzing a real-world call detail record (CDR) dataset including calls from refugees and locals in Turkey throughout 2017. Specifically, we exploit the similarity between refugees' and locals' spatial and temporal behaviors, in terms of communication and mobility in order to assess integration dynamics. Together with the already known methods for data analysis, we use a novel computational approach to analyze spatio-temporal patterns: computational stigmergy, a bio-inspired scalar and temporal aggregation of samples. Computational stigmergy associates each sample with a virtual pheromone deposit (mark). Marks in spatiotemporal proximity are aggregated into functional structures called trails, which summarize the spatiotemporal patterns in data and allow computing the similarity between different patterns. According to our results, collective mobility and behavioral similarity with locals have great potential as measures of integration, since they are: 1) correlated with the amount of interaction with locals; 2) an effective proxy for refugee's economic capacity, and thus refugee's potential employment; and 3) able to capture events that may disrupt the integration phenomena, such as social tension.
Antonio L. Alfeo, Mario G. C. A. Cimino, Bruno Lepri, Alex Pentland, Gigliola Vaglini
IEEE Trans. Comput. Soc. Syst.3
2018 Weak Nodes Detection in Urban Transport Systems: Planning for Resilience in Singapore
abstract
The availability of massive data-sets describing human mobility offers the possibility to design simulation tools to monitor and improve the resilience of transport systems in response to traumatic events such as natural and man-made disasters (e.g., floods, terrorist attacks, etc. . . ). In this perspective, we propose ACHILLES, an application to models people's movements in a given transport mode through a multiplex network representation based on mobility data. ACHILLES is a web-based application which provides an easy-to-use interface to explore the mobility fluxes and the connectivity of every urban zone in a city, as well as to visualize changes in the transport system resulting from the addition or removal of transport modes, urban zones, and single stops. Notably, our application allows the user to assess the overall resilience of the transport network by identifying its weakest node, i.e. Urban Achilles Heel, with reference to the ancient Greek mythology. To demonstrate the impact of ACHILLES for humanitarian aid we consider its application to a real-world scenario by exploring human mobility in Singapore in response to flood prevention.
Michele Ferretti, Gianni Barlacchi, Luca Pappalardo, Lorenzo Lucchini, Bruno Lepri
DSAA5
2018 The Economic Value of Neighborhoods: Predicting Real Estate Prices from the Urban Environment
abstract
Housing costs have a significant impact on individuals, families, businesses, and governments. Recently, online companies such as Zillow have developed proprietary systems that provide automated estimates of housing prices without the immediate need of professional appraisers. Yet, our understanding of what drives the value of houses is very limited. In this paper, we use multiple sources of data to entangle the economic contribution of the neighborhood's characteristics such as walkability and security perception. We also develop and release a framework able to now-cast housing prices from Open data, without the need for historical transactions. Experiments involving 70,000 houses in 8 Italian cities highlight that the neighborhood's vitality and walkability seem to drive more than 20% of the housing value. Moreover, the use of this information improves the nowcast by 60%. Hence, the use of property's surroundings' characteristics can be an invaluable resource to appraise the economic and social value of houses after neighborhood changes and, potentially, anticipate gentrification.
Marco De Nadai, Bruno Lepri
DSAA2
2018 Social Bridges in Urban Purchase Behavior
abstract
The understanding and modeling of human purchase behavior in city environment can have important implications in the study of urban economy and in the design and organization of cities. In this article, we study human purchase behavior at the community level and argue that people who live in different communities but work at close-by locations could act as “social bridges” between the respective communities and that they are correlated with similarity in community purchase behavior. We provide empirical evidence by studying millions of credit card transaction records for tens of thousands of individuals in a city environment during a period of three months. More specifically, we show that the number of social bridges between communities is a much stronger indicator of similarity in their purchase behavior than traditionally considered factors such as income and sociodemographic variables. Our findings also suggest that such an effect varies across different merchant categories, that the presence of female customers in social bridges is a stronger indicator compared to that of their male counterparts, and that there seems to be a geographical constraint for this effect, all of which may have implications in the studies of urban economy and data-driven urban planning.
Xiaowen Dong 0001, Yoshihiko Suhara, Burçin Bozkaya, Vivek K. Singh 0001, Bruno Lepri, Alex Pentland
ACM Trans. Intell. Syst. Technol.5
2018 A Stigmergy-Based Analysis of City Hotspots to Discover Trends and Anomalies in Urban Transportation Usage
abstract
A key aspect of a sustainable urban transportation system is the effectiveness of transportation policies. To be effective, a policy has to consider a broad range of elements, such as pollution emission, traffic flow, and human mobility. Due to the complexity and variability of these elements in the urban area, to produce effective policies remains a very challenging task. With the introduction of the smart city paradigm, a widely available amount of data can be generated in urban spaces. Such data can be a fundamental source of knowledge to improve policies because they can reflect the sustainability issues underlying the city. In this context, we propose an approach to exploit urban positioning data based on stigmergy, a bio-inspired mechanism providing scalar and temporal aggregation of samples. By employing stigmergy, samples in proximity with each other are aggregated into a functional structure called trail. The trail summarizes relevant dynamics in data and allows matching them, providing a measure of their similarity. Moreover, this mechanism can be specialized to unfold specific dynamics. Specifically, we identify high-density urban areas (i.e. hotspots), analyze their activity over time, and unfold anomalies. Furthermore, by matching activity patterns, a continuous measure of the dissimilarity with respect to the typical activity pattern is provided. This measure can be used by policy makers to evaluate the effect of policies and change them dynamically. As a case study, we analyze taxi trip data gathered in Manhattan from 2013 to 2015.
Antonio L. Alfeo, Mario G. C. A. Cimino, Sara Egidi, Bruno Lepri, Gigliola Vaglini
IEEE Trans. Intell. Transp. Syst.4
2017 Profilio: Psychometric Profiling to Boost Social Media Advertising
abstract
Profilio Company is a startup in its early business development stage that has developed a profiling solution for the field of paid social media advertising. In particular, the solution is designed for the enrichment of Customer Relationship Managements data and for segmentation of customer audiences. Three different Proof of Concepts with different clients have showed that the solution reduces the costs of paid social media advertising in different settings and with different advertising targets, especially starting from large audiences. In this paper we report the details about Profilio's business idea, the development of Profilio's technologies and the results of the Proof of Concepts.
Fabio Celli, Pietro Zani Massani, Bruno Lepri
ACM Multimedia3
2017 What your Facebook Profile Picture Reveals about your Personality
abstract
People spend considerable effort managing the impressions they give others. Social psychologists have shown that people manage these impressions differently depending upon their personality. Facebook and other social media provide a new forum for this fundamental process; hence, understanding people's behaviour on social media could provide interesting insights on their personality. In this paper we investigate automatic personality recognition from Facebook profile pictures. We analyze the effectiveness of four families of visual features and we discuss some human interpretable patterns that explain the personality traits of the individuals. For example, extroverts and agreeable individuals tend to have warm colored pictures and to exhibit many faces in their portraits, mirroring their inclination to socialize; while neurotic ones have a prevalence of pictures of indoor places. Then, we propose a classification approach to automatically recognize personality traits from these visual features. Finally, we compare the performance of our classification approach to the one obtained by human raters and we show that computer-based classifications are significantly more accurate than averaged human-based classifications for Extraversion and Neuroticism.
Cristina Segalin, Fabio Celli, Luca Polonio, Michal Kosinski, David Stillwell, Nicu Sebe, Marco Cristani, Bruno Lepri
ACM Multimedia8
2017 Structural Semantic Models for Automatic Analysis of Urban Areas
Gianni Barlacchi, Alberto Rossi, Bruno Lepri, Alessandro Moschitti
ECML/PKDD (3)3
2017 Anonymous or Not? Understanding the Factors Affecting Personal Mobile Data Disclosure
abstract
The wide adoption of mobile devices and social media platforms have dramatically increased the collection and sharing of personal information. More and more frequently, users are called to make decisions concerning the disclosure of their personal information. In this study, we investigate the factors affecting users’ choices toward the disclosure of their personal data, including not only their demographic and self-reported individual characteristics, but also their social interactions and their mobility patterns inferred from months of mobile phone data activity. We report the findings of a field study conducted with a community of 63 subjects provided with (i) a smart-phone and (ii) a Personal Data Store (PDS) enabling them to control the disclosure of their data. We monitor the sharing behavior of our participants through the PDS and evaluate the contribution of different factors affecting their disclosing choices of location and social interaction data. Our analysis shows that social interaction inferred by mobile phones is an important factor revealing willingness to share, regardless of the data type. In addition, we provide further insights on the individual traits relevant to the prediction of sharing behavior.
Christos Perentis, Michele Vescovi, Chiara Leonardi, Corrado Moiso, Mirco Musolesi, Fabio Pianesi, Bruno Lepri
ACM Trans. Internet Techn.7
2016 Are Safer Looking Neighborhoods More Lively?: A Multimodal Investigation into Urban Life
abstract
Policy makers, urban planners, architects, sociologists, and economists are interested in creating urban areas that are both lively and safe. But are the safety and liveliness of neighborhoods independent characteristics? Or are they just two sides of the same coin? In a world where people avoid unsafe looking places, neighborhoods that look unsafe will be less lively, and will fail to harness the natural surveillance of human activity. But in a world where the preference for safe looking neighborhoods is small, the connection between the perception of safety and liveliness will be either weak or nonexistent. In this paper we explore the connection between the levels of activity and the perception of safety of neighborhoods in two major Italian cities by combining mobile phone data (as a proxy for activity or liveliness) with scores of perceived safety estimated using a Convolutional Neural Network trained on a dataset of Google Street View images scored using a crowdsourced visual perception survey. We find that: (i) safer looking neighborhoods are more active than what is expected from their population density, employee density, and distance to the city centre; and (ii) that the correlation between appearance of safety and activity is positive, strong, and significant, for females and people over 50, but negative for people under 30, suggesting that the behavioral impact of perception depends on the demographic of the population. Finally, we use occlusion techniques to identify the urban features that contribute to the appearance of safety, finding that greenery and street facing windows contribute to a positive appearance of safety (in agreement with Oscar Newman's defensible space theory). These results suggest that urban appearance modulates levels of human activity and, consequently, a neighborhood's rate of natural surveillance.
Marco De Nadai, Radu L. Vieriu, Gloria Zen, Stefan Dragicevic, Nikhil Naik 0003, Michele Caraviello, César A. Hidalgo 0001, Nicu Sebe, Bruno Lepri
ACM Multimedia9
2016 The Death and Life of Great Italian Cities: A Mobile Phone Data Perspective
abstract
The Death and Life of Great American Cities was written in 1961 and is now one of the most influential book in city planning. In it, Jane Jacobs proposed four conditions that promote life in a city. However, these conditions have not been empirically tested until recently. This is mainly because it is hard to collect data about "city life". The city of Seoul recently collected pedestrian activity through surveys at an unprecedented scale, with an effort spanning more than a decade, allowing researchers to conduct the first study successfully testing Jacobs's conditions.
Marco De Nadai, Jacopo Staiano, Roberto Larcher, Nicu Sebe, Daniele Quercia, Bruno Lepri
WWW6
2016 SALSA: A Novel Dataset for Multimodal Group Behavior Analysis
abstract
Studying free-standing conversational groups (FCGs) in unstructured social settings (e.g., cocktail party ) is gratifying due to the wealth of information available at the group (mining social networks) and individual (recognizing native behavioral and personality traits) levels. However, analyzing social scenes involving FCGs is also highly challenging due to the difficulty in extracting behavioral cues such as target locations, their speaking activity and head/body pose due to crowdedness and presence of extreme occlusions. To this end, we propose SALSA, a novel dataset facilitating multimodal and Synergetic sociAL Scene Analysis, and make two main contributions to research on automated social interaction analysis: (1) SALSA records social interactions among 18 participants in a natural, indoor environment for over 60 minutes, under the poster presentation and cocktail party contexts presenting difficulties in the form of low-resolution images, lighting variations, numerous occlusions, reverberations and interfering sound sources; (2) To alleviate these problems we facilitate multimodal analysis by recording the social interplay using four static surveillance cameras and sociometric badges worn by each participant, comprising the microphone, accelerometer, bluetooth and infrared sensors. In addition to raw data, we also provide annotations concerning individuals' personality as well as their position, head, body orientation and F-formation information over the entire event duration. Through extensive experiments with state-of-the-art approaches, we show (a) the limitations of current methods and (b) how the recorded multiple cues synergetically aid automatic analysis of social interactions. SALSA is available at http://tev.fbk.eu/salsa.
Xavier Alameda-Pineda, Jacopo Staiano, Subramanian Ramanathan, Ligia Maria Batrinca, Elisa Ricci 0001, Bruno Lepri, Oswald Lanz, Nicu Sebe
IEEE Trans. Pattern Anal. Mach. Intell.6
2016 Multimodal Personality Recognition in Collaborative Goal-Oriented Tasks
abstract
Incorporating research on personality recognition into computers, both from a cognitive as well as an engineering perspective, would facilitate the interactions between humans and machines. Previous attempts on personality recognition have focused on a variety of different corpora (ranging from text to audiovisual data), scenarios (interviews, meetings), channels of communication (audio, video, text), and different subsets of personality traits (out of the five ones from the Big Five Model). Our study uses simple acoustic and visual nonverbal features extracted from multimodal data, which have been recorded in previously uninvestigated scenarios, and consider all five personality traits and not just a subset. First, we look at the human-machine interaction scenario, where we introduce the display of different “collaboration levels.” Second, we look at the contribution of the human-human interaction (HHI) scenario on the emergence of personality traits. Investigating the HHI scenario creates a stronger basis for future human-agents interactions. Our goal is to study, from a computational approach, the emergence degree of the five personality traits in these two scenarios. The results demonstrate the relevance of each of the two scenarios when it comes to the degree of emergence of certain traits and the feasibility to automatically recognize personality under different conditions.
Ligia Maria Batrinca, Nadia Mana, Bruno Lepri, Nicu Sebe, Fabio Pianesi
IEEE Trans. Multim.3
2016 The role of personality in shaping social networks and mediating behavioral change
Bruno Lepri, Jacopo Staiano, Erez Shmueli, Fabio Pianesi, Alex Pentland
User Model. User Adapt. Interact.1
2015 If You Are Happy and You Know It, Say "I'm Here": Investigating Parents' Location-Sharing Preferences
Paolo Massa, Chiara Leonardi, Bruno Lepri, Fabio Pianesi, Massimo Zancanaro
INTERACT (3)3
2015 Predicting and Understanding Urban Perception with Convolutional Neural Networks
abstract
Cities' visual appearance plays a central role in shaping human perception and response to the surrounding urban environment. For example, the visual qualities of urban spaces affect the psychological states of their inhabitants and can induce negative social outcomes. Hence, it becomes critically important to understand people's perceptions and evaluations of urban spaces. Previous works have demonstrated that algorithms can be used to predict high level attributes of urban scenes (e.g. safety, attractiveness, uniqueness), accurately emulating human perception. In this paper we propose a novel approach for predicting the perceived safety of a scene from Google Street View Images. Opposite to previous works, we formulate the problem of learning to predict high level judgments as a ranking task and we employ a Convolutional Neural Network (CNN), significantly improving the accuracy of predictions over previous methods. Interestingly, the proposed CNN architecture relies on a novel pooling layer, which permits to automatically discover the most important areas of the images for predicting the concept of perceived safety. An extensive experimental evaluation, conducted on the publicly available Place Pulse dataset, demonstrates the advantages of the proposed approach over state-of-the-art methods.
Lorenzo Porzi, Samuel Rota Bulò, Bruno Lepri, Elisa Ricci 0001
ACM Multimedia3
2014 Money walks: a human-centric study on the economics of personal mobile data
abstract
In the context of a myriad of mobile apps which collect personally identifiable information (PII) and a prospective market place of personal data, we investigate a user-centric monetary valuation of mobile PII. During a 6-week long user study in a living lab deployment with 60 participants, we collected their daily valuations of 4 categories of mobile PII (communication, e.g. phonecalls made/received, applications, e.g. time spent on different apps, location and media, e.g. photos taken) at three levels of complexity (individual data points, aggregated statistics and processed, i.e. meaningful interpretations of the data). In order to obtain honest valuations, we employ a reverse second price auction mechanism. Our findings show that the most sensitive and valued category of personal information is location. We report statistically significant associations between actual mobile usage, personal dispositions, and bidding behavior. Finally, we outline key implications for the design of mobile services and future markets of personal data.
Jacopo Staiano, Nuria Oliver, Bruno Lepri, Rodrigo de Oliveira, Michele Caraviello, Nicu Sebe
UbiComp3
2014 Once Upon a Crime: Towards Crime Prediction from Demographics and Mobile Data
abstract
In this paper, we present a novel approach to predict crime in a geographic space from multiple data sources, in particular mobile phone and demographic data. The main contribution of the proposed approach lies in using aggregated and anonymized human behavioral data derived from mobile network activity to tackle the crime prediction problem. While previous research efforts have used either background historical knowledge or offenders' profiling, our findings support the hypothesis that aggregated human behavioral data captured from the mobile network infrastructure, in combination with basic demographic information, can be used to predict crime. In our experimental results with real crime data from London we obtain an accuracy of almost 70% when predicting whether a specific area in the city will be a crime hotspot or not. Moreover, we provide a discussion of the implications of our findings for data-driven crime analysis.
Andrey Bogomolov, Bruno Lepri, Jacopo Staiano, Nuria Oliver, Fabio Pianesi, Alex Pentland
ICMI2
2014 Daily Stress Recognition from Mobile Phone Data, Weather Conditions and Individual Traits
abstract
Research has proven that stress reduces quality of life and causes many diseases. For this reason, several researchers devised stress detection systems based on physiological parameters. However, these systems require that obtrusive sensors are continuously carried by the user. In our paper, we propose an alternative approach providing evidence that daily stress can be reliably recognized based on behavioral metrics, derived from the user's mobile phone activity and from additional indicators, such as the weather conditions (data pertaining to transitory properties of the environment) and the personality traits (data concerning permanent dispositions of individuals). Our multifactorial statistical model, which is person-independent, obtains the accuracy score of 72.28% for a 2-class daily stress recognition problem. The model is efficient to implement for most of multimedia applications due to highly reduced low-dimensional feature space (32d). Moreover, we identify and discuss the indicators which have strong predictive power.
Andrey Bogomolov, Bruno Lepri, Michela Ferron, Fabio Pianesi, Alex Pentland
ACM Multimedia2
2014 Automatic Personality and Interaction Style Recognition from Facebook Profile Pictures
abstract
In this paper, we address the issue of personality and interaction style recognition from profile pictures in Facebook. We recruited volunteers among Facebook users and collected a dataset of profile pictures, labeled with gold standard self-assessed personality and interaction style labels. Then, we exploited a bag-of-visual-words technique to extract features from pictures. Finally, different machine learning approaches were used to test the effectiveness of these features in predicting personality and interaction style traits. Our good results show that this task is very promising, because profile pictures convey a lot of information about a user and are directly connected to impression formation and identity management.
Fabio Celli, Elia Bruni, Bruno Lepri
ACM Multimedia3
2014 The Workshop on Computational Personality Recognition 2014
abstract
The Workshop on Computational Personality Recognition aims to define the state-of-the-art in the field and to provide tools for future standard evaluations in personality recognition tasks. In the WCPR14 we released two different datasets: one of Youtube Vlogs and one of Mobile Phone interactions. We structured the workshop in two tracks: an open shared task, where participants can do any kind of experiment, and a competition. We also distinguished two tasks: A) personality recognition from multimedia data, and B) personality recognition from text only. In this paper we discuss the results of the workshop.
Fabio Celli, Bruno Lepri, Joan-Isaac Biel, Daniel Gatica-Perez, Giuseppe Riccardi, Fabio Pianesi
ACM Multimedia2
2014 Sensing, Understanding, and Shaping Social Behavior
abstract
The ability to understand social systems through the aid of computational tools is central to the emerging field of computational social systems. Such understanding can answer epistemological questions on human behavior in a data-driven manner, and provide prescriptive guidelines for persuading humans to undertake certain actions in real-world social scenarios. The growing number of works in this subfield has the potential to impact multiple walks of human life including health, wellness, productivity, mobility, transportation, education, shopping, and sustenance. The contribution of this paper is twofold. First, we provide a functional survey of recent advances in sensing, understanding, and shaping human behavior, focusing on real-world behavior of users as measured using passive sensors. Second, we present a case study on how trust, which is an important building block of computational social systems, can be quantified, sensed, and applied to shape human behavior. Our findings suggest that:1) trust can be operationalized and predicted via computational methods (passive sensing and network analysis) and 2) trust has a significant impact on social persuasion; in fact, it was found to be significantly more effective than the closeness of ties in determining the amount of behavior change.
Erez Shmueli, Vivek K. Singh 0001, Bruno Lepri, Alex Pentland
IEEE Trans. Comput. Soc. Syst.3
2013 Inferring social activities with mobile sensor networks
abstract
While our daily activities usually involve interactions with others, the current methods on activity recognition do not often exploit the relationship between social interactions and human activity. This paper addresses the problem of interpreting social activity from human interactions captured by mobile sensing networks. Our first goal is to discover different social activities such as chatting with friends from interaction logs and then characterize them by the set of people involved, and the time and location of the occurring event. Our second goal is to perform automatic labeling of the discovered activities using predefined semantic labels such as coffee breaks, weekly meetings, or random discussions. Our analysis was conducted on a real-life interaction network sensed with Bluetooth and infrared sensors of about fifty subjects who carried sociometric badges over 6 weeks. We show that the proposed system reliably recognized coffee breaks with 99% accuracy, while weekly meetings were recognized with 88% accuracy.
Trinh Minh Tri Do, Kyriaki Kalimeri, Bruno Lepri, Fabio Pianesi, Daniel Gatica-Perez
ICMI3
2013 Going beyond traits: multimodal classification of personality states in the wild
abstract
Recent studies in social and personality psychology introduced the notion of personality states conceived as concrete behaviors that can be described as having the same contents as traits. Our paper is a first step towards addressing automatically this new perspective. In particular, we will focus on the classification of excerpts of social behavior into personality states corresponding to the Big Five traits, rather than focusing on the more traditional goal of using those behaviors to directly infer about the personality traits of the person producing them. The multimodal behavioral cues we exploit were obtained by means of the Sociometric Badges worn by people working at a research institution for a period of six weeks. We investigate the effectiveness of cues concerning acted social behaviors as well as of other situational characteristics for the sake of personality state classification. The encouraging results show that our classifiers always, and sometimes greatly, improve the performances of a random baseline classifier (from 1.5 to 1.8 better than chance). At a general level, we believe that these results support the proposed shift from the classification of personality traits to the classification of personality states.
Kyriaki Kalimeri, Bruno Lepri, Fabio Pianesi
ICMI2
2013 Modeling Functional Roles Dynamics in Small Group Interactions
abstract
The paper addresses the automatic recognition of social and task-oriented functional roles in small-group meetings, focusing on several properties: a) the importance of non-linguistic behaviors, b) the relative time-consistency of the social roles played by a given person during the course of a meeting, and c) the interplays and mutual constraints among the roles enacted by the different participants in a social encounter. In particular, this paper proposes that the Influence Model framework can address these properties of functional roles, and compares the performance obtained by this framework to the performances of models that consider only property (a) (SVM), and to those that address both (a) and (b) (HMM). The results obtained confirm our expectations: the classification of social functional roles improves if models account for temporal dependencies among the roles played by the same subject, for the time properties of the roles played by each individual, and for the mutual constraints among the roles of different group members. The two versions of the Influence Model (IM and newIM), which encode all three properties together, outperform both the SVM and the HMM on most of the figures of merit used. Of particular interest is the capability of the Influence Model to obtain good or very good results on the less-populated classes-Orienteer and Seeker for the task area, and Attacker and Supporter for the socio-emotional area.
Wen Dong 0001, Bruno Lepri, Fabio Pianesi, Alex Pentland
IEEE Trans. Multim.2
2012 Friends don't lie: inferring personality traits from social network structure
abstract
In this work, we investigate the relationships between social network structure and personality; we assess the performances of different subsets of structural network features, and in particular those concerned with ego-networks, in predicting the Big-5 personality traits. In addition to traditional survey-based data, this work focuses on social networks derived from real-life data gathered through smartphones. Besides showing that the latter are superior to the former for the task at hand, our results provide a fine-grained analysis of the contribution the various feature sets are able to provide to personality classification, along with an assessment of the relative merits of the various networks exploited.
Jacopo Staiano, Bruno Lepri, Nadav Aharony, Fabio Pianesi, Nicu Sebe, Alex Pentland
UbiComp2
2012 Multimodal recognition of personality traits in human-computer collaborative tasks
abstract
The user's personality in Human-Computer Interaction (HCI) plays an important role for the overall success of the interaction. The present study focuses on automatically recognizing the Big Five personality traits from 2-5 min long videos, in which the computer interacts using different levels of collaboration, in order to elicit the manifestation of these personality traits. Emotional Stability and Extraversion are the easiest traits to automatically detect under the different collaborative settings: all the settings for Emotional Stability and intermediate and fully-non collaborative settings for Extraversion. Interestingly, Agreeableness and Conscientiousness can be detected only under a moderately non-collaborative setting. Finally, our task does not seem to activate the full range of dispositions for Creativity.
Ligia Maria Batrinca, Bruno Lepri, Nadia Mana, Fabio Pianesi
ICMI2
2012 Modeling dominance effects on nonverbal behaviors using granger causality
abstract
In this paper we modeled the effects that dominant people might induce on the nonverbal behavior (speech energy and body motion) of the other meeting participants using Granger causality technique. Our initial hypothesis that more dominant people have generalized higher influence was not validated when using the DOME-AMI corpus as data source. However, from the correlational analysis some interesting patterns emerged: contradicting our initial hypothesis dominant individuals are not accounting for the majority of the causal flow in a social interaction. Moreover, they seem to have more intense causal effects as their causal density was significantly higher. Finally dominant individuals tend to respond to the causal effects more often with complementarity than with mimicry.
Kyriaki Kalimeri, Bruno Lepri, Oya Aran, Dinesh Babu Jayagopi, Daniel Gatica-Perez, Fabio Pianesi
ICMI2
2012 Do Linguistic Style and Readability of Scientific Abstracts Affect their Virality?
Marco Guerini, Alberto Pepe, Bruno Lepri
ICWSM3
2012 Connecting Meeting Behavior with Extraversion - A Systematic Study
abstract
This work investigates the suitability of medium-grained meeting behaviors, namely, speaking time and social attention, for automatic classification of the Extraversion personality trait. Experimental results confirm that these behaviors are indeed effective for the automatic detection of Extraversion. The main findings of our study are that: 1) Speaking time and (some forms of) social gaze are effective indicators of Extraversion, 2) classification accuracy is affected by the amount of time for which meeting behavior is observed, 3) independently considering only the attention received by the target from peers is insufficient, and 4) distribution of social attention of peers plays a crucial role.
Bruno Lepri, Subramanian Ramanathan, Kyriaki Kalimeri, Jacopo Staiano, Fabio Pianesi, Nicu Sebe
IEEE Trans. Affect. Comput.1
2011 Please, tell me about yourself: automatic personality assessment using short self-presentations
abstract
Personality plays an important role in the way people manage the images they convey in self-presentations and employment interviews, trying to affect the other"s first impressions and increase effectiveness. This paper addresses the automatically detection of the Big Five personality traits from short (30-120 seconds) self-presentations, by investigating the effectiveness of 29 simple acoustic and visual non-verbal features. Our results show that Conscientiousness and Emotional Stability/Neuroticism are the best recognizable traits. The lower accuracy levels for Extraversion and Agreeableness are explained through the interaction between situational characteristics and the differential activation of the behavioral dispositions underlying those traits.
Ligia Maria Batrinca, Nadia Mana, Bruno Lepri, Fabio Pianesi, Nicu Sebe
ICMI3
2011 Automatic modeling of personality states in small group interactions
abstract
In this paper, we target the automatic recognition of personality states in a meeting scenario employing visual and acoustic features. The social psychology literature has coined the name personality state to refer to a specific behavioral episode wherein a person behaves as more or less introvert/extrovert, neurotic or open to experience, etc. Personality traits can then be reconstructed as density distributions over personality states. Different machine learning approaches were used to test the effectiveness of the selected features in modeling the dynamics of personality states.
Jacopo Staiano, Bruno Lepri, Subramanian Ramanathan, Nicu Sebe, Fabio Pianesi
ACM Multimedia2
2011 Modeling the co-evolution of behaviors and social relationships using mobile phone data
abstract
The co-evolution of social relationships and individual behavior in time and space has important implications, but is poorly understood because of the difficulty closely tracking the everyday life of a complete community. We offer evidence that relationships and behavior co-evolve in a student dormitory, based on monthly surveys and location tracking through resident cellular phones over a period of nine months. We demonstrate that a Markov jump process could capture the co-evolution in terms of the rates at which residents visit places and friends.
Wen Dong 0001, Bruno Lepri, Alex Pentland
MUM2
2010 What is happening now? Detection of activities of daily living from simple visual features
Bruno Lepri, Nadia Mana, Alessandro Cappelletti, Fabio Pianesi, Massimo Zancanaro
Pers. Ubiquitous Comput.1
2009 Automatic prediction of individual performance from "thin slices" of social behavior
abstract
This paper targets the automatic detection of individual performances in group tasks by means of short sequences, "thin slices", of nonverbal behavior. We designed our task as a classification one. We also investigated the relevance of social context in our task and the effectiveness of our feature selection.
Bruno Lepri, Nadia Mana, Alessandro Cappelletti, Fabio Pianesi
ACM Multimedia1
2009 Modeling the Personality of Participants During Group Interactions
Bruno Lepri, Nadia Mana, Alessandro Cappelletti, Fabio Pianesi, Massimo Zancanaro
UMAP1
2008 Multimodal recognition of personality traits in social interactions
abstract
This paper targets the automatic detection of personality traits in a meeting environment by means of audio and visual features; information about the relational context is captured by means of acoustic features designed to that purpose. Two personality traits are considered: Extraversion (from the Big Five) and the Locus of Control. The classification task is applied to thin slices of behaviour, in the form of 1-minute sequences. SVM were used to test the performances of several training and testing instance setups, including a restricted set of audio features obtained through feature selection. The outcomes improve considerably over existing results, provide evidence about the feasibility of the multimodal analysis of personality, the role of social context, and pave the way to further studies addressing different features setups and/or targeting different personality traits.
Fabio Pianesi, Nadia Mana, Alessandro Cappelletti, Bruno Lepri, Massimo Zancanaro
ICMI4
2008 Multimodal support to group dynamics
Fabio Pianesi, Massimo Zancanaro, Elena Not, Chiara Leonardi, Vera Falcon, Bruno Lepri
Pers. Ubiquitous Comput.6
2007 Using the influence model to recognize functional roles in meetings
abstract
In this paper, an influence model is used to recognize functional roles played during meetings. Previous works on the same corpus demonstrated a high recognition accuracy using SVMs with RBF kernels. In this paper, we discuss the problems of that approach, mainly over-fitting, the curse of dimensionality and the inability to generalize to different group configurations. We present results obtained with an influence modeling method that avoid these problems and ensures both greater robustness and generalization capability.
Wen Dong 0001, Bruno Lepri, Alessandro Cappelletti, Alex Pentland, Fabio Pianesi, Massimo Zancanaro
ICMI2
2006 Automatic detection of group functional roles in face to face interactions
abstract
In this paper, we discuss a machine learning approach to automatically detect functional roles played by participants in a face to face interaction. We shortly introduce the coding scheme we used to classify the roles of the group members and the corpus we collected to assess the coding scheme reliability as well as to train statistical systems for automatic recognition of roles. We then discuss a machine learning approach based on multi-class SVM to automatically detect such roles by employing simple features of the visual and acoustical scene. The effectiveness of the classification is better than the chosen baselines and although the results are not yet good enough for a real application, they demonstrate the feasibility of the task of detecting group functional roles in face to face interactions.
Massimo Zancanaro, Bruno Lepri, Fabio Pianesi
ICMI2