Antonio Origlia

dblp:68/3553 · DBLP profile ↗
← Back
31ranked-venue papers
12as first author
13since 2021 · last 2026
0000-0002-8635-1623ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 6 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 13 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Advanced Visual Interfaces for Cultural Heritage (AVI-CH) 2026
abstract
Cultural Heritage (CH) is a challenging domain of application for novel Information and Communication Technologies (ICT), where visualization plays a major role in enhancing the visitors’ experience, whether onsite or online. Technology-supported natural human-computer interaction is a key factor in enabling access to CH assets. Following the successful workshops at AVI 2016, AVI 2018, AVI 2020, AVI 2022, and AVI 2024, the goal of the AVI-CH 2026 workshop is to bring together researchers and practitioners interested in presenting and discussing the potential of state-of-the-art advanced visual interfaces in enhancing our daily cultural heritage experience. This year, the workshop theme is Hybrid Reasoning Methods behind Interactive Cultural Heritage Applications. Seven papers were accepted, covering topics ranging from augmented and virtual reality for heritage sites, knowledge graph exploration, conversational AI architectures, digital preservation of multimedia art, computer vision for historical audio recordings, cross-device visitor behaviour in virtual museum tours, and LLM-driven adaptive content in VR museums.
Antonio Origlia, Julia Sheidin, Tsvi Kuflik, Niccolò Pretto
AVI1
2025 A Bayesian framework for learning proactive robot behaviour in assistive tasks
abstract
Abstract Socially assistive robots represent a promising tool in assistive contexts for improving people’s quality of life and well-being through social, emotional, cognitive, and physical support. However, the effectiveness of interactions heavily relies on the robots’ ability to adapt to the needs of the assisted individuals and to offer support proactively, before it is explicitly requested. Previous work has primarily focused on defining the actions the robot should perform, rather than considering when to act and how confident it should be in a given situation. To address this gap, this paper introduces a new data-driven framework that involves a learning pipeline, consisting of two phases, with the ultimate goal of training an algorithm based on Influence Diagrams. The proposed assistance scenario involves a sequential memory game, where the robot autonomously learns what assistance to provide when to intervene, and with what confidence to take control. The results from a user study showed that the proactive behaviour of the robot had a positive impact on the users’ game performance. Users obtained higher scores, made fewer mistakes, and requested less assistance from the robot. The study also highlighted the robot’s ability to provide assistance tailored to users’ specific needs and anticipate their requests.
Antonio Andriella, Ilenia Cucciniello, Antonio Origlia, Silvia Rossi 0002
User Model. User Adapt. Interact.3
2024 Advanced Visual Interfaces and Interactions in Cultural Heritage (AVICH 2024)
abstract
Cultural Heritage (CH) is a challenging domain of application for novel Information and Communication Technologies (ICT), where visualization plays a major role in enhancing the visitors’ experience, whether onsite or online. Technology-supported natural human-computer interaction is a key factor in enabling access, both on-site and online, to CH assets. Recent advances in ICT enhance visitors’ access to online collections as well as their CH experience onsite, bringing even wider audiences than those who visit the physical museums. The range of visualization devices – from tiny smartwatch screens, through large situated public displays, to the latest generation of immersive Head-Mounted Displays – together with the increasing availability of real-time 3D rendering technologies for online and mobile devices and, recently, Internet of Things (IoT) approaches, Social Robotics, requires exploring how they can be applied successfully in CH. Following the successful workshops at AVI 2016, AVI 2018, AVI 2020, AVI 2022 and AVI 2024 and a large number of recent events and projects focusing on CH, the goal of the workshop is to bring together researchers and practitioners interested in presenting and discussing the potential of state of the art advanced visual interfaces in enhancing our daily cultural heritage experience. This year the workshop theme will be: AVI for accessibility in CH.
Tsvi Kuflik, Berardina De Carolis, Cristina Gena, Antonio Origlia, Julia Sheidin
AVI4
2024 On the Use of Plausible Arguments in Explainable Conversational AI
Martina Di Bratto, Maria Di Maro, Antonio Origlia
INTERSPEECH3
2024 A Linguistically Motivated Approach to Hybrid Conversational AI with the FANTASIA Plugin
abstract
This paper presents an installation showcasing the basic concepts of the FANTASIA tool and the Interaction Model built around it. The tool itself can be used in a variety of ways, through the modular set of components it provides to extend the capabilities of a powerful engine for Real Time Interactive 3D applications, the Unreal Engine. The Interaction Model represents how FANTASIA can be used to build linguistically motivated models for dialogue management, going beyond the use of statistical learning only. Specifically, an explicit separation between language modelling capabilities and decision making will be described for movie recommendations.
Antonio Origlia, Maria Di Maro
IVA1
2024 Though this be hesitant, yet there is method in 't: Effects of disfluency patterns in neural speech synthesis for cultural heritage presentations
abstract
This study presents the results of two perception experiments aimed at evaluating the effect that specific patterns of disfluencies have on people listening to synthetic speech. We consider the particular case of Cultural Heritage presentations and propose a linguistic model to support the positioning of disfluencies throughout the utterances in the Italian language. A state-of-the-art speech synthesizer, based on Deep Neural Networks, is used to prepare a set of experimental stimuli and two different experiments are presented to provide both subjective evaluations and behavioural assessments from human subjects. Results show that synthetic utterances including disfluencies, predicted by a linguistic model, are identified as more natural and that the presence of disfluencies benefits the listeners’ recall of the provided information.
Loredana Schettino, Antonio Origlia, Francesco Cutugno
Comput. Speech Lang.2
2024 Exploring emergent syllables in end-to-end automatic speech recognizers through model explainability technique
abstract
Abstract Automatic speech recognition systems based on end-to-end models (E2E-ASRs) can achieve comparable performance to conventional ASR systems while reproducing all their essential parts automatically, from speech units to the language model. However, they hide the underlying perceptual processes modelled, if any, and they have lower adaptability to multiple application contexts, and, furthermore, they require powerful hardware and an extensive amount of training data. Model-explainability techniques can explore the internal dynamics of these ASR systems and possibly understand and explain the processes conducting to their decisions and outputs. Understanding these processes can help enhance ASR performance and reduce the required training data and hardware significantly. In this paper, we probe the internal dynamics of three E2E-ASRs pre-trained for English by building an acoustic-syllable boundary detector for Italian and Spanish based on the E2E-ASRs’ internal encoding layer outputs. We demonstrate that the shallower E2E-ASR layers spontaneously form a rhythmic component correlated with prominent syllables, central in human speech processing. This finding highlights a parallel between the analysed E2E-ASRs and human speech recognition. Our results contribute to the body of knowledge by providing a human-explainable insight into behaviours encoded in popular E2E-ASR systems.
Vincenzo Norman Vitale, Francesco Cutugno, Antonio Origlia, Gianpaolo Coro
Neural Comput. Appl.3
2024 Linguistics-based dialogue simulations to evaluate argumentative conversational recommender systems
abstract
Abstract Conversational recommender systems aim at recommending the most relevant information for users based on textual or spoken dialogues, through which users can communicate their preferences to the system more efficiently. Argumentative conversational recommender systems represent a kind of deliberation dialogue in which participants share their specific beliefs in the respective representations of the common ground, to act towards a common goal. The goal of such systems is to present appropriate supporting arguments to their recommendations to show the interlocutor that a specific item corresponds to their manifested interests. Here, we present a cross-disciplinary argumentation-based conversational recommender model based on cognitive pragmatics. We also present a dialogue simulator to investigate the quality of the theoretical background. We produced a set of synthetic dialogues based on a computational model implementing the linguistic theory and we collected human evaluations about the plausibility and efficiency of these dialogues. Our results show that the synthetic dialogues obtain high scores concerning their naturalness and the selection of the supporting arguments.
Martina Di Bratto, Antonio Origlia, Maria Di Maro, Sabrina Mennella
User Model. User Adapt. Interact.2
2023 Increasing visitors attention with introductory portal technology to complex cultural sites
abstract
We investigate the effectiveness of visually impacting interfaces which were located next to the entrance of the San Martino Charterhouse in Naples (Italy), by using high quality 3D reconstructions annotated with semantic information. Semantic data were used to develop a gesture based orienting interface, for the generation of camera movements and pointing gestures for a 3D avatar. Using the Visitor Employed Photography protocol, we collected data about visitors noticing details and discovered that those exposed to the informative systems were able to detect more details than visitors who did not experience it.
Antonio Origlia, Maria Laura Chiacchio, Marco Grazioso, Francesco Cutugno
Int. J. Hum. Comput. Stud.1
2022 AVI-CH 2022: Workshop on Advanced Visual Interfaces and Interactions in Cultural Heritage
abstract
AVI-CH is the 14th workshop in the series of PATCH workshops, since 2007 and the 4th in a row at AVI. It is the meeting place for researchers and practitioners focusing on the application of advanced information and communication technology (ICT) in cultural heritage with a specific focus on user interfaces, visualization and interaction. This year, eight papers were submitted by researchers from Greece, Italy and Israel. All were accepted.
Angeliki Antoniou, Berardina De Carolis, Tsvi Kuflik, Antonio Origlia, George E. Raptis, Cristina Gena
AVI4
2022 A Multi-source Graph Representation of the Movie Domain for Recommendation Dialogues Analysis
abstract
In dialogue analysis, characterising named entities in the domain of interest is relevant in order to understand how people are making use of them for argumentation purposes. The movie recommendation domain is a frequently considered case study for many applications and by linguistic studies and, since many different resources have been collected throughout the years to describe it, a single database combining all these data sources is a valuable asset for cross-disciplinary investigations. We propose an integrated graph-based structure of multiple resources, enriched with the results of the application of graph analytics approaches to provide an encompassing view of the domain and of the way people talk about it during the recommendation task. While we cannot distribute the final resource because of licensing issues, we share the code to assemble and process it once the reference data have been obtained from the original sources.
Antonio Origlia, Martina Di Bratto, Maria Di Maro, Sabrina Mennella
LREC1
2022 Developing Embodied Conversational Agents in the Unreal Engine: The FANTASIA Plugin
abstract
The fast-evolving industry of games pushes the technology behind audio, graphics and controllers to evolve at a very high speed. The front-end of interactive experiences such as games sets standards that must be met also by research teams working on Artificial Intelligence. It is, therefore, important, to devise strategies to let researchers concentrate on interaction models while being up to date with interface design standards. FANTASIA is designed to extend the functionalities offered by a high-profile industrial game engine, the Unreal Engine, to obtain an advanced development environment to develop Embodied Conversational Agents. The FANTASIA plugin is open source and is freely available on GitHub https://github.com/antori82/FANTASIA together with tutorial series on a dedicated Youtube channel.
Antonio Origlia, Martina Di Bratto, Maria Di Maro, Sabrina Mennella
ACM Multimedia1
2022 On the Impact of Location-related Terms in Neural Embeddings for Content Similarity Measures in Cultural Heritage Recommender Systems
Antonio Origlia, Sergio Di Martino
W2GIS1
2020 AVI2CH 2020: Workshop on Advanced Visual Interfaces and Interactions in Cultural Heritage
abstract
AVI2CH is a meeting place for researchers and practitioners focusing on the application of advanced information and communication technology in cultural heritage (CH) with a specific focus on user interfaces, visualization and interaction. It builds on a series of PATCH workshops, since 2007 including three at AVI and also a series of European workshops on cultural informatics. Eleven papers range from novel interfaces in museums to wider community engagement; all share a common mission to ensure that the latest digital technology helps preserve the past in ways that enrich the lives of current and future generations
Angeliki Antoniou, Berardina De Carolis, George E. Raptis, Cristina Gena, Tsvi Kuflik, Alan J. Dix, Antonio Origlia, Giorgos Lepouras
AVI7
2020 Smart Parking: Using a Crowd of Taxis to Sense On-Street Parking Space Availability
abstract
Monitoring the occupancy of on-street parking spaces on a city-wide scale is still an open issue. Past research demonstrated the viability of parking crowd-sensing by means of the standard on-board sensors of probe vehicles, foreseeing the use of high-mileage vehicles, like taxis. Nevertheless, the achievable spatio-temporal sensing coverage has never been deeply investigated. In this paper, we investigate the suitability of taxi fleets of different sizes to crowd-sense on-street parking availability. We considered 579 road segments in San Francisco (USA), covered both by sensors of the SFpark project and by the GPS traces of 536 taxis. For each of these segments, we computed the taxi transit frequencies, representing the achievable coverage by vehicles equipped with sensors detecting empty parking spots. By combining these frequencies with parking occupancy data coming from SFpark, we estimated the potential quality of crowd-sensed on-street parking information for different fleet sizes. Moreover, we investigated the impact of different misdetection amounts, and Kalman filters to handle them. The results show that a total of 300 taxis can crowd-sense on-street parking availability with an error of up to ±1 stall in 86% of the cases. Moreover, the quality of the sensors is as important as the fleet size (300 taxis with 10% probability of misreadings provide availability information comparable to 486 taxis with 16% probability), while the use of Kalman filters did not lead to statistically significant improvements. In conclusion, the traffic management authorities should consider parking crowd-sensing via probe vehicles as a promising alternative to the expensive deployment of the static parking sensors.
Fabian Bock, Sergio Di Martino, Antonio Origlia
IEEE Trans. Intell. Transp. Syst.3
2019 FANTASIA: a framework for advanced natural tools and applications in social, interactive approaches
Antonio Origlia, Francesco Cutugno, Antonio Rodà, Piero Cosi, Claudio Zmarich
Multim. Tools Appl.1
2018 AVI-CH 2018: Advanced Visual Interfaces for Cultural Heritage
abstract
Cultural Heritage (CH) is a challenging domain of application for novel Information and Communication Technologies (ICT), where visualization plays a major role in enhancing visitors' experience, either onsite or online. Technology-supported natural human-computer interaction is a key factor in enabling access to CH assets. Advances in ICT ease visitors to access collections online and better experience CH onsite. The range of visualization devices - from tiny smart watch screens and wall-size large situated public displays to the latest generation of immersive head-mounted displays - together with the increasing availability of real-time 3D rendering technologies for online and mobile devices and, recently, Internet of Things (IoT) approaches, require exploring how they can be applied successfully in CH. Following the successful workshop at AVI 2016 and the large numbers of recent events and projects focusing on CH and, considering that 2018 has been declared the European Year of Cultural Heritage, the goal of the workshop is to bring together researchers and practitioners interested in presenting and discussing the potential use of state-of-the-art advanced visual interfaces in enhancing our daily CH experience.
Berardina De Carolis, Cristina Gena, Tsvi Kuflik, Antonio Origlia, George E. Raptis
AVI4
2018 From Linguistic Linked Open Data to Multimodal Natural Interaction: A Case Study
abstract
We present here the conversion of Linguistic Linked Open Data into Semantic Maps to be used to produce contents in a set of technological applications for Cultural Heritage. The paper describes the architectural data collection and annotation procedure adopted in the Cultural Heritage Orienting Multimodal Experiences (CHROME) project (PRIN 2015 funded by Italian University and Research Ministry). Such data will be used in Multimodal Dialogue Systems to obtain precise information about Architectural Heritage, by means of pointing gestures or verbal requests. In particular, we design conversational agents accessing fine-detailed semantic data linked to available 3D models of historical buildings. The starting point of our scientific approach is the Getty Vocabulary on Art & Architecture Thesaurus, integrated with the Getty Thesaurus of Geographic Names (TGN) and the Union List of Artist Names (ULAN). These data are related to 3D mesh of the considered buildings in order to associate abstract concepts to architectural elements. In the field of 3D architectural investigation, a significant amount of research has been conducted to allow domain experts to represent semantic data while keeping spatial references. We will discuss how this will make it possible to support multimodal user interaction and generate Cultural Heritage presentations.
Marco Grazioso, Valeria Cera, Maria Di Maro, Antonio Origlia, Francesco Cutugno
IV4
2017 An adaptive neuro-fuzzy inference system for the qualitative study of perceptual prominence in linguistics
abstract
This paper explores the applications of fuzzy logic inference systems as an instrument to perform linguistic analysis in the domain of prosodic prominence. Understanding how acoustic features interact to make a linguistic unit be perceived as more relevant than the surrounding ones is generally needed to study the cognitive processes needed for speech understanding. It also has technological applications in the field of speech recognition and synthesis. We present a first experiment to show how fuzzy inference systems, being characterised by their capability to provide detailed insight about the models obtained through supervised learning can help investigate the complex relationships among acoustic features linked to prominence perception.
Autilia Vitiello, Giovanni Acampora, Francesco Cutugno, Petra Wagner, Antonio Origlia
FUZZ-IEEE5
2016 Combining Energy and Cross-Entropy Analysis for Nuclear Segments Detection
Antonio Origlia, Francesco Cutugno
INTERSPEECH1
2015 An analysis of perceptual cues in robot group selection tasks
abstract
The aim of the proposed investigation is to provide the users with the capability of creating robot teams “on- the-fly” using grouping strategies expressed through speech. Our working hypothesis is that people are inclined to assemble objects into macro-entities, or groups, according to perceptual principles. We observed the real linguistic utterances used by individuals in a testing environment, showing that the type of robots and their mutual arrangements can affect both the choice of elements to form a team, and the way such choice is made. Moreover, we provide an initial insight for the capabilities needed by a robot for reasoning about its membership in a team.
Alessandra Rossi 0001, Mariacarla Staffa, Antonio Origlia, Silvia Rossi 0002
RO-MAN3
2014 Continuous emotion recognition with phonetic syllables
Antonio Origlia, Francesco Cutugno, Vincenzo Galatà
Speech Commun.1
2013 CoWME: a general framework to evaluate cognitive workload during multimodal interaction
abstract
Evaluating human machine interaction in the case of multimodal systems is often a difficult task involving the monitoring of multiple sources, data fusion and results interpretation. While subtasks are highly dependent on the specific goal of the application and on the available interaction modalities, it is possible to formalize this workflow into a standard process and to consider a generic measure to estimate the ease of use of a specific application. In this work, we present CoWME, a modular software architecture describing multimodal human machine interaction evaluation, from data collection to final evaluation, in a formal way, in terms of cognitive workload. Communication protocols between modules are described in XML while data fusion is delegated to a configurable rule engine. An interface module is introduced between the monitoring modules and the rule engine to collect and summarize data streams for cognitive workload evaluation. We present a deployment example showing how this architecture is deployed by monitoring an interactive session with an Android application taking into account stressed speech detection, mydriasis and touch analysis.
Davide Maria Calandra, Antonio Caso, Francesco Cutugno, Antonio Origlia, Silvia Rossi 0002
ICMI4
2013 A dynamic tonal perception model for optimal pitch stylization
Antonio Origlia, Giovanni Abete, Francesco Cutugno
Comput. Speech Lang.1
2012 Investigating syllabic prominence with Conditional Random Fields and Latent-Dynamic Conditional Random Fields
Francesco Cutugno, Enrico Leone, Bogdan Ludusan, Antonio Origlia
INTERSPEECH4
2012 W-PhAMT: A web tool for phonetic multilevel timeline visualization
Francesco Cutugno, Vincenza Anna Leano, Antonio Origlia
LREC3
2012 Prosomarker: a prosodic analysis tool based on optimal pitch stylization and automatic syllabi fication
Antonio Origlia, Iolanda Alfano
LREC1
2012 From speech to personality: mapping voice quality and intonation into personality differences
abstract
From a cognitive point of view, personality perception corresponds to capturing individual differences and can be thought of as positioning the people around us in an ideal personality space. The more similar the personality of two individuals, the closer their position in the space. This work shows that the mutual position of two individuals in the personality space can be inferred from prosodic features. The experiments, based on ordinal regression techniques, have been performed over a corpus of 640 speech samples comprising 322 individuals assessed in terms of personality traits by 11 human judges, which is the largest database of this type in the literature. The results show that the mutual position of two individuals can be predicted with up to 80% accuracy.
Gelareh Mohammadi, Antonio Origlia, Maurizio Filippone, Alessandro Vinciarelli
ACM Multimedia2
2012 Attentional and emotional regulation in human-robot interaction
abstract
In this paper, we propose a human-robot interaction system that exploits emotion and attention to regulate and adapt the robotic interactive behavior. In particular, we will focus on the relation between arousal, predictability, and attentional allocation considering as a case study a robotic manipulator interacting with a human operator. We rely on a frequency based model of attention allocation and a 4-dimensional model of emotion. The experiment reported in this paper explores the effectiveness of an attentional regulation mechanisms modulated by arousal and predictability values extracted from the human voice. The collected results show that the attentional modulation, mediated by basic emotional speech features, provides a natural and computationally light regulation mechanism for coordinating the robotic behaviors.
Salvatore Iengo, Antonio Origlia, Mariacarla Staffa, Alberto Finzi
RO-MAN2
2011 On the Use of the Rhythmogram for Automatic Syllabic Prominence Detection
Bogdan Ludusan, Antonio Origlia, Francesco Cutugno
INTERSPEECH2
2011 A Divide et impera Algorithm for Optimal Pitch Stylization
Antonio Origlia, Giovanni Abete, Francesco Cutugno, Iolanda Alfano, Renata Savy, Bogdan Ludusan
INTERSPEECH1