EDBT 2026 Demo / reviewers in the wild / expert
Gianpaolo Coro
dblp:42/6318
· DBLP profile ↗
13ranked-venue papers
7as first author
5since 2021 · last 2024
0000-0001-7232-191XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 4 first-author · 3 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Exploring emergent syllables in end-to-end automatic speech recognizers through model explainability techniqueabstractAbstract Automatic speech recognition systems based on end-to-end models (E2E-ASRs) can achieve comparable performance to conventional ASR systems while reproducing all their essential parts automatically, from speech units to the language model. However, they hide the underlying perceptual processes modelled, if any, and they have lower adaptability to multiple application contexts, and, furthermore, they require powerful hardware and an extensive amount of training data. Model-explainability techniques can explore the internal dynamics of these ASR systems and possibly understand and explain the processes conducting to their decisions and outputs. Understanding these processes can help enhance ASR performance and reduce the required training data and hardware significantly. In this paper, we probe the internal dynamics of three E2E-ASRs pre-trained for English by building an acoustic-syllable boundary detector for Italian and Spanish based on the E2E-ASRs’ internal encoding layer outputs. We demonstrate that the shallower E2E-ASR layers spontaneously form a rhythmic component correlated with prominent syllables, central in human speech processing. This finding highlights a parallel between the analysed E2E-ASRs and human speech recognition. Our results contribute to the body of knowledge by providing a human-explainable insight into behaviours encoded in popular E2E-ASR systems. Vincenzo Norman Vitale, Francesco Cutugno, Antonio Origlia, Gianpaolo Coro |
Neural Comput. Appl. | 4 |
| 2023 | Virtual research environments co-creation: The D4Science experienceabstractAbstract Virtual research environments are systems called to serve the needs of their designated communities of practice. Every community of practice is a group of people dynamically aggregated by the willingness to collaborate to address a given research question. The virtual research environment provides its users with seamless access to the resources of interest (namely, data and services) no matter what and where they are. Developing a virtual research environment thus to guarantee its uptake from the community of practice is a challenging task. In this article, we advocate how the co‐creation driven approach promoted by D4Science has proven to be effective. In particular, we present the co‐creation options supported, discuss how diverse communities of practice have exploited these options, and give some usage indicators on the created VREs. Massimiliano Assante, Leonardo Candela, Donatella Castelli, Roberto Cirillo, Gianpaolo Coro, Andrea Dell'Amico, Luca Frosini, Lucio Lelii, Marco Lettere, Francesco Mangiacrapa, Pasquale Pagano, Giancarlo Panichi, Tommaso Piccioli, Fabio Sinibaldi |
Concurr. Comput. Pract. Exp. | 5 |
| 2023 | A self-training automatic infant-cry detectorabstractAbstract Infant cry is one of the first distinctive and informative life signals observed after birth. Neonatologists and automatic assistive systems can analyse infant cry to early-detect pathologies. These analyses extensively use reference expert-curated databases containing annotated infant-cry audio samples. However, these databases are not publicly accessible because of their sensitive data. Moreover, the recorded data can under-represent specific phenomena or the operational conditions required by other medical teams. Additionally, building these databases requires significant investments that few hospitals can afford. This paper describes an open-source workflow for infant-cry detection, which identifies audio segments containing high-quality infant-cry samples with no other overlapping audio events (e.g. machine noise or adult speech). It requires minimal training because it trains an LSTM-with-self-attention model on infant-cry samples automatically detected from the recorded audio through cluster analysis and HMM classification. The audio signal processing uses energy and intonation acoustic features from 100-ms segments to improve spectral robustness to noise. The workflow annotates the input audio with intervals containing infant-cry samples suited for populating a database for neonatological and early diagnosis studies. On 16 min of hospital phone-audio recordings, it reached sufficient infant-cry detection accuracy in 3 neonatal care environments (nursery—69%, sub-intensive—82%, intensive—77%) involving 20 infants subject to heterogeneous cry stimuli, and had substantial agreement with an expert’s annotation. Our workflow is a cost-effective solution, particularly suited for a sub-intensive care environment, scalable to monitor from one to many infants. It allows a hospital to build and populate an extensive high-quality infant-cry database with a minimal investment. Gianpaolo Coro, Serena Bardelli, Armando Cuttano, Rosa T. Scaramuzzo, Massimiliano Ciantelli |
Neural Comput. Appl. | 1 |
| 2021 | Realizing virtual research environments for the agri-food community: The AGINFRA PLUS experienceabstractAbstract The enhancements in IT solutions and the open science movement are injecting changes in the practices dealing with data collection, collation, processing, analytics, and publishing in all the domains, including agri‐food. However, in implementing these changes one of the major issues faced by the agri‐food researchers is the fragmentation of the “assets” to be exploited when performing research tasks, for example, data of interest are heterogeneous and scattered across several repositories, the tools modelers rely on are diverse and often make use of limited computing capacity, the publishing practices are various and rarely aim at making available the “whole story” including datasets, processes, and results. This paper presents the AGINFRA PLUS endeavor to overcome these limitations by providing researchers in three designated communities with Virtual Research Environments facilitating the use of the “assets” of interest and promote collaboration. Massimiliano Assante, Alice Boizet, Leonardo Candela, Donatella Castelli, Roberto Cirillo, Gianpaolo Coro, Enol Fernández-del-Castillo, Matthias Filter, Luca Frosini, Teodor Georgiev, George Kakaletris, Panagis Katsivelis, Rob Knapen, Lucio Lelii, Rob M. Lokers, Francesco Mangiacrapa, Nikos Manouselis, Pasquale Pagano, Giancarlo Panichi, Lyubomir Penev, Fabio Sinibaldi |
Concurr. Comput. Pract. Exp. | 6 |
| 2021 | NLPHub: An e-Infrastructure-based text mining hubabstractSummary Text mining involves a set of processes that analyze text to extract high‐quality information. Among its large number of applications, there are experiments that tackle big data challenges using complex system architectures. However, text mining approaches are neither easy to discover and use nor easily combinable by end‐users. Furthermore, they should be contextualized within new approaches to science (eg, Open Science) that ensure longevity and reuse of methods and results. This article presents NLPHub, a distributed system that orchestrates and combines several state‐of‐the‐art text mining services that recognize spatiotemporal events, keywords, and a large set of named entities. NLPHub adopts an Open Science approach, which fosters the reproducibility, repeatability, and reusability of methods and results, by using an e‐Infrastructure supporting data‐intensive Science. NLPHub adds Open Science‐compliance to the connected services through the use of representational standards for services and computations. It also manages heterogeneous service access policies and enables collaboration and sharing facilities. This article reports a performance assessment based on an annotated corpus of named entities, which demonstrates that NLPHub can improve the performance of the single‐integrated processes by cleverly combining their output. Gianpaolo Coro, Giancarlo Panichi, Pasquale Pagano, Erico Perrone |
Concurr. Comput. Pract. Exp. | 1 |
| 2019 | Distinguishing Violinists and Pianists Based on Their Brain Signals
Gianpaolo Coro, Giulio Masetti, Philipp Bonhoeffer, Michael Betcher |
ICANN (1) | 1 |
| 2019 | Reconstructing 3D virtual environments within a collaborative e-infrastructureabstractSummary Sets of two‐dimensional images are insufficient to capture the development in time and space of three‐dimensional structures. The 2D “flattening” of photographs results in a significant loss of features especially if the photos were taken by one person. Automatically collecting and aligning photos in order to render 3D structures from 2D images without specialised equipment is currently a complex process that requires specialist knowledge with often limited results. In this paper, an Open Science oriented workflow is proposed where an on‐line file system is used to share photos of an object or an environment and to produce a virtual reality scene as a navigable 3D reconstruction that can be shared with other people. Our workflow is based on a distributed e‐Infrastructure and overcomes common limitations of other approaches by having all the used technology integrated on the same platform and by not requiring specialist knowledge. A performance evaluation of the 3D reconstruction process embedded in the workflow is reported against a commercial software and an open‐source software in terms of computational efficiency and reconstruction accuracy, and three marine science use cases are reported to show potential applications of the workflow. Gianpaolo Coro, Marco Palma, Anton Ellenbroek, Giancarlo Panichi, Thiviya Nair, Pasquale Pagano |
Concurr. Comput. Pract. Exp. | 1 |
| 2019 | The gCube system: Delivering Virtual Research Environments as-a-Service
Massimiliano Assante, Leonardo Candela, Donatella Castelli, Roberto Cirillo, Gianpaolo Coro, Luca Frosini, Lucio Lelii, Francesco Mangiacrapa, Valentina Marioli, Pasquale Pagano, Giancarlo Panichi, Costantino Perciante, Fabio Sinibaldi |
Future Gener. Comput. Syst. | 5 |
| 2019 | Enacting open science by D4Science
Massimiliano Assante, Leonardo Candela, Donatella Castelli, Roberto Cirillo, Gianpaolo Coro, Luca Frosini, Lucio Lelii, Francesco Mangiacrapa, Pasquale Pagano, Giancarlo Panichi, Fabio Sinibaldi |
Future Gener. Comput. Syst. | 5 |
| 2017 | Cloud computing in a distributed e-infrastructure using the web processing service standardabstractSummary New Science paradigms have recently evolved to promote open publication of scientific findings as well as multi‐disciplinary collaborative approaches to scientific experimentation. These approaches can face modern scientific challenges but must deal with large quantities of data produced by industrial and scientific experiments. These data, so‐called Big Data, require to introduce new computer science systems to help scientists cooperate, extract information, and possibly produce new knowledge out of the data. E‐infrastructures are distributed computer systems that foster collaboration between users and can embed distributed and parallel processing systems to manage big data. However, in order to meet modern Science requirements, e‐Infrastructures impose several requirements to computational systems in turn, eg, being economically sustainable, managing community‐provided processes, using standard representations for processes and data, managing big data size and heterogeneous representations, supporting reproducible Science, collaborative experimentation, and cooperative online environments, managing security and privacy for data and services. In this paper, we present a cloud computing system (gCube DataMiner) that meets these requirements and operates in an e‐Infrastructure, while sharing characteristics with state‐of‐the‐art cloud computing systems. To this aim, DataMiner uses the web processing service standard of the open geospatial consortium and introduces features like collaborative experimental spaces, automatic installation of processes and services on top of a flexible and sustainable cloud computing architecture. We compare DataMiner with another mature cloud computing system and highlight the benefits our system brings, the new paradigms requirements it satisfies, and the applications that can be developed based on this system. Gianpaolo Coro, Giancarlo Panichi, Paolo Scarponi, Pasquale Pagano |
Concurr. Comput. Pract. Exp. | 1 |
| 2016 | Species distribution modeling in the cloudabstractSummary Species distribution modeling is a process aiming at computationally predicting the distribution of species in geographic areas on the basis of environmental parameters including climate data. Such a quantitative approach has a lot of potentialities in many areas that include setting up conservation priorities, testing biogeographic hypotheses, and assessing the impact of accelerated land use. To further promote the diffusion of such an approach, it is fundamental to develop a flexible, comprehensive, and robust environment capable of enabling practitioners and communities of practice to produce species distribution models more efficiently. A promising way to build such an environment is offered by modern infrastructures promoting the sharing of resources, including hardware, software, data, and services. This paper describes an approach to species distribution modeling based on a Hybrid Data Infrastructure that can offer a rich array of data and data management services by leveraging other infrastructures (including Cloud). It discusses the whole set of services needed to support the phases of such a complex process including access to occurrence records and environmental parameters and the processing of such information to predict the probability of a species’ occurrence in given areas.Copyright © 2013 John Wiley & Sons, Ltd. Leonardo Candela, Donatella Castelli, Gianpaolo Coro, Pasquale Pagano, Fabio Sinibaldi |
Concurr. Comput. Pract. Exp. | 3 |
| 2015 | Parallelizing the execution of native data mining algorithms for computational biologyabstractSummary Data mining is being increasingly used in biology. Biologists are adopting prototyping languages, like R and Matlab, to facilitate the application of data mining algorithms to their data. As a result, their scripts are becoming increasingly complex and also require frequent updates. Application to large datasets becomes impractical and the time‐to‐paper increases. Furthermore, even if there are various systems that can be used to efficiently process large datasets, for example, using Cloud and High Performance Computing, they usually require procedures to be translated into specific languages or to be adapted to a certain computing platform. Such modifications can speed up the processing, but translation is not automatic, especially in complex cases, and can require a large amount of programming effort and accurate validation. In this paper, we propose an approach to parallelize data mining procedures in the form of compiled software or R scripts developed by biology communities of practice. Our approach requires minimal alteration of the original code. In many cases, there is no need for code modification. Furthermore, it allows for fast updating when a new version is ready. We clarify the constraints and the benefits of our method and report a practical use case to demonstrate such benefits compared with a standard execution. Our approach relies on a distributed network of web services and ultimately exposes the algorithms as‐a‐Service, to be invoked by remote thin clients. Copyright © 2014 John Wiley & Sons, Ltd. Gianpaolo Coro, Leonardo Candela, Pasquale Pagano, Angela Italiano, Loredana Liccardo |
Concurr. Comput. Pract. Exp. | 1 |
| 2007 | Speech recognition with factorial-HMM syllabic acoustic modelsabstractApproaches in Automatic Speech Recognition based on classic acoustic models seem not to exploit all the information lying in a speech signal; furthermore decoding procedures have real time constraints preventing the system to achieve optimal alignment between acoustic models and signal. In this paper, we present an approach to speech recognition in which Factorial Hidden Markov Models (FHMM) are used as syllabic acoustic models. An alignment algorithm is used for unit decoding. As applicative domain we choose numbers (range 0-999,999) uttered in Italian. Syllabic accuracy in our model is 84.81%, correctness on numbers is 77.74%. Aim of the experiment is to show that the performances of FHMMs lie in the ability to retrieve the presence of two different temporal dynamics in a speech segments: the former with a quasi-segmental timing, the latter presenting a quasi-syllabic trend. Moreover, we evaluate a unit decoding process based on a dynamic programming algorithm in order to exploit the acoustic models performances at best. Gianpaolo Coro, Francesco Cutugno, Fulvio Caropreso |
INTERSPEECH | 1 |