Massimo Piccardi

dblp:p/MassimoPiccardi · DBLP profile ↗
← Back
98ranked-venue papers
7as first author
23since 2021 · last 2026
0000-0001-9250-6604ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 42 · 3 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 41 · 5 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 since 2021Systems, architecture and hardware · 6 · 1 first-authorHuman-computer interaction and ubiquitous computing · 6Computer networks · 3 · 1 since 2021Security and privacy · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ViMedCSS: A Vietnamese Medical Code-Switching Speech Dataset & Benchmark
abstract
Code-switching (CS), which is when Vietnamese speech uses English words like drug names or procedures, is a common phenomenon in Vietnamese medical communication. This creates challenges for Automatic Speech Recognition (ASR) systems, especially in low-resource languages like Vietnamese. Current most ASR systems struggle to recognize correctly English medical terms within Vietnamese sentences, and no benchmark addresses this challenge. In this paper, we construct a 34-hour Vietnamese Medical Code-Switching Speech dataset (ViMedCSS) containing 16,576 utterances. Each utterance includes at least one English medical term drawn from a curated bilingual lexicon covering five medical topics. Using this dataset, we evaluate several state-of-the-art ASR models and examine different specific fine-tuning strategies for improving medical term recognition to investigate the best approach to solve in the dataset. Experimental results show that Vietnamese-optimized models perform better on general segments, while multilingual pretraining helps capture English insertions. The combination of both approaches yields the best balance between overall and code-switched accuracy. This work provides the first benchmark for Vietnamese medical code-switching and offers insights into effective domain adaptation for low-resource, multilingual ASR systems.
Tung X. Nguyen, Nhu Vo, Giang-Son Nguyen, Duy Mai Hoang, Chien Dinh Huynh, Inigo Jauregi Unanue, Massimo Piccardi, Wray L. Buntine, Dung D. Le
LREC7
2026 Training Objectives and Evaluation Metrics for Counterfactual Story Rewriting
abstract
Counterfactual story rewriting is an important task of natural language processing that expects a model to comprehend a story and a statement contradicting a part of the story, and suitably rewrite the ending of the story in a way that well reflects the counterfactual statement. This task is particularly challenging due to the fact that the model is not expected to rewrite the original ending from scratch, but only make the minimal, selective changes required to incorporate the countering elements. As such, conventional training objectives that train the models to predict whole reference sentences may fail to capture the nuances of this task. Analogously, standard evaluation metrics that weigh all tokens equally may not be able to discriminate effectively between more correct and less correct predictions. For these reasons, in this article we propose novel training objectives and evaluation metrics that mirror this task more closely, and train and evaluate two Flan-T5 transformer models accordingly. Experiments carried out over a popular counterfactual story rewriting dataset show that the proposed T5 models have achieved significant performance improvements over their respective baselines, and have also outperformed configurations of GPT-3.5, GPT-4o and Gemini 2.0 in most metrics. In addition, a qualitative analysis of the predictions and a heatmap visualization suggest that the modified training objective has been able to focus the model over the required counterfactual elements.
Amelie Girard, Inigo Jauregi Unanue, Massimo Piccardi
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2025 DAT-CRF: Improving the Directed Acyclic Transformer with CRF Integration
abstract
Non-autoregressive transformers (NATs) have demonstrated significant potential in reducing decoding latency for language generation tasks. However, “vanilla” NATs often struggle to effectively capture the sequential structure of generated text. To address this limitation, the recently proposed Directed Acyclic Transformer (DAT) introduces a graph structure into the decoder, explicitly modeling state transitions within the graph. Although DAT has achieved impressive performance, it typically requires large graph sizes to attain optimal results, which are highly memory-intensive, thereby limiting its applicability to tasks involving lengthy output sentences such as document-level machine translation. To mitigate this issue, we propose reducing DAT’s reliance on large graph sizes by coupling its state transition model with a more robust observation model—a Conditional Random Field (CRF). The CRF inherently models pairwise transitions between output tokens, enabling the model to capture dependencies without relying on large graphs. In addition, unlike other NAT-CRF models where the NAT and CRF modules operate independently, our approach is the first to jointly decode both modules, permitting joint optimality of the inferred graph and the output tokens. Experimental results on both sentence-level and document-level machine translation show that this modification substantially improves the baseline DAT in both lexical and semantic metrics, while retaining near-parity of decoding speed.
Inigo Jauregi Unanue, Massimo Piccardi
ECAI3
2025 SPAD - A Secure and Privacy-Preserving Distributed Analytics Framework
Imran Makhdoom, Mehran Abolhasan, Justin Lipman, Daniel Robert Franklin, Massimo Piccardi
ICBC5
2025 Defeating Eavesdropping Attacks with Inter-Cell Interference and Deep Reinforcement Learning
abstract
This paper introduces a novel joint user association and resource allocation framework to efficiently deal with eavesdropping attacks without requiring prior information about eavesdroppers. Specifically, the co-channel interference when reusing resource blocks is leveraged to disrupt the signal reception at eavesdroppers. To maximize the secure area, defined as the area where eavesdroppers cannot wiretap the channel due to co-channel interference, we first formulate the system by using the Markov decision process to capture the dynamics and uncertainty of mobile users and wireless communications. Then, a deep reinforcement learning (DRL)-based approach is proposed to obtain the joint optimal user association and resource allocation policy to utilize the co-channel interference and maximize the secure area. Extensive simulation results demonstrate that by intelligently associating users and allocating resource blocks to them, our proposed solution can help to effectively defeat eavesdropping attacks without requiring prior information of eavesdroppers which may not be readily available in practice. In addition, the proposed DRL-based algorithm can converge to the optimal policy quickly and achieve better performance compared to existing solutions.
Nguyen Van Huynh, Diep N. Nguyen, Lorenzo Mucchi, Stefano Caputo, Massimo Piccardi, Dinh Thai Hoang, Eryk Dutkiewicz
WCNC5
2025 Empowering large language models for automated clinical assessment with generation-augmented retrieval and hierarchical chain-of-thought
abstract
BACKGROUND: Understanding and extracting valuable information from electronic health records (EHRs) is important for improving healthcare delivery and health outcomes. Large language models (LLMs) have demonstrated significant proficiency in natural language understanding and processing, offering promises for automating the typically labor-intensive and time-consuming analytical tasks with EHRs. Despite the active application of LLMs in the healthcare setting, many foundation models lack real-world healthcare relevance. Applying LLMs to EHRs is still in its early stage. To advance this field, in this study, we pioneer a generation-augmented prompting paradigm "GAPrompt" to empower generic LLMs for automated clinical assessment, in particular, quantitative stroke severity assessment, using data extracted from EHRs. METHODS: The GAPrompt paradigm comprises five components: (i) prompt-driven selection of LLMs, (ii) generation-augmented construction of a knowledge base, (iii) summary-based generation-augmented retrieval (SGAR); (iv) inferencing with a hierarchical chain-of-thought (HCoT), and (v) ensembling of multiple generations. RESULTS: GAPrompt addresses the limitations of generic LLMs in clinical applications in a progressive manner. It efficiently evaluates the applicability of LLMs in specific tasks through LLM selection prompting, enhances their understanding of task-specific knowledge from the constructed knowledge base, improves the accuracy of knowledge and demonstration retrieval via SGAR, elevates LLM inference precision through HCoT, enhances generation robustness, and reduces hallucinations of LLM via ensembling. Experiment results demonstrate the capability of our method to empower LLMs to automatically assess EHRs and generate quantitative clinical assessment results. CONCLUSION: Our study highlights the applicability of enhancing the capabilities of foundation LLMs in medical domain-specific tasks, i.e., automated quantitative analysis of EHRs, addressing the challenges of labor-intensive and often manually conducted quantitative assessment of stroke in clinical practice and research. This approach offers a practical and accessible GAPrompt paradigm for researchers and industry practitioners seeking to leverage the power of LLMs in domain-specific applications. Its utility extends beyond the medical domain, applicable to a wide range of fields.
Zhanzhong Gu, Wenjing Jia, Massimo Piccardi, Ping Yu 0004
Artif. Intell. Medicine3
2024 XVD: Cross-Vocabulary Differentiable Training for Generative Adversarial Attacks
abstract
An adversarial attack to a text classifier consists of an input that induces the classifier into an incorrect class prediction, while retaining all the linguistic properties of correctly-classified examples. A popular class of adversarial attacks exploits the gradients of the victim classifier to train a dedicated generative model to produce effective adversarial examples. However, this training signal alone is not sufficient to ensure other desirable properties of the adversarial attacks, such as similarity to non-adversarial examples, linguistic fluency, grammaticality, and so forth. For this reason, in this paper we propose a novel training objective which leverages a set of pretrained language models to promote such properties in the adversarial generation. A core component of our approach is a set of vocabulary-mapping matrices which allow cascading the generative model to any victim or component model of choice, while retaining differentiability end-to-end. The proposed approach has been tested in an ample set of experiments covering six text classification datasets, two victim models, and four baselines. The results show that it has been able to produce effective adversarial attacks, outperforming the compared generative approaches in a majority of cases and proving highly competitive against established token-replacement approaches.
Tom Roth, Inigo Jauregi Unanue, Alsharif Abuadbba, Massimo Piccardi
LREC/COLING4
2024 Improving Vietnamese-English Medical Machine Translation
abstract
Machine translation for Vietnamese-English in the medical domain is still an under-explored research area. In this paper, we introduce MedEV—a high-quality Vietnamese-English parallel dataset constructed specifically for the medical domain, comprising approximately 360K sentence pairs. We conduct extensive experiments comparing Google Translate, ChatGPT (gpt-3.5-turbo), state-of-the-art Vietnamese-English neural machine translation models and pre-trained bilingual/multilingual sequence-to-sequence models on our new MedEV dataset. Experimental results show that the best performance is achieved by fine-tuning “vinai-translate” for each translation direction. We publicly release our dataset to promote further research.
Nhu Vo, Dat Quoc Nguyen, Dung D. Le, Massimo Piccardi, Wray L. Buntine
LREC/COLING4
2024 SumTra: A Differentiable Pipeline for Few-Shot Cross-Lingual Summarization
abstract
Jacob Parnell, Inigo Jauregi Unanue, Massimo Piccardi. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Jacob Parnell, Inigo Jauregi Unanue, Massimo Piccardi
NAACL-HLT3
2024 LayerGLAT: A Flexible Non-autoregressive Transformer for Single-Pass and Multi-pass Prediction
Inigo Jauregi Unanue, Massimo Piccardi
ECML/PKDD (2)3
2024 PrivySeC: A secure and privacy-compliant distributed framework for personal data sharing in IoT ecosystems
abstract
The contemporary era experiences an unprecedented dependence on data generated by individuals via an array of interconnected devices constituting the Internet of Things (IoT). The information amassed through IoT devices serves many objectives, including prescriptive analytics and predictive maintenance, preemptive healthcare measures, disaster mitigation, operational efficiency, and increased yield. In contrast, most applications or systems that rely on user-generated data to fulfill their business objectives face challenges in adhering to privacy protocols. Consequently, users are exposed to many privacy risks. Such infringements upon privacy provisions give rise to apprehensions regarding the authenticity of the processed data. Hence, this paper presents weaknesses and challenges in current practices and proposes “PrivySeC,” a distributed ledger technology (DLT) based framework for privacy-preserving and secure sharing of personally and non-personally identifiable information. The security analysis indicates that the proposed solution ensures data privacy by design and complies with most of the requirements mandated by various privacy regulations. Similarly, PrivySeC promises low transaction latency and provides high throughput. Although we have created a privacy-preserving solution for sharing smart farm data, it can be customized to meet the specific privacy requirements of individual applications.
Imran Makhdoom, Mehran Abolhasan, Justin Lipman, Massimo Piccardi, Daniel Robert Franklin
Blockchain Res. Appl.4
2023 Improving Machine Translation and Summarization with the Sinkhorn Divergence
Inigo Jauregi Unanue, Massimo Piccardi
PAKDD (4)3
2023 I2Map: IoT Device Attestation Using Integrity Map
abstract
The reliability of any IoT system's operation depends upon the accuracy of sensor data. Numerous sensing devices are also embedded in autonomous systems. The adversary can manipulate the output of an IoT device, such as a speed sensor or a temperature sensor, by altering its hardware, software modules, network parameters, or device configuration. These unauthorized changes may affect the legitimate operation of the autonomous system and cause a malfunction or a safety hazard. In addition, the corrupt sensor data may also affect the machine learning models by introducing false training data or biasing the model towards certain decisions to cause the autonomous system to make incorrect decisions. The existing techniques mostly perform memory attestation or ensure the secure execution of an application. In addition, current approaches have unrealistic assumptions about adversaries’ capabilities and rely on trusted parties to initiate and run the attestation protocol. No existing technique provides all the required security features, including protection against return-oriented programming and network attacks, including interference, rainbow, and physical compromise. Hence, this research presents a secure and efficient version of a unique hybrid attestation scheme "I2Map" that detects a malicious or malfunctioning IoT device based on an integrity map. The performance analysis infers that I2Map performs better in transaction commit time and transaction costs with more device parameters than its predecessor.
Imran Makhdoom, Mehran Abolhasan, Justin Lipman, Daniel Robert Franklin, Massimo Piccardi
TrustCom5
2023 Detecting compromised IoT devices: Existing techniques, challenges, and a way forward
Imran Makhdoom, Mehran Abolhasan, Daniel Robert Franklin, Justin Lipman, Massimo Piccardi, Negin Shariati
Comput. Secur.6
2023 T3L: Translate-and-Test Transfer Learning for Cross-Lingual Text Classification
abstract
Abstract Cross-lingual text classification leverages text classifiers trained in a high-resource language to perform text classification in other languages with no or minimal fine-tuning (zero/ few-shots cross-lingual transfer). Nowadays, cross-lingual text classifiers are typically built on large-scale, multilingual language models (LMs) pretrained on a variety of languages of interest. However, the performance of these models varies significantly across languages and classification tasks, suggesting that the superposition of the language modelling and classification tasks is not always effective. For this reason, in this paper we propose revisiting the classic “translate-and-test” pipeline to neatly separate the translation and classification stages. The proposed approach couples 1) a neural machine translator translating from the targeted language to a high-resource language, with 2) a text classifier trained in the high-resource language, but the neural machine translator generates “soft” translations to permit end-to-end backpropagation during fine-tuning of the pipeline. Extensive experiments have been carried out over three cross-lingual text classification datasets (XNLI, MLDoc, and MultiEURLEX), with the results showing that the proposed approach has significantly improved performance over a competitive baseline.
Inigo Jauregi Unanue, Gholamreza Haffari, Massimo Piccardi
Trans. Assoc. Comput. Linguistics3
2023 Topic-Based Unsupervised and Supervised Dictionary Induction
abstract
Word translation is a natural language processing task that provides translation between the words of a source and a target language. As a task, it reduces to the induction of a bilingual dictionary, which is typically performed by aligning word embeddings of the source language to word embeddings of the target language. To date, all the existing approaches have focused on performing a single, global alignment in word embedding space. However, semantic differences between the various languages, in addition to differences in the content of the corpora used for training the word embeddings, can hinder the effectiveness of such a global alignment. For this reason, in this article we propose conducting the alignment between the source and target embedding spaces by multiple mappings at topic level. The experimental results show that our approach has been able to achieve an average accuracy improvement of +3.30 percentage points over a state-of-the-art approach in unsupervised dictionary induction from languages as diverse as German, French, Italian, Spanish, Finnish, Turkish, and Chinese to English, and +3.95 points average improvement in supervised dictionary induction.
Massimo Piccardi
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2022 A Multi-Document Coverage Reward for RELAXed Multi-Document Summarization
abstract
Multi-document summarization (MDS) has made significant progress in recent years, in part facilitated by the availability of new, dedicated datasets and capacious language models.However, a standing limitation of these models is that they are trained against limited references and with plain maximum-likelihood objectives.As for many other generative tasks, reinforcement learning (RL) offers the potential to improve the training of MDS models; yet, it requires a carefully-designed reward that can ensure appropriate leverage of both the reference summaries and the input documents.For this reason, in this paper we propose fine-tuning an MDS baseline with a reward that balances a reference-based metric such as ROUGE with coverage of the input documents.To implement the approach, we utilize RELAX (Grathwohl et al., 2018), a contemporary gradient estimator which is both low-variance and unbiased, and we fine-tune the baseline in a few-shot style for both stability and computational efficiency.Experimental results over the Multi-News and WCEP MDS datasets show significant improvements of up to +0.95 pp average ROUGE score and +3.17 pp METEOR score over the baseline, and competitive results with the literature.In addition, they show that the coverage of the input documents is increased, and evenly across all documents.
Jacob Parnell, Inigo Jauregi Unanue, Massimo Piccardi
ACL (1)3
2022 Neural Topic Model Training with the REBAR Gradient Estimator
abstract
Topic modelling is an important approach of unsupervised machine learning that allows automatically extracting the main “topics” from large collections of documents. In addition, topic modelling is able to identify the topic proportions of each individual document, which can be helpful for organizing the collections. Many topic modelling algorithms have been proposed to date, including several that leverage advanced techniques such as variational inference and deep autoencoders. However, to date topic modelling has made limited use of reinforcement learning, a framework that has obtained vast success in many other unsupervised learning tasks. For this reason, in this article we propose training a neural topic model using a reinforcement learning objective and minimizing the objective with the recently-proposed REBAR gradient estimator. Experiments performed over two probing datasets have shown that the proposed model has achieved improvements over all the compared models in terms of both model perplexity and topic coherence, and produced topics that appear qualitatively informative and consistent.
Amit Kumar 0031, Nazanin Esmaili, Massimo Piccardi
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2022 BiLSTM-SSVM: Training the BiLSTM with a Structured Hinge Loss for Named-Entity Recognition
abstract
Building on the achievements of the BiLSTM-CRF in named-entity recognition (NER), this paper introduces the BiLSTM-SSVM, an equivalent neural model where training is performed using a structured hinge loss. The typical loss functions used for evaluating NER are entity-level variants of the$F_1$F1score such as the CoNLL and MUC losses. Unfortunately, the common loss function used for training NER - the cross entropy - is only loosely related to the evaluation losses. For this reason, in this paper we propose a training approach for the BiLSTM-CRF that leverages a hinge loss bounding the CoNLL loss from above. In addition, we present a mixed hinge loss that bounds either the CoNLL loss or the Hamming loss based on the density of entity tokens in each sentence. The experimental results over four benchmark languages (English, German, Spanish and Dutch) show that training with the mixed hinge loss has led to small but consistent improvements over the cross entropy across all languages and four different evaluation measures.
Hanieh Poostchi, Massimo Piccardi
IEEE Trans. Big Data2
2021 A REINFORCEd Variational Autoencoder Topic Model
Amit Kumar 0031, Nazanin Esmaili, Massimo Piccardi
ICONIP (5)3
2021 Improving Adversarial Text Generation with n-Gram Matching
Massimo Piccardi
PACLIC2
2021 Multichannel mixture models for time-series analysis and classification of engagement with multiple health services: An application to psychology and physiotherapy utilization patterns after traffic accidents
Nazanin Esmaili, Quinlan D. Buchlak, Massimo Piccardi, Bernie Kruger, Federico Girosi
Artif. Intell. Medicine3
2021 An Embedding-Based Topic Model for Document Classification
abstract
Topic modeling is an unsupervised learning task that discovers the hidden topics in a collection of documents. In turn, the discovered topics can be used for summarizing, organizing, and understanding the documents in the collection. Most of the existing techniques for topic modeling are derivatives of the Latent Dirichlet Allocation which uses a bag-of-word assumption for the documents. However, bag-of-words models completely dismiss the relationships between the words. For this reason, this article presents a two-stage algorithm for topic modelling that leverages word embeddings and word co-occurrence. In the first stage, we determine the topic-word distributions by soft-clustering a random set of embedded n -grams from the documents. In the second stage, we determine the document-topic distributions by sampling the topics of each document from the topic-word distributions. This approach leverages the distributional properties of word embeddings instead of using the bag-of-words assumption. Experimental results on various data sets from an Australian compensation organization show the remarkable comparative effectiveness of the proposed algorithm in a task of document classification.
Sattar Seifollahi, Massimo Piccardi, Alireza Jolfaei
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2020 Leveraging Discourse Rewards for Document-Level Neural Machine Translation
abstract
Document-level machine translation focuses on the translation of entire documents from a source to a target language.It is widely regarded as a challenging task since the translation of the individual sentences in the document needs to retain aspects of the discourse at document level.However, document-level translation models are usually not trained to explicitly ensure discourse quality.Therefore, in this paper we propose a training approach that explicitly optimizes two established discourse metrics, lexical cohesion (LC) and coherence (COH), by using a reinforcement learning objective.Experiments over four different language pairs and three translation domains have shown that our training approach has been able to achieve more cohesive and coherent document translations than other competitive approaches, yet without compromising the faithfulness to the reference translation.In the case of the Zh-En language pair, our method has achieved an improvement of 2.46 percentage points (pp) in LC and 1.17 pp in COH over the runner-up, while at the same time improving 0.63
Inigo Jauregi Unanue, Nazanin Esmaili, Gholamreza Haffari, Massimo Piccardi
COLING4
2020 Learning Neural Textual Representations for Citation Recommendation
abstract
With the rapid growth of the scientific literature, manually selecting appropriate citations for a paper is becoming increasingly challenging and time-consuming. While several approaches for automated citation recommendation have been proposed in the recent years, effective document representations for citation recommendation are still elusive to a large extent. For this reason, in this paper we propose a novel approach to citation recommendation which leverages a deep sequential representation of the documents (Sentence-BERT) cascaded with Siamese and triplet networks in a submodular scoring function. To the best of our knowledge, this is the first approach to combine deep representations and submodular selection for a task of citation recommendation. Experiments have been carried out using a popular benchmark dataset - the ACL Anthology Network corpus - and evaluated against baselines and a state-of-the-art approach using metrics such as the MRR and F1@ k score. The results show that the proposed approach has been able to outperform all the compared approaches in every measured metric.
Binh Thanh Kieu, Inigo Jauregi Unanue, Son Bao Pham, Hieu Xuan Phan, Massimo Piccardi
ICPR5
2020 Controlled Text Generation with Adversarial Learning
abstract
In recent years, generative adversarial networks (GANs) have started to attain promising results also in natural language generation.However, the existing models have paid limited attention to the semantic coherence of the generated sentences.For this reason, in this paper we propose a novel network -the Controlled TExt generation Relational Memory GAN (CTERM-GAN) -that uses an external input to influence the coherence of sentence generation.The network is composed of three main components: a generator based on a Relational Memory conditioned on the external input; a syntactic discriminator which learns to discriminate between real and generated sentences; and a semantic discriminator which assesses the coherence with the external conditioning.Our experiments on six probing datasets have showed that the model has been able to achieve interesting results, retaining or improving the syntactic quality of the generated sentences while significantly improving their semantic coherence with the given input.
Federico Betti 0001, Giorgia Ramponi, Massimo Piccardi
INLG3
2019 Taxonomy-Based Feature Extraction for Document Classification, Clustering and Semantic Analysis
Sattar Seifollahi, Massimo Piccardi
CICLing (2)2
2019 A simulated annealing-based maximum-margin clustering algorithm
abstract
Abstract Maximum‐margin clustering is an extension of the support vector machine (SVM) to clustering. It partitions a set of unlabeled data into multiple groups by finding hyperplanes with the largest margins. Although existing algorithms have shown promising results, there is no guarantee of convergence of these algorithms to global solutions due to the nonconvexity of the optimization problem. In this paper, we propose a simulated annealing‐based algorithm that is able to mitigate the issue of local minima in the maximum‐margin clustering problem. The novelty of our algorithm is twofold, ie, (i) it comprises a comprehensive cluster modification scheme based on simulated annealing, and (ii) it introduces a new approach based on the combination of k‐means++ and SVM at each step of the annealing process. More precisely, k‐means++ is initially applied to extract subsets of the data points. Then, an unsupervised SVM is applied to improve the clustering results. Experimental results on various benchmark data sets (of up to over a million points) give evidence that the proposed algorithm is more effective at solving the clustering problem than a number of popular clustering algorithms.
Sattar Seifollahi, Adil M. Bagirov, Ehsan Zare Borzeshi, Massimo Piccardi
Comput. Intell.4
2018 SMGKM: An Efficient Incremental Algorithm for Clustering Document Collections
Adil M. Bagirov, Sattar Seifollahi, Massimo Piccardi, Ehsan Zare Borzeshi, Bernie Kruger
CICLing (2)3
2018 BiLSTM-CRF for Persian Named-Entity Recognition ArmanPersoNERCorpus: the First Entity-Annotated Persian Dataset
Hanieh Poostchi, Ehsan Zare Borzeshi, Massimo Piccardi
LREC3
2018 English-Basque Statistical and Neural Machine Translation
Inigo Jauregi Unanue, Lierni Garmendia Arratibel, Ehsan Zare Borzeshi, Massimo Piccardi
LREC4
2018 Minimum-risk temporal alignment of videos
Massimo Piccardi
Multim. Tools Appl.2
2018 AdOn HDP-HMM: An Adaptive Online Model for Segmentation and Classification of Sequential Data
abstract
Recent years have witnessed an increasing need for the automated classification of sequential data, such as activities of daily living, social media interactions, financial series, and others. With the continuous flow of new data, it is critical to classify the observations on-the-fly and without being limited by a predetermined number of classes. In addition, a model should be able to update its parameters in response to a possible evolution in the distributions of the classes. This compelling problem, however, does not seem to have been adequately addressed in the literature, since most studies focus on offline classification over predefined class sets. In this paper, we present a principled solution for this problem based on an adaptive online system leveraging Markov switching models and hierarchical Dirichlet process priors. This adaptive online approach is capable of classifying the sequential data over an unlimited number of classes while meeting the memory and delay constraints typical of streaming contexts. In this paper, we introduce an adaptive "learning rate" that is responsible for balancing the extent to which the model retains its previous parameters or adapts to new observations. Experimental results on stationary and evolving synthetic data and two video data sets, TUM Assistive Kitchen and collated Weizmann, show a remarkable performance in terms of segmentation and classification, particularly for sequences from evolutionary distributions and/or those containing previously unseen classes.
Ava Bargi, Massimo Piccardi
IEEE Trans. Neural Networks Learn. Syst.3
2018 Sequential Labeling With Structural SVM Under Nondecomposable Losses
abstract
Sequential labeling addresses the classification of sequential data, which are widespread in fields as diverse as computer vision, finance, and genomics. The model traditionally used for sequential labeling is the hidden Markov model (HMM), where the sequence of class labels to be predicted is encoded as a Markov chain. In recent years, HMMs have benefited from minimum-loss training approaches, such as the structural support vector machine (SSVM), which, in many cases, has reported higher classification accuracy. However, the loss functions available for training are restricted to decomposable cases, such as the 0-1 loss and the Hamming loss. In many practical cases, other loss functions, such as those based on the $F_{1}$ measure, the precision/recall break-even point, and the average precision (AP), can describe desirable performance more effectively. For this reason, in this paper, we propose a training algorithm for SSVM that can minimize any loss based on the classification contingency table, and we present a training algorithm that minimizes an AP loss. Experimental results over a set of diverse and challenging data sets (TUM Kitchen, CMU Multimodal Activity, and Ozone Level Detection) show that the proposed training algorithms achieve significant improvements of the $F_{1}$ measure and AP compared with the conventional SSVM, and their performance is in line with or above that of other state-of-the-art sequential labeling approaches.
Guopeng Zhang, Massimo Piccardi, Ehsan Zare Borzeshi
IEEE Trans. Neural Networks Learn. Syst.2
2017 Minimum-Risk Structured Learning of Video Summarization
abstract
Video summarization is an important multimedia task for applications such as video indexing and retrieval, video surveillance, human-computer interaction and video "storyboarding". In this paper, we present a new approach for automatic summarization of video collections that leverages a structured minimum-risk classifier and efficient submodular inference. To test the accuracy of the predicted summaries we utilize a recently-proposed measure (V-JAUNE) that considers both the content and frame order of the original video. Qualitative and quantitative tests over two action video datasets - the ACE and the MSR DailyActivity3D datasets - show that the proposed approach delivers more accurate summaries than the compared minimum-risk and syntactic approaches.
Fairouz Hussein, Massimo Piccardi
ISM2
2017 Dissimilarity-based action recognition with the pair hidden Markov support vector machine
abstract
Human action recognition in video is highly challenging due to the substantial variations in motion performance, recording settings and inter-personal differences. Most current research focuses on the extraction of effective features and the design of suitable classifiers. Conversely, in this paper we tackle this problem by a dissimilarity-based approach where classification is performed in terms of minimum distance from templates. To measure the dissimilarity between any two action instances, we propose leveraging the Pair Hidden Markov Support Vector Machine (PHMM-SSVM) that was recently proposed for tasks of video alignment. The main advantages of PHMM-SSVM are its ability to learn optimal alignment models from training sets of manually-aligned action pairs and provide alignment scores that can be used for action classification. The experimental results over two popular action datasets show that the proposed approach has been capable of achieving an accuracy higher than many existing methods and comparable to a state-of-the-art algorithm.
Massimo Piccardi
MMSP2
2017 Recurrent neural networks with specialized word embeddings for health-domain named-entity recognition
Inigo Jauregi Unanue, Ehsan Zare Borzeshi, Massimo Piccardi
J. Biomed. Informatics3
2017 Prototype-based budget maintenance for tracking in depth videos
Sari Awwad, Massimo Piccardi
Multim. Tools Appl.2
2017 V-JAUNE: A Framework for Joint Action Recognition and Video Summarization
abstract
Video summarization and action recognition are two important areas of multimedia video analysis. While these two areas have been tackled separately to date, in this article, we present a latent structural SVM framework to recognize the action and derive the summary of a video in a joint, simultaneous fashion. Efficient inference is provided by a submodular score function that accounts for the action and summary jointly. In this article, we also define a novel measure to evaluate the quality of a predicted video summary against the annotations of multiple annotators. Quantitative and qualitative results over two challenging action datasets—the ACE and MSR DailyActivity3D datasets—show that the proposed joint approach leads to higher action recognition accuracy and equivalent or better summary quality than comparable approaches that perform these tasks separately.
Fairouz Hussein, Massimo Piccardi
ACM Trans. Multim. Comput. Commun. Appl.2
2016 PersoNER: Persian Named-Entity Recognition
abstract
Named-Entity Recognition (NER) is still a challenging task for languages with low digital resources. The main difficulties arise from the scarcity of annotated corpora and the consequent problematic training of an effective NER pipeline. To abridge this gap, in this paper we target the Persian language that is spoken by a population of over a hundred million people world-wide. We first present and provide ArmanPerosNERCorpus, the first manually-annotated Persian NER corpus. Then, we introduce PersoNER, an NER pipeline for Persian that leverages a word embedding and a sequential max-margin classifier. The experimental results show that the proposed approach is capable of achieving interesting MUC7 and CoNNL scores while outperforming two alternatives based on a CRF and a recurrent neural network.
Hanieh Poostchi, Ehsan Zare Borzeshi, Mohammad Abdous, Massimo Piccardi
COLING4
2016 Joint action recognition and summarization by sub-modular inference
abstract
Action recognition and video summarization are two important multimedia tasks that are useful for applications such as video indexing and retrieval, video surveillance, human-computer interaction and home intelligence. While many approaches exist in the literature for these two tasks, to date they have always been addressed separately. Instead, in this paper we move from the assumption that these two tasks should be tackled as a joint objective: on the one hand, action recognition can drive the selection of meaningful and informative summaries; on the other, recognizing actions from a summary rather than the entire video can in principle reduce noise and prove more accurate. To this aim, we propose a novel approach for joint action recognition-summarization based on the performing latent structural SVM framework, together with an efficient algorithm for inferring the action and the summary based on the property of sub-modularity. Experimental results on a challenging benchmark, MSR Dai-lyActivity3D, show that the approach is capable of achieving remarkable action recognition accuracy while providing appealing video summaries.
Fairouz Hussein, Sari Awwad, Massimo Piccardi
ICASSP3
2016 A pair hidden Markov support vector machine for alignment of human actions
abstract
Alignment of human actions in videos is an important task for applications such as action comparison and classification. While well-established algorithms such as dynamic time warping are available for this task, they still heavily rely on basic linear cost models and heuristic parameter tuning. In this paper we propose a novel framework that combines the flexibility of the pair hidden Markov model (PHMM) with the effective parameter training of the structural support vector machine (SSVM). The framework extends the scoring function of SSVM to capture the similarity of two input sequences and introduces suitable feature and loss functions. The proposed approach is evaluated against state-of-the-art algorithms such as dynamic time warping (DTW) and canonical time warping (CTW) on pairs of human actions from the Weizmann and Olympic Sports datasets. The experimental results show that the proposed approach is capable of achieving an accuracy improvement of over 7 percentage points over the runner-up on both datasets.
Massimo Piccardi
ICME2
2016 Static action recognition by efficient greedy inference
abstract
Action recognition from a single image is an important task for applications such as image annotation, robotic navigation, video surveillance and several others. Existing methods for recognizing actions from still images mainly rely on either bag-of-feature representations or pose estimation from articulated body-part models. However, the relationship between the action and the containing image is still substantially unexplored. Actually, the presence of given objects or specific backgrounds is likely to provide informative clues for the recognition of the action. For this reason, in this paper we propose approaching action recognition by first partitioning the entire image into superpixels, and then using their latent classes as attributes of the action. The action class is predicted based on a graphical model composed of measurements from each superpixel and a fully-connected graph of superpixel classes. The model is learned using a latent structural SVM approach, and an efficient, greedy algorithm is proposed to provide inference over the graph. Differently from most existing methods, the proposed approach does not require annotation of the actor (usually provided as a bounding box). Experimental results over the challenging Stanford 40 Action dataset have reported an impressive mean average precision of 72.3%, the highest achieved to date.
Shaukat R. Abidi, Massimo Piccardi, Mary-Anne Williams
WACV2
2016 Big data meets multimedia analytics
Tat-Seng Chua, Xiangjian He, Weifeng Liu 0001, Massimo Piccardi, Yonggang Wen 0001, Dacheng Tao
Signal Process.4
2015 Local Depth Patterns for Tracking in Depth Videos
abstract
Conventional video tracking operates over RGB or grey-level data which contain significant clues for the identification of the targets. While this is often desirable in a video surveillance context, use of video tracking in privacy-sensitive environments such as hospitals and care facilities is often perceived as intrusive. Therefore, in this work we present a tracker that provides effective target tracking based solely on depth data. The proposed tracker is an extension of the popular Struck algorithm which leverages a structural SVM framework for tracking. The main contributions of this work are novel depth features based on local depth patterns and a heuristic for effectively handling occlusions. Experimental results over the challenging Princeton Tracking Benchmark (PTB) dataset report a remarkable accuracy compared to the original Stuck tracker and other state-of-the-art trackers using depth and RGB data.
Sari Awwad, Fairouz Hussein, Massimo Piccardi
ACM Multimedia3
2015 Tracking people under heavy occlusions by layered data association
Zui Zhang, Óscar Pérez, Massimo Piccardi
Multim. Tools Appl.3
2015 Structural SVM with Partial Ranking for Activity Segmentation and Classification
abstract
Structural SVM is an extension of the support vector machine for the joint prediction of structured labels from multiple measurements. Following a large margin principle, the training of structural SVM ensures that the ground-truth labeling of each sample receives a score higher than that of any other labeling. However, no specific score ranking is imposed among the other labelings. In this letter, we extend the standard constraint set of structural SVM with constraints between “almost-correct” labelings and less desirable ones to obtain a partial-ranking structural SVM (PR-SSVM) approach. Experimental results on action segmentation and classification with two challenging datasets (the TUM Kitchen mocap dataset and the CMU-MMAC video dataset) show that the proposed method achieves better detection and false alarm rates and higher F1 scores than both the conventional structural SVM and a comparable unstructured predictor. The proposed method also achieves higher accuracy than the state of the art on these datasets in excess of 14 and 31 percentage points, respectively.
Guopeng Zhang, Massimo Piccardi
IEEE Signal Process. Lett.2
2014 A Non-parametric Conditional Factor Regression Model for Multi-Dimensional Input and Response
abstract
In this paper, we propose a non-parametric conditional factor regression (NCFR) model for domains with multi-dimensional input and response. NCFR enhances linear regression in two ways: a) introducing low-dimensional latent factors leading to dimensionality reduction and b) integrating the Indian Buffet Process as prior for the latent layer to dynamically derive an optimal number of sparse factors. Thanks to IBP’s enhancements to the latent factors, NCFR can significantly avoid over-fitting even in the case of a very small sample size compared to the dimensionality. Experimental results on three diverse datasets comparing NCRF to a few baseline alternatives give evidence of its robust learning, remarkable predictive performance, good mixing and computational efficiency.
Ava Bargi, Zoubin Ghahramani, Massimo Piccardi
AISTATS4
2014 Complex event recognition by latent temporal models of concepts
abstract
Complex event recognition is an expanding research area aiming to recognize entities of high-level semantics in videos. Typical approaches exploit the so-called “bags” of spatiotemporal features such as STIP, ISA and DTF-HOG; yet, more recently, the notion of concept has emerged as an alternative, intermediate representation with greater descriptive power, and “bags of concepts” have been used for recognition. In this paper we argue that concepts in an event tend to articulate over a discernible temporal structure and we exploit a temporal model using the scores of concept detectors as measurements. In addition, we propose several heuristics to improve the initialization of the model's latent states and take advantage of the time-sparsity of the concepts. Experimental results on videos from the challenging TRECVID MED 2012 dataset show that the proposed approach achieves an improvement in average precision of 8.92% over comparable bags of concepts, thus validating the use of temporal structure over concepts for complex event recognition.
Ehsan Zare Borzeshi, Afshin Dehghan, Massimo Piccardi, Mubarak Shah
ICIP3
2014 Sequential labeling with structural SVM under the F1 loss
abstract
Sequential labeling addresses the classification of sequential data and is of increasing importance for the classification and segmentation of video data. The model traditionally used for sequential labeling is the hidden Markov model where the sequence of class labels to be predicted is encoded as a Markov chain. In recent years, hidden Markov models and other structural models have benefited from minimum-loss training approaches which in many cases lead to greater classification accuracy. However, the loss functions available for training are restricted to decomposable cases such as the zero-one loss and the Hamming loss. Other useful losses such as the F1loss, equal error rates and others are not available for sequential labeling. For this reason, in this paper we propose a training algorithm that can cater for the F1loss and any other loss function based on the contingency table. Experimental results over the challenging TUM Kitchen Dataset depicting human actions in a kitchen scenario show that the proposed training approach leads to significant improvement of different performance metrics such as the classification accuracy (4.3 percentage points) and the F1measure (8.9 percentage points).
Guopeng Zhang, Massimo Piccardi
ICIP2
2014 An Infinite Adaptive Online Learning Model for Segmentation and Classification of Streaming Data
abstract
In recent years, the desire and need to understand streaming data has been increasing. Along with the constant flow of data, it is critical to classify and segment the observations on-the-fly without being limited to a rigid number of classes. In other words, the system needs to be adaptive to the streaming data and capable of updating its parameters to comply with natural changes. This interesting problem, however, is poorly addressed in the literature, as many of the common studies focus on offline classification over a pre-defined class set. In this paper, we propose a novel adaptive online system based on Markov switching models with hierarchical Dirichlet process priors. This infinite adaptive online approach is capable of segmenting and classifying the streaming data over infinite classes, while meeting the memory and delay constraints of streaming contexts. The model is further enhanced by a 'predictive batching' mechanism, that is able to divide the flowing data into batches of variable size, imitating the ground-truth segments. Experiments on two video datasets show significant performance of the proposed approach in frame-level accuracy, segmentation recall and precision, while determining the accurate number of classes in acceptable computational time.
Ava Bargi, Massimo Piccardi
ICPR3
2014 Special issue on background modeling for foreground detection in real-world dynamic scenes
Thierry Bouwmans, Jordi Gonzàlez 0001, Caifeng Shan, Massimo Piccardi, Larry Davis 0001
Mach. Vis. Appl.4
2014 Training Initialization of Hidden Markov Models in Human Action Recognition
abstract
Human action recognition in video is often approached by means of sequential probabilistic models as they offer a natural match to the temporal dimension of the actions. However, effective estimation of the models' parameters is critical if one wants to achieve significant recognition accuracy. Parameter estimation is typically performed over a set of training data by maximizing objective functions such as the data likelihood or the conditional likelihood. However, such functions are nonconvex in nature and subject to local maxima. This problem is major since any solution algorithm (expectation-maximization, gradient ascent, variational methods and others) requires an arbitrary initialization and can only find a corresponding local maximum. Exhaustive search is otherwise impossible since the number of local maxima is unknown. While no theoretical solutions are available for this problem, the only practicable mollification is to repeat training with different initializations until satisfactory cross-validation accuracy is attained. Such a process is overall empirical and highly time-consuming. In this paper, we propose two methods for one-off initialization of hidden Markov models achieving interesting tradeoffs between accuracy and training time. Experiments over three challenging human action video datasets (Weizmann, MuHAVi and Hollywood Human Actions) and with various feature sets measured from the frames (STIP descriptors, projection histograms, notable contour points) prove that the proposed one-off initializations are capable of achieving accuracy above the average of repeated random initializations and comparable to the best. In addition, the methods proposed are not restricted solely to human action recognition as they suit time series classification as a general problem.
Zia Moghaddam, Massimo Piccardi
IEEE Trans Autom. Sci. Eng.2
2013 Towards simultaneous place classification and object detection based on conditional random field with multiple cues
abstract
Simultaneous place classification and object detection (SPCOD) is an algorithm which is able to categorize the environment (place) and detect the objects presented in the environment. Although both place classification and object detection problems have been in discussion in literature, as a concept SPCOD is still in its early stage of research. Focusing mainly on the discrimination ability of SPCOD, in this paper we have proposed a pairwise conditional random field (CRF) framework to integrate mature techniques on laser data based place classification and vision based off-the-shelf object descriptor. Extensive experimental results on a public data set demonstrate the effectiveness of the proposed method.
Lei Shi 0013, Sarath Kodagoda, Massimo Piccardi
IROS3
2013 Discriminative prototype selection methods for graph embedding
Ehsan Zare Borzeshi, Massimo Piccardi, Kaspar Riesen, Horst Bunke
Pattern Recognit.2
2013 Joint Action Segmentation and Classification by an Extended Hidden Markov Model
abstract
Hidden Markov models (HMMs) provide joint segmentation and classification of sequential data by efficient inference algorithms and have therefore been employed in fields as diverse as speech recognition, document processing, and genomics. However, conventional HMMs do not suit action segmentation in video due to the nature of the measurements which are often irregular in space and time, high dimensional and affected by outliers. For this reason, in this paper we present a joint action segmentation and classification approach based on an extended model: the hidden Markov model for multiple, irregular observations (HMM-MIO). Experiments performed over a concatenated version of the popular KTH action dataset and the challenging CMU multi-modal activity dataset (CMU-MMAC) report accuracies comparable to or higher than those of a bag-of-features approach, showing the usefulness of improved sequential models for joint action segmentation and classification tasks.
Ehsan Zare Borzeshi, Óscar Pérez, Massimo Piccardi
IEEE Signal Process. Lett.4
2011 Robust density modelling using the student's t-distribution for human action recognition
abstract
The extraction of human features from videos is often inaccurate and prone to outliers. Such outliers can severely affect density modelling when the Gaussian distribution is used as the model since it is highly sensitive to outliers. The Gaussian distribution is also often used as base component of graphical models for recognising human actions in the videos (hidden Markov model and others) and the presence of outliers can significantly affect the recognition accuracy. In contrast, the Student's t-distribution is more robust to outliers and can be exploited to improve the recognition rate in the presence of abnormal data. In this paper, we present an HMM which uses mixtures of t-distributions as observation probabilities and show how experiments over two well-known datasets (Weizmann, MuHAVi) reported a remarkable improvement in classification accuracy.
Zia Moghaddam, Massimo Piccardi
ICIP2
2011 MLiT: mixtures of Gaussians under linear transformations
Ahmed Fawzi Otoom, Hatice Gunes, Óscar Pérez, Massimo Piccardi
Pattern Anal. Appl.4
2010 Histogram-Based Training Initialisation of Hidden Markov Models for Human Action Recognition
abstract
Human action recognition is often addressed by use of latent-state models such as the hidden Markov model and similar graphical models. As such models require Expectation-Maximisation training, arbitrary choices must be made for training initialisation, with major impact on the final recognition accuracy. In this paper, we propose a histogram-based deterministic initialisation and compare it with both random and a time-based deterministic initialisations. Experiments on a human action dataset show that the accuracy of the proposed method proved higher than that of the other tested methods.
Zia Moghaddam, Massimo Piccardi
AVSS2
2009 Mixtures of Normalized Linear Projections
Ahmed Fawzi Otoom, Óscar Pérez, Hatice Gunes, Massimo Piccardi
ACIVS4
2009 An efficient Bayesian framework for on-line action recognition
abstract
On-line action recognition from a continuous stream of actions is still an open problem with fewer solutions proposed compared to time-segmented action recognition. The most challenging task is to classify the current action while finding its time boundaries at the same time. In this paper we propose an approach capable of performing on-line action segmentation and recognition by means of batteries of HMM taking into account all the possible time boundaries and action classes. A suitable Bayesian normalization is applied to make observation sequences of different length comparable and computational optimizations are introduce to achieve real-time performances. Results on a well known action dataset prove the efficacy of the proposed method.
Roberto Vezzani, Massimo Piccardi, Rita Cucchiara
ICIP2
2009 Head detection for video surveillance based on categorical hair and skin colour models
abstract
We propose a new robust head detection algorithm that is capable of handling significantly different conditions in terms of viewpoint, tilt angle, scale and resolution. To this aim, we built a new model for the head based on appearance distributions and shape constraints. We construct a categorical model for hair and skin, separately, and train the models for four categories of hair (brown, red, blond and black) and three categories of skin representing the different illumination conditions (bright, standard and dark). The shape constraint fits an elliptical model to the candidate region and compares its parameters with priors based on human anatomy. The experimental results validate the usability of the proposed algorithm in various video surveillance and multimedia applications.
Zui Zhang, Hatice Gunes, Massimo Piccardi
ICIP3
2009 Automatic Temporal Segment Detection and Affect Recognition From Face and Body Display
abstract
Psychologists have long explored mechanisms with which humans recognize other humans' affective states from modalities, such as voice and face display. This exploration has led to the identification of the main mechanisms, including the important role played in the recognition process by the modalities' dynamics. Constrained by the human physiology, the temporal evolution of a modality appears to be well approximated by a sequence of temporal segments called onset, apex, and offset. Stemming from these findings, computer scientists, over the past 15 years, have proposed various methodologies to automate the recognition process. We note, however, two main limitations to date. The first is that much of the past research has focused on affect recognition from single modalities. The second is that even the few multimodal systems have not paid sufficient attention to the modalities' dynamics: The automatic determination of their temporal segments, their synchronization to the purpose of modality fusion, and their role in affect recognition are yet to be adequately explored. To address this issue, this paper focuses on affective face and body display, proposes a method to automatically detect their temporal segments or phases, explores whether the detection of the temporal phases can effectively support recognition of affective states, and recognizes affective states based on phase synchronization/alignment. The experimental results obtained show the following: 1) affective face and body displays are simultaneous but not strictly synchronous; 2) explicit detection of the temporal phases can improve the accuracy of affect recognition; 3) recognition from fused face and body modalities performs better than that from the face or the body modality alone; and 4) synchronized feature-level fusion achieves better performance than decision-level fusion.
Hatice Gunes, Massimo Piccardi
IEEE Trans. Syst. Man Cybern. Part B2
2008 Tracking People in Crowds by a Part Matching Approach
abstract
The major difficulty in human tracking is the problem raised by challenging occlusions where the target person is repeatedly and extensively occluded by either the background or another moving object. These types of occlusions may cause significant changes in the person¿s shape, appearance or motion, thus making the data association problem extremely difficult to solve. Unlike most of the existing methods for human tracking that handle occlusions by data association of the complete human body, in this paper we propose a method that tracks people under challenging spatial occlusions based on body part tracking. The human model we propose consists of five body parts with six degrees of freedom and each part is represented by a rich set of features. The tracking is solved using a layered data association approach, direct comparison between features (feature layer) and subsequently matching between parts of the same bodies (part layer) lead to a final decision for the global match (global layer). Experimental results have confirmed the effectiveness of the proposed method.
Zui Zhang, Hatice Gunes, Massimo Piccardi
AVSS3
2008 Feature extraction techniques for abandoned object classification in video surveillance
abstract
We address the problem of abandoned object classification in video surveillance. Our aim is to determine (i) which feature extraction technique proves more useful for accurate object classification in a video surveillance context (scale invariant image transform (SIFT) keypoints vs. geometric primitive features), and (ii) how the resulting features affect classification accuracy and false positive rates for different classification schemes used. Objects are classified into four different categories: bag (s), person (s), trolley (s), and group (s) of people. Our experimental results show that the highest recognition accuracy and the lowest false alarm rate are achieved by building a classifier based on our proposed set of statistics of geometric primitives' features. Moreover, classification performance based on this set of features proves to be more invariant across different learning algorithms.
Ahmed Fawzi Otoom, Hatice Gunes, Massimo Piccardi
ICIP3
2008 An accurate algorithm for head detection based on XYZ and HSV hair and skin color models
abstract
Head detection in images and videos plays an important role in a wide range of computer vision and multimedia applications. In this paper, we propose a new head detection algorithm that is capable of handling significantly variable conditions in terms of viewpoint (i.e. frontal, profile, back view, from −180 degrees to +180 degrees), tilt angle (i.e. from horizontal to aerial), scale and resolution. To this aim, we built a new model for the head based on appearance distributions and shape constraints. The appearance distribution models the colors of hair and skin by sets of Gaussian mixtures in the XYZ and HSV color spaces. The shape constraint fits an elliptical model to the candidate region and compares its parameters with priors based on the human anatomy. This presents a pixel-level measurement of accuracy for the proposed algorithm both prior and after applying the spatial constraints referenced by the elliptical model. The excellent accuracy at both levels confirms the accuracy of the appearance model and the appropriateness of the spatial and topological process.
Zui Zhang, Hatice Gunes, Massimo Piccardi
ICIP3
2008 Maximum-likelihood dimensionality reduction in gaussian mixture models with an application to object classification
abstract
Accurate classification of objects of interest for video surveillance is difficult due to occlusions, deformations and variable views/illumination. The adopted feature sets tend to overcome these issues by including many and complementary features; however, their large dimensionality poses an intrinsic challenge to the classification task. In this paper, we present a novel technique providing maximum-likelihood dimensionality reduction in Gaussian mixture models for classification. The technique, called hereafter mixture of maximum-likelihood normalized projections (mixture of ML-NP), was used in this work to classify a 44-dimensional data set into 4 classes (bag, trolley, single person, group of people). The accuracy achieved on an independent test set is 98% vs. 80% of the runner-up (MultiBoost/AdaBoost).
Massimo Piccardi, Hatice Gunes, Ahmed Fawzi Otoom
ICPR1
2008 Comparative performance analysis of feature sets for abandoned object classification
abstract
Accurate classification of abandoned objects is crucial in video surveillance systems. In this paper, we experiment with different validation techniques (hold-out and 10-fold cross validation), with the aim of determining which feature set proves more useful for accurate object classification in a video surveillance context (scale invariant image transform (SIFT) keypoints vs. geometric primitive features). Moreover, we show how the resulting features affect classification performance across different classifiers. We also further analyze the best performing classifier in order to have better understanding of its classification results. Objects are classified into four different categories: bag (s), person (s), trolley (s), and group (s) of people. Our experimental results show that the highest recognition accuracy and the lowest false alarm rate are achieved by building a classifier based on our proposed set of statistics of geometric primitives' features. This set of features maximizes inter-class separation and simplifies the classification process. Classification based on this set of features thus outperforms the second best approach based on SIFT keypoint histograms by providing on average 22% higher recognition accuracy and 7% lower false alarm rate.
Ahmed Fawzi Otoom, Hatice Gunes, Massimo Piccardi
SMC3
2007 A framework for track matching across disjoint cameras using robust shape and appearance features
abstract
This paper presents a framework based on robust shape and appearance features for matching the various tracks generated by a single individual moving within a surveillance system. Each track is first automatically analysed in order to detect and remove the frames affected by large segmentation errors and drastic changes in illumination. The object's features computed over the remaining frames prove more robust and capable of supporting correct matching of tracks even in the case of significantly disjointed camera views. The shape and appearance features used include a height estimate as well as illumination-tolerant colour representation of the individual's global colours and the colours of the upper and lower portions of clothing. The results of a test from a real surveillance system show that the combination of these four features can provide a probability of matching as high as 91 percent with 5 percent probability of false alarms under views which have significantly differing illumination levels and suffer from significant segmentation errors in as many as 1 in 4 frames.
Christopher S. Madden, Massimo Piccardi
AVSS2
2007 Hidden Markov Models with Kernel Density Estimation of Emission Probabilities and their Use in Activity Recognition
abstract
In this paper, we present a modified hidden Markov model with emission probabilities modelled by kernel density estimation and its use for activity recognition in videos. In the proposed approach, kernel density estimation of the emission probabilities is operated simultaneously with that of all the other model parameters by an adapted Baum-Welch algorithm. This allows us to retain maximum-likelihood estimation while overcoming the known limitations of mixture of Gaussians in modelling certain probability distributions. Experiments on activity recognition have been performed on ground-truthed data from the CAVIAR video surveillance database and reported in the paper. The error on the training and validation sets with kernel density estimation remains around 14-16% while for the conventional Gaussian mixture approach varies between 15 and 24%, strongly depending on the initial values chosen for the parameters. Overall, kernel density estimation proves capable of providing more flexible modelling of the emission probabilities and, unlike Gaussian mixtures, does not suffer from being highly parametric and of difficult initialisation.
Massimo Piccardi, Óscar Pérez
CVPR1
2007 Bi-modal emotion recognition from expressive face and body gestures
Hatice Gunes, Massimo Piccardi
J. Netw. Comput. Appl.2
2007 Tracking people across disjoint camera views by an illumination-tolerant appearance representation
Christopher S. Madden, Eric Dahai Cheng, Massimo Piccardi
Mach. Vis. Appl.3
2006 Mitigating the Effects of Variable Illumination for Tracking across Disjoint Camera Views
abstract
Tracking people by their appearance across disjoint camera views is challenging since appearance may vary significantly across such views. This problem has been tackled in the past by computing intensity transfer functions between each camera pair during an initial training stage. However, in real-life situations, intensity transfer functions depend not only on the camera pair, but also on the actual illumination at pixel-wise resolution and may prove impractical to estimate to a satisfactory extent. For this reason, in this paper we propose an appearance representation for people tracking capable of coping with the typical illumination changes occurring in a surveillance scenario. Our appearance representation is based on an online K-means color clustering algorithm, a fixed, data-dependent intensity transformation, and the incremental use of frames. Moreover, a similarity measurement is proposed to match the appearance representations of any two given moving objects along sequences of frames. Experimental results presented in this paper show that the proposed methods provides a viable while effective approach for tracking people across disjoint camera views in typical surveillance scenarios.
Eric Dahai Cheng, Christopher S. Madden, Massimo Piccardi
AVSS3
2006 Video Surveillance at the Beginning of the Third Millennium: The Viewpoint of Research, Industry, Government Bodies, Research Funding Agencies and the Community
Massimo Piccardi
AVSS1
2006 Matching of Objects Moving Across Disjoint Cameras
abstract
Matching of single individuals as they move across disjoint camera views is a challenging task in video surveillance. In this paper, we present a novel algorithm capable of matching single individuals in such a scenario based on appearance features. In order to reduce the variable illumination effects in a typical disjoint camera environment, a cumulative color histogram transformation is first applied to the segmented moving object. Then, an incremental major color spectrum histogram representation (IMCSHR) is used to represent the appearance of a moving object and cope with small pose changes occurring along the track. An IMCHSR-based similarity measurement algorithm is also proposed to measure the similarity of any two segmented moving objects. A final step of post-matching integration along the object's track is eventually applied. Experimental results show that the proposed approach proved capable of providing correct matching in typical situations.
Eric Dahai Cheng, Massimo Piccardi
ICIP2
2006 Creating and Annotating Affect Databases from Face and Body Display: A Contemporary Survey
abstract
Databases containing representative samples of human multi-modal expressive behavior are needed for the development of affect recognition systems. However, at present publicly-available databases exist mainly for single expressive modalities such as facial expressions, static and dynamic hand postures, and dynamic hand gestures. Only recently, a first bimodal affect database consisting of expressive face and upper-body display has been released. To foster development of affect recognition systems, this paper presents a comprehensive survey of the current state-of-the art in affect database creation from face and body display and elicits the requirements of an ideal multi-modal affect database.
Hatice Gunes, Massimo Piccardi
SMC2
2006 Assessing facial beauty through proportion analysis by image processing and supervised learning
Hatice Gunes, Massimo Piccardi
Int. J. Hum. Comput. Stud.2
2005 Fusing Face and Body Display for Bi-modal Emotion Recognition: Single Frame Analysis and Multi-frame Post Integration
Hatice Gunes, Massimo Piccardi
ACII2
2005 Track matching over disjoint camera views based on an incremental major color spectrum histogram
abstract
Matching tracks from a single individual across disjoint camera views is a challenging task in video surveillance. In this paper, a major color spectrum histogram representation (MCSHR) is introduced to represent a moving object by using a normalized distance between two points in the RGB space. Then, an incremental MCSHR is proposed to cope with small pose changes and segmentation errors occurring along the track. Finally, a similarity measurement algorithm is proposed based on the incremental MCSHR to measure the similarity of any two tracked moving objects. The proposed similarity measurement algorithm proved capable of measuring the similarity of the two moving objects accurately. Experimental results show that with three to five frames integration, the proposed incremental MCSHR algorithm can make matching more robust and reliable than single-frame matching, especially for small pose changes. The matching performance is not obviously improved instead when the number of integration is more than five. The similarity of a same moving object in two different tracks has been improved from 92% to 95% with the integration number increased from three to five, while two different moving objects have been easily discriminated. The proposed algorithm can be used to match tracks from single individuals in camera networks, which do not provide full coverage of the monitored space.
Massimo Piccardi, Eric Dahai Cheng
AVSS1
2005 An FPGA generator for multipoint distributed random variables (abstract only)
abstract
Multi-point distributed random variables whose moments match those of a Gaussian random variable up to a certain order play an important role in Monte Carlo simulations of weak approximations of stochastic differential equations. In applications such as finance, where real time execution is required, there is a strong need for highly efficient implementations. In this paper a fast and flexible dedicated hardware solution on a Field Programmable Gate Array (FPGA) is presented. A comparative performance analysis between a software-only and the proposed hardware solution demonstrates that the FPGA solution is bottleneck-free, retains the flexibility of the software solution and significantly increases the computational efficiency.
Nicola Bruti Liberati, Eckhard Platen, Filippo Martini, Massimo Piccardi
FPGA4
2005 Affect recognition from face and body: early fusion vs. late fusion
abstract
This paper presents an approach to automatic visual emotion recognition from two modalities: face and body. Firstly, individual classifiers are trained from individual modalities. Secondly, we fuse facial expression and affective body gesture information first at a feature-level, in which the data from both modalities are combined before classification, and later at a decision-level, in which we integrate the outputs of the monomodal systems by the use of suitable criteria. We then evaluate these two fusion approaches, in terms of performance over monomodal emotion recognition based on facial expression modality only. In the experiments performed the emotion classification using the two modalities achieved a better recognition accuracy outperforming the classification using the individual facial modality. Moreover, fusion at the feature-level proved better recognition than fusion at the decision-level.
Hatice Gunes, Massimo Piccardi
SMC2
2004 Mean-shift background image modelling
abstract
Background modelling is widely used in computer vision for the detection of foreground objects in a frame sequence. The more accurate the background model, the more correct is the detection of the foreground objects. In this paper, we present an approach to background modelling based on a mean-shift procedure. The mean shift vector convergence properties enable the system to achieve reliable background modelling. In addition, histogram-based computation and the new concept of local basins of attraction allow us to meet the stringent real-time requirements of video processing.
Massimo Piccardi, Tony Jan
ICIP1
2004 Automated classification of female facial beauty by image analysis and supervised learning
abstract
The fact that perception of facial beauty may be a universal concept has long been debated amongst psychologists and anthropologists. In this paper, we performed experiments to evaluate the extent of beauty universality by asking a number of diverse human referees to grade a same collection of female facial images. Results obtained show that the different individuals gave similar votes, thus well supporting the concept of beauty universality. We then trained an automated classifier using the human votes as the ground truth and used it to classify an independent test set of facial images. The high accuracy achieved proves that this classifier can be used as a general, automated tool for objective classification of female facial beauty. Potential applications exist in the entertainment industry and plastic surgery.
Hatice Gunes, Massimo Piccardi, Tony Jan
VCIP2
2004 Neighbor cache prefetching for multimedia image and video processing
abstract
Cache performance is strongly influenced by the type of locality embodied in programs. In particular, multimedia programs handling images and videos are characterized by a bidimensional spatial locality, which is not adequately exploited by standard caches. In this paper we propose novel cache prefetching techniques for image data, called neighbor prefetching, able to improve exploitation of bidimensional spatial locality. A performance comparison is provided against other assessed prefetching techniques on a multimedia workload (with MPEG-2 and MPEG-4 decoding, image processing, and visual object segmentation), including a detailed evaluation of both the miss rate and the memory access time. Results prove that neighbor prefetching achieves a significant reduction in the time due to delayed memory cycles (more than 97% on MPEG-4 with respect to 75% of the second performing technique). This reduction leads to a substantial speedup on the overall memory access time (up to 140% for MPEG-4). Performance has been measured with the PRIMA trace-driven simulator, specifically devised to support cache prefetching.
Rita Cucchiara, Massimo Piccardi, Andrea Prati 0001
IEEE Trans. Multim.2
2003 Improving Data Prefetching Efficacy in Multimedia Applications
Rita Cucchiara, Andrea Prati 0001, Massimo Piccardi
Multim. Tools Appl.3
2003 Detecting Moving Objects, Ghosts, and Shadows in Video Streams
abstract
Background subtraction methods are widely exploited for moving object detection in videos in many applications, such as traffic monitoring, human motion capture, and video surveillance. How to correctly and efficiently model and update the background model and how to deal with shadows are two of the most distinguishing and challenging aspects of such approaches. The article proposes a general-purpose method that combines statistical assumptions with the object-level knowledge of moving objects, apparent objects (ghosts), and shadows acquired in the processing of the previous frames. Pixels belonging to moving objects, ghosts, and shadows are processed differently in order to supply an object-based selective update. The proposed approach exploits color information for both background subtraction and shadow detection to improve object segmentation and background update. The approach proves fast, flexible, and precise in terms of both pixel accuracy and reactivity to background changes.
Rita Cucchiara, Costantino Grana, Massimo Piccardi, Andrea Prati 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2001 An application of machine learning and statistics to defect detection
Rita Cucchiara, Paola Mello, Massimo Piccardi, Fabrizio Riguzzi
Intell. Data Anal.3
2000 Focus based Feature Extraction for Pallets Recognition
abstract
Visual recognition for object grasping is a well-known challenge for robot automation in industrial applications. A typical example is pallet recognition in industrial environment for pick-and-place automated process. The aim of vision and reasoning algorithms is to help robots in choosing the best pallets holes location. This work proposes an application-based approach, which full all requirements, dealing with every kind of occlusions and light situ-ations possible. Even some meaning noise (or meaning misunderstand-ing) is considered. A pallet model, with limited degrees of freedom, is de-scribed and, starting from it, a complete approach to pallet recognition is out-lined. In the model we dene both virtual and real corners, that are geomet-rical object proprieties computed by different image analysis operators. Real corners are perceived by processing brightness information directly from the image, while virtual corners are inferred at a higher level of abstraction. A nal reasoning stage selects the best solution tting the model. Experimental results and performance are reported in order to demonstrate the suitability of the proposed approach. 1
Rita Cucchiara, Massimo Piccardi, Andrea Prati 0001
BMVC2
2000 Image analysis and rule-based reasoning for a traffic monitoring system
abstract
The paper presents an approach for detecting vehicles in urban traffic scenes by means of rule-based reasoning on visual data. The strength of the approach is its formal separation between the low-level image processing modules and the high-level module, which provides a general-purpose knowledge-based framework for tracking vehicles in the scene. The image-processing modules extract visual data from the scene by spatio-temporal analysis during daytime, and by morphological analysis of headlights at night. The high-level module is designed as a forward chaining production rule system, working on symbolic data, i.e., vehicles and their attributes (area, pattern, direction, and others) and exploiting a set of heuristic rules tuned to urban traffic conditions. The synergy between the artificial intelligence techniques of the high-level and the low-level image analysis techniques provides the system with flexibility and robustness.
Rita Cucchiara, Massimo Piccardi, Paola Mello
IEEE Trans. Intell. Transp. Syst.2
1999 Constraint Propagation and Value Acquisition: Why we should do it Interactively
Evelina Lamma, Paola Mello, Michela Milano, Rita Cucchiara, Marco Gavanelli, Massimo Piccardi
IJCAI6
1999 Eliciting visual primitives for detection of elongated shapes
Rita Cucchiara, Massimo Piccardi
Image Vis. Comput.2
1998 Exploiting image processing locality in cache pre-fetching
abstract
Emerging trends in computer design attempt to include specific solutions for handling images also in general-purpose computers, because of the current spread of multimedia, image processing and computer graphics applications. In this context, we propose hardware pre-fetching techniques specific for caching images: the main issue we state is that most algorithms working on images exhibit a 2D spatial locality that is not taken into account in current cache organization and data access strategies. To this aim we propose an adaptive local pre-fetching for the image data type; this technique, mirroring the two-dimensional spatial locality of image processing algorithms, results in being more efficient than other approaches, such as sequential pre-fetching and adaptive pre-fetching. Performance is evaluated on different classes of image processing algorithms, namely raster-scan and propagative algorithms, common in computer vision and multimedia applications.
Rita Cucchiara, Massimo Piccardi
HiPC2
1998 A real-time hardware implementation of the hough transform
Rita Cucchiara, Giovanni Neri, Massimo Piccardi
J. Syst. Archit.3
1997 Exploiting Symbolic Learning in Visual Inspection
Massimo Piccardi, Rita Cucchiara, Michele Bariani, Paola Mello
IDA1
1997 Block processing on multiprocessor DSPs for multimedia applications
abstract
The paper explores software development for multiprocessor DSPs for data parallel local algorithms. These algorithms are very common in multimedia applications, such as filtering, compression and so on. Multiprocessor DSPs are very attractive for this application since they offer performance typical of parallel machines together with limited cost. The paper provides performance analysis and software design issues according to different data partitioning models. As a case study, performance evaluation has been carried out on the Multimedia Video Processor from Texas Instrument.
Rita Cucchiara, Alessandro Callipo, Massimo Piccardi
MMSP3
1996 Detection of luminosity profiles of elongated shapes
abstract
A novel technique for identifying elongated shapes in grey-scale images is presented. The method provides the detection and identification of elongated shapes not only modelling their principal direction, but also reconstructing the transversal luminosity profile. The approach is proposed starting from the gradient-weighted Hough transform, endowing the Hough space with more complete information about the luminance gradient of the image. This paper presents and discusses the algorithm devised to implement the method on discretized data. As an example of application, we present results on images from mechanical pieces, where real and false defects are discriminated through effective reconstruction of their luminosity profile.
Rita Cucchiara, Massimo Piccardi
ICIP (3)2
1995 Detection of Circular Objects by Wave Propagation on a Mesh-Connected Computer
Rita Cucchiara, Luigi Di Stefano, Massimo Piccardi
J. Parallel Distributed Comput.3
1993 Processing of variable size images on a cellular array: Performance analysis with the Abingdon Cross Benchmark
abstract
Handling a continuous flow of variable size images is a requirement for real time computer vision machines. A modular system based on a small size SIMD cellular array of 1-bit processing elements has been developed with this goal in mind and it is now evaluated against the Abingdon Cross Benchmark specifications. The benchmark tests the combination of algorithms and architecture and generates a quality factor expressed as the ratio of the image lateral size and the processing time. The examined machine supports an efficient means to automatically partition, process and reconstruct images larger than the array size. The authors briefly describe the system, discuss the selected algorithms and present performance results and estimates for several system configurations.>
Massimo Piccardi, Luigi Di Stefano, Rita Cucchiara, Tullio Salmon Cinotti
ASAP1