EDBT 2026 Demo / reviewers in the wild / expert
Massimo Tistarelli
dblp:44/1095
· DBLP profile ↗
76ranked-venue papers
22as first author
18since 2021 · last 2026
0000-0002-3406-3048ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 54 · 19 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 33 · 8 first-author · 7 since 2021Security and privacy · 8 · 1 first-author · 4 since 2021Systems, architecture and hardware · 5 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | 30 Years of Face Recognition Research: A Vision Ahead
Massimo Tistarelli |
ICPRAM | 1 |
| 2026 | Self-attention as the backbone: A survey on Vision TransformersabstractThe human cognitive system efficiently processes complex environments by focusing on salient regions, inspiring attention mechanisms that amplify critical information. Deep learning models integrate attention to enhance performance across tasks, with self-attention gaining prominence through its remarkable success in Transformer-based natural language processing. This success propelled its adoption in computer vision, where Transformer-based models surpass conventional deep networks on multiple benchmarks. Recently, large-scale pretraining has enabled vision foundation models, where self-attention acts as a scalable backbone for learning transferable visual representations across diverse tasks. By modeling global dependencies without heavy inductive biases, self-attention supports flexible and expressive representation learning. However, Vision Transformers (ViTs) rely on large parameter counts and high computational budgets, and their quadratic complexity with respect to input size restricts scalability and real-world deployment. Consequently, substantial research has focused on reducing computational overhead while preserving or improving performance through self-attention optimization. In parallel, State Space Models (SSMs) have emerged as an alternative paradigm, replacing attention with linear-time recurrence for efficient long-range modeling. This survey provides a comprehensive review of advanced backbone techniques that improve ViTs through novel self-attention designs or hybrid integrations, mainly for image classification, including their transfer to multimodal large language models, generative frameworks, and emerging SSM-based architectures, alongside ViT explainability approaches. We first outline self-attention and Transformers, then introduce a taxonomy of self-attention variants. Finally, we systematically compare attention architectures in terms of design, accuracy, and computational efficiency, and discuss current limitations and future research directions. Ahmad Waseem, Pietro Ruiu, Seth Nixon, Andrea Lagorio, Massimo Tistarelli |
Comput. Vis. Image Underst. | 5 |
| 2026 | Depth-induced bipolar neural collapse for privacy-preserving face verification
Yen-Lung Lai, Wun-She Yap, Bok-Min Goi, Zhe Jin 0001, Massimo Tistarelli |
Neural Networks | 5 |
| 2026 | HiTMM: Generative Temporal Masked Modeling of Human Interactive MotionsabstractWe have recently seen some progress in the current field of human-human interaction generation. However, directly generating complex two-person interactive motions remains a significant challenge. Meanwhile, these models typically employ two independent timelines when generating motions for interactive scenarios involving two individuals. This design overlooks the temporal dependencies between motions at each timestep and fails to account for the roles of active and reactive participants during the generation process, often resulting in unrealistic and unnatural motions. In this work, we propose HiTMM, a novel framework for Human interaction generation based on Temporal Masked Modeling. HiTMM first decomposes the human interaction into two separate single-person motions. Individual motions within the interaction belong to the same type, enabling them to be mapped to a shared latent space through a coarse-to-fine approach that produces multi-layer discrete tokens. We then arrange all tokens of the two interacting individuals along a shared timeline. Subsequently, we employ a masked transformer and a residual transformer to model the base-layer and rest-layer motion tokens. Both the base-layer and rest-layer motion tokens are arranged along a single timeline, allowing the model to explicitly capture the temporal order and initiating role embedded in the sequence, where the first individual's motion initiates the interaction. Note that, our model utilizes a shared temporal representation, making it capable of performing temporal editing on specific regions within human interaction sequences. Experimental results show that our model achieves an FID of 5.017 on the InterHuman dataset, surpassing the current state-of-the-art model (vs 5.154 for InterMask), and an FID of 0.373 on the InterX dataset (vs 0.399 for InterMask). Zicheng Jiao, Yunlian Sun, Hongwen Zhang 0001, Jinhui Tang 0001, Massimo Tistarelli |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | Uniss-MDF: A Multidimensional Face dataset for assessing face analysis on the moveabstractMultidimensional 2D-3D face analysis has demonstrated a strong potential for human identification in several application domains. The combined, synergic use of 2D and 3D data from human faces can counteract typical limitations in 2D face recognition, while improving both accuracy and robustness in identification. On the other hand, current mobile devices, often equipped with depth cameras and high performance computing resources, offer a powerful and practical tool to better investigate new models to jointly process real 2D and 3D face data. However, recent concerns related to privacy of individuals and the collection, storage and processing of personally identifiable biometric information have diminished the availability of public face recognition datasets. Uniss-MDF (Uniss-MultiDimensional Face) represents the first collection of combined 2D-3D data of human faces captured with a mobile device. Over 76,000 depth images and videos are captured from over 100 subjects, in both controlled and uncontrolled conditions, over two sessions. The features of Uniss-MDF are extensively compared with existing 2D-3D face datasets. The reported statistics underscore the value of the dataset as a versatile resource for researchers in face recognition on the move and for a wide range of applications. Notably, it is the sole 2D-3D facial dataset using data from a mobile device that includes both 2D and 3D synchronized sequences acquired in controlled and uncontrolled conditions. The Uniss-MDF dataset and the proposed experimental protocols with baseline results provide a new platform to compare processing models for novel research avenues in advanced face analysis on the move. Pietro Ruiu, Marinella Cadoni, Andrea Lagorio, Seth Nixon, Filippo Casu, Massimo Farina, Mauro Fadda, Giuseppe A. Trunfio, Massimo Tistarelli, Enrico Grosso |
Comput. Vis. Image Underst. | 9 |
| 2025 | Rethinking Contemporary Deep Learning Techniques for Error Correction in Biometric Data
Yen-Lung Lai, Xingbo Dong, Zhe Jin 0001, Wei Jia 0001, Massimo Tistarelli, Xuejun Li 0001 |
Int. J. Comput. Vis. | 5 |
| 2025 | IFViT: Interpretable Fixed-Length Representation for Fingerprint Matching via Vision TransformerabstractDetermining dense feature points on fingerprints used in constructing deep fixed-length representations for accurate matching, particularly at the pixel level, is of significant interest. To explore the interpretability of fingerprint matching, we propose a multi-stage interpretable fingerprint matching network, namely Interpretable Fixed-length Representation for Fingerprint Matching via Vision Transformer (IFViT), which consists of two primary modules. The first module, an interpretable dense registration module, establishes a Vision Transformer (ViT)-based Siamese Network to capture long-range dependencies and the global context in fingerprint pairs. It provides interpretable dense pixel-wise correspondences of feature points for fingerprint alignment and enhances the interpretability in the subsequent matching stage. The second module takes into account both local and global representations of the aligned fingerprint pair to achieve an interpretable fixed-length representation extraction and matching. It employs the ViTs trained in the first module with the additional fully connected layer and retrains them to simultaneously produce the discriminative fixed-length representation and interpretable dense pixel-wise correspondences of feature points. Extensive experimental results on diverse publicly available fingerprint databases demonstrate that the proposed framework not only exhibits superior performance on dense registration and matching but also significantly promotes the interpretability in deep fixed-length representations-based fingerprint matching. Honghui Chen, Xingbo Dong, Zheng Lin 0001, Iman Yi Liao, Massimo Tistarelli, Zhe Jin 0001 |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2024 | A Multi-Task Adversarial Attack against Face AuthenticationabstractDeep learning-based identity management systems, such as face authentication systems, are vulnerable to adversarial attacks. However, existing attacks are typically designed for single-task purposes, which means they are tailored to exploit vulnerabilities unique to the individual target rather than being adaptable for multiple users or systems. This limitation makes them unsuitable for certain attack scenarios, such as morphing, universal, transferable, and counterattacks. In this article, we propose a multi-task adversarial attack algorithm called MTADV that are adaptable for multiple users or systems. By interpreting these scenarios as multi-task attacks, MTADV is applicable to both single- and multi-task attacks, and feasible in the white- and gray-box settings. Furthermore, MTADV is effective against various face datasets, including LFW, CelebA, and CelebA-HQ, and can work with different deep learning models, such as FaceNet, InsightFace, and CurricularFace. Importantly, MTADV retains its feasibility as a single-task attack targeting a single user/system. To the best of our knowledge, MTADV is the first adversarial attack method that can target all of the aforementioned scenarios in one algorithm. Hanrui Wang 0005, Shuo Wang 0012, Cunjian Chen, Massimo Tistarelli, Zhe Jin 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | ScratchHOI: Training Human-Object Interaction Detectors from ScratchabstractTransformer-based approaches have exhibited outstanding performances in the field of human-object interaction (HOI) detection. However, these approaches rely on underlying object detectors that have undergone large-scale pre-trainings on the ImageNet and MS-COCO dataset. This limits the potential of unique architectural designs and induces a learning bias, causing ineffective HOI representation learning. In this paper, we propose ScratchHOI, a transformer-based method for human-object interaction detection that can be trained from scratch, eliminating the need for pre-trained object detectors. ScratchHOI employs dynamic and static affinity-based feature aggregation for processing local and long-range visual information. Additional techniques are also employed to improve detection performance, such as dynamic and interactive anchor refinement for objects and interactions. Experiments on the HICO-Det dataset show that ScratchHOI achieves competitive performance against other state-of-the-art approaches over a variety of different evaluation measures. JunYi Lim, Vishnu Monn Baskaran, Joanne Mun-Yee Lim, Ricky Sutopo, Koksheik Wong, Massimo Tistarelli |
ICIP | 6 |
| 2023 | Editorial for special section 'ICPR 2020'
Massimo Tistarelli |
Pattern Recognit. Lett. | 1 |
| 2023 | Breaking Free From Entropy's Shackles: Cosine Distance-Sensitive Error Correction for Reliable Biometric CryptographyabstractBiometric cryptosystems present a promising avenue for secure authentication; however, the efficiency and security of such systems can be hindered by errors in biometric data. To address this challenge, existing systems employ error-correction codes, but often fail to consider the distribution of biometric sources, potentially leading to an underestimation of the system’s security. In response to this issue, we propose a novel algorithm pair, designated as ENCODE and DECODE, which facilitates direct codeword generation from biometric samples. Our approach accounts for the distribution of biometric sources, thereby providing a more accurate estimation of system security compared to traditional methods. Our proposed algorithm pair generates codewords that maintain interpretability and are sensitive to the cosine distance between original biometric samples. This similarity metric is particularly well-suited for high-dimensional data analysis and enables a precise assessment of system performance. We have rigorously established the correctness of our algorithm pair, and empirical results illustrate its efficacy in tolerating distance between codewords while preserving accuracy in cosine distance-sensitive contexts. This approach has the potential to significantly improve the efficiency and security of biometric cryptosystems, rendering them more appropriate for daily cryptographic applications. Yen-Lung Lai, Xingbo Dong, Zhe Jin 0001, Massimo Tistarelli, Wun-She Yap, Bok-Min Goi |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2023 | ERNet: An Efficient and Reliable Human-Object Interaction Detection NetworkabstractHuman-Object Interaction (HOI) detection recognizes how persons interact with objects, which is advantageous in autonomous systems such as self-driving vehicles and collaborative robots. However, current HOI detectors are often plagued by model inefficiency and unreliability when making a prediction, which consequently limits its potential for real-world scenarios. In this paper, we address these challenges by proposing ERNet, an end-to-end trainable convolutional-transformer network for HOI detection. The proposed model employs an efficient multi-scale deformable attention to effectively capture vital HOI features. We also put forward a novel detection attention module to adaptively generate semantically rich instance and interaction tokens. These tokens undergo pre-emptive detections to produce initial region and vector proposals that also serve as queries which enhances the feature refinement process in the transformer decoders. Several impactful enhancements are also applied to improve the HOI representation learning. Additionally, we utilize a predictive uncertainty estimation framework in the instance and interaction classification heads to quantify the uncertainty behind each prediction. By doing so, we can accurately and reliably predict HOIs even under challenging scenarios. Experiment results on the HICO-Det, V-COCO, and HOI-A datasets demonstrate that the proposed model achieves state-of-the-art performance in detection accuracy and training efficiency. Codes are publicly available at https://github.com/Monash-CyPhi-AI-Research-Lab/ernet. JunYi Lim, Vishnu Monn Baskaran, Joanne Mun-Yee Lim, Koksheik Wong, John See, Massimo Tistarelli |
IEEE Trans. Image Process. | 6 |
| 2022 | RBECA: A regularized Bi-partitioned entropy component analysis for human face recognition
Arindam Kar, Debapriya Banik, Debotosh Bhattacharjee, Massimo Tistarelli |
Expert Syst. Appl. | 4 |
| 2022 | Gender and ethnicity recognition based on visual attention-driven deep architectures
Souad Khellat-Kihel, Jawad Muhammad, Zhenan Sun, Massimo Tistarelli |
J. Vis. Commun. Image Represent. | 4 |
| 2022 | Alignment-Robust Cancelable Biometric Scheme for Iris VerificationabstractIn this paper, we propose a histogram of oriented gradient inspired cancelable biometrics - Random Augmented Histogram of Gradients (R∙HoG) for iris template protection. The proposed R∙HoG is built upon on two main components: 1) column vector random augmentation and 2) gradient orientation grouping mechanisms to transform the unaligned irisCode feature into the alignment-robust cancelable template. The alignment-robust property of the proposed R∙HoG enables the fast template comparison which is crucial for an efficient authentication process. Experiments were performed on CASIA-IrisV3-Internal and CASIA-IrisV4-Thousand datasets. The results demonstrate the proposed R∙HoG could achieve acceptable verification performance in both datasets. Other than that, the irreversibility and security properties are studied based on major security and privacy attacks in biometric system. Lastly, results from the benchmarking evaluation framework show the proposed method is satisfying the unlinkability property. Ming Jie Lee, Zhe Jin 0001, Shiuan-Ni Liang, Massimo Tistarelli |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2021 | Similarity-based Gray-box Adversarial Attack Against Deep Face RecognitionabstractThe majority of adversarial attack techniques perform well against deep face recognition when the full knowledge of the system is revealed (white-box). However, such techniques act unsuccessfully in the gray-box setting where the face templates are unknown to the attackers. In this work, we propose a similarity-based gray-box adversarial attack (SGADV) technique with a newly developed objective function. SGADV utilizes the dissimilarity score to produce the optimized adversarial example, i.e., similarity-based adversarial attack. This technique applies to both white-box and gray-box attacks against authentication systems that determine genuine or imposter users using the dissimilarity score. To validate the effectiveness of SGADV, we conduct extensive experiments on face datasets of LFW, CelebA, and CelebA-HQ against deep face recognition models of FaceNet and InsightFace in both white-box and gray-box settings. The results suggest that the proposed method significantly outperforms the existing adversarial attack techniques in the gray-box setting. We hence summarize that the similarity-base approaches to develop the adversarial example could satisfactorily cater to the gray-box attack scenarios for de-authentication. Hanrui Wang 0003, Shuo Wang 0012, Zhe Jin 0001, Yandan Wang, Cunjian Chen, Massimo Tistarelli |
FG | 6 |
| 2021 | Appearance-based passenger counting in cluttered scenes with lateral movement compensation
Ricky Sutopo, Joanne Mun-Yee Lim, Vishnu Monn Baskaran, Koksheik Wong, Massimo Tistarelli, Heng Fui Liau |
Neural Comput. Appl. | 5 |
| 2021 | Efficient Known-Sample Attack for Distance-Preserving Hashing Biometric Template Protection SchemesabstractThe rapid deployment of biometric authentication systems raises concern over user privacy and security. A biometric template protection scheme emerges as a solution to protect individual biometric templates stored in a database. Among all available protection schemes, a template protection scheme that relies on distance-preserving hashing has received much attention due to its simplicity and efficiency in offering privacy protection while archiving decent authentication performance. In this work, we introduce an efficient attack called known sample attack and demonstrate that most state-of-art template protection schemes that utilize distance-preserving hashing can be compromised in practice (within few seconds), especially when the output is significantly smaller than the original input sample size. These findings further motivated our subsequent work in proposing a secure authentication mechanism to resist such an attack with proper study over the distribution of the input samples. Furthermore, we conducted revocability, unlinkability analysis to demonstrate the satisfactory of general biometric template protection requirements; and showed the resistance of various security and privacy attacks, i.e., false acceptance attack, and attack via record multiplicity. Yen-Lung Lai, Zhe Jin 0001, Koksheik Wong, Massimo Tistarelli |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2020 | Cross-spectrum Face Recognition Using Subspace Projection HashingabstractCross-spectrum face recognition, e.g. visible to thermal matching, remains a challenging task due to the large variation originated from different domains. This paper proposed a subspace projection hashing (SPH) to enable the cross-spectrum face recognition task. The intrinsic idea behind SPH is to project the features from different domains onto a common subspace, where matching the faces from different domains can be accomplished. Notably, we proposed a new loss function that can (i) preserve both inter-domain and intra-domain similarity; (ii) regularize a scaled-up pairwise distance between hashed codes, to optimize projection matrix. Three datasets, Wiki, EURECOM VIS-TH paired face and TDFace are adopted to evaluate the proposed SPH. The experimental results indicate that the proposed SPH outperforms the original linear subspace ranking hashing (LSRH) in the benchmark dataset (Wiki) and demonstrates a reasonably good performance for visible-thermal, visible-near-infrared face recognition, therefore suggests the feasibility and effectiveness of the proposed SPH. Hanrui Wang 0003, Xingbo Dong, Zhe Jin 0001, Jean-Luc Dugelay, Massimo Tistarelli |
ICPR | 5 |
| 2020 | Quality-based Representation for Unconstrained Face RecognitionabstractSignificant advances have been achieved in face recognition in the last decade thanks to the development of deep learning methods. However, recognizing faces captured in uncontrolled environments is still a challenging problem for the scientific community. In these scenarios, the performance of most of existing deep learning based methods abruptly falls, due to the bad quality of the face images. In this work, we propose to use an activation map to represent the quality information in a face image. Different face regions are analyzed to determine their quality and then only those regions with good quality are used to perform the recognition using a given deep face model. For experimental evaluation, in order to simulate unconstrained environments, three challenging databases, with different variations in appearance, were selected: the Labeled Faces in the Wild Database, the Celebrities in Frontal-Profile in the Wild Database, and the AR Database. Three deep face models were used to evaluate the proposal on these databases and in all cases, the use of the proposed activation map allows the improvement of the recognition rates obtained by the original models in a range from 0.3 up to 31%. The obtained results experimentally demonstrated that the proposal is able to select those face areas with higher discriminative power and enough identifying information, while ignores the ones with spurious information. Nelson Méndez Llanes, Katy Castillo-Rosado, Heydi Mendez Vazquez, Massimo Tistarelli |
ICPR | 4 |
| 2020 | Facial Age Synthesis With Label Distribution-Guided Generative Adversarial NetworkabstractThe existing research work on facial age synthesis has been mostly focused on long-term aging (e.g., over an age span of 10 years or more). In this paper, we employ generative adversarial networks (GANs) as a tool to investigate age synthesis over different age spans. Compared with long-term aging, short-term age synthesis suffers from the reduced amount of available training data, which can severely hinder the model training. We conduct a series of experiments to validate this. To facilitate short-term age synthesis, we further propose label distribution-guided generative adversarial network (ldGAN), where each sample is associated with an age label distribution (ALD) rather than a single age group. Accordingly, each sample can contribute not only to the learning of its own age group but also to neighbouring groups' learning. This is useful when addressing short-term aging to cope with the reduced amount of training data. In addition, unlike one-hot encoding which treats age groups as independent from one another, ldGAN can well capture the correlation among different age groups, so that smooth aging sequences can be achieved. The ALD model is integrated into GAN with a two-step process. Firstly, instead of the traditional one-hot encoding, ALD is applied as the condition of the generator. Secondly, we add a sequence of label distribution learners on top of several multi-scale discriminators, with the aim of minimizing the label distribution learning loss when optimizing both the generator and discriminators. Both qualitative and quantitative evaluations are conducted to assess ldGAN's ability in dealing with two core issues of face aging, i.e., aging effect generation and identity preservation. The obtained experimental results demonstrate the effectiveness of ldGAN in both learning short-term aging patterns and coping with the lack of training data. Yunlian Sun, Jinhui Tang 0001, Xiangbo Shu, Zhenan Sun, Massimo Tistarelli |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2020 | Facial Age and Expression Synthesis Using Ordinal Ranking Adversarial NetworksabstractFacial image synthesis has been extensively studied, for a long time, in both computer graphics and computer vision. Particularly, the synthesis of face images with varying ages, expressions and poses has received an increasing attention owing to several real-world applications. In this paper, facial age and expression synthesis are addressed. While previous and current research papers on facial age synthesis mostly adopt an age span of 10 years, this paper investigates face aging with a shorter time span. For expression synthesis, given a neutral face, we work on synthesizing faces with varying expression intensities (e.g., from zero to high). Note that both human ages and expression intensities are inherently ordinal. To fully exploit this ordinal nature, we devise ordinal ranking generative adversarial networks (ranking GAN). For each face, a one-hot label is assigned to define its age range/expression intensity. By exploiting the relative order information among age ranges/expression intensities, a binary ranking vector is further computed for each face. In ranking GAN, one-hot labels are used as the condition of the generator for synthesizing faces with target age groups/expression intensities. Moreover, we add a sequence of cost-sensitive ordinal rankers on top of several multi-scale discriminators, with the aim of minimizing age/intensity rank estimation loss when optimizing both the generator and discriminators. In order to evaluate the proposed ranking GAN, extensive experiments are carried out on several public face databases. As demonstrated by the experimental testing, this ranking scheme performs well even when the amount of available labeled training data is limited. The reported experimental results well demonstrate the effectiveness of ranking GAN on synthesizing face aging sequences and faces with varying expression intensities. Yunlian Sun, Jinhui Tang 0001, Zhenan Sun, Massimo Tistarelli |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2019 | Foveated Vision for Deepface Recognition
Souad Khellat-Kihel, Andrea Lagorio, Massimo Tistarelli |
CIARP | 3 |
| 2019 | Face Recognition on Mobile Devices Based on Frames Selection
Nelson Méndez Llanes, Katy Castillo-Rosado, Heydi Mendez Vazquez, Souad Khellat-Kihel, Massimo Tistarelli |
CIARP | 5 |
| 2018 | On Combining Face Local Appearance and Geometrical Features for Race Classification
Fabiola Becerra-Riera, Nelson Méndez Llanes, Annette Morales-González, Heydi Mendez Vazquez, Massimo Tistarelli |
CIARP | 5 |
| 2018 | Context awareness in biometric systems and methods: State of the art and future scenarios
Michele Nappi, Stefano Ricciardi, Massimo Tistarelli |
Image Vis. Comput. | 3 |
| 2018 | Super-resolution for biometrics: A comprehensive survey
Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan, Massimo Tistarelli, Mark S. Nixon |
Pattern Recognit. | 4 |
| 2017 | Age and gender classification using local appearance descriptors from facial componentsabstractFace analysis and recognition systems have shown to be a valuable tool for forensic examiners. Particularly, the automatic estimation of age and gender from face images, can be useful in a wide range of forensic applications. In this work we propose to use a local appearance descriptor in a component-based way, to classify age and gender from face images. We subdivide a face image into regions of interest based on automatically detected landmarks, and represent them by using Histograms of Oriented Gradient (HOG). The representations obtained from different face regions are feeded to Support Vector Machine (SVM) classifiers to estimate the age and gender of the person in the image. Experimental analysis show the good results of this component-based approach, and its additional benefits when face images are affected by occlusions. Fabiola Becerra-Riera, Heydi Mendez Vazquez, Annette Morales-González, Massimo Tistarelli |
IJCB | 4 |
| 2017 | Special issue on Best of Biometrics 2015
Massimo Tistarelli, J. Ross Beveridge, Patrick J. Flynn, Michele Nappi |
Image Vis. Comput. | 1 |
| 2016 | 2D Recurrent Neural Networks for Robust Visual Tracking of Non-Rigid Bodies
Giovanni Luca Masala, Bruno Golosio, Massimo Tistarelli, Enrico Grosso |
EANN | 3 |
| 2016 | Deceiving faces: When plastic surgery challenges face recognition
Michele Nappi, Stefano Ricciardi, Massimo Tistarelli |
Image Vis. Comput. | 3 |
| 2015 | On soft biometrics
Mark S. Nixon, Paulo Lobato Correia, Kamal Nasrollahi, Thomas B. Moeslund, Abdenour Hadid, Massimo Tistarelli |
Pattern Recognit. Lett. | 6 |
| 2014 | On the Use of Discriminative Cohort Score Normalization for Unconstrained Face RecognitionabstractFacial imaging has been largely addressed for automatic personal identification, in a variety of different environments. However, automatic face recognition becomes very challenging whenever the acquisition conditions are unconstrained. In this paper, a picture-specific cohort normalization approach, based on polynomial regression, is proposed to enhance the robustness of face matching under challenging conditions. A careful analysis is presented to better understand the actual discriminative power of a given cohort set. In particular, it is shown that the cohort polynomial regression alone conveys some discriminative information on the matching face pair, which is just marginally worse than the raw matching score. The influence of the cohort set size in the matching accuracy is also investigated. Further, tests performed on the Face Recognition Grand Challenge ver 2 database and the labeled faces in the wild database allowed to determine the relation between the quality of the cohort samples and cohort normalization performance. Experimental results obtained from the LFW data set demonstrate the effectiveness of the proposed approach to improve the recognition accuracy in unconstrained face acquisition scenarios. Massimo Tistarelli, Yunlian Sun, Norman Poh |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2012 | Customizing biometric authentication systems via discriminative score calibrationabstractThere is mounting evidence about the benefit of tailoring a biometric authentication system to each user by postprocessing the system output at the score level, also known as client-specific score normalisation. Examples of these procedures are Z-norm and F-norm. These procedures can calibrate the uneven hypothesis space such that the dispropotionate false acceptance and false rejection errors are reduced after the calibration. The interest in studying these schemes is that they are applicable to any biometric authentication system regardless of the underlying biometric modality, and furthermore, potentially be extended to object recognition framed as a verification problem. We propose to further improve these procedures by adding additional client-specific terms that cannot be incorporated easily in their respective existing form. Experiments carried out on 13 face and speech systems show that both variants systematically outperform their respective score normalisation scheme (Z-norm or F-norm). Norman Poh, Massimo Tistarelli |
CVPR | 2 |
| 2011 | Automated quality control of printed flasks and bottles
Enrico Grosso, Andrea Lagorio, Massimo Tistarelli |
Mach. Vis. Appl. | 3 |
| 2010 | MCC: A baseline algorithm for fingerprint verification in FVC-onGoingabstractThis paper describes an improved version of the MCC fingerprint matching approach. An in-depth error analysis allowed us to point out the weakest points of the original MCC and to design: i) a more effective minutiae pair selection and ii) a more distortion-tolerant relaxation. The parameters of the new version have been tuned over a new larger dataset and the final algorithm has been evaluated on FVC-onGoing. The results show that MCC compares favorably with some of the most accurate commercial algorithms published in FVC-onGoing. Raffaele Cappelli, Matteo Ferrara, Davide Maltoni, Massimo Tistarelli |
ICARCV | 4 |
| 2010 | Sparse representations and Random Projections for robust and cancelable biometricsabstractIn recent years, the theories of Sparse Representation (SR) and Compressed Sensing (CS) have emerged as powerful tools for efficiently processing data in non-traditional ways. An area of promise for these theories is biométrie identification. In this paper, we review the role of sparse representation and CS for efficient biométrie identification. Algorithms to perform identification from face and iris data are reviewed. By applying Random Projections it is possible to purposively hide the biométrie data within a template. This procedure can be effectively employed for securing and protecting personal biométrie data against theft. Some of the most compelling challenges and issues that confront research in biometrics using sparse representations and CS are also addressed. Vishal M. Patel, Rama Chellappa, Massimo Tistarelli |
ICARCV | 3 |
| 2010 | The Multiscenario Multienvironment BioSecure Multimodal Database (BMDB)abstractA new multimodal biometric database designed and acquired within the framework of the European BioSecure Network of Excellence is presented. It is comprised of more than 600 individuals acquired simultaneously in three scenarios: 1) over the Internet, 2) in an office environment with desktop PC, and 3) in indoor/outdoor environments with mobile portable hardware. The three scenarios include a common part of audio/video data. Also, signature and fingerprint data have been acquired both with desktop PC and mobile portable hardware. Additionally, hand and iris data were acquired in the second scenario using desktop PC. Acquisition has been conducted by 11 European institutions. Additional features of the BioSecure Multimodal Database (BMDB) are: two acquisition sessions, several sensors in certain modalities, balanced gender and age distributions, multimodal realistic scenarios with simple and quick tasks per modality, cross-European diversity, availability of demographic data, and compatibility with other multimodal databases. The novel acquisition conditions of the BMDB allow us to perform new challenging research and evaluation of either monomodal or multimodal biometric systems, as in the recent BioSecure Multimodal Evaluation campaign. A description of this campaign including baseline results of individual modalities from the new database is also given. The database is expected to be available for research purposes through the BioSecure Association during 2008. Javier Ortega-Garcia, Julian Fierrez, Fernando Alonso-Fernandez, Javier Galbally, Manuel R. Freire, Joaquín González-Rodríguez, Carmen García-Mateo, José Luis Alba-Castro, Elisardo González-Agulla, Enrique Otero Muras, Sonia Garcia-Salicetti, Lorène Allano, Van-Bao Ly, Bernadette Dorizzi, Josef Kittler, Thirimachos Bourlai, Norman Poh, Farzin Deravi, Ming W. R. Ng, Michael C. Fairhurst, Jean Hennebert, Andreas Humm, Massimo Tistarelli, Linda Brodo, Jonas Richiardi, Andrzej Drygajlo, Harald Ganster, Federico Sukno, Sri-Kaushik Pavani, Alejandro F. Frangi, Lale Akarun, Arman Savran |
IEEE Trans. Pattern Anal. Mach. Intell. | 23 |
| 2009 | Image and vision computing journal special issue on multimodal biometrics
Massimo Tistarelli, Josef Bigün |
Image Vis. Comput. | 1 |
| 2009 | Dynamic face recognition: From human to machine vision
Massimo Tistarelli, Manuele Bicego, Enrico Grosso |
Image Vis. Comput. | 1 |
| 2008 | Automatic Detection of Adverse Weather Conditions in Traffic ScenesabstractVisual surveillance in outdoor environments requires the monitoring of both objects and events. The analysis is generally driven by the target application which, in turn, determines the set of relevant events and objects to be analyzed. In this paper we concentrate on the analysis of outdoor scenes, in particular for vehicle traffic control. In this scenario, the analysis of weather conditions is considered to signal particular and potentially dangerous situations like the presence of snow, fog, or heavy rain. The developed system uses a statistical framework based on the mixture of Gaussians to identify changes both in the spatial and temporal frequencies which characterize specific meteorological events. Several experiments performed on standard databases and real scenes demonstrate the applicability of the proposed approach. Andrea Lagorio, Enrico Grosso, Massimo Tistarelli |
AVSS | 3 |
| 2008 | Graph application on face for personal authentication and recognitionabstractThis paper presents a novel face recognition technique with graph topology drawn on scale invariant feature transform (SIFT) features and is compared with all the available well known techniques on SIFT features, and elastic bunch graph matching (EBGM) technique drawn on gabor wavelet feature. IITK face database is used for evaluation purpose. Test results show that the proposed graph matching technique will be an appropriate one for face recognition. Finally, the test results have been compared with the results found on BANCA face database following MC protocol only. Dakshina Ranjan Kisku, Ajita Rattani, Massimo Tistarelli, Phalguni Gupta |
ICARCV | 3 |
| 2008 | Robust fusion using boosting and transduction for component-based face recognitionabstractFace recognition performance depends upon the input variability as encountered during biometric data capture including occlusion and disguise. The challenge met in this paper is to expand the scope and utility of biometrics by discarding unwarranted assumptions regarding the completeness and quality of the data captured. Towards that end we propose a model-free and non-parametric component-based face recognition strategy with robust decisions for data fusion that are driven by transduction and boosting. The conceptual framework draws support throughout from discriminative methods using likelihood ratios. It links at the conceptual level forensics and biometrics, while at the implementation level it links the Bayesian framework and statistical learning theory (SLT). Feature selection of local patch instances and their corresponding high-order combinations, exemplar-based clustering (of patches) as components including the sharing (of exemplars) among components, and finally decision-making regarding authentication using boosting driven by components that play the role of weak-learners, are implemented in a similar fashion using transduction driven by a strangeness measure akin to typicality. The feasibility, reliability, and utility of the proposed open set face recognition architecture vis-à-vis adverse image capture conditions are illustrated using FRGC data. The potential for future developments concludes the paper. Fayin Li, Harry Wechsler, Massimo Tistarelli |
ICARCV | 3 |
| 2008 | Distinctiveness of faces: A computational approachabstractThis paper develops and demonstrates an original approach to face-image analysis based on identifying distinctive areas of each individual's face by its comparison to others in the population. The method differs from most others—that we refer as unary —where salient regions are defined by analyzing only images of the same individual. We extract a set of multiscale patches from each face image before projecting them into a common feature space. The degree of “distinctiveness” of any patch depends on its distance in feature space from patches mapped from other individuals. First a pairwise analysis is developed and then a simple generalization to the multiple-face case is proposed. A perceptual experiment, involving 45 observers, indicates the method to be fairly compatible with how humans mark faces as distinct. A quantitative example of face authentication is also performed in order to show the essential role played by the distinctive information. A comparative analysis shows that performance of our n-ary approach is as good as several contemporary unary, or binary, methods, while tapping a complementary source of information. Furthermore, we show it can also provide a useful degree of illumination invariance. Manuele Bicego, Enrico Grosso, Andrea Lagorio, Gavin Brelstaff, Linda Brodo, Massimo Tistarelli |
ACM Trans. Appl. Percept. | 6 |
| 2006 | Recognizing People's Faces: from Human to Machine VisionabstractAs confirmed by recent neurophysiological studies, the use of dynamic information is extremely important for humans in visual perception of biological forms and motion. Apart from the mere computation of the visual motion of the viewed objects, the motion itself conveys far more information, which helps understanding the scene. This paper provides an overview and some new insights on the use of dynamic visual information for face recognition. In this context, not only physical features emerge in the face representation, but also behavioral features should be accounted. While physical features are obtained from the subject's face appearance, behavioral features are obtained from the individual motion and articulation of the face. In order to capture both the face appearance and the face dynamics, a dynamical face model based on a combination of hidden Markov models is presented. The number of states (or facial expressions) are automatically determined from the data by unsupervised clustering of expressions of faces in the video. The underlying architecture closely recalls the neural patterns activated in the perception of moving faces. Preliminary results on real video image data show the feasibility of the proposed approach Massimo Tistarelli, Manuele Bicego, Enrico Grosso |
ICARCV | 1 |
| 2004 | Guest Editorial Introduction to the Special Issue on Image- and Video-Based Biometrics
Xiaoou Tang, Songde Ma, Lawrence O'Gorman, Massimo Tistarelli |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2004 | Guest Editorial: Introduction to the Special Issue on Image- and Video-Based Biometrics - Part II
Xiaoou Tang, Songde Ma, Lawrence O'Gorman, Massimo Tistarelli |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2003 | On testing methods for biometric authenticationabstractThe use of biometric data for user authentication and/or recognition is now a reality. On the other hand, there is still a strong need for new technologies to overpass intrinsic limitations of already "established" techniques. This not only requires to devise new algorithms but to determine the real potential and limitations of existing techniques. This is possible only devising standard testing and assessment procedures based on statistical observations of the outputs of the system. In order to define better standard evaluation process, a system based on space-variant iconic image matching is described and the validation procedure defined. It turns out that all methods based on the same biometric measurements have the same intrinsic limitations, which can be only overcome by the adoption of a multi-modal or multi-algorithmic approach. Enrico Grosso, Massimo Tistarelli |
ICME | 2 |
| 2002 | Techniques for iconic image-based biometricsabstractAlgorithms based on face images are quite appealing for the possibility to easily adapt and tailor a system to many application domains. A system for personal identity verification and also recognition is presented. The core engine is a standard correlation-based matcher performed on iconic representations of face images. Through standard statistical tests of the recognition results obtained from two different data sets the actual physical limits of the pattern matcher are clearly shown. Successively also other aspects are taken into account, related to the feature space, allowing us to greatly improve the system performance reaching almost 100% correct recognition. Andrea Lagorio, Enrico Grosso, Massimo Tistarelli |
ICME (2) | 3 |
| 2000 | Log-polar Stereo for Anthropomorphic Robots
Enrico Grosso, Massimo Tistarelli |
ECCV (1) | 2 |
| 2000 | Special issue on facial image analysis
Massimo Tistarelli |
Image Vis. Comput. | 1 |
| 2000 | Active vision-based face authentication
Massimo Tistarelli, Enrico Grosso |
Image Vis. Comput. | 1 |
| 1997 | Active face recognition with a hybrid approach
Massimo Tistarelli, Enrico Grosso |
Pattern Recognit. Lett. | 1 |
| 1996 | Multiple Constraints to Compute Optical FlowabstractThe computation of the optical flow field from an image sequence requires the definition of constraints on the temporal change of image features. In this paper, we consider the implications of using multiple constraints in the computational schema. In the first step, it is shown that differential constraints correspond to an implicit feature tracking. Therefore, the best results (either in terms of measurement accuracy, and speed in the computation) are obtained by selecting and applying the constraints which are best "tuned" to the particular image feature under consideration. Considering also multiple image points not only allows us to obtain a (locally) better estimate of the velocity field, but also to detect erroneous measurements due to discontinuities in the velocity field. Moreover, by hypothesizing a constant acceleration motion model, also the derivatives of the optical flow are computed. Several experiments are presented from real image sequences. Massimo Tistarelli |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1995 | Computation of Coherent Optical Flow by Using Multiple ConstraintsabstractThe optical flow constitutes one of the most widely adopted representations to define and characterize the evolution of image features over time. In order to compute the velocity field, it is necessary to define a set of constraints on the temporal change of image features. We consider the implications in using multiple constraints arising from multiple data points. The first step is the analysis of differential constraints and how they can be applied, locally, to compute the image velocity. This analysis allows to relate each constraint to a particular gray level pattern. This approach is extended to multiple image points, allowing also the characterization of the temporal behaviour of the image features and to detect erroneous measurements due to occlusions, depth discontinuities or shadows. Several experiments are presented from real image sequences. > Massimo Tistarelli |
ICCV | 1 |
| 1995 | Active/space-variant object recognition
Massimo Tistarelli |
Image Vis. Comput. | 1 |
| 1995 | Active/Dynamic Stereo VisionabstractVisual navigation is a challenging issue in automated robot control. In many robot applications, like object manipulation in hazardous environments or autonomous locomotion, it is necessary to automatically detect and avoid obstacles while planning a safe trajectory. In this context the detection of corridors of free space along the robot trajectory is a very important capability which requires nontrivial visual processing. In most cases it is possible to take advantage of the active control of the cameras. In this paper we propose a cooperative schema in which motion and stereo vision are used to infer scene structure and determine free space areas. Binocular disparity, computed on several stereo images over time, is combined with optical flow from the same sequence to obtain a relative-depth map of the scene. Both the time to impact and depth scaled by the distance of the camera from the fixation point in space are considered as good, relative measurements which are based on the viewer, but centered on the environment. The need for calibrated parameters is considerably reduced by using an active control strategy. The cameras track a point in space independently of the robot motion and the full rotation of the head, which includes the unknown robot motion, is derived from binocular image data. The feasibility of the approach in real robotic applications is demonstrated by several experiments performed on real image data acquired from an autonomous vehicle and a prototype camera head.> Enrico Grosso, Massimo Tistarelli |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1995 | Active/Dynamic Stereo Vision
Enrico Grosso, Massimo Tistarelli |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1994 | Recognition by using an active/space-variant sensorabstractThe problem of object recognition is addressed. In the literature this task has been generally considered in a "passive" perspective, where everything is static and there is no definite relation between the object and its environment. We propose an "active" approach for object recognition, based on the capability of the observer to move and give a better description of the object under consideration and also to take advantage of the relations between the objects and the environment. This can be accomplished at the task level and at the sensor level. The face recognition problem, based on the face-space approach, is considered to demonstrate the advantage of adopting an active retina to sample the face, build a database and perform the recognition task. By using an active space-variant retina the size of the database is considerably reduced and consequently the processing time for recognition. A comparative experiment using the active and static approach is presented.> Massimo Tistarelli |
CVPR | 1 |
| 1994 | Multiple Constraints for Optical Flow
Massimo Tistarelli |
ECCV (1) | 1 |
| 1994 | Application of optical flow for automated overtaking controlabstractIn this paper the problem of automated control of a road vehicle is addressed. In particular, a system is presented where optical flow techniques are applied to monitor overtaking maneuvers of other vehicles coming from the rear of the car. During car driving, most visual information is conveyed by the motion perceived within the visual field. The proposed approach for overtaking control is based on the computation of the optical flow (or the normal flow) from an image stream acquired from a camera sensor mounted on board of the vehicle. The method applies for a road vehicle and constitutes a building block of a software and hardware architecture devised to provide helpful information as an aid to the driver. Several experiments are presented from real image sequences.> Massimo Tistarelli, Francesco Guarnotta, Danilo Rizzieri, Federico Tarocchi |
WACV | 1 |
| 1993 | Active/dynamic stereo: a general frameworkabstractA general framework for dynamic stereo is proposed. It is shown that a binocular system, able to actively control the gaze of the cameras, can profitably exploit stereo and motion analysis to compute the relative orientation and time-to-impact with respect to the environment. The cooperative schema, in which motion and stereo vision are combined to take under control the temporal evolution of disparity, is particularly suitable for robot navigation or for tasks involving a robotic head mounted on a mobile platform. The framework takes into account tilt and pan motion of the cameras, and generic translation of the binocular system. An experiment performed on a stereo image sequence from a prototype robotic head is presented.> Enrico Grosso, Massimo Tistarelli |
CVPR | 2 |
| 1993 | On the Advantages of Polar and Log-Polar Mapping for Direct Estimation of Time-To-Impact from Optical FlowabstractThe application of an anthropomorphic retina-like visual sensor and the advantages of polar and log-polar mapping for visual navigation are investigated. It is demonstrated that the motion equations that relate the egomotion and/or the motion of the objects in the scene to the optical flow are considerably simplified if the velocity is represented in a polar or log-polar coordinate system, as opposed to a Cartesian representation. The analysis is conducted for tracking egomotion but is then generalized to arbitrary sensor and object motion. The main result stems from the abundance of equations that can be written directly that relate the polar or log-polar optical flow with the time to impact. Experiments performed on images acquired from real scenes are presented.> Massimo Tistarelli, Giulio Sandini |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1992 | Active/Dynamic Stereo for Navigation
Enrico Grosso, Massimo Tistarelli, Giulio Sandini |
ECCV | 2 |
| 1992 | Dynamic aspects in active vision
Massimo Tistarelli, Giulio Sandini |
CVGIP Image Underst. | 1 |
| 1991 | Dynamic stereo in visual navigationabstractA cooperative schema is proposed in which binocular disparity, computed on several stereo images over time, is combined with optical flow from the same sequence to obtain a relative-depth map of the scene. Both time-to-impact and depth scaled by the distance of the camera from the fixation point in space are considered as good, relative measurements which are based on the viewer (but centered on the environment). Two experiments, performed on image sequences from real scenes, are presented.> Massimo Tistarelli, Enrico Grosso, Giulio Sandini |
CVPR | 1 |
| 1991 | On the usefulness of optical flow for robotic part manipulationabstractThis paper describes experience gained in using optical flow fields, arising from constrained camera motion, to estimate range information, for robotic part manipulation. Several key issues in robot vision are addressed. These include the computation of optical flow, computation of depth, interpolation of depth values, model construction, object recognition, and pose estimation. The paper provides an example of the approach and evaluates the system's performance in the context of its applicability to part manipulation.> Kenneth M. Dawson, Massimo Tistarelli, David Vernon |
IROS | 2 |
| 1991 | Visual monitoring of robot actionsabstractPresent a perspective in relation to the use of vision for the control of robots and for the monitoring of robot actions. In order to explain the approach some experiments have been performed in the field of object manipulation and robot navigation. In particular the monitoring of pushing and tapping actions by means of dynamic measurements performed on image sequences are reported. The experiments show that, not only it seems possible to detect unexpected events, but it is possible to control the arm or vehicle trajectory in such a way to achieve a complex task like the pushing of an unknown object along a predetermined trajectory.> Francesca Gandolfo, Massimo Tistarelli, Giulio Sandini |
IROS | 2 |
| 1991 | Analysis of multidimensional images on the Connection Machine systemabstractAbstract The Connection Machine (CM) has been demonstrated to be an efficient and fast computational engine for the solution of many problems related to image processing. The high‐level parallelism of the CM naturally fits to many large‐scale data intensive applications. In this paper the implementation of parallel algorithms for the analysis of multidimensional images on the CM is presented. Different aspects in the analysis of multidimensional images are considered. In the field of artificial vision, the implementation of algorithms for the filtering of image sequences (both in space and time) and the estimation of the optical flow is described and some results in terms of accuracy and computation time are presented. The processing of three‐dimensional images is investigated in the field of biomedical engineering. In this case the goal is the development of algorithms for the 3‐D reconstruction of human body segments and their visualization. The parallel implementations exploit the fine grain parallelism allowed by the CM, processing each point of the data on a different processor. This mechanism is allowed by the possibility of dynamically reconfiguring the connectivity of the CM nodes and of defining a huge number of virtual processors. Moreover, as the CM processors operate on one‐bit data, it is possible to tune the number of bits for each data point to match the accuracy required by the application. Giampiero Marcenaro, Massimo Tistarelli |
Concurr. Pract. Exp. | 2 |
| 1990 | On The Estimation Of Depth From Motion Using An Anthropomorphic Visual Sensor
Massimo Tistarelli, Giulio Sandini |
ECCV | 1 |
| 1990 | Robot navigation using an anthropomorphic visual sensorabstractThe use of an anthropomorphic, retinalike visual sensor for navigation tasks is investigated. The main advantage, besides the topological scaling and rotation invariance, stems from the considerable data reduction obtained with nonuniform sampling, in conjunction with high resolution in the part of the field of view corresponding to the focus of attention. Active movements are also considered to be a beneficial feature, solving the depth-from-motion problem and maintaining a 3-D representation of the viewed scene. For short range navigation, a tracking egomotion strategy is adopted which greatly simplifies the motion equations and complements the characteristics of the retinal sensor. An algorithm for the computation of depth from motion is developed for image sequences acquired with the retinal sensor, and an error analysis is carried out to determine the uncertainty of range measurements. An experiment is presented in which depth maps are computed from a sequence of images sampled with the retinalike sensor, building a volumetric representation of the scene.> Massimo Tistarelli, Giulio Sandini |
ICRA | 1 |
| 1990 | Estimation of depth from motion using an anthropomorphic visual sensor
Massimo Tistarelli, Giulio Sandini |
Image Vis. Comput. | 1 |
| 1990 | Active Tracking Strategy for Monocular Depth Inference over Multiple FramesabstractThe extraction of depth information from a sequence of images is investigated. An algorithm that exploits the constraint imposed by active motion of the camera is described. Within this framework, in order to facilitate measurement of the navigation parameters, a constrained egomotion strategy was adopted in which the position of the fixation point is stabilized during the navigation (in an anthropomorphic fashion). This constraint reduces the dimensionality of the parameter space without increasing the complexity of the equations. A further distinctive point is the use of two sampling rates: the faster (related to the computation of the instantaneous optical flow) is fast enough to allow the local operator to sense the passing edge (or, in other words, to allow the tracking of moving contour points), while the slower (used to perform the triangulation procedure necessary to derive depth) is slow enough to provide a sufficiently large baseline for triangulation. Experimental results on real image sequences are presented.> Giulio Sandini, Massimo Tistarelli |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1990 | Using camera motion to estimate range for robotic parts manipulationabstractA technique is described for determining a depth map of parts in bins using optical flow derived from camera motion. Simple programmed camera motions are generated by mounting the camera on the robot end effector and directing the effector along a known path. The results achieved using two simple trajectories, where one is along the optical axis and the other is in rotation about a fixation point, are detailed. Optical flow is estimated by computing the time derivative of a sequence of images, i.e. by forming differences between two successive images and, in particular, matching between contours in images that have been generated from the zero crossings of Laplacian of Gaussian-filtered images. Once the flow field has been determined, a depth map is computed utilizing the parameters of the known camera trajectory. Empirical results are presented for a calibration object and two bins of parts; these are compared with the theoretical precision of the technique, and it is demonstrated that a ranging accuracy on the order of two parts in 100 is achievable.> David Vernon, Massimo Tistarelli |
IEEE Trans. Robotics Autom. | 2 |
| 1989 | 3D object reconstruction using stereo and motionabstractThe extraction of reliable range data from images is investigated, considering, as a possible solution, the integration of different sensor modalities. Two different algorithms are used to obtain independent estimates of depth from a sequence of stereo images. The results are integrated on the basis of the uncertainty of each measure. The stereo algorithm uses a coarse-to-fine control strategy to compute disparity. An algorithm for depth-from-motion is used, exploiting the constraint imposed by active motion of the cameras. To obtain a 3D description of the objects, the motion of the cameras is purposefully controlled, in such a manner as to move around the objects in view while the gaze is directed toward a fixed point in space. This egomotion strategy, which is similar to that adopted by the human visuomotor system, allows a better exploration of partially occluded objects and simplifies the motion equations. When tested on real scenes, the algorithm demonstrated a low sensitivity to image noise, mainly due to the integration of independent measures. An experiment performed on a real scene containing several objects is presented.> Enrico Grosso, Giulio Sandini, Massimo Tistarelli |
IEEE Trans. Syst. Man Cybern. | 3 |
| 1986 | Analysis of object motion and camera motion in real scenesabstractIn this paper a contour based motion estimation technique is presented applied to the measurement of the optic-flow field from image sequences. In particular two experimental situations are considered: the former, in which a fixed camera analyzes a scene with some moving objects in order to detect their trajectories; the latter, in which a sensor moves in a static environment, determining the depth map of the "world". For the first case, we compute for every contour the prevalent direction of movement. For the second, we solve the problem of recovering the 3D structure of the worldby computing the depth field from a sequence of images acquired during a motion of the camera, controlled in order to mantain the fixation point (i.e. the point projected over the image plane) still during the entire sequence. The algorithm has proved to be very robust in real scene analysis and an acceptable time for computation is required. Some experimental results are presented. Giulio Sandini, Vincenzo Tagliasco, Massimo Tistarelli |
ICRA | 3 |