William Robson Schwartz

dblp:78/6398 · also William Schwartz 0001 · DBLP profile ↗
← Back
101ranked-venue papers
11as first author
15since 2021 · last 2025
0000-0003-1449-8834ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 75 · 9 first-author · 6 since 2021Artificial intelligence and machine learning · 55 · 6 first-author · 9 since 2021Security and privacy · 9 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 7 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2025 Axial Sphere Loss: Encouraging Open-Space Risk Minimization in Face Identification Tasks
abstract
Open-set face recognition challenges biometric systems by requiring them to identify registered subjects while rejecting unregistered individuals. This task is particularly demanding in watchlist scenarios, where biometric systems must focus on subjects of interest and disregard irrelevant faces. To address real-world face applications, this study associates quickly trainable adaptation networks with a logit-and-distance-based cost function that explores non-gallery samples in favor of minimizing the open-space risk. These negative instances are either specified in dataset protocols or synthetically built at training time. The proposed Axial Sphere Loss (ASL) shifts each class into pre-defined regions in the latent space and mutually pushes non-gallery samples toward the space origin, forming spherical containers around each class template at inference time. We show that training an adapter network with ASL does not hinder closed-set recognition scores but significantly boosts open-set identification rates, achieving state-of-the-art performance on three well-known face benchmarks, namely, LFW, IJB-C, and UCCS datasets.
Rafael Henrique Vareto, William Robson Schwartz
FG2
2025 Unsupervised Analysis of Cyclist Performance for Route Segmentation and Ranking
Rensso Mora Colque, William Robson Schwartz
ICINCO (1)2
2025 VESSA: Video-based objEct-centric Self-Supervised Adaptation for Visual Foundation Models
abstract
Foundation models have advanced computer vision by enabling strong performance across diverse tasks through large-scale pretraining and supervised fine-tuning. However, they may underperform in domains with distribution shifts and scarce labels, where supervised fine-tuning may be infeasible. While continued self-supervised learning for model adaptation is common for generative language models, this strategy has not proven effective for vision-centric encoder models. To address this challenge, we introduce a novel formulation of self-supervised fine-tuning for vision foundation models, where the model is adapted to a new domain without requiring annotations, leveraging only short multi-view object-centric videos. Our method is referred to as VESSA: **V**ideo-based obj**E**ct-centric **S**elf-**S**upervised **A**daptation for visual foundation models. VESSA's training technique is based on a self-distillation paradigm, where it is critical to carefully tune prediction heads and deploy parameter-efficient adaptation techniques – otherwise, the model may quickly forget its pretrained knowledge and reach a degraded state. VESSA benefits significantly from multi-view object observations sourced from different frames in an object-centric video, efficiently learning robustness to varied capture conditions, without the need of annotations. Through comprehensive experiments with 3 vision foundation models on 2 datasets, VESSA demonstrates consistent improvements in downstream classification tasks, compared to the base models and previous adaptation methods. Code is publicly available at https://github.com/jesimonbarreto/VESSA.
Jesimon Barreto Santos, Carlos Antônio Caetano Jr., André Araújo 0001, William Robson Schwartz
NeurIPS4
2025 A Principled Benchmark for Seismic Data Segmentation
abstract
In recent years, several deep learning techniques have been applied to the problem of seismic facies segmentation. However, there is a lack of an authoritative protocol for evaluating such models, so that the comparison between results becomes compromised. This paper proposes a principled benchmark for lithofacies segmentation based on the public seismic volumes from the F3 Netherlands, Penobscot and Parihaka datasets, along with standard metrics for performance assessment. The utility of the benchmark is illustrated by the evaluation of the U-Net DeconvNet and SegNet encoder-decoder architectures. The goal is to offer a framework which will enable researchers to compare different methods and to develop more effective strategies for the segmentation problem.
Gabriel L. Canguçu, Leonardo M. S. Jorge, Gabriela T. Barreto, Thales H. Silva, Luiz A. Lima, Walace S. Caldas, Carlos G. S. Tavares, William Robson Schwartz, Pedro O. S. Vaz-de-Melo, Alexei M. C. Machado
IEEE Trans. Geosci. Remote. Sens.8
2024 Watchlist Challenge: 3rd Open-set Face Detection and Identification
abstract
In the current landscape of biometrics and surveillance, the ability to accurately recognize faces in uncontrolled settings is paramount. The Watchlist Challenge addresses this critical need by focusing on face detection and open-set identification in real-world surveillance scenarios. This paper presents a comprehensive evaluation of participating algorithms, using the enhanced UnConstrained College Students (UCCS) dataset with new evaluation protocols. In total, four participants submitted four face detection and nine open-set face recognition systems. The evaluation demonstrates that while detection capabilities are generally robust, closed-set identification performance varies significantly, with models pre-trained on large-scale datasets showing superior performance. However, open-set scenarios require further improvement, especially at higher true positive identification rates, i.e., lower thresholds.
Furkan Kasim, Terrance E. Boult, Rensso Mora Colque, Bernardo Biesseck, Rafael O. Ribeiro, Jan Schlüter, Tomás Repák, Rafael Henrique Vareto, David Menotti, William Robson Schwartz, Manuel Günther
IJCB10
2024 Open-set face recognition with maximal entropy and Objectosphere loss
Rafael Henrique Vareto, Yu Linghu, Terrance E. Boult, William Robson Schwartz, Manuel Günther
Image Vis. Comput.4
2023 Diurnal Pain Classification in Critically Ill Patients using Machine Learning on Accelerometry and Analgesic Data
abstract
Quantifying pain in patients admitted to intensive care units (ICUs) is challenging due to the increased prevalence of communication barriers in this patient population. Previous research has posited a positive correlation between pain and physical activity in critically ill patients. In this study, we advance this hypothesis by building machine learning classifiers to examine the ability of accelerometer data collected from daily wearables to predict self-reported pain levels experienced by patients in the ICU. We trained multiple Machine Learning (ML) models, including Logistic Regression, CatBoost, and XG-Boost, on statistical features extracted from the accelerometer data combined with previous pain measurements and patient demographics. Following previous studies that showed a change in pain sensitivity in ICU patients at night, we performed the task of pain classification separately for daytime and nighttime pain reports. In the pain versus no-pain classification setting, logistic regression gave the best classifier in daytime (AUC: 0.72, F1-score: 0.72), and CatBoost gave the best classifier at nighttime (AUC: 0.82, F1-score: 0.82). Performance of logistic regression dropped to 0.61 AUC, 0.62 F1-score (mild vs. moderate pain, nighttime), and CatBoost's performance was similarly affected with 0.61 AUC, 0.60 F1-score (moderate vs. severe pain, daytime). The inclusion of analgesic information benefited the classification between moderate and severe pain. SHAP analysis was conducted to find the most significant features in each setting. It assigned the highest importance to accelerometer-related features on all evaluated settings but also showed the contribution of the other features such as age and medications in specific contexts. In conclusion, accelerometer data combined with patient demographics and previous pain measurements can be used to screen painful from painless episodes in the ICU and can be combined with analgesic information to provide moderate classification between painful episodes of different severities.
Jessica Sena, Sabyasachi Bandyopadhyay, Mohammad Tahsin Mostafiz, Andrea Davidson, Ziyuan Guan, Jesimon Barreto Santos, Tezcan Baslanti, Patrick James Tighe, Azra Bihorac, William Robson Schwartz, Parisa Rashidi
BIBM10
2023 Super-resolution of license plate images using attention modules and sub-pixel convolution layers
Valfride Nascimento, Rayson Laroca, Jorge de A. Lambert, William Robson Schwartz, David Menotti
Comput. Graph.4
2021 MoRe: A Large-Scale Motorcycle Re-Identification Dataset
abstract
Motorcycles are often related to transit and criminal issues due to its abundance in the transit. Despite its importance, motorcycles are a seldom addressed problem in the computer vision community. We credit this problem to the lack of large-scale datasets and strong baseline models. Therefore, we present the first large-scale Motorcycles Re-Identification (MoRe) dataset. MoRe consists of 3,827 individuals (i.e., the set of motorbikes and motorcyclist) captured by ten surveillance cameras placed in Brazil's urban traffic scenarios. Furthermore, we evaluate a deep learning model trained using well-known training tricks from the object re-identification literature to present a strong baseline for the motorcycle re-identification (ReID) problem. More importantly, we highlight some crucial problems in this topic as the influence of distractors and the domain shift. Experimental results demonstrate the effectiveness of the strong baseline model with an increase of at least 19.27 p.p. in the rank-1 when compared to the state-of-the-art in the BPReID dataset. Finally, we present some insights regarding the information learned by the strong baseline model when computing the similarities between motorcycle images.1
Augusto Figueiredo, Johnata Brayan, Renan Oliveira Reis, Raphael Felipe de Carvalho Prates, William Robson Schwartz
WACV5
2021 Covariance-free Partial Least Squares: An Incremental Dimensionality Reduction Method
abstract
Dimensionality reduction plays an important role in computer vision problems since it reduces computational cost and is often capable of yielding more discriminative data representation. In this context, Partial Least Squares (PLS) has presented notable results in tasks such as image classification and neural network optimization. However, PLS is infeasible on large datasets, such as ImageNet, because it requires all the data to be in memory in advance, which is often impractical due to hardware limitations. Additionally, this requirement prevents us from employing PLS on streaming applications where the data are being continuously generated. Motivated by this, we propose a novel incremental PLS, named Covariance-free Incremental Partial Least Squares (CIPLS), which learns a low-dimensional representation of the data using a single sample at a time. In contrast to other state-of-the-art approaches, instead of adopting a partially-discriminative or SGD-based model, we extend Nonlinear Iterative Partial Least Squares (NI-PALS) - the standard algorithm used to compute PLS - for incremental processing. Among the advantages of this approach are the preservation of discriminative information across all components, the possibility of employing its score matrices for feature selection, and its computational efficiency. We validate CIPLS on face verification and image classification tasks, where it outperforms several other incremental dimensionality reduction techniques. In the context of feature selection, CIPLS achieves comparable results when compared to state-of-the-art techniques.
Artur Jordão, Maiko M. I. Lie, Victor C. de Melo, William Robson Schwartz
WACV4
2021 Human activity recognition based on smartphone and wearable sensors using multiscale DCNN ensemble
Jessica Sena, Jesimon Barreto Santos, Carlos Antônio Caetano Jr., Guilherme Cramer, William Robson Schwartz
Neurocomputing5
2021 A content-based late fusion approach applied to pedestrian detection
Jessica Sena, Artur Jordão, William Robson Schwartz
J. Vis. Commun. Image Represent.3
2021 Simple and efficient pose-based gait recognition method for challenging environments
Vitor de Lima, Victor C. de Melo, William Robson Schwartz
Pattern Anal. Appl.3
2021 Correction to: Simple and efficient pose-based gait recognition method for challenging environments
Vitor de Lima, Victor C. de Melo, William Robson Schwartz
Pattern Anal. Appl.3
2021 Face spoofing detection via ensemble of classifiers toward low-power devices
Rafael Henrique Vareto, William Robson Schwartz
Pattern Anal. Appl.2
2020 Face Attributes as Cues for Deep Face Recognition Understanding
abstract
Deeply learned representations are the state-of-the-art descriptors for face recognition methods. These representations encode latent features that are difficult to explain, compromising the confidence and interpretability of their predictions. Most attempts to explain deep features are visualization techniques that are often open to interpretation. Instead of relying only on visualizations, we use the outputs of hidden layers to predict face attributes. The obtained performance is an indicator of how well the attribute is implicitly learned in that layer of the network. Using a variable selection technique, we also analyze how these semantic concepts are distributed inside each layer, establishing the precise location of relevant neurons for each attribute. According to our experiments, gender, eyeglasses and hat usage can be predicted with over 96% accuracy even when only a single neural output is used to predict each attribute. These performances are less than 3 percentage points lower than the ones achieved by deep supervised face attribute networks. In summary, our experiments show that, inside DCNNs optimized for face identification, there exists latent neurons encoding face attributes almost as accurately as DCNNs optimized for these attributes.
Matheus Alves Diniz, William Robson Schwartz
FG2
2020 The Swax Benchmark: Attacking Biometric Systems with Wax Figures
abstract
A face spoofing attack occurs when an intruder attempts to impersonate someone who carries a gainful authentication clearance. It is a trending topic due to the increasing demand for biometric authentication on mobile devices, high-security areas, among others. This work introduces a new database named Sense Wax Attack dataset (SWAX), comprised of real human and wax figure images and videos that endorse the problem of face spoofing detection. The dataset consists of more than 1800 face images and 110 videos of 55 people/waxworks, arranged in training, validation and test sets with a large range in expression, illumination and pose variations. Experiments performed with baseline methods show that despite the progress in recent years, advanced spoofing methods are still vulnerable to high-quality violation attempts.
Rafael Henrique Vareto, Araceli Marcia Saldanha, William Robson Schwartz
ICASSP3
2020 Unconstrained Face Identification using Ensembles trained on Clustered Data
abstract
Open-set face recognition describes a scenario where unknown subjects, unseen during training stage, appear on test time. Not only it requires methods that accurately identify individuals of interest, but also demands approaches that effectively deal with unfamiliar faces. This work details a scalable open-set face identification approach to galleries composed of hundreds and thousands of subjects. It is composed of clustering and ensemble of binary learning algorithms that estimates when query face samples belong to the face gallery and then retrieves their correct identity. The approach selects the most suitable gallery subjects and use the ensemble to improve prediction performance. We carry out experiments on well-known LFW and YTF benchmarks. Results show that competitive performance can be achieved even when targeting scalability.
Rafael Henrique Vareto, William Robson Schwartz
IJCB2
2020 Bubblenet: A Disperse Recurrent Structure To Recognize Activities
abstract
This paper presents an approach to perform human activity recognition in videos through the employment of a deep recurrent network, taking as inputs appearance and optical flow information. Our method proposes a novel architecture named BubbleNET, which is based on a recurrent layer dispersed into several modules (referred to as bubbles) along with an attention mechanism based on squeeze-and-excitation strategy, responsible to modulate each bubble contribution. Thereby, we intend to gather information from fundamentally correlated segments of the input data, creating a signature of components that characterize each activity. Our experiments, conducted on widely employed activity recognition datasets, support the existence of these signatures, evidenced by maps of bubble activations for every class of the datasets. To compare the approach to literature methods, mean accuracy is taken into account, for which BubbleNET obtained 97.62%, 91.70% and 82.60% on UCF-101, YUP++ and HMDB-51 datasets, respectively, being placed among state-of-the-art methods.
Igor L. O. Bastos, Victor C. de Melo, William Robson Schwartz
ICIP3
2020 Cycleptz: The Learning-Based Control Method For Master-Slave Camera Systems
abstract
A master-slave setup consists of fixed and PTZ cameras monitoring a scenario to provide high-resolution images of target regions. Most of the works in literature focus on the unrealistic setting with a single person and single ground plane. Differently, in this work we proposed the CyclePTZ, a learning-based method that learns the mappings between master-slave and slave-master using a cycle-consistent neural network. While the master-slave mapping is used to control the PTZ cameras, the slave-master works as a supplementary supervisor for network training and to perform hard sample mining. More importantly, as both functions are learned simultaneously using the cycle loss, the CyclePTZ is able to learn a better mapping between fixed and PTZ cameras. Experimental results demonstrate that the proposed CyclePTZ is able to follow targets in multiple ground planes and to record corresponding points with multiple people in the scene. We compare the proposed method with the current literature in real-time experiments that demonstrate favorable performance of CyclePTZ (i.e., reducing the target center error in 10.92 percentage points).
Renan Oliveira Reis, Igor Dias, Raphael Felipe de Carvalho Prates, William Robson Schwartz
ICIP4
2020 Stage-Wise Neural Architecture Search
abstract
Modern convolutional networks such as ResNet and NASNet have achieved state-of-the-art results in many computer vision applications. These architectures consist of stages, which are sets of layers that operate on representations in the same resolution. It has been demonstrated that increasing the number of layers in each stage improves the prediction ability of the network. However, the resulting architecture becomes computationally expensive in terms of floating point operations, memory requirements and inference time. Thus, significant human effort is necessary to evaluate different trade-offs between depth and performance. To handle this problem, recent works have proposed to automatically design high-performance architectures, mainly by means of neural architecture search (NAS). Current NAS strategies analyze a large set of possible candidate architectures and, hence, require vast computational resources and take many GPUs days. Motivated by this, we propose a NAS approach to efficiently design accurate and low-cost convolutional architectures and demonstrate that an efficient strategy for designing these architectures is to learn the depth stage-by-stage. For this purpose, our approach increases depth incrementally in each stage taking into account its importance, such that stages with low importance are kept shallow while stages with high importance become deeper. We conduct experiments on the CIFAR and different versions of ImageNet datasets, where we show that architectures discovered by our approach achieve better accuracy and efficiency than human-designed architectures. Additionally, we show that architectures discovered on CIFAR-10 can be successfully transferred to large datasets. Compared to previous NAS approaches, our method is substantially more efficient, as it evaluates one order of magnitude fewer models and yields architectures on par with the state-of-the-art.
Artur Jordão, Fernando Akio, Maiko M. I. Lie, William Robson Schwartz
ICPR4
2020 Deep network compression based on partial least squares
Artur Jordão, Fernando Yamada, William Robson Schwartz
Neurocomputing3
2019 SkeleMotion: A New Representation of Skeleton Joint Sequences based on Motion Information for 3D Action Recognition
abstract
Due to the availability of large-scale skeleton datasets, 3D human action recognition has recently called the attention of computer vision community. Many works have focused on encoding skeleton data as skeleton image representations based on spatial structure of the skeleton joints, in which the temporal dynamics of the sequence is encoded as variations in columns and the spatial structure of each frame is represented as rows of a matrix. To further improve such representations, we introduce a novel skeleton image representation to be used as input of Convolutional Neural Networks (CNNs), named SkeleMotion. The proposed approach encodes the temporal dynamics by explicitly computing the magnitude and orientation values of the skeleton joints. Different temporal scales are employed to compute motion values to aggregate more temporal dynamics to the representation making it able to capture long-range joint interactions involved in actions as well as filtering noisy motion values. Experimental results demonstrate the effectiveness of the proposed representation on 3D action recognition outperforming the state-of-the-art on NTU RGB+D 120 dataset.
Carlos Antônio Caetano Jr., Jessica Sena, François Brémond, Jefersson A. dos Santos, William Robson Schwartz
AVSS5
2019 Multi-task Learning for Low-Resolution License Plate Recognition
Gabriel Resende Gonçalves, Matheus Alves Diniz, Rayson Laroca, David Menotti, William Robson Schwartz
CIARP5
2019 Gait Recognition Using Pose Estimation and Signal Processing
Vitor de Lima, William Robson Schwartz
CIARP2
2019 Face Spoofing Detection on Low-Power Devices Using Embeddings with Spatial and Frequency-Based Descriptors
Rafael Henrique Vareto, Matheus Alves Diniz, William Robson Schwartz
CIARP3
2019 Point-Placement Techniques and Temporal Self-Similarity Maps for Visual Analysis of Surveillance Videos
abstract
The manual analysis of surveillance videos is unfeasible due to the excessive amount of data, the associated subjectivity, or the eventual presence of distracting noise. Automatic summarization approaches provide little/no user interaction, limiting his/her comprehension regarding the involved phenomena. Visual analytics techniques represent a potential tool for such analysis, providing video representations that clearly communicate their content, potentially revealing patterns that may represent events of interest. In this paper, we present a methodology for visual analysis of surveillance video that combines point-placement techniques and Temporal Self-similarity Maps (TSSMs) to reveal the events occurrence structure and to enhance the comprehension of their temporal properties. Experiments in several surveillance scenarios demonstrate that our proposed methodology provides an effective events summarization and the exploration of both the structure of each event and the relationship among them, allowing the security agent to filter/explore those that represent potential alert situations.
Gilson Mendes, Jose Gustavo Paiva, William Robson Schwartz
IV (1)3
2019 Magnitude-Orientation Stream network and depth information applied to activity recognition
Carlos Antônio Caetano Jr., Victor C. de Melo, François Brémond, Jefersson A. dos Santos, William Robson Schwartz
J. Vis. Commun. Image Represent.5
2019 Kernel cross-view collaborative representation based classification for person re-identification
Raphael Felipe de Carvalho Prates, William Robson Schwartz
J. Vis. Commun. Image Represent.2
2019 Dynamic Multicontext Segmentation of Remote Sensing Images Based on Convolutional Networks
abstract
Semantic segmentation requires methods capable of learning high-level features while dealing with large volume of data. Toward such goal, convolutional networks can learn specific and adaptable features based on the data. However, these networks are not capable of processing a whole remote sensing image, given its huge size. To overcome such limitation, the image is processed using fixed size patches. The definition of the input patch size is usually performed empirically (evaluating several sizes) or imposed (by network constraint). Both strategies suffer from drawbacks and could not lead to the best patch size. To alleviate this problem, several works exploited multicontext information by combining networks or layers. This process increases the number of parameters, resulting in a more difficult model to train. In this paper, we propose a novel technique to perform semantic segmentation of remote sensing images that exploits a multicontext paradigm without increasing the number of parameters while defining, in training time, the best patch size. The main idea is to train a dilated network with distinct patch sizes, allowing it to capture multicontext characteristics from heterogeneous contexts. While processing these varying patches, the network provides a score for each patch size, helping in the definition of the best size for the current scenario. A systematic evaluation of the proposed algorithm is conducted using four high-resolution remote sensing data sets with very distinct properties. Our results show that the proposed algorithm provides improvements in pixelwise classification accuracy when compared to the state-of-the-art methods.
Keiller Nogueira, Mauro Dalla Mura, Jocelyn Chanussot, William Robson Schwartz, Jefersson A. dos Santos
IEEE Trans. Geosci. Remote. Sens.4
2018 MORA: A Generative Approach to Extract Spatiotemporal Information Applied to Gesture Recognition
abstract
Gestures are related to a non-verbal language used on the interaction between subjects. Due to its applicability in several contexts, gesture recognition has been investigated by different researches, often investing on the capture of motion and appearance on videos. However, most of these methods do not properly explore the well-defined gesture temporal structure and are not suitable to deal with an increasing number of classes. Thus, we propose the Multi-Output Recurrent Autoencoders (MORA), an approach that relies on the representation of each gesture class independently. MORA employs a specific autoencoder model per class, composed by convolutional (3D) and a Gated Recurrent Unit (GRU) layer, what allows spatiotemporal information extraction and scalability in terms of number of classes. To validate MORA, experiments are conducted on SKIG and ChaLearn IsoGD datasets, for which the approach achieved accuracies comparable to state-of-the-art methods.
Igor L. O. Bastos, Victor C. de Melo, Gabriel Resende Gonçalves, William Robson Schwartz
AVSS4
2018 AVSS Challenges 2018 Soft Biometric Retrieval Using Deep Multi-Task Network
abstract
In surveillance, humans are the agents performing actions to change the states in the scene. They are the main focus in the surveillance systems and, therefore, the design of processing methods focusing on humans is extremely important to identify a person and determine his/her role in the scene. Thus, one of the goals of smart surveillance systems is to address the automatic video understanding by applying computer vision techniques to automatically detect specific humans in video streams based on attributes, such as soft biometrics. For that purpose, this work proposes an approach that receives a set of textual attributes as a query and searches for people by matching those attributes in a gallery of images, as defined in the challenge Semantic Person Retrieval in Surveillance Using Soft Biometrics Challenge, proposed in AVSS 2018. We address this problem with a multitask learning approach hypothesizing that the attributes available for the query are highly related to each other that could be learned together in the same network as different tasks. Finally, we performed both qualitative and quantitative experimental evaluations, indicating a promising direction for the proposed approach.
Gabriel Resende Gonçalves, Antonio C. Nazare, Matheus Alves Diniz, Luiz Eduardo Coelho Lima, William Robson Schwartz
AVSS5
2018 Content-Based Multi-Camera Video Alignment using Accelerometer Data
abstract
Video alignment is an important task for environments with distributed multiple cameras, making possible, for instance, to verify the exact instant that an event happened on different views. The alignment comprises establishing a temporal correspondence among frames captured by different video cameras. Many works have been proposed to solve this problem when the cameras present overlapping of Field of View (FOV) or are located close to each other. However, in this work, we present a novel approach to perform video alignment for cameras without overlapping FOV (i.e., the cameras might be located on different floors of a building). The method employs the sensor data generated by a smartphone (synchronized to a time server), to align multiple videos by finding a temporal matching between the videos which captured a person in the scene and the signal of the smartphone accelerometer carried by this individual, providing the exact time that a movement was performed. To the best of our knowledge, this is the first attempt at performing such type of alignment. According to experimental results, the proposed approach was able to align multiple videos with a length of 30 minutes with errors as low as 160 ms.
Antonio C. Nazare, Filipe de Oliveira Costa, William Robson Schwartz
AVSS3
2018 Neural Network Control for Active Cameras Using Master-Slave Setup
abstract
The use of active cameras has increased to perform tasks such as tracking and biometrics at distance. Furthermore, recent efforts have focused on the master-slave setup, which is composed of fixed and PTZ cameras. Although, there are many works regarding active camera control, there is no standard way to compare different control approaches once the experiment cannot be reproduced. Thus, in this work, besides the proposition of a novel learning-based approach to the master-slave setup, we also propose an experimental setup that allows a fair comparison between different methods. The proposed control method learns corresponding points between the fixed and the PTZ cameras using a neural network. The novel experimental setup places two PTZ cameras side-by-side with a very similar view so that two different algorithms can be executed simultaneously. The experiments show that the proposed method is better than literature method when the focus is centralizing a target at the PTZ view.
Renan Oliveira Reis, Igor Dias, William Robson Schwartz
AVSS3
2018 Face Verification: Strategies for Employing Deep Models
abstract
Features extracted with deep learning have now achieved state-of-the-art results in many tasks. However, to reuse a learned deep model, transfer learning with fine-tuning needs to be employed, which requires to re-train the whole model or part of it to extract useful features in the new domain. This step is burdensome and requires heavy computing power. Therefore, this work investigates alternatives in transfer-learning that do not involve performing fine-tuning for a model with the new domain. Namely, we explore the correlation of depth and scale in deep models, and look for the layer/scale that yields the best results for the new domain, we also explore metrics for the verification task, using locally connected convolutions to learn distance metrics. Our experiments use a model pre-trained in face identification and adapt it to the face verification task with different data, but still on the face domain. We achieve 96.65% mean accuracy on the Labeled Faces in the Wild dataset and 93.12% mean accuracy on the Youtube Faces dataset which are in the state-of-the-art.
Ricardo Barbosa Kloss, Artur Jordão, William Robson Schwartz
FG3
2018 Latent HyperNet: Exploring the Layers of Convolutional Neural Networks
abstract
Since Convolutional Neural Networks (ConvNets) are able to simultaneously learn features and classifiers to discriminate different categories of activities, recent works have employed ConvNets approaches to perform human activity recognition (HAR) based on wearable sensors, allowing the removal of expensive human work and expert knowledge. However, these approaches have their power of discrimination limited mainly by the large number of parameters that compose the network and the reduced number of samples available for training. Inspired by this, we propose an accurate and robust approach, referred to as Latent HyperNet (LHN). The LHN uses feature maps from early layers (hyper) and projects them, individually, onto a low dimensionality (latent) space. Then, these latent features are concatenated and presented to a classifier. To demonstrate the robustness and accuracy of the LHN, we evaluate it using four different network architectures in five publicly available HAR datasets based on wearable sensors, which vary in the sampling rate and number of activities. We experimentally demonstrate that the proposed LHN is able to capture rich information, improving the results regarding the original ConvNets. Furthermore, the method outperforms existing state-of-the-art methods, on average, by 5.1 percentage points.
Artur Jordão, Ricardo Barbosa Kloss, William Robson Schwartz
IJCNN3
2018 A Robust Real-Time Automatic License Plate Recognition Based on the YOLO Detector
abstract
Automatic License Plate Recognition (ALPR) has been a frequent topic of research due to many practical applications. However, many of the current solutions are still not robust in real-world situations, commonly depending on many constraints. This paper presents a robust and efficient ALPR system based on the state-of-the-art YOLO object detector. The Convolutional Neural Networks (CNNs) are trained and finetuned for each ALPR stage so that they are robust under different conditions (e.g., variations in camera, lighting, and background). Specially for character segmentation and recognition, we design a two-stage approach employing simple data augmentation tricks such as inverted License Plates (LPs) and flipped characters. The resulting ALPR approach achieved impressive results in two datasets. First, in the SSIG dataset, composed of 2,000 frames from 101 vehicle videos, our system achieved a recognition rate of 93.53% and 47 Frames Per Second (FPS), performing better than both Sighthound and OpenALPR commercial systems (89.80% and 93.03%, respectively) and considerably outperforming previous results (81.80%). Second, targeting a more realistic scenario, we introduce a larger public dataset1dataset, designed to ALPR. This dataset contains 150 videos and 4,500 frames captured when both camera and vehicles are moving and also contains different types of vehicles (cars, motorcycles, buses and trucks). In our proposed dataset, the trial versions of commercial systems achieved recognition rates below 70%. On the other hand, our system performed better, with recognition rate of 78.33% and 35 FPS.The UFPR-ALPR dataset is publicly available to the research community at https://web.inf.ufpr.br/vri/databases/ufpr-alpr/ subject to privacy restrictions.
Rayson Laroca, Evair Severo, Luiz Antonio Zanlorensi, Luiz Eduardo Soares de Oliveira, Gabriel Resende Gonçalves, William Robson Schwartz, David Menotti
IJCNN6
2018 Preface of Special Issue on Data Representation and Representation Learning for Video Analysis
William Robson Schwartz, Larry Davis 0001
Pattern Recognit. Lett.1
2018 Learning Deep Off-the-Person Heart Biometrics Representations
abstract
Since the beginning of the new millennium, the electrocardiogram (ECG) has been studied as a biometric trait for security systems and other applications. Recently, with devices such as smartphones and tablets, the acquisition of ECG signal in the off-the-person category has made this biometric signal suitable for real scenarios. In this paper, we introduce the usage of deep learning techniques, specifically convolutional networks, for extracting useful representation for heart biometrics recognition. Particularly, we investigate the learning of feature representations for heart biometrics through two sources: on the raw heartbeat signal and on the heartbeat spectrogram. We also introduce heartbeat data augmentation techniques, which are very important to generalization in the context of deep learning approaches. Using the same experimental setup for six methods in the literature, we show that our proposal achieves state-of-the-art results in the two off-the-person publicly available databases.
Eduardo José da S. Luz, Gladston J. P. Moreira, Luiz Eduardo Soares de Oliveira, William Robson Schwartz, David Menotti
IEEE Trans. Inf. Forensics Secur.4
2017 Pyramidal Zernike Over Time: A Spatiotemporal Feature Descriptor Based on Zernike Moments
Igor L. O. Bastos, Larissa Rocha Soares, William Robson Schwartz
CIARP3
2017 Boosted Projection: An Ensemble of Transformation Models
Ricardo Barbosa Kloss, Artur Jordão, William Robson Schwartz
CIARP3
2017 Noisy Character Recognition Using Deep Convolutional Neural Networks
Sirlene Peixoto, Gabriel Resende Gonçalves, Andrea Bianchi, Alceu De S. Brito, William Robson Schwartz, David Menotti
CIARP5
2017 Towards open-set face recognition using hashing functions
abstract
Face Recognition is one of the most relevant problems in computer vision as we consider its importance to areas such as surveillance, forensics and psychology. Furthermore, open-set face recognition has a large room for improvement since only few researchers have focused on it. In fact, a real-world recognition system has to cope with several unseen individuals and determine whether a given face image is associated with a subject registered in a gallery of known individuals. In this work, we combine hashing functions and classification methods to estimate when probe samples are known (i.e., belong to the gallery set). We carry out experiments with partial least squares and neural networks and show how response value histograms tend to behave for known and unknown individuals whenever we test a probe sample. In addition, we conduct experiments on FRGCv1, PubFig83 and VGGFace to show that our method continues effective regardless of the dataset difficulty.
Rafael Henrique Vareto, Samira Silva, Filipe de Oliveira Costa, William Robson Schwartz
IJCB4
2017 Combination techniques for hyperspectral image interpretation
abstract
In this work, we propose two main contributions to hyperspectral image interpretation. Firstly, while the traditional Weighted Linear Combination optimized by Genetic Algorithms (WLC-GA) [1] intends to give more discriminant power to those classification approaches contributing the most, we extend it to make a fine tuning over the class probabilities within the combination process. Then, we compare both methods (WLC-GA and its extension) with a more complex non-linear meta learning strategy called Stacked Generalization in which Support Vector Machines with Radial Basis Function kernel was used as combiner [2]. The experimental results, considering two widely used data sets, the Indian Pines and the Pavia University, are conducted in three different scenarios. Results show that both WLC-GA and its extended version achieve the best overall accuracy, and the proposed classification approach overcomes the accuracies of the other traditional ones used in this study.
Andrey Bicalho Santos, Arnaldo de Albuquerque Araújo, Jefersson A. dos Santos, William Robson Schwartz, David Menotti
IGARSS4
2017 Histograms of Optical Flow Orientation and Magnitude and Entropy to Detect Anomalous Events in Videos
abstract
This paper presents an approach for detecting anomalous events in videos with crowds. The main goal is to recognize patterns that might lead to an anomalous event. An anomalous event might be characterized by the deviation from the normal or usual, but not necessarily in an undesirable manner, e.g., an anomalous event might just be different from normal but not a suspicious event from the surveillance point of view. One of the main challenges of detecting such events is the difficulty to create models due to their unpredictability and their dependency on the context of the scene. Based on these challenges, we present a model that uses general concepts, such as orientation, velocity, and entropy to capture anomalies. Using such a type of information, we can define models for different cases and environments. Assuming images captured from a single static camera, we propose a novel spatiotemporal feature descriptor, calledhistograms of optical flow orientation and magnitude and entropy, based on optical flow information. To determine the normality or abnormality of an event, the proposed model is composed of training and test steps. In the training, we learn the normal patterns. Then, during test, events are described and if they differ significantly from the normal patterns learned, they are considered as anomalous. The experimental results demonstrate that our model can handle different situations and is able to recognize anomalous events with success. We use the well-known UCSD and Subway data sets and introduce a new data set, namely, Badminton.
Rensso Mora Colque, Carlos Antônio Caetano Jr., Matheus Toledo Lustosa de Andrade, William Robson Schwartz
IEEE Trans. Circuits Syst. Video Technol.4
2016 Kernel Partial Least Squares for person re-identification
abstract
Person re-identification (Re-ID) keeps the same identity for a person as he moves along an area with nonoverlapping surveillance cameras. Re-ID is a challenging task due to appearance changes caused by different camera viewpoints, occlusion and illumination conditions. While robust and discriminative descriptors are obtained combining texture, shape and color features in a high-dimensional representation, the achievement of accuracy and efficiency demands dimensionality reduction methods. At this paper, we propose variations of Kernel Partial Least Squares (KPLS) that simultaneously reduce the dimensionality and increase the discriminative power. The Cross-View KPLS (X-KPLS) and KPLS Mode A capture cross-view discriminative information and are successful for unsupervised and supervised Re-ID. Experimental results demonstrate that X-KPLS presents equal or higher matching results when compared to other methods in literature at PRID450S.
Raphael Felipe de Carvalho Prates, Marina Oliveira, William Robson Schwartz
AVSS3
2016 Oblique random forest based on partial least squares applied to pedestrian detection
abstract
The increasing popularity of approaches based on random forest in computer vision tasks is due to its simplicity and flexibility with complex data. Random forest is a set of decision trees that can be divided in two subsets according to the view of the feature descriptors provided as input: orthogonal and oblique. In the former, the feature space is separated orthogonally (axis-aligned) by a single feature at a time. In the latter, it separates the space by oriented hyperplanes, which usually provides better data modeling. This work proposes a novel oblique random forest associated with Partial Least Squares to perform the oblique split. We validate the proposed approach, referred to as oRF-PLS, on the challenge INRIA Person dataset. Experimental results demonstrate that the proposed method outperforms traditional state-of-the-art detectors. In addition, we demonstrate that PLS is a more suitable choice to build oblique random forest than SVM, being faster and producing more accurate forests.
Artur Jordão, William Robson Schwartz
ICIP2
2016 Predominant color name indexing structure for person re-identification
abstract
The automation of surveillance systems is important to allow real-time analysis of critical events, crime investigation and prevention. A crucial step in the surveillance systems is the person re-identification (Re-ID) which aims at maintaining the identity of agents in non-overlapping camera networks. Most of the works in literature compare a test sample against the entire gallery, restricting the scalability. We address this problem employing multiple indexing lists obtained by color name descriptors extracted from part-based models using our proposed Predominant Color Name (PCN) indexing structure. PCN is a flexible indexing structure that relates features to gallery images without the need of labelled training images and can be integrated with existing supervised and unsupervised person Re-ID frameworks. Experimental results demonstrate that the proposed approach outperforms indexation based on unsupervised clustering methods such as k-means and c-means. Furthermore, PCN reduces the computational efforts with a minimum performance degradation. For instance, when indexing 50% and 75% of the gallery images, we observed a reduction in AUC curve of 0.01 and 0.08, respectively, when compared to indexing the entire gallery.
Raphael Felipe de Carvalho Prates, Cristianne R. S. Dutra, William Robson Schwartz
ICIP3
2016 Optical Flow Co-occurrence Matrices: A novel spatiotemporal feature descriptor
abstract
Suitable feature representation is essential for performing video analysis and understanding in applications within the smart surveillance domain. In this paper, we propose a novel spatiotemporal feature descriptor based on co-occurrence matrices computed from the optical flow magnitude and orientation. Our method, called Optical Flow Co-occurrence Matrices (OFCM), extracts a robust set of measures known as Haralick features to describe the flow patterns by measuring meaningful properties such as contrast, entropy and homogeneity of co-occurrence matrices to capture local space-time characteristics of the motion through the neighboring optical flow magnitude and orientation. We evaluate the proposed method on the action recognition problem by applying a visual recognition pipeline involving bag of local spatiotemporal features and SVM classification. The experimental results, carried on three well-known datasets (KTH, UCF Sports and HMDB51), demonstrate that OFCM outperforms the results achieved by several widely employed spatiotemporal feature descriptors such as HOF, HOG3D and MBH, indicating its suitability to be used as video representation.
Carlos Antônio Caetano Jr., Jefersson A. dos Santos, William Robson Schwartz
ICPR3
2016 A late fusion approach to combine multiple pedestrian detectors
abstract
Pedestrian detection is a well-known problem in Computer Vision. To improve detection, several feature descriptors have been proposed and combined. However, there are cases where the most powerful features fail to discriminate between false positives similar to the human body structure and actual true positives, which is a critical problem for applications such as surveillance, driving assistance and robotics. To address this issue, we propose a novel approach to combine results of distinct pedestrian detectors by reinforcing the human hypothesis. The method is able to reduce the confidence of the false positives due to the lack of spatial consensus when multiple detectors are considered. Our experimental validation, performed on three pedestrian detection benchmarks, INRIA person, ETH and Caltech pedestrian dataset, demonstrates that the proposed approach, referred to as Spatial Consensus (SC), outperforms the state-of-the-art on INRIA and ETH datasets and achieves comparable results on the Caltech dataset.
Artur Jordão, Jessica Sena, William Robson Schwartz
ICPR3
2016 Learning to semantically segment high-resolution remote sensing images
abstract
Land cover classification is a task that requires methods capable of learning high-level features while dealing with high volume of data. Overcoming these challenges, Convolutional Networks (ConvNets) can learn specific and adaptable features depending on the data while, at the same time, learn classifiers. In this work, we propose a novel technique to automatically perform pixel-wise land cover classification. To the best of our knowledge, there is no other work in the literature that perform pixel-wise semantic segmentation based on data-driven feature descriptors for high-resolution remote sensing images. The main idea is to exploit the power of ConvNet feature representations to learn how to semantically segment remote sensing images. First, our method learns each label in a pixel-wise manner by taking into account the spatial context of each pixel. In a predicting phase, the probability of a pixel belonging to a class is also estimated according to its spatial context and the learned patterns. We conducted a systematic evaluation of the proposed algorithm using two remote sensing datasets with very distinct properties. Our results show that the proposed algorithm provides improvements when compared to traditional and state-of-the-art methods that ranges from 5 to 15% in terms of accuracy.
Keiller Nogueira, Mauro Dalla Mura, Jocelyn Chanussot, William Robson Schwartz, Jefersson A. dos Santos
ICPR4
2016 Kernel Hierarchical PCA for person re-identification
abstract
Person re-identification (Re-ID) maintains a global identity for an individual while he moves along a large area covered by multiple cameras. Re-ID enables a multi-camera monitoring of individual activity that is critical for surveillance systems. However, the low-resolution images combined with the different poses, illumination conditions and camera viewpoints make person Re-ID a challenging problem. To reach a higher matching performance, state-of-the-art methods map the data to a nonlinear feature space where they learn a cross-view matching function using training data. Kernel PCA is a statistical method that learns a common subspace that captures most of the variability of samples using a small number of vector basis. However, Kernel PCA disregards that images were captured by distinct cameras, a critical problem in person Re-ID. Differently, Hierarchical PCA (HPCA) captures a consensus projection between multiblock data (e.g, two camera views), but it is a linear model. Therefore, we propose the Kernel Hierarchical PCA (Kernel HPCA) to tackle camera transition and dimensionality reduction in a unique framework. To the best of our knowledge, this is the first work to propose a kernel extension to the multiblock HPCA method. Experimental results demonstrate that Kernel HPCA reaches a matching performance comparable with state-of-the-art nonlinear subspace learning methods at PRID450S and VIPeR datasets. Furthermore, Kernel HPCA reaches a better combination of subspace learning and dimensionality requiring significantly lower subspace dimensions.
Raphael Felipe de Carvalho Prates, William Robson Schwartz
ICPR2
2016 A scalable and flexible framework for smart video surveillance
Antonio C. Nazare, William Robson Schwartz
Comput. Vis. Image Underst.2
2016 A mid-level video representation based on binary descriptors: A case study for pornography detection
Carlos Antônio Caetano Jr., Sandra Eliza Fontes de Avila, William Robson Schwartz, Silvio Jamil Ferzoli Guimarães, Arnaldo de Albuquerque Araújo
Neurocomputing3
2016 Partial least squares for face hashing
Cassio E. dos Santos, Ewa Kijak, Guillaume Gravier, William Robson Schwartz
Neurocomputing4
2015 Coffee Crop Recognition Using Multi-scale Convolutional Neural Networks
Keiller Nogueira, William Robson Schwartz, Jefersson A. dos Santos
CIARP2
2015 A Study on Low-Cost Representations for Image Feature Extraction on Mobile Devices
Ramon F. Pessoa, William Robson Schwartz, Jefersson A. dos Santos
CIARP2
2015 A Non-parametric Approach to Detect Changes in Aerial Images
Marco Túlio Alves N. Rodrigues, Daniel Balbino de Mesquita, Erickson R. Nascimento, William Robson Schwartz
CIARP4
2015 CBRA: Color-based ranking aggregation for person re-identification
abstract
The problem of automatically tracking a pedestrian within camera networks with non-overlapping field-of-view, known as person re-identification, is a challenging task with still suboptimal results. Different features have been proposed in the literature, specially colors which achieved the best results when fused in a unique feature representation. Despite being better than considering individually, the fusion still does not explores all the feature discriminative power. Therefore, we propose the use of rank aggregation to improve the results. In this paper, we address the person re-identification problem using a Color-based Ranking Aggregation (CBRA) method, which explores different feature representations to obtain complementary ranking lists and combine them using the Stuart ranking aggregation method. The obtained experimental results demonstrate a great improvement in state-of-the-art, reaching top-1 rank recognition rates of 50.0% and 56.9% in the ViPER and PRTD450S data sets, respectively.
Raphael Felipe de Carvalho Prates, William Robson Schwartz
ICIP2
2015 Hyperspectral image interpretation based on partial least squares
abstract
Remote sensed hyperspectral images have been used for many purposes and have become one of the most important tools in remote sensing. Due to the large amount of available bands, e.g., a few hundreds, the feature extraction step plays an important role for hyperspectral images interpretation. In this paper, we extend a well-know feature extraction method called Extended Morphological Profile (EMP) which encodes spatial and spectral information by using Partial Least Squares (PLS) to emphasize the importance of the more discriminative features. PLS is employed twice in our proposal, i.e., to the EMP features and to the raw spatial information, which are then concatenated to be further interpreted by the SVM classifier. Our experiments in two well-known data sets, the Indian Pines and Pavia University, have shown that our proposal outperforms the accuracy of classification methods employing EMP and other baseline feature extraction methods with different classifiers.
Andrey Bicalho Santos, Arnaldo de Albuquerque Araújo, William Robson Schwartz, David Menotti
ICIP3
2015 Classification schemes based on Partial Least Squares for face identification
Gerson de Paulo Carlos, Hélio Pedrini, William Robson Schwartz
J. Vis. Commun. Image Represent.3
2015 A topology-based approach to computing neighborhood-of-interest points using the Morse complex
Ricardo Dutra da Silva, William Robson Schwartz, Hélio Pedrini, Jesus Pulido, Bernd Hamann
J. Vis. Commun. Image Represent.2
2015 Deep Representations for Iris, Face, and Fingerprint Spoofing Detection
abstract
Biometrics systems have significantly improved person identification and authentication, playing an important role in personal, national, and global security. However, these systems might be deceived (or spoofed) and, despite the recent advances in spoofing detection, current solutions often rely on domain knowledge, specific biometric reading systems, and attack types. We assume a very limited knowledge about biometric spoofing at the sensor to derive outstanding spoofing detection systems for iris, face, and fingerprint modalities based on two deep learning approaches. The first approach consists of learning suitable convolutional network architectures for each domain, whereas the second approach focuses on learning the weights of the network via back propagation. We consider nine biometric spoofing benchmarks - each one containing real and fake samples of a given biometric modality and attack type - and learn deep representations for each benchmark by combining and contrasting the two learning approaches. This strategy not only provides better comprehension of how these approaches interplay, but also creates systems that exceed the best known results in eight out of the nine benchmarks. The results strongly indicate that spoofing detection systems based on convolutional networks can be robust to attacks already known and possibly adapted, with little effort, to image-based attacks that are yet to come.
David Menotti, Giovani Chiachia, Allan Pinto, William Robson Schwartz, Hélio Pedrini, Alexandre X. Falcão, Anderson Rocha 0001
IEEE Trans. Inf. Forensics Secur.4
2015 Using Visual Rhythms for Detecting Video-Based Facial Spoof Attacks
abstract
Spoofing attacks or impersonation can be easily accomplished in a facial biometric system wherein users without access privileges attempt to authenticate themselves as valid users, in which an impostor needs only a photograph or a video with facial information of a legitimate user. Even with recent advances in biometrics, information forensics and security, vulnerability of facial biometric systems against spoofing attacks is still an open problem. Even though several methods have been proposed for photo-based spoofing attack detection, attacks performed with videos have been vastly overlooked, which hinders the use of the facial biometric systems in modern applications. In this paper, we present an algorithm for video-based spoofing attack detection through the analysis of global information which is invariant to content, since we discard video contents and analyze content-independent noise signatures present in the video related to the unique acquisition processes. Our approach takes advantage of noise signatures generated by the recaptured video to distinguish between fake and valid access videos. For that, we use the Fourier spectrum followed by the computation of video visual rhythms and the extraction of different characterization methods. For evaluation, we consider the novel unicamp video-attack database, which comprises 17 076 videos composed of real access and spoofing attack videos. In addition, we evaluate the proposed method using the replay-attack database, which contains photo-based and video-based face spoofing attacks.
Allan Pinto, William Robson Schwartz, Hélio Pedrini, Anderson Rocha 0001
IEEE Trans. Inf. Forensics Secur.2
2015 Face Spoofing Detection Through Visual Codebooks of Spectral Temporal Cubes
abstract
Despite important recent advances, the vulnerability of biometric systems to spoofing attacks is still an open problem. Spoof attacks occur when impostor users present synthetic biometric samples of a valid user to the biometric system seeking to deceive it. Considering the case of face biometrics, a spoofing attack consists in presenting a fake sample (e.g., photograph, digital video, or even a 3D mask) to the acquisition sensor with the facial information of a valid user. In this paper, we introduce a low cost and software-based method for detecting spoofing attempts in face recognition systems. Our hypothesis is that during acquisition, there will be inevitable artifacts left behind in the recaptured biometric samples allowing us to create a discriminative signature of the video generated by the biometric sensor. To characterize these artifacts, we extract time-spectral feature descriptors from the video, which can be understood as a low-level feature descriptor that gathers temporal and spectral information across the biometric sample and use the visual codebook concept to find mid-level feature descriptors computed from the low-level ones. Such descriptors are more robust for detecting several kinds of attacks than the low-level ones. The experimental results show the effectiveness of the proposed method for detecting different types of attacks in a variety of scenarios and data sets, including photos, videos, and 3D masks.
Allan Pinto, Hélio Pedrini, William Robson Schwartz, Anderson Rocha 0001
IEEE Trans. Image Process.3
2015 An Approach to Supporting Incremental Visual Data Classification
abstract
Automatic data classification is a computationally intensive task that presents variable precision and is considerably sensitive to the classifier configuration and to data representation, particularly for evolving data sets. Some of these issues can best be handled by methods that support users' control over the classification steps. In this paper, we propose a visual data classification methodology that supports users in tasks related to categorization such as training set selection; model creation, application and verification; and classifier tuning. The approach is then well suited for incremental classification, present in many applications with evolving data sets. Data set visualization is accomplished by means of point placement strategies, and we exemplify the method through multidimensional projections and Neighbor Joining trees. The same methodology can be employed by a user who wishes to create his or her own ground truth (or perspective) from a previously unlabeled data set. We validate the methodology through its application to categorization scenarios of image and text data sets, involving the creation, application, verification, and adjustment of classification models.
Jose Gustavo Paiva, William Robson Schwartz, Hélio Pedrini, Rosane Minghim
IEEE Trans. Vis. Comput. Graph.2
2014 Detection of Groups of People in Surveillance Videos Based on Spatio-Temporal Clues
Rensso Mora Colque, Guillermo Cámara Chávez, William Robson Schwartz
CIARP3
2014 Person Re-Identification Based on Weighted Indexing Structures
Cristianne R. S. Dutra, Matheus Castro Rocha, William Robson Schwartz
CIARP3
2014 Scalable Feature Extraction for Visual Surveillance
Antonio C. Nazare, Renato Ferreira 0001, William Robson Schwartz
CIARP3
2014 An Adaptive Vehicle License Plate Detection at Higher Matching Degree
Raphael Felipe de Carvalho Prates, Guillermo Cámara Chávez, William Robson Schwartz, David Menotti
CIARP3
2014 Spatial Pyramid Matching for Finger Spelling Recognition in Intensity Images
Samira Silva, William Robson Schwartz, Guillermo Cámara Chávez
CIARP2
2014 Change detection based on features invariant to monotonic transforms and spatial constrained matching
abstract
Discovering regions that have changed in a set of images acquired from a scene at different times and possibly from different view points and cameras is a crucial step for many image processing applications. Remote sensing, visual surveillance, medical diagnosis and treatment, civil infrastructure, and underwater sensing are some examples of such applications. This work proposes a novel approach to detect changes automatically without a learning step by using image analysis techniques and segmentation based on superpixels. Unlike most common approaches, which are pixel-based, we present an approach that combines super-pixel extraction, hierarchical clustering and segment matching. The experimental results show the effectiveness of the proposed approach comparing it a background subtraction technique, demonstrating the robustness of our algorithm to illumination variations, non-uniform attenuations, atmospheric absorption and swaying trees.
Marco Túlio Alves N. Rodrigues, Luciano O. Milen, Erickson R. Nascimento, William Robson Schwartz
ICASSP4
2014 An Optimized Sliding Window Approach to Pedestrian Detection
abstract
While a large number of surveillance cameras available nowadays provide a safe environment, the huge amount of data generated by them prevents a manual processing, requiring the application of automated methods to understand the scene. However, the majority of the currently available methods are still unable to process this amount of data in real time, mainly those focusing on pedestrian detection. To optimize pedestrian detection methods, this work proposes a novel approach that performs a random filtering supported by the Maximum Search Problem theorem to select a very small number from all possible detection windows. Although the random filtering is able to select regions that capture every person on an image, some windows can cover only parts of a person, diminishing the accuracy. To solve that, a regression is applied to adjust the windows to the person's location. The computational cost reduction comes from the fact that the proposed approach does not need to perform any processing while selecting windows, differently from cascades of rejection that must evaluate at least simple features for every window. The experiments performed using a pedestrian detection based on Partial Least Squares show that the approach is effective in both accuracy and computational cost reduction.
Victor C. de Melo, Samir Leao, David Menotti, William Robson Schwartz
ICPR4
2014 Smart surveillance framework: A versatile tool for video analysis
abstract
Computer Vision problems applied to visual surveillance have been studied for several years aiming at finding accurate and efficient solutions, required to allow the execution of surveillance systems in real environments. The main goal of such systems is to analyze the scene focusing on the detection and recognition of suspicious activities performed by humans in the scene, so that the security personnel can pay closer attention to these preselected activities. To accomplish that, several problems have to be solved first, for instance background subtraction, person detection, tracking and re-identification, face recognition, and action recognition. Even though each of these problems have been researched in the past decades, they are hardly considered in a sequence, each one is usually solved individually. However, in a real surveillance scenarios, the aforementioned problems have to be solved in sequence considering only videos as the input. Aiming at the direction of evaluating approaches in more realistic scenarios, this work proposes a framework called Smart Surveillance Framework (SSF), to allow researchers to implement their solutions to the above problems as a sequence of processing modules that communicate through a shared memory.
Antonio C. Nazare, Cassio E. dos Santos, Renato Ferreira 0001, William Robson Schwartz
WACV4
2014 Evaluating the use of ECG signal in low frequencies as a biometry
Eduardo José da S. Luz, David Menotti, William Robson Schwartz
Expert Syst. Appl.3
2013 Person Re-identification Using Partial Least Squares Appearance Modeling
Gabriel Lorencetti Prado, William Robson Schwartz, Hélio Pedrini
CIARP (2)2
2013 Fast pedestrian detection based on a partial least squares cascade
abstract
In applications such as surveillance, pedestrian detection can be seen as a filtering stage which will locate the objects of interest so that higher level tasks, such as recognition, re-identification, action and activity recognition, can be performed considering only those objects. Therefore, it is imperative that the pedestrian detection task presents low computational cost. Several methods have been proposed to detect pedestrians in images and videos. However, a remaining challenge is to detect pedestrians with high accuracy at a very low computational cost. Towards accomplishing the goal of reducing the costs for pedestrian detection, we propose a cascade of rejection based on Partial Least Squares (PLS) and the variable selection method Variable Importance in Projection (VIP) combined with the propagation of latent variables through the stages. Our results show that the method reduces the computational cost by increasing the number of rejected background samples in earlier stages of the cascade.
Victor C. de Melo, Samir Leao, Mario Fernando Montenegro Campos, David Menotti, William Robson Schwartz
ICIP5
2013 A data-driven detection optimization framework
William Robson Schwartz, Victor C. de Melo, Hélio Pedrini, Larry Davis 0001
Neurocomputing1
2013 Multi-scale gray level co-occurrence matrices for texture description
Fernando Roberti de Siqueira, William Robson Schwartz, Hélio Pedrini
Neurocomputing2
2013 Adaptive edge-preserving image denoising using wavelet transforms
Ricardo Dutra da Silva, Rodrigo Minetto, William Robson Schwartz, Hélio Pedrini
Pattern Anal. Appl.3
2012 Person-Specific Subspace Analysis for Unconstrained Familiar Face Identification
Giovani Chiachia, Nicolas Pinto, William Robson Schwartz, Anderson Rocha 0001, Alexandre X. Falcão, David D. Cox
BMVC3
2012 Scalable people re-identification based on a one-against-some classification scheme
abstract
People re-identification is a problem of increasing interest in computer vision, mainly in applications such as video surveillance and dynamic environment monitoring. However, the large amount of data captured from multiple cameras, the large number of agents involved and poor acquisition conditions make it a difficult problem to solve. Recent works have shown that the use of multiple feature extraction methods combined by a weighting technique considering a one-against-all classification scheme provide accurate results for applications such as face recognition and appearance-based modeling. However, to enroll new subjects, all models need to be rebuild, which results in an increasingly computational time. To reduce this problem, this work proposes a classification scheme, called one-against-some, to allow scalable enrollment of new individuals without reducing the accuracy when compared to the one-against-all classification scheme.
William Robson Schwartz
ICIP1
2012 Scalar image interest point detection and description based on discrete Morse theory and geometric descriptors
abstract
The use of scalar data has arisen in many image applications which obtain data from simulated experiments, range scanners and photogrammetry. Their different nature of acquisition influences the type of processing and the analysis that may be undertaken in tasks such as registration, retrieval and recognition of structures or objects. This paper presents a new method for detecting and describing scale invariant descriptors over dense scalar data sets. The detection of interest points for each image is accomplished using an approach based on the Morse theory and descriptors are computed exploring the geometry of the data. Experiments are performed to demonstrate the effectiveness of the proposed method.
Ricardo Dutra da Silva, William Robson Schwartz, Hélio Pedrini
ICIP2
2012 EDVD - Enhanced descriptor for visual and depth data
Erickson R. Nascimento, William Robson Schwartz, Mario Fernando Montenegro Campos
ICPR2
2012 Distance matrices as invariant features for classifying MoCap data
Antônio Wilson Vieira, Thomas Lewiner, William Robson Schwartz, Mario Fernando Montenegro Campos
ICPR3
2012 BRAND: A robust appearance and depth descriptor for RGB-D images
abstract
This work introduces a novel descriptor called Binary Robust Appearance and Normals Descriptor (BRAND), that efficiently combines appearance and geometric shape information from RGB-D images, and is largely invariant to rotation and scale transform. The proposed approach encodes point information as a binary string providing a descriptor that is suitable for applications that demand speed performance and low memory consumption. Results of several experiments demonstrate that as far as precision and robustness are concerned, BRAND achieves improved results when compared to state of the art descriptors based on texture, geometry and combination of both information. We also demonstrate that our descriptor is robust and provides reliable results in a registration task even when a sparsely textured and poorly illuminated scene is used.
Erickson R. Nascimento, Gabriel L. Oliveira, Mario Fernando Montenegro Campos, Antônio Wilson Vieira, William Robson Schwartz
IROS5
2012 A complementary local feature descriptor for face identification
abstract
In many descriptors, spatial intensity transforms are often packed into a histogram or encoded into binary strings to be insensitive to local misalignment and compact. Discriminative information, however, might be lost during the process as a trade-off. To capture the lost pixel-wise local information, we propose a new feature descriptor, Circular Center Symmetric-Pairs of Pixels (CCS-POP). It concatenates the symmetric pixel differences centered at a pixel position along various orientations with various radii; it is a generalized form of Local Binary Patterns, its variants and Pairs-of-Pixels (POP). Combining CCS-POP with existing descriptors achieves better face identification performance on FRGC Ver. 1.0 and FERET datasets compared to state-of-the-art approaches.
William Robson Schwartz, Huimin Guo, Larry Davis 0001
WACV2
2012 Semi-Supervised Dimensionality Reduction based on Partial Least Squares for Visual Analysis of High Dimensional Data
abstract
Abstract Dimensionality reduction is employed for visual data analysis as a way to obtaining reduced spaces for high dimensional data or to mapping data directly into 2D or 3D spaces. Although techniques have evolved to improve data segregation on reduced or visual spaces, they have limited capabilities for adjusting the results according to user's knowledge. In this paper, we propose a novel approach to handling both dimensionality reduction and visualization of high dimensional data, taking into account user's input. It employs Partial Least Squares (PLS), a statistical tool to perform retrieval of latent spaces focusing on the discriminability of the data. The method employs a training set for building a highly precise model that can then be applied to a much larger data set very effectively. The reduced data set can be exhibited using various existing visualization techniques. The training data is important to code user's knowledge into the loop. However, this work also devises a strategy for calculating PLS reduced spaces when no training data is available. The approach produces increasingly precise visual mappings as the user feeds back his or her knowledge and is capable of working with small and unbalanced training sets.
Jose Gustavo Paiva, William Robson Schwartz, Hélio Pedrini, Rosane Minghim
Comput. Graph. Forum2
2012 Face Identification Using Large Feature Sets
abstract
With the goal of matching unknown faces against a gallery of known people, the face identification task has been studied for several decades. There are very accurate techniques to perform face identification in controlled environments, particularly when large numbers of samples are available for each face. However, face identification under uncontrolled environments or with a lack of training data is still an unsolved problem. We employ a large and rich set of feature descriptors (with more than 70,000 descriptors) for face identification using partial least squares to perform multichannel feature weighting. Then, we extend the method to a tree-based discriminative structure to reduce the time required to evaluate probe samples. The method is evaluated on Facial Recognition Technology (FERET) and Face Recognition Grand Challenge (FRGC) data sets. Experiments show that our identification method outperforms current state-of-the-art results, particularly for identifying faces acquired across varying conditions.
William Robson Schwartz, Huimin Guo, Larry Davis 0001
IEEE Trans. Image Process.1
2011 Local Response Context Applied to Pedestrian Detection
William Robson Schwartz, Larry Davis 0001, Hélio Pedrini
CIARP1
2011 Competition on counter measures to 2-D facial spoofing attacks
abstract
Spoofing identities using photographs is one of the most common techniques to attack 2-D face recognition systems. There seems to exist no comparative studies of different techniques using the same protocols and data. The motivation behind this competition is to compare the performance of different state-of-the-art algorithms on the same database using a unique evaluation method. Six different teams from universities around the world have participated in the contest. Use of one or multiple techniques from motion, texture analysis and liveness detection appears to be the common trend in this competition. Most of the algorithms are able to clearly separate spoof attempts from real accesses. The results suggest the investigation of more complex attacks.
Murali Mohan Chakka, André Anjos, Sébastien Marcel, Roberto Tronci, Daniele Muntoni, Gianluca Fadda, Maurizio Pili, Nicola Sirena, Gabriele Murgia, Marco Ristori, Fabio Roli, Dong Yi, Zhen Lei 0001, Stan Z. Li, William Robson Schwartz, Anderson Rocha 0001, Hélio Pedrini, Javier Lorenzo-Navarro, Modesto Castrillón-Santana, Jukka Komulainen, Abdenour Hadid, Matti Pietikäinen
IJCB17
2011 Face verification using large feature sets and one shot similarity
abstract
We present a method for face verification that combines Partial Least Squares (PLS) and the One-Shot similarity model[28]. First, a large feature set combining shape, texture and color information is used to describe a face. Then PLS is applied to reduce the dimensionality of the feature set with multi-channel feature weighting. This provides a discriminative facial descriptor. PLS regression is used to compute the similarity score of an image pair by One-Shot learning. Given two feature vector representing face images, the One-Shot algorithm learns discriminative models exclusively for the vectors being compared. A small set of unlabeled images, not containing images belonging to the people being compared, is used as a reference (negative) set. The approach is evaluated on the Labeled Face in the Wild (LFW) benchmark and shows very comparable results to the state-of-the-art methods (achieving 86.12% classification accuracy) while maintaining simplicity and good generalization ability.
Huimin Guo, William Robson Schwartz, Larry Davis 0001
IJCB2
2011 Face spoofing detection through partial least squares and low-level descriptors
abstract
Personal identity verification based on biometrics has received increasing attention since it allows reliable authentication through intrinsic characteristics, such as face, voice, iris, fingerprint, and gait. Particularly, face recognition techniques have been used in a number of applications, such as security surveillance, access control, crime solving, law enforcement, among others. To strengthen the results of verification, biometric systems must be robust against spoofing attempts with photographs or videos, which are two common ways of bypassing a face recognition system. In this paper, we describe an anti-spoofing solution based on a set of low-level feature descriptors capable of distinguishing between 'live' and 'spoof images and videos. The proposed method explores both spatial and temporal information to learn distinctive characteristics between the two classes. Experiments conducted to validate our solution with datasets containing images and videos show results comparable to state-of-the-art approaches.
William Robson Schwartz, Anderson Rocha 0001, Hélio Pedrini
IJCB1
2011 A novel feature descriptor based on the shearlet transform
abstract
Problems such as image classification, object detection and recognition rely on low-level feature descriptors to represent visual information. Several feature extraction methods have been proposed, including the Histograms of Oriented Gradients (HOG), which captures edge information by analyzing the distribution of intensity gradients and their directions. In addition to directions, the analysis of edge at different scales provides valuable information. Shearlet transforms provide a general framework for analyzing and representing data with anisotropic information at multiple scales. As a consequence, signal singularities, such as edges, can be precisely detected and located in images. Based on the idea of employing histograms to estimate the distribution of edge orientations and on the accurate multi-scale analysis provided by shearlet transforms, we propose a feature descriptor called Histograms of Shearlet Coefficients (HSC). Experimental results comparing HOG with HSC show that HSC provides significantly better results for the problems of texture classification and face identification.
William Robson Schwartz, Ricardo Dutra da Silva, Larry Davis 0001, Hélio Pedrini
ICIP1
2010 A Robust and Scalable Approach to Face Identification
William Robson Schwartz, Huimin Guo, Larry Davis 0001
ECCV (6)1
2009 Human detection using partial least squares analysis
abstract
Significant research has been devoted to detecting people in images and videos. In this paper we describe a human detection method that augments widely used edge-based features with texture and color information, providing us with a much richer descriptor set. This augmentation results in an extremely high-dimensional feature space (more than 170,000 dimensions). In such high-dimensional spaces, classical machine learning algorithms such as SVMs are nearly intractable with respect to training. Furthermore, the number of training samples is much smaller than the dimensionality of the feature space, by at least an order of magnitude. Finally, the extraction of features from a densely sampled grid structure leads to a high degree of multicollinearity. To circumvent these data characteristics, we employ Partial Least Squares (PLS) analysis, an efficient dimensionality reduction technique, one which preserves significant discriminative information, to project the data onto a much lower dimensional subspace (20 dimensions, reduced from the original 170,000). Our human detection system, employing PLS analysis over the enriched descriptor set, is shown to outperform state-of-the-art techniques on three varied datasets including the popular INRIA pedestrian dataset, the low-resolution gray-scale DaimlerChrysler pedestrian dataset, and the ETHZ pedestrian dataset consisting of full-length videos of crowded scenes.
William Robson Schwartz, Aniruddha Kembhavi, David Harwood, Larry Davis 0001
ICCV1
2006 Textured Image Segmentation Based on Spatial Dependence using a Markov Random Field Model
abstract
Image segmentation is a primary step in many computer vision tasks. Although many segmentation methods have been proposed in the last decades, there is no generic method that can be applied in a great variety of images. This work presents a new image segmentation method using texture features extracted by wavelet transforms combined with spatial dependence modeled by a Markov random field (MRF). The method initially produces a coarse segmentation, which is refined through a relaxation method based on a new energy function. A set of textured images is used to demonstrate the effectiveness of the proposed method.
William Robson Schwartz, Hélio Pedrini
ICIP1
2004 Texture classification based on spatial dependence features using co-occurrence matrices and markov random fields
abstract
This paper presents a method for classification of textures based on features obtained from co-occurrence matrices and Markov random fields. Two steps are performed to classify the images. Initially, the method recognizes the homogeneous regions (object interior) in the image. Regions consisting of dissimilar elements (transition between objects) are then properly identified and classified. Experimental results demonstrate the robustness of the method in terms of variation in region size and number of parameters.
William Robson Schwartz, Hélio Pedrini
ICIP1
2002 Fitting smooth surfaces to scattered 3D data using piecewise quadratic approximation
abstract
The approximation of surfaces to scattered data is an important problem encountered in a variety of scientific applications, such as reverse engineering, computer vision, computer graphics, and terrain modeling. This paper describes an automatic method for constructing smooth surfaces defined as a network of curved triangular patches. The method starts with a coarse mesh approximating the surface through triangular elements covering the boundary of the domain, then iteratively adds new points from the data set until a specified error tolerance is achieved. The resulting surface over the triangular mesh is represented by piecewise polynomial patches possessing C/sup 1/ continuity. The method has been implemented and tested on a number of real data sets.
Oliver van Kaick, Murilo V. G. da Silva, Hélio Pedrini, William Robson Schwartz
ICIP (1)4
2001 Topographic Feature Identification Based on Triangular Meshes
Hélio Pedrini, William Robson Schwartz
CAIP2
2001 Automatic extraction of topographic features using adaptive triangular meshes
abstract
A method is described for the extraction of morphological information from images approximated by triangular meshes. Topographic features such as peaks, pits, ridges, valleys, and planar regions are considered the basic descriptive surface elements and are defined in terms of the local organization of the triangles in the mesh. The approach is suitable for image analysis tasks, simplifying object recognition and scene interpretation. Several images have been used to demonstrate the performance of the proposed method.
Hélio Pedrini, William Robson Schwartz, W. Randolph Franklin
ICIP (3)2