EDBT 2026 Demo / reviewers in the wild / expert
Anderson Rocha 0001
dblp:81/2446 · also Anderson R. Rocha 0001, Anderson de Rezende Rocha
· DBLP profile ↗
135ranked-venue papers
8as first author
41since 2021 · last 2026
0000-0002-4236-8212ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 63 · 4 first-author · 12 since 2021Artificial intelligence and machine learning · 45 · 4 first-author · 15 since 2021Security and privacy · 35 · 1 first-author · 14 since 2021Human-computer interaction and ubiquitous computing · 9 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GS-Checker: Tampering Localization for 3D Gaussian SplattingabstractRecent advances in editing technologies for 3D Gaussian Splatting (3DGS) have made it simple to manipulate 3D scenes. However, these technologies raise concerns about potential malicious manipulation of 3D content. To avoid such malicious applications, localizing tampered regions becomes crucial. In this paper, we propose GS-Checker, a novel method for locating tampered areas in 3DGS models. Our approach integrates a 3D tampering attribute into the 3D Gaussian parameters to indicate whether the Gaussian has been tampered. Additionally, we design a 3D contrastive mechanism by comparing the similarity of key attributes between 3D Gaussians to seek tampering cues at 3D level. Furthermore, we introduce a cyclic optimization strategy to refine the 3D tampering attribute, enabling more accurate tampering localization. Notably, our approach does not require expensive 3D labels for supervision. Extensive experimental results demonstrate the effectiveness of our proposed method to locate the tampered 3DGS area. Haoliang Han, Ziyuan Luo, Anderson Rocha 0001, Renjie Wan |
AAAI | 4 |
| 2026 | LithoVision: a novel approach for carbonate rock classification in the Brazilian pre-saltabstractAbstract Classifying lithotypes from a reservoir is crucial for assessing drilling, production, and operational risks. Traditionally, accurately identifying the rock lithologies has been a manual, time-consuming, and error-prone process. Recent advancements in deep learning-based and computer vision have shown promising results in classifying rock lithology. However, many state-of-the-art techniques struggle when working with limited, imbalanced, and fine-grained rock datasets, revealing room for improvement. To address these limitations, this paper proposes an innovative strategy that integrates the Vision Transformer and the Central Difference Convolution architectures to classify the major lithology classes in carbonate rock images. The proposed method combines an Encoder Transformer with a Convolutional Neural Network to learn invariant representations of carbonate textures, eliminating manual feature extraction and performing effectively for small and large datasets. We conducted extensive experimental evaluations on two real-world datasets, demonstrating that: (1) our approach is competitive on fine-grained rock datasets; (2) it surpasses most baseline methods in most evaluation scenarios. Specifically, our proposal outperforms the second-best method (ViT) by 8% in the whole-image classification, and exceeds the best CNN model (VGG) by at least 24% in patches classification, considering the balanced accuracy metric. Michael Cruz 0001, Soroor Salavati, Anderson Rocha 0001, Jeferson Santos, Leonardo Borghi |
Neural Comput. Appl. | 4 |
| 2026 | Adversarially robust multimedia watermarking via data-centric optimization
Ziyuan Luo, Qi Song 0003, Haoliang Li, Anderson Rocha 0001, Renjie Wan |
Pattern Recognit. | 4 |
| 2026 | Open-Set Deepfake Detection: A Parameter-Efficient Adaptation Method With Forgery Style MixtureabstractOpen-set face forgery detection poses significant security threats and presents substantial challenges for existing detection models. These detectors primarily have two limitations: they cannot generalize across unknown forgery domains or inefficiently adapt to new data. To address these issues, we introduce an approach that is both general and parameter-efficient for face forgery detection. Our method builds on the assumption that different forgery source domains exhibit distinct style statistics. Specifically, we design a forgery-style-mixture formulation that augments the diversity of forgery source domains, enhancing the model’s generalizability across unseen domains. In addition, previous methods typically require fully fine-tuning pretrained networks, consuming substantial time and computational resources. Drawing on recent advancements in vision transformers (ViT) for face forgery detection, we develop a parameter-efficient ViT-based detection model that includes lightweight forgery feature extraction modules and enables the model to extract global and local forgery clues simultaneously. We only optimize the inserted lightweight modules during training, maintaining the original ViT structure with its pre-trained weights. This training strategy effectively preserves the informative pre-trained knowledge while flexibly adapting the model to the task of Deepfake detection. Extensive experimental results demonstrate that the designed model achieves state-of-the-art generalizability with significantly reduced trainable parameters, representing an important step toward open-set Deepfake detection in the wild. Chenqi Kong, Anwei Luo, Peijun Bao, Haoliang Li, Renjie Wan, Zengwei Zheng, Anderson Rocha 0001, Alex Chichung Kot |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2026 | XNet: Enhancing Physical Activity Intensity Assessment With Attentional Multidomain Fusion and Visual AnalyticsabstractSedentary behavior (SB) is a major global health concern, necessitating accurate physical activity (PA) intensity monitoring. Conventional machine-learning (ML) methods using accelerometers struggle to generalize due to variability across populations, sensors, and activities, leading to inconsistent real-world performance. This study presents XNet, a dual-domain deep learning (DL) model for classifying PA intensity and estimating energy expenditure. XNet features a hierarchical multihead architecture that independently extracts temporal and frequency features from multiple sensors, then integrates them via a novel attentional feature fusion (AFF) module applied in two stages: first aggregating sensor features, then fusing domain embeddings. This hierarchical approach outperforms single-stage fusion and provides interpretable attention weights revealing sensor and domain contributions. Frequency-domain features are essential for generalization: in cross-dataset evaluations, XNet achieved the highest F1-score of 70.5 while maintaining robust sedentary detection (88% TPR), and in open-set scenarios, it achieved an F1-score of 77.0, surpassing all DL and hand-crafted baselines. We validated XNet on multiple public datasets and a new dataset of 105 participants. Furthermore, our analysis shows that lightweight 1D-convolutional spectral encoders yield better out-of-distribution generalization than transformer and graph attention (GAT) network alternatives, while benchmarking confirms that AFF outperforms nine fusion strategies in balancing accuracy, efficiency, and robustness to sensor failure. The model adapts to physiological signals (heart rate and ECG) and exhibits low inference latency (~25 ms), making it suitable for on-device deployment. A complementary visual analytics framework uses attention weights to facilitate expert auditing, thereby promoting transparent and equitable health monitoring. Alexis Mendoza, Emely Pujólli da Silva, Didier Augusto Vega-Oliveros, T. Frota de Souza, Aurea Soriano-Vargas, M. Uchida, Anderson Rocha 0001 |
IEEE Trans. Cybern. | 7 |
| 2026 | MantleMark: Migrating Watermarks From Multi-View Images to Radiance Fields via Frequency ModulationabstractMulti-view images are essential for modern radiance field reconstruction methods like Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS). While image watermarking is a crucial data protection and ownership verification technique, it faces unprecedented challenges in multi-view scenarios. Traditional 2D watermarking techniques often fail to maintain detectability in rendered views, while existing 3D watermarking methods are typically limited to specific reconstruction methods and require access to the reconstruction process. To address these limitations, we propose MantleMark, a watermarking framework that migrates watermarks from multi-view images to radiance fields via frequency modulation. Our key insight is constructing a mantle-like Frequency-domain Watermarking Representation in 3D frequency space, which can be projected to create view-dependent watermarking patterns. Relying upon the Fourier Projection-Slice Theorem, we embed these patterns through magnitude spectrum modulation in the image frequency domain, enabling watermarks to migrate into 3D representations. This approach ensures watermark detectability in rendered views regardless of the reconstruction methods used by adversaries. Extensive experiments demonstrate that our method achieves robust watermark detection while maintaining high visual quality across various radiance field-based reconstruction methods. Ziyuan Luo, Jun Liu 0036, Haoliang Li, Anderson Rocha 0001, Renjie Wan |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | WeiDetect: Weibull distribution-based defense against poisoning attacks in federated learning for network intrusion detection systems
K. M. Sameera, P. Vinod 0001, Anderson Rocha 0001, Rafidha Rehiman K. A., Mauro Conti |
J. Inf. Secur. Appl. | 3 |
| 2025 | Adaptive loss optimization for enhanced learning performance: application to image-based rock classification
Soroor Salavati, Pedro Ribeiro Mendes Júnior, Anderson Rocha 0001 |
Neural Comput. Appl. | 3 |
| 2025 | Pixel-Inconsistency Modeling for Image Manipulation LocalizationabstractDigital image forensics plays a crucial role in image authentication and manipulation localization. Despite the progress powered by deep neural networks, existing forgery localization methodologies exhibit limitations when deployed to unseen datasets and perturbed images (i.e., lack of generalization and robustness to real-world applications). To circumvent these problems and aid image integrity, this paper presents a generalized and robust manipulation localization model through the analysis of pixel inconsistency artifacts. The rationale is grounded on the observation that most image signal processors (ISP) involve the demosaicing process, which introduces pixel correlations in pristine images. Moreover, manipulating operations, including splicing, copy-move, and inpainting, directly affect such pixel regularity. We, therefore, first split the input image into several blocks and design masked self-attention mechanisms to model the global pixel dependency in input images. Simultaneously, we optimize another local pixel dependency stream to mine local manipulation clues within input forgery images. In addition, we design novel Learning-to-Weight Modules (LWM) to combine features from the two streams, thereby enhancing the final forgery localization performance. To improve the training process, we propose a novel Pixel-Inconsistency Data Augmentation (PIDA) strategy, driving the model to focus on capturing inherent pixel-level artifacts instead of mining semantic forgery traces. This work establishes a comprehensive benchmark integrating 16 representative detection models across 12 datasets. Extensive experiments show that our method successfully extracts inherent pixel-inconsistency forgery fingerprints and achieve state-of-the-art generalization and robustness performances in image manipulation localization. Chenqi Kong, Anwei Luo, Shiqi Wang 0001, Haoliang Li, Anderson Rocha 0001, Alex Chichung Kot |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | The NeRF Signature: Codebook-Aided Watermarking for Neural Radiance FieldsabstractNeural Radiance Fields (NeRF) have been gaining attention as a significant form of 3D content representation. With the proliferation of NeRF-based creations, the need for copyright protection has emerged as a critical issue. Although some approaches have been proposed to embed digital watermarks into NeRF, they often neglect essential model-level considerations and incur substantial time overheads, resulting in reduced imperceptibility and robustness, along with user inconvenience. In this paper, we extend the previous criteria for image watermarking to the model level and propose NeRF Signature, a novel watermarking method for NeRF. We employ a Codebook-aided Signature Embedding (CSE) that does not alter the model structure, thereby maintaining imperceptibility and enhancing robustness at the model level. Furthermore, after optimization, any desired signatures can be embedded through the CSE, and no fine-tuning is required when NeRF owners want to use new binary signatures. Then, we introduce a joint pose-patch encryption watermarking strategy to hide signatures into patches rendered from a specific viewpoint for higher robustness. In addition, we explore a Complexity-Aware Key Selection (CAKS) scheme to embed signatures in high visual complexity patches to enhance imperceptibility. The experimental results demonstrate that our method outperforms other baseline methods in terms of imperceptibility and robustness. Ziyuan Luo, Anderson Rocha 0001, Boxin Shi, Qing Guo 0005, Haoliang Li, Renjie Wan |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Self-Rationalization in the Wild: A Large-scale Out-of-Distribution Evaluation on NLI-related tasksabstractAbstract Free-text explanations are expressive and easy to understand, but many datasets lack annotated explanation data, making it challenging to train models for explainable predictions. To address this, we investigate how to use existing explanation datasets for self-rationalization and evaluate models’ out-of-distribution (OOD) performance. We fine-tune T5-Large and OLMo-7B models and assess the impact of fine-tuning data quality, the number of fine-tuning samples, and few-shot selection methods. The models are evaluated on 19 diverse OOD datasets across three tasks: natural language inference (NLI), fact-checking, and hallucination detection in abstractive summarization. For the generated explanation evaluation, we conduct a human study on 13 selected models and study its correlation with the Acceptability score (T5-11B) and three other LLM-based reference-free metrics. Human evaluation shows that the Acceptability score correlates most strongly with human judgments, demonstrating its effectiveness in evaluating free-text explanations. Our findings reveal: 1) few annotated examples effectively adapt models for OOD explanation generation; 2) compared to sample selection strategies, fine-tuning data source has a larger impact on OOD performance; and 3) models with higher label prediction accuracy tend to produce better explanations, as reflected by higher Acceptability scores.1 Jing Yang 0031, Max Glockner, Anderson Rocha 0001, Iryna Gurevych |
Trans. Assoc. Comput. Linguistics | 3 |
| 2025 | FakeBench: Probing Explainable Fake Image Detection via Large Multimodal ModelsabstractThe ability to distinguish whether an image is generated by artificial intelligence (AI) is a crucial ingredient in human intelligence, usually accompanied by a complex and dialectical forensic and reasoning process. However, current fake image detection models and databases focus on binary classification without understandable explanations for the general populace. This weakens the credibility of authenticity judgment and may conceal potential model biases. Meanwhile, large multimodal models (LMMs) have exhibited immense vision-language capabilities on various tasks, bringing the potential for explainable fake image detection. Therefore, we pioneer the probe of LMMs for explainable fake image detection by presenting a multimodal database encompassing descriptions of textual authenticity, the FakeBench. For construction, we first introduce a fine-grained taxonomy of generative visual forgery concerning human perception, based on which we collect forgery descriptions in human natural language with a human-in-the-loop strategy. FakeBench examines LMMs with four evaluation criteria: detection, reasoning, explanation and fine-grained forgery analysis, to obtain deeper insights into image authenticity-relevant capabilities. Experiments on various LMMs confirm their merits and demerits in different aspects of fake image detection tasks. This research presents a paradigm shift towards transparency for the fake image detection area and reveals the need for greater emphasis on forensic elements in visual-language research and AI risk control. FakeBench will be available athttps://github.com/Yixuan423/FakeBench Xuelin Liu, Xiaoyang Wang 0009, Bu-Sung Lee, Shiqi Wang 0001, Anderson Rocha 0001, Weisi Lin |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2025 | Image Provenance Analysis via Graph Encoding With Vision TransformerabstractRecent advances in AI-powered image editing tools have significantly lowered the barrier to image modification, raising pressing security concerns those related to spreading misinformation and disinformation on social platforms. Image provenance analysis is crucial in this context, as it identifies relevant images within a database and constructs a relationship graph by mining hidden manipulation and transformation cues, thereby providing concrete evidence chains. This paper introduces a novel end-to-end deep learning framework designed to explore the structural information of provenance graphs. Our proposed method distinguishes from previous approaches in two main ways. First, unlike earlier methods that rely on prior knowledge and have limited generalizability, our framework relies upon a patch attention mechanism to capture image provenance clues for local manipulations and global transformations, thereby enhancing graph construction performance. Second, while previous methods primarily focus on identifying tampering traces only between image pairs, they often overlook the hidden information embedded in the topology of the provenance graph. Our approach aligns the model training objectives with the final graph construction task, incorporating the overall structural information of the graph into the training process. We integrate graph structure information with the attention mechanism, enabling precise determination of the direction of transformation. Experimental results show the superiority of the proposed method over previous approaches, underscoring its effectiveness in addressing the challenges of image provenance analysis. Keyang Zhang, Chenqi Kong, Shiqi Wang 0001, Anderson Rocha 0001, Haoliang Li |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | Introduction to the Special Issue on Security and Privacy of Avatar in MetaverseabstractThe Metaverse is a 3D interactive virtual community that has gained significant attention in academia, business, and industry as a potential future internet paradigm. In this space, avatars serve as key elements, acting as the primary means of human interaction. Avatars are expected to be created using real data, tailored to users' preferences, and controlled in real-time through signals from wearable devices. Avatars allow users to feel as though they are extensions of their own bodies, creating an immersive experience that blurs the line between virtual and real compared to other virtual communities. On the other hand, the avatar faces serious security and privacy problems, especially when people and the law/regulation are increasingly less tolerant of security and privacy, such as copyright, false identity detection, dataset security, authentication, and content tampering. This special issue collects 15 papers reporting the recent developments of security and privacy of avatar in metaverse. For the Avatar Copyright Protection. "A Self-Defense Copyright Protection Scheme for NFT Image Art Based on Information Embedding" addresses copyright issues related to avatars produced in the Metaverse and proposes a copyright protection scheme that not only enables tracking and verification of avatar content transactions but also validates the legality of the source and ownership of the avatar content. "Invisible Adversarial Watermarking: A Novel Security Mechanism for Enhancing Copyright Protection" addresses the potential for unauthorized access and use of image datasets used to generate avatars and proposes an image protection method that combines adversarial perturbations with invisible watermarks. This approach not only prevents illegal use of the image datasets but also enables effective tracking of data copyright. In "FaceDefend: Copyright Protection to Prevent Face Embezzle, " the authors propose a solution to the misuse problem arising from the theft of real facial image data used in avatar generation, based on defensive strategies. This approach effectively ensures copyright protection for real facial data. For the False Identity Detection for Avatars. The authors of "Audio-Visual Contrastive Pre-train for Face Forgery Detection" address the issue of potential facial privacy breaches due to the realism of avatars in virtual worlds, which can lead. Yushu Zhang 0001, William Puech, Anderson Rocha 0001, Rongxing Lu, Stefano Cresci, Roberto Di Pietro |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Open-Set Deepfake Detection To Fight The UnknownabstractIn this paper, we design a new open-set method to detect deepfakes that does not assume information about the techniques behind the deepfakes generation. Contrary to existing methods, which build upon known telltales left by the deepfake creation process, we assume no prior knowledge about the sample generation, thus presenting a method for blind deepfake detection, a necessary step toward true generalization. Our methodology relies upon unsupervised learning, open-set formulations for each discovered group and, finally, relevance tests through extreme-value theory and isolation forest formulations. The results indicate that the proposed open-set technique is competitive with state-of-the-art closed-set deepfake detection methods. As a notable outcome, we achieved an AUC = 0.807, which is 5.57% higher than the baseline architecture trained using a closed-set approach. Finally, we believe our efforts herein are just a first attempt tackling this difficult problem and discuss some additional improvements for practical deployment of such systems. Michael Macedo Diniz, Anderson Rocha 0001 |
ICASSP | 2 |
| 2024 | A multi-modal approach for mixed-frequency time series forecasting
Leopoldo Lusquino Filho, Rafael de Oliveira Werneck, Manuel Castro 0002, Pedro Ribeiro Mendes Júnior, Augusto Lustosa, Marcelo Ferreira Zampieri, Oscar Linares, Renato Moura, Elayne Morais, Murilo Amaral, Soroor Salavati, Ashish Loomba, Ahmed Esmin, Maiara Moreira Gonçalves, Denis J. Schiozer, Alessandra Davólio, Anderson Rocha 0001 |
Neural Comput. Appl. | 18 |
| 2024 | Few-shot learning and modeling of 3D reservoir properties for predicting oil reservoir production
Gabriel Cirac Mendes Souza, Guilherme Daniel Avansi, Jeanfranco Farfan, Denis J. Schiozer, Anderson Rocha 0001 |
Neural Comput. Appl. | 5 |
| 2024 | M$^{3}$3FAS: An Accurate and Robust MultiModal Mobile Face Anti-Spoofing SystemabstractFace presentation attacks (FPA), also known as face spoofing, have brought increasing concerns to the public through various malicious applications, such as financial fraud and privacy leakage. Therefore, safeguarding face recognition systems against FPA is of utmost importance. Although existing learning-based face anti-spoofing (FAS) models can achieve outstanding detection performance, they lack generalization capability and suffer significant performance drops in unforeseen environments. Many methodologies seek to use auxiliary modality data (e.g., depth and infrared maps) during the presentation attack detection (PAD) to address this limitation. However, these methods can be limited since (1) they require specific sensors such as depth and infrared cameras for data capture, which are rarely available on commodity mobile devices, and (2) they cannot work properly in practical scenarios when either modality is missing or of poor quality. In this paper, we devise an accurate and robustMultiModalMobileFaceAnti-Spoofing system namedM$^{3}$FASto overcome the issues above. The primary innovation of this work lies in the following aspects: (1) To achieve robust PAD, our system combines visual and auditory modalities using three commonly available sensors: camera, speaker, and microphone; (2) We design a novel two-branch neural network with three hierarchical feature aggregation modules to perform cross-modal feature fusion; (3). We propose a multi-head training strategy, allowing the model to output predictions from the vision, acoustic, and fusion heads, resulting in a more flexible PAD. Extensive experiments have demonstrated the accuracy, robustness, and flexibility of M$^{3}$FAS under various challenging experimental settings. The source code and dataset are available at:https://github.com/ChenqiKONG/M3FAS/. Chenqi Kong, Kexin Zheng, Yibing Liu, Shiqi Wang 0001, Anderson Rocha 0001, Haoliang Li |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2024 | Robust Domain Misinformation Detection via Multi-Modal Feature AlignmentabstractSocial media misinformation harms individuals and societies and is potentialized by fast-growing multi-modal content (i.e., texts and images), which accounts for higher “credibility” than text-only news pieces. Although existing supervised misinformation detection methods have obtained acceptable performances in key setups, they may require large amounts of labeled data from various events, which can be time-consuming and tedious. In turn, directly training a model by leveraging a publicly available dataset may fail to generalize due to domain shifts between the training data (a.k.a. source domains) and the data from target domains. Most prior work on domain shift focuses on a single modality (e.g., text modality) and ignores the scenario where sufficient unlabeled target domain data may not be readily available in an early stage. The lack of data often happens due to the dynamic propagation trend (i.e., the number of posts related to fake news increases slowly before catching the public attention). We propose a novel robust domain and cross-modal approach (RDCM) for multi-modal misinformation detection. It reduces the domain shift by aligning the joint distribution of textual and visual modalities through an inter-domain alignment module and bridges the semantic gap between both modalities through a cross-modality alignment module. We also propose a framework that simultaneously considers application scenarios of domain generalization (in which the target domain data is unavailable) and domain adaptation (in which unlabeled target domain data is available). Evaluation results on two public multi-modal misinformation detection datasets (Pheme and Twitter Datasets) evince the superiority of the proposed model. Hui Liu 0036, Wenya Wang 0001, Anderson Rocha 0001, Haoliang Li |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2023 | AG-ReID 2023: Aerial-Ground Person Re-identification Challenge ResultsabstractPerson re-identification (Re-ID) on aerial-ground platforms has emerged as an intriguing topic within computer vision, presenting a plethora of unique challenges. Highflying altitudes of aerial cameras make persons appear differently in terms of viewpoints, poses, and resolution compared to the images of the same person viewed from ground cameras. Despite its potential, few algorithms have been developed for person re-identification on aerial-ground data, mainly due to the absence of comprehensive datasets. In response, we have collected a large-scale dataset and organized the Aerial-Ground person Re-IDentification Challenge (AG-ReID2023) to foster advancements in the field. The dataset comprises 100,502 images with 1,615 unique identities, including 51,530 training images featuring 807 identities. The test set is divided into two subsets: Aerial to Ground (808 ids, 4,348 query images, 19,259 gallery images) and Ground to Aerial (808 ids, 4,151 query images, 21,214 gallery images). In addition, we manually annotate individuals with their matching IDs across cameras and provide 15 soft attribute labels. The AG-ReID2023 Challenge in conjunction with the 7thIEEE International Joint Conference on Biometrics (IJCB) has garnered interest from numerous institutes, resulting in the submission of five distinct algorithms. We provide an in-depth examination of the evaluation outcomes and present our findings from the contest. For additional details, kindly refer to the official website1.1https://agreid23.github.io. Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan, Feng Liu 0037, Xiaoming Liu 0002, Arun Ross, Dana Michalski, Debayan Deb, Mahak Kothari, Manisha Saini, Dawei Du, Scott McCloskey, Gabriel Bertocco, Fernanda A. Andaló, Terrance E. Boult, Anderson Rocha 0001, Haidong Zhu, Zhaoheng Zheng, Ramakant Nevatia, Zaigham A. Randhawa, Sinan Sabri, Gianfranco Doretto |
IJCB | 17 |
| 2023 | Deep hierarchical distillation proxy-oil modeling for heterogeneous carbonate reservoirs
Gabriel Cirac Mendes Souza, Jeanfranco Farfan, Guilherme Daniel Avansi, Denis J. Schiozer, Anderson Rocha 0001 |
Eng. Appl. Artif. Intell. | 5 |
| 2023 | Authorship Attribution of Social Media MessagesabstractThe world is facing a new era in which social media communication plays a fundamental role in people’s lives. Along with proven benefits, several collateral drawbacks have risen, one being the widespread of false information with malicious intents, oftentimes using anonymous or false identities. Fighting this problem is challenging, especially when considering the nature of text messages involved on social media platforms: a sea of small messages and a myriad of users. Attributing the authorship of such messages is an ambitious endeavor; nevertheless, it is a way to fight this undesired disinformation scenario. In this work, we tackle the problem of authorship attribution of tiny messages, but, different from what has been done with longer texts, we rely upon data-driven approaches, avoiding handcraft features and harnessing recent advances of deep neural networks in the field of pattern recognition. By modeling small texts employed in social media as unidimensional signals, we propose a deep learning model to project these messages onto a manifold suitable for the task of authorship attribution. We provide two state-of-the-art solutions tailored for different setups and strategies for the scenario of authorship verification. These advances were possible, thanks to three additional contributions: an updated dataset based on the Twitter® platform, new sanitization techniques to improve the quality of the training data, and novel visual analytics techniques to help the development of authorship attribution solutions. Antonio Theophilo, Romain Giot, Anderson Rocha 0001 |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2023 | Leveraging Ensembles and Self-Supervised Learning for Fully-Unsupervised Person Re-Identification and Text Authorship AttributionabstractLearning from fully-unlabeled data is challenging in Multimedia Forensics problems, such as Person Re-Identification and Text Authorship Attribution. Recent self-supervised learning methods have shown to be effective when dealing with fully-unlabeled data in cases where the underlying classes have significant semantic differences, as intra-class distances are substantially lower than inter-class distances. However, this is not the case for forensic applications in which classes have similar semantics and the training and test sets have disjoint identities. General self-supervised learning methods might fail to learn discriminative features in this scenario, thus requiring more robust strategies. We propose a strategy to tackle Person Re-Identification and Text Authorship Attribution by enabling learning from unlabeled data even when samples from different classes are not prominently diverse. We propose a novel ensemble-based clustering strategy whereby clusters derived from different configurations are combined to generate a better grouping for the data samples in a fully-unsupervised way. This strategy allows clusters with different densities and higher variability to emerge, reducing intra-class discrepancies without requiring the burden of finding an optimal configuration per dataset. We also consider different Convolutional Neural Networks for feature extraction and subsequent distance computations between samples. We refine these distances by incorporating context and grouping them to capture complementary information. Our method is robust across both tasks, with different data modalities, and outperforms state-of-the-art methods with a fully-unsupervised solution without any labeling or human intervention. Gabriel Bertocco, Antonio Theophilo, Fernanda A. Andaló, Anderson Rocha 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2022 | Explainable Artificial Intelligence for Authorship Attribution on Social MediaabstractOne of the major modern threats to society is the propagation of misinformation — fake news, science denialism, hate speech — fueled by social media’s widespread adoption. On the leading social platforms, millions of automated and fake profiles exist only for this purpose. One step to mitigate this problem is verifying the authenticity of profiles, which proves to be an infeasible task to be done manually. Recent data-driven methods accurately tackle this problem by performing automatic authorship attribution, although an important aspect is often overlooked: model interpretability. Is it possible to make the decision process of such methods transparent and interpretable for social media content considering its specificities? In this work, we extend upon LIME — a model-agnostic interpretability technique — to improve the explanations of the state-of-the-art methods for authorship attribution on social media posts. Our extension allows us to employ the same input representation of the model as interpretable features, identifying important elements for the authorship process. We also allow coping with the lack of perturbed samples in the scenario of short messages. Finally, we show qualitative and quantitative evidence of these findings. Antonio Theophilo, Rafael Padilha, Fernanda A. Andaló, Anderson Rocha 0001 |
ICASSP | 4 |
| 2022 | Explainable Fact-Checking Through Question AnsweringabstractMisleading or false information has been creating chaos in some places around the world. To mitigate this issue, many researchers have proposed automated fact-checking methods to fight the spread of fake news. However, most methods cannot explain the reasoning behind their decisions, failing to build trust between machines and humans using such technology. Trust is essential for fact-checking to be applied in the real world. Here, we address fact-checking explainability through question answering. In particular, we propose generating questions and answers from claims and answering the same questions from evidence. We also propose an answer comparison model with an attention mechanism attached to each question. Leveraging question answering as a proxy, we break down automated fact-checking into several steps — this separation aids models’ explainability as it allows for more detailed analysis of their decision-making processes. Experimental results show that the proposed model can achieve state-of-the-art performance while providing reasonable explainable capabilities1. Jing Yang 0031, Didier Augusto Vega-Oliveros, Tais Seibt, Anderson Rocha 0001 |
ICASSP | 4 |
| 2022 | Cross-dataset emotion recognition from facial expressions through convolutional neural networks
William Dias, Fernanda A. Andaló, Rafael Padilha, Gabriel Bertocco, Waldir R. de Almeida, Paula Dornhofer Paro Costa, Anderson Rocha 0001 |
J. Vis. Commun. Image Represent. | 7 |
| 2022 | Detect and Locate: Exposing Face Manipulation by Semantic- and Noise-Level TelltalesabstractThe technological advancements of deep learning have enabled sophisticated face manipulation schemes, raising severe trust issues and security concerns in modern society. Generally speaking, detecting manipulated faces and locating the potentially altered regions are challenging tasks. Herein, we propose a conceptually simple but effective method to efficiently detect forged faces in an image while simultaneously locating the manipulated regions. The proposed scheme relies on a segmentation map that delivers meaningful high-level semantic information clues about the image. Furthermore, a noise map is estimated, playing a complementary role in capturing low-level clues and subsequently empowering decision-making. Finally, the features from these two modules are combined to distinguish fake faces. Extensive experiments show that the proposed model achieves state-of-the-art detection accuracy and remarkable localization performance. Chenqi Kong, Baoliang Chen, Haoliang Li, Shiqi Wang 0001, Anderson Rocha 0001, Sam Kwong |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2022 | Beyond the Pixel World: A Novel Acoustic-Based Face Anti-Spoofing System for Smartphonesabstract2D face presentation attacks are one of the most notorious and pervasive face spoofing types, which have caused pressing security issues to facial authentication systems. While RGB-based face anti-spoofing (FAS) models have proven to counter the face spoofing attack effectively, most existing FAS models suffer from the overfitting problem (i.e., lack generalization capability to data collected from an unseen environment). Recently, many models have been devoted to capturing auxiliary information (e.g., depth and infrared maps) to achieve a more robust face liveness detection performance. However, these methods require expensive sensors and cost extra hardware to capture the specific modality information, limiting their applications in practical scenarios. To tackle these problems, we devise a novel and cost-effective FAS system based on the acoustic modality, named Echo-FAS, which employs the crafted acoustic signal as the probe to perform face liveness detection. We first propose to build a large-scale, high-diversity, and acoustic-based FAS database, Echo-Spoof. Then, based upon Echo-Spoof, we propose designing a novel two-branch framework that combines the global and local frequency clues of input signals to distinguish inputs, live vs. spoofing faces accurately. The devised Echo-FAS comprises the following three merits: (1) It only needs one available speaker and microphone as sensors while not requiring any expensive hardware; (2) It can successfully capture the 3D geometrical information of input queries and achieve a remarkable face anti-spoofing performance; and (3) It can be handily allied with other RGB-based FAS models to mitigate the overfitting problem in the RGB modality and make the FAS model more accurate and robust. Our proposed Echo-FAS provides new insights regarding the development of FAS systems for mobile devices. Chenqi Kong, Kexin Zheng, Shiqi Wang 0001, Anderson Rocha 0001, Haoliang Li |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2022 | Content-Aware Detection of Temporal Metadata ManipulationabstractMost pictures shared online are accompanied by temporal metadata (i.e., the day and time they were taken), which makes it possible to associate an image content with real-world events. Maliciously manipulating this metadata can convey a distorted version of reality. In this work, we present the emerging problem of detecting timestamp manipulation. We propose an end-to-end approach to verify whether the purported time of capture of an outdoor image is consistent with its content and geographic location. We consider manipulations done in the hour and/or month of capture of a photograph. The central idea is the use of supervised consistency verification, in which we predict the probability that the image content, capture time, and geographical location are consistent. We also include a pair of auxiliary tasks, which can be used to explain the network decision. Our approach improves upon previous work on a large benchmark dataset, increasing the classification accuracy from 59.0% to 81.1%. We perform an ablation study that highlights the importance of various components of the method, showing what types of tampering are detectable using our approach. Finally, we demonstrate how the proposed method can be employed to estimate a possible time-of-capture in scenarios in which the timestamp is missing from the metadata. Rafael Padilha, Tawfiq Salem, Scott Workman, Fernanda A. Andaló, Anderson Rocha 0001, Nathan Jacobs |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2022 | Link Prediction Based on Stochastic Information DiffusionabstractLink prediction (LP) in networks aims at determining future interactions among elements; it is a critical machine-learning tool in different domains, ranging from genomics to social networks to marketing, especially in e-commerce recommender systems. Although many LP techniques have been developed in the prior art, most of them consider only static structures of the underlying networks, rarely incorporating the network's information flow. Exploiting the impact of dynamic streams, such as information diffusion, is still an open research topic for LP. Information diffusion allows nodes to receive information beyond their social circles, which, in turn, can influence the creation of new links. In this work, we analyze the LP effects through two diffusion approaches, susceptible-infected-recovered and independent cascade. As a result, we propose the progressive-diffusion (PD) method for LP based on nodes' propagation dynamics. The proposed model leverages a stochastic discrete-time rumor model centered on each node's propagation dynamics. It presents low-memory and low-processing footprints and is amenable to parallel and distributed processing implementation. Finally, we also introduce an evaluation metric for LP methods considering both the information diffusion capacity and the LP accuracy. Experimental results on a series of benchmarks attest to the proposed method's effectiveness compared with the prior art in both criteria. Didier Augusto Vega-Oliveros, Liang Zhao 0001, Anderson Rocha 0001, Lilian Berton |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Open-Set Support Vector MachinesabstractOften, when dealing with real-world recognition problems, we do not need, and often cannot have, knowledge of the entire set of possible classes that might appear during operational testing. In such cases, we need to think of robust classification methods able to deal with the “unknown” and properly reject samples belonging to classes never seen during training. Notwithstanding, existing classifiers to date were mostly developed for the closed-set scenario, i.e., the classification setup in which it is assumed that all test samples belong to one of the classes with which the classifier was trained. In the open-set scenario, however, a test sample can belong to none of the known classes and the classifier must properly reject it by classifying it as unknown. In this work, we extend upon the well-known support vector machines (SVMs) classifier and introduce the open-set SVMs (OSSVMs), which is suitable for recognition in open-set setups. OSSVM balances the empirical risk and the risk of the unknown and ensures that the region of the feature space in which a test sample would be classified as known (one of the known classes) is always bounded, ensuring a finite risk of the unknown. In this work, we also highlight the properties of the SVM classifier related to the open-set scenario, and provide necessary and sufficient conditions for an RBF SVM to have bounded open-space risk. Pedro Ribeiro Mendes Júnior, Terrance E. Boult, Jacques Wainer, Anderson Rocha 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2022 | Convolutional Neural Network Formulation to Compare 4-D Seismic and Reservoir Simulation ModelsabstractWe propose a novel approach based on Deep Learning methods using a convolutional neural networks (CNNs) to compare observed four-dimensional (4-D) seismic data with reservoir simulation model results and select the simulation models with the best match. Our approach advantages are twofold: 1) it is not a pixel-based approach, as it learns spatial features through training examples and 2) it does not require the simulation model and 4-D seismic maps to be in the same domain, enabling a direct comparison of maps in the seismic domain with maps in the numerical simulation domain. Additionally, we present a new methodology to quantitatively assess the performance of each method relative to the expected results. For this purpose, we build an annotated dataset containing the maps that best match the 4-D seismic in different realistic scenarios while taking into account the specialists’ knowledge. We apply our methodology to evaluate if a method is able to perform according to what specialists expect, considering the annotated data. We also propose an improvement to traditional methods using a statistical method to convert maps of different physical properties into a common property, which can be used as an alternative to avoid forward and/or inversion procedures commonly used to bring 4-D seismic and simulation data to the same domain. Finally, we present an extensive comparison, qualitative and quantitative, of our approaches with other methods in the literature, such as the traditional pixelwise strategy for maps in the same domain, and methods based on binarization of maps in different domains. The results show that our proposed methods are capable of more accurately identifying the best-matched models, according to the specialists’ answers. Klaus Rollmann, Aurea Soriano-Vargas, Forlan Almeida, Alessandra Davólio, Denis J. Schiozer, Anderson Rocha 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 6 |
| 2021 | Authentication of Vincent van Gogh's Work
Lucas David 0001, Hélio Pedrini, Zanoni Dias, Anderson Rocha 0001 |
CAIP (2) | 4 |
| 2021 | Semi-Supervised Feature Embedding for Data Sanitization in Real-World EventsabstractWith the rapid growth of data sharing through social media networks, determining relevant data items concerning a particular subject becomes paramount. We address the issue of establishing which images represent an event of interest through a semi-supervised learning technique. The method learns consistent and shared features related to an event (from a small set of examples) to propagate them to an unlabeled set. We investigate the behavior of five image feature representations considering low- and high-level features and their combinations. We evaluate the effectiveness of the feature embedding approach on five collected datasets from real-world events. Bahram Lavi, José Nascimento, Anderson Rocha 0001 |
ICASSP | 3 |
| 2021 | Exploring multiobjective training in multiclass classification
Marcos M. Raimundo, Thalita F. Drumond, Alan Caio R. Marques, Christiano Lyra, Anderson Rocha 0001, Fernando J. Von Zuben |
Neurocomputing | 5 |
| 2021 | Harnessing high-level concepts, visual, and auditory features for violence detection in videos
Bruno Peixoto, Bahram Lavi, Zanoni Dias, Anderson Rocha 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2021 | Temporally sorting images from real-world events
Rafael Padilha, Fernanda A. Andaló, Bahram Lavi, Luís A. M. Pereira, Anderson Rocha 0001 |
Pattern Recognit. Lett. | 5 |
| 2021 | Unsupervised and Self-Adaptative Techniques for Cross-Domain Person Re-IdentificationabstractPerson Re-Identification (ReID) across non-overlapping cameras is a challenging task, and most works in prior art rely on supervised feature learning from a labeled dataset to match the same person in different views. However, it demands the time-consuming task of labeling the acquired data, prohibiting its fast deployment in forensic scenarios. Unsupervised Domain Adaptation (UDA) emerges as a promising alternative, as it performs feature adaptation from a model trained on a source to a target domain without identity-label annotation. However, most UDA-based methods rely upon a complex loss function with several hyper-parameters, hindering the generalization to different scenarios. Moreover, as UDA depends on the translation between domains, it is crucial to select the most reliable data from the unseen domain, avoiding error propagation caused by noisy examples on the target data — an often overlooked problem. In this sense, we propose a novel UDA-based ReID method that optimizes a simple loss function with only one hyper-parameter and takes advantage of triplets of samples created by a new offline strategy based on the diversity of cameras within a cluster. This new strategy adapts and regularizes the model, avoiding overfitting the target domain. We also introduce a new self-ensembling approach, which aggregates weights from different iterations to create a final model, combining knowledge from distinct moments of the adaptation. For evaluation, we consider three well-known deep learning architectures and combine them for the final decision. The proposed method does not use person re-ranking nor any identity label on the target domain and outperforms state-of-the-art techniques, with a much simpler setup, on the Market to Duke, the challenging Market1501 to MSMT17, and Duke to MSMT17 adaptation scenarios. Gabriel Bertocco, Fernanda A. Andaló, Anderson Rocha 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2021 | Transformation-Aware Embeddings for Image ProvenanceabstractA dramatic rise in the flow of manipulated image content on the Internet has led to a prompt response from the media forensics research community. New mitigation efforts leverage cutting-edge data-driven strategies and increasingly incorporate usage of techniques from computer vision and machine learning to detect and profile the space of image manipulations. This paper addresses Image Provenance Analysis, which aims at discovering relationships among different manipulated image versions that share content. One important task in provenance analysis, like most visual understanding problems, is establishing a visual description and dissimilarity computation method that connects images that share full or partial content. But the existing handcrafted or learned descriptors - generally appropriate for tasks such as object recognition - may not sufficiently encode the subtle differences between near-duplicate image variants, which significantly characterize the provenance of any image. This paper introduces a novel data-driven learning-based approach that provides the context for ordering images that have been generated from a single image source through various transformations. Our approach learns transformation-aware embeddings using weak supervision via composited transformations and a rank-based Edit Sequence Loss. To establish the effectiveness of the proposed approach, comparisons are made with state-of-the-art handcrafted and deep-learning-based descriptors, as well as image matching approaches. Further experimentation validates the proposed approach in the context of image provenance analysis and improves upon existing approaches. Aparna Bharati, Daniel Moreira, Patrick J. Flynn, Anderson Rocha 0001, Kevin W. Bowyer, Walter J. Scheirer |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2021 | Manifold Learning for Real-World Event UnderstandingabstractInformation coming from social media is vital to the understanding of the dynamics involved in multiple events such as terrorist attacks and natural disasters. With the spread and popularization of cameras and the means to share content through social networks, an event can be followed through many different lenses and vantage points. However, social media data present numerous challenges, and frequently it is necessary a great deal of data cleaning and filtering techniques to separate what is related to the depicted event from contents otherwise useless. In a previous effort of ours, we decomposed events into representative components aiming at describing vital details of an event to characterize its defining moments. However, the lack of minimal supervision to guide the combination of representative components somehow limited the performance of the method. In this paper, we extend upon our prior work and present a learning-from-data method for dynamically learning the contribution of different components for a more effective event representation. The method relies upon just a few training samples (few-shot learning), which can be easily provided by an investigator. The obtained results on real-world datasets show the effectiveness of the proposed ideas. Caroline Mazini Rodrigues, Aurea Soriano-Vargas, Bahram Lavi, Anderson Rocha 0001, Zanoni Dias |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2021 | Fast Local Spatial Verification for Feature-Agnostic Large-Scale Image RetrievalabstractImages from social media can reflect diverse viewpoints, heated arguments, and expressions of creativity, adding new complexity to retrieval tasks. Researchers working on Content-Based Image Retrieval (CBIR) have traditionally tuned their algorithms to match filtered results with user search intent. However, we are now bombarded with composite images of unknown origin, authenticity, and even meaning. With such uncertainty, users may not have an initial idea of what the search query results should look like. For instance, hidden people, spliced objects, and subtly altered scenes can be difficult for a user to detect initially in a meme image, but may contribute significantly to its composition. It is pertinent to design systems that retrieve images with these nuanced relationships in addition to providing more traditional results, such as duplicates and near-duplicates - and to do so with enough efficiency at large scale. We propose a new approach for spatial verification that aims at modeling object-level regions using image keypoints retrieved from an image index, which is then used to accurately weight small contributing objects within the results, without the need for costly object detection steps. We call this method the Objects in Scene to Objects in Scene (OS2OS) score, and it is optimized for fast matrix operations, which can run quickly on either CPUs or GPUs. It performs comparably to state-of-the-art methods on classic CBIR problems (Oxford 5K, Paris 6K, and Google-Landmarks), and outperforms them in emerging retrieval tasks such as image composite matching in the NIST MFC2018 dataset and meme-style imagery from Reddit. Joel Brogan, Aparna Bharati, Daniel Moreira, Anderson Rocha 0001, Kevin W. Bowyer, Patrick J. Flynn, Walter J. Scheirer |
IEEE Trans. Image Process. | 4 |
| 2020 | Improving the Chronological Sorting of Images through Occlusion: A Study on the Notre-Dame Cathedral FireabstractWe live in a connected society in which no major event - from music concerts to terrorist attempts - happens without being recorded by a smartphone and shared instantly to the world. This flow of generated data is often unstructured and carries no reliable information with respect to time of capture of such media pieces. Consequently, posterior reconstruction, understanding and fact-checking of that event are hindered if the data is not properly organized. In this work, we train a data-driven method to chronologically sort images originated from a real event, the Notre-Dame Cathedral fire, which broke out on April 15th, 2019. Our network leverages visual clues - such as the destruction of the cathedral's structure or the evolution of the fire - to position an image in time. We investigate several occlusion strategies to improve classification accuracy, generalization and explainability of our method. Besides comparing the performance of each strategy, we evaluate their activation maps, i.e., the important regions considered for classification by each method. Rafael Padilha, Fernanda A. Andaló, Anderson Rocha 0001 |
ICASSP | 3 |
| 2020 | Multimodal Violence Detection in VideosabstractEffective tools for detection of violence are highly demanded, specially when dealing with video streams. Such tools have a wide range of applications, from forensics and law enforcement to parental control over the ever increasing amount of videos available online. Prior studies showed that deep learning has great potential in detecting violence, but focuses on detecting violence in general, or only specific cases of violent behavior. While the concept of violence is broad and highly subjective, simpler concepts such as fights, explosions, and gunshots, convey the idea of violence while being more objective. Even though different concepts relate to this same broader idea of violence, they differ widely in relation to whether or not they convey the idea of movement, the presence of a specific object, or even if they generate distinctive sounds. In this study, we propose to analyze different concepts related to violence and how to better describe these concepts exploring visual and auditory cues in order to reach a robust method to detect violence. Bruno Peixoto, Bahram Lavi, Paolo Bestagini, Zanoni Dias, Anderson Rocha 0001 |
ICASSP | 5 |
| 2020 | EMET: Embeddings from Multilingual-Encoder Transformer for Fake News DetectionabstractIn the last few years, social media networks have changed human life experience and behavior as it has broken down communication barriers, allowing ordinary people to actively produce multimedia content on a massive scale. On this wise, the information dissemination in social media platforms becomes increasingly common. However, misinformation is propagated with the same facility and velocity as real news, though it can result in irreversible damage to an individual or society at large. Solving this problem is not a trivial task, considering the reduced size of the text messages usually posted on these communication vehicles. This paper proposes an end-to-end framework called EMET to classify the reliability of small messages posted on social media platforms. Our method leverages text-embeddings from multilingual-encoder transformers that take into consideration the semantic knowledge from preceding trustworthy news and the use of the reader's reactions to detect misleading content. Our findings demonstrated the value of user interaction and prior information to check social media post's credibility. Stephane Schwarz, Antonio Theophilo, Anderson Rocha 0001 |
ICASSP | 3 |
| 2020 | A stochastic learning-from-data approach to the history-matching problem
Cristina C. B. Cavalcante, Célio Maschio, Denis J. Schiozer, Anderson Rocha 0001 |
Eng. Appl. Artif. Intell. | 4 |
| 2020 | Manifold learning for user profiling and identity verification using motion sensors
Geise Santos, Paulo Henrique Pisani, Roberto Leyva, Chang-Tsun Li, Tiago Fernandes Tavares, Anderson Rocha 0001 |
Pattern Recognit. | 6 |
| 2020 | Ground-to-Aerial Viewpoint Localization via Landmark Graphs MatchingabstractThe capability of associating an image to its geographical location is a significant concern in journalism and digital forensics. Given the availability of geo-tagged satellite imagery for most of the Earth's surface, retrieving the location of a generic picture can be addressed as a cross-view image matching between aerial and ground views. In this paper, we outline some initial steps toward the development of a fully-unsupervised algorithm for ground-to-aerial image matching, exploiting the view-invariant adjacency relationships of the landmarks appearing in both views. We introduce a graph-based strategy that, given a set of pre-extracted landmarks, localizes the viewpoint of a ground-level 360-degree image within a broad aerial view of the same area, by matching the respective landmark graphs according to a specifically designed likelihood model. Sebastiano Verde, Thiago Resek, Simone Milani, Anderson Rocha 0001 |
IEEE Signal Process. Lett. | 4 |
| 2020 | Leveraging Shape, Reflectance and Albedo From Shading for Face Presentation Attack DetectionabstractPresentation attack detection is a challenging problem that aims at exposing an impostor user seeking to deceive the authentication system. In facial biometrics systems, this kind of attack is performed using a photograph, video, or 3D mask containing the biometric information of a genuine identity. In this paper, we propose a novel approach to detecting face presentation attacks based on intrinsic properties of the scene such as albedo, depth, and reflectance properties of the facial surfaces, which were recovered through a shape-from-shading (SfS) algorithm. To extract meaningful patterns from the different maps obtained with the SfS algorithm, we designed a novel shallow CNN architecture for learning features useful to the presentation attack detection (PAD). We performed several experiments considering the intra- and inter-dataset evaluation protocols. The obtained results showed the effectiveness of the proposed method considering several types of photo- and video-based presentation attacks, and in the cross-sensor scenario, besides achieving competitive results for the inter-dataset evaluation protocol. Allan Pinto, Siome Goldenstein, Alexandre M. Ferreira, Tiago J. Carvalho, Hélio Pedrini, Anderson Rocha 0001 |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2019 | Open-Set Web Genre Identification Using Distributional Features and Nearest Neighbors Distance Ratio
Dimitrios A. Pritsos, Anderson Rocha 0001, Efstathios Stamatatos |
ECIR (2) | 2 |
| 2019 | Score-based Learning for Relevance Prediction in Image Similarity SearchabstractPredicting the performance of queries when labels are not present has been a recurring problem faced in information retrieval systems. Beyond its clear importance, it can also be applied to aid post-retrieval optimization approaches such as re-ranking or rank-aggregation. However, most post-retrieval performance prediction approaches to retrieval systems rely on generating a single effectiveness value of performance for queries. We propose an alternative method to assess the performance of systems reliant on similarity search, which consists of predicting the individual relevance of ranked results according to the distribution of similarity scores of a given query compared to instances in a collection. The idea is that relationships between the ith ranked score and other scores of the rank can be leveraged to generate features which, in turn, are used to classify ranked objects according to their relevance to the query. We propose a positional classification scheme, in conjunction with simple and fast score-based features to predict the relevance of the top-10 results of a similarity search rank. Our results in nine scenarios, comprising three different large image datasets, show good prediction accuracy for the top-10 results, with the advantage of being amenable suitable to deploy at query time. Alberto A. de Oliveira, Anderson Rocha 0001 |
ICASSP | 2 |
| 2019 | Toward Subjective Violence Detection in VideosabstractViolence detection in videos aims to identify whether a violent action occurred within a video stream. Effective tools for intelligent video analysis are highly demanded, specially to determine violence in video streams. Such solution could have applications in detecting inappropriate behaviors in video feeds, aiding law-enforcement in forensic cases, protecting children from accessing inappropriate online content and helping parents making informed decisions about what their kids should watch. Prior art on violence detection, particularly recently proposed deep learning based ones, seeks to identify violence in videos as a whole, without considering breaking down the subject into some of its underlying concepts. In this paper, we explore a different methodology of violence detection, which relies upon two deep neural network (DNNs) frameworks to learn spatial-temporal information on video clips under different scenarios - subjective- and conceptual-based. We leverage deep feature representations for each specific concept, and aggregate them by training a shallow neural network as a binary-classification problem to describe violence as a whole. Finally, we show that using more specific concepts is an intuitive and effective solution, besides being complementary to form a more robust definition of violence. Bruno Peixoto, Bahram Lavi, João Paulo Pereira Martin, Sandra Eliza Fontes de Avila, Zanoni Dias, Anderson Rocha 0001 |
ICASSP | 6 |
| 2019 | Detection of Real-world Fights in Surveillance VideosabstractCCTVs have since long been used to enforce security, e.g. to detect fights arising from many different situations. But their effectiveness is questionable, because they rely on continuous and specialized human supervision, demanding automated solutions. Previous work are either too superficial (classification of short-clips) or unrealistic (movies, sports, fake fights). None performed detection of actual fights on long duration CCTV recordings. In this work, we tackle this problem by firstly proposing CCTV-Fights1, a novel and challenging dataset containing 1,000 videos of real fights, with more than 8 hours of annotated CCTV footage. Then we propose a pipeline, on which we assess the impact of different feature extractors, through Two-stream CNN, 3D CNN and a local interest point descriptor, as well as different classifiers, such as end-to-end CNN, LSTM and SVM. Results confirm how challenging the problem is, and highlight the importance of explicit motion information to improve performance. Mauricio Perez, Alex Chichung Kot, Anderson Rocha 0001 |
ICASSP | 3 |
| 2019 | A Needle in a Haystack? Harnessing Onomatopoeia and User-specific Stylometrics for Authorship Attribution of Micro-messagesabstractThe world is facing a new era in which social media communication plays a fundamental role in people's lives. Along with irrefutable benefits, several collateral drawbacks have risen, one being the wide spread of false information with malicious intents, what is now commonly called "Fake News". The fight against this problem is not easy, especially when taking into account the nature of text messages involved on social media platforms (a sea of small messages and myriad users). In this work, we cope with the challenging problem of authorship attribution of small text messages posted on social media platforms. Differently from what has been done with longer texts, we rely upon data-driven approaches, exploiting recent advances of deep neural networks in the field of pattern recognition. By viewing small texts usually employed in social media as unidimensional signals, we devise modern deep-learning techniques tailored for this kind of data to find the author of these posts with promising results. Antonio Theophilo, Luís A. M. Pereira, Anderson Rocha 0001 |
ICASSP | 3 |
| 2019 | Deep Face Verification for Spherical ImagesabstractOver the years, several problems regarding the analysis of face images have been addressed, including face detection, recognition, identification, and verification. The advent of Convolutional Neural Networks (CNNs) gave rise to a drastic improvement on state-of-the-art performances for these problems. With the increasing popularity of 360ocameras, the demand for models to extract relevant information from spherical images has also emerged. However, traditional CNNs, originally designed for planar images, are typically not suitable for spherical images, as it is necessary to project these spherical images onto a plane, leading to severe distortions. This work presents a method for face verification on spherical images that relies upon CNNs to extract features for training a binary classifier, as well as two new face datasets with spherical images. The effectiveness of our method is assessed through a comparative analysis with relevant planar and spherical CNNs. Marcos V. M. Cirne, Fernanda A. Andaló, Rafael Dias, Thiago Resek, Gabriel Bertocco, Ricardo da Silva Torres, Anderson Rocha 0001 |
ICIP | 7 |
| 2019 | A Prnu-Based Method to Expose Video Device Compositions in Open-Set SetupsabstractAs the diffusion of altered video sequences over social networks and the internet can have severe consequences (e.g., fake news spreading, false accusations, etc.), the forensic research community has started actively working toward the development of methodologies tailored to assess the authenticity and integrity of video sequences. In this work, we focus on the problem of spotting video compilations composed by temporally concatenating sequences acquired with different devices. To solve this problem, we leverage trace characteristics of each video recording device left on each video at acquisition time. Specifically, we propose a method leveraging Photo-Response Non-Uniformity (PRNU) traces along with a binary classifier in order to understand which frames of a video have been acquired with the same device used to record the first few frames. Results show that the proposed solution outperforms baselines based on standard PRNU correlation and thresholding tests. Experiments have been carried out in an open-set scenario and show promising results. Pedro Ribeiro Mendes Júnior, Luca Bondi, Paolo Bestagini, Anderson Rocha 0001, Stefano Tubaro |
ICIP | 4 |
| 2019 | Detection and Synchronization of Video Sequences for Event ReconstructionabstractWith an ever-growing amount of unexpected menaces in crowded places such as terrorist attacks, it is paramount to develop techniques to aid investigators reconstructing all details about an event of interest. To extract reliable information about the event, all kinds of available clues must be jointly exploited. As a matter of fact, today's sources of information are plenty and varied, as important events affecting many people are typically documented by different sources. Both witnesses' smartphones and security cameras can provide valuable information coming from multiple viewpoints and time instants - "the eyes of the crowd". In this paper, we focus on the specific problem of automatically detecting and temporally synchronizing videos depicting the same event of interest. Videos can be either near-duplicates (i.e., edited copies of the same original source) or sequences shot by different users from different vantage points. The proposed method relies upon a video fingerprinting technique capable of describing how video semantic content evolves in time. The solution does not assume a priori information about cameras location, and it only exploits visual cues, not relying on audio channels. Giuliano Pinheiro, Marcos V. M. Cirne, Paolo Bestagini, Stefano Tubaro, Anderson Rocha 0001 |
ICIP | 5 |
| 2019 | Beyond Pixels: Image Provenance Analysis Leveraging MetadataabstractCreative works, whether paintings or memes, follow unique journeys that result in their final form. Understanding these journeys, a process known as "provenance analysis," provides rich insights into the use, motivation, and authenticity underlying any given work. The application of this type of study to the expanse of unregulated content on the Internet is what we consider in this paper. Provenance analysis provides a snapshot of the chronology and validity of content as it is uploaded, re-uploaded, and modified over time. Although still in its infancy, automated provenance analysis for online multimedia is already being applied to different types of content. Most current works seek to build provenance graphs based on the shared content between images or videos. This can be a computationally expensive task, especially when considering the vast influx of content that the Internet sees every day. Utilizing non-content-based information, such as timestamps, geotags, and camera IDs can help provide important insights into the path a particular image or video has traveled during its time on the Internet without large computational overhead. This paper tests the scope and applicability of metadata-based inferences for provenance graph construction in two different scenarios: digital image forensics and cultural analytics. Aparna Bharati, Daniel Moreira, Joel Brogan, Patricia Hale, Kevin W. Bowyer, Patrick J. Flynn, Anderson Rocha 0001, Walter J. Scheirer |
WACV | 7 |
| 2019 | A data-driven approach to referable diabetic retinopathy detection
Ramon Pires, Sandra Eliza Fontes de Avila, Jacques Wainer, Eduardo Valle, Michael D. Abràmoff, Anderson Rocha 0001 |
Artif. Intell. Medicine | 6 |
| 2019 | A continuous learning algorithm for history matching
Cristina C. B. Cavalcante, Célio Maschio, Antonio Alberto S. Santos, Denis J. Schiozer, Anderson Rocha 0001 |
Eng. Appl. Artif. Intell. | 5 |
| 2019 | Relevance prediction in similarity-search systems using extreme value theory
Alberto A. de Oliveira, Eric Oakley, Ricardo da Silva Torres, Anderson Rocha 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2019 | Ensemble of Multi-View Learning Classifiers for Cross-Domain Iris Presentation Attack DetectionabstractThe adoption of large-scale iris recognition systems around the world has brought to light the importance of detecting presentation attack images (textured contact lenses and printouts). This paper presents a new approach in iris presentation attack detection (PAD) by exploring combinations of convolutional neural networks (CNNs) and transformed input spaces through binarized statistical image features (BSIFs). Our method combines lightweight CNNs to classify multiple BSIF views of the input image. Following explorations on complementary input spaces leading to more discriminative features to detect presentation attacks, we also propose an algorithm to select the best (and most discriminative) predictors for the task at hand. An ensemble of predictors makes use of their expected individual performances to aggregate their results into a final prediction. Results show that this technique improves on the current state of the art in iris PAD, outperforming the winner of LivDet-Iris 2017 competition both for intra- and cross-dataset scenarios, and illustrating the very difficult nature of the cross-dataset scenario. Andrey Kuehlkamp, Allan Pinto, Anderson Rocha 0001, Kevin W. Bowyer, Adam Czajka |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2018 | Breaking down violence: A deep-learning strategy to model and classify violence in videosabstractDetecting violence in videos through automatic means is significant for law enforcement and analysis of surveillance cameras with the intent of maintaining public safety. Moreover, it may be a great tool for protecting children from accessing inappropriate content and help parents make a better informed decision about what their kids should watch. However, this is a challenging problem since the very definition of violence is broad and highly subjective. Hence, detecting such nuances from videos with no human supervision is not only technical, but also a conceptual problem. With this in mind, we explore how to better describe the idea of violence for a convolutional neural network by breaking it into more objective and concrete parts. Initially, our method uses independent networks to learn features for more specific concepts related to violence, such as fights, explosions, blood, etc. Then we use these features to classify each concept and later fuse them in a meta-classification to describe violence. We also explore how to represent time-based events in still-images as network inputs; since many violent acts are described in terms of movement. We show that using more specific concepts is an intuitive and effective solution, besides being complementary to form a more robust definition of violence. When compared to other methods for violence detection, this approach holds better classification quality while using only automatic features. Bruno Peixoto, Sandra Eliza Fontes de Avila, Zanoni Dias, Anderson Rocha 0001 |
ARES | 4 |
| 2018 | Image Splicing Detection Through Illumination Inconsistencies and Deep LearningabstractFake news and deep fakes have been making social and mainstream media headlines. At the same time, engaged scientists strive for finding ways to detect forgeries and suspicious manipulations using even the subtlest clues. In this vein, this work proposes a new method for detecting photographic splicing by bringing together the high representation power of Illuminant Maps and Convolutional Neural Networks as a way of learning directly from available training data the most important hints of a forgery. This work propose a method that eliminates the laborious feature engineering process, allow locate forgery region and yields a classification accuracy of more than 96%, outperforming state-of-the-art methods in different datasets. The potential uses of the proposed method is further highlighted by analyzing some suspicious real-world photographs that recently broke the news. Thales Pomari, Guilherme C. S. Ruppert, Edmar R. S. De Rezende, Anderson Rocha 0001, Tiago J. Carvalho |
ICIP | 4 |
| 2018 | Leveraging ontologies and machine-learning techniques for malware analysis into Android permissions ecosystems
Luiz C. Navarro, Alexandre K. W. Navarro, André Ricardo Abed Grégio, Anderson Rocha 0001, Ricardo Dahab |
Comput. Secur. | 4 |
| 2018 | Kuaa: A unified framework for design, deployment, execution, and recommendation of machine learning experiments
Rafael de Oliveira Werneck, Waldir R. de Almeida, Bernardo V. Stein, Daniel V. Pazinato, Pedro Ribeiro Mendes Júnior, Otávio A. B. Penatti, Anderson Rocha 0001, Ricardo da Silva Torres |
Future Gener. Comput. Syst. | 7 |
| 2018 | Connecting the dots: Toward accountable machine-learning printer attribution methods
Luiz C. Navarro, Alexandre K. W. Navarro, Anderson Rocha 0001, Ricardo Dahab |
J. Vis. Commun. Image Represent. | 3 |
| 2018 | Data-driven multimedia forensics and security
Anderson Rocha 0001, Shujun Li 0001, C.-C. Jay Kuo, Alessandro Piva, Jiwu Huang |
J. Vis. Commun. Image Represent. | 1 |
| 2018 | Leveraging deep neural networks to fight child pornography in the age of social media
Paulo Vitorino, Sandra Eliza Fontes de Avila, Mauricio Perez, Anderson Rocha 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2018 | Learning Generalized Deep Feature Representation for Face Anti-SpoofingabstractIn this paper, we propose a novel framework leveraging the advantages of the representational ability of deep learning and domain generalization for face spoofing detection. In particular, the generalized deep feature representation is achieved by taking both spatial and temporal information into consideration, and a 3D convolutional neural network architecture tailored for the spatial-temporal input is proposed. The network is first initialized by training with augmented facial samples based on cross-entropy loss and further enhanced with a specifically designed generalization loss, which coherently serves as the regularization term. The training samples from different domains can seamlessly work together for learning the generalized feature representation by manipulating their feature distribution distances. We evaluate the proposed framework with different experimental setups using various databases. Experimental results indicate that our method can learn more discriminative and generalized information compared with the state-of-the-art methods. Haoliang Li, Peisong He, Shiqi Wang 0001, Anderson Rocha 0001, Xinghao Jiang, Alex Chichung Kot |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2018 | Image Provenance Analysis at ScaleabstractPrior art has shown it is possible to estimate, through image processing and computer vision techniques, the types and parameters of transformations that have been applied to the content of individual images to obtain new images. Given a large corpus of images and a query image, an interesting further step is to retrieve the set of original images whose content is present in the query image, as well as the detailed sequences of transformations that yield the query image given the original images. This is a problem that recently has received the name of image provenance analysis. In these times of public media manipulation (e.g., fake news and meme sharing), obtaining the history of image transformations is relevant for fact checking and authorship verification, among many other applications. This article presents an end-to-end processing pipeline for image provenance analysis, which works at real-world scale. It employs a cutting-edge image filtering solution that is custom-tailored for the problem at hand, as well as novel techniques for obtaining the provenance graph that expresses how the images, as nodes, are ancestrally connected. A comprehensive set of experiments for each stage of the pipeline is provided, comparing the proposed solution with state-of-the-art results, employing previously published datasets. In addition, this work introduces a new dataset of real-world provenance cases from the social media site Reddit, along with baseline results. Aparna Bharati, Joel Brogan, Allan Pinto, Michael Parowski, Kevin W. Bowyer, Patrick J. Flynn, Anderson Rocha 0001, Walter J. Scheirer |
IEEE Trans. Image Process. | 7 |
| 2017 | A competition on generalized software-based face presentation attack detection in mobile scenariosabstractIn recent years, software-based face presentation attack detection (PAD) methods have seen a great progress. However, most existing schemes are not able to generalize well in more realistic conditions. The objective of this competition is to evaluate and compare the generalization performances of mobile face PAD techniques under some real-world variations, including unseen input sensors, presentation attack instruments (PAI) and illumination conditions, on a larger scale OULU-NPU dataset using its standard evaluation protocols and metrics. Thirteen teams from academic and industrial institutions across the world participated in this competition. This time typical liveness detection based on physiological signs of life was totally discarded. Instead, every submitted system relies practically on some sort of feature representation extracted from the face and/or background regions using hand-crafted, learned or hybrid descriptors. Interesting results and findings are presented and discussed in this paper. Zinelabidine Boulkenafet, Jukka Komulainen, Zahid Akhtar, Azeddine Benlamoudi, Djamel Samai, Salah Eddine Bekhouche, Abdelkrim Ouafi, Fadi Dornaika, Abdelmalik Taleb-Ahmed, Fei Peng 0001, L. B. Zhang, Min Long 0003, Shruti Bhilare, Vivek Kanhangad, Artur Costa-Pazo, Esteban Vázquez-Fernández, Daniel Pérez-Cabo, J. J. Moreira-Perez, Daniel González-Jiménez, Amir Mohammadi, Sushil Bhattacharjee, Sébastien Marcel, Svetlana Volkova, N. Abe, X. Feng, Z. Xia, Rui Shao 0001, Pong C. Yuen, Waldir R. de Almeida, Fernanda A. Andaló, Rafael Padilha, Gabriel Bertocco, William Dias, Jacques Wainer, Ricardo da Silva Torres, Anderson Rocha 0001, Marcus A. Angeloni, Guilherme Folego, Alan Godoy, Abdenour Hadid |
IJCB | 41 |
| 2017 | U-Phylogeny: Undirected provenance graph construction in the wildabstractDeriving relationships between images and tracing back their history of modifications are at the core of Multimedia Phylogeny solutions, which aim to combat misinformation through doctored visual media. Nonetheless, most recent image phylogeny solutions cannot properly address cases of forged composite images with multiple donors, an area known as multiple parenting phylogeny (MPP). This paper presents a preliminary undirected graph construction solution for MPP, without any strict assumptions. The algorithm is underpinned by robust image representative keypoints and different geometric consistency checks among matching regions in both images to provide regions of interest for direct comparison. The paper introduces a novel technique to geometrically filter the most promising matches as well as to aid in the shared region localization task. The strength of the approach is corroborated by experiments with real-world cases, with and without image distractors (unrelated cases). Aparna Bharati, Daniel Moreira, Allan Pinto, Joel Brogan, Kevin W. Bowyer, Patrick J. Flynn, Walter J. Scheirer, Anderson Rocha 0001 |
ICIP | 8 |
| 2017 | Spotting the difference: Context retrieval and analysis for improved forgery detection and localizationabstractAs image tampering becomes ever more sophisticated and commonplace, the need for image forensics algorithms that can accurately and quickly detect forgeries grows. In this paper, we revisit the ideas of image querying and retrieval to provide clues to better localize forgeries. We propose a method to perform large-scale image forensics on the order of one million images using the help of an image search algorithm and database to gather contextual clues as to where tampering may have taken place. In this vein, we introduce five new strongly invariant image comparison methods and test their effectiveness under heavy noise, rotation, and color space changes. Lastly, we show the effectiveness of these methods compared to passive image forensics using Nimble [1], a new, state-of-the-art dataset from the National Institute of Standards and Technology (NIST). Joel Brogan, Paolo Bestagini, Aparna Bharati, Allan Pinto, Daniel Moreira, Kevin W. Bowyer, Patrick J. Flynn, Anderson Rocha 0001, Walter J. Scheirer |
ICIP | 8 |
| 2017 | Provenance filtering for multimedia phylogenyabstractDeparting from traditional digital forensics modeling, which seeks to analyze single objects in isolation, multimedia phylogeny analyzes the evolutionary processes that influence digital objects and collections over time. One of its integral pieces is provenance filtering, which consists of searching a potentially large pool of objects for the most related ones with respect to a given query, in terms of possible ancestors (donors or contributors) and descendants. In this paper, we propose a two-tiered provenance filtering approach to find all the potential images that might have contributed to the creation process of a given query q. In our solution, the first (coarse) tier aims to find the most likely “host” images - the major donor or background - contributing to a composite/doctored image. The search is then refined in the second tier, in which we search for more specific (potentially small) parts of the query that might have been extracted from other images and spliced into the query image. Experimental results with a dataset containing more than a million images show that the two-tiered solution underpinned by the context of the query is highly useful for solving this difficult task. Allan Pinto, Daniel Moreira, Aparna Bharati, Joel Brogan, Kevin W. Bowyer, Patrick J. Flynn, Walter J. Scheirer, Anderson Rocha 0001 |
ICIP | 8 |
| 2017 | Temporal Robust Features for Violence DetectionabstractAutomatically detecting violence in videos is paramount for enforcing the law and providing the society with better policies for safer public places. In addition, it may be essential for protecting minors from accessing inappropriate contents on-line, and for helping parents choose suitable movie titles for their children. However, this is an open problem as the very definition of violence is subjective and may vary from one society to another. Detecting such nuances from video footages with no human supervision is very challenging. Clearly, when designing a computer-aided solution to this problem, we need to think of efficient (quickly harness large troves of data) and effective detection methods (robustly filter what needs special attention and further analysis). In this vein, we explore a content description method for violence detection founded upon temporal robust features that quickly grasp video sequences, automatically classifying violent videos. The used method also holds promise for fast and effective classification of other recognition tasks (e.g., pornography and other inappropriate material). When compared to more complex counterparts for violence detection, the method shows similar classification quality while being several times more efficient in terms of runtime and memory footprint. Daniel Moreira, Sandra Eliza Fontes de Avila, Mauricio Perez, Daniel Moraes, Vanessa Testoni, Eduardo Valle, Siome Goldenstein, Anderson Rocha 0001 |
WACV | 8 |
| 2017 | Video pornography detection through deep learning techniques and motion information
Mauricio Perez, Sandra Eliza Fontes de Avila, Daniel Moreira, Daniel Moraes, Vanessa Testoni, Eduardo Valle, Siome Goldenstein, Anderson Rocha 0001 |
Neurocomputing | 8 |
| 2017 | Nearest neighbors distance ratio open-set classifier
Pedro Ribeiro Mendes Júnior, Roberto Souza 0001, Rafael de Oliveira Werneck, Bernardo V. Stein, Daniel V. Pazinato, Waldir R. de Almeida, Otávio A. B. Penatti, Ricardo da Silva Torres, Anderson Rocha 0001 |
Mach. Learn. | 9 |
| 2017 | New dissimilarity measures for image phylogeny reconstruction
Filipe de Oliveira Costa, Alberto A. de Oliveira, Pasquale Ferrara, Zanoni Dias, Siome Goldenstein, Anderson Rocha 0001 |
Pattern Anal. Appl. | 6 |
| 2017 | A Kinect-Based Wearable Face Recognition System to Aid Visually Impaired UsersabstractIn this paper, we introduce a real-time face recognition (and announcement) system targeted at aiding the blind and low-vision people. The system uses a Microsoft Kinect sensor as a wearable device, performs face detection, and uses temporal coherence along with a simple biometric procedure to generate a sound associated with the identified person, virtualized at his/her estimated 3-D location. Our approach uses a variation of the K-nearest neighbors algorithm over histogram of oriented gradient descriptors dimensionally reduced by principal component analysis. The results show that our approach, on average, outperforms traditional face recognition methods while requiring much less computational resources (memory, processing power, and battery life) when compared with existing techniques in the literature, deeming it suitable for the wearable hardware constraints. We also show the performance of the system in the dark, using depth-only information acquired with Kinect's infrared camera. The validation uses a new dataset available for download, with 600 videos of 30 people, containing variation of illumination, background, and movement patterns. Experiments with existing datasets in the literature are also considered. Finally, we conducted user experience evaluations on both blindfolded and visually impaired users, showing encouraging results. Laurindo de Sousa Britto Neto, Felipe Grijalva, Vanessa Regina Margareth Lima Maike, Luiz Martini, Dinei A. F. Florêncio, Maria Cecília Calani Baranauskas, Anderson Rocha 0001, Siome Goldenstein |
IEEE Trans. Hum. Mach. Syst. | 7 |
| 2017 | Data-Driven Feature Characterization Techniques for Laser Printer AttributionabstractLaser printer attribution is an increasing problem with several applications, such as pointing out the ownership of crime proofs and authentication of printed documents. However, as commonly proposed methods for this task are based on custom-tailored features, they are limited by modeling assumptions about printing artifacts. In this paper, we explore solutions able to learn discriminant-printing patterns directly from the available data during an investigation, without any further feature engineering, proposing the first approach based on deep learning to laser printer attribution. This allows us to avoid any prior assumption about printing artifacts that characterize each printer, thus highlighting almost invisible and difficult printer footprints generated during the printing process. The proposed approach merges, in a synergistic fashion, convolutional neural networks (CNNs) applied on multiple representations of multiple data. Multiple representations, generated through different pre-processing operations, enable the use of the small and lightweight CNNs whilst the use of multiple data enable the use of aggregation procedures to better determine the provenance of a document. Experimental results show that the proposed method is robust to noisy data and outperforms existing counterparts in the literature for this problem. Anselmo Ferreira, Luca Bondi, Luca Baroffio, Paolo Bestagini, Jiwu Huang, Jefersson A. dos Santos, Stefano Tubaro, Anderson Rocha 0001 |
IEEE Trans. Inf. Forensics Secur. | 8 |
| 2017 | Authorship Attribution for Social Media ForensicsabstractThe veil of anonymity provided by smartphones with pre-paid SIM cards, public Wi-Fi hotspots, and distributed networks like Tor has drastically complicated the task of identifying users of social media during forensic investigations. In some cases, the text of a single posted message will be the only clue to an author's identity. How can we accurately predict who that author might be when the message may never exceed 140 characters on a service like Twitter? For the past 50 years, linguists, computer scientists, and scholars of the humanities have been jointly developing automated methods to identify authors based on the style of their writing. All authors possess peculiarities of habit that influence the form and content of their written works. These characteristics can often be quantified and measured using machine learning algorithms. In this paper, we provide a comprehensive review of the methods of authorship attribution that can be applied to the problem of social media forensics. Furthermore, we examine emerging supervised learning-based methods that are effective for small sample sizes, and provide step-by-step explanations for several scalable approaches as instructional case studies for newcomers to the field. We argue that there is a significant need in forensics for new authorship attribution algorithms that can exploit context, can process multi-modal data, and are tolerant to incomplete knowledge of the space of all possible authors at training time. Anderson Rocha 0001, Walter J. Scheirer, Christopher W. Forstall, Thiago Cavalcante, Antonio Theophilo, Bingyu Shen 0001, Ariadne Carvalho, Efstathios Stamatatos |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2017 | Beyond Lesion-Based Diabetic Retinopathy: A Direct Approach for ReferralabstractDiabetic retinopathy (DR) is the leading cause of blindness in adults, but can be managed if detected early. Automated DR screening helps by indicating which patients should be referred to the doctor. However, current techniques of automated screening still depend too much on the detection of individual lesions. In this study, we bypass lesion detection, and directly train a classifier for DR referral. Additional novelties are the use of state-of-the-art mid-level features for the retinal images: BossaNova and Fisher Vector. Those features extend the classical Bags of Visual Words and greatly improve the accuracy of complex classification tasks. The proposed technique for direct referral is promising, achieving an area under the curve of 96.4%, thus, reducing the classification error by almost 40% over the current state of the art, held by lesion-based techniques. Ramon Pires, Sandra Eliza Fontes de Avila, Herbert F. Jelinek, Jacques Wainer, Eduardo Valle, Anderson Rocha 0001 |
IEEE J. Biomed. Health Informatics | 6 |
| 2016 | From impressionism to expressionism: Automatically identifying van Gogh's paintingsabstractCurators, art historians, and connoisseurs are often interested in determining the authorship of paintings. Machine learning and image processing techniques can assist in this task by providing non-invasive, automatic, and objective methods. In this work, we study the automatic identification of Vincent van Gogh's paintings using a Convolutional Neural Network that extracts discriminative visual patterns of a painter directly from images, and a machine learning classifier allied with a fusion method in the final decision process. We divide each painting into non-overlapping patches, classify them individually, and then aggregate the outcomes for the final response. We find out that using the patch with highest confidence score leads to the best result, outperforming the traditional voting scheme. We also contribute with a new and public dataset for van Gogh painting identification. Guilherme Folego, Otavio Gomes, Anderson Rocha 0001 |
ICIP | 3 |
| 2016 | Multi-directional and multi-scale perturbation approaches for blind forensic median filtering detectionabstractThe forensic detection of median filtering has recently attracted the attention of the research community, mainly because of the median filtering potential uses for tampering and concealing image tampering traces in digital images. In this paper, we propose multi-scale and multi-perturbation soluti ons that build a highly discriminative feature space, which highlights the artifacts of median filtering by means of image quality measures. The proposed methods achieve promising results when validated with a series of real-world test cases, comprising different image compression levels, resolutions, and also a cross-dataset validation protocol. Anselmo Ferreira, Jefersson A. dos Santos, Anderson Rocha 0001 |
Intell. Data Anal. | 3 |
| 2016 | Rate-energy-accuracy optimization of convolutional architectures for face recognition
Luca Bondi, Luca Baroffio, Matteo Cesana, Marco Tagliasacchi, Giovani Chiachia, Anderson Rocha 0001 |
J. Vis. Commun. Image Represent. | 6 |
| 2016 | Low false positive learning with support vector machines
Daniel Moraes, Jacques Wainer, Anderson Rocha 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2016 | Time series-based classifier fusion for fine-grained plant species recognition
Fábio Augusto Faria, Jurandy Almeida, Bruna Alberton, Leonor Patricia C. Morellato, Anderson Rocha 0001, Ricardo da Silva Torres |
Pattern Recognit. Lett. | 5 |
| 2016 | Illuminant-Based Transformed Spaces for Image ForensicsabstractIn this paper, we explore transformed spaces, represented by image illuminant maps, to propose a methodology for selecting complementary forms of characterizing visual properties for an effective and automated detection of image forgeries. We combine statistical telltales provided by different image descriptors that explore color, shape, and texture features. We focus on detecting image forgeries containing people and present a method for locating the forgery, specifically, the face of a person in an image. Experiments performed on three different open-access data sets show the potential of the proposed method for pinpointing image forgeries containing people. In the two first data sets (DSO-1 and DSI-1), the proposed method achieved a classification accuracy of 94% and 84%, respectively, a remarkable improvement when compared with the state-of-the-art methods. Finally, when evaluating the third data set comprising questioned images downloaded from the Internet, we also present a detailed analysis of target images. Tiago Jose de Carvalho, Fábio Augusto Faria, Hélio Pedrini, Ricardo da Silva Torres, Anderson Rocha 0001 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2016 | Manifold Learning and Spectral Clustering for Image Phylogeny ForestsabstractThe ever-increasing number of gadgets being used to create digital content, as well as the easiness in sharing, editing, and republishing this content, brings the problem of dealing with a large amount of digital objects (e.g., images or videos) whose content is very similar. Some issues faced by investigators of digital crimes when analyzing this type of data include finding the original source of a suspect image, and the responsible for first publishing it. It is also challenging to determine how these objects are related to each other. Recent efforts in developing algorithms to find automatically the underlying relationship among groups of digital media objects with similar content have been explored in the multimedia phylogeny field. A tree structure is used to represent the relationship among these objects, inspired by the phylogenetic trees in biology. Discovering whether these objects came from the same source or from different sources is fundamentally a clustering problem: 1) related objects belong to the same cluster (tree) and 2) unrelated objects should fit in different clusters. In this paper, we address the problem of finding these clusters in sets of semantically similar images, prior to tree reconstruction. We propose the combination of manifold learning and spectral clustering approaches, which have been successfully used in different applications embedding the original data into a lower, but meaningful, dimensional space. Experiments with more than 40 000 test cases show that the proposed approach improves the accuracy in finding the correct number of trees in the set, as well as the reconstruction of the phylogeny trees. Marina A. Oikawa, Zanoni Dias, Anderson Rocha 0001, Siome Goldenstein |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2016 | Multiple Parenting Phylogeny Relationships in Digital ImagesabstractRecently, several studies have been concerned with modeling the parenthood relationships between near duplicates in a set of images. Two images share a parenthood relationship if one is obtained by applying transformations to the other. However, this is not the only form of parenting that can exist among images. An image might be a composition created through the combination of the semantic information existent in two or more source images, establishing a relationship between the sources and the composite. The problem of identifying these relations in a set containing near-duplicate subsets of source and composition images is referred to as multiple parenting phylogeny. Thus far, researchers tackled this problem with a three-step solution: 1) separation of near-duplicate groups; 2) classification of the relations between the groups; and 3) identification of the images used to create the original composition. In this work, we extend upon this framework by introducing key improvements, such as better identification of when two images share content, and improved ways to compare this content. In addition, we also introduce a new realistic professionally created data set of compositions involving multiple parenting relationships. The method we present in this paper is properly evaluated through quantitative metrics, established for assessing the accuracy in finding multiple parenting relationships. Finally, we discuss some particularities of the framework, such as the importance of an accurate reconstruction of phylogenies and the method's behavior when dealing with more complex compositions. Alberto A. de Oliveira, Pasquale Ferrara, Alessia De Rosa, Alessandro Piva, Mauro Barni, Siome Goldenstein, Zanoni Dias, Anderson Rocha 0001 |
IEEE Trans. Inf. Forensics Secur. | 8 |
| 2016 | Behavior Knowledge Space-Based Fusion for Copy-Move Forgery DetectionabstractThe detection of copy-move image tampering is of paramount importance nowadays, mainly due to its potential use for misleading the opinion forming process of the general public. In this paper, we go beyond traditional forgery detectors and aim at combining different properties of copy-move detection approaches by modeling the problem on a multiscale behavior knowledge space, which encodes the output combinations of different techniques as a priori probabilities considering multiple scales of the training data. Afterward, the conditional probabilities missing entries are properly estimated through generative models applied on the existing training data. Finally, we propose different techniques that exploit the multi-directionality of the data to generate the final outcome detection map in a machine learning decision-making fashion. Experimental results on complex data sets, comparing the proposed techniques with a gamut of copy-move detection approaches and other fusion methodologies in the literature, show the effectiveness of the proposed method and its suitability for real-world applications. Anselmo Ferreira, Siovani Cintra Felipussi, Carlos Alfaro 0002, Pablo Fonseca, John E. Vargas-Munoz, Jefersson A. dos Santos, Anderson Rocha 0001 |
IEEE Trans. Image Process. | 7 |
| 2016 | Pixel-Level Tissue Classification for Ultrasound ImagesabstractBACKGROUND: Pixel-level tissue classification for ultrasound images, commonly applied to carotid images, is usually based on defining thresholds for the isolated pixel values. Ranges of pixel values are defined for the classification of each tissue. The classification of pixels is then used to determine the carotid plaque composition and, consequently, to determine the risk of diseases (e.g., strokes) and whether or not a surgery is necessary. The use of threshold-based methods dates from the early 2000s but it is still widely used for virtual histology. METHODOLOGY/PRINCIPAL FINDINGS: We propose the use of descriptors that take into account information about a neighborhood of a pixel when classifying it. We evaluated experimentally different descriptors (statistical moments, texture-based, gradient-based, local binary patterns, etc.) on a dataset of five types of tissues: blood, lipids, muscle, fibrous, and calcium. The pipeline of the proposed classification method is based on image normalization, multiscale feature extraction, including the proposal of a new descriptor, and machine learning classification. We have also analyzed the correlation between the proposed pixel classification method in the ultrasound images and the real histology with the aid of medical specialists. CONCLUSIONS/SIGNIFICANCE: The classification accuracy obtained by the proposed method with the novel descriptor in the ultrasound tissue images (around 73%) is significantly above the accuracy of the state-of-the-art threshold-based methods (around 54%). The results are validated by statistical tests. The correlation between the virtual and real histology confirms the quality of the proposed approach showing it is a robust ally for the virtual histology in ultrasound images. Daniel V. Pazinato, Bernardo V. Stein, Waldir R. de Almeida, Rafael de Oliveira Werneck, Pedro Ribeiro Mendes Júnior, Otávio A. B. Penatti, Ricardo da Silva Torres, Fabio H. Menezes, Anderson Rocha 0001 |
IEEE J. Biomed. Health Informatics | 9 |
| 2015 | Phylogeny reconstruction for misaligned and compressed video sequencesabstractIn the last few years, the amount of videos distributed online has dramatically increased due to the popularity of media sharing platforms (e.g., YouTube, Vimeo, etc.). However, distributed videos are often edited copies of original content, typically referred to as near duplicates. In this paper, we face the problem of reconstructing a video phylogeny tree, i.e., given a set of near-duplicate videos, we want to reconstruct the relationships between every pair of videos to detect which one generated the others and trace back their evolution history. Solving this problem is of paramount importance when the first published video within a set is sought, e.g., to solve copyright infringement cases or to pinpoint criminal impersonation online. The technique we propose exploits the same rationale of previous works in the field of image and video phylogeny. However, we embed in the commonly used pipeline of operations the possibility of dealing with temporally misaligned and encoded video sequences, thus making the method applicable to user-generated videos shared on online platforms. Results computed on a wide dataset of video sequences highlight the importance of taking care of both coding and misalignment in the reconstruction pipeline. Filipe de Oliveira Costa, Silvia Lameri, Paolo Bestagini, Zanoni Dias, Anderson Rocha 0001, Marco Tagliasacchi, Stefano Tubaro |
ICIP | 5 |
| 2015 | Going deeper into copy-move forgery detection: Exploring image telltales via multi-scale analysis and voting processes
Ewerton Silva, Tiago Jose de Carvalho, Anselmo Ferreira, Anderson Rocha 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2015 | Deep Representations for Iris, Face, and Fingerprint Spoofing DetectionabstractBiometrics systems have significantly improved person identification and authentication, playing an important role in personal, national, and global security. However, these systems might be deceived (or spoofed) and, despite the recent advances in spoofing detection, current solutions often rely on domain knowledge, specific biometric reading systems, and attack types. We assume a very limited knowledge about biometric spoofing at the sensor to derive outstanding spoofing detection systems for iris, face, and fingerprint modalities based on two deep learning approaches. The first approach consists of learning suitable convolutional network architectures for each domain, whereas the second approach focuses on learning the weights of the network via back propagation. We consider nine biometric spoofing benchmarks - each one containing real and fake samples of a given biometric modality and attack type - and learn deep representations for each benchmark by combining and contrasting the two learning approaches. This strategy not only provides better comprehension of how these approaches interplay, but also creates systems that exceed the best known results in eight out of the nine benchmarks. The results strongly indicate that spoofing detection systems based on convolutional networks can be robust to attacks already known and possibly adapted, with little effort, to image-based attacks that are yet to come. David Menotti, Giovani Chiachia, Allan Pinto, William Robson Schwartz, Hélio Pedrini, Alexandre X. Falcão, Anderson Rocha 0001 |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2015 | Using Visual Rhythms for Detecting Video-Based Facial Spoof AttacksabstractSpoofing attacks or impersonation can be easily accomplished in a facial biometric system wherein users without access privileges attempt to authenticate themselves as valid users, in which an impostor needs only a photograph or a video with facial information of a legitimate user. Even with recent advances in biometrics, information forensics and security, vulnerability of facial biometric systems against spoofing attacks is still an open problem. Even though several methods have been proposed for photo-based spoofing attack detection, attacks performed with videos have been vastly overlooked, which hinders the use of the facial biometric systems in modern applications. In this paper, we present an algorithm for video-based spoofing attack detection through the analysis of global information which is invariant to content, since we discard video contents and analyze content-independent noise signatures present in the video related to the unique acquisition processes. Our approach takes advantage of noise signatures generated by the recaptured video to distinguish between fake and valid access videos. For that, we use the Fourier spectrum followed by the computation of video visual rhythms and the extraction of different characterization methods. For evaluation, we consider the novel unicamp video-attack database, which comprises 17 076 videos composed of real access and spoofing attack videos. In addition, we evaluate the proposed method using the replay-attack database, which contains photo-based and video-based face spoofing attacks. Allan Pinto, William Robson Schwartz, Hélio Pedrini, Anderson Rocha 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2015 | Face Spoofing Detection Through Visual Codebooks of Spectral Temporal CubesabstractDespite important recent advances, the vulnerability of biometric systems to spoofing attacks is still an open problem. Spoof attacks occur when impostor users present synthetic biometric samples of a valid user to the biometric system seeking to deceive it. Considering the case of face biometrics, a spoofing attack consists in presenting a fake sample (e.g., photograph, digital video, or even a 3D mask) to the acquisition sensor with the facial information of a valid user. In this paper, we introduce a low cost and software-based method for detecting spoofing attempts in face recognition systems. Our hypothesis is that during acquisition, there will be inevitable artifacts left behind in the recaptured biometric samples allowing us to create a discriminative signature of the video generated by the biometric sensor. To characterize these artifacts, we extract time-spectral feature descriptors from the video, which can be understood as a low-level feature descriptor that gathers temporal and spectral information across the biometric sample and use the visual codebook concept to find mid-level feature descriptors computed from the low-level ones. Such descriptors are more robust for detecting several kinds of attacks than the low-level ones. The experimental results show the effectiveness of the proposed method for detecting different types of attacks in a variety of scenarios and data sets, including photos, videos, and 3D masks. Allan Pinto, Hélio Pedrini, William Robson Schwartz, Anderson Rocha 0001 |
IEEE Trans. Image Process. | 4 |
| 2014 | Large-Scale Micro-Blog Authorship Attribution: Beyond Simple Feature Engineering
Thiago Cavalcante, Anderson Rocha 0001, Ariadne Carvalho |
CIARP | 2 |
| 2014 | A Multiscale and Multi-Perturbation Blind Forensic Technique for Median Detecting
Anselmo Ferreira, Anderson Rocha 0001 |
CIARP | 2 |
| 2014 | Who is my parent? Reconstructing video sequences from partially matching shotsabstractNowadays, a significant fraction of the available video content is created by reusing already existing online videos. In these cases, the source video is seldom reused as is. Conversely, it is typically time clipped to extract only a subset of the original frames, and other transformations are commonly applied (e.g., cropping, logo insertion, etc.). In this paper, we analyze a pool of videos related to the same event or topic. We propose a method that aims at automatically reconstructing the content of the original source videos, i.e., the parent sequences, by splicing together sets of near-duplicate shots seemingly extracted from the same parent sequence. The result of the analysis shows how content is reused, thus revealing the intent of content creators, and enables us to reconstruct a parent sequence also when it is no longer available online. In doing so, we make use of a robust-hash algorithm that allows us to detect whether groups of frames are near-duplicates. Based on that, we developed an algorithm to automatically find near-duplicate matchings between multiple parts of multiple sequences. All the near-duplicate parts are finally temporally aligned to reconstruct the parent sequence. The proposed method is validated with both synthetic and real world datasets downloaded from YouTube. Silvia Lameri, Paolo Bestagini, Andrea Melloni, Simone Milani, Anderson Rocha 0001, Marco Tagliasacchi, Stefano Tubaro |
ICIP | 5 |
| 2014 | Multiple parenting identification in image phylogenyabstractImage phylogeny deals with tracing back parent-child relationships among near duplicates, images that share the same semantic content. This approach results in a visual structure showing the inheritance of semantic content among images, called phylogeny tree. In this paper, we extend upon the image phylogeny's original formulation, which considers that an image may inherit content from only a single parent, to deal with situations whereby an image may inherit it from multiple different parents. Our objective is to find the multiple parenting relationships in a set of images, a problem which we refer to as multiple parenting phylogeny. The proposed solution works by first identifying near-duplicate groups and reconstructing their phylogenies; then among the found groups we determine the one(s) representing the composition images; finally, we detect the parenting relations between those compositions and the source images used to create them. Alberto A. de Oliveira, Pasquale Ferrara, Alessia De Rosa, Alessandro Piva, Mauro Barni, Siome Goldenstein, Zanoni Dias, Anderson Rocha 0001 |
ICIP | 8 |
| 2014 | Automatic identification of fruit flies (Diptera: Tephritidae)
Fábio Augusto Faria, P. Perre, Roberto A. Zucchi, Leonardo Ré Jorge, T. M. Lewinsohn, Anderson Rocha 0001, Ricardo da Silva Torres |
J. Vis. Commun. Image Represent. | 6 |
| 2014 | Open set source camera attribution and device linking
Filipe de Oliveira Costa, Ewerton Silva, Michael Eckmann, Walter J. Scheirer, Anderson Rocha 0001 |
Pattern Recognit. Lett. | 5 |
| 2014 | Visual words dictionaries and fusion techniques for searching people through textual and visual attributes
Junior Fabian, Ramon Pires, Anderson Rocha 0001 |
Pattern Recognit. Lett. | 3 |
| 2014 | A framework for selection and fusion of pattern classifiers in multimedia recognition
Fábio Augusto Faria, Jefersson A. dos Santos, Anderson Rocha 0001, Ricardo da Silva Torres |
Pattern Recognit. Lett. | 3 |
| 2014 | A multiple camera methodology for automatic localization and tracking of futsal players
Erikson Freitas de Morais, Anselmo Ferreira, Sergio Augusto Cunha, Ricardo M. L. Barros, Anderson Rocha 0001, Siome Goldenstein |
Pattern Recognit. Lett. | 5 |
| 2014 | Learning Person-Specific Representations From Faces in the WildabstractHumans are natural face recognition experts, far out-performing current automated face recognition algorithms, especially in naturalistic, “in the wild” settings. However, a striking feature of human face recognition is that we are dramatically better at recognizing highly familiar faces, presumably because we can leverage large amounts of past experience with the appearance of an individual to aid future recognition. Meanwhile, the analogous situation in automated face recognition, where a large number of training examples of an individual are available, has been largely underexplored, in spite of the increasing relevance of this setting in the age of social media. Inspired by these observations, we propose to explicitly learn enhanced face representations on a per-individual basis, and we present two methods enabling this approach. By learning and operating within person-specific representations, we are able to significantly outperform the previous state-of-the-art on PubFig83, a challenging benchmark for familiar face recognition in the wild, using a novel method for learning representations in deep visual hierarchies. We suggest that such person-specific representations aid recognition by introducing an intermediate form of regularization to the problem. Giovani Chiachia, Alexandre X. Falcão, Nicolas Pinto, Anderson Rocha 0001, David D. Cox |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2014 | Image Phylogeny Forests ReconstructionabstractToday, a simple search for an image on the Web can return thousands of related images. Some results are exact copies, some are variants (or near-duplicates) of the same digital image, and others are unrelated. Although we can recognize some of these images as being semantically similar, it is not as straightforward to find which image is the original. It is not easy either to find the chain of transformations used to create each modified version. There are several approaches in the literature to identify near-duplicate images, as well as to reconstruct their relational structure. For the latter, a common representation uses the parent-child relationship, allowing us to visualize the evolution of modifications as a phylogeny tree. However, most of the approaches are restricted to the case of finding the tree of evolution of the near-duplicates, with few works dealing with sets of trees. Since one set of near-duplicates can contain n independent subsets, it is necessary to reconstruct not only one phylogeny tree, but several trees that will compose a phylogeny forest. In this paper, through the analysis of the state-of-the-art image phylogeny algorithms, we introduce a novel approach to deal with phylogeny forests, based on different combinations of these algorithms, aiming at improving their reconstruction accuracy. We analyze the effectiveness of each combination and evaluate our method with more than 40 000 testing cases, using quantitative metrics. Filipe de Oliveira Costa, Marina A. Oikawa, Zanoni Dias, Siome Goldenstein, Anderson Rocha 0001 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2014 | Multiclass From Binary: Expanding One-Versus-All, One-Versus-One and ECOC-Based ApproachesabstractRecently, there has been a lot of success in the development of effective binary classifiers. Although many statistical classification techniques have natural multiclass extensions, some, such as the support vector machines, do not. The existing techniques for mapping multiclass problems onto a set of simpler binary classification problems run into serious efficiency problems when there are hundreds or even thousands of classes, and these are the scenarios where this paper's contributions shine. We introduce the concept of correlation and joint probability of base binary learners. We learn these properties during the training stage, group the binary leaner's based on their independence and, with a Bayesian approach, combine the results to predict the class of a new instance. Finally, we also discuss two additional strategies: one to reduce the number of required base learners in the multiclass classification, and another to find new base learners that might best complement the existing set. We use these two new procedures iteratively to complement the initial solution and improve the overall performance. This paper has two goals: finding the most discriminative binary classifiers to solve a multiclass problem and keeping up the efficiency, i.e., small number of base learners. We validate and compare the method with a diverse set of methods of the literature in several public available datasets that range from small (10 to 26 classes) to large multiclass problems (1000 classes) always using simple reproducible scenarios. Anderson Rocha 0001, Siome Goldenstein |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2013 | Exploring heuristic and optimum branching algorithms for image phylogeny
Zanoni Dias, Siome Goldenstein, Anderson Rocha 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2013 | Computer generated images vs. digital photographs: A synergetic feature and classifier combination approach
Eric Tokuda, Hélio Pedrini, Anderson Rocha 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2013 | Toward Open Set RecognitionabstractTo date, almost all experimental evaluations of machine learning-based recognition algorithms in computer vision have taken the form of "closed set" recognition, whereby all testing classes are known at training time. A more realistic scenario for vision applications is "open set" recognition, where incomplete knowledge of the world is present at training time, and unknown classes can be submitted to an algorithm during testing. This paper explores the nature of open set recognition and formalizes its definition as a constrained minimization problem. The open set recognition problem is not well addressed by existing algorithms because it requires strong generalization. As a step toward a solution, we introduce a novel "1-vs-set machine," which sculpts a decision space from the marginal distances of a 1-class or binary SVM with a linear kernel. This methodology applies to several different applications in computer vision where open set recognition is a challenging problem, including object recognition and face verification. We consider both in this work, with large scale cross-dataset experiments performed over the Caltech 256 and ImageNet sets, as well as face matching experiments performed over the Labeled Faces in the Wild set. The experiments highlight the effectiveness of machines adapted for open set evaluation compared to existing 1-class and binary SVMs for the same tasks. Walter J. Scheirer, Anderson Rocha 0001, Archana Sapkota, Terrance E. Boult |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2013 | Exposing Digital Image Forgeries by Illumination Color ClassificationabstractFor decades, photographs have been used to document space-time events and they have often served as evidence in courts. Although photographers are able to create composites of analog pictures, this process is very time consuming and requires expert knowledge. Today, however, powerful digital image editing software makes image modifications straightforward. This undermines our trust in photographs and, in particular, questions pictures as evidence for real-world events. In this paper, we analyze one of the most common forms of photographic manipulation, known as image composition or splicing. We propose a forgery detection method that exploits subtle inconsistencies in the color of the illumination of images. Our approach is machine-learning-based and requires minimal user interaction. The technique is applicable to images containing two or more people and requires no expert interaction for the tampering decision. To achieve this, we incorporate information from physics- and statistical-based illuminant estimators on image regions of similar material. From these illuminant estimates, we extract texture- and edge-based features which are then provided to a machine-learning approach for automatic decision-making. The classification performance using an SVM meta-fusion classifier is promising. It yields detection rates of 86% on a new benchmark dataset consisting of 200 images, and 83% on 50 images that were collected from the Internet. Tiago Jose de Carvalho, Christian Riess, Elli Angelopoulou, Hélio Pedrini, Anderson Rocha 0001 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2012 | Person-Specific Subspace Analysis for Unconstrained Familiar Face Identification
Giovani Chiachia, Nicolas Pinto, William Robson Schwartz, Anderson Rocha 0001, Alexandre X. Falcão, David D. Cox |
BMVC | 4 |
| 2012 | Data fusion for multi-lesion Diabetic Retinopathy detectionabstractScreening of Diabetic Retinopathy (DR) with timely treatment prevents blindness. Several researchers have focused their work on the development of computer-aided lesion-specific detectors. Combining detectors is a complex task as frequently the detectors have different properties and constraints and are not designed under a unified framework. We extend our previous work for detecting DR lesions based on points of interest and visual words to include additional detectors for the most common DR lesions and investigate fusion techniques to combine different classifiers for classification of normal or signs of diabetic retinopathy. The combination methods show promising results and shed light on the possible advantages of combining complementary lesion detectors for the DR diagnosis problem. Herbert F. Jelinek, Ramon Pires, Rafael Padilha, Siome Goldenstein, Jacques Wainer, Terry Bossomaier, Anderson Rocha 0001 |
CBMS | 7 |
| 2012 | Automatic visual dictionary generation through Optimum-Path Forest clusteringabstractImage categorization by means of bag of visual words has received increasing attention by the image processing and vision communities in the last years. In these approaches, each image is represented by invariant points of interest which are mapped to a Hilbert Space representing a visual dictionary which aims at comprising the most discriminative features in a set of images. Notwithstanding, the main problem of such approaches is to find a compact and representative dictionary. Finding such representative dictionary automatically with no user intervention is an even more difficult task. In this paper, we propose a method to automatically find such dictionary by employing a recent developed graph-based clustering algorithm called Optimum-Path Forest, which does not make any assumption about the visual dictionary's size and is more efficient and effective than the state-of-the-art techniques used for dictionary generation. Luis C. S. Afonso, João Paulo Papa, Luciene P. Papa, Aparecido Nilceu Marana, Anderson Rocha 0001 |
ICIP | 5 |
| 2012 | Descriptor correlation analysis for remote sensing image multi-scale classification
Jefersson A. dos Santos, Fábio Augusto Faria, Ricardo da Silva Torres, Anderson Rocha 0001, Philippe Henri Gosselin, Sylvie Philipp-Foliguet, Alexandre X. Falcão |
ICPR | 4 |
| 2012 | Automatic fusion of region-based classifiers for coffee crop recognitionabstractCoffee crop recognition in remote sensing images is a complex task. It poses several challenges due to different spectral responses and texture patterns that can be extracted from coffee regions. This paper presents a novel framework for combining different classifiers using support vector machine technique (SVM), which try to learn with each one of classifiers previews experiences (meta-learning). We investigate the combination of seven learning methods and seven image descriptors aiming at creating low-cost classifiers for coffee crops recognition. The objective is to provide an effective mechanism for coffee crop recognition by fusion of region-based classifiers in remote sensing images. The experiments showed that the proposed framework for fusion of classifiers produces better results than the traditional majority voting fusion approach and all base classifiers tested. Fábio Augusto Faria, Jefersson A. dos Santos, Ricardo da Silva Torres, Anderson Rocha 0001, Alexandre X. Falcão |
IGARSS | 4 |
| 2012 | How Far do We Get Using Machine Learning Black-Boxes?abstractWith several good research groups actively working in machine learning (ML) approaches, we have now the concept of self-containing machine learning solutions that oftentimes work out-of-the-box leading to the concept of ML black-boxes. Although it is important to have such black-boxes helping researchers to deal with several problems nowadays, it comes with an inherent problem increasingly more evident: we have observed that researchers and students are progressively relying on ML black-boxes and, usually, achieving results without knowing the machinery of the classifiers. In this regard, this paper discusses the use of machine learning black-boxes and poses the question of how far we can get using these out-of-the-box solutions instead of going deeper into the machinery of the classifiers. The paper focuses on three aspects of classifiers: (1) the way they compare examples in the feature space; (2) the impact of using features with variable dimensionality; and (3) the impact of using binary classifiers to solve a multi-class problem. We show how knowledge about the classifier's machinery can improve the results way beyond out-of-the-box machine learning solutions. Anderson Rocha 0001, João Paulo Papa, Luis A. A. Meira |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2012 | Image Phylogeny by Minimal Spanning TreesabstractNowadays, digital content is widespread and also easily redistributable, either lawfully or unlawfully. Images and other digital content can also mutate as they spread out. For example, after images are posted on the Internet, other users can copy, resize and/or re-encode them and then repost their versions, thereby generating similar but not identical copies. While it is straightforward to detect exact image duplicates, this is not the case for slightly modified versions. In the last decade, some researchers have successfully focused on the design and deployment of near-duplicate detection and recognition systems to identify the cohabiting versions of a given document in the wild. Those efforts notwithstanding, only recently have there been the first attempts to go beyond the detection of near-duplicates to find the structure of evolution within a set of images. In this paper, we tackle and formally define the problem of identifying these image relationships within a set of near-duplicate images, what we call Image Phylogeny Tree (IPT), due to its natural analogy with biological systems. The mechanism of building IPTs aims at finding the structure of transformations and their parameters if necessary, among a near-duplicate image set, and has immediate applications in security and law-enforcement, forensics, copyright enforcement, and news tracking services. We devise a method for calculating an asymmetric dissimilarity matrix from a set of near-duplicate images and formally introduce an efficient algorithm to build IPTs from such a matrix. We validate our approach with more than 625000 test cases, including both synthetic and real data, and show that when using an appropriate dissimilarity function we can obtain good IPT reconstruction even when some pieces of information are missing. We also evaluate our solution when there are more than one near-duplicate sets in the pool of analysis and compare to other recent related approaches in the literature. Zanoni Dias, Anderson Rocha 0001, Siome Goldenstein |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2012 | Learning for Meta-RecognitionabstractIn this paper, we consider meta-recognition, an approach for postrecognition score analysis, whereby a prediction of matching accuracy is made from an examination of the tail of the scores produced by a recognition algorithm. This is a general approach that can be applied to any recognition algorithm producing distance or similarity scores. In practice, meta-recognition can be implemented in two different ways: a statistical fitting algorithm based on the extreme value theory, and a machine learning algorithm utilizing features computed from the raw scores. While the statistical algorithm establishes a strong theoretical basis for meta-recognition, the machine learning algorithm is more accurate in its predictions in all of our assessments. In this paper, we present a study of the machine learning algorithm and its associated features for the purpose of building a highly accurate meta-recognition system for security and surveillance applications. Through the use of feature- and decision-level fusion, we achieve levels of accuracy well beyond those of the statistical algorithm, as well as the popular “cohort” model for postrecognition score analysis. In addition, we also explore the theoretical question of why machine learning-based algorithms tend to outperform statistical meta-recognition and provide a partial explanation. We show that our proposed methods are effective for a variety of different recognition applications across security and forensics-oriented computer vision, including biometrics, object recognition, and content-based image retrieval. Walter J. Scheirer, Anderson Rocha 0001, Jonathan Parris, Terrance E. Boult |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2011 | Competition on counter measures to 2-D facial spoofing attacksabstractSpoofing identities using photographs is one of the most common techniques to attack 2-D face recognition systems. There seems to exist no comparative studies of different techniques using the same protocols and data. The motivation behind this competition is to compare the performance of different state-of-the-art algorithms on the same database using a unique evaluation method. Six different teams from universities around the world have participated in the contest. Use of one or multiple techniques from motion, texture analysis and liveness detection appears to be the common trend in this competition. Most of the algorithms are able to clearly separate spoof attempts from real accesses. The results suggest the investigation of more complex attacks. Murali Mohan Chakka, André Anjos, Sébastien Marcel, Roberto Tronci, Daniele Muntoni, Gianluca Fadda, Maurizio Pili, Nicola Sirena, Gabriele Murgia, Marco Ristori, Fabio Roli, Dong Yi, Zhen Lei 0001, Stan Z. Li, William Robson Schwartz, Anderson Rocha 0001, Hélio Pedrini, Javier Lorenzo-Navarro, Modesto Castrillón-Santana, Jukka Komulainen, Abdenour Hadid, Matti Pietikäinen |
IJCB | 18 |
| 2011 | Person-specific face representation for recognitionabstractMost face recognition methods rely on a common feature space to represent the faces, in which the face aspects that better distinguish among all the persons are emphasized. This strategy may be inadequate to represent more appropriate aspects of a specific person's face, since there may be some aspects that are good at distinguishing only a given person from the others. Based on this idea and sup- ported by some findings in the human perception of faces, we propose a face recognition framework that associates a feature space to each person that we intend to recognize. Such feature spaces are conceived to underline the discriminating face aspects of the persons they represent. In order to recognize a probe, we match it to the gallery in all the feature spaces and fuse the results to establish the identity. With the help of an algorithm that we devise, the Discriminant Patch Selection, we were capable of carrying out experiments to intuitively compare the traditional approaches with the person-specific representation. In the performed experiments, the person-specific face representation always resulted in a better identification of the faces. Giovani Chiachia, Alexandre X. Falcão, Anderson Rocha 0001 |
IJCB | 3 |
| 2011 | Face spoofing detection through partial least squares and low-level descriptorsabstractPersonal identity verification based on biometrics has received increasing attention since it allows reliable authentication through intrinsic characteristics, such as face, voice, iris, fingerprint, and gait. Particularly, face recognition techniques have been used in a number of applications, such as security surveillance, access control, crime solving, law enforcement, among others. To strengthen the results of verification, biometric systems must be robust against spoofing attempts with photographs or videos, which are two common ways of bypassing a face recognition system. In this paper, we describe an anti-spoofing solution based on a set of low-level feature descriptors capable of distinguishing between 'live' and 'spoof images and videos. The proposed method explores both spatial and temporal information to learn distinctive characteristics between the two classes. Experiments conducted to validate our solution with datasets containing images and videos show results comparable to state-of-the-art approaches. William Robson Schwartz, Anderson Rocha 0001, Hélio Pedrini |
IJCB | 2 |
| 2011 | Image categorization through optimum path forest and visual wordsabstractDifferent from the first attempts to solve the image categorization problem (often based on global features), recently, several researchers have been tackling this research branch through a new vantage point - using features around locally invariant interest points and visual dictionaries. Although several advances have been done in the visual dictionaries literature in the past few years, a problem we still need to cope with is calculation of the number of representative words in the dictionary. Therefore, in this paper we introduce a new solution for automatically finding the number of visual words in an N-Way image categorization problem by means of supervised pattern classification based on optimum-path forest. João Paulo Papa, Anderson Rocha 0001 |
ICIP | 2 |
| 2011 | Face liveness detection under bad illumination conditionsabstractSpoofing face recognition systems with photos or videos of someone else is not difficult. Sometimes, all one needs is to display a picture on a laptop monitor or a printed photograph to the biometric system. In order to detect this kind of spoofs, in this paper we present a solution that works either with printed or LCD displayed photographs, even under bad illumination conditions without extra-devices or user involvement. Tests conducted on large databases show good improvements of classification accuracy as well as true positive and false positive rates compared to the state-of-the-art. Bruno Peixoto, Carolina Michelassi, Anderson Rocha 0001 |
ICIP | 3 |
| 2011 | Eye specular highlights telltales for digital forensics: A machine learning approachabstractAmong the possible forms of photographic fabrication and manipulation, there is an increasing number of composite pictures containing people. With such compositions, it is very common to see politicians depicted side-by-side with criminals during election campaigns, or even Hollywood superstars relationships being wrecked by allegedly affairs depicted in gossip magazines. Thinking about this problem, in this paper we analyze telltales obtained from highlights in the eyes of every person standing in a picture in order to decide whether or not those people were really together at the moment of such image acquisition. We validate our approach with a data set containing realistic photographic compositions, as well as authentic unchanged pictures. As a result, our proposed extension improves the classification accuracy of the state-of-art solution in more than 20%. Priscila Saboia, Tiago Jose de Carvalho, Anderson Rocha 0001 |
ICIP | 3 |
| 2011 | Meta-Recognition: The Theory and Practice of Recognition Score AnalysisabstractIn this paper, we define meta-recognition, a performance prediction method for recognition algorithms, and examine the theoretical basis for its postrecognition score analysis form through the use of the statistical extreme value theory (EVT). The ability to predict the performance of a recognition system based on its outputs for each match instance is desirable for a number of important reasons, including automatic threshold selection for determining matches and nonmatches, and automatic algorithm selection or weighting for multi-algorithm fusion. The emerging body of literature on postrecognition score analysis has been largely constrained to biometrics, where the analysis has been shown to successfully complement or replace image quality metrics as a predictor. We develop a new statistical predictor based upon the Weibull distribution, which produces accurate results on a per instance recognition basis across different recognition problems. Experimental results are provided for two different face recognition algorithms, a fingerprint recognition algorithm, a SIFT-based object recognition system, and a content-based image retrieval system. Walter J. Scheirer, Anderson Rocha 0001, Ross J. Micheals, Terrance E. Boult |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2010 | Robust Fusion: Extreme Value Theory for Recognition Score Normalization
Walter J. Scheirer, Anderson Rocha 0001, Ross J. Micheals, Terrance E. Boult |
ECCV (3) | 2 |
| 2010 | Label propagation through neuronal synchronyabstractSemi-Supervised Learning (SSL) is a machine learning research area aiming the development of techniques which are able to take advantage from both labeled and unlabeled samples. Additionally, most of the times where SSL techniques can be deployed, only a small portion of samples in the data set is labeled. To deal with such situations in a straightforward fashion, in this paper we introduce a semi-supervised learning approach based on neuronal synchrony in a network of coupled integrate-and-fire neurons. For that, we represent the input data set as a graph and model each of its nodes by an integrate-and-fire neuron. Thereafter, we propagate the class labels from the seed samples to unlabeled samples through the graph by means of the emerging synchronization dynamics. Experimentations on synthetic and real data show that the introduced technique achieves good classification results regardless the feature space distribution or geometrical shape. Marcos G. Quiles, Liang Zhao 0001, Fabricio A. Breve, Anderson Rocha 0001 |
IJCNN | 4 |
| 2010 | Progressive randomization: Seeing the unseen
Anderson Rocha 0001, Siome Goldenstein |
Comput. Vis. Image Underst. | 1 |
| 2008 | Efficient and Flexible Cluster-and-Search for CBIR
Anderson Rocha 0001, Jurandy Almeida, Mario A. Nascimento, Ricardo da Silva Torres, Siome Goldenstein |
ACIVS | 1 |
| 2007 | PR: More than Meets the EyeabstractIn this paper, we introduce a new image descriptor for broad Image Categorization, the Progressive Randomization (PR) that uses perturbations on the values of the Least Significant Bits (LSB) of images. We show that different classes of images have a distinct behavior under our methodology and that using statistical descriptors of LSB occurrences and enough training examples, the method already performs as well or better than comparable existing techniques in the literature. With few training examples PR still has good separability and its accuracy increases with the size of the training set. We validate our method using four image databases with different categories. Anderson Rocha 0001, Siome Goldenstein |
ICCV | 1 |
| 2006 | A Linear-Time Approach for Image Segmentation Using Graph-Cut Measures
Alexandre X. Falcão, Paulo André Vechiatto Miranda, Anderson Rocha 0001 |
ACIVS | 3 |
| 2006 | Progressive Randomization for SteganalysisabstractIn this paper, we describe a new methodology to detect the presence of hidden digital content in the least significant bits (LSB) of images. We introduce the progressive randomization (PR) technique that captures statistical artifacts inserted during the hiding process. Our technique is a progressive application of LSB modifying transformations that receives an image as input, and returns n images that only differ in the LSB from the initial image. Each step of the progressive randomization approach represents a possible content-hiding scenario with increasing size, and increasing LSB entropy. We validate our method with 20,000 real, non-synthetic images. Using only statistical descriptors of LSB occurrences, our method already performs as well or better than comparable techniques in the literature Anderson Rocha 0001, Siome Goldenstein |
MMSP | 1 |