EDBT 2026 Demo / reviewers in the wild / expert
Gustavo Carneiro 0001
dblp:53/3609 · also Gustavo H. M. B. Carneiro
· DBLP profile ↗
159ranked-venue papers
32as first author
68since 2021 · last 2026
0000-0002-5571-6220ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 107 · 21 first-author · 43 since 2021Artificial intelligence and machine learning · 82 · 20 first-author · 36 since 2021Applied, interdisciplinary, general and emerging computing · 45 · 6 first-author · 23 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Coverage-Constrained Human-AI Cooperation with Multiple ExpertsabstractHuman-AI cooperative classification (HAI-CC) aims to develop hybrid intelligent systems that enhance decision-making in various high-stakes real-world scenarios by leveraging both human expertise and AI capabilities. Current HAI-CC methods primarily focus on learning-to-defer (L2D), where decisions are deferred to human experts when AI is not confident, and learning-to-complement (L2C), where AI and human experts make predictions cooperatively. However, existing research in both L2D and L2C has not effectively been explored under diverse expert knowledge to improve decision-making, particularly when constrained by the operation cost of human involvement. In this paper, we address this research gap by proposing the Coverage-constrained Learning to Defer and Complement with Specific Experts (CL2DC) method. In particular, CL2DC assesses input data before making final decisions through either AI prediction alone or by deferring to or complementing a specific human expert. Furthermore, we propose a coverage-constrained optimisation to control the cooperation cost, ensuring it approximates a target probability for AI-only selection. This approach enables an effective assessment of system performance within a specified budget. Comprehensive evaluations on both synthetic and real-world datasets demonstrate that CL2DC achieves superior performance compared to state-of-the-art HAI-CC methods. Zheng Zhang 0046, Cuong Nguyen 0006, Kevin Wells, Thanh-Toan Do, David Rosewarne, Gustavo Carneiro 0001 |
AAAI | 6 |
| 2026 | Layer-Wise High-Impact Parameter Ratio Optimization in Post-Training Quantization for Large Language ModelsabstractCuong Pham, Anh Dung Hoang, Cuong C. Nguyen, Trung Le, Gustavo Carneiro, Thanh-Toan Do. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Cuong Pham 0007, Dung Anh Hoang, Cuong Nguyen 0006, Trung Le 0001, Gustavo Carneiro 0001, Thanh-Toan Do |
ACL (1) | 5 |
| 2026 | Reciprocal Teaching: Dynamic Multi-Model Teacher-Student Learning for Multiple Noisy AnnotationsabstractAs datasets grow, expert-based annotation becomes impractical, making crowdsourcing a scalable alternative. In crowdsourcing, samples are typically annotated by multiple workers and aggregated via majority voting, which ignores annotator-specific biases and introduces noisy labels that impair downstream models. Traditional multi-rater methods attempt to model annotator biases but often overfit with many classes or few, noisy annotators. Learning with Noisy Labels (LNL) methods offer more robust strategies for handling noisy labels, but their assumption of a single noisy label per sample makes extending them to multi-annotator settings non-trivial. To bridge this gap, we propose the Reciprocal Teacher-student Learning from Multi-rater Noisy Annotation (RETINA), which trains annotator-specific models and employs a dynamic teacher–student process to separate clean from noisy samples. Progress in multi-rater learning has also been limited by benchmarks with few classes, fixed noise rates, and no control over annotators. To address this, we introduce the Synthetic MRL (SynMRL) benchmark that contains many classes and controllable noise and annotator settings for systematic evaluation. Experiments on synthetic and real-world data show that RETINA outperforms existing multi-rater methods, particularly in high-noise, low-annotator, many-class settings. Wenjie Ai, Cuong Nguyen 0006, Adrian Hilton 0001, Gustavo Carneiro 0001 |
WACV | 4 |
| 2026 | Unpaired multi-modal multi-label learning for detecting endometriosis signsabstractEndometriosis is a widespread gynecological disorder causing severe pain and infertility, with diagnosis currently relying on slow, costly, and risky laparoscopy. This highlights the critical need for non-invasive imaging diagnostics using transvaginal ultrasound (TVUS) and magnetic resonance imaging (MRI). A key challenge is that patients typically receive only one scan modality in practice, despite TVUS and MRI offering differing diagnostic strengths for endometriosis signs like Pouch of Douglas (POD) obliteration and bowel nodules (BN). Previous work partially addressed this challenge by leveraging unpaired multi-modal data for detecting a single marker: Pouch of Douglas (POD) obliteration. However, this is restrictive because endometriosis signs, such as POD obliteration and bowel nodules (BN), often provide correlated diagnostic cues. Capturing these correlations is essential for accurate detection of endometriosis imaging signs, particularly when combined with multi-modal learning, as each modality offers complementary strengths for different signs. To overcome these limitations, we propose EndoFusion, a novel unpaired multi-modal, multi-label learning framework that enables the detection of POD and BN from TVUS and MRI. Our approach introduces three key innovations: (1) label-based pairing, mixup, and cross-modal feature exchange for robust single-modality inference; (2) Dynamic Mutual Knowledge Distillation (DMKD), which adaptively selects teachers using a worst-student-oriented strategy for effective cross-modal transfer; and (3) label correlations modeling with multi-head attention and a specialized loss to handle imbalance and boost accuracy. This design ensures that knowledge from the superior modality and from co-occurring signs is effectively transferred, mitigating modality-specific weaknesses and improving robustness in imaging sign detection. Experiments on our endometriosis dataset show that our method significantly outperforms all comparison methods, achieving an average AUC of 0.827 (95% CI: 0.790-0.861) when evaluated using single-modality inference. These results represent an initial proof-of-concept toward multi-modal, non-invasive assessment of selected endometriosis imaging signs from MRI and TVUS. Hu Wang 0005, Yutong Xie 0001, Minh-Son To, Steven Knox, Mathew Leonardi, George Condous, Jodie Avery, Louise Hull, Gustavo Carneiro 0001 |
Artif. Intell. Medicine | 10 |
| 2026 | PASS: Peer-agreement based sample selection for training with instance dependent noisy labelsabstractDeep learning encounters significant challenges in the form of noisy-label samples, which can cause the overfitting of trained models. A primary challenge in learning with noisy-label (LNL) techniques is their ability to differentiate between hard samples (clean-label samples near the decision boundary) and instance-dependent noisy (IDN) label samples to allow these samples to be treated differently during training. Existing methodologies to identify IDN samples, including the small-loss hypothesis and feature-based selection, have demonstrated limited efficacy, thus impeding their effectiveness in dealing with real-world label noise. We present Peer-Agreement-based Sample Selection (PASS), a novel approach that utilises three classifiers, where a consensus-driven agreement between two models accurately differentiates between clean and noisy-label IDN samples to train the third model. In contrast to current techniques, PASS is specifically designed to address the complexities of IDN, where noise patterns are correlated with instance features. Our approach seamlessly integrates with existing LNL algorithms to enhance the accuracy of detecting both noisy and clean samples. Comprehensive experiments conducted on simulated benchmarks (CIFAR-100 and Red mini-ImageNet) and real-world datasets (Animal-10N, CIFAR-N, Clothing1M, and mini-WebVision) demonstrated that PASS substantially improved the performance of multiple state-of-the-art methods. This technique achieves superior classification accuracy, particularly in scenarios with high noise levels. 1 Arpit Garg, Cuong Nguyen 0006, Rafael Felix, Thanh-Toan Do, Gustavo Carneiro 0001 |
Image Vis. Comput. | 5 |
| 2026 | Learning to complement with multiple humans
Zheng Zhang 0046, Cuong Nguyen 0006, Kevin Wells, Thanh-Toan Do, Gustavo Carneiro 0001 |
Pattern Recognit. | 5 |
| 2025 | CLOC: Contrastive Learning for Ordinal Classification with Multi-Margin N-pair LossabstractIn ordinal classification, misclassifying neighboring ranks is common, yet the consequences of these errors are not the same. For example, misclassifying benign tumor categories is less consequential, compared to an error at the pre-cancerous to cancerous threshold, which could profoundly influence treatment choices. Despite this, existing ordinal classification methods do not account for the varying importance of these margins, treating all neighboring classes as equally significant. To address this limitation, we propose CLOC, a new margin-based contrastive learning method for ordinal classification that learns an ordered representation based on the optimization of multiple margins with a novel multi-margin n-pair loss (MMNP). CLOC enables flexible decision boundaries across key adjacent categories, facilitating smooth transitions between classes and reducing the risk of overfitting to biases present in the training data. We provide empirical discussion regarding the properties of MMNP and show experimental results on five real-world image datasets (Adience, Historical Colour Image Dating, Knee Osteoarthritis, Indian Diabetic Retinopathy Image, and Breast Carcinoma Subtyping) and one synthetic dataset simulating clinical decision bias. Our results demonstrate that CLOC outperforms existing ordinal classification methods and show the interpretability and controllability of CLOC in learning meaningful, ordered representations that align with clinical and practical needs. Dileepa Pitawela, Gustavo Carneiro 0001, Hsiang-Ting Chen |
CVPR | 2 |
| 2025 | Probabilistic Learning to Defer: Handling Missing Expert Annotations and Controlling Workload DistributionabstractRecent progress in machine learning research is gradually shifting its focus towards *human-AI cooperation* due to the advantages of exploiting the reliability of human experts and the efficiency of AI models. One of the promising approaches in human-AI cooperation is *learning to defer* (L2D), where the system analyses the input data and decides to make its own decision or defer to human experts. Although L2D has demonstrated state-of-the-art performance, in its standard setting, L2D entails a severe limitation: all human experts must annotate the whole training dataset of interest, resulting in a time-consuming and expensive annotation process that can subsequently influence the size and diversity of the training set. Moreover, the current L2D does not have a principled way to control workload distribution among human experts and the AI classifier, which is critical to optimise resource allocation. We, therefore, propose a new probabilistic modelling approach inspired by the mixture-of-experts, where the Expectation - Maximisation algorithm is leverage to address the issue of missing expert's annotations. Furthermore, we introduce a constraint, which can be solved efficiently during the E-step, to control the workload distribution among human experts and the AI classifier. Empirical evaluation on synthetic and real-world datasets shows that our proposed probabilistic approach performs competitively, or surpasses previously proposed methods assessed on the same benchmarks. Cuong Nguyen 0006, Thanh-Toan Do, Gustavo Carneiro 0001 |
ICLR | 3 |
| 2025 | Risk Estimation of Knee Osteoarthritis Progression via Predictive Multi-task Modelling from Efficient Diffusion Model Using X-Ray Images
Adrian Hilton 0001, Gustavo Carneiro 0001 |
MICCAI (14) | 3 |
| 2025 | A Novel Perspective for Multi-Modal Multi-Label Skin Lesion ClassificationabstractThe efficacy of deep learning-based Computer-Aided Diagnosis (CAD) methods for skin diseases relies on analyzing multiple data modalities (i.e., clinical+dermoscopic images, and patient metadata) and addressing the challenges of multi-label classification. Current approaches tend to rely on limited multi-modal techniques and treat the multi-label problem as a multiple multi-class problem, overlooking issues related to imbalanced learning and multi-label correlation. This paper introduces the innovative Skin Lesion Classifier, utilizing a Multi-modal Multilabel TransFormer-based model (SkinM2Former). For multi-modal analysis, we introduce the Tri-Modal Cross-attention Transformer (TMCT) that fuses the three image and metadata modalities at various feature levels of a transformer encoder. For multi-label classification, we introduce a multi-head attention (MHA) module to learn multi-label correlations, complemented by an optimisation that handles multi-label and imbalanced learning problems. SkinM2Former achieves a mean average accuracy of 77.27% and a mean diagnostic accuracy of 77.85% on the public Derm7pt dataset, outperforming state-of-the-art (SOTA) methods. Yutong Xie 0001, Hu Wang 0005, Jodie Avery, Louise Hull, Gustavo Carneiro 0001 |
WACV | 6 |
| 2025 | Leveraging labelled data knowledge: A cooperative rectification learning network for semi-supervised 3D medical image segmentation
Yanyan Wang 0007, Kechen Song, Yuyuan Liu, Yunhui Yan, Gustavo Carneiro 0001 |
Medical Image Anal. | 6 |
| 2025 | Mixture of Gaussian-Distributed Prototypes With Generative Modelling for Interpretable and Trustworthy Image RecognitionabstractPrototypical-part methods, e.g., ProtoPNet, enhance interpretability in image recognition by linking predictions to training prototypes, thereby offering intuitive insights into their decision-making. Existing methods, which rely on a point-based learning of prototypes, typically face two critical issues: 1) the learned prototypes have limited representation power and are not suitable to detect Out-of-Distribution (OoD) inputs, reducing their decision trustworthiness; and 2) the necessary projection of the learned prototypes back into the space of training images causes a drastic degradation in the predictive performance. Furthermore, current prototype learning adopts an aggressive approach that considers only the most active object parts during training, while overlooking sub-salient object regions which still hold crucial classification information. In this paper, we present a new generative paradigm to learn prototype distributions, termed as Mixture of Gaussian-distributed Prototypes (MGProto). The distribution of prototypes from MGProto enables both interpretable image classification and trustworthy recognition of OoD inputs. The optimisation of MGProto naturally projects the learned prototype distributions back into the training image space, thereby addressing the performance degradation caused by prototype projection. Additionally, we develop a novel and effective prototype mining strategy that considers not only the most active but also sub-salient object parts. To promote model compactness, we further propose to prune MGProto by removing prototypes with low importance priors. Experiments on CUB-200-2011, Stanford Cars, Stanford Dogs, and Oxford-IIIT Pets datasets show that MGProto achieves state-of-the-art image recognition and OoD detection performances, while providing encouraging interpretability results. Chong Wang 0012, Yuanhong Chen, Fengbei Liu, Yuyuan Liu, Davis J. McCarthy, Helen Frazer, Gustavo Carneiro 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2025 | ANNE: Adaptive Nearest Neighbours and Eigenvector-based sample selection for robust learning with noisy labels
Filipe R. Cordeiro, Gustavo Carneiro 0001 |
Pattern Recognit. | 2 |
| 2025 | Progressive Mining and Dynamic Distillation of Hierarchical Prototypes for Disease Classification and LocalisationabstractConstructing effective representation of lesions is essential for disease classification and localization in medical image analysis. Prototype-based models address this by leveraging visual prototypes to capture representative lesion patterns, yet effectively handling the complexity of diverse lesion characteristics remains a critical challenge, as they typically rely on single-level, fixed-size prototypes and suffer from prototype redundancy. In this paper, we present HierProtoPNet, a new prototype-based framework designed to handle the complexity of lesions in medical images. HierProtoPNet leverages hierarchical visual prototypes across different semantic feature granularities to effectively capture diverse lesion patterns. To prevent redundancy and increase utility of the prototypes, we devise a novel prototype mining paradigm to progressively discover semantically distinct prototypes, offering multi-level complementary analysis of lesions. Also, we introduce a dynamic knowledge distillation strategy that allows transferring essential classification information across hierarchical levels, thereby improving generalisation performance. Comprehensive experiments show that HierProtoPNet achieves state-of-the-art classification performances in three benchmarks: binary breast cancer screening, multi-class retinal disease diagnosis, and multi-label chest X-ray classification. Quantitative assessments also illustrate HierProtoPNet's significant advantages in weakly-supervised disease localisation and segmentation. Chong Wang 0012, Fengbei Liu, Yuanhong Chen, Chun Fung Kwok, Michael Elliott, Carlos A. Peña-Solórzano, Davis J. McCarthy, Helen Frazer, Gustavo Carneiro 0001 |
IEEE J. Biomed. Health Informatics | 9 |
| 2025 | Self-Correcting ClusteringabstractThe incorporation of target distribution significantly enhances the success of deep clustering. However, most of the related deep clustering methods suffer from two drawbacks: (1) manually-designed target distribution functions with uncertain performance and (2) cluster misassignment accumulation. To address these issues, aSelf-CorrectingClustering (Self-CC) framework is proposed. In Self-CC, a robust target distribution solver (RTDS) is designed to automatically predict the target distribution and alleviate the adverse influence of misassignments. Specifically, RTDS divides the high confidence samples selected according to the cluster assignments predicted by a clustering module into labeled samples with correct pseudo labels and unlabeled samples of possible misassignments by modeling its training loss distribution. With the divided data, RTDS can be trained in a semi-supervised way. The critical hyperparameter which controls the semi-supervised training process can be set adaptively by estimating the distribution property of misassignments in the pseudo-label space with the support of a theoretical analysis. The target distribution can be predicted by the well-trained RTDS automatically, optimizing the clustering module and correcting misassignments in the cluster assignments. The clustering module and RTDS mutually promote each other forming a positive feedback loop. Extensive experiments on four benchmark datasets demonstrate the effectiveness of the proposed Self-CC. Hanxuan Wang, Zixuan Wang 0012, Yuxuan Yan, Gustavo Carneiro 0001, Zhen Wang 0004 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2025 | Translation Consistent Semi-Supervised Segmentation for 3D Medical Imagesabstract3D medical image segmentation methods have been successful, but their dependence on large amounts of voxel-level annotated data is a disadvantage that needs to be addressed given the high cost to obtain such annotation. Semi-supervised learning (SSL) solves this issue by training models with a large unlabelled and a small labelled dataset. The most successful SSL approaches are based on consistency learning that minimises the distance between model responses obtained from perturbed views of the unlabelled data. These perturbations usually keep the spatial input context between views fairly consistent, which may cause the model to learn segmentation patterns from the spatial input contexts instead of the foreground objects. In this paper, we introduce the Translation Consistent Co-training (TraCoCo) which is a consistency learning SSL method that perturbs the input data views by varying their spatial input context, allowing the model to learn segmentation patterns from foreground objects. Furthermore, we propose a new Confident Regional Cross entropy (CRC) loss, which improves training convergence and keeps the robustness to co-training pseudo-labelling mistakes. Our method yields state-of-the-art (SOTA) results for several 3D data benchmarks, such as the Left Atrium (LA), Pancreas-CT (Pancreas), and Brain Tumor Segmentation (BraTS19). Our method also attains best results on a 2D-slice benchmark, namely the Automated Cardiac Diagnosis Challenge (ACDC), further demonstrating its effectiveness. Our code, training logs and checkpoints are available at https://github.com/yyliu01/ TraCoCo. Yuyuan Liu, Yu Tian 0001, Chong Wang 0012, Yuanhong Chen, Fengbei Liu, Vasileios Belagiannis, Gustavo Carneiro 0001 |
IEEE Trans. Medical Imaging | 7 |
| 2025 | Cross- and Intra-Image Prototypical Learning for Multi-Label Disease Diagnosis and InterpretationabstractRecent advances in prototypical learning have shown remarkable potential to provide useful decision interpretations associating activation maps and predictions with class-specific training prototypes. Such prototypical learning has been well-studied for various single-label diseases, but for quite relevant and more challenging multi-label diagnosis, where multiple diseases are often concurrent within an image, existing prototypical learning models struggle to obtain meaningful activation maps and effective class prototypes due to the entanglement of the multiple diseases. In this paper, we present a novel Cross- and Intra-image Prototypical Learning (CIPL) framework, for accurate multi-label disease diagnosis and interpretation from medical images. CIPL takes advantage of common cross-image semantics to disentangle the multiple diseases when learning the prototypes, allowing a comprehensive understanding of complicated pathological lesions. Furthermore, we propose a new two-level alignment-based regularisation strategy that effectively leverages consistent intra-image information to enhance interpretation robustness and predictive performance. Extensive experiments show that our CIPL attains the state-of-the-art (SOTA) classification accuracy in two public multi-label benchmarks of disease diagnosis: thoracic radiography and fundus images. Quantitative interpretability results show that CIPL also has superiority in weakly-supervised thoracic disease localisation over other leading saliency- and prototype-based explanation methods. Chong Wang 0012, Fengbei Liu, Yuanhong Chen, Helen Frazer, Gustavo Carneiro 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2024 | Unraveling Instance Associations: A Closer Look for Audio-Visual SegmentationabstractAudio-visual segmentation (AVS) is a challenging task that involves accurately segmenting sounding objects based on audio-visual cues. The effectiveness of audio-visual learning critically depends on achieving accurate cross-modal alignment between sound and visual objects. Successful audio-visual learning requires two essential components: 1) a challenging dataset with high-quality pixel-level multiclass annotated images associated with audio files, and 2) a model that can establish strong links between audio information and its corresponding visual object. However, these requirements are only partially addressed by current methods, with training sets containing biased audio-visual data, and models that generalise poorly beyond this biased training set. In this work, we propose a new cost-effective strategy to build challenging and relatively unbiased high-quality audio-visual segmentation benchmarks. We also propose a new informative sample mining method for audio-visual supervised contrastive learning to leverage discriminative contrastive samples to enforce cross-modal understanding. We show empirical results that demonstrate the effectiveness of our benchmark. Furthermore, experiments conducted on existing AVS datasets and on our new benchmark show that our method achieves state-of-the-art (SOTA) segmentation accuracy11This work was supported by Australian Research Council through grant FT190100525. Yuanhong Chen, Yuyuan Liu, Hu Wang 0005, Fengbei Liu, Chong Wang 0012, Helen Frazer, Gustavo Carneiro 0001 |
CVPR | 7 |
| 2024 | CPM: Class-Conditional Prompting Machine for Audio-Visual Segmentation
Yuanhong Chen, Chong Wang 0012, Yuyuan Liu, Hu Wang 0005, Gustavo Carneiro 0001 |
ECCV (10) | 5 |
| 2024 | Instance-Dependent Noisy-Label Learning with Graphical Model Based Noise-Rate Estimation
Arpit Garg, Cuong Nguyen 0006, Rafael Felix, Thanh-Toan Do, Gustavo Carneiro 0001 |
ECCV (4) | 5 |
| 2024 | ItTakesTwo: Leveraging Peer Representations for Semi-supervised LiDAR Semantic Segmentation
Yuyuan Liu, Yuanhong Chen, Hu Wang 0005, Vasileios Belagiannis, Ian D. Reid 0001, Gustavo Carneiro 0001 |
ECCV (1) | 6 |
| 2024 | MetaAug: Meta-data Augmentation for Post-training Quantization
Cuong Pham 0007, Hoang Anh Dung, Cuong Nguyen 0006, Trung Le 0001, Dinh Q. Phung, Gustavo Carneiro 0001, Thanh-Toan Do |
ECCV (27) | 6 |
| 2024 | Bayesian Detector Combination for Object Detection with Crowdsourced Annotations
Zhi Qin Tan, Olga Isupova, Gustavo Carneiro 0001, Xiatian Zhu, Yunpeng Li 0001 |
ECCV (63) | 3 |
| 2024 | Learning to Complement and to Defer to Multiple Users
Zheng Zhang 0046, Wenjie Ai, Kevin Wells, David Rosewarne, Thanh-Toan Do, Gustavo Carneiro 0001 |
ECCV (56) | 6 |
| 2024 | Effects of Primary Capsule Shapes and Sizes in Capsule Networks
William Tapper, Gustavo Carneiro 0001, Mohammad Hussein, Phillip Evans, Spencer Angus Thomas |
ICPR (5) | 2 |
| 2024 | Frequency Attention for Knowledge DistillationabstractKnowledge distillation is an attractive approach for learning compact deep neural networks, which learns a lightweight student model by distilling knowledge from a complex teacher model. Attention-based knowledge distillation is a specific form of intermediate feature-based knowledge distillation that uses attention mechanisms to encourage the student to better mimic the teacher. However, most of the previous attention-based distillation approaches perform attention in the spatial domain, which primarily affects local regions in the input image. This may not be sufficient when we need to capture the broader context or global information necessary for effective knowledge transfer. In frequency domain, since each frequency is determined from all pixels of the image in spatial domain, it can contain global information about the image. Inspired by the benefits of the frequency domain, we propose a novel module that functions as an attention mechanism in the frequency domain. The module consists of a learnable global filter that can adjust the frequencies of student’s features under the guidance of the teacher’s features, which encourages the student’s features to have patterns similar to the teacher’s features. We then propose an enhanced knowledge review-based distillation model by leveraging the proposed frequency attention module. The extensive experiments with various teacher and student architectures on image classification and object detection benchmark datasets show that the proposed approach outperforms other knowledge distillation methods. Cuong Pham 0007, Van-Anh Nguyen, Trung Le 0001, Dinh Q. Phung, Gustavo Carneiro 0001, Thanh-Toan Do |
WACV | 5 |
| 2024 | BRAIxDet: Learning to detect malignant breast lesion with incomplete annotations
Yuanhong Chen, Yuyuan Liu, Chong Wang 0012, Michael Elliott, Chun Fung Kwok, Carlos A. Peña-Solórzano, Yu Tian 0001, Fengbei Liu, Helen Frazer, Davis J. McCarthy, Gustavo Carneiro 0001 |
Medical Image Anal. | 11 |
| 2024 | Diabetic foot ulcers segmentation challenge report: Benchmark and analysisabstractMonitoring the healing progress of diabetic foot ulcers is a challenging process. Accurate segmentation of foot ulcers can help podiatrists to quantitatively measure the size of wound regions to assist prediction of healing status. The main challenge in this field is the lack of publicly available manual delineation, which can be time consuming and laborious. Recently, methods based on deep learning have shown excellent results in automatic segmentation of medical images, however, they require large-scale datasets for training, and there is limited consensus on which methods perform the best. The 2022 Diabetic Foot Ulcers segmentation challenge was held in conjunction with the 2022 International Conference on Medical Image Computing and Computer Assisted Intervention, which sought to address these issues and stimulate progress in this research domain. A training set of 2000 images exhibiting diabetic foot ulcers was released with corresponding segmentation ground truth masks. Of the 72 (approved) requests from 47 countries, 26 teams used this data to develop fully automated systems to predict the true segmentation masks on a test set of 2000 images, with the corresponding ground truth segmentation masks kept private. Predictions from participating teams were scored and ranked according to their average Dice similarity coefficient of the ground truth masks and prediction masks. The winning team achieved a Dice of 0.7287 for diabetic foot ulcer segmentation. This challenge has now entered a live leaderboard stage where it serves as a challenging benchmark for diabetic foot ulcer segmentation. Moi Hoon Yap, Bill Cassidy, Michal Byra, Ting-Yu Liao, Huahui Yi, Adrian Galdran, Yung-Han Chen, Raphael Brüngel, Sven Koitka, Christoph M. Friedrich, Yu-Wen Lo, Ching-Hui Yang, Kang Li 0004, Qicheng Lao, Miguel Ángel González Ballester, Gustavo Carneiro 0001, Yi-Jen Ju, Juinn-Dar Huang, Joseph Pappachan, Neil D. Reeves, Vishnu Chandrabalan, Darren Dancey, Connah Kendrick |
Medical Image Anal. | 16 |
| 2024 | AIROGS: Artificial Intelligence for Robust Glaucoma Screening ChallengeabstractThe early detection of glaucoma is essential in preventing visual impairment. Artificial intelligence (AI) can be used to analyze color fundus photographs (CFPs) in a cost-effective manner, making glaucoma screening more accessible. While AI models for glaucoma screening from CFPs have shown promising results in laboratory settings, their performance decreases significantly in real-world scenarios due to the presence of out-of-distribution and low-quality images. To address this issue, we propose the Artificial Intelligence for Robust Glaucoma Screening (AIROGS) challenge. This challenge includes a large dataset of around 113,000 images from about 60,000 patients and 500 different screening centers, and encourages the development of algorithms that are robust to ungradable and unexpected input data. We evaluated solutions from 14 teams in this paper and found that the best teams performed similarly to a set of 20 expert ophthalmologists and optometrists. The highest-scoring team achieved an area under the receiver operating characteristic curve of 0.99 (95% CI: 0.98-0.99) for detecting ungradable images on-the-fly. Additionally, many of the algorithms showed robust performance when tested on three other publicly available datasets. These results demonstrate the feasibility of robust AI-enabled glaucoma screening. Coen de Vente, Koen A. Vermeer, Nicolas Jaccard, He Wang 0016, Hongyi Sun, Firas Khader, Daniel Truhn, Temirgali Aimyshev, Yerkebulan Zhanibekuly, Tien-Dung Le, Adrian Galdran, Miguel Ángel González Ballester, Gustavo Carneiro 0001, Devika R. G., Hrishikesh Panikkasseril Sethumadhavan, Densen Puthussery, Hong Liu 0007, Zekang Yang, Satoshi Kondo, Satoshi Kasai, Ashritha Durvasula, Jónathan Heras, Miguel Ángel Zapata, Teresa Araujo, Guilherme Aresta, Hrvoje Bogunovic, Mustafa Arikan, Yeong Chan Lee, Hyun Bin Cho, Yoon Ho Choi, Abdul Qayyum 0002, Muhammad Imran Razzak, Bram van Ginneken, Hans G. Lemij, Clara I. Sánchez |
IEEE Trans. Medical Imaging | 13 |
| 2024 | An Interpretable and Accurate Deep-Learning Diagnosis Framework Modeled With Fully and Semi-Supervised Reciprocal LearningabstractThe deployment of automated deep-learning classifiers in clinical practice has the potential to streamline the diagnosis process and improve the diagnosis accuracy, but the acceptance of those classifiers relies on both their accuracy and interpretability. In general, accurate deep-learning classifiers provide little model interpretability, while interpretable models do not have competitive classification accuracy. In this paper, we introduce a new deep-learning diagnosis framework, called InterNRL, that is designed to be highly accurate and interpretable. InterNRL consists of a student-teacher framework, where the student model is an interpretable prototype-based classifier (ProtoPNet) and the teacher is an accurate global image classifier (GlobalNet). The two classifiers are mutually optimised with a novel reciprocal learning paradigm in which the student ProtoPNet learns from optimal pseudo labels produced by the teacher GlobalNet, while GlobalNet learns from ProtoPNet's classification performance and pseudo labels. This reciprocal learning paradigm enables InterNRL to be flexibly optimised under both fully- and semi-supervised learning scenarios, reaching state-of-the-art classification performance in both scenarios for the tasks of breast cancer and retinal disease diagnosis. Moreover, relying on weakly-labelled training images, InterNRL also achieves superior breast cancer localisation and brain tumour segmentation results than other competing methods. Chong Wang 0012, Yuanhong Chen, Fengbei Liu, Michael Elliott, Chun Fung Kwok, Carlos A. Peña-Solórzano, Helen Frazer, Davis J. McCarthy, Gustavo Carneiro 0001 |
IEEE Trans. Medical Imaging | 9 |
| 2023 | Multi-Modal Learning with Missing Modality via Shared-Specific Feature ModellingabstractThe missing modality issue is critical but non-trivial to be solved by multi-modal models. Current methods aiming to handle the missing modality problem in multi-modal tasks, either deal with missing modalities only during evaluation or train separate models to handle specific missing modality settings. In addition, these models are designed for specific tasks, so for example, classification models are not easily adapted to segmentation tasks and vice versa. In this paper, we propose the Shared-Specific Feature Modelling (ShaSpec) method that is considerably simpler and more effective than competing approaches that address the issues above. ShaSpec is designed to take advantage of all available input modalities during training and evaluation by learning shared and specific features to better represent the input data. This is achieved from a strategy that relies on auxiliary tasks based on distribution alignment and domain classification, in addition to a residual feature fusion procedure. Also, the design simplicity of ShaSpec enables its easy adaptation to multiple tasks, such as classification and segmentation. Experiments are conducted on both medical image segmentation and computer vision classification, with results indicating that ShaSpec outperforms competing methods by a large margin. For instance, on BraTS2018, ShaSpec improves the SOTA by more than 3% for enhancing tumour, 5% for tumour core and 3% for whole tumour.11This work received funding from the Australian Government the through Medical Research Futures Fund: Primary Health Care Research Data Infrastructure Grant 2020 and from Endometriosis Australia. G.C. was supported by Australian Research Council through grant FT190100525. Hu Wang 0005, Yuanhong Chen, Congbo Ma, Jodie Avery, Louise Hull, Gustavo Carneiro 0001 |
CVPR | 6 |
| 2023 | BoMD: Bag of Multi-label Descriptors for Noisy Chest X-ray ClassificationabstractDeep learning methods have shown outstanding classification accuracy in medical imaging problems, which is largely attributed to the availability of large-scale datasets manually annotated with clean labels. However, given the high cost of such manual annotation, new medical imaging classification problems may need to rely on machine-generated noisy labels extracted from radiology reports. Indeed, many Chest X-Ray (CXR) classifiers have been modelled from datasets with noisy labels, but their training procedure is in general not robust to noisy-label samples, leading to sub-optimal models. Furthermore, CXR datasets are mostly multi-label, so current multi-class noisy-label learning methods cannot be easily adapted. In this paper, we propose a new method designed for noisy multi-label CXR learning, which detects and smoothly re-labels noisy samples from the dataset to be used in the training of common multi-label classifiers. The proposed method optimises a bag of multi-label descriptors (BoMD) to promote their similarity with the semantic descriptors produced by language models from multi-label image annotations. Our experiments on noisy multi-label training sets and clean testing sets show that our model has state-of-the-art accuracy and robustness in many CXR multi-label classification benchmarks, including a new benchmark that we propose to systematically assess noisy multi-label methods. Code is available at https://github.com/cyh-0/BoMD. Yuanhong Chen, Fengbei Liu, Hu Wang 0005, Chong Wang 0012, Yuyuan Liu, Yu Tian 0001, Gustavo Carneiro 0001 |
ICCV | 7 |
| 2023 | Residual Pattern Learning for Pixel-wise Out-of-Distribution Detection in Semantic SegmentationabstractSemantic segmentation models classify pixels into a set of known ("in-distribution") visual classes. When deployed in an open world, the reliability of these models depends on their ability to not only classify in-distribution pixels but also to detect out-of-distribution (OoD) pixels. Historically, the poor OoD detection performance of these models has motivated the design of methods based on model re-training using synthetic training images that include OoD visual objects. Although successful, these re-trained methods have two issues: 1) their in-distribution segmentation accuracy may drop during re-training, and 2) their OoD detection accuracy does not generalise well to new contexts outside the training set (e.g., from city to country context). In this paper, we mitigate these issues with: (i) a new residual pattern learning (RPL) module that assists the segmentation model to detect OoD pixels with minimal deterioration to inlier segmentation accuracy; and (ii) a novel context-robust contrastive learning (CoroCL) that enforces RPL to robustly detect OoD pixels in various contexts. Our approach improves by around 10% FPR and 7% AuPRC previous state-of-the-art in Fishyscapes, Segment-Me-If-You-Can, and RoadAnomaly datasets. Yuyuan Liu, Choubo Ding, Yu Tian 0001, Guansong Pang, Vasileios Belagiannis, Ian D. Reid 0001, Gustavo Carneiro 0001 |
ICCV | 7 |
| 2023 | Learning Support and Trivial Prototypes for Interpretable Image ClassificationabstractPrototypical part network (ProtoPNet) methods have been designed to achieve interpretable classification by associating predictions with a set of training prototypes, which we refer to as trivial prototypes because they are trained to lie far from the classification boundary in the feature space. Note that it is possible to make an analogy between ProtoPNet and support vector machine (SVM) given that the classification from both methods relies on computing similarity with a set of training points (i.e., trivial prototypes in ProtoPNet, and support vectors in SVM). However, while trivial prototypes are located far from the classification boundary, support vectors are located close to this boundary, and we argue that this discrepancy with the well-established SVM theory can result in ProtoPNet models with inferior classification accuracy. In this paper, we aim to improve the classification of ProtoPNet with a new method to learn support prototypes that lie near the classification boundary in the feature space, as suggested by the SVM theory. In addition, we target the improvement of classification results with a new model, named ST-ProtoPNet, which exploits our support prototypes and the trivial prototypes to provide more effective classification. Experimental results on CUB-200-2011, Stanford Cars, and Stan-ford Dogs datasets demonstrate that ST-ProtoPNet achieves state-of-the-art classification accuracy and interpretability results. We also show that the proposed support prototypes tend to be better localised in the object of interest rather than in the background region. Chong Wang 0012, Yuyuan Liu, Yuanhong Chen, Fengbei Liu, Yu Tian 0001, Davis J. McCarthy, Helen Frazer, Gustavo Carneiro 0001 |
ICCV | 8 |
| 2023 | Multi-Head Multi-Loss Model Calibration
Adrian Galdran, Johan Verjans, Gustavo Carneiro 0001, Miguel Ángel González Ballester |
MICCAI (3) | 3 |
| 2023 | Learnable Cross-modal Knowledge Distillation for Multi-modal Learning with Missing Modality
Hu Wang 0005, Congbo Ma, Jodie Avery, Louise Hull, Gustavo Carneiro 0001 |
MICCAI (4) | 7 |
| 2023 | Model and Feature Diversity for Bayesian Neural Networks in Mutual LearningabstractBayesian Neural Networks (BNNs) offer probability distributions for model parameters, enabling uncertainty quantification in predictions. However, they often underperform compared to deterministic neural networks. Utilizing mutual learning can effectively enhance the performance of peer BNNs. In this paper, we propose a novel approach to improve BNNs performance through deep mutual learning. The proposed approaches aim to increase diversity in both network parameter distributions and feature distributions, promoting peer networks to acquire distinct features that capture different characteristics of the input, which enhances the effectiveness of mutual learning. Experimental results demonstrate significant improvements in the classification accuracy, negative log-likelihood, and expected calibration error when compared to traditional mutual learning for BNNs. Van Cuong Pham, Cuong Nguyen 0006, Trung Le 0001, Dinh Q. Phung, Gustavo Carneiro 0001, Thanh-Toan Do |
NeurIPS | 5 |
| 2023 | Knowing What to Label for Few Shot Microscopy Image Cell SegmentationabstractIn microscopy image cell segmentation, it is common to train a deep neural network on source data, containing different types of microscopy images, and then fine-tune it using a support set comprising a few randomly selected and annotated training target images. In this paper, we argue that the random selection of unlabelled training target images to be annotated and included in the support set may not enable an effective fine-tuning process, so we propose a new approach to optimise this image selection process. Our approach involves a new scoring function to find informative unlabelled target images. In particular, we propose to measure the consistency in the model predictions on target images against specific data augmentations. However, we observe that the model trained with source datasets does not reliably evaluate consistency on target images. To alleviate this problem, we propose novel self-supervised pretext tasks to compute the scores of unlabelled target images. Finally, the top few images with the least consistency scores are added to the support set for oracle (i.e., expert) annotation and later used to fine-tune the model to the target images. In our evaluations that involve the segmentation of five different types of cell images, we demonstrate promising results on several target test sets compared to the random selection approach as well as other selection approaches, such as Shannon's entropy and Monte-Carlo dropout. Youssef Dawoud, Arij Bouazizi, Katharina Ernst, Gustavo Carneiro 0001, Vasileios Belagiannis |
WACV | 4 |
| 2023 | Instance-Dependent Noisy Label Learning via Graphical ModellingabstractNoisy labels are unavoidable yet troublesome in the ecosystem of deep learning because models can easily overfit them. There are many types of label noise, such as symmetric, asymmetric and instance-dependent noise (IDN), with IDN being the only type that depends on image information. Such dependence on image information makes IDN a critical type of label noise to study, given that labelling mistakes are caused in large part by insufficient or ambiguous information about the visual classes present in images. Aiming to provide an effective technique to address IDN, we present a new graphical modelling approach called InstanceGM, that combines discriminative and generative models. The main contributions of InstanceGM are: i) the use of the continuous Bernoulli distribution to train the generative model, offering significant training advantages, and ii) the exploration of a state-of-the-art noisy-label discriminative classifier to generate clean labels from instance-dependent noisy-label samples. InstanceGM is competitive with current noisy-label learning approaches, particularly in IDN benchmarks using synthetic and real-world datasets, where our method shows better accuracy than the competitors in most experiments1. Arpit Garg, Cuong Nguyen 0006, Rafael Felix, Thanh-Toan Do, Gustavo Carneiro 0001 |
WACV | 5 |
| 2023 | Bootstrapping the Relationship Between Images and Their Clean and Noisy LabelsabstractMany state-of-the-art noisy-label learning methods rely on learning mechanisms that estimate the samples’ clean labels during training and discard their original noisy labels. However, this approach prevents the learning of the relationship between images, noisy labels and clean labels, which has been shown to be useful when dealing with instance-dependent label noise problems. Further-more, methods that do aim to learn this relationship re-quire cleanly annotated subsets of data, as well as distillation or multi-faceted models for training. In this paper, we propose a new training algorithm that relies on a simple model to learn the relationship between clean and noisy labels without the need for a cleanly labelled subset of data. Our algorithm follows a 3-stage process, namely: 1) self-supervised pre-training followed by an early-stopping training of the classifier to confidently predict clean labels for a subset of the training set; 2) use the clean set from stage (1) to bootstrap the relationship between images, noisy labels and clean labels, which we exploit for effective relabelling of the remaining training set using semi-supervised learning; and 3) supervised training of the classifier with all relabelled samples from stage (2). By learning this relationship, we achieve state-of-the-art performance in asymmetric and instance-dependent label noise problems1. Code is available at https://github.com/btsmart/bootstrapping-label-noise. Brandon Smart, Gustavo Carneiro 0001 |
WACV | 2 |
| 2023 | Self-supervised pseudo multi-class pre-training for unsupervised anomaly detection and segmentation in medical images
Yu Tian 0001, Fengbei Liu, Guansong Pang, Yuanhong Chen, Yuyuan Liu, Johan Verjans, Rajvinder Singh, Gustavo Carneiro 0001 |
Medical Image Anal. | 8 |
| 2023 | PAC-Bayes Meta-Learning With Implicit Task-Specific PosteriorsabstractWe introduce a new and rigorously-formulated PAC-Bayes meta-learning algorithm that solves few-shot learning. Our proposed method extends the PAC-Bayes framework from a single-task setting to the meta-learning multiple-task setting to upper-bound the error evaluated on any, even unseen, tasks and samples. We also propose a generative-based approach to estimate the posterior of task-specific model parameters more expressively compared to the usual assumption based on a multivariate normal distribution with a diagonal covariance matrix. We show that the models trained with our proposed meta-learning algorithm are well-calibrated and accurate, with state-of-the-art calibration errors while still being competitive on classification results on few-shot classification (mini-ImageNet and tiered-ImageNet) and regression (multi-modal task-distribution regression) benchmarks. Cuong Nguyen 0006, Thanh-Toan Do, Gustavo Carneiro 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | LongReMix: Robust learning with high confidence samples in a noisy label environment
Filipe R. Cordeiro, Ragav Sachdeva, Vasileios Belagiannis, Ian D. Reid 0001, Gustavo Carneiro 0001 |
Pattern Recognit. | 5 |
| 2023 | ScanMix: Learning from Severe Label Noise via Semantic Clustering and Semi-Supervised LearningabstractWe propose a new training algorithm, ScanMix, that explores semantic clustering and semi-supervised learning (SSL) to allow superior robustness to severe label noise and competitive robustness to non-severe label noise problems, in comparison to the state of the art (SOTA) methods. ScanMix is based on the expectation maximisation framework, where the E-step estimates the latent variable to cluster the training images based on their appearance and classification results, and the M-step optimises the SSL classification and learns effective feature representations via semantic clustering. We present a theoretical result that shows the correctness and convergence of ScanMix, and an empirical result that shows that ScanMix has SOTA results on CIFAR-10/-100 (with symmetric, asymmetric and semantic label noise), Red Mini-ImageNet (from the Controlled Noisy Web Labels), Clothing1M and WebVision. In all benchmarks with severe label noise, our results are competitive to the current SOTA. Ragav Sachdeva, Filipe R. Cordeiro, Vasileios Belagiannis, Ian D. Reid 0001, Gustavo Carneiro 0001 |
Pattern Recognit. | 5 |
| 2023 | BowelNet: Joint Semantic-Geometric Ensemble Learning for Bowel Segmentation From Both Partially and Fully Labeled CT ImagesabstractAccurate bowel segmentation is essential for diagnosis and treatment of bowel cancers. Unfortunately, segmenting the entire bowel in CT images is quite challenging due to unclear boundary, large shape, size, and appearance variations, as well as diverse filling status within the bowel. In this paper, we present a novel two-stage framework, named BowelNet, to handle the challenging task of bowel segmentation in CT images, with two stages of 1) jointly localizing all types of the bowel, and 2) finely segmenting each type of the bowel. Specifically, in the first stage, we learn a unified localization network from both partially- and fully-labeled CT images to robustly detect all types of the bowel. To better capture unclear bowel boundary and learn complex bowel shapes, in the second stage, we propose to jointly learn semantic information (i.e., bowel segmentation mask) and geometric representations (i.e., bowel boundary and bowel skeleton) for fine bowel segmentation in a multi-task learning scheme. Moreover, we further propose to learn a meta segmentation network via pseudo labels to improve segmentation accuracy. By evaluating on a large abdominal CT dataset, our proposed BowelNet method can achieve Dice scores of 0.764, 0.848, 0.835, 0.774, and 0.824 in segmenting the duodenum, jejunum-ileum, colon, sigmoid, and rectum, respectively. These results demonstrate the effectiveness of our proposed BowelNet framework in segmenting the entire bowel from CT images. Chong Wang 0012, Zhiming Cui 0001, Miaofei Han, Gustavo Carneiro 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 5 |
| 2022 | Deep One-Class Classification via Interpolated Gaussian DescriptorabstractOne-class classification (OCC) aims to learn an effective data description to enclose all normal training samples and detect anomalies based on the deviation from the data description. Current state-of-the-art OCC models learn a compact normality description by hyper-sphere minimisation, but they often suffer from overfitting the training data, especially when the training set is small or contaminated with anomalous samples. To address this issue, we introduce the interpolated Gaussian descriptor (IGD) method, a novel OCC model that learns a one-class Gaussian anomaly classifier trained with adversarially interpolated training samples. The Gaussian anomaly classifier differentiates the training samples based on their distance to the Gaussian centre and the standard deviation of these distances, offering the model a discriminability w.r.t. the given samples during training. The adversarial interpolation is enforced to consistently learn a smooth Gaussian descriptor, even when the training data is small or contaminated with anomalous samples. This enables our model to learn the data description based on the representative normal samples rather than fringe or anomalous samples, resulting in significantly improved normality description. In extensive experiments on diverse popular benchmarks, including MNIST, Fashion MNIST, CIFAR10, MVTec AD and two medical datasets, IGD achieves better detection accuracy than current state-of-the-art models. IGD also shows better robustness in problems with small or contaminated training sets. Yuanhong Chen, Yu Tian 0001, Guansong Pang, Gustavo Carneiro 0001 |
AAAI | 4 |
| 2022 | Perturbed and Strict Mean Teachers for Semi-supervised Semantic SegmentationabstractConsistency learning using input image, feature, or network perturbations has shown remarkable results in semi-supervised semantic segmentation, but this approach can be seriously affected by inaccurate predictions of unlabelled training images. There are two consequences of these inaccurate predictions: 1) the training based on the “strict” cross-entropy (CE) loss can easily overfit prediction mistakes, leading to confirmation bias; and 2) the perturbations applied to these inaccurate predictions will use potentially erroneous predictions as training signals, degrading consistency learning. In this paper, we address the prediction accuracy problem of consistency learning methods with novel extensions of the mean-teacher (MT) model, which include a new auxiliary teacher, and the replacement of MT's mean square error (MSE) by a stricter confidence-weighted cross-entropy (Conf-CE) loss. The accurate prediction by this model allows us to use a challenging combination of network, input data and feature perturbations to improve the consistency learning generalisation, where the feature perturbations consist of a new adversarial perturbation. Results on public benchmarks show that our approach achieves remarkable improvements over the previous SOTA methods in the field.11Supported by Australian Research Council through grants DP180103232 and FT190100525. Our code is available at https://github.com/yyliu01/PS-MT. Yuyuan Liu, Yu Tian 0001, Yuanhong Chen, Fengbei Liu, Vasileios Belagiannis, Gustavo Carneiro 0001 |
CVPR | 6 |
| 2022 | ACPL: Anti-curriculum Pseudo-labelling for Semi-supervised Medical Image ClassificationabstractEffective semi-supervised learning (SSL) in medical image analysis (MIA) must address two challenges: 1) work effectively on both multi-class (e.g., lesion classification) and multi-label (e.g., multiple-disease diagnosis) problems, and 2) handle imbalanced learning (because of the high variance in disease prevalence). One strategy to explore in SSL MIA is based on the pseudo labelling strategy, but it has a few shortcomings. Pseudo-labelling has in general lower accuracy than consistency learning, it is not specifically design for both multi-class and multi-label problems, and it can be challenged by imbalanced learning. In this paper, unlike traditional methods that select confident pseudo label by threshold, we propose a new SSL algorithm, called anti-curriculum pseudo-labelling (ACPL), which introduces novel techniques to select informative unlabelled samples, improving training balance and allowing the model to work for both multi-label and multi-class problems, and to estimate pseudo labels by an accurate ensemble of classifiers (improving pseudo label accuracy). We run extensive experiments to evaluate ACPL on two public medical image classification benchmarks: Chest X-Ray 14 for thorax disease multi-label classification and ISIC2018 for skin lesion multi-class classification. Our method outperforms previous SOTA SSL methods on both datasets11Supported by Australian Research Council through grants DP180103232 and FT190100525.22Code is available at https://github.com/FBLADL/ACPL. Fengbei Liu, Yu Tian 0001, Yuanhong Chen, Yuyuan Liu, Vasileios Belagiannis, Gustavo Carneiro 0001 |
CVPR | 6 |
| 2022 | Pixel-Wise Energy-Biased Abstention Learning for Anomaly Segmentation on Complex Urban Driving Scenes
Yu Tian 0001, Yuyuan Liu, Guansong Pang, Fengbei Liu, Yuanhong Chen, Gustavo Carneiro 0001 |
ECCV (39) | 6 |
| 2022 | Uncertainty-Aware Multi-modal Learning via Cross-Modal Random Network Prediction
Hu Wang 0005, Yuanhong Chen, Congbo Ma, Jodie Avery, Louise Hull, Gustavo Carneiro 0001 |
ECCV (37) | 7 |
| 2022 | Mixup-Based Deep Metric Learning Approaches for Incomplete SupervisionabstractDeep learning architectures have achieved promising results in different areas (e.g., medicine, agriculture, and security). However, using those powerful techniques in many real applications becomes challenging due to the large labeled collections required during training. Several works have pursued solutions to overcome it by proposing strategies that can learn more for less, e.g., weakly and semi-supervised learning approaches. As these approaches do not usually address memorization and sensitivity to adversarial examples, this paper presents three deep metric learning approaches combined with Mixup for incomplete-supervision scenarios. We show that some state-of-the-art approaches in metric learning might not work well in such scenarios. Moreover, the proposed approaches outperform most of them in different datasets. Luiz H. Buris, Daniel C. G. Pedronette, João Paulo Papa, Jurandy Almeida, Gustavo Carneiro 0001, Fábio Augusto Faria |
ICIP | 5 |
| 2022 | Multi-view Local Co-occurrence and Global Consistency Learning Improve Mammogram Classification Generalisation
Yuanhong Chen, Hu Wang 0005, Chong Wang 0012, Yu Tian 0001, Fengbei Liu, Yuyuan Liu, Michael Elliott, Davis J. McCarthy, Helen Frazer, Gustavo Carneiro 0001 |
MICCAI (3) | 10 |
| 2022 | Test Time Transform Prediction for Open Set Histopathological Image Recognition
Adrian Galdran, Katherine Jane Hewitt, Narmin Ghaffari Laleh, Jakob Nikolas Kather, Gustavo Carneiro 0001, Miguel Ángel González Ballester |
MICCAI (2) | 5 |
| 2022 | Censor-Aware Semi-supervised Learning for Survival Time Prediction from Medical Images
Renato Hermoza, Gabriel Maicas, Jacinto C. Nascimento, Gustavo Carneiro 0001 |
MICCAI (8) | 4 |
| 2022 | NVUM: Non-volatile Unbiased Memory for Robust Medical Image Classification
Fengbei Liu, Yuanhong Chen, Yu Tian 0001, Yuyuan Liu, Chong Wang 0012, Vasileios Belagiannis, Gustavo Carneiro 0001 |
MICCAI (3) | 7 |
| 2022 | Contrastive Transformer-Based Multiple Instance Learning for Weakly Supervised Polyp Frame Detection
Yu Tian 0001, Guansong Pang, Fengbei Liu, Yuyuan Liu, Chong Wang 0012, Yuanhong Chen, Johan Verjans, Gustavo Carneiro 0001 |
MICCAI (3) | 8 |
| 2022 | Knowledge Distillation to Ensemble Global and Interpretable Prototype-Based Mammogram Classification Models
Chong Wang 0012, Yuanhong Chen, Yuyuan Liu, Yu Tian 0001, Fengbei Liu, Davis J. McCarthy, Michael Elliott, Helen Frazer, Gustavo Carneiro 0001 |
MICCAI (3) | 9 |
| 2021 | Domain Generalisation with Domain Augmented Supervised Contrastive Learning (Student Abstract)abstractDomain generalisation (DG) methods address the problem of domain shift, when there is a mismatch between the distributions of training and target domains. Data augmentation approaches have emerged as a promising alternative for DG. However, data augmentation alone is not sufficient to achieve lower generalisation errors. This project proposes a new method that combines data augmentation and domain distance minimisation to address the problems associated with data augmentation and provide a guarantee on the learning performance, under an existing framework. Empirically, our method outperforms baseline results on DG benchmarks. Hoang-Son Le, Rini Akmeliawati, Gustavo Carneiro 0001 |
AAAI | 3 |
| 2021 | PropMix: Hard Sample Filtering and Proportional MixUp for Learning with Noisy Labels
Filipe R. Cordeiro, Vasileios Belagiannis, Ian D. Reid 0001, Gustavo Carneiro 0001 |
BMVC | 4 |
| 2021 | Weakly-supervised Video Anomaly Detection with Robust Temporal Feature Magnitude LearningabstractAnomaly detection with weakly supervised video-level labels is typically formulated as a multiple instance learning (MIL) problem, in which we aim to identify snippets containing abnormal events, with each video represented as a bag of video snippets. Although current methods show effective detection performance, their recognition of the positive instances, i.e., rare abnormal snippets in the abnormal videos, is largely biased by the dominant negative instances, especially when the abnormal events are subtle anomalies that exhibit only small differences compared with normal events. This issue is exacerbated in many methods that ignore important video temporal dependencies. To address this issue, we introduce a novel and theoretically sound method, named Robust Temporal Feature Magnitude learning (RTFM), which trains a feature magnitude learning function to effectively recognise the positive instances, substantially improving the robustness of the MIL approach to the negative instances from abnormal videos. RTFM also adapts dilated convolutions and self-attention mechanisms to capture long- and short-range temporal dependencies to learn the feature magnitude more faithfully. Extensive experiments show that the RTFM-enabled MIL model (i) outperforms several state-of-the-art methods by a large margin on four benchmark data sets (ShanghaiTech, UCF-Crime, XD-Violence and UCSD-Peds) and (ii) achieves significantly improved subtle anomaly discriminability and sample efficiency. Yu Tian 0001, Guansong Pang, Yuanhong Chen, Rajvinder Singh, Johan Verjans, Gustavo Carneiro 0001 |
ICCV | 6 |
| 2021 | Balanced-MixUp for Highly Imbalanced Medical Image Classification
Adrian Galdran, Gustavo Carneiro 0001, Miguel Ángel González Ballester |
MICCAI (5) | 2 |
| 2021 | 3D Semantic Mapping from Arthroscopy Using Out-of-Distribution Pose and Depth and In-Distribution Segmentation Training
Yaqub Jonmohamadi, Shahnewaz Ali, Fengbei Liu, Jonathan Roberts 0001, Ross Crawford, Gustavo Carneiro 0001, Ajay K. Pandey |
MICCAI (2) | 6 |
| 2021 | Constrained Contrastive Distribution Learning for Unsupervised Anomaly Detection and Localisation in Medical Images
Yu Tian 0001, Guansong Pang, Fengbei Liu, Yuanhong Chen, Seon-Ho Shin, Johan Verjans, Rajvinder Singh, Gustavo Carneiro 0001 |
MICCAI (5) | 8 |
| 2021 | Self-Supervised Lesion Change Detection and Localisation in Longitudinal Multiple Sclerosis Brain Imaging
Minh-Son To, Ian G. Sarno, Chee Chong, Mark Jenkinson, Gustavo Carneiro 0001 |
MICCAI (7) | 5 |
| 2021 | Probabilistic task modelling for meta-learningabstractWe propose probabilistic task modelling – a generative probabilistic model for collections of tasks used in meta-learning. The proposed model combines variational auto-encoding and latent Dirichlet allocation to model each task as a mixture of Gaussian distribution in an embedding space. Such modelling provides an explicit representation of a task through its task-theme mixture. We present an efficient approximation inference technique based on variational inference method for empirical Bayes parameter estimation. We perform empirical evaluations to validate the task uncertainty and task distance produced by the proposed method through correlation diagrams of the prediction accuracy on testing tasks. We also carry out experiments of task selection in meta-learning to demonstrate how the task relatedness inferred from the proposed model help to facilitate meta-learning algorithms. Cuong Nguyen 0006, Thanh-Toan Do, Gustavo Carneiro 0001 |
UAI | 3 |
| 2021 | EvidentialMix: Learning with Combined Open-set and Closed-set Noisy LabelsabstractThe efficacy of deep learning depends on large-scale data sets that have been carefully curated with reliable data acquisition and annotation processes. However, acquiring such large-scale data sets with precise annotations is very expensive and time-consuming, and the cheap alternatives often yield data sets that have noisy labels. The field has addressed this problem by focusing on training models under two types of label noise: 1) closed-set noise, where some training samples are incorrectly annotated to a training label other than their known true class; and 2) open-set noise, where the training set includes samples that possess a true class that is (strictly) not contained in the set of known training labels. In this work, we study a new variant of the noisy label problem that combines the open-set and closed-set noisy labels, and introduce a benchmark evaluation to assess the performance of training algorithms under this setup. We argue that such problem is more general and better reflects the noisy label scenarios in practice. Furthermore, we propose a novel algorithm, called EvidentialMix, that addresses this problem and compare its performance with the state-of-the-art methods for both closed-set and open-set noise on the proposed benchmark. Our results show that our method produces superior classification results and better feature representations than previous state-of-the-art methods. The code is available at https:/github.com/ragavsachdeva/EvidentialMix. Ragav Sachdeva, Filipe R. Cordeiro, Vasileios Belagiannis, Ian D. Reid 0001, Gustavo Carneiro 0001 |
WACV | 5 |
| 2021 | Artificial intelligence for the diagnosis of lymph node metastases in patients with abdominopelvic malignancy: A systematic review and meta-analysis
Sergei Bedrikovetski, Nagendra N. Dudi-Venkata, Gabriel Maicas, Hidde M. Kroon, Warren Seow, Gustavo Carneiro 0001, James W. Moore, Tarik Sammour |
Artif. Intell. Medicine | 6 |
| 2021 | LOW: Training deep neural networks by learning optimal sample weights
Carlos Santiago, Catarina Barata, Michele Sasdelli, Gustavo Carneiro 0001, Jacinto C. Nascimento |
Pattern Recognit. | 4 |
| 2020 | Augmentation Network for Generalised Zero-Shot Learning
Rafael Felix, Michele Sasdelli, Ian D. Reid 0001, Gustavo Carneiro 0001 |
ACCV (4) | 4 |
| 2020 | Deep Metric Learning Meets Deep Clustering: An Novel Unsupervised Approach for Feature Embedding
Binh X. Nguyen, Binh D. Nguyen, Gustavo Carneiro 0001, Erman Tjiputra, Quang D. Tran, Thanh-Toan Do |
BMVC | 3 |
| 2020 | Self-Supervised Monocular Trained Depth Estimation Using Self-Attention and Discrete Disparity VolumeabstractMonocular depth estimation has become one of the most studied applications in computer vision, where the most accurate approaches are based on fully supervised learning models. However, the acquisition of accurate and large ground truth data sets to model these fully supervised methods is a major challenge for the further development of the area. Self-supervised methods trained with monocular videos constitute one the most promising approaches to mitigate the challenge mentioned above due to the wide-spread availability of training data. Consequently, they have been intensively studied, where the main ideas explored consist of different types of model architectures, loss functions, and occlusion masks to address non-rigid motion. In this paper, we propose two new ideas to improve self-supervised monocular trained depth estimation: 1) self-attention, and 2) discrete disparity prediction. Compared with the usual localised convolution operation, self-attention can explore a more general contextual information that allows the inference of similar disparity values at non-contiguous regions of the image. Discrete disparity prediction has been shown by fully supervised methods to provide a more robust and sharper depth estimation than the more common continuous disparity prediction, besides enabling the estimation of depth uncertainty. We show that the extension of the state-of-the-art self-supervised monocular trained depth estimator Monodepth2 with these two ideas allows us to design a model that produces the best results in the field in KITTI 2015 and Make3D, closing the gap with respect self-supervised stereo training and fully supervised approaches. Adrian Johnston, Gustavo Carneiro 0001 |
CVPR | 2 |
| 2020 | Creating Classifier Ensembles through Meta-heuristic Algorithms for Aerial Scene ClassificationabstractConvolutional Neural Networks (CNN) have been being widely employed to solve the challenging remote sensing task of aerial scene classification. Nevertheless, it is not straightforward to find single CNN models that can solve all aerial scene classification tasks, allowing the development of a better alternative, which is to fuse CNN-based classifiers into an ensemble. However, an appropriate choice of the classifiers that will belong to the ensemble is a critical factor, as it is unfeasible to employ all the possible classifiers in the literature. Therefore, this work proposes a novel framework based on meta-heuristic optimization for creating optimized ensembles in the context of aerial scene classification. The experimental results were performed across nine meta-heuristic algorithms and three aerial scene literature datasets, being compared in terms of effectiveness (accuracy), efficiency (execution time), and behavioral performance in different scenarios. Our results suggest that the Univariate Marginal Distribution Algorithm shows more effective and efficient results than other commonly used meta-heuristic algorithms, such as Genetic Programming and Particle Swarm Optimization. Álvaro R. Ferreira, Gustavo H. Rosa, João Paulo Papa, Gustavo Carneiro 0001, Fábio Augusto Faria |
ICPR | 4 |
| 2020 | Region Proposals for Saliency Map Refinement for Weakly-Supervised Disease Localisation and Classification
Renato Hermoza, Gabriel Maicas, Jacinto C. Nascimento, Gustavo Carneiro 0001 |
MICCAI (6) | 4 |
| 2020 | Self-supervised Depth Estimation to Regularise Semantic Segmentation in Knee Arthroscopy
Fengbei Liu, Yaqub Jonmohamadi, Gabriel Maicas, Ajay K. Pandey, Gustavo Carneiro 0001 |
MICCAI (1) | 5 |
| 2020 | Few-Shot Anomaly Detection for Polyp Frames from Colonoscopy
Yu Tian 0001, Gabriel Maicas, Leonardo Z. C. T. Pu, Rajvinder Singh, Johan Verjans, Gustavo Carneiro 0001 |
MICCAI (6) | 6 |
| 2020 | Probabilistic Object Detection: Definition and EvaluationabstractWe introduce Probabilistic Object Detection, the task of detecting objects in images and accurately quantifying the spatial and semantic uncertainties of the detections. Given the lack of methods capable of assessing such probabilistic object detections, we present the new Probability-based Detection Quality measure (PDQ). Unlike AP-based measures, PDQ has no arbitrary thresholds and rewards spatial and label quality, and foreground/background separation quality while explicitly penalising false positive and false negative detections. We contrast PDQ with existing mAP and moLRP measures by evaluating state-of-the-art detectors and a Bayesian object detector based on Monte Carlo Dropout. Our experiments indicate that conventional object detectors tend to be spatially overconfident and thus perform poorly on the task of probabilistic object detection. Our paper aims to encourage the development of new object detection approaches that provide detections with accurately estimated spatial and label uncertainties and are of critical importance for deployment on robots and embodied AI systems in the real world. David Hall 0003, Feras Dayoub, John Skinner, Dimity Miller, Peter I. Corke, Gustavo Carneiro 0001, Anelia Angelova, Niko Sünderhauf |
WACV | 7 |
| 2020 | Uncertainty in Model-Agnostic Meta-Learning using Variational InferenceabstractWe introduce a new, rigorously-formulated Bayesian meta-learning algorithm that learns a probability distribution of model parameter prior for few-shot learning. The proposed algorithm employs a gradient-based variational inference to infer the posterior of model parameters for a new task. Our algorithm can be applied to any model architecture and can be implemented in various machine learning paradigms, including regression and classification. We show that the models trained with our proposed meta-learning algorithm are well calibrated and accurate, with state-of-the-art calibration and classification results on three few-shot classification benchmarks (Om- niglot, mini-ImageNet and tiered-ImageNet), and competitive results in a multi-modal task-distribution regression. Cuong Nguyen 0006, Thanh-Toan Do, Gustavo Carneiro 0001 |
WACV | 3 |
| 2020 | Special Issue on Deep Learning for Robotic Vision
Anelia Angelova, Gustavo Carneiro 0001, Niko Sünderhauf, Jürgen Leitner |
Int. J. Comput. Vis. | 2 |
| 2020 | Deep learning uncertainty and confidence calibration for the five-class polyp classification from colonoscopy
Gustavo Carneiro 0001, Leonardo Z. C. T. Pu, Rajvinder Singh, Alastair D. Burt |
Medical Image Anal. | 1 |
| 2020 | Siam-U-Net: encoder-decoder siamese network for knee cartilage tracking in ultrasound images
Matteo Dunnhofer, Maria Antico, Fumio Sasazawa, Yu Takeda, Saskia Camps, Niki Martinel, Christian Micheloni, Gustavo Carneiro 0001, Davide Fontanarosa |
Medical Image Anal. | 8 |
| 2020 | Approximate Fisher Information Matrix to Characterize the Training of Deep Neural NetworksabstractIn this paper, we introduce a novel methodology for characterizing the performance of deep learning networks (ResNets and DenseNet) with respect to training convergence and generalization as a function of mini-batch size and learning rate for image classification. This methodology is based on novel measurements derived from the eigenvalues of the approximate Fisher information matrix, which can be efficiently computed even for high capacity deep models. Our proposed measurements can help practitioners to monitor and control the training process (by actively tuning the mini-batch size and learning rate) to allow for good training convergence and generalization. Furthermore, the proposed measurements also allow us to show that it is possible to optimize the training process with a new dynamic sampling training approach that continuously and automatically change the mini-batch size and learning rate during the training process. Finally, we show that the proposed dynamic sampling training approach has a faster training time and a competitive classification accuracy compared to the current state of the art. Zhibin Liao, Tom Drummond, Ian D. Reid 0001, Gustavo Carneiro 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2020 | One Shot Segmentation: Unifying Rigid Detection and Non-Rigid Segmentation Using Elastic RegularizationabstractThis paper proposes a novel approach for the non-rigid segmentation of deformable objects in image sequences, which is based on one-shot segmentation that unifies rigid detection and non-rigid segmentation using elastic regularization. The domain of application is the segmentation of a visual object that temporally undergoes a rigid transformation (e.g., affine transformation) and a non-rigid transformation (i.e., contour deformation). The majority of segmentation approaches to solve this problem are generally based on two steps that run in sequence: a rigid detection, followed by a non-rigid segmentation. In this paper, we propose a new approach, where both the rigid and non-rigid segmentation are performed in a single shot using a sparse low-dimensional manifold that represents the visual object deformations. Given the multi-modality of these deformations, the manifold partitions the training data into several patches, where each patch provides a segmentation proposal during the inference process. These multiple segmentation proposals are merged using the classification results produced by deep belief networks (DBN) that compute the confidence on each segmentation proposal. Thus, an ensemble of DBN classifiers is used for estimating the final segmentation. Compared to current methods proposed in the field, our proposed approach is advantageous in four aspects: (i) it is a unified framework to produce rigid and non-rigid segmentations; (ii) it uses an ensemble classification process, which can help the segmentation robustness; (iii) it provides a significant reduction in terms of the number of dimensions of the rigid and non-rigid segmentations search spaces, compared to current approaches that divide these two problems; and (iv) this lower dimensionality of the search space can also reduce the need for large annotated training sets to be used for estimating the DBN models. Experiments on the problem of left ventricle endocardial segmentation from ultrasound images, and lip segmentation from frontal facial images using the extended Cohn-Kanade (CK+) database, demonstrate the potential of the methodology through qualitative and quantitative evaluations, and the ability to reduce the search and training complexities without a significant impact on the segmentation accuracy. Jacinto C. Nascimento, Gustavo Carneiro 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | A Theoretically Sound Upper Bound on the Triplet Loss for Improving the Efficiency of Deep Distance Metric LearningabstractWe propose a method that substantially improves the efficiency of deep distance metric learning based on the optimization of the triplet loss function. One epoch of such training process based on a na¨ıve optimization of the triplet loss function has a run-time complexity O(N^3), where N is the number of training samples. Such optimization scales poorly, and the most common approach proposed to address this high complexity issue is based on sub-sampling the set of triplets needed for the training process. Another approach explored in the field relies on an ad-hoc linearization (in terms of N) of the triplet loss that introduces class centroids, which must be optimized using the whole training set for each mini-batch – this means that a na¨ıve implementation of this approach has run-time complexity O(N^2). This complexity issue is usually mitigated with poor, but computationally cheap, approximate centroid optimization methods. In this paper, we first propose a solid theory on the linearization of the triplet loss with the use of class centroids, where the main conclusion is that our new linear loss represents a tight upper-bound to the triplet loss. Furthermore, based on the theory above, we propose a training algorithm that no longer requires the centroid optimization step, which means that our approach is the first in the field with a guaranteed linear run-time complexity. We show that the training of deep distance metric learning methods using the proposed upper-bound is substantially faster than triplet-based methods, while producing competitive retrieval accuracy results on benchmark datasets (CUB-200-2011 and CAR196). Thanh-Toan Do, Toan Tran 0002, Ian D. Reid 0001, Tuan Hoang, Gustavo Carneiro 0001 |
CVPR | 6 |
| 2019 | Bayesian Generative Active Deep LearningabstractDeep learning models have demonstrated outstanding performance in several problems, but their training process tends to require immense amounts of computational and human resources for training and labeling, constraining the types of problems that can be tackled. Therefore, the design of effective training methods that require small labeled training sets is an important research direction that will allow a more effective use of resources. Among current approaches designed to address this issue, two are particularly interesting: data augmentation and active learning. Data augmentation achieves this goal by artificially generating new training points, while active learning relies on the selection of the “most informative” subset of unlabeled training samples to be labelled by an oracle. Although successful in practice, data augmentation can waste computational resources because it indiscriminately generates samples that are not guaranteed to be informative, and active learning selects a small subset of informative samples (from a large un-annotated set) that may be insufficient for the training process. In this paper, we propose a Bayesian generative active deep learning approach that combines active learning with data augmentation – we provide theoretical and empirical evidence (MNIST, CIFAR-$\{10,100\}$, and SVHN) that our approach has more efficient training and better classification results than data augmentation and active learning. Toan Tran 0002, Thanh-Toan Do, Ian D. Reid 0001, Gustavo Carneiro 0001 |
ICML | 4 |
| 2019 | Pre and post-hoc diagnosis and interpretation of malignancy from breast DCE-MRI
Gabriel Maicas, Andrew P. Bradley, Jacinto C. Nascimento, Ian D. Reid 0001, Gustavo Carneiro 0001 |
Medical Image Anal. | 5 |
| 2018 | Multi-modal Cycle-Consistent Generalized Zero-Shot Learning
Rafael Felix, Ian D. Reid 0001, Gustavo Carneiro 0001 |
ECCV (6) | 4 |
| 2018 | Bayesian Semantic Instance Segmentation in Open Set World
Trung Pham, Thanh-Toan Do, Gustavo Carneiro 0001, Ian D. Reid 0001 |
ECCV (10) | 4 |
| 2018 | Training Medical Image Analysis Systems like Radiologists
Gabriel Maicas, Andrew P. Bradley, Jacinto C. Nascimento, Ian D. Reid 0001, Gustavo Carneiro 0001 |
MICCAI (1) | 5 |
| 2017 | Smart Mining for Deep Metric LearningabstractTo solve deep metric learning problems and producing feature embeddings, current methodologies will commonly use a triplet model to minimise the relative distance between samples from the same class and maximise the relative distance between samples from different classes. Though successful, the training convergence of this triplet model can be compromised by the fact that the vast majority of the training samples will produce gradients with magnitudes that are close to zero. This issue has motivated the development of methods that explore the global structure of the embedding and other methods that explore hard negative/positive mining. The effectiveness of such mining methods is often associated with intractable computational requirements. In this paper, we propose a novel deep metric learning method that combines the triplet model and the global structure of the embedding space. We rely on a smart mining procedure that produces effective training samples for a low computational cost. In addition, we propose an adaptive controller that automatically adjusts the smart mining hyper-parameters and speeds up the convergence of the training process. We show empirically that our proposed method allows for fast and more accurate training of triplet ConvNets than other competing mining methods. Additionally, we show that our method achieves new state-of-the-art embedding results for CUB-200-2011 and Cars196 datasets. Ben Harwood, Gustavo Carneiro 0001, Ian D. Reid 0001, Tom Drummond |
ICCV | 3 |
| 2017 | Mass segmentation in mammograms: A cross-sensor comparison of deep and tailored featuresabstractThrough the years, several CAD systems have been developed to help radiologists in the hard task of detecting signs of cancer in mammograms. In these CAD systems, mass segmentation plays a central role in the decision process. In the literature, mass segmentation has been typically evaluated in a intra-sensor scenario, where the methodology is designed and evaluated in similar data. However, in practice, acquisition systems and PACS from multiple vendors abound and current works fails to take into account the differences in mammogram data in the performance evaluation. In this work it is argued that a comprehensive assessment of the mass segmentation methods requires the design and evaluation in datasets with different properties. To provide a more realistic evaluation, this work proposes: a) improvements to a state of the art method based on tailored features and a graph model; b) a head-to-head comparison of the improved model with recently proposed methodologies based in deep learning and structured prediction on four reference databases, performing a cross-sensor evaluation. The results obtained support the assertion that the evaluation methods from the literature are optimistically biased when evaluated on data gathered from exactly the same sensor and/or acquisition protocol. Jaime S. Cardoso 0001, Neeraj Dhungel, Gustavo Carneiro 0001, Andrew P. Bradley |
ICIP | 4 |
| 2017 | Deep Reinforcement Learning for Active Breast Lesion Detection from DCE-MRI
Gabriel Maicas, Gustavo Carneiro 0001, Andrew P. Bradley, Jacinto C. Nascimento, Ian D. Reid 0001 |
MICCAI (3) | 2 |
| 2017 | A Bayesian Data Augmentation Approach for Learning Deep ModelsabstractData augmentation is an essential part of the training process applied to deep learning models. The motivation is that a robust training process for deep learning models depends on large annotated datasets, which are expensive to be acquired, stored and processed. Therefore a reasonable alternative is to be able to automatically generate new annotated training samples using a process known as data augmentation. The dominant data augmentation approach in the field assumes that new training samples can be obtained via random geometric or appearance transformations applied to annotated training samples, but this is a strong assumption because it is unclear if this is a reliable generative model for producing new training samples. In this paper, we provide a novel Bayesian formulation to data augmentation, where new annotated training points are treated as missing variables and generated based on the distribution learned from the training set. For learning, we introduce a theoretically sound algorithm --- generalised Monte Carlo expectation maximisation, and demonstrate one possible implementation via an extension of the Generative Adversarial Network (GAN). Classification results on MNIST, CIFAR-10 and CIFAR-100 show the better performance of our proposed method compared to the current dominant data augmentation approach mentioned above --- the results also show that our approach produces better classification results than similar GAN models. Toan Tran 0002, Trung Pham, Gustavo Carneiro 0001, Lyle John Palmer, Ian D. Reid 0001 |
NIPS | 3 |
| 2017 | A deep learning approach for the analysis of masses in mammograms with minimal user intervention
Neeraj Dhungel, Gustavo Carneiro 0001, Andrew P. Bradley |
Medical Image Anal. | 2 |
| 2017 | Combining deep learning and level set for the automated segmentation of the left ventricle of the heart from cardiac cine magnetic resonance
Tuan Anh Ngo, Gustavo Carneiro 0001 |
Medical Image Anal. | 3 |
| 2017 | A deep convolutional neural network module that promotes competition of multiple-size filters
Zhibin Liao, Gustavo Carneiro 0001 |
Pattern Recognit. | 2 |
| 2017 | Improving the performance of pedestrian detectors using convolutional learning
David Ribeiro, Jacinto C. Nascimento, Alexandre Bernardino, Gustavo Carneiro 0001 |
Pattern Recognit. | 4 |
| 2017 | Deep Learning on Sparse Manifolds for Faster Object SegmentationabstractWe propose a new combination of deep belief networks and sparse manifold learning strategies for the 2D segmentation of non-rigid visual objects. With this novel combination, we aim to reduce the training and inference complexities while maintaining the accuracy of machine learning-based non-rigid segmentation methodologies. Typical non-rigid object segmentation methodologies divide the problem into a rigid detection followed by a non-rigid segmentation, where the low dimensionality of the rigid detection allows for a robust training (i.e., a training that does not require a vast amount of annotated images to estimate robust appearance and shape models) and a fast search process during inference. Therefore, it is desirable that the dimensionality of this rigid transformation space is as small as possible in order to enhance the advantages brought by the aforementioned division of the problem. In this paper, we propose the use of sparse manifolds to reduce the dimensionality of the rigid detection space. Furthermore, we propose the use of deep belief networks to allow for a training process that can produce robust appearance models without the need of large annotated training sets. We test our approach in the segmentation of the left ventricle of the heart from ultrasound images and lips from frontal face images. Our experiments show that the use of sparse manifolds and deep belief networks for the rigid detection stage leads to segmentation results that are as accurate as the current state of the art, but with lower search complexity and training processes that require a small amount of annotated training data. Jacinto C. Nascimento, Gustavo Carneiro 0001 |
IEEE Trans. Image Process. | 2 |
| 2017 | Evaluation of Three Algorithms for the Segmentation of Overlapping Cervical CellsabstractIn this paper, we introduce and evaluate the systems submitted to the first Overlapping Cervical Cytology Image Segmentation Challenge, held in conjunction with the IEEE International Symposium on Biomedical Imaging 2014. This challenge was organized to encourage the development and benchmarking of techniques capable of segmenting individual cells from overlapping cellular clumps in cervical cytology images, which is a prerequisite for the development of the next generation of computer-aided diagnosis systems for cervical cancer. In particular, these automated systems must detect and accurately segment both the nucleus and cytoplasm of each cell, even when they are clumped together and, hence, partially occluded. However, this is an unsolved problem due to the poor contrast of cytoplasm boundaries, the large variation in size and shape of cells, and the presence of debris and the large degree of cellular overlap. The challenge initially utilized a database of 16 high-resolution ( ×40 magnification) images of complex cellular fields of view, in which the isolated real cells were used to construct a database of 945 cervical cytology images synthesized with a varying number of cells and degree of overlap, in order to provide full access of the segmentation ground truth. These synthetic images were used to provide a reliable and comprehensive framework for quantitative evaluation on this segmentation problem. Results from the submitted methods demonstrate that all the methods are effective in the segmentation of clumps containing at most three cells, with overlap coefficients up to 0.3. This highlights the intrinsic difficulty of this challenge and provides motivation for significant future improvement. Gustavo Carneiro 0001, Andrew P. Bradley, Daniela Ushizima, Masoud S. Nosrati, Andrea G. C. Bianchi, Cláudia M. Carneiro, Ghassan Hamarneh |
IEEE J. Biomed. Health Informatics | 2 |
| 2017 | Automated Analysis of Unregistered Multi-View Mammograms With Deep LearningabstractWe describe an automated methodology for the analysis of unregistered cranio-caudal (CC) and medio-lateral oblique (MLO) mammography views in order to estimate the patient's risk of developing breast cancer. The main innovation behind this methodology lies in the use of deep learning models for the problem of jointly classifying unregistered mammogram views and respective segmentation maps of breast lesions (i.e., masses and micro-calcifications). This is a holistic methodology that can classify a whole mammographic exam, containing the CC and MLO views and the segmentation maps, as opposed to the classification of individual lesions, which is the dominant approach in the field. We also demonstrate that the proposed system is capable of using the segmentation maps generated by automated mass and micro-calcification detection systems, and still producing accurate results. The semi-automated approach (using manually defined mass and micro-calcification segmentation maps) is tested on two publicly available data sets (INbreast and DDSM), and results show that the volume under ROC surface (VUS) for a 3-class problem (normal tissue, benign, and malignant) is over 0.9, the area under ROC curve (AUC) for the 2-class "benign versus malignant" problem is over 0.9, and for the 2-class breast screening problem (malignancy versus normal/benign) is also over 0.9. For the fully automated approach, the VUS results on INbreast is over 0.7, and the AUC for the 2-class "benign versus malignant" problem is over 0.78, and the AUC for the 2-class breast screening is 0.86. Gustavo Carneiro 0001, Jacinto C. Nascimento, Andrew P. Bradley |
IEEE Trans. Medical Imaging | 1 |
| 2017 | Automatic Quantification of Tumour Hypoxia From Multi-Modal Microscopy Images Using Weakly-Supervised Learning MethodsabstractIn recently published clinical trial results, hypoxia-modified therapies have shown to provide more positive outcomes to cancer patients, compared with standard cancer treatments. The development and validation of these hypoxia-modified therapies depend on an effective way of measuring tumor hypoxia, but a standardized measurement is currently unavailable in clinical practice. Different types of manual measurements have been proposed in clinical research, but in this paper we focus on a recently published approach that quantifies the number and proportion of hypoxic regions using high resolution (immuno-)fluorescence (IF) and hematoxylin and eosin (HE) stained images of a histological specimen of a tumor. We introduce new machine learning-based methodologies to automate this measurement, where the main challenge is the fact that the clinical annotations available for training the proposed methodologies consist of the total number of normoxic, chronically hypoxic, and acutely hypoxic regions without any indication of their location in the image. Therefore, this represents a weakly-supervised structured output classification problem, where training is based on a high-order loss function formed by the norm of the difference between the manual and estimated annotations mentioned above. We propose four methodologies to solve this problem: 1) a naive method that uses a majority classifier applied on the nodes of a fixed grid placed over the input images; 2) a baseline method based on a structured output learning formulation that relies on a fixed grid placed over the input images; 3) an extension to this baseline based on a latent structured output learning formulation that uses a graph that is flexible in terms of the amount and positions of nodes; and 4) a pixel-wise labeling based on a fully-convolutional neural network. Using a data set of 89 weakly annotated pairs of IF and HE images from eight tumors, we show that the quantitative results of methods (3) and (4) above are equally competitive and superior to the naive (1) and baseline (2) methods. All proposed methodologies show high correlation values with respect to the clinical annotations. Gustavo Carneiro 0001, Tingying Peng, Christine Bayer, Nassir Navab |
IEEE Trans. Medical Imaging | 1 |
| 2016 | Learning Local Image Descriptors with Deep Siamese and Triplet Convolutional Networks by Minimizing Global Loss FunctionsabstractRecent innovations in training deep convolutional neural network (ConvNet) models have motivated the design of new methods to automatically learn local image descriptors. The latest deep ConvNets proposed for this task consist of a siamese network that is trained by penalising misclassification of pairs of local image patches. Current results from machine learning show that replacing this siamese by a triplet network can improve the classification accuracy in several problems, but this has yet to be demonstrated for local image descriptor learning. Moreover, current siamese and triplet networks have been trained with stochastic gradient descent that computes the gradient from individual pairs or triplets of local image patches, which can make them prone to overfitting. In this paper, we first propose the use of triplet networks for the problem of local image descriptor learning. Furthermore, we also propose the use of a global loss that minimises the overall classification error in the training set, which can improve the generalisation capability of the model. Using the UBC benchmark dataset for comparing local image descriptors, we show that the triplet network produces a more accurate embedding than the siamese network in terms of the UBC dataset errors. Moreover, we also demonstrate that a combination of the triplet and global losses produces the best embedding in the field, using this triplet network. Finally, we also show that the use of the central-surround siamese network trained with the global loss produces the best result of the field on the UBC dataset. Gustavo Carneiro 0001, Ian D. Reid 0001 |
CVPR | 2 |
| 2016 | Unsupervised CNN for Single View Depth Estimation: Geometry to the Rescue
Ravi Garg, Gustavo Carneiro 0001, Ian D. Reid 0001 |
ECCV (8) | 3 |
| 2016 | CRISTAL: Adapting Workplace Training to the Real World Context with an Intelligent Simulator for Radiology Trainees
Hope Lee, Amali Weerasinghe, Jayden Barnes, Luke Oakden-Rayner, William Gale, Gustavo Carneiro 0001 |
ITS | 6 |
| 2016 | The Automated Learning of Deep Features for Breast Mass Classification from Mammograms
Neeraj Dhungel, Gustavo Carneiro 0001, Andrew P. Bradley |
MICCAI (2) | 2 |
| 2016 | On the importance of normalisation layers in deep learning with piecewise linear activation unitsabstractDeep feedforward neural networks with piecewise linear activations are currently producing the state-of-the-art results in several public datasets (e.g., CIFAR-10, CIFAR-100, MNIST, and SVHN). The combination of deep learning models and piecewise linear activation functions allows for the estimation of exponentially complex functions with the use of a large number of subnetworks specialized in the classification of similar input examples. During the training process, these subnetworks avoid overfitting with an implicit regularization scheme based on the fact that they must share their parameters with other subnetworks. Using this framework, we have made an empirical observation that can improve even more the performance of such models. We notice that these models assume a balanced initial distribution of data points with respect to the domain of the piecewise linear activation function. If that assumption is violated, then the piecewise linear activation units can degenerate into purely linear activation units, which can result in a significant reduction of their capacity to learn complex functions. Furthermore, as the number of model layers increases, this unbalanced initial distribution makes the model ill-conditioned. Therefore, we propose the introduction of batch normalisation units into deep feedforward neural networks with piecewise linear activations, which drives a more balanced use of these activation units, where each region of the activation function is trained with a relatively large proportion of training samples. Also, this batch normalisation promotes the pre-conditioning of very deep learning models. We show that by introducing maxout and batch normalisation units to the network in network model results in a model that produces classification results that are better than or comparable to the current state of the art in CIFAR-10, CIFAR-100, MNIST, and SVHN datasets. Zhibin Liao, Gustavo Carneiro 0001 |
WACV | 2 |
| 2015 | Robust Optimization for Deep RegressionabstractConvolutional Neural Networks (ConvNets) have successfully contributed to improve the accuracy of regression-based methods for computer vision tasks such as human pose estimation, landmark localization, and object detection. The network optimization has been usually performed with L2 loss and without considering the impact of outliers on the training process, where an outlier in this context is defined by a sample estimation that lies at an abnormal distance from the other training sample estimations in the objective space. In this work, we propose a regression model with ConvNets that achieves robustness to such outliers by minimizing Tukey's biweight function, an M-estimator robust to outliers, as the loss function for the ConvNet. In addition to the robust loss, we introduce a coarse-to-fine model, which processes input images of progressively higher resolutions for improving the accuracy of the regressed values. In our experiments, we demonstrate faster convergence and better generalization of our robust loss function for the tasks of human pose estimation and age estimation from face images. We also show that the combination of the robust loss function with the coarse-to-fine model produces comparable or better results than current state-of-the-art approaches in four publicly available human pose estimation datasets. Vasileios Belagiannis, Christian Rupprecht 0001, Gustavo Carneiro 0001, Nassir Navab |
ICCV | 3 |
| 2015 | Weakly-Supervised Structured Output Learning with Flexible and Latent Graphs Using High-Order Loss FunctionsabstractWe introduce two new structured output models that use a latent graph, which is flexible in terms of the number of nodes and structure, where the training process minimises a high-order loss function using a weakly annotated training set. These models are developed in the context of microscopy imaging of malignant tumours, where the estimation of the number and proportion of classes of microcirculatory supply units (MCSU) is important in the assessment of the efficacy of common cancer treatments (an MCSU is a region of the tumour tissue supplied by a microvessel). The proposed methodologies take as input multimodal microscopy images of a tumour, and estimate the number and proportion of MCSU classes. This estimation is facilitated by the use of an underlying latent graph (not present in the manual annotations), where each MCSU is represented by a node in this graph, labelled with the MCSU class and image location. The training process uses the manual weak annotations available, consisting of the number of MCSU classes per training image, where the training objective is the minimisation of a high-order loss function based on the norm of the error between the manual and estimated annotations. One of the models proposed is based on a new flexible latent structure support vector machine (FLSSVM) and the other is based on a deep convolutional neural network (DCNN) model. Using a dataset of 89 weakly annotated pairs of multimodal images from eight tumours, we show that the quantitative results from DCNN are superior, but the qualitative results from FLSSVM are better and both display high correlation values regarding the number and proportion of MCSU classes compared to the manual annotations. Gustavo Carneiro 0001, Tingying Peng, Christine Bayer, Nassir Navab |
ICCV | 1 |
| 2015 | Automatic detection of necrosis, normoxia and hypoxia in tumors from multimodal cytological imagesabstractThe efficacy of cancer treatments (e.g., radiotherapy, chemotherapy, etc.) has been observed to critically depend on the proportion of hypoxic regions (i.e., a region deprived of adequate oxygen supply) in tumor tissue, so it is important to estimate this proportion from histological samples. Medical imaging data can be used to classify tumor tissue regions into necrotic or vital and then the vital tissue into normoxia (i.e., a region receiving a normal level of oxygen), chronic or acute hypoxia. Currently, this classification is a lengthy manual process performed using (immuno-)fluorescence (IF) and hematoxylin and eosin (HE) stained images of a histological specimen, which requires an expertise that is not widespread in clinical practice. In this paper, we propose a fully automated way to detect and classify tumor tissue regions into necrosis, normoxia, chronic hypoxia and acute hypoxia using IF and HE images from the same histological specimen. Instead of relying on any single classification methodology, we propose a principled combination of the following current state-of-the-art classifiers in the field: Adaboost, support vector machine, random forest and convolutional neural networks. Results show that on average we can successfully detect and classify more than 87% of the tumor tissue regions correctly. This automated system for estimating the proportion of chronic and acute hypoxia could provide clinicians with valuable information on assessing the efficacy of cancer treatments. Gustavo Carneiro 0001, Tingying Peng, Christine Bayer, Nassir Navab |
ICIP | 1 |
| 2015 | Deep structured learning for mass segmentation from mammogramsabstractIn this paper, we present a novel method for the segmentation of breast masses from mammograms exploring structured and deep learning. Specifically, using structured support vector machine (SSVM), we formulate a model that combines different types of potential functions, including one that classifies image regions using deep learning. Our main goal with this work is to show the accuracy and efficiency improvements that these relatively new techniques can provide for the segmentation of breast masses from mammograms. We also propose an easily reproducible quantitative analysis to assess the performance of breast mass segmentation methodologies based on widely accepted accuracy and running time measurements on public datasets, which will facilitate further comparisons for this segmentation problem. In particular, we use two publicly available datasets (DDSM-BCRP and INbreast) and propose the computation of the running time taken for the methodology to produce a mass segmentation given an input image and the use of the Dice index to quantitatively measure the segmentation accuracy. For both databases, we show that our proposed methodology produces competitive results in terms of accuracy and running time. Neeraj Dhungel, Gustavo Carneiro 0001, Andrew P. Bradley |
ICIP | 2 |
| 2015 | The use of deep learning features in a hierarchical classifier learned with the minimization of a non-greedy loss function that delays gratificationabstractRecently, we have observed the traditional feature representations are being rapidly replaced by the deep learning representations, which produce significantly more accurate classification results when used together with the linear classifiers. However, it is widely known that non-linear classifiers can generally provide more accurate classification but at a higher computational cost involved in their training and testing procedures. In this paper, we propose a new efficient and accurate non-linear hierarchical classification method that uses the aforementioned deep learning representations. In essence, our classifier is based on a binary tree, where each node is represented by a linear classifier trained using a loss function that minimizes the classification error in a non-greedy way, in addition to postponing hard classification problems to further down the tree. In comparison with linear classifiers, our training process increases only marginally the training and testing time complexities, while showing competitive classification accuracy results. In addition, our method is shown to generalize better than shallow non-linear classifiers. Empirical validation shows that the proposed classifier produces more accurate classification results when compared to several linear and non-linear classifiers on Pascal VOC07 database. Zhibin Liao, Gustavo Carneiro 0001 |
ICIP | 2 |
| 2015 | Towards reduction of the training and search running time complexities for non-rigid object segmentationabstractThe problem of non-rigid object segmentation is formulated in a two-stage approach in Machine Learning based methodologies. In the first stage, the automatic initialization problem is solved by the estimation of a rigid shape of the object. In the second stage, the non-rigid segmentation is performed. The rational behind this strategy, is that the rigid detection can be performed at lower dimensional space than the original contour space. In this paper, we explore this idea and propose the use of manifolds to reduce even more the dimensionality of the rigid transformation space (first stage) of current state-of-the-art top-down segmentation methodologies. Also, we propose the use of deep belief networks to allow for a training process capable to produce robust appearance models. Experiments in lips segmentation from frontal face images are conducted to testify the performance of the proposed algorithm. Jacinto C. Nascimento, Gustavo Carneiro 0001 |
ICIP | 2 |
| 2015 | Lung segmentation in chest radiographs using distance regularized level set and deep-structured learning and inferenceabstractComputer-aided diagnosis of digital chest X-ray (CXR) images critically depends on the automated segmentation of the lungs, which is a challenging problem due to the presence of strong edges at the rib cage and clavicle, the lack of a consistent lung shape among different individuals, and the appearance of the lung apex. From recently published results in this area, hybrid methodologies based on a combination of different techniques (e.g., pixel classification and deformable models) are producing the most accurate lung segmentation results. In this paper, we propose a new methodology for lung segmentation in CXR using a hybrid method based on a combination of distance regularized level set and deep structured inference. This combination brings together the advantages of deep learning methods (robust training with few annotated samples and top-down segmentation with structured inference and learning) and level set methods (use of shape and appearance priors and efficient optimization techniques). Using the publicly available Japanese Society of Radiological Technology (JSRT) dataset, we show that our approach produces the most accurate lung segmentation results in the field. In particular, depending on the initialization used, our methodology produces an average accuracy on JSTR that varies from 94.8% to 98.5%. Tuan Anh Ngo, Gustavo Carneiro 0001 |
ICIP | 2 |
| 2015 | Unregistered Multiview Mammogram Analysis with Pre-trained Deep Learning Models
Gustavo Carneiro 0001, Jacinto C. Nascimento, Andrew P. Bradley |
MICCAI (3) | 1 |
| 2015 | Deep Learning and Structured Prediction for the Segmentation of Mass in Mammograms
Neeraj Dhungel, Gustavo Carneiro 0001, Andrew P. Bradley |
MICCAI (1) | 2 |
| 2015 | An Improved Joint Optimization of Multiple Level Set Functions for the Segmentation of Overlapping Cervical CellsabstractIn this paper, we present an improved algorithm for the segmentation of cytoplasm and nuclei from clumps of overlapping cervical cells. This problem is notoriously difficult because of the degree of overlap among cells, the poor contrast of cell cytoplasm and the presence of mucus, blood, and inflammatory cells. Our methodology addresses these issues by utilizing a joint optimization of multiple level set functions, where each function represents a cell within a clump, that have both unary (intracell) and pairwise (intercell) constraints. The unary constraints are based on contour length, edge strength, and cell shape, while the pairwise constraint is computed based on the area of the overlapping regions. In this way, our methodology enables the analysis of nuclei and cytoplasm from both free-lying and overlapping cells. We provide a systematic evaluation of our methodology using a database of over 900 images generated by synthetically overlapping images of free-lying cervical cells, where the number of cells within a clump is varied from 2 to 10 and the overlap coefficient between pairs of cells from 0.1 to 0.5. This quantitative assessment demonstrates that our methodology can successfully segment clumps of up to 10 cells, provided the overlap between pairs of cells is <;0.2. Moreover, if the clump consists of three or fewer cells, then our methodology can successfully segment individual cells even when the overlap is ~0.5. We also evaluate our approach quantitatively and qualitatively on a set of 16 extended depth of field images, where we are able to segment a total of 645 cells, of which only ~10% are free-lying. Finally, we demonstrate that our method of cell nuclei segmentation is competitive when compared with the current state of the art. Gustavo Carneiro 0001, Andrew P. Bradley |
IEEE Trans. Image Process. | 2 |
| 2014 | Non-rigid Segmentation Using Sparse Low Dimensional Manifolds and Deep Belief NetworksabstractIn this paper, we propose a new methodology for segmenting non-rigid visual objects, where the search procedure is onducted directly on a sparse low-dimensional manifold, guided by the classification results computed from a deep belief network. Our main contribution is the fact that we do not rely on the typical sub-division of segmentation tasks into rigid detection and non-rigid delineation. Instead, the non-rigid segmentation is performed directly, where points in the sparse low-dimensional can be mapped to an explicit contour representation in image space. Our proposal shows significantly smaller search and training complexities given that the dimensionality of the manifold is much smaller than the dimensionality of the search spaces for rigid detection and non-rigid delineation aforementioned, and that we no longer require a two-stage segmentation process. We focus on the problem of left ventricle endocardial segmentation from ultrasound images, and lip segmentation from frontal facial images using the extended Cohn-Kanade (CK+) database. Our experiments show that the use of sparse low dimensional manifolds reduces the search and training complexities of current segmentation approaches without a significant impact on the segmentation accuracy shown by state-of-the-art approaches. Jacinto C. Nascimento, Gustavo Carneiro 0001 |
CVPR | 2 |
| 2014 | Fully Automated Non-rigid Segmentation with Distance Regularized Level Set Evolution Initialized and Constrained by Deep-Structured InferenceabstractWe propose a new fully automated non-rigid segmentation approach based on the distance regularized level set method that is initialized and constrained by the results of a structured inference using deep belief networks. This recently proposed level-set formulation achieves reasonably accurate results in several segmentation problems, and has the advantage of eliminating periodic re-initializations during the optimization process, and as a result it avoids numerical errors. Nevertheless, when applied to challenging problems, such as the left ventricle segmentation from short axis cine magnetic ressonance (MR) images, the accuracy obtained by this distance regularized level set is lower than the state of the art. The main reasons behind this lower accuracy are the dependence on good initial guess for the level set optimization and on reliable appearance models. We address these two issues with an innovative structured inference using deep belief networks that produces reliable initial guess and appearance model. The effectiveness of our method is demonstrated on the MICCAI 2009 left ventricle segmentation challenge, where we show that our approach achieves one of the most competitive results (in terms of segmentation accuracy) in the field. Tuan Anh Ngo, Gustavo Carneiro 0001 |
CVPR | 2 |
| 2013 | Top-Down Segmentation of Non-rigid Visual Objects Using Derivative-Based Search on Sparse ManifoldsabstractThe solution for the top-down segmentation of non rigid visual objects using machine learning techniques is generally regarded as too complex to be solved in its full generality given the large dimensionality of the search space of the explicit representation of the segmentation contour. In order to reduce this complexity, the problem is usually divided into two stages: rigid detection and non-rigid segmentation. The rationale is based on the fact that the rigid detection can be run in a lower dimensionality space (i.e., less complex and faster) than the original contour space, and its result is then used to constrain the non-rigid segmentation. In this paper, we propose the use of sparse manifolds to reduce the dimensionality of the rigid detection search space of current state-of-the-art top-down segmentation methodologies. The main goals targeted by this smaller dimensionality search space are the decrease of the search running time complexity and the reduction of the training complexity of the rigid detector. These goals are attainable given that both the search and training complexities are function of the dimensionality of the rigid search space. We test our approach in the segmentation of the left ventricle from ultrasound images and lips from frontal face images. Compared to the performance of state-of-the-art non-rigid segmentation system, our experiments show that the use of sparse manifolds for the rigid detection leads to the two goals mentioned above. Jacinto C. Nascimento, Gustavo Carneiro 0001 |
CVPR | 2 |
| 2013 | Combining a bottom up and top down classifiers for the segmentation of the left ventricle from cardiac imageryabstractThe segmentation of anatomical structures is a crucial first stage of most medical imaging analysis procedures. A primary example is the segmentation of the left ventricle (LV), from cardiac imagery. Accuracy in the segmentation often requires a considerable amount of expert intervention and guidance which are expensive. Thus, automating the segmentation is welcome, but difficult because of the LV shape variability within and across individuals. To cope with this difficulty, the algorithm should have the skills to interpret the shape of the anatomical structure (i.e. LV shape) using distinct kinds of information, (i.e. different views of the same feature space). These different views will ascribe to the algorithm a more general capability that surely allows for the robustness in the segmentation accuracy. In this paper, we propose an on-line co-training algorithm using a bottom-up and top-down classifiers (each one having a different view of the data) to perform the segmentation of the LV. In particular, we consider a setting in which the LV shape can be partitioned into two distinct views and use a co-training as a way to boost each of the classifiers, thus providing a principled way to use both views together. We testify the usefulness of the approach on a public data base illustrating that the approach compares favorably with other recent proposed methodologies. Jacinto C. Nascimento, Gustavo Carneiro 0001 |
ICIP | 2 |
| 2013 | Left ventricle segmentation from cardiac MRI combining level set methods with deep belief networksabstractThis paper introduces a new semi-automated methodology combining a level set method with a top-down segmentation produced by a deep belief network for the problem of left ventricle segmentation from cardiac magnetic resonance images (MRI). Our approach combines the level set advantages that uses several a priori facts about the object to be segmented (e.g., smooth contour, strong edges, etc.) with the knowledge automatically learned from a manually annotated database (e.g., shape and appearance of the object to be segmented). The use of deep belief networks is justified because of its ability to learn robust models with few annotated images and its flexibility that allowed us to adapt it to a top-down segmentation problem. We demonstrate that our method produces competitive results using the database of the MICCAI grand challenge on left ventricle segmentation from cardiac MRI images, where our methodology produces results on par with the best in the field in each one of the measures used in that challenge (perpendicular distance, Dice metric, and percentage of good detections). Therefore, we conclude that our proposed methodology is one of the most competitive approaches in the field. Tuan Anh Ngo, Gustavo Carneiro 0001 |
ICIP | 2 |
| 2013 | Automated Nucleus and Cytoplasm Segmentation of Overlapping Cervical Cells
Gustavo Carneiro 0001, Andrew P. Bradley |
MICCAI (1) | 2 |
| 2013 | Combining Multiple Dynamic Models and Deep Learning Architectures for Tracking the Left Ventricle Endocardium in Ultrasound DataabstractWe present a new statistical pattern recognition approach for the problem of left ventricle endocardium tracking in ultrasound data. The problem is formulated as a sequential importance resampling algorithm such that the expected segmentation of the current time step is estimated based on the appearance, shape, and motion models that take into account all previous and current images and previous segmentation contours produced by the method. The new appearance and shape models decouple the affine and nonrigid segmentations of the left ventricle to reduce the running time complexity. The proposed motion model combines the systole and diastole motion patterns and an observation distribution built by a deep neural network. The functionality of our approach is evaluated using a dataset of diseased cases containing 16 sequences and another dataset of normal cases comprised of four sequences, where both sets present long axis views of the left ventricle. Using a training set comprised of diseased and healthy cases, we show that our approach produces more accurate results than current state-of-the-art endocardium tracking methods in two test sequences from healthy subjects. Using three test sequences containing different types of cardiopathies, we show that our method correlates well with interuser statistics produced by four cardiologists. Gustavo Carneiro 0001, Jacinto C. Nascimento |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2013 | Artistic Image Analysis Using Graph-Based Learning ApproachesabstractWe introduce a new methodology for the problem of artistic image analysis, which among other tasks, involves the automatic identification of visual classes present in an art work. In this paper, we advocate the idea that artistic image analysis must explore a graph that captures the network of artistic influences by computing the similarities in terms of appearance and manual annotation. One of the novelties of our methodology is the proposed formulation that is a principled way of combining these two similarities in a single graph. Using this graph, we show that an efficient random walk algorithm based on an inverted label propagation formulation produces more accurate annotation and retrieval results compared with the following baseline algorithms: bag of visual words, label propagation, matrix completion, and structural learning. We also show that the proposed approach leads to a more efficient inference and training procedures. This experiment is run on a database containing 988 artistic images (with 49 visual classification problems divided into a multiclass problem with 27 classes and 48 binary problems), where we show the inference and training running times, and quantitative comparisons with respect to several retrieval and annotation performance measures. Gustavo Carneiro 0001 |
IEEE Trans. Image Process. | 1 |
| 2012 | The use of on-line co-training to reduce the training set size in pattern recognition methods: Application to left ventricle segmentation in ultrasoundabstractThe use of statistical pattern recognition models to segment the left ventricle of the heart in ultrasound images has gained substantial attention over the last few years. The main obstacle for the wider exploration of this methodology lies in the need for large annotated training sets, which are used for the estimation of the statistical model parameters. In this paper, we present a new on-line co-training methodologythat reduces the need for large training sets for such parameter estimation. Our approach learns the initial parameters of two different models using a small manually annotated training set. Then, given each frame of a test sequence, the methodology not only produces the segmentation of the current frame, but it also uses the results of both classifiers to retrain each other incrementally. This on-line aspect of our approach has the advantages of producing segmentation results and retraining the classifiers on the fly as frames of a test sequence are presented, but it introduces a harder learning setting compared to the usual off-line co-training, where the algorithm has access to the whole set of un-annotated training samples from the beginning. Moreover, we introduce the use of the following new types of classifiers in the co-training framework: deep belief network and multiple model probabilistic data association. We show that our method leads to a fully automatic left ventricle segmentation system that achieves state-of-the-art accuracy on a public database with training sets containing at least twenty annotated images. Gustavo Carneiro 0001, Jacinto C. Nascimento |
CVPR | 1 |
| 2012 | Artistic Image Classification: An Analysis on the PRINTART Database
Gustavo Carneiro 0001, Nuno Pinho da Silva, Alessio Del Bue, João Paulo Costeira |
ECCV (4) | 1 |
| 2012 | In Defence of RANSAC for Outlier Rejection in Deformable Registration
Quoc-Huy Tran, Tat-Jun Chin, Gustavo Carneiro 0001, Michael S. Brown, David Suter |
ECCV (4) | 3 |
| 2012 | On-line re-training and segmentation with reduction of the training set: Application to the left ventricle detection in ultrasound imagingabstractThe segmentation of the left ventricle (LV) still constitutes an active research topic in medical image processing field. The problem is usually tackled using pattern recognition methodologies. The main difficulty with pattern recognition methods is its dependence of a large manually annotated training sets for a robust learning strategy. However, in medical imaging, it is difficult to obtain such large annotated data. In this paper, we propose an on-line semi-supervised algorithm capable of reducing the need of large training sets. The main difference regarding semi-supervised techniques is that, the proposed framework provides both an on-line retraining and segmentation, instead of on-line retraining and off-line segmentation. Our proposal is applied to a fully automatic LV segmentation with substantially reduced training sets while maintaining good segmentation accuracy. Jacinto C. Nascimento, Gustavo Carneiro 0001 |
ICIP | 2 |
| 2012 | Transparent and scalable terminal mobility for vehicular networks
Gustavo Carneiro 0001, Pedro Fortuna, Jaime Dias, Manuel Ricardo 0001 |
Comput. Networks | 1 |
| 2012 | The Segmentation of the Left Ventricle of the Heart From Ultrasound Data Using Deep Learning Architectures and Derivative-Based Search MethodsabstractWe present a new supervised learning model designed for the automatic segmentation of the left ventricle (LV) of the heart in ultrasound images. We address the following problems inherent to supervised learning models: 1) the need of a large set of training images; 2) robustness to imaging conditions not present in the training data; and 3) complex search process. The innovations of our approach reside in a formulation that decouples the rigid and nonrigid detections, deep learning methods that model the appearance of the LV, and efficient derivative-based search algorithms. The functionality of our approach is evaluated using a data set of diseased cases containing 400 annotated images (from 12 sequences) and another data set of normal cases comprising 80 annotated images (from two sequences), where both sets present long axis views of the LV. Using several error measures to compute the degree of similarity between the manual and automatic segmentations, we show that our method not only has high sensitivity and specificity but also presents variations with respect to a gold standard (computed from the manual annotations of two experts) within interuser variability on a subset of the diseased cases. We also compare the segmentations produced by our approach and by two state-of-the-art LV segmentation models on the data set of normal cases, and the results show that our approach produces segmentations that are comparable to these two approaches using only 20 training images and increasing the training set to 400 images causes our approach to be generally more accurate. Finally, we show that efficient search methods reduce up to tenfold the complexity of the method while still producing competitive segmentations. In the future, we plan to include a dynamical model to improve the performance of the algorithm, to use semisupervised learning methods to reduce even more the dependence on rich and large training sets, and to design a shape model less dependent on the training set. Gustavo Carneiro 0001, Jacinto C. Nascimento, António Freitas |
IEEE Trans. Image Process. | 1 |
| 2011 | Incremental on-line semi-supervised learning for segmenting the left ventricle of the heart from ultrasound dataabstractRecently, there has been an increasing interest in the investigation of statistical pattern recognition models for the fully automatic segmentation of the left ventricle (LV) of the heart from ultrasound data. The main vulnerability of these models resides in the need of large manually annotated training sets for the parameter estimation procedure. The issue is that these training sets need to be annotated by clinicians, which makes this training set acquisition process quite expensive. Therefore, reducing the dependence on large training sets is important for a more extensive exploration of statistical models in the LV segmentation problem. In this paper, we present a novel incremental on-line semi-supervised learning model that reduces the need of large training sets for estimating the parameters of statistical models. Compared to other semi-supervised techniques, our method yields an on-line incremental re-training and segmentation instead of the off-line incremental re-training and segmentation more commonly found in the literature. Another innovation of our approach is that we use a statistical model based on deep learning architectures, which are easily adapted to this on-line incremental learning framework. We show that our fully automatic LV segmentation method achieves state-of-the-art accuracy with training sets containing less than twenty annotated images. Gustavo Carneiro 0001, Jacinto C. Nascimento |
ICCV | 1 |
| 2011 | Reducing the training set using semi-supervised self-training algorithm for segmenting the left ventricle in ultrasound imagesabstractStatistical pattern recognition models are one of the core research topics in the segmentation of the left ventricle of the heart from ultrasound data. The underlying statistical model usually relies on a complex model for the shape and appearance of the left ventricle whose parameters can be learned using a manually segmented data set. Unfortunately, this complex requires a large number of parameters that can be robustly learned only if the training set is sufficiently large. The difficulty in obtaining large training sets is currently a major roadblock for the further exploration of statistical models in medical image analysis. In this paper, we present a novel semi-supervised self-training model that reduces the need of large training sets for estimating the parameters of statistical models. This model is initially trained with a small set of manually segmented images, and for each new test sequence, the system re-estimates the model parameters incrementally without any further manual intervention. We show that state-of-the-art segmentation results can be achieved with training sets containing 50 annotated examples. Jacinto C. Nascimento, Gustavo Carneiro 0001 |
ICIP | 2 |
| 2011 | Graph-based methods for the automatic annotation and retrieval of art printsabstractThe analysis of images taken from cultural heritage artifacts is an emerging area of research in the field of information retrieval. Current methodologies are focused on the analysis of digital images of paintings for the tasks of forgery detection and style recognition. In this paper, we introduce a graph-based method for the automatic annotation and retrieval of digital images of art prints. Such method can help art historians analyze printed art works using an annotated database of digital images of art prints. The main challenge lies in the fact that art prints generally have limited visual information. The results show that our approach produces better results in a weakly annotated database of art prints in terms of annotation and retrieval performance compared to state-of-the-art approaches based on bag of visual words. Gustavo Carneiro 0001 |
ICMR | 1 |
| 2010 | The automatic design of feature spaces for local image descriptors using an ensemble of non-linear feature extractorsabstractThe design of feature spaces for local image descriptors is an important research subject in computer vision due to its applicability in several problems, such as visual classification and image matching. In order to be useful, these descriptors have to present a good trade off between discriminating power and robustness to typical image deformations. The feature spaces of the most useful local descriptors have been manually designed based on the goal above, but this design often limits the use of these descriptors for some specific matching and visual classification problems. Alternatively, there has been a growing interest in producing feature spaces by an automatic combination of manually designed feature spaces, or by an automatic selection of feature spaces and spatial pooling methods, or by the use of distance metric learning methods. While most of these approaches are usually applied to specific matching or classification problems, where test classes are the same as training classes, a few works aim at the general feature transform problem where the training classes are different from the test classes. The hope in the latter works is the automatic design of a universal feature space for local descriptor matching, which is the topic of our work. In this paper, we propose a new incremental method for learning automatically feature spaces for local descriptors. The method is based on an ensemble of non-linear feature extractors trained in relatively small and random classification problems with supervised distance metric learning techniques. Results on two widely used public databases show that our technique produces competitive results in the field. Gustavo Carneiro 0001 |
CVPR | 1 |
| 2010 | Multiple dynamic models for tracking the left ventricle of the heart from ultrasound data using particle filters and deep learning architecturesabstractThe problem of automatic tracking and segmentation of the left ventricle (LV) of the heart from ultrasound images can be formulated with an algorithm that computes the expected segmentation value in the current time step given all previous and current observations using a filtering distribution. This filtering distribution depends on the observation and transition models, and since it is hard to compute the expected value using the whole parameter space of segmentations, one has to resort to Monte Carlo sampling techniques to compute the expected segmentation parameters. Generally, it is straightforward to compute probability values using the filtering distribution, but it is hard to sample from it, which indicates the need to use a proposal distribution to provide an easier sampling method. In order to be useful, this proposal distribution must be carefully designed to represent a reasonable approximation for the filtering distribution. In this paper, we introduce a new LV tracking and segmentation algorithm based on the method described above, where our contributions are focused on a new transition and observation models, and a new proposal distribution. Our tracking and segmentation algorithm achieves better overall results on a previously tested dataset used as a benchmark by the current state-of-the-art tracking algorithms of the left ventricle of the heart from ultrasound images. Gustavo Carneiro 0001, Jacinto C. Nascimento |
CVPR | 1 |
| 2010 | Efficient search methods and deep belief networks with particle filtering for non-rigid tracking: Application to lip trackingabstractPattern recognition methods have become a powerful tool for segmentation in the sense that they are capable of automatically building a segmentation model from training images. However, they present several difficulties, such as requirement of a large set of training data, robustness to imaging conditions not present in the training set, and complexity of the search process. In this paper we tackle the second problem by using a deep belief network learning architecture, and the third problem by resorting to efficient searching algorithms. As an example, we illustrate the performance of the algorithm in lip segmentation and tracking in video sequences. Quantitative comparison using different strategies for the search process are presented. We also compare our approach to a state-of-the-art segmentation and tracking algorithm. The comparison show that our algorithm produces competitive segmentation results and that efficient search strategies reduce ten times the run-complexity. Jacinto C. Nascimento, Gustavo Carneiro 0001 |
ICIP | 2 |
| 2010 | A Comparative Study on the Use of an Ensemble of Feature Extractors for the Automatic Design of Local Image DescriptorsabstractThe use of an ensemble of feature spaces trained with distance metric learning methods has been empirically shown to be useful for the task of automatically designing local image descriptors. In this paper, we present a quantitative analysis which shows that in general, nonlinear distance metric learning methods provide better results than linear methods for automatically designing local image descriptors. In addition, we show that the learned feature spaces present better results than state of- the-art hand designed features in benchmark quantitative comparisons. We discuss the results and suggest relevant problems for further investigation. Gustavo Carneiro 0001 |
ICPR | 1 |
| 2010 | The Fusion of Deep Learning Architectures and Particle Filtering Applied to Lip TrackingabstractThis work introduces a new pattern recognition model for segmenting and tracking lip contours in video sequences. We formulate the problem as a general nonrigid object tracking method, where the computation of the expected segmentation is based on a filtering distribution. This is a difficult task because one has to compute the expected value using the whole parameter space of segmentation. As a result, we compute the expected segmentation using sequential Monte Carlo sampling methods, where the filtering distribution is approximated with a proposal distribution to be used for sampling. The key contribution of this paper is the formulation of this proposal distribution using a new observation model based on deep belief networks and a new transition model. The efficacy of the model is demonstrated in publicly available databases of video sequences of people talking and singing. Our method produces results comparable to state-of-the-art models, but showing potential to be more robust to imaging conditions. Gustavo Carneiro 0001, Jacinto C. Nascimento |
ICPR | 1 |
| 2009 | Fast and Robust 3-D MRI Brain Structure Segmentation
Michael Wels, Yefeng Zheng 0001, Gustavo Carneiro 0001, Martin Huber 0001, Joachim Hornegger, Dorin Comaniciu |
MICCAI (1) | 3 |
| 2009 | The quantitative characterization of the distinctiveness and robustness of local image descriptors
Gustavo Carneiro 0001, Allan Douglas Jepson |
Image Vis. Comput. | 1 |
| 2009 | Minimum Bayes error features for visual recognition
Gustavo Carneiro 0001, Nuno Vasconcelos |
Image Vis. Comput. | 1 |
| 2008 | Semantic-based indexing of fetal anatomies from 3-D ultrasound data using global/semi-local context and sequential samplingabstractThe use of 3-D ultrasound data has several advantages over 2-D ultrasound for fetal biometric measurements, such as considerable decrease in the examination time, possibility of post-exam data processing by experts and the ability to produce 2-D views of the fetal anatomies in orientations that cannot be seen in common 2-D ultrasound exams. However, the search for standardized planes and the precise localization of fetal anatomies in ultrasound volumes are hard and time consuming processes even for expert physicians and sonographers. The relative low resolution in ultrasound volumes, small size of fetus anatomies and inter-volume position, orientation and size variability make this localization problem even more challenging. In order to make the plane search and fetal anatomy localization problems completely automatic, we introduce a novel principled probabilistic model that combines discriminative and generative classifiers with contextual information and sequential sampling. We implement a system based on this model, where the user queries consist of semantic keywords that represent anatomical structures of interest. After queried, the system automatically displays standardized planes and produces biometric measurements of the fetal anatomies. Experimental results on a held-out test set show that the automatic measurements are within the inter-user variability of expert users. It resolves for position, orientation and size of three different anatomies in less than 10 seconds in a dual-core computer running at 1.7 GHz. Gustavo Carneiro 0001, Fernando Amat, Bogdan Georgescu, Sara Good, Dorin Comaniciu |
CVPR | 1 |
| 2008 | A Discriminative Model-Constrained Graph Cuts Approach to Fully Automated Pediatric Brain Tumor Segmentation in 3-D MRI
Michael Wels, Gustavo Carneiro 0001, Alexander Aplas, Martin Huber 0001, Joachim Hornegger, Dorin Comaniciu |
MICCAI (1) | 2 |
| 2008 | Detection and Measurement of Fetal Anatomies from Ultrasound Images using a Constrained Probabilistic Boosting TreeabstractWe propose a novel method for the automatic detection and measurement of fetal anatomical structures in ultrasound images. This problem offers a myriad of challenges, including: difficulty of modeling the appearance variations of the visual object of interest, robustness to speckle noise and signal dropout, and large search space of the detection procedure. Previous solutions typically rely on the explicit encoding of prior knowledge and formulation of the problem as a perceptual grouping task solved through clustering or variational approaches. These methods are constrained by the validity of the underlying assumptions and usually are not enough to capture the complex appearances of fetal anatomies. We propose a novel system for fast automatic detection and measurement of fetal anatomies that directly exploits a large database of expert annotated fetal anatomical structures in ultrasound images. Our method learns automatically to distinguish between the appearance of the object of interest and background by training a constrained probabilistic boosting tree classifier. This system is able to produce the automatic segmentation of several fetal anatomies using the same basic detection algorithm. We show results on fully automatic measurement of biparietal diameter (BPD), head circumference (HC), abdominal circumference (AC), femur length (FL), humerus length (HL), and crown rump length (CRL). Notice that our approach is the first in the literature to deal with the HL and CRL measurements. Extensive experiments (with clinical validation) show that our system is, on average, close to the accuracy of experts in terms of segmentation and obstetric measurements. Finally, this system runs under half second on a standard dual-core PC computer. Gustavo Carneiro 0001, Bogdan Georgescu, Sara Good, Dorin Comaniciu |
IEEE Trans. Medical Imaging | 1 |
| 2007 | A probabilistic, hierarchical, and discriminant framework for rapid and accurate detection of deformable anatomic structureabstractWe propose a probabilistic, hierarchical, and discriminant (PHD) framework for fast and accurate detection of deformable anatomic structures from medical images. The PHD framework has three characteristics. First, it integrates distinctive primitives of the anatomic structures at global, segmental, and landmark levels in a probabilistic manner. Second, since the configuration of the anatomic structures lies in a high-dimensional parameter space, it seeks the best configuration via a hierarchical evaluation of the detection probability that quickly prunes the search space. Finally, to separate the primitive from the background, it adopts a discriminative boosting learning implementation. We apply the PHD framework for accurately detecting various deformable anatomic structures from M- mode and Doppler echocardiograms in about a second. Shaohua Kevin Zhou, Gustavo Carneiro 0001, John Jackson, M. Brendel, Costas Simopoulos, Joanne Otsuki, Dorin Comaniciu |
ICCV | 3 |
| 2007 | Automatic Fetal Measurements in Ultrasound Using Constrained Probabilistic Boosting Tree
Gustavo Carneiro 0001, Bogdan Georgescu, Sara Good, Dorin Comaniciu |
MICCAI (2) | 1 |
| 2007 | Supervised Learning of Semantic Classes for Image Annotation and RetrievalabstractA probabilistic formulation for semantic image annotation and retrieval is proposed. Annotation and retrieval are posed as classification problems where each class is defined as the group of database images labeled with a common semantic label. It is shown that, by establishing this one-to-one correspondence between semantic labels and semantic classes, a minimum probability of error annotation and retrieval are feasible with algorithms that are 1) conceptually simple, 2) computationally efficient, and 3) do not require prior semantic segmentation of training images. In particular, images are represented as bags of localized feature vectors, a mixture density estimated for each image, and the mixtures associated with all images annotated with a common semantic label pooled into a density estimate for the corresponding semantic class. This pooling is justified by a multiple instance learning argument and performed efficiently with a hierarchical extension of expectation-maximization. The benefits of the supervised formulation over the more complex, and currently popular, joint modeling of semantic label and visual feature distributions are illustrated through theoretical arguments and extensive experiments. The supervised formulation is shown to achieve higher accuracy than various previously published methods at a fraction of their computational cost. Finally, the proposed method is shown to be fairly robust to parameter tuning. Gustavo Carneiro 0001, Antoni B. Chan, Pedro J. Moreno 0001, Nuno Vasconcelos |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2007 | Flexible Spatial Configuration of Local Image FeaturesabstractLocal image features have been designed to be informative and repeatable under rigid transformations and illumination deformations. Even though current state-of-the-art local image features present a high degree of repeatability, their local appearance alone usually does not bring enough discriminative power to support a reliable matching, resulting in a relatively high number of mismatches in the correspondence set formed during the data association procedure. As a result, geometric filters, commonly based on global spatial configuration, have been used to reduce this number of mismatches. However, this approach presents a trade off between the effectiveness to reject mismatches and the robustness to non-rigid deformations. In this paper, we propose two geometric filters, based on semilocal spatial configuration of local features, that are designed to be robust to non-rigid deformations and to rigid transformations, without compromising its efficacy to reject mismatches. We compare our methods to the Hough transform, which is an efficient and effective mismatch rejection step based on global spatial configuration of features. In these comparisons, our methods are shown to be more effective in the task of rejecting mismatches for rigid transformations and non-rigid deformations at comparable time complexity figures. Finally, we demonstrate how to integrate these methods in a probabilistic recognition system such that the final verification step uses not only the similarity between features, but also their semi-local configuration. Gustavo Carneiro 0001, Allan Douglas Jepson |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2006 | Weakly Supervised Top-down Image SegmentationabstractThere has recently been significant interest in top-down image segmentation methods, which incorporate the recognition of visual concepts as an intermediate step of segmentation. This work addresses the problem of top-down segmentation with weak supervision. Under this framework, learning does not require a set of manually segmented examples for each concept of interest, but simply a weakly labeled training set. This is a training set where images are annotated with a set of keywords describing their contents, but visual concepts are not explicitly segmented and no correspondence is specified between keywords and image regions. We demonstrate, both analytically and empirically, that weakly supervised segmentation is feasible when certain conditions hold. We also propose a simple weakly supervised segmentation algorithm that extends state-of-theart bottom-up segmentation methods in the direction of perceptually meaningful segmentation1. Manuela Vasconcelos, Nuno Vasconcelos, Gustavo Carneiro 0001 |
CVPR (1) | 3 |
| 2006 | Sparse Flexible Models of Local Features
Gustavo Carneiro 0001 |
ECCV (3) | 1 |
| 2005 | The Distinctiveness, Detectability, and Robustness of Local Image FeaturesabstractWe introduce a new method that characterizes typical local image features (e.g., SIFT, phase feature) in terms of their distinctiveness, detectability, and robustness to image deformations. This is useful for the task of classifying local image features in terms of those three properties. The importance of this classification process for a recognition system using local features is as follows: a) reduce the recognition time due to a smaller number of features present in the test image and in the database of model features; b) improve the recognition accuracy since only the most useful features for the recognition task are kept in the model database; and c) increase the scalability of the recognition system given the smaller number of features per model. A discriminant classifier is trained to select well behaved feature points. A regression network is then trained to provide quantitative models of the detection distributions for each selected feature point. It is important to note that both the classifier and the regression network use image data alone as their input. Experimental results show that the use of these trained networks not only improves the performance of our recognition system, but it also significantly reduces the computation time for the recognition process. Gustavo Carneiro 0001, Allan Douglas Jepson |
CVPR (2) | 1 |
| 2005 | Formulating Semantic Image Annotation as a Supervised Learning ProblemabstractWe introduce a new method to automatically annotate and retrieve images using a vocabulary of image semantics. The novel contributions include a discriminant formulation of the problem, a multiple instance learning solution that enables the estimation of concept probability distributions without prior image segmentation, and a hierarchical description of the density of each image class that enables very efficient training. Compared to current methods of image annotation and retrieval, the one now proposed has significantly smaller time complexity and better recognition performance. Specifically, its recognition complexity is O(C/spl times/R), where C is the number of classes (or image annotations) and R is the number of image regions, while the best results in the literature have complexity O(T/spl times/R), where T is the number of training images. Since the number of classes grows substantially slower than that of training images, the proposed method scales better during training, and processes test images faster This is illustrated through comparisons in terms of complexity, time, and recognition performance with current state-of-the-art methods. Gustavo Carneiro 0001, Nuno Vasconcelos |
CVPR (2) | 1 |
| 2005 | Robust header compression in 4G networks with QoS supportabstractThe 4th generation of mobile communication networks uses heterogeneous wireless technologies. Voice and low quality video flows consist of very small IP packets, for which the standard RTP/UDP/IP headers constitute a significant overhead. Considering that the radio resources are scarce, this overhead may be unacceptable and the adoption of header compression mechanisms is desirable. This paper presents a solution for including RoHC header compression mechanisms within 4G networks; the solution is combined with the mechanisms required to provide QoS to the real time flows Pedro Fortuna, Gustavo Carneiro 0001, Manuel Ricardo 0001 |
PIMRC | 2 |
| 2005 | A database centric view of semantic image annotation and retrievalabstractWe introduce a new model for semantic annotation and retrieval from image databases. The new model is based on a probabilistic formulation that poses annotation and retrieval as classification problems, and produces solutions that are optimal in the minimum probability of error sense. It is also database centric, by establishing a one-to-one mapping between semantic classes and the groups of database images that share the associated semantic labels. In this work we show that, under the database centric probabilistic model, optimal annotation and retrieval can be implemented with algorithms that are conceptually simple, computationally efficient, and do not require prior semantic segmentation of training images. Due to its simplicity, the annotation and retrieval architecture is also amenable to sophisticated parameter tuning, a property that is exploited to investigate the role of feature selection in the design of optimal annotation and retrieval systems. Finally, we demonstrate the benefits of simply establishing a one-to-one mapping between keywords and the states of the semantic classification problem over the more complex, and currently popular, joint modeling of keyword and visual feature distributions. The database centric probabilistic retrieval model is compared to existing semantic labeling and retrieval methods, and shown to achieve higher accuracy than the previously best published results, at a fraction of their computational cost. Gustavo Carneiro 0001, Nuno Vasconcelos |
SIGIR | 1 |
| 2004 | Flexible Spatial Models for Grouping Local Image Features
Gustavo Carneiro 0001, Allan Douglas Jepson |
CVPR (2) | 1 |
| 2003 | Multi-scale Phase-based Local FeaturesabstractLocal feature methods suitable for image feature based object recognition and for the estimation of motion and structure are composed of two steps, namely the 'where' and 'what' steps. The 'where' step (e.g., interest point detector) must select image points that are robustly localizable under common image deformations and whose neighborhoods are relatively informative. The 'what' step (e.g., local feature extractor) then provides a representation of the image neighborhood that is semi-invariant to image deformations, but distinctive enough to provide model identification. We present a quantitative evaluation of both the 'where' and the 'what' steps for three recent local feature methods: a) phase-based local features (Carneiro and Jepson, 2002), b) differential invariants (Schmid and Mohr, 1997), and c) the scale invariant feature transform (SIFT) (Lowe, 1999). Moreover, in order to make the phase-based approach more comparable to the other two approaches, we also introduce a new form of multi-scale interest point detector to be used for its 'where' step. The results show that the phase-based local features lead to better performance than the other two approaches when dealing with common illumination changes, 2D rotation, and sub-pixel translation. On the other hand, the phase-based local features are somewhat more sensitive to scale and large shear changes than the other two methods. Finally, we demonstrate the viability of the phase-based local feature in a simple object recognition system. Gustavo Carneiro 0001, Allan Douglas Jepson |
CVPR (1) | 1 |
| 2002 | Phase-Based Local Features
Gustavo Carneiro 0001, Allan Douglas Jepson |
ECCV (1) | 1 |
| 2002 | What Is the Role of Independence for Visual Recognition?
Nuno Vasconcelos, Gustavo Carneiro 0001 |
ECCV (1) | 2 |
| 2002 | Support of IP QoS over UMTS networksabstractThe paper presents an end-to-end quality of service (QoS) architecture suitable for IP communications scenarios that include UMTS access networks. The rationale for the architecture is justified and its main features are described, notably the QoS management functions on the terminal equipment, the mapping between IP and UMTS QoS parameters and the negotiation of these parameters. Manuel Ricardo 0001, Jaime Dias, Gustavo Carneiro 0001, José Ruela |
PIMRC | 3 |
| 1999 | CONTROLAB MUFA: A Multi-Level Fusion Architecture for Intelligent Navigation of a TelerobotabstractThis paper proposes a multi-level fusion architecture (MUFA) for controlling the navigation of 12 tele-commanded autonomous guided vehicle (AGV). The architecture combines ideas derived from the fundamental concepts of sensor fusion and distributed intelligence. The focus of the work is the development of an intelligent navigation system for a tricycle drive AGV with the ability to move autonomously within any office environment following instructions issued by client stations connected to the office network and reacting accordingly to different situations found in the real world. The modules which integrate the MUFA architecture are discussed and results of some simulation experiments are presented. Eliana P. L. Aude, Gustavo Carneiro 0001, Henrique Serdeira, Julio T. C. Silveira, Mario F. Martins, Ernesto P. Lopes |
ICRA | 2 |