VLDB 2026 Research / reviewers in the wild / expert
Kevin McGuinness
dblp:87/1938
· DBLP profile ↗
64ranked-venue papers
5as first author
26since 2021 · last 2025
0000-0003-1336-6477ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 52 · 3 first-author · 19 since 2021Artificial intelligence and machine learning · 28 · 2 first-author · 14 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Enhancing Brain Disease Diagnosis with XAI: A Review of Recent StudiesabstractThe area of eXplainable Artificial Intelligence (XAI) has shown remarkable progress in the past few years, with the aim of enhancing the transparency and interpretability of the Machine Learning (ML) and Deep Learning (DL) models. This review article presents an in-depth review of the current state-of-the-art XAI techniques applied to the diagnosis of brain diseases. The challenges encountered by traditional ML and DL models within this domain are thoroughly examined, emphasizing the pivotal role of XAI in providing the transparency and interpretability of these models. Furthermore, this article presents a comprehensive survey of the XAI methodologies used for making diagnoses of various brain disorders. Recent studies utilizing XAI for diagnosing a range of brain illnesses, including Alzheimer, brain tumors, dementia, Parkinson, multiple sclerosis, autism, epilepsy, and stroke, are critically reviewed. Finally, the limitations inherent in current XAI techniques are discussed, along with prospective avenues for future research. The key goal of this study is to provide researchers with a roadmap that shows the potential of XAI techniques in improving the interpretability and transparency of DL and ML algorithms for the diagnosis of brain diseases, while also delineating the challenges that require concerted research efforts. Nighat Bibi, Jane Courtney, Kevin McGuinness |
ACM Trans. Comput. Heal. | 3 |
| 2024 | An Accurate Detection Is Not All You Need to Combat Label Noise in Web-Noisy Datasets
Paul Albert, Jack Valmadre, Eric Arazo Sanchez, Tarun Krishna, Noel E. O'Connor, Kevin McGuinness |
ECCV (49) | 6 |
| 2024 | Learning Saliency From FixationsabstractWe present a novel approach for saliency prediction in images, leveraging parallel decoding in transformers to learn saliency solely from fixation maps. Models typically rely on continuous saliency maps, to overcome the difficulty of optimizing for the discrete fixation map. We attempt to replicate the experimental setup that generates saliency datasets. Our approach treats saliency prediction as a direct set prediction problem, via a global loss that enforces unique fixations prediction through bipartite matching and a transformer encoder-decoder architecture. By utilizing a fixed set of learned fixation queries, the cross-attention reasons over the image features to directly output the fixation points, distinguishing it from other modern saliency predictors. Our approach, named Saliency TRansformer (SalTR), achieves metric scores on par with state-of-the-art approaches on the Salicon and MIT300 benchmarks. Y. A. Dahou Djilali, Kevin McGuinness, Noel E. O'Connor |
WACV | 2 |
| 2023 | Vision Transformers are Inherently Saliency Learners
Y. A. Dahou Djilali, Kevin McGuinness, Noel E. O'Connor |
BMVC | 2 |
| 2023 | Unifying Synergies between Self-supervised Learning and Dynamic Computation
Tarun Krishna, Ayush K. Rai, Alexandru Drimbarean, Eric Arazo Sanchez, Paul Albert, Alan F. Smeaton, Kevin McGuinness, Noel E. O'Connor |
BMVC | 7 |
| 2023 | Self-Supervised and Semi-Supervised Polyp Segmentation using Synthetic DataabstractEarly detection of colorectal polyps is of utmost importance for their treatment and for colorectal cancer prevention. Computer vision techniques have the potential to aid professionals in the diagnosis stage, where colonoscopies are manually carried out to examine the entirety of the patient's colon. The main challenge in medical imaging is the lack of data, and a further challenge specific to polyp segmentation approaches is the difficulty of manually labeling the available data: the annotation process for segmentation tasks is very time-consuming. While most recent approaches address the data availability challenge with sophisticated techniques to better exploit the available labeled data, few of them explore the self-supervised or semi-supervised paradigm, where the amount of labeling required is greatly reduced. To address both challenges, we leverage synthetic data and propose an end-to-end model for polyp segmentation that integrates real and synthetic data to artificially increase the size of the datasets and aid the training when unlabeled samples are available. Concretely, our model, PI-CUT-Seg, transforms synthetic images with an image-to-image translation module and combines the resulting images with real images to train a segmentation model, where we use model predictions as pseudolabels to better leverage unlabeled samples. Additionally, we propose PL-CUT-Seg+, an improved version of the model that incorporates targeted regularization to address the domain gap between real and synthetic images. The models are evaluated on standard benchmarks for polyp segmentation and reach state-of-the-art results in the self- and semi-supervised setups. Enric Moreu, Eric Arazo Sanchez, Kevin McGuinness, Noel E. O'Connor |
IJCNN | 3 |
| 2023 | Is your noise correction noisy? PLS: Robustness to label noise with two stage detectionabstractDesigning robust algorithms capable of training accurate neural networks on uncurated datasets from the web has been the subject of much research as it reduces the need for time consuming human labor. The focus of many previous research contributions has been on the detection of different types of label noise; however, this paper proposes to improve the correction accuracy of noisy samples once they have been detected. In many state-of-the-art contributions, a two phase approach is adopted where the noisy samples are detected before guessing a corrected pseudo-label in a semi-supervised fashion. The guessed pseudo-labels are then used in the supervised objective without ensuring that the label guess is likely to be correct. This can lead to confirmation bias, which reduces the noise robustness. Here we propose the pseudo-loss, a simple metric that we find to be strongly correlated with pseudo-label correctness on noisy samples. Using the pseudo-loss, we dynamically down weight under-confident pseudo-labels throughout training to avoid confirmation bias and improve the network accuracy. We additionally propose to use a confidence guided contrastive objective that learns robust representation on an interpolated objective between class bound (supervised) for confidently corrected samples and unsupervised representation for under-confident label corrections. Experiments demonstrate the state-of-the-art performance of our Pseudo-Loss Selection (PLS) algorithm on a variety of benchmark datasets including curated data synthetically corrupted with in-distribution and out-of-distribution noise, and two real world web noise datasets. Our experiments are fully reproducible github.com/PaulAlbert31/PLS. Paul Albert, Eric Arazo Sanchez, Tarun Krishna, Noel E. O'Connor, Kevin McGuinness |
WACV | 5 |
| 2023 | TinyHD: Efficient Video Saliency Prediction with Heterogeneous Decoders using Hierarchical Maps DistillationabstractVideo saliency prediction has recently attracted attention of the research community, as it is an upstream task for several practical applications. However, current solutions are particurly computationally demanding, especially due to the wide usage of spatio-temporal 3D convolutions. We observe that, while different model architectures achieve similar performance on benchmarks, visual variations between predicted saliency maps are still significant. Inspired by this intuition, we propose a lightweight model that employs multiple simple heterogeneous decoders and adopts several practical approaches to improve accuracy while keeping computational costs low, such as hierarchical multi-map knowledge distillation, multi-output saliency prediction, unlabeled auxiliary datasets and channel reduction with teacher assistant supervision. Our approach achieves saliency prediction accuracy on par or better than state-of-the-art methods on DFH1K, UCF-Sports and Hollywood2 benchmarks, while enhancing significantly the efficiency of the model. Feiyan Hu, Simone Palazzo, Federica Proietto Salanitri, Giovanni Bellitto, Morteza Moradi 0001, Concetto Spampinato, Kevin McGuinness |
WACV | 7 |
| 2023 | Motion Aware Self-Supervision for Generic Event Boundary DetectionabstractThe task of Generic Event Boundary Detection (GEBD) aims to detect moments in videos that are naturally perceived by humans as generic and taxonomy-free event boundaries. Modeling the dynamically evolving temporal and spatial changes in a video makes GEBD a difficult problem to solve. Existing approaches involve very complex and sophisticated pipelines in terms of architectural design choices, hence creating a need for more straightforward and simplified approaches. In this work, we address this issue by revisiting a simple and effective self-supervised method and augment it with a differentiable motion feature learning module to tackle the spatial and temporal diversities in the GEBD task. We perform extensive experiments on the challenging Kinetics-GEBD and TAPOS datasets to demonstrate the efficacy of the proposed approach compared to the other self-supervised state-of-the-art methods. We also show that this simple self-supervised approach learns motion features without any explicit motion-specific pretext task. Our results can be reproduced on $github$. Ayush K. Rai, Tarun Krishna, Julia Dietlmeier, Kevin McGuinness, Alan F. Smeaton, Noel E. O'Connor |
WACV | 4 |
| 2023 | Fast and robust video-based exercise classification via body pose tracking and scalable multivariate time series classifiers
Antonio Bevilacqua, Thach Le Nguyen, Feiyan Hu, Kevin McGuinness, Martin O'Reilly 0001, Darragh Whelan, Brian Caulfield 0001, Georgiana Ifrim |
Data Min. Knowl. Discov. | 5 |
| 2023 | Joint one-sided synthetic unpaired image translation and segmentation for colorectal cancer preventionabstractAbstract Deep learning has shown excellent performance in analysing medical images. However, datasets are difficult to obtain due privacy issues, standardization problems, and lack of annotations. We address these problems by producing realistic synthetic images using a combination of 3D technologies and generative adversarial networks. We propose CUT‐seg, a joint training where a segmentation model and a generative model are jointly trained to produce realistic images while learning to segment polyps. We take advantage of recent one‐sided translation models because they use significantly less memory, allowing us to add a segmentation model in the training loop. CUT‐seg performs better, is computationally less expensive, and requires less real images than other memory‐intensive image translation approaches that require two stage training. Promising results are achieved on five real polyp segmentation datasets using only one real image and zero real annotations. As a part of this study we release Synth‐Colon, an entirely synthetic dataset that includes 20,000 realistic colon images and additional details about depth and 3D geometry: https://enric1994.github.io/synth-colon Enric Moreu, Eric Arazo Sanchez, Kevin McGuinness, Noel E. O'Connor |
Expert Syst. J. Knowl. Eng. | 3 |
| 2023 | Dexterous robotic manipulation using deep reinforcement learning and knowledge transfer for complex sparse reward-based tasksabstractAbstract This paper describes a deep reinforcement learning (DRL) approach that won Phase 1 of the Real Robot Challenge (RRC) 2021, and then extends this method to a more difficult manipulation task. The RRC consisted of using a TriFinger robot to manipulate a cube along a specified positional trajectory, but with no requirement for the cube to have any specific orientation. We used a relatively simple reward function, a combination of a goal‐based sparse reward and a distance reward, in conjunction with Hindsight Experience Replay (HER) to guide the learning of the DRL agent (Deep Deterministic Policy Gradient [DDPG]). Our approach allowed our agents to acquire dexterous robotic manipulation strategies in simulation. These strategies were then deployed on the real robot and outperformed all other competition submissions, including those using more traditional robotic control techniques, in the final evaluation stage of the RRC. Here we extend this method, by modifying the task of Phase 1 of the RRC to require the robot to maintain the cube in a particular orientation, while the cube is moved along the required positional trajectory. The requirement to also orient the cube makes the agent less able to learn the task through blind exploration due to increased problem complexity. To circumvent this issue, we make novel use of a Knowledge Transfer (KT) technique that allows the strategies learned by the agent in the original task (which was agnostic to cube orientation) to be transferred to this task (where orientation matters). KT allowed the agent to learn and perform the extended task in the simulator, which improved the average positional deviation from 0.134 to 0.02 m, and average orientation deviation from 142° to 76° during evaluation. This KT concept shows good generalization properties and could be applied to any actor‐critic learning algorithm. Francisco Roldan, Robert McCarthy, David Cordova Bulens, Kevin McGuinness, Noel E. O'Connor, Manuel Wüthrich, Felix Widmaier, Stefan Bauer, Stephen James Redmond |
Expert Syst. J. Knowl. Eng. | 5 |
| 2022 | BaseTransformers: Attention over base data-points for One Shot Learning
Mayug Maniparambil, Kevin McGuinness, Noel E. O'Connor |
BMVC | 2 |
| 2022 | A Fine Grained Quality Assessment of Video Anomaly DetectionabstractIn this paper we propose a new approach to assess the performance of video anomaly detection algorithms. Inspired by the COCO metrics we propose a quartile based quality assessment of video anomaly detection to have a detailed breakdown of algorithm performance. The proposed assessment divides the detection into five categories based on the measurement quartiles of the position, scale and motion magnitude of anomalies. A weighted precision is introduced in the average precision calculation such that the frame-level average precision reported in categories can be compared to each other regardless of the baseline of the precision-recall curve in every category. Kevin McGuinness, Joseph Antony, Noel E. O'Connor |
CBMI | 2 |
| 2022 | Embedding Contrastive Unsupervised Features to Cluster In- And Out-of-Distribution Noise in Corrupted Image Datasets
Paul Albert, Eric Arazo Sanchez, Noel E. O'Connor, Kevin McGuinness |
ECCV (31) | 4 |
| 2022 | Pixinwav: Residual Steganography for Hiding Pixels in AudioabstractSteganography comprises the mechanics of hiding data in a host media that may be publicly available. While previous works focused on unimodal setups (e.g., hiding images in images, or hiding audio in audio), PixInWav targets the multimodal case of hiding images in audio. To this end, we propose a novel residual architecture operating on top of short-time discrete cosine transform (STDCT) audio spectrograms. Among our results, we find that the residual steganography setup we propose allows an encoding of the hidden image that is independent from the host audio without compromising quality. Accordingly, while previous works require both host and hidden signals to hide a signal, PixInWav can encode images offline—which can be later hidden, in a residual fashion, into any audio signal. Margarita Geleta, Cristina Punti, Kevin McGuinness, Jordi Pons, Cristian Canton, Xavier Giró-i-Nieto |
ICASSP | 3 |
| 2022 | Evaluation of Automatically Generated Video Captions Using Vision and Language ModelsabstractVision and language models are easily transferred to other tasks. In particular, they have been demonstrated to work well in the evaluation of automatic image captioning. This has made it possible to evaluate systems without the need for references or additional information apart from the image and the caption. However, these models do not provide a straightforward way of evaluating videos. In this paper, we propose using these models for video captioning evaluation. We explore the use of both single image-based evaluation and different methods to include data from multiple frames. Experiments demonstrate that using clustering methods to select a few frames to compute the final score gives an excellent correlation with human judgment. The bias in the human annotations can also influence the metric, so we propose filtering the human assessments to discard outliers and improve the evaluation process. Luis Lebron Casas, Yvette Graham, Noel E. O'Connor, Kevin McGuinness |
ICIP | 4 |
| 2022 | BERTHA: Video Captioning Evaluation Via Transfer-Learned Human AssessmentabstractEvaluating video captioning systems is a challenging task as there are multiple factors to consider; for instance: the fluency of the caption, multiple actions happening in a single scene, and the human bias of what is considered important. Most metrics try to measure how similar the system generated captions are to a single or a set of human-annotated captions. This paper presents a new method based on a deep learning model to evaluate these systems. The model is based on BERT, which is a language model that has been shown to work well in multiple NLP tasks. The aim is for the model to learn to perform an evaluation similar to that of a human. To do so, we use a dataset that contains human evaluations of system generated captions. The dataset consists of the human judgments of the captions produces by the system participating in various years of the TRECVid video to text task. BERTHA obtain favourable results, outperforming the commonly used metrics in some setups. Luis Lebron Casas, Yvette Graham, Kevin McGuinness, Konstantinos Kouramas, Noel E. O'Connor |
LREC | 3 |
| 2022 | Addressing out-of-distribution label noise in webly-labelled dataabstractA recurring focus of the deep learning community is towards reducing the labeling effort. Data gathering and annotation using a search engine is a simple alternative to generating a fully human-annotated and human-gathered dataset. Although web crawling is very time efficient, some of the retrieved images are unavoidably noisy, i.e. incorrectly labeled. Designing robust algorithms for training on noisy data gathered from the web is an important research perspective that would render the building of datasets easier. In this paper we conduct a study to understand the type of label noise to expect when building a dataset using a search engine. We review the current limitations of state-of-the-art methods for dealing with noisy labels for image classification tasks in the case of web noise distribution. We propose a simple solution to bridge the gap with a fully clean dataset using Dynamic Softening of Out-of-distribution Samples (DSOS), which we design on corrupted versions of the CIFAR-100 dataset, and compare against state-of-the-art algorithms on the web noise perturbated MiniImageNet and Stanford datasets and on real label noise datasets: WebVision 1.0 and Clothing1M. Our work is fully reproducible https://git.io/JKGcj. Paul Albert, Diego Ortego, Eric Arazo Sanchez, Noel E. O'Connor, Kevin McGuinness |
WACV | 5 |
| 2021 | How Important is Importance Sampling for Deep Budgeted Training?
Eric Arazo Sanchez, Diego Ortego, Paul Albert, Noel E. O'Connor, Kevin McGuinness |
BMVC | 5 |
| 2021 | Multi-Objective Interpolation Training for Robustness To Label NoiseabstractDeep neural networks trained with standard cross-entropy loss memorize noisy labels, which degrades their performance. Most research to mitigate this memorization proposes new robust classification loss functions. Conversely, we propose a Multi-Objective Interpolation Training (MOIT) approach that jointly exploits contrastive learning and classification to mutually help each other and boost performance against label noise. We show that standard supervised contrastive learning degrades in the presence of label noise and propose an interpolation training strategy to mitigate this behavior. We further propose a novel label noise detection method that exploits the robust feature representations learned via contrastive learning to estimate per-sample soft-labels whose disagreements with the original labels accurately identify noisy samples. This detection allows treating noisy samples as unlabeled and training a classifier in a semi-supervised manner to prevent noise memorization and improve representation learning. We further propose MOIT+, a refinement of MOIT by fine-tuning on detected clean samples. Hyperparameter and ablation studies verify the key components of our method. Experiments on synthetic and real-world noise benchmarks demonstrate that MOIT/MOIT+ achieves state-of-the-art results. Code is available at https://git.io/JI40X. Diego Ortego, Eric Arazo Sanchez, Paul Albert, Noel E. O'Connor, Kevin McGuinness |
CVPR | 5 |
| 2021 | Unsupervised Contrastive Learning of Sound Event RepresentationsabstractSelf-supervised representation learning can mitigate the limitations in recognition tasks with few manually labeled data but abundant unlabeled data—a common scenario in sound event research. In this work, we explore unsupervised contrastive learning as a way to learn sound event representations. To this end, we propose to use the pretext task of contrasting differently augmented views of sound events. The views are computed primarily via mixing of training examples with unrelated backgrounds, followed by other data augmentations. We analyze the main components of our method via ablation experiments. We evaluate the learned representations using linear evaluation, and in two in-domain downstream sound event classification tasks, namely, using limited manually labeled data, and using noisy labeled data. Our results suggest that unsupervised contrastive pre-training can mitigate the impact of data scarcity and increase robustness against noisy labels. Eduardo Fonseca, Diego Ortego, Kevin McGuinness, Noel E. O'Connor, Xavier Serra |
ICASSP | 3 |
| 2021 | Rethinking 360° Image Visual Attention Modelling with Unsupervised LearningabstractDespite the success of self-supervised representation learning on planar data, to date it has not been studied on 360° images. In this paper, we extend recent advances in contrastive learning to learn latent representations that are sufficiently invariant to be highly effective for spherical saliency prediction as a downstream task. We argue that omni-directional images are particularly suited to such an approach due to the geometry of the data domain. To verify this hypothesis, we design an unsupervised framework that effectively maximizes the mutual information between the different views from both the equator and the poles. We show that the decoder is able to learn good quality saliency distributions from the encoder embeddings. Our model compares favorably with fully-supervised learning methods on the Salient360!, VR-EyeTracking and Sitzman datasets. This performance is achieved using an encoder that is trained in a completely unsupervised way and a relatively lightweight supervised decoder (3.8 × fewer parameters in the case of the ResNet50 encoder). We believe that this combination of supervised and unsupervised learning is an important step toward flexible formulations of human visual attention. The results can be reproduced on GitHub Y. A. Dahou Djilali, Tarun Krishna, Kevin McGuinness, Noel E. O'Connor |
ICCV | 3 |
| 2021 | ReLaB: Reliable Label Bootstrapping for Semi-Supervised LearningabstractReducing the amount of labels required to train convolutional neural networks without performance degradation is key to effectively reduce human annotation efforts. We propose Reliable Label Bootstrapping (ReLaB), an unsupervised preprossessing algorithm which improves the performance of semi-supervised algorithms in extremely low supervision settings. Given a dataset with few labeled samples, we first learn meaningful self-supervised, latent features for the data. Second, a label propagation algorithm propagates the known labels on the unsupervised features, effectively labeling the full dataset in an automatic fashion. Third, we select a subset of correctly labeled (reliable) samples using a label noise detection algorithm. Finally, we train a semi-supervised algorithm on the extended subset. We show that the selection of the network architecture and the self-supervised algorithm are important factors to achieve successful label propagation and demonstrate that ReLaB substantially improves semi-supervised learning in scenarios of very limited supervision on image classification benchmarks such as CIFAR-10, CIFAR-100 and mini-ImageNet. We reach average error rates of 22.34 with 1 random labeled sample per class on CIFAR-10 and lower this error to 8.46 when the labeled sample in each class is highly representative. Our work is fully reproducible: https://github.com/PaulAlbert31/ReLaB. Paul Albert, Diego Ortego, Eric Arazo Sanchez, Noel E. O'Connor, Kevin McGuinness |
IJCNN | 5 |
| 2021 | Evaluating Contrastive Models for Instance-based Image RetrievalabstractIn this work, we evaluate contrastive models for the task of image retrieval. We hypothesise that models that are learned to encode semantic similarity among instances via discriminative learning should perform well on the task of image retrieval, where relevancy is defined in terms of instances of the same object. Through our extensive evaluation, we find that representations from models trained using contrastive methods perform on-par with (and outperforms) a pre-trained supervised baseline trained on the ImageNet labels in retrieval tasks under various configurations. This is remarkable given that the contrastive models require no explicit supervision. Thus, we conclude that these models can be used to bootstrap base models to build more robust image retrieval engines. Tarun Krishna, Kevin McGuinness, Noel E. O'Connor |
ICMR | 2 |
| 2021 | Restricted Boltzmann machine as an aggregation technique for binary descriptors
Szymon Sobczak, Rafal Kapela, Kevin McGuinness, Aleksandra Swietlicka, Dariusz Pazderski, Noel E. O'Connor |
Vis. Comput. | 3 |
| 2020 | Investigating Class-Level Difficulty Factors In Multi-Label Classification ProblemsabstractThis work investigates the use of class-level difficulty factors in multi-label classification problems for the first time. Four class-level difficulty factors are proposed: frequency, visual variation, semantic abstraction, and class co-occurrence. Once computed for a given multi-label classification dataset, these difficulty factors are shown to have several potential applications including the prediction of class-level performance across datasets and the improvement of predictive performance through difficulty weighted optimisation. Significant improvements to mAP and AUC performance are observed for two challenging multi-label datasets (WWW Crowd and Visual Genome) with the inclusion of difficulty weighted optimisation. The proposed technique does not require any additional computational complexity during training or inference and can be extended over time with inclusion of other class-level difficulty factors. Mark Marsden, Kevin McGuinness, Joseph Antony, Haolin Wei, Milan D. Redzic, Zhilan Hu, Alan F. Smeaton, Noel E. O'Connor |
ICME | 2 |
| 2020 | How important are faces for person re-identification?abstractThis paper investigates the dependence of existing state-of-the-art person re-identification models on the presence and visibility of human faces. We apply a face detection and blurring algorithm to create anonymized versions of several popular person re-identification datasets including Market1501, DukeMTMC-reID, CUHK03, Viper, and Airport. Using a cross-section of existing state-of-the-art models that range in accuracy and computational efficiency, we evaluate the effect of this anonymization on re-identification performance using standard metrics. Perhaps surprisingly, the effect on mAP is very small, and accuracy is recovered by simply training on the anonymized versions of the data rather than the original data. These findings are consistent across multiple models and datasets. These results indicate that datasets can be safely anonymized by blurring faces without significantly impacting the performance of person reidentification systems, and may allow for the release of new richer re-identification datasets where previously there were privacy or data protection concerns. Julia Dietlmeier, Joseph Antony, Kevin McGuinness, Noel E. O'Connor |
ICPR | 3 |
| 2020 | FastSal: a Computationally Efficient Network for Visual Saliency PredictionabstractThis paper focuses on the problem of visual saliency prediction, predicting regions of an image that tend to attract human visual attention, under a constrained computational budget. We modify and test various recent efficient convolutional neural network architectures like EfficientNet and MobileNetV2 and compare them with existing state-of-the-art saliency models such as SalGAN and DeepGaze II both in terms of standard accuracy metrics like Area Under Curve (AUC) and Normalized Scanpath Saliency (NSS), and in terms of the computational complexity and model size. We find that MobileNetV2 makes an excellent backbone for a visual saliency model and can be effective even without a complex decoder. We also show that knowledge transfer from a more computationally expensive model like DeepGaze II can be achieved via pseudo-labelling an unlabelled dataset, and that this approach gives result on-par with many state-of-the-art algorithms with a fraction of the computational cost and model size. Feiyan Hu, Kevin McGuinness |
ICPR | 2 |
| 2020 | Towards Robust Learning with Different Label Noise DistributionsabstractNoisy labels are an unavoidable consequence of labeling processes and detecting them is an important step towards preventing performance degradations in Convolutional Neural Networks. Discarding noisy labels avoids a harmful memorization, while the associated image content can still be exploited in a semi-supervised learning (SSL) setup. Clean samples are usually identified using the small loss trick, i.e. they exhibit a low loss. However, we show that different noise distributions make the application of this trick less straightforward and propose to continuously relabel all images to reveal a discriminative loss against multiple distributions. SSL is then applied twice, once to improve the clean-noisy detection and again for training the final model. We design an experimental setup based on ImageNet32/64 for better understanding the consequences of representation learning with differing label noise distributions and find that non-uniform out-of-distribution noise better resembles real-world noise and that in most cases intermediate features are not affected by label noise corruption. Experiments in CIFAR-10/100, ImageNet32/64 and WebVision (real-world noise) demonstrate that the proposed label noise Distribution Robust Pseudo-Labeling (DRPL) approach gives substantial improvements over recent state-of-the-art. Code is available at https://git.io/JJ0PV. Diego Ortego, Eric Arazo Sanchez, Paul Albert, Noel E. O'Connor, Kevin McGuinness |
ICPR | 5 |
| 2020 | Pseudo-Labeling and Confirmation Bias in Deep Semi-Supervised LearningabstractSemi-supervised learning, i.e. jointly learning from labeled and unlabeled samples, is an active research topic due to its key role on relaxing human supervision. In the context of image classification, recent advances to learn from unlabeled samples are mainly focused on consistency regularization methods that encourage invariant predictions for different perturbations of unlabeled samples. We, conversely, propose to learn from unlabeled data by generating soft pseudo-labels using the network predictions. We show that a naive pseudo-labeling overfits to incorrect pseudo-labels due to the so-called confirmation bias and demonstrate that mixup augmentation and setting a minimum number of labeled samples per mini-batch are effective regularization techniques for reducing it. The proposed approach achieves state-of-the-art results in CIFAR-10/100, SVHN, and Mini-ImageNet despite being much simpler than other methods. These results demonstrate that pseudo-labeling alone can outperform consistency regularization methods, while the opposite was supposed in previous work. Source code is available at https://git.io/fjQsC. Eric Arazo Sanchez, Diego Ortego, Paul Albert, Noel E. O'Connor, Kevin McGuinness |
IJCNN | 5 |
| 2020 | A Smart-Site-Survey System using Image-based 3D Metric Reconstruction and Interactive Panorama VisualizationabstractThis work presents a so-called Smart Site Survey (SSS) system that provides an efficient, web-based platform for virtual inspection of remote sites with absolute 3D metrics. Traditional manual surveying requires sending surveyors and specialised measuring tools to the targeted scene, which takes time and requires significant human resource, and often includes human error. The proposed system provides an automated site survey tool. Sample indoor scenes including offices, storage rooms, and laboratory are used for testing purposes, and highly precise virtual scenes are restored, with the measurement accuracy of 1%, i.e. an error ±1.5cm to a 150cm length. This is comparable or superior to existing works or commercial products. Sha Yu, Kevin McGuinness, Patricia Moore, David Azcona, Noel E. O'Connor |
ACM Multimedia | 2 |
| 2019 | Simple vs complex temporal recurrences for video saliency prediction
Panagiotis Linardos, Eva Mohedano, Juan José Nieto 0002, Noel E. O'Connor, Xavier Giró-i-Nieto, Kevin McGuinness |
BMVC | 6 |
| 2019 | On guiding video object segmentationabstractThis paper presents a novel approach for segmenting moving objects in unconstrained environments using guided convolutional neural networks. This guiding process relies on foreground masks from independent algorithms (i.e. state-of-the-art algorithms) to implement an attention mechanism that incorporates the spatial location of foreground and background to compute their separated representations. Our approach initially extracts two kinds of features for each frame using colour and optical flow information. Such features are combined following a multiplicative scheme to benefit from their complementarity. These unified colour and motion features are later processed to obtain the separated foreground and background representations. Then, both independent representations are concatenated and decoded to perform foreground segmentation. Experiments conducted on the challenging DAVIS 2016 dataset demonstrate that our guided representations not only outperform non-guided, but also recent and top-performing video object segmentation algorithms. Diego Ortego, Kevin McGuinness, Juan C. SanMiguel, Eric Arazo Sanchez, José María Martínez Sanchez, Noel E. O'Connor |
CBMI | 2 |
| 2019 | Wav2Pix: Speech-conditioned Face Generation Using Generative Adversarial NetworksabstractSpeech is a rich biometric signal that contains information about the identity, gender and emotional state of the speaker. In this work, we explore its potential to generate face images of a speaker by conditioning a Generative Adversarial Network (GAN) with raw speech input. We propose a deep neural network that is trained from scratch in an end-to-end fashion, generating a face directly from the raw speech waveform without any additional identity information (e.g reference image or one-hot encoding). Our model is trained in a self-supervised approach by exploiting the audio and visual signals naturally aligned in videos. With the purpose of training from video data, we present a novel dataset collected for this work, with high-quality videos of youtubers with notable expressiveness in both the speech and visual signals. Amanda Cardoso Duarte, Francisco Roldan, Miquel Tubau, Janna Escur, Santiago Pascual, Amaia Salvador, Eva Mohedano, Kevin McGuinness, Jordi Torres, Xavier Giró-i-Nieto |
ICASSP | 8 |
| 2019 | Unsupervised Label Noise Modeling and Loss CorrectionabstractDespite being robust to small amounts of label noise, convolutional neural networks trained with stochastic gradient methods have been shown to easily fit random labels. When there are a mixture of correct and mislabelled targets, networks tend to fit the former before the latter. This suggests using a suitable two-component mixture model as an unsupervised generative model of sample loss values during training to allow online estimation of the probability that a sample is mislabelled. Specifically, we propose a beta mixture to estimate this probability and correct the loss by relying on the network prediction (the so-called bootstrapping loss). We further adapt mixup augmentation to drive our approach a step further. Experiments on CIFAR-10/100 and TinyImageNet demonstrate a robustness to label noise that substantially outperforms recent state-of-the-art. Source code is available at https://git.io/fjsvE and Appendix at https://arxiv.org/abs/1904.11238. Eric Arazo Sanchez, Diego Ortego, Paul Albert, Noel E. O'Connor, Kevin McGuinness |
ICML | 5 |
| 2019 | Exploring the Impact of Training Data Bias on Automatic Generation of Video Captions
Alan F. Smeaton, Yvette Graham, Kevin McGuinness, Noel E. O'Connor, Seán Quinn, Eric Arazo Sanchez |
MMM (1) | 3 |
| 2019 | Few-shot hypercolumn-based mitochondria segmentation in cardiac and outer hair cells in focused ion beam-scanning electron microscopy (FIB-SEM) data
Julia Dietlmeier, Kevin McGuinness, Sandra Rugonyi, Teresa Wilson, Alfred L. Nuttall, Noel E. O'Connor |
Pattern Recognit. Lett. | 2 |
| 2018 | Saliency Weighted Convolutional Features for Instance SearchabstractThis work explores attention models to weight the contribution of local convolutional representations for the instance search task. We present a retrieval framework based on bags of local convolutional features (BLCF) that benefits from saliency weighting to build an efficient image representation. The use of human visual attention models (saliency) allows significant improvements in retrieval performance without the need to conduct region analysis or spatial verification, and without requiring any feature fine tuning. We investigate the impact of different saliency models, finding that higher performance on saliency benchmarks does not necessarily equate to improved performance when used in instance search tasks. The proposed approach outperforms the state-of-the-art on the challenging INSTRE benchmark by a large margin, and provides similar performance on the Oxford and Paris benchmarks compared to more complex methods that use off-the-shelf representations. Source code is publicly available at https:llgithub.com/imatge-upc/salbow. Eva Mohedano, Kevin McGuinness, Xavier Giró-i-Nieto, Noel E. O'Connor |
CBMI | 2 |
| 2018 | People, Penguins and Petri Dishes: Adapting Object Counting Models to New Visual Domains and Object Types Without ForgettingabstractIn this paper we propose a technique to adapt a convolutional neural network (CNN) based object counter to additional visual domains and object types while still preserving the original counting function. Domain-specific normalisation and scaling operators are trained to allow the model to adjust to the statistical distributions of the various visual domains. The developed adaptation technique is used to produce a singular patch-based counting regressor capable of counting various object types including people, vehicles, cell nuclei and wildlife. As part of this study a challenging new cell counting dataset in the context of tissue culture and patient diagnosis is constructed. This new collection, referred to as the Dublin Cell Counting (DCC) dataset, is the first of its kind to be made available to the wider computer vision community. State-of-the-art object counting performance is achieved in both the Shanghaitech (parts A and B) and Penguins datasets while competitive performance is observed on the TRANCOS and Modified Bone Marrow (MBM) datasets, all using a shared counting model. Mark Marsden, Kevin McGuinness, Suzanne Little, Ciara E. Keogh, Noel E. O'Connor |
CVPR | 2 |
| 2018 | Demonstration of an Open Source Framework for Qualitative Evaluation of CBIR SystemsabstractEvaluating image retrieval systems in a quantitative way, for example by computing measures like mean average precision, allows for objective comparisons with a ground-truth. However, in cases where ground-truth is not available, the only alternative is to collect feedback from a user. Thus, qualitative assessments become important to better understand how the system works. Visualizing the results could be, in some scenarios, the only way to evaluate the results obtained and also the only opportunity to identify that a system is failing. This necessitates developing a User Interface (UI) for a Content Based Image Retrieval (CBIR) system that allows visualization of results and improvement via capturing user relevance feedback. A well-designed UI facilitates understanding of the performance of the system, both in cases where it works well and perhaps more importantly those which highlight the need for improvement. Our open-source system implements three components to facilitate researchers to quickly develop these capabilities for their retrieval engine. We present: a web-based user interface to visualize retrieval results and collect user annotations; a server that simplifies connection with any underlying CBIR system; and a server that manages the search engine data. The software itself is described in a separate submission to the ACM MM Open Source Software Competition. Paula Gómez Duran, Eva Mohedano, Kevin McGuinness, Xavier Giró-i-Nieto, Noel E. O'Connor |
ACM Multimedia | 3 |
| 2018 | A Text Recognition and Retrieval System for e-Business Image Management
Kevin McGuinness, Noel E. O'Connor |
MMM (2) | 2 |
| 2018 | Scanpath and saliency prediction on 360 degree images
Marc Assens, Xavier Giró-i-Nieto, Kevin McGuinness, Noel E. O'Connor |
Signal Process. Image Commun. | 3 |
| 2017 | ResnetCrowd: A residual deep learning architecture for crowd counting, violent behaviour detection and crowd density level classificationabstractIn this paper we propose ResnetCrowd, a deep residual architecture for simultaneous crowd counting, violent behaviour detection and crowd density level classification. To train and evaluate the proposed multi-objective technique, a new 100 image dataset referred to as Multi Task Crowd is constructed. This new dataset is the first computer vision dataset fully annotated for crowd counting, violent behaviour detection and density level classification. Our experiments show that a multi-task approach boosts individual task performance for all tasks and most notably for violent behaviour detection which receives a 9% boost in ROC curve AUC (Area under the curve). The trained ResnetCrowd model is also evaluated on several additional benchmarks highlighting the superior generalisation of crowd analysis models trained for multiple objectives. Mark Marsden, Kevin McGuinness, Suzanne Little, Noel E. O'Connor |
AVSS | 2 |
| 2016 | Shallow and Deep Convolutional Networks for Saliency PredictionabstractThe prediction of salient areas in images has been traditionally addressed with hand-crafted features based on neuroscience principles. This paper, however, addresses the problem with a completely data-driven approach by training a convolutional neural network (convnet). The learning process is formulated as a minimization of a loss function that measures the Euclidean distance of the predicted saliency map with the provided ground truth. The recent publication of large datasets of saliency prediction has provided enough data to train end-to-end architectures that are both fast and accurate. Two designs are proposed: a shallow convnet trained from scratch, and a another deeper solution whose first three layers are adapted from another network trained for classification. To the authors' knowledge, these are the first end-to-end CNNs trained and tested for the purpose of saliency prediction. Junting Pan, Elisa Sayrol, Xavier Giró-i-Nieto, Kevin McGuinness, Noel E. O'Connor |
CVPR | 4 |
| 2016 | Reduction of false alarms triggered by spiders/cobwebs in surveillance camera networksabstractThe percentage of false alarms caused by spiders in automated surveillance can range from 20-50%. False alarms increase the workload of surveillance personnel validating the alarms and the maintenance labor cost associated with regular cleaning of webs. We propose a novel, cost effective method to detect false alarms triggered by spiders/webs in surveillance camera networks. This is accomplished by building a spider classifier intended to be a part of the surveillance video processing pipeline. The proposed method uses a feature descriptor obtained by early fusion of blur and texture. The approach is sufficiently efficient for real-time processing and yet comparable in performance with more computationally costly approaches like SIFT with bag of visual words aggregation. The proposed method can eliminate 98.5% of false alarms caused by spiders in a data set supplied by an industry partner, with a false positive rate of less than 1%. Ramya Hebbalaguppe, Kevin McGuinness, Jogile Kuklyte, Rami Albatal, Cem Direkoglu, Noel E. O'Connor |
ICIP | 2 |
| 2016 | Holistic features for real-time crowd behaviour anomaly detectionabstractThis paper presents a new approach to crowd behaviour anomaly detection that uses a set of efficiently computed, easily interpretable, scene-level holistic features. This low-dimensional descriptor combines two features from the literature: crowd collectiveness and crowd conflict, with two newly developed crowd features: mean motion speed and a new formulation of crowd density. Two different anomaly detection approaches are investigated using these features. When only normal training data is available we use a Gaussian Mixture Model (GMM) for outlier detection. When both normal and abnormal training data is available we use a Support Vector Machine (SVM) for binary classification. We evaluate on two crowd behaviour anomaly detection datasets, achieving both state-of-the-art classification performance on the violent-flows dataset as well as better than real-time processing performance (40 frames per second). Mark Marsden, Kevin McGuinness, Suzanne Little, Noel E. O'Connor |
ICIP | 2 |
| 2016 | Quantifying radiographic knee osteoarthritis severity using deep convolutional neural networksabstractThis paper proposes a new approach to automatically quantify the severity of knee osteoarthritis (OA) from radiographs using deep convolutional neural networks (CNN). Clinically, knee OA severity is assessed using Kellgren & Lawrence (KL) grades, a five point scale. Previous work on automatically predicting KL grades from radiograph images were based on training shallow classifiers using a variety of hand engineered features. We demonstrate that classification accuracy can be significantly improved using deep convolutional neural network models pre-trained on ImageNet and fine-tuned on knee OA images. Furthermore, we argue that it is more appropriate to assess the accuracy of automatic knee OA severity predictions using a continuous distance-based evaluation metric like mean squared error than it is to use classification accuracy. This leads to the formulation of the prediction of KL grades as a regression problem and further improves accuracy. Results on a dataset of X-ray images and KL grades from the Osteoarthritis Initiative (OAI) show a sizable improvement over the current state-of-the-art. Joseph Antony, Kevin McGuinness, Noel E. O'Connor, Kieran Moran |
ICPR | 2 |
| 2016 | Smart Stadium for Smarter Living: Enriching the Fan ExperienceabstractRapid urbanization has led to more people residing in cities than ever before, and projections estimate that 64% of the global population will be urban by 2050. Cities are beginning to explore Smart City initiatives to reduce expenses and complexities while increasing efficiency and quality of life for its citizens. To achieve this goal, advances in technology and policies are needed together with rethinking traditional solutions to transportation, safety, sustainability, among other priority areas. We propose the use of a Smart Stadium as a 'living laboratory' to identify, deploy and test Internet of Things technologies and Smart City solutions in an environment small enough to practically trial but large enough to evaluate effectiveness and scalability. The Smart Stadium for Smarter Living initiative brings together Arizona State University, Dublin City University, Intel Corporation, Gaelic Athletic Association, Sun Devil Stadium and Croke Park to explore smart environment solutions. Sethuraman Panchanathan, Shayok Chakraborty, Troy McDaniel, Matt Bunch, Noel E. O'Connor, Suzanne Little, Kevin McGuinness, Mark Marsden |
ISM | 7 |
| 2016 | Bags of Local Convolutional Features for Scalable Instance SearchabstractThis work proposes a simple instance retrieval pipeline based on encoding the convolutional features of CNN using the bag of words aggregation scheme (BoW). Assigning each local array of activations in a convolutional layer to a visual word produces an assignment map, a compact representation that relates regions of an image with a visual word. We use the assignment map for fast spatial reranking, obtaining object localizations that are used for query expansion. We demonstrate the suitability of the BoW representation based on local CNN features for instance retrieval, achieving competitive performance on the Oxford and Paris buildings benchmarks. We show that our proposed system for CNN feature aggregation with BoW outperforms state-of-the-art techniques using sum pooling at a subset of the challenging TRECVid INS benchmark. Eva Mohedano, Kevin McGuinness, Noel E. O'Connor, Amaia Salvador, Ferran Marqués, Xavier Giró-i-Nieto |
ICMR | 2 |
| 2016 | Learning Multiple Views with Orthogonal Denoising Autoencoders
TengQi Ye, Tianchun Wang, Kevin McGuinness, Cathal Gurrin |
MMM (1) | 3 |
| 2015 | Improving spatial codification in semantic segmentationabstractThis paper explores novel approaches for improving the spatial codification for the pooling of local descriptors to solve the semantic segmentation problem. We propose to partition the image into three regions for each object to be described: Figure, Border and Ground. This partition aims at minimizing the influence of the image context on the object description and vice versa by introducing an intermediate zone around the object contour. Furthermore, we also propose a richer visual descriptor of the object by applying a Spatial Pyramid over the Figure region. Two novel Spatial Pyramid configurations are explored: Cartesian-based and crown-based Spatial Pyramids. We test these approaches with state-of-the-art techniques and show that they improve the Figure-Ground based pooling in the Pascal VOC 2011 and 2012 semantic segmentation challenges. Carles Ventura, Xavier Giró-i-Nieto, Verónica Vilaplana, Kevin McGuinness, Ferran Marqués, Noel E. O'Connor |
ICIP | 4 |
| 2015 | Exploring EEG for Object Detection and RetrievalabstractThis paper explores the potential for using Brain Computer Interfaces (BCI) as a relevance feedback mechanism in content-based image retrieval. Several experiments are performed using a rapid serial visual presentation (RSVP) of images at different rates (5Hz and 10Hz) on 8 users with different degrees of familiarization with BCI and the dataset. We compare the feedback from the BCI and mouse-based interfaces in a subset of TRECVid images, finding that, when users have limited time to annotate the images, both interfaces are comparable in performance. Comparing our best users in a retrieval task, we found that EEG-based relevance feedback can outperform mouse-based feedback. Eva Mohedano, Kevin McGuinness, Graham Healy, Noel E. O'Connor, Alan F. Smeaton, Amaia Salvador, Sergi Porta, Xavier Giró-i-Nieto |
ICMR | 2 |
| 2015 | Improving object segmentation by using EEG signals and rapid serial visual presentation
Eva Mohedano, Graham Healy, Kevin McGuinness, Xavier Giró-i-Nieto, Noel E. O'Connor, Alan F. Smeaton |
Multim. Tools Appl. | 3 |
| 2014 | Object Segmentation in Images using EEG SignalsabstractThis paper explores the potential of brain-computer interfaces in segmenting objects from images. Our approach is centered around designing an effective method for displaying the image parts to the users such that they generate measurable brain reactions. When an image region, specifically a block of pixels, is displayed we estimate the probability of the block containing the object of interest using a score based on EEG activity. After several such blocks are displayed, the resulting probability map is binarized and combined with the GrabCut algorithm to segment the image into object and background regions. This study shows that BCI and simple EEG analysis are useful in locating object boundaries in images. Eva Mohedano, Graham Healy, Kevin McGuinness, Xavier Giró-i-Nieto, Noel E. O'Connor, Alan F. Smeaton |
ACM Multimedia | 3 |
| 2014 | Average Precision: Good Guide or False Friend to Multimedia Search Effectiveness?
Robin Aly, Dolf Trieschnigg, Kevin McGuinness, Noel E. O'Connor, Franciska de Jong |
MMM (2) | 3 |
| 2014 | Audio-Visual Classification Video Browser
David Scott, Rami Albatal, Kevin McGuinness, Esra Acar, Frank Hopfgartner, Cathal Gurrin, Noel E. O'Connor, Alan F. Smeaton |
MMM (2) | 4 |
| 2013 | Improved graph cut segmentation by learning a contrast model on the flyabstractThis paper describes an extension to the graph cut interactive image segmentation algorithm based on a novel approach to addressing the well known small cut problem. The approach uses a generative contrast model to weight interaction potentials. The model attempts to capture the expected changes in color between adjacent pixels in the unlabeled area of the image using the adjacent pixels in the user interactions as training data. We compare our approach to the standard graph cuts algorithm and show that the contrast model allows a user to achieve a more accurate segmentation with fewer interactions. We additionally introduce a variant of the approach based on superpixels that further enhances performance but reduces computational complexity to ensure instant feedback for optimal user experience. Kevin McGuinness, Noel E. O'Connor |
ICIP | 1 |
| 2013 | The AXES PRO video search systemabstractWe demonstrate a multimedia content information retrieval engine developed for audiovisual digital libraries targeted at media professionals. It is the first of three multimedia IR systems being developed by the AXES project. The system brings together traditional text IR and state-of-the-art content indexing and retrieval technologies to allow users to search and browse digital libraries in novel ways. Key features include: metadata and ASR search and filtering, on-the-fly visual concept classification (categories, faces, places, and logos), and similarity search (instances and faces). Kevin McGuinness, Noel E. O'Connor, Robin Aly, Franciska de Jong, Ken Chatfield, Omkar M. Parkhi, Relja Arandjelovic, Andrew Zisserman, Matthijs Douze, Cordelia Schmid |
ICMR | 1 |
| 2013 | DCU at MMM 2013 Video Browser Showdown
David Scott, Jinlin Guo, Cathal Gurrin, Frank Hopfgartner, Kevin McGuinness, Noel E. O'Connor, Alan F. Smeaton, Yang Yang 0076 |
MMM (2) | 5 |
| 2012 | Efficient Storage and Decoding of SURF Feature Points
Kevin McGuinness, Kealan McCusker, Neil O'Hare, Noel E. O'Connor |
MMM | 1 |
| 2011 | Toward automated evaluation of interactive segmentation
Kevin McGuinness, Noel E. O'Connor |
Comput. Vis. Image Underst. | 1 |
| 2010 | A comparative evaluation of interactive segmentation algorithms
Kevin McGuinness, Noel E. O'Connor |
Pattern Recognit. | 1 |
| 2008 | Incorporating spatio-temporal mid-level features in a region segmentation algorithm for video sequencesabstractSegmentation algorithms traditionally employ low-level features to divide images into different regions that show a certain degree of homogeneity. However, low-level features, spatial or temporal, are not always reliable when processing real-world video sequences, because of issues like illuminations or complex backgrounds. Furthermore, real world objects can be composed of different regions with heterogeneous features. Although the inclusion of motion can mitigate some of these effects, many problems are still present. This paper proposes the utilization of some spatio-temporal mid-level features that are related, on the one hand, to geometric properties of real objects and, on the other, to well-known motion patterns. Specifically, the proposed algorithm uses a mid-level module that controls the subsequent segmentation using these kinds of features. Some experiments and evaluations show that the inclusion of mid-level features can help to obtain perceptually more meaningful segmentations, thus resulting in regions that are closer to semantic concepts. Iván González-Díaz 0001, Kevin McGuinness, Tomasz Adamek, Noel E. O'Connor, Fernando Díaz-de-María |
ICIP | 2 |