VLDB 2026 Research / reviewers in the wild / expert
Diego Ortego
dblp:136/0273
· DBLP profile ↗
19ranked-venue papers
11as first author
8since 2021 · last 2026
0000-0002-1011-3610ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 7 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 9 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Large Language Models Meet Extreme Multi-label Classification: Scaling and Multi-modal Framework
Diego Ortego, Marlon Rodríguez, Mario Almagro-Cádiz, Kunal Dahiya, David Jiménez, Juan C. SanMiguel |
AAAI | 1 |
| 2025 | Prototypical Extreme Multi-label Classification with a Dynamic Margin LossabstractKunal Dahiya, Diego Ortego, David Jimenez-Cabello. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Kunal Dahiya, Diego Ortego, David Jimenez-Cabello |
NAACL (Long Papers) | 2 |
| 2023 | LEA: Improving Sentence Similarity Robustness to Typos Using Lexical Attention BiasabstractTextual noise, such as typos or abbreviations, is a well-known issue that penalizes vanilla Transformers for most downstream tasks. We show that this is also the case for sentence similarity, a fundamental task in multiple domains, e.g. matching, retrieval or paraphrasing. Sentence similarity can be approached using cross-encoders, where the two sentences are concatenated in the input allowing the model to exploit the inter-relations between them. Previous works addressing the noise issue mainly rely on data augmentation strategies, showing improved robustness when dealing with corrupted samples that are similar to the ones used for training. However, all these methods still suffer from the token distribution shift induced by typos. In this work, we propose to tackle textual noise by equipping cross-encoders with a novel LExical-aware Attention module (LEA) that incorporates lexical similarities between words in both sentences. By using raw text similarities, our approach avoids the tokenization shift problem obtaining improved robustness. We demonstrate that the attention bias introduced by LEA helps cross-encoders to tackle complex scenarios with textual noise, specially in domains with short-text descriptions and limited context. Experiments using three popular Transformer encoders in five e-commerce datasets for product matching show that LEA consistently boosts performance under the presence of noise, while remaining competitive on the original (clean) splits. We also evaluate our approach in two datasets for textual entailment and paraphrasing showing that LEA is robust to typos in domains with longer sentences and more natural context. Additionally, we thoroughly analyze several design choices in our approach, providing insights about the impact of the decisions made and fostering future research in cross-encoders dealing with typos. Mario Almagro-Cádiz, Emilio J. Almazán, Diego Ortego, David Jiménez |
KDD | 3 |
| 2022 | Addressing out-of-distribution label noise in webly-labelled dataabstractA recurring focus of the deep learning community is towards reducing the labeling effort. Data gathering and annotation using a search engine is a simple alternative to generating a fully human-annotated and human-gathered dataset. Although web crawling is very time efficient, some of the retrieved images are unavoidably noisy, i.e. incorrectly labeled. Designing robust algorithms for training on noisy data gathered from the web is an important research perspective that would render the building of datasets easier. In this paper we conduct a study to understand the type of label noise to expect when building a dataset using a search engine. We review the current limitations of state-of-the-art methods for dealing with noisy labels for image classification tasks in the case of web noise distribution. We propose a simple solution to bridge the gap with a fully clean dataset using Dynamic Softening of Out-of-distribution Samples (DSOS), which we design on corrupted versions of the CIFAR-100 dataset, and compare against state-of-the-art algorithms on the web noise perturbated MiniImageNet and Stanford datasets and on real label noise datasets: WebVision 1.0 and Clothing1M. Our work is fully reproducible https://git.io/JKGcj. Paul Albert, Diego Ortego, Eric Arazo Sanchez, Noel E. O'Connor, Kevin McGuinness |
WACV | 2 |
| 2021 | How Important is Importance Sampling for Deep Budgeted Training?
Eric Arazo Sanchez, Diego Ortego, Paul Albert, Noel E. O'Connor, Kevin McGuinness |
BMVC | 2 |
| 2021 | Multi-Objective Interpolation Training for Robustness To Label NoiseabstractDeep neural networks trained with standard cross-entropy loss memorize noisy labels, which degrades their performance. Most research to mitigate this memorization proposes new robust classification loss functions. Conversely, we propose a Multi-Objective Interpolation Training (MOIT) approach that jointly exploits contrastive learning and classification to mutually help each other and boost performance against label noise. We show that standard supervised contrastive learning degrades in the presence of label noise and propose an interpolation training strategy to mitigate this behavior. We further propose a novel label noise detection method that exploits the robust feature representations learned via contrastive learning to estimate per-sample soft-labels whose disagreements with the original labels accurately identify noisy samples. This detection allows treating noisy samples as unlabeled and training a classifier in a semi-supervised manner to prevent noise memorization and improve representation learning. We further propose MOIT+, a refinement of MOIT by fine-tuning on detected clean samples. Hyperparameter and ablation studies verify the key components of our method. Experiments on synthetic and real-world noise benchmarks demonstrate that MOIT/MOIT+ achieves state-of-the-art results. Code is available at https://git.io/JI40X. Diego Ortego, Eric Arazo Sanchez, Paul Albert, Noel E. O'Connor, Kevin McGuinness |
CVPR | 1 |
| 2021 | Unsupervised Contrastive Learning of Sound Event RepresentationsabstractSelf-supervised representation learning can mitigate the limitations in recognition tasks with few manually labeled data but abundant unlabeled data—a common scenario in sound event research. In this work, we explore unsupervised contrastive learning as a way to learn sound event representations. To this end, we propose to use the pretext task of contrasting differently augmented views of sound events. The views are computed primarily via mixing of training examples with unrelated backgrounds, followed by other data augmentations. We analyze the main components of our method via ablation experiments. We evaluate the learned representations using linear evaluation, and in two in-domain downstream sound event classification tasks, namely, using limited manually labeled data, and using noisy labeled data. Our results suggest that unsupervised contrastive pre-training can mitigate the impact of data scarcity and increase robustness against noisy labels. Eduardo Fonseca, Diego Ortego, Kevin McGuinness, Noel E. O'Connor, Xavier Serra |
ICASSP | 2 |
| 2021 | ReLaB: Reliable Label Bootstrapping for Semi-Supervised LearningabstractReducing the amount of labels required to train convolutional neural networks without performance degradation is key to effectively reduce human annotation efforts. We propose Reliable Label Bootstrapping (ReLaB), an unsupervised preprossessing algorithm which improves the performance of semi-supervised algorithms in extremely low supervision settings. Given a dataset with few labeled samples, we first learn meaningful self-supervised, latent features for the data. Second, a label propagation algorithm propagates the known labels on the unsupervised features, effectively labeling the full dataset in an automatic fashion. Third, we select a subset of correctly labeled (reliable) samples using a label noise detection algorithm. Finally, we train a semi-supervised algorithm on the extended subset. We show that the selection of the network architecture and the self-supervised algorithm are important factors to achieve successful label propagation and demonstrate that ReLaB substantially improves semi-supervised learning in scenarios of very limited supervision on image classification benchmarks such as CIFAR-10, CIFAR-100 and mini-ImageNet. We reach average error rates of 22.34 with 1 random labeled sample per class on CIFAR-10 and lower this error to 8.46 when the labeled sample in each class is highly representative. Our work is fully reproducible: https://github.com/PaulAlbert31/ReLaB. Paul Albert, Diego Ortego, Eric Arazo Sanchez, Noel E. O'Connor, Kevin McGuinness |
IJCNN | 2 |
| 2020 | Towards Robust Learning with Different Label Noise DistributionsabstractNoisy labels are an unavoidable consequence of labeling processes and detecting them is an important step towards preventing performance degradations in Convolutional Neural Networks. Discarding noisy labels avoids a harmful memorization, while the associated image content can still be exploited in a semi-supervised learning (SSL) setup. Clean samples are usually identified using the small loss trick, i.e. they exhibit a low loss. However, we show that different noise distributions make the application of this trick less straightforward and propose to continuously relabel all images to reveal a discriminative loss against multiple distributions. SSL is then applied twice, once to improve the clean-noisy detection and again for training the final model. We design an experimental setup based on ImageNet32/64 for better understanding the consequences of representation learning with differing label noise distributions and find that non-uniform out-of-distribution noise better resembles real-world noise and that in most cases intermediate features are not affected by label noise corruption. Experiments in CIFAR-10/100, ImageNet32/64 and WebVision (real-world noise) demonstrate that the proposed label noise Distribution Robust Pseudo-Labeling (DRPL) approach gives substantial improvements over recent state-of-the-art. Code is available at https://git.io/JJ0PV. Diego Ortego, Eric Arazo Sanchez, Paul Albert, Noel E. O'Connor, Kevin McGuinness |
ICPR | 1 |
| 2020 | Pseudo-Labeling and Confirmation Bias in Deep Semi-Supervised LearningabstractSemi-supervised learning, i.e. jointly learning from labeled and unlabeled samples, is an active research topic due to its key role on relaxing human supervision. In the context of image classification, recent advances to learn from unlabeled samples are mainly focused on consistency regularization methods that encourage invariant predictions for different perturbations of unlabeled samples. We, conversely, propose to learn from unlabeled data by generating soft pseudo-labels using the network predictions. We show that a naive pseudo-labeling overfits to incorrect pseudo-labels due to the so-called confirmation bias and demonstrate that mixup augmentation and setting a minimum number of labeled samples per mini-batch are effective regularization techniques for reducing it. The proposed approach achieves state-of-the-art results in CIFAR-10/100, SVHN, and Mini-ImageNet despite being much simpler than other methods. These results demonstrate that pseudo-labeling alone can outperform consistency regularization methods, while the opposite was supposed in previous work. Source code is available at https://git.io/fjQsC. Eric Arazo Sanchez, Diego Ortego, Paul Albert, Noel E. O'Connor, Kevin McGuinness |
IJCNN | 2 |
| 2019 | On guiding video object segmentationabstractThis paper presents a novel approach for segmenting moving objects in unconstrained environments using guided convolutional neural networks. This guiding process relies on foreground masks from independent algorithms (i.e. state-of-the-art algorithms) to implement an attention mechanism that incorporates the spatial location of foreground and background to compute their separated representations. Our approach initially extracts two kinds of features for each frame using colour and optical flow information. Such features are combined following a multiplicative scheme to benefit from their complementarity. These unified colour and motion features are later processed to obtain the separated foreground and background representations. Then, both independent representations are concatenated and decoded to perform foreground segmentation. Experiments conducted on the challenging DAVIS 2016 dataset demonstrate that our guided representations not only outperform non-guided, but also recent and top-performing video object segmentation algorithms. Diego Ortego, Kevin McGuinness, Juan C. SanMiguel, Eric Arazo Sanchez, José María Martínez Sanchez, Noel E. O'Connor |
CBMI | 1 |
| 2019 | Unsupervised Label Noise Modeling and Loss CorrectionabstractDespite being robust to small amounts of label noise, convolutional neural networks trained with stochastic gradient methods have been shown to easily fit random labels. When there are a mixture of correct and mislabelled targets, networks tend to fit the former before the latter. This suggests using a suitable two-component mixture model as an unsupervised generative model of sample loss values during training to allow online estimation of the probability that a sample is mislabelled. Specifically, we propose a beta mixture to estimate this probability and correct the loss by relying on the network prediction (the so-called bootstrapping loss). We further adapt mixup augmentation to drive our approach a step further. Experiments on CIFAR-10/100 and TinyImageNet demonstrate a robustness to label noise that substantially outperforms recent state-of-the-art. Source code is available at https://git.io/fjsvE and Appendix at https://arxiv.org/abs/1904.11238. Eric Arazo Sanchez, Diego Ortego, Paul Albert, Noel E. O'Connor, Kevin McGuinness |
ICML | 2 |
| 2019 | Hierarchical Improvement of Foreground Segmentation Masks in Background SubtractionabstractA plethora of algorithms have been defined for foreground segmentation, a fundamental stage for many computer vision applications. In this paper, we propose a post-processing framework to improve the foreground segmentation performance of background subtraction algorithms. We define a hierarchical framework for extending segmented foreground pixels to undetected foreground object areas and for removing erroneously segmented foreground. First, we create a motion-aware hierarchical image segmentation of each frame that prevents merging foreground and background image regions. Then, we estimate the quality of the foreground mask through the fitness of the binary regions in the mask and the hierarchy of segmented regions. Finally, the improved foreground mask is obtained as an optimal labeling by jointly exploiting foreground quality and spatial color relations in a pixel-wise fully connected conditional random field. Experiments are conducted over four large and heterogeneous data sets with varied challenges (CDNET2014, LASIESTA, SABS, and BMC) demonstrating the capability of the proposed framework to improve background subtraction results. Diego Ortego, Juan C. SanMiguel, José María Martínez Sanchez |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2017 | Stand-alone quality estimation of background subtraction algorithms
Diego Ortego, Juan C. SanMiguel, José María Martínez Sanchez |
Comput. Vis. Image Underst. | 1 |
| 2016 | Rejection based multipath reconstruction for background estimation in SBMnet 2016 datasetabstractBackground Estimation in video consists in extracting a foreground-free image from a set of training frames. In this paper, we overview a temporal-spatial block-level approach for background estimation in video and present their results in the SBMnet dataset. First, the employed approach uses a Temporal Analysis module to obtain a compact representation of the training data that is later clustered by a threshold-free technique to generate background candidates at each block location. Then, a Spatial Analysis module iteratively reconstructs the background using a multipath reconstruction guided by background smoothness constraints. The experimental results in the SBMnet dataset demonstrates the utility of the employed approach against stationary objects and its weaknesses when motion information is involved. Diego Ortego, Juan C. SanMiguel, José María Martínez Sanchez |
ICPR | 1 |
| 2016 | Rejection based multipath reconstruction for background estimation in video sequences with stationary objectsabstractBackground estimation in video consists in extracting a foreground-free image from a set of training frames. Moving and stationary objects may affect the background visibility, thus invalidating the assumption of many related literature where background is the temporal dominant data. In this paper, we present a temporal-spatial block-level approach for background estimation in video to cope with moving and stationary objects. First, a Temporal Analysis module obtains a compact representation of the training data by motion filtering and dimensionality reduction. Then, a threshold-free hierarchical clustering determines a set of candidates to represent the background for each spatial location (block). Second, a Spatial Analysis module iteratively reconstructs the background using these candidates. For each spatial location , multiple reconstruction hypotheses (paths) are explored to obtain its neighboring locations by enforcing inter-block similarities and intra-block homogeneity constraints in terms of color discontinuity, color dissimilarity and variability. The experimental results show that the proposed approach outperforms the related state-of-the-art over challenging video sequences in presence of moving and stationary objects. Diego Ortego, Juan C. SanMiguel, José María Martínez Sanchez |
Comput. Vis. Image Underst. | 1 |
| 2015 | Long-Term Stationary Object Detection Based on Spatio-Temporal Change DetectionabstractWe present a block-wise approach to detect stationary objects based on spatio-temporal change detection. First, block candidates are extracted by filtering out consecutive blocks containing moving objects. Then, an online clustering approach groups similar blocks at each spatial location over time via statistical variation of pixel ratios. The stability changes are identified by analyzing the relationships between the most repeated clusters at regular sampling instants. Finally, stationary objects are detected as those stability changes that exceed an alarm time and have not been visualized before. Unlike previous approaches making use of Background Subtraction, the proposed approach does not require foreground segmentation and provides robustness to illumination changes, crowds and intermittent object motion. The experiments over an heterogeneous dataset demonstrate the ability of the proposed approach for short- and long-term operation while overcoming challenging issues. Diego Ortego, Juan C. SanMiguel, José María Martínez Sanchez |
IEEE Signal Process. Lett. | 1 |
| 2014 | Multi-feature stationary foreground detection for crowded video-surveillanceabstractWe propose a novel approach for stationary foreground detection in crowds based on the spatio-temporal evolution of multiple features. A generic framework is presented to detect stationarity where history images model the spatio-temporal feature patterns. A feature is proposed based on structural information over each pixel neighborhood for dealing with shadows and illumination changes. A multifeature detector is composed by combining the history images of three features (namely, foreground, motion and structural information) to estimate the foreground stationarity over time, which is later thresholded to detect stationary regions. Experimental results over challenging video-surveillance sequences show the improvement of the proposed approach against related work as structural information reduces false detections, which are common in crowded places. Diego Ortego, Juan C. SanMiguel |
ICIP | 1 |
| 2013 | Stationary foreground detection for video-surveillance based on foreground and motion history imagesabstractStationary foreground detection is a common stage in many video-surveillance applications. In this paper, we propose an approach for stationary foreground detection in video based on the spatio-temporal variation of foreground and motion data. Foreground data are obtained by Background Subtraction to detect regions of interest. Motion data allows to filter out the moving regions and it is estimated using median filters over sliding windows. Spatiotemporal patterns of both data are computed through history images and the final detection is obtained using a two-threshold scheme that considers motion activity. Partial visibility of stationary foreground for short-time intervals is handled to increase robustness. The results over challenging video-surveillance sequences show an improvement of the proposed approach against the related work. Diego Ortego, Juan C. SanMiguel |
AVSS | 1 |