VLDB 2026 Research / reviewers in the wild / expert
Jose Caballero
dblp:118/9828
· DBLP profile ↗
15ranked-venue papers
5as first author
0since 2021 · last 2019
0000-0003-3285-8928ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-authorArtificial intelligence and machine learning · 4 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
3 papers |
Image and video processing · 85% Image and video coding · 15% | |
| Artificial intelligence
4 papers |
Generative modeling · 65% Deep learning architectures and training · 19% Probabilistic and Bayesian machine learning · 16% |
Topics — the 13 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training
convolutional neural network |
0.3 | 2 | 2017 | Real-Time Single Image and Video Super-Resolution Using an Efficient Sub-Pixel Convolutional Neural Network · CVPR 2016 Real-Time Video Super-Resolution with Spatio-Temporal Networks and Motion Compensation · CVPR 2017 |
Machine learning › Generative modeling
diffusion model |
0.3 | 1 | 2017 | Amortised MAP Inference for Image Super-resolution · ICLR 2017 |
Machine learning › Generative modeling
generative adversarial network |
0.3 | 1 | 2017 | Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network · CVPR 2017 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
MAP inference |
0.3 | 1 | 2017 | Amortised MAP Inference for Image Super-resolution · ICLR 2017 |
Machine learning › Generative modeling › image reconstruction
super-resolution |
0.3 | 1 | 2017 | Amortised MAP Inference for Image Super-resolution · ICLR 2017 |
Machine learning › Generative modeling › generative adversarial network › image-to-image translation
super-resolution GAN |
0.3 | 1 | 2017 | Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network · CVPR 2017 |
Image and video processing › super-resolution
image super-resolution |
0.3 | 1 | 2017 | Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network · CVPR 2017 |
Image and video coding › video compression
motion compensation |
0.3 | 1 | 2017 | Real-Time Video Super-Resolution with Spatio-Temporal Networks and Motion Compensation · CVPR 2017 |
Image and video processing
perceptual loss |
0.3 | 1 | 2017 | Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network · CVPR 2017 |
Image and video processing › super-resolution › image super-resolution
perceptual super-resolution |
0.3 | 1 | 2017 | Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network · CVPR 2017 |
Image and video processing › super-resolution
video super-resolution |
0.3 | 1 | 2017 | Real-Time Video Super-Resolution with Spatio-Temporal Networks and Motion Compensation · CVPR 2017 |
Image and video processing › super-resolution › image super-resolution
single image super-resolution |
0.2 | 1 | 2016 | Real-Time Single Image and Video Super-Resolution Using an Efficient Sub-Pixel Convolutional Neural Network · CVPR 2016 |
Image and video processing
super-resolution |
0.2 | 1 | 2016 | Real-Time Single Image and Video Super-Resolution Using an Efficient Sub-Pixel Convolutional Neural Network · CVPR 2016 |
Methods — techniques the papers use, named apart from their topics
sub-pixel convolution · 0.6spatial transformer · 0.6residual network · 0.6content loss · 0.6adversarial loss · 0.63d convolution · 0.6upscaling filters · 0.5sub-pixel convolution layer · 0.5variational inference · 0.3amortized inference · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | Convolutional Recurrent Neural Networks for Dynamic MR Image ReconstructionabstractAccelerating the data acquisition of dynamic magnetic resonance imaging leads to a challenging ill-posed inverse problem, which has received great interest from both the signal processing and machine learning communities over the last decades. The key ingredient to the problem is how to exploit the temporal correlations of the MR sequence to resolve aliasing artifacts. Traditionally, such observation led to a formulation of an optimization problem, which was solved using iterative algorithms. Recently, however, deep learning-based approaches have gained significant popularity due to their ability to solve general inverse problems. In this paper, we propose a unique, novel convolutional recurrent neural network architecture which reconstructs high quality cardiac MR images from highly undersampled k-space data by jointly exploiting the dependencies of the temporal sequences as well as the iterative nature of the traditional optimization algorithms. In particular, the proposed architecture embeds the structure of the traditional iterative algorithms, efficiently modeling the recurrence of the iterative reconstruction stages by using recurrent hidden connections over such iterations. In addition, spatio-temporal dependencies are simultaneously learnt by exploiting bidirectional recurrent hidden connections across time sequences. The proposed method is able to learn both the temporal dependence and the iterative reconstruction process effectively with only a very small number of parameters, while outperforming current MR reconstruction methods in terms of reconstruction accuracy and speed. Chen Qin, Jo Schlemper, Jose Caballero, Anthony N. Price, Joseph V. Hajnal, Daniel Rueckert |
IEEE Trans. Medical Imaging | 3 |
| 2018 | Anatomically Constrained Neural Networks (ACNNs): Application to Cardiac Image Enhancement and SegmentationabstractIncorporation of prior knowledge about organ shape and location is key to improve performance of image analysis approaches. In particular, priors can be useful in cases where images are corrupted and contain artefacts due to limitations in image acquisition. The highly constrained nature of anatomical objects can be well captured with learning-based techniques. However, in most recent and promising techniques such as CNN-based segmentation it is not obvious how to incorporate such prior knowledge. State-of-the-art methods operate as pixel-wise classifiers where the training objectives do not incorporate the structure and inter-dependencies of the output. To overcome this limitation, we propose a generic training strategy that incorporates anatomical prior knowledge into CNNs through a new regularisation model, which is trained end-to-end. The new framework encourages models to follow the global anatomical properties of the underlying anatomy (e.g. shape, label structure) via learnt non-linear representations of the shape. We show that the proposed approach can be easily adapted to different analysis tasks (e.g. image enhancement, segmentation) and improve the prediction accuracy of the state-of-the-art models. The applicability of our approach is shown on multi-modal cardiac data sets and public benchmarks. In addition, we demonstrate how the learnt deep models of 3-D shapes can be interpreted and used as biomarkers for classification of cardiac pathologies. Ozan Oktay, Enzo Ferrante, Konstantinos Kamnitsas, Mattias P. Heinrich, Wenjia Bai, Jose Caballero, Stuart A. Cook, Antonio M. Simoes Monteiro de Marvao, Timothy Dawes, Declan P. O'Regan, Bernhard Kainz, Ben Glocker, Daniel Rueckert |
IEEE Trans. Medical Imaging | 6 |
| 2018 | A Deep Cascade of Convolutional Neural Networks for Dynamic MR Image ReconstructionabstractInspired by recent advances in deep learning, we propose a framework for reconstructing dynamic sequences of 2-D cardiac magnetic resonance (MR) images from undersampled data using a deep cascade of convolutional neural networks (CNNs) to accelerate the data acquisition process. In particular, we address the case where data are acquired using aggressive Cartesian undersampling. First, we show that when each 2-D image frame is reconstructed independently, the proposed method outperforms state-of-the-art 2-D compressed sensing approaches, such as dictionary learning-based MR image reconstruction, in terms of reconstruction error and reconstruction speed. Second, when reconstructing the frames of the sequences jointly, we demonstrate that CNNs can learn spatio-temporal correlations efficiently by combining convolution and data sharing approaches. We show that the proposed method consistently outperforms state-of-the-art methods and is capable of preserving anatomical structure more faithfully up to 11-fold undersampling. Moreover, reconstruction is very fast: each complete dynamic sequence can be reconstructed in less than 10 s and, for the 2-D case, each image frame can be reconstructed in 23 ms, enabling real-time applications. Jo Schlemper, Jose Caballero, Joseph V. Hajnal, Anthony N. Price, Daniel Rueckert |
IEEE Trans. Medical Imaging | 2 |
| 2017 | Real-Time Video Super-Resolution with Spatio-Temporal Networks and Motion CompensationabstractConvolutional neural networks have enabled accurate image super-resolution in real-time. However, recent attempts to benefit from temporal correlations in video super-resolution have been limited to naive or inefficient architectures. In this paper, we introduce spatio-temporal sub-pixel convolution networks that effectively exploit temporal redundancies and improve reconstruction accuracy while maintaining real-time speed. Specifically, we discuss the use of early fusion, slow fusion and 3D convolutions for the joint processing of multiple consecutive video frames. We also propose a novel joint motion compensation and video super-resolution algorithm that is orders of magnitude more efficient than competing methods, relying on a fast multi-resolution spatial transformer module that is end-to-end trainable. These contributions provide both higher accuracy and temporally more consistent videos, which we confirm qualitatively and quantitatively. Relative to single-frame models, spatio-temporal networks can either reduce the computational cost by 30% whilst maintaining the same quality or provide a 0.2dB gain for a similar computational cost. Results on publicly available datasets demonstrate that the proposed algorithms surpass current state-of-the-art performance in both accuracy and efficiency. Jose Caballero, Christian Ledig, Andrew P. Aitken, Alejandro Acosta, Johannes Totz, Wenzhe Shi |
CVPR | 1 |
| 2017 | Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial NetworkabstractDespite the breakthroughs in accuracy and speed of single image super-resolution using faster and deeper convolutional neural networks, one central problem remains largely unsolved: how do we recover the finer texture details when we super-resolve at large upscaling factors? The behavior of optimization-based super-resolution methods is principally driven by the choice of the objective function. Recent work has largely focused on minimizing the mean squared reconstruction error. The resulting estimates have high peak signal-to-noise ratios, but they are often lacking high-frequency details and are perceptually unsatisfying in the sense that they fail to match the fidelity expected at the higher resolution. In this paper, we present SRGAN, a generative adversarial network (GAN) for image super-resolution (SR). To our knowledge, it is the first framework capable of inferring photo-realistic natural images for 4x upscaling factors. To achieve this, we propose a perceptual loss function which consists of an adversarial loss and a content loss. The adversarial loss pushes our solution to the natural image manifold using a discriminator network that is trained to differentiate between the super-resolved images and original photo-realistic images. In addition, we use a content loss motivated by perceptual similarity instead of similarity in pixel space. Our deep residual network is able to recover photo-realistic textures from heavily downsampled images on public benchmarks. An extensive mean-opinion-score (MOS) test shows hugely significant gains in perceptual quality using SRGAN. The MOS scores obtained with SRGAN are closer to those of the original high-resolution images than to those obtained with any state-of-the-art method. Christian Ledig, Lucas Theis, Ferenc Huszar, Jose Caballero, Andrew Cunningham, Alejandro Acosta, Andrew P. Aitken, Alykhan Tejani, Johannes Totz, Wenzhe Shi |
CVPR | 4 |
| 2017 | Amortised MAP Inference for Image Super-resolution
Casper Kaae Sønderby, Jose Caballero, Lucas Theis, Wenzhe Shi, Ferenc Huszar |
ICLR | 2 |
| 2016 | Real-Time Single Image and Video Super-Resolution Using an Efficient Sub-Pixel Convolutional Neural NetworkabstractRecently, several models based on deep neural networks have achieved great success in terms of both reconstruction accuracy and computational performance for single image super-resolution. In these methods, the low resolution (LR) input image is upscaled to the high resolution (HR) space using a single filter, commonly bicubic interpolation, before reconstruction. This means that the super-resolution (SR) operation is performed in HR space. We demonstrate that this is sub-optimal and adds computational complexity. In this paper, we present the first convolutional neural network (CNN) capable of real-time SR of 1080p videos on a single K2 GPU. To achieve this, we propose a novel CNN architecture where the feature maps are extracted in the LR space. In addition, we introduce an efficient sub-pixel convolution layer which learns an array of upscaling filters to upscale the final LR feature maps into the HR output. By doing so, we effectively replace the handcrafted bicubic filter in the SR pipeline with more complex upscaling filters specifically trained for each feature map, whilst also reducing the computational complexity of the overall SR operation. We evaluate the proposed approach using images and videos from publicly available datasets and show that it performs significantly better (+0.15dB on Images and +0.39dB on Videos) and is an order of magnitude faster than previous CNN-based methods. Wenzhe Shi, Jose Caballero, Ferenc Huszar, Johannes Totz, Andrew P. Aitken, Rob Bishop, Daniel Rueckert |
CVPR | 2 |
| 2016 | Multi-input Cardiac Image Super-Resolution Using Convolutional Neural Networksabstract3D cardiac MR imaging enables accurate analysis of cardiac morphology and physiology. However, due to the requirements for long acquisition and breath-hold, the clinical routine is still dominated by multi-slice 2D imaging, which hamper the visualization of anatomy and quantitative measurements as relatively thick slices are acquired. As a solution, we propose a novel image super-resolution (SR) approach that is based on a residual convolutional neural network (CNN) model. It reconstructs high resolution 3D volumes from 2D image stacks for more accurate image analysis. The proposed model allows the use of multiple input data acquired from different viewing planes for improved performance. Experimental results on 1233 cardiac short and long-axis MR image stacks show that the CNN model outperforms state-of-the-art SR methods in terms of image quality while being computationally efficient. Also, we show that image segmentation and motion tracking benefits more from SR-CNN when it is used as an initial upscaling method than conventional interpolation methods for the subsequent analysis. Ozan Oktay, Wenjia Bai, Matthew C. H. Lee, Ricardo Guerrero, Konstantinos Kamnitsas, Jose Caballero, Antonio M. Simoes Monteiro de Marvao, Stuart A. Cook, Declan P. O'Regan, Daniel Rueckert |
MICCAI (3) | 6 |
| 2016 | Standardized Evaluation System for Left Ventricular Segmentation Algorithms in 3D EchocardiographyabstractReal-time 3D Echocardiography (RT3DE) has been proven to be an accurate tool for left ventricular (LV) volume assessment. However, identification of the LV endocardium remains a challenging task, mainly because of the low tissue/blood contrast of the images combined with typical artifacts. Several semi and fully automatic algorithms have been proposed for segmenting the endocardium in RT3DE data in order to extract relevant clinical indices, but a systematic and fair comparison between such methods has so far been impossible due to the lack of a publicly available common database. Here, we introduce a standardized evaluation framework to reliably evaluate and compare the performance of the algorithms developed to segment the LV border in RT3DE. A database consisting of 45 multivendor cardiac ultrasound recordings acquired at different centers with corresponding reference measurements from three experts are made available. The algorithms from nine research groups were quantitatively evaluated and compared using the proposed online platform. The results showed that the best methods produce promising results with respect to the experts' measurements for the extraction of clinical indices, and that they offer good segmentation precision in terms of mean distance error in the context of the experts' variability range. The platform remains open for new submissions. Olivier Bernard 0001, Johan G. Bosch, Brecht Heyde, Martino Alessandrini, Daniel Barbosa 0001, Sorina Camarasu-Pop, Frederic Cervenansky, Sébastien Valette, Oana Mirea, Michaël Bernier, Pierre-Marc Jodoin, Jaime Santo Domingos, Richard V. Stebbing, Kevin Keraudren, Ozan Oktay, Jose Caballero, Daniel Rueckert, Fausto Milletari, Seyed-Ahmad Ahmadi, Erik Smistad, Frank Lindseth, Maartje van Stralen, Örjan Smedby, Erwan Donal, Mark Monaghan, Alex Papachristidis, Marcel L. Geleijnse, Elena Galli, Jan D'hooge |
IEEE Trans. Medical Imaging | 16 |
| 2015 | Fast Reconstruction of Accelerated Dynamic MRI Using Manifold Kernel Regression
Kanwal K. Bhatia, Jose Caballero, Anthony N. Price, Ying Sun 0001, Joseph V. Hajnal, Daniel Rueckert |
MICCAI (3) | 2 |
| 2014 | Application-Driven MRI: Joint Reconstruction and Segmentation from Undersampled MRI Data
Jose Caballero, Wenjia Bai, Anthony N. Price, Daniel Rueckert, Joseph V. Hajnal |
MICCAI (1) | 1 |
| 2014 | Dictionary Learning and Time Sparsity for Dynamic MR Data ReconstructionabstractThe reconstruction of dynamic magnetic resonance data from an undersampled k-space has been shown to have a huge potential in accelerating the acquisition process of this imaging modality. With the introduction of compressed sensing (CS) theory, solutions for undersampled data have arisen which reconstruct images consistent with the acquired samples and compliant with a sparsity model in some transform domain. Fixed basis transforms have been extensively used as sparsifying transforms in the past, but recent developments in dictionary learning (DL) have been shown to outperform them by training an overcomplete basis that is optimal for a particular dataset. We present here an iterative algorithm that enables the application of DL for the reconstruction of cardiac cine data with Cartesian undersampling. This is achieved with local processing of spatio-temporal 3D patches and by independent treatment of the real and imaginary parts of the dataset. The enforcement of temporal gradients is also proposed as an additional constraint that can greatly accelerate the convergence rate and improve the reconstruction for high acceleration rates. The method is compared to and shown to systematically outperform k- t FOCUSS, a successful CS method that uses a fixed basis transform. Jose Caballero, Anthony N. Price, Daniel Rueckert, Joseph V. Hajnal |
IEEE Trans. Medical Imaging | 1 |
| 2013 | Cardiac Image Super-Resolution with Global Correspondence Using Multi-Atlas PatchMatch
Wenzhe Shi, Jose Caballero, Christian Ledig, Xiahai Zhuang, Wenjia Bai, Kanwal K. Bhatia, Antonio M. Simoes Monteiro de Marvao, Timothy Dawes, Declan P. O'Regan, Daniel Rueckert |
MICCAI (3) | 2 |
| 2012 | Spike sorting at sub-Nyquist ratesabstractSpike sorting relies on the ability to establish the temporal occurrence of action potentials and their relation to specific neurons. Neural information is intrinsically compressible and as such suitable for sparse sampling. Potentially, this should allow for the use of multi-channel recordings, which is particularly advantageous to improve spike sorting. In this paper we propose a novel algorithm capable of sampling neural data at sub-Nyquist rates, yielding the same performance for spike sorting as traditional schemes. Jose Caballero, Jose Antonio Uriguen, Simon R. Schultz, Pier Luigi Dragotti |
ICASSP | 1 |
| 2012 | Dictionary Learning and Time Sparsity in Dynamic MRI
Jose Caballero, Daniel Rueckert, Joseph V. Hajnal |
MICCAI (1) | 1 |