Santiago Lopez Tapia

dblp:191/0937 · also Santiago López-Tapia · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
6since 2021 · last 2025
0000-0003-2090-7446ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 first-author
YearPublicationVenuePosition
2025 Advancing Limited-Angle CT Reconstruction through Diffusion-Based Sinogram Completion
abstract
Limited Angle Computed Tomography (LACT) often faces significant challenges due to missing angular information. Unlike previous methods that operate in the image domain, we propose a new method that focuses on sinogram inpainting. We leverage MR-SDEs, a variant of diffusion models that characterize the diffusion process with mean-reverting stochastic differential equations, to fill in missing angular data at the projection level. Furthermore, by combining distillation with constraining the output of the model using the pseudo-inverse of the inpainting matrix, the diffusion process is accelerated and done in a step, enabling efficient and accurate sinogram completion. A subsequent post-processing module back-projects the inpainted sinogram into the image domain and further refines the reconstruction, effectively suppressing artifacts while preserving critical structural details. Quantitative experimental results demonstrate that the proposed method achieves state-of-the-art performance in both perceptual and fidelity quality, offering a promising solution for LACT reconstruction in scientific and clinical applications.
Santiago Lopez Tapia, Aggelos K. Katsaggelos
ICIP2
2024 Real-World Atmospheric Turbulence Correction Via Domain Adaptation
abstract
Atmospheric turbulence, a common phenomenon in daily life, is primarily caused by the uneven heating of the Earth’s surface. This phenomenon results in distorted and blurred acquired images or videos and can significantly impact downstream vision tasks, particularly those that rely on capturing clear, stable images or videos from outdoor environments, such as accurately detecting or recognizing objects. Therefore, people have proposed ways to simulate atmospheric turbulence and designed effective deep learning-based methods to remove the atmospheric turbulence effect. However, these synthesized turbulent images can not cover all the range of real-world turbulence effects. Though the models have achieved great performance for synthetic scenarios, there always exists a performance drop when applied to real-world cases. Moreover, reducing real-world turbulence is a more challenging task as there are no clean ground truth counter-parts provided to the models during training. In this paper, we propose a real-world atmospheric turbulence mitigation model under a domain adaptation framework, which links the supervised simulated atmospheric turbulence correction with the unsupervised real-world atmospheric turbulence correction. We will show our proposed method enhances performance in real-world atmospheric turbulence scenarios, improving both image quality and downstream vision tasks.
Xijun Wang 0003, Santiago Lopez Tapia, Aggelos K. Katsaggelos
ICIP2
2024 Real-Time Lightweight Video Super-Resolution With RRED-Based Perceptual Constraint
abstract
Real-time video services are gaining popularity in our daily life, yet limited network bandwidth can constrain the delivered video quality. Video Super Resolution (VSR) technology emerges as a key solution to enhance user experience by reconstructing high-resolution (HR) videos. The existing real-time VSR frameworks have primarily emphasized spatial quality metrics like PSNR and SSIM, which often lack consideration of temporal coherence, a critical factor for accurately reflecting the overall quality of super-resolved videos. Inspired by Video Quality Assessment (VQA) strategies, we propose a dual-frame training framework and a lightweight multi-branch network to address VSR processing in real time. Such designs thoroughly leverage the spatio-temporal correlations between consecutive frames so as to ensure efficient video restoration. Furthermore, we incorporate ST-RRED, a powerful VQA approach that separately measures spatial and temporal consistency aligning with human perception principles, into our loss functions. This guides us to synthesize quality-aware perceptual features across both space and time for realistic reconstruction. Our model demonstrates remarkable efficiency, achieving near real-time processing of 4K videos. Compared to the state-of-the-art lightweight model MRVSR, ours is more compact and faster, 60% smaller in size (0.483M vs. 1.21M parameters), and 106% quicker (96.44fps vs. 46.7fps on 1080p frames), with significantly improved perceptual quality.
Xinyi Wu 0001, Santiago Lopez Tapia, Xijun Wang 0003, Rafael Molina 0001, Aggelos K. Katsaggelos
IEEE Trans. Circuits Syst. Video Technol.2
2023 Deep Robust Image Restoration Using the Moore-Penrose Blur Inverse
abstract
This paper proposes a deep learning model for robust image restoration when the degradation is not precisely known. We show how the Moore-Penrose pseudo-inverse of a blur convolution operator can be approximated by a Wiener filter’s impulse response. The image restoration problem is then cast as the learning of a residual on the frequencies where the blurring filter is zero which, when added to the Wiener restoration, will satisfy the image formation model. A Dynamic Filter Network removes artifacts introduced by inaccurate blur estimations and other image formation model inconsistencies. The experiments conducted on synthetic and real image datasets assert the performance and robustness of the proposed method and show its superiority to existing ones.
Santiago Lopez Tapia, Javier Mateos, Rafael Molina 0001, Aggelos K. Katsaggelos
ICIP1
2023 Variational Deep Atmospheric Turbulence Correction for Video
abstract
This paper presents a novel variational deep-learning approach for video atmospheric turbulence correction. We modify and tailor a Nonlinear Activation Free Network to video restoration. By including it in a variational inference framework, we boost the model’s performance and stability. This is achieved through conditioning the model on features extracted by a variational autoencoder (VAE). Furthermore, we enhance these features by making the encoder of the VAE include information pertinent to the image formation via a new loss based on the prediction of parameters of the geometrical distortion and the spatially variant blur responsible for the video sequence degradation. Experiments on a comprehensive synthetic video dataset demonstrate the effectiveness and reliability of the proposed method and validate its superiority compared to existing state-of-the-art approaches.
Santiago Lopez Tapia, Xijun Wang 0003, Aggelos K. Katsaggelos
ICIP1
2021 Fast and Robust Cascade Model for Multiple Degradation Single Image Super-Resolution
abstract
Single Image Super-Resolution (SISR) is one of the low-level computer vision problems that has received increased attention in the last few years. Current approaches are primarily based on harnessing the power of deep learning models and optimization techniques to reverse the degradation model. Owing to its hardness, isotropic blurring or Gaussians with small anisotropic deformations have been mainly considered. Here, we widen this scenario by including large non-Gaussian blurs that arise in real camera movements. Our approach leverages the degradation model and proposes a new formulation of the Convolutional Neural Network (CNN) cascade model, where each network sub-module is constrained to solve a specific degradation: deblurring or upsampling. A new densely connected CNN-architecture is proposed where the output of each sub-module is restricted using some external knowledge to focus it on its specific task. As far we know, this use of domain-knowledge to module-level is a novelty in SISR. To fit the finest model, a final sub-module takes care of the residual errors propagated by the previous sub-modules. We check our model with three state-of-the-art (SOTA) datasets in SISR and compare the results with the SOTA models. The results show that our model is the only one able to manage our wider set of deformations. Furthermore, our model overcomes all current SOTA methods for a standard set of deformations. In terms of computational load, our model also improves on the two closest competitors in terms of efficiency. Although the approach is non-blind and requires an estimation of the blur kernel, it shows robustness to blur kernel estimation errors, making it a good alternative to blind models.
Santiago Lopez Tapia, Nicolas Pérez de la Blanca
IEEE Trans. Image Process.1
2019 Spatially Adaptive Losses for Video Super-resolution with GANs
abstract
Deep Learning techniques and more specifically Generative Adversarial Networks (GANs) have recently been used for solving the video super-resolution (VSR) problem. In some of the published works, feature-based perceptual losses have also been used, resulting in promising results. While there has been work in the literature incorporating temporal information into the loss function, studies which make use of the spatial activity to improve GAN models are still lacking. Towards this end, this paper aims to train a GAN guided by a spatially adaptive loss function. Experimental results demonstrate that the learned model achieves improved results with sharper images, fewer artifacts and less noise.
Xijun Wang 0003, Alice Lucas, Santiago Lopez Tapia, Xinyi Wu 0001, Rafael Molina 0001, Aggelos K. Katsaggelos
ICASSP3
2019 Gan-Based Video Super-Resolution With Direct Regularized Inversion of the Low-Resolution Formation Model
abstract
While high and ultra high definition displays are becoming popular, most of the available content has been acquired at much lower resolutions. In this work we propose to pseudo-invert with regularization the image formation model using GANs and perceptual losses. Our model, which does not require the use of motion compensation, utilizes explicitly the low resolution image formation model and additionally introduces two feature losses which are used to obtain perceptually improved high resolution images. The experimental validation shows that our approach outperforms current video super resolution learning based models.
Santiago Lopez Tapia, Alice Lucas, Rafael Molina 0001, Aggelos K. Katsaggelos
ICIP1
2019 Efficient Fine-Tuning of Neural Networks for Artifact Removal in Deep Learning for Inverse Imaging Problems
abstract
While Deep Neural Networks trained for solving inverse imaging problems (such as super-resolution, denoising, or inpainting tasks) regularly achieve new state-of-the-art restoration performance, this increase in performance is often accompanied with undesired artifacts generated in their solution. These artifacts are usually specific to the type of neural network architecture, training, or test input image used for the inverse imaging problem at hand. In this paper, we propose a fast, efficient post-processing method for reducing these artifacts. Given a test input image and its known image formation model, we fine-tune the parameters of the trained network and iteratively update them using a data consistency loss. We show that in addition to being efficient and applicable to large variety of problems, our post-processing through fine-tuning approach enhances the solution originally provided by the neural network by maintaining its restoration quality while reducing the observed artifacts, as measured qualitatively and quantitatively.
Alice Lucas, Santiago Lopez Tapia, Rafael Molina 0001, Aggelos K. Katsaggelos
ICIP2
2019 Deep CNNs for Object Detection Using Passive Millimeter Sensors
abstract
Passive millimeter wave images (PMMWIs) can be used to detect and localize objects concealed under clothing. Unfortunately, the quality of the acquired images and the unknown position, shape, and size of the hidden objects render these tasks challenging. In this paper, we discuss a deep learning approach to this detection/localization problem. The effect of the nonstationary acquisition noise on different architectures is analyzed and discussed. A comparison with shallow architectures is also presented. The achieved detection accuracy defines a new state of the art in object detection on PMMWIs. The low computational training and testing costs of the solution allow its use in real-time applications.
Santiago Lopez Tapia, Rafael Molina 0001, Nicolas Pérez de la Blanca
IEEE Trans. Circuits Syst. Video Technol.1
2019 Generative Adversarial Networks and Perceptual Losses for Video Super-Resolution
abstract
Video super-resolution (VSR) has become one of the most critical problems in video processing. In the deep learning literature, recent works have shown the benefits of using adversarial-based and perceptual losses to improve the performance on various image restoration tasks; however, these have yet to be applied for video super-resolution. In this paper, we propose a generative adversarial network (GAN)-based formulation for VSR. We introduce a new generator network optimized for the VSR problem, named VSRResNet, along with new discriminator architecture to properly guide VSRResNet during the GAN training. We further enhance our VSR GAN formulation with two regularizers, a distance loss in feature-space and pixel-space, to obtain our final VSRResFeatGAN model. We show that pre-training our generator with the mean-squared-error loss only quantitatively surpasses the current state-of-the-art VSR models. Finally, we employ the PercepDist metric to compare the state-of-the-art VSR models. We show that this metric more accurately evaluates the perceptual quality of SR solutions obtained from neural networks, compared with the commonly used PSNR/SSIM metrics. Finally, we show that our proposed model, the VSRResFeatGAN model, outperforms the current state-of-the-art SR models, both quantitatively and qualitatively.
Alice Lucas, Santiago Lopez Tapia, Rafael Molina 0001, Aggelos K. Katsaggelos
IEEE Trans. Image Process.2
2018 Generative Adversarial Networks and Perceptual Losses for Video Super-Resolution
abstract
Recent research on image super-resolution (SR) has shown that the use of perceptual losses such as feature-space loss functions and adversarial training can greatly improve the perceptual quality of the resulting SR output. In this paper, we extend the use of these perceptual-focused approaches for image SR to that of video SR. We design a 15-block residual neural network, VSRResNet, which is pre-trained on a the traditional mean -squared -error (MSE) loss and later fine-tuned with a feature-space loss function in an adversarial setting. We show that our proposed system, VSRRes-FeatGAN, produces super-resolved frames of much higher perceptual quality than those provided by the MSE-based model.
Alice Lucas, Aggelos K. Katsaggelos, Santiago Lopez Tapia, Rafael Molina 0001
ICIP3
2018 Using machine learning to detect and localize concealed objects in passive millimeter-wave images
Santiago Lopez Tapia, Rafael Molina 0001, Nicolas Pérez de la Blanca
Eng. Appl. Artif. Intell.1