VLDB 2026 Research / reviewers in the wild / expert
Luís Ducla Soares
dblp:88/3657
· DBLP profile ↗
28ranked-venue papers
12as first author
5since 2021 · last 2025
0000-0001-9738-639XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 11 first-author · 5 since 2021Artificial intelligence and machine learning · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | WaveE2VID: Frequency-Aware Event-Based Video ReconstructionabstractEvent cameras, which detect local brightness changes instead of capturing full-frame images, offer high temporal resolution and low latency. Although existing convolutional neural networks (CNNs) and transformer-based methods for event-based video reconstruction have achieved impressive results, they suffer from high computational costs due to their linear operations. These methods often require 10M-30M parameters and inference times of 30-110 ms per forward pass at a resolution of 640 × 480 on modern GPUs. Furthermore, to reduce computational costs, these methods apply CNN-based downsampling, which leads to the loss of fine details. To address these challenges, we propose an efficient hybrid model, WaveE2VID, which combines the frequency-domain analysis of the wavelet transform with the spatio-temporal context modeling of a deep convolutional recurrent network. Our model achieves 50% faster inference speed and lower GPU memory usage than CNN and transformer-based methods, maintaining reconstruction performance on par with state-of-the-art approaches across benchmark datasets. Ramna Maqsood, Paulo J. L. Nunes, Caroline Conti, Luís Ducla Soares |
ICIP | 4 |
| 2025 | Swinscale-LFVS: Parallel Feature Integration for Light Field View SynthesisabstractLight Field (LF) view synthesis aims to synthesize a dense set of LF views from a sparse set of input views. Although many recent learning-based methods have shown promising results in this task, they often rely on deep residual networks or on multiple LF representations to extract dense features, without fully exploiting the geometric structure of the LFs. In this paper, we introduce SwinScale-LFVS, a novel framework that combines the strengths of the Swin Transformer and the Multi-Scale Convolutional Network in parallel streams. The first stream uses a Swin Transformer to model local and global features using a geometry-aware Angular Mutual Self Attention (AMSA) network, and the second stream uses multi-scale 3D convolutions to extract dense features and to ensure spatial-angular consistency in synthesized LF views. The outputs from these streams are integrated and processed by an LF View Synthesis (LFVS) network to synthesize high-quality dense LF views. Extensive experiments show that SwinScale-LFVS outperforms existing methods on both real-world and synthetic datasets. The code is publicly available at https://github.com/MSP-IUL/SwinScale-LFVS. Paulo J. L. Nunes, Caroline Conti, Luís Ducla Soares |
ICIP | 4 |
| 2025 | LFVS-Mamba: State-Space Model for Light Field View SynthesisabstractLight Field View Synthesis (LFVS) methods using Convolutional Neural Networks (CNNs) and Vision Transformers (VTs) have been extensively studied: CNNs excel at learning local spatial features via hierarchical receptive fields but cannot capture long-range global dependencies, while VTs inherently model global context through self-attention at the cost of quadratic computation and memory complexity. To address these issues, we propose LFVS-Mamba, which integrates a State-Space Module (SSM) with a Selective Scanning Mechanism to efficiently capture long-range dependencies. LFVS-Mamba processes 2D slices of the 4D LF to fully exploit spatial context, complementary angular information, and depth cues. The LFVS-Mamba comprises three modules to progressively synthesize dense LFs: (i) Shallow Feature Extraction (SFE), (ii) Spatial-Angular Depth Feature Extraction (SADFE), and (iii) Angular Upsampling (AU). Experimental results on standard LF benchmarks demonstrate that LFVS-Mamba consistently outperforms existing methods. Paulo J. L. Nunes, Caroline Conti, Luís Ducla Soares |
VCIP | 4 |
| 2024 | Light Field View Synthesis Using Deformable Convolutional Neural NetworksabstractLight Field (LF) imaging has emerged as a technology that can simultaneously capture both intensity values and directions of light rays from real-world scenes. Densely sampled LFs are drawing increased attention for their wide application in 3D reconstruction, depth estimation, and digital refocusing. In order to synthesize additional views to obtain a LF with higher angular resolution, many learning-based methods have been proposed. This paper follows a similar approach to Liu et al. [1] but using deformable convolutions to improve the view synthesis performance and depth-wise separable convolutions to reduce the amount of model parameters. The proposed framework consists of two main modules: i) a multi-representation view synthesis module to extract features from different LF representations of the sparse LF, and ii) a geometry-aware refinement module to synthesize a dense LF by exploring the structural characteristics of the corresponding sparse LF. Experimental results over various benchmarks demonstrate the superiority of the proposed method when compared to state-of-the-art ones. The code is available at https://github.com/MSP-IUL/deformable_lfvs. Paulo J. L. Nunes, Caroline Conti, Luís Ducla Soares |
PCS | 4 |
| 2023 | Hyperpixels: Flexible 4D Over-Segmentation for Dense and Sparse Light Fieldsabstract4D Light Field (LF) imaging, since it conveys both spatial and angular scene information, can facilitate computer vision tasks and generate immersive experiences for end-users. A key challenge in 4D LF imaging is to flexibly and adaptively represent the included spatio-angular information to facilitate subsequent computer vision applications. Recently, image over-segmentation into homogenous regions with perceptually meaningful information has been exploited to represent 4D LFs. However, existing methods assume densely sampled LFs and do not adequately deal with sparse LFs with large occlusions. Furthermore, the spatio-angular LF cues are not fully exploited in the existing methods. In this paper, the concept of hyperpixels is defined and a flexible, automatic, and adaptive representation for both dense and sparse 4D LFs is proposed. Initially, disparity maps are estimated for all views to enhance over-segmentation accuracy and consistency. Afterwards, a modified weighted K -means clustering using robust spatio-angular features is performed in 4D Euclidean space. Experimental results on several dense and sparse 4D LF datasets show competitive and outperforming performance in terms of over-segmentation accuracy, shape regularity and view consistency against state-of-the-art methods. Maryam Hamad, Caroline Conti, Paulo J. L. Nunes, Luís Ducla Soares |
IEEE Trans. Image Process. | 4 |
| 2018 | Using transfer learning for classification of gait pathologies
Tanmay T. Verlekar, Paulo Lobato Correia, Luís Ducla Soares |
BIBM | 3 |
| 2018 | Gait recognition in the wild using shadow silhouettes
Tanmay T. Verlekar, Luís Ducla Soares, Paulo Lobato Correia |
Image Vis. Comput. | 2 |
| 2018 | Light field image coding with jointly estimated self-similarity bi-prediction
Caroline Conti, Paulo J. L. Nunes, Luís Ducla Soares |
Signal Process. Image Commun. | 3 |
| 2018 | Light Field Coding With Field-of-View Scalability and Exemplar-Based Interlayer PredictionabstractLight field imaging based on microlens arrays-a.k.a. holoscopic, plenoptic, and integral imaging-has currently risen up as a feasible and prospective technology for future image and video applications. However, deploying actual light field applications will require identifying more powerful representations and coding solutions that support arising new manipulation and interaction functionalities. In this context, this paper proposes a novel scalable coding solution that supports a new type of scalability, referred to as field-of-view scalability. The proposed scalable coding solution comprises a base layer compliant with the High Efficiency Video Coding (HEVC) standard, complemented by one or more enhancement layers that progressively allow richer versions of the same light field content in terms of content manipulation and interaction possibilities. In addition, to achieve high-compression performance in the enhancement layers, novel exemplar-based interlayer coding tools are also proposed, namely: 1) a direct prediction based on exemplar texture samples from lower layers and 2) an interlayer compensated prediction using a reference picture that is built relying on an exemplar-based algorithm for texture synthesis. Experimental results demonstrate the advantages of the proposed scalable coding solution to cater to users with different preferences/requirements in terms of interaction functionalities, while providing better rate-distortion performance (independently of the optical setup used for acquisition) compared to HEVC and other scalable light field coding solutions in the literature. Caroline Conti, Luís Ducla Soares, Paulo J. L. Nunes |
IEEE Trans. Multim. | 2 |
| 2016 | HEVC-based 3D holoscopic video coding using self-similarity compensated prediction
Caroline Conti, Luís Ducla Soares, Paulo J. L. Nunes |
Signal Process. Image Commun. | 2 |
| 2013 | Inter-Layer Prediction Scheme for Scalable 3-D Holoscopic Video CodingabstractHoloscopic imaging has recently become a prospective glassless 3-D technology conquering the attention of researchers seeking more realistic depth-illusion approaches. However, backward compatibility with legacy displays is crucial to progressively introduce this technology into the consumer market and to efficiently deliver 3-D holoscopic content to end-users. Therefore, this letter proposes a new display scalable coding solution for 3-D holoscopic based on an inter-layer prediction scheme that exploits the redundancy between multiview and 3-D holoscopic content representations. Experimental results show that this inter-layer prediction scheme integrated into the High Efficiency Video Coding (HEVC) is advantageous, always outperforming the simulcast approach. Caroline Conti, Paulo J. L. Nunes, Luís Ducla Soares |
IEEE Signal Process. Lett. | 3 |
| 2012 | New HEVC prediction modes for 3D holoscopic video codingabstractHoloscopic imaging is an advantageous solution for glassless 3D video systems, which promises to revolutionize the 3D market in the near future. Besides freeing the user from wearing any viewing device, it supports full motion parallax, improving this way the users' viewing experience. However, in order to provide 3D holoscopic content with convenient visual quality in terms of resolution and 3D perception, ultra-high resolution acquisition and display devices are required. Consequently, efficient video coding tools to deal with this large amount of data become of paramount importance. The recent standardization project called High Efficiency Video Coding (HEVC) addresses the requirements of high resolution video coding, but does not yet address the specific characteristics of 3D holoscopic content. To remedy this situation, this paper proposes to incorporate new prediction modes in HEVC to explore the particular structure of 3D holoscopic content, in order to further improve the performance of HEVC for this type of content. Experimental results, based on the HEVC test model version 4.0 are presented and clearly show the advantages of using this approach. Caroline Conti, Paulo J. L. Nunes, Luís Ducla Soares |
ICIP | 3 |
| 2011 | Spatial prediction based on self-similarity compensation for 3D holoscopic image and video codingabstractHoloscopic imaging, also known as integral imaging, provides a solution for glassless 3D, and is promising to change the market for 3D television. To start, this paper briefly describes the general concepts of holoscopic imaging, focusing mainly on the spatial correlations inherent to this new type of content, which appear due to the micro-lens array that is used for both acquisition and display. The micro-images that are formed behind each micro-lens, from which only one pixel is viewed from a given observation point, have a high cross-correlation between them, which can be exploited for coding. A novel scheme for spatial prediction, exploring the particular arrangement of holoscopic images, is proposed. The proposed scheme can be used for both still image coding and intra-coding of video. Experimental results based on an H.264/AVC video codec modified to handle 3D holoscopic images and video are presented, showing the superior performance of this approach. Caroline Conti, João Lino, Paulo J. L. Nunes, Luís Ducla Soares, Paulo Lobato Correia |
ICIP | 4 |
| 2011 | Distributed source coding for securing a hand-based biometric recognition systemabstractDistributed source coding (DSC) is typically used to compress information from multiple correlated sources that do not communicate with each other. In this paper, DSC principles are used to secure biometric data in a multimodal hand-based recognition system, introducing a novel approach for Log-Likelihood Ratio initialization and a new binarization technique. The proposed biometric recognition system relies on a new architecture using three hand-based biometric traits: palmprint, finger surface and hand geometry, being the latter used as a soft biometric to accelerate the identification process. Promising results are achieved in terms of recognition accuracy, speed and security when performing identification on large databases, like the UST Hand Image Database. Mauricio Ramalho, Paulo Lobato Correia, Luís Ducla Soares |
ICIP | 3 |
| 2009 | Automatic and adaptive network-aware macroblock intra refresh for error-resilient H.264/AVC video codingabstractIn this paper, an automatic and adaptive network-aware macroblock Intra coding refresh method is proposed. It adaptively selects the amount of gracefully forced Intra macroblocks and the amount of cyclic Intra refresh (CIR) macroblocks based on the actual network error conditions, in terms of packet loss rate, an the target encoding bit rate. With the proposed method, the error robustness of H.264/AVC bitstreams can be significantly increased by efficiently taking into account the actual rate-distortion impact of Intra coding macroblock mode decisions, while simultaneously guaranteeing that errors do not propagate endlessly by selecting an adequate amount of CIR macroblocks per frame according to the network packet loss rate and the encoding target bit rate. Paulo J. L. Nunes, Luís Ducla Soares, Fernando Pereira 0001 |
ICIP | 2 |
| 2008 | Error resilient macroblock rate control for H.264/AVC video codingabstractIn this paper, an error resilient rate control scheme for the H.264/AVC standard is proposed. This scheme differs from traditional rate control schemes in that macroblock mode decisions are not made only to minimize their rate-distortion cost, but also take into account that the bitstream will have to be transmitted through an error-prone network. Since channel errors will probably occur, error propagation due to predictive coding should be mitigated by adequate Intra coding refreshes. The proposed scheme works by comparing the rate-distortion cost of coding a macroblock in Intra and Inter modes: if the cost of Intra coding is only slightly larger than the cost of Inter coding, the coding mode is changed to Intra, thus reducing error propagation. Additionally, cyclic Intra refresh is also applied to guarantee that all macroblocks are eventually refreshed. The proposed scheme outperforms the H.264/AVC reference software, for typical test sequences, for error-free transmission and several packet loss rates. Paulo J. L. Nunes, Luís Ducla Soares, Fernando Pereira 0001 |
ICIP | 2 |
| 2006 | Spatio-Temporal Scene Level Error Concealment for Shape and Texture Data in Segmented Video ContentabstractIn this paper, a novel shape and texture error concealment technique for segmented object-based video scenes is proposed. This technique is different from existing concealment techniques because it considers not only the corrupted video objects to be concealed, but also the context/scene in which they are inserted. In the proposed technique, concealment is done by using information from the current time instant as well as from the past. The obtained results suggest that the use of this technique significantly improves the subjective visual impact of scenes on the end-user, when compared to independent concealment of video objects. Luís Ducla Soares, Fernando Pereira 0001 |
ICIP | 1 |
| 2006 | Temporal shape error concealment by global motion compensation with local refinementabstractThis paper presents an original temporal shape error concealment technique based on a combination of global and local motion compensation. For this technique, which is especially useful for object-based video applications in error-prone environments (e.g., mobile networks), it is assumed that the shape of the corrupted object at hand is in the form of a binary alpha plane and some of the shape data is missing due to channel errors. To conceal the corrupted shape, the decoder first assumes that a global motion model can describe the shape changes in consecutive time instants. This way, based on locally estimated global motion parameters, the decoder attempts to conceal the corrupted alpha plane by global motion compensating the shape data from the previous time instant. Afterwards, since a global motion model cannot perfectly describe all alpha plane changes, a local motion refinement is applied to improve the concealment in areas of the object with significant local motion. Luís Ducla Soares, Fernando Pereira 0001 |
IEEE Trans. Image Process. | 1 |
| 2004 | Motion-based shape error concealment for object-based videoabstractIn this paper, an original motion-based shape error concealment technique, especially useful for object-based video applications in error-prone environments such as mobile networks, is proposed. It is assumed that the shape of the corrupted object at hand is in the form of a binary alpha plane and some of the shape data is missing due to channel errors. To conceal the corrupted shape, the decoder starts by assuming that the alpha plane changes in consecutive time instants can be described by a global motion model. This way, based on locally estimated global motion parameters, the decoder tries to conceal the corrupted alpha plane by global motion compensating the shape data from the previous time instant. Then, since not all alpha plane changes can be perfectly described by global motion, an additional local motion refinement is applied to deal with areas of the object that have significant motion. Luís Ducla Soares, Fernando Pereira 0001 |
ICIP | 1 |
| 2004 | Spatial shape error concealment for object-based image and video codingabstractIn this paper, an original spatial shape error-concealment technique, to be used in the context of object-based image and video coding schemes, is proposed. In this technique, it is assumed that the shape of the corrupted object at hand is in the form of a binary alpha plane, in which some of the shape data is missing due to channel errors. From this alpha plane, a contour corresponding to the border of the object can be extracted. However, due to errors, some parts of the contour will be missing and, therefore, the contour will be broken. The proposed technique relies on the interpolation of the missing contours with Bézier curves, which is done based on the available surrounding contours. After all the missing parts of the contour have been interpolated, the concealed alpha plane can be easily reconstructed from the fully recovered contour and used instead of the erroneous one improving the final subjective impact. Luís Ducla Soares, Fernando Pereira 0001 |
IEEE Trans. Image Process. | 1 |
| 2004 | Adaptive shape and texture intra refreshment schemes for improved error resilience in object-based video codingabstractVideo encoders may use several techniques to improve error resilience. In particular, for video encoders that rely on predictive (inter) coding to remove temporal redundancy, intra coding refreshment is especially useful to stop temporal error propagation when errors occur in the transmission or storage of the coded streams, since these errors may cause the decoded quality to decay very rapidly. In the context of object-based video coding, intra coding refreshment can be applied to both the shape and texture data. In this paper, novel shape and texture intra refreshment schemes are proposed which can be used by object-based video encoders, such as MPEG-4 video encoders, independently or combined. These schemes allow to adaptively determine when the shape and texture of the various video objects in a scene should be refreshed in order to maximize the decoded video quality for a certain total bit rate. Luís Ducla Soares, Fernando Pereira 0001 |
IEEE Trans. Image Process. | 1 |
| 2003 | Refreshment need metrics for improved shape and texture object-based resilient video codingabstractVideo encoders may use several techniques to improve error resilience. In particular, for video encoders that rely on predictive (inter) coding to remove temporal redundancy, intra coding refreshment is especially useful to stop error propagation when errors occur in the transmission or storage of the coded streams, which can cause the decoded quality to decay very rapidly. In the context of object-based video coding, the video encoder can apply intra coding refreshment to both the shape and the texture data. In this paper, shape refreshment need and texture refreshment need metrics are proposed which can be used by object-based video encoders, notably MPEG-4 video encoders, to determine when the shape and the texture of the various video objects in the scene should be refreshed in order to improve the decoded video quality, e.g., for a given bitrate. Luís Ducla Soares, Fernando Pereira 0001 |
IEEE Trans. Image Process. | 1 |
| 2002 | Shape refreshment need metric for object-based resilient video codingabstractAlthough there are several techniques that video encoders may use to improve error resilience, it is largely recognized that intra coding refreshment plays a major role. This technique is especially useful for video encoders that rely on predictive (inter) coding to remove temporal redundancy because, in these conditions, the decoded quality can decay very rapidly due to error propagation if errors occur in the transmission or storage of the coded streams. Therefore, in order to avoid error propagation for too long a time, the encoder can use a coding refreshment scheme to refresh the decoding process and stop (spatial and temporal) error propagation. In the context of object-based video coding, the video encoder can apply intra coding refreshment to both the shape and the texture data. A shape refreshment need metric is proposed which can be used by object-based video encoders, notably MPEG-4 video encoders, to determine when the shape of a given video object should be refreshed in order to improve the decoded video quality. Luís Ducla Soares, Fernando Pereira 0001 |
ICIP (1) | 1 |
| 2002 | Texture refreshment need metric for resilient object-based video codingabstractVideo encoders may use several techniques to improve error resilience. In particular, for video encoders that rely on predictive (inter) coding to remove temporal redundancy, intra coding refreshment is especially useful to stop error propagation when errors occur in the transmission or storage of the coded streams, which can cause the decoded quality to decay very rapidly. In object-based video coders, intra coding refreshment can be applied to both shape and texture data. A texture refreshment need metric is proposed which can be used by object-based video encoders, notably MPEG-4 video encoders, to determine when the texture of the various video objects in a scene should be refreshed in order to improve the decoded video quality, e.g. for a certain amount of bitrate resources. Luís Ducla Soares, Fernando Pereira 0001 |
ICME (2) | 1 |
| 2000 | Influence of Encoder Parameters on the Decoded Video Quality for MPEG-4 over W-CDMA Mobile NetworksabstractThe MPEG-4 standard provides error resilience tools that can significantly increase the decoded video quality when using error prone media to transmit or store the encoded data. However, the use of these tools introduces extra redundancy and overhead in the bitstreams, which means that the decoded video quality can be severely affected if no careful configuration of the relevant parameters is done while taking into account the error characteristics of the channel being used. In this paper, a videotelephony system over a W-CDMA mobile network is used to study the influence of the encoding parameters on the decoded video quality, notably its behavior and optimization. Luís Ducla Soares, Satoru Adachi, Fernando Pereira 0001 |
ICIP | 1 |
| 1999 | Error resilience and concealment performance for MPEG-4 frame-based video coding
Luís Ducla Soares, Fernando Pereira 0001 |
Signal Process. Image Commun. | 1 |
| 1998 | An Alternative to the MPEG-4 Object-based Error Resilient Video Syntax
Luís Ducla Soares, Fernando Pereira 0001 |
ICIP (3) | 1 |
| 1998 | MPEG-4: a flexible coding standard for the emerging mobile multimedia applicationsabstractThis paper analyses the relevance and performance of the emerging MPEG-4 audiovisual coding standard for emerging mobile multimedia applications. Some results are presented for one of the MPEG-4 profiles targeting mobile scenarios. Luís Ducla Soares, Fernando Pereira 0001 |
PIMRC | 1 |