VLDB 2026 Research / reviewers in the wild / expert
Paulo J. L. Nunes
dblp:231/2380
· DBLP profile ↗
23ranked-venue papers
7as first author
6since 2021 · last 2025
0000-0003-3982-5723ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 7 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | WaveE2VID: Frequency-Aware Event-Based Video ReconstructionabstractEvent cameras, which detect local brightness changes instead of capturing full-frame images, offer high temporal resolution and low latency. Although existing convolutional neural networks (CNNs) and transformer-based methods for event-based video reconstruction have achieved impressive results, they suffer from high computational costs due to their linear operations. These methods often require 10M-30M parameters and inference times of 30-110 ms per forward pass at a resolution of 640 × 480 on modern GPUs. Furthermore, to reduce computational costs, these methods apply CNN-based downsampling, which leads to the loss of fine details. To address these challenges, we propose an efficient hybrid model, WaveE2VID, which combines the frequency-domain analysis of the wavelet transform with the spatio-temporal context modeling of a deep convolutional recurrent network. Our model achieves 50% faster inference speed and lower GPU memory usage than CNN and transformer-based methods, maintaining reconstruction performance on par with state-of-the-art approaches across benchmark datasets. Ramna Maqsood, Paulo J. L. Nunes, Caroline Conti, Luís Ducla Soares |
ICIP | 2 |
| 2025 | Swinscale-LFVS: Parallel Feature Integration for Light Field View SynthesisabstractLight Field (LF) view synthesis aims to synthesize a dense set of LF views from a sparse set of input views. Although many recent learning-based methods have shown promising results in this task, they often rely on deep residual networks or on multiple LF representations to extract dense features, without fully exploiting the geometric structure of the LFs. In this paper, we introduce SwinScale-LFVS, a novel framework that combines the strengths of the Swin Transformer and the Multi-Scale Convolutional Network in parallel streams. The first stream uses a Swin Transformer to model local and global features using a geometry-aware Angular Mutual Self Attention (AMSA) network, and the second stream uses multi-scale 3D convolutions to extract dense features and to ensure spatial-angular consistency in synthesized LF views. The outputs from these streams are integrated and processed by an LF View Synthesis (LFVS) network to synthesize high-quality dense LF views. Extensive experiments show that SwinScale-LFVS outperforms existing methods on both real-world and synthetic datasets. The code is publicly available at https://github.com/MSP-IUL/SwinScale-LFVS. Paulo J. L. Nunes, Caroline Conti, Luís Ducla Soares |
ICIP | 2 |
| 2025 | LFVS-Mamba: State-Space Model for Light Field View SynthesisabstractLight Field View Synthesis (LFVS) methods using Convolutional Neural Networks (CNNs) and Vision Transformers (VTs) have been extensively studied: CNNs excel at learning local spatial features via hierarchical receptive fields but cannot capture long-range global dependencies, while VTs inherently model global context through self-attention at the cost of quadratic computation and memory complexity. To address these issues, we propose LFVS-Mamba, which integrates a State-Space Module (SSM) with a Selective Scanning Mechanism to efficiently capture long-range dependencies. LFVS-Mamba processes 2D slices of the 4D LF to fully exploit spatial context, complementary angular information, and depth cues. The LFVS-Mamba comprises three modules to progressively synthesize dense LFs: (i) Shallow Feature Extraction (SFE), (ii) Spatial-Angular Depth Feature Extraction (SADFE), and (iii) Angular Upsampling (AU). Experimental results on standard LF benchmarks demonstrate that LFVS-Mamba consistently outperforms existing methods. Paulo J. L. Nunes, Caroline Conti, Luís Ducla Soares |
VCIP | 2 |
| 2024 | Light Field View Synthesis Using Deformable Convolutional Neural NetworksabstractLight Field (LF) imaging has emerged as a technology that can simultaneously capture both intensity values and directions of light rays from real-world scenes. Densely sampled LFs are drawing increased attention for their wide application in 3D reconstruction, depth estimation, and digital refocusing. In order to synthesize additional views to obtain a LF with higher angular resolution, many learning-based methods have been proposed. This paper follows a similar approach to Liu et al. [1] but using deformable convolutions to improve the view synthesis performance and depth-wise separable convolutions to reduce the amount of model parameters. The proposed framework consists of two main modules: i) a multi-representation view synthesis module to extract features from different LF representations of the sparse LF, and ii) a geometry-aware refinement module to synthesize a dense LF by exploring the structural characteristics of the corresponding sparse LF. Experimental results over various benchmarks demonstrate the superiority of the proposed method when compared to state-of-the-art ones. The code is available at https://github.com/MSP-IUL/deformable_lfvs. Paulo J. L. Nunes, Caroline Conti, Luís Ducla Soares |
PCS | 2 |
| 2023 | Hyperpixels: Flexible 4D Over-Segmentation for Dense and Sparse Light Fieldsabstract4D Light Field (LF) imaging, since it conveys both spatial and angular scene information, can facilitate computer vision tasks and generate immersive experiences for end-users. A key challenge in 4D LF imaging is to flexibly and adaptively represent the included spatio-angular information to facilitate subsequent computer vision applications. Recently, image over-segmentation into homogenous regions with perceptually meaningful information has been exploited to represent 4D LFs. However, existing methods assume densely sampled LFs and do not adequately deal with sparse LFs with large occlusions. Furthermore, the spatio-angular LF cues are not fully exploited in the existing methods. In this paper, the concept of hyperpixels is defined and a flexible, automatic, and adaptive representation for both dense and sparse 4D LFs is proposed. Initially, disparity maps are estimated for all views to enhance over-segmentation accuracy and consistency. Afterwards, a modified weighted K -means clustering using robust spatio-angular features is performed in 4D Euclidean space. Experimental results on several dense and sparse 4D LF datasets show competitive and outperforming performance in terms of over-segmentation accuracy, shape regularity and view consistency against state-of-the-art methods. Maryam Hamad, Caroline Conti, Paulo J. L. Nunes, Luís Ducla Soares |
IEEE Trans. Image Process. | 3 |
| 2021 | Light field image coding with flexible viewpoint scalability and random accessabstractThis paper proposes a novel light field image compression approach with viewpoint scalability and random access functionalities. Although current state-of-the-art image coding algorithms for light fields already achieve high compression ratios, there is a lack of support for such functionalities, which are important for ensuring compatibility with different displays/capturing devices, enhanced user interaction and low decoding delay. The proposed solution enables various encoding profiles with different flexible viewpoint scalability and random access capabilities, depending on the application scenario. When compared to other state-of-the-art methods, the proposed approach consistently presents higher bitrate savings (44% on average), namely when compared to pseudo-video sequence coding approach based on HEVC. Moreover, the proposed scalable codec also outperforms MuLE and WaSP verification models, achieving average bitrate saving gains of 37% and 47%, respectively. The various flexible encoding profiles proposed add fine control to the image prediction dependencies, which allow to exploit the tradeoff between coding efficiency and the viewpoint random access, consequently, decreasing the maximum random access penalties that range from 0.60 to 0.15, for lenslet and HDCA light fields. Ricardo J. S. Monteiro, Nuno M. M. Rodrigues, Sérgio M. M. de Faria, Paulo J. L. Nunes |
Signal Process. Image Commun. | 4 |
| 2018 | Light field image coding with jointly estimated self-similarity bi-prediction
Caroline Conti, Paulo J. L. Nunes, Luís Ducla Soares |
Signal Process. Image Commun. | 2 |
| 2018 | Light Field Coding With Field-of-View Scalability and Exemplar-Based Interlayer PredictionabstractLight field imaging based on microlens arrays-a.k.a. holoscopic, plenoptic, and integral imaging-has currently risen up as a feasible and prospective technology for future image and video applications. However, deploying actual light field applications will require identifying more powerful representations and coding solutions that support arising new manipulation and interaction functionalities. In this context, this paper proposes a novel scalable coding solution that supports a new type of scalability, referred to as field-of-view scalability. The proposed scalable coding solution comprises a base layer compliant with the High Efficiency Video Coding (HEVC) standard, complemented by one or more enhancement layers that progressively allow richer versions of the same light field content in terms of content manipulation and interaction possibilities. In addition, to achieve high-compression performance in the enhancement layers, novel exemplar-based interlayer coding tools are also proposed, namely: 1) a direct prediction based on exemplar texture samples from lower layers and 2) an interlayer compensated prediction using a reference picture that is built relying on an exemplar-based algorithm for texture synthesis. Experimental results demonstrate the advantages of the proposed scalable coding solution to cater to users with different preferences/requirements in terms of interaction functionalities, while providing better rate-distortion performance (independently of the optical setup used for acquisition) compared to HEVC and other scalable light field coding solutions in the literature. Caroline Conti, Luís Ducla Soares, Paulo J. L. Nunes |
IEEE Trans. Multim. | 3 |
| 2016 | HEVC-based 3D holoscopic video coding using self-similarity compensated prediction
Caroline Conti, Luís Ducla Soares, Paulo J. L. Nunes |
Signal Process. Image Commun. | 3 |
| 2013 | Inter-Layer Prediction Scheme for Scalable 3-D Holoscopic Video CodingabstractHoloscopic imaging has recently become a prospective glassless 3-D technology conquering the attention of researchers seeking more realistic depth-illusion approaches. However, backward compatibility with legacy displays is crucial to progressively introduce this technology into the consumer market and to efficiently deliver 3-D holoscopic content to end-users. Therefore, this letter proposes a new display scalable coding solution for 3-D holoscopic based on an inter-layer prediction scheme that exploits the redundancy between multiview and 3-D holoscopic content representations. Experimental results show that this inter-layer prediction scheme integrated into the High Efficiency Video Coding (HEVC) is advantageous, always outperforming the simulcast approach. Caroline Conti, Paulo J. L. Nunes, Luís Ducla Soares |
IEEE Signal Process. Lett. | 2 |
| 2012 | New HEVC prediction modes for 3D holoscopic video codingabstractHoloscopic imaging is an advantageous solution for glassless 3D video systems, which promises to revolutionize the 3D market in the near future. Besides freeing the user from wearing any viewing device, it supports full motion parallax, improving this way the users' viewing experience. However, in order to provide 3D holoscopic content with convenient visual quality in terms of resolution and 3D perception, ultra-high resolution acquisition and display devices are required. Consequently, efficient video coding tools to deal with this large amount of data become of paramount importance. The recent standardization project called High Efficiency Video Coding (HEVC) addresses the requirements of high resolution video coding, but does not yet address the specific characteristics of 3D holoscopic content. To remedy this situation, this paper proposes to incorporate new prediction modes in HEVC to explore the particular structure of 3D holoscopic content, in order to further improve the performance of HEVC for this type of content. Experimental results, based on the HEVC test model version 4.0 are presented and clearly show the advantages of using this approach. Caroline Conti, Paulo J. L. Nunes, Luís Ducla Soares |
ICIP | 2 |
| 2011 | Spatial prediction based on self-similarity compensation for 3D holoscopic image and video codingabstractHoloscopic imaging, also known as integral imaging, provides a solution for glassless 3D, and is promising to change the market for 3D television. To start, this paper briefly describes the general concepts of holoscopic imaging, focusing mainly on the spatial correlations inherent to this new type of content, which appear due to the micro-lens array that is used for both acquisition and display. The micro-images that are formed behind each micro-lens, from which only one pixel is viewed from a given observation point, have a high cross-correlation between them, which can be exploited for coding. A novel scheme for spatial prediction, exploring the particular arrangement of holoscopic images, is proposed. The proposed scheme can be used for both still image coding and intra-coding of video. Experimental results based on an H.264/AVC video codec modified to handle 3D holoscopic images and video are presented, showing the superior performance of this approach. Caroline Conti, João Lino, Paulo J. L. Nunes, Luís Ducla Soares, Paulo Lobato Correia |
ICIP | 3 |
| 2010 | Automatic MPEG-4 sprite coding - Comparison of integrated object segmentation algorithms
Alexander Glantz, Andreas Krutz, Thomas Sikora, Paulo J. L. Nunes, Fernando Pereira 0001 |
Multim. Tools Appl. | 4 |
| 2009 | Automatic and adaptive network-aware macroblock intra refresh for error-resilient H.264/AVC video codingabstractIn this paper, an automatic and adaptive network-aware macroblock Intra coding refresh method is proposed. It adaptively selects the amount of gracefully forced Intra macroblocks and the amount of cyclic Intra refresh (CIR) macroblocks based on the actual network error conditions, in terms of packet loss rate, an the target encoding bit rate. With the proposed method, the error robustness of H.264/AVC bitstreams can be significantly increased by efficiently taking into account the actual rate-distortion impact of Intra coding macroblock mode decisions, while simultaneously guaranteeing that errors do not propagate endlessly by selecting an adequate amount of CIR macroblocks per frame according to the network packet loss rate and the encoding target bit rate. Paulo J. L. Nunes, Luís Ducla Soares, Fernando Pereira 0001 |
ICIP | 1 |
| 2009 | Joint Rate Control Algorithm for Low-Delay MPEG-4 Object-Based Video EncodingabstractThis paper proposes an improved rate control algorithm for jointly encoding multiple arbitrarily shaped video objects in the context of low-delay MPEG-4 compliant video coding. The algorithm provides adequate mechanisms for dealing with deviations between the ideal and the actual behavior of video scene encoders, notably: 1) compensation mechanisms (e.g., rate control decisions) that are able to track these deviations and compensate them to allow a stable and efficient operation of the encoder, and 2) adaptation mechanisms (e.g., estimation of model parameters) that are able to instantaneously represent the actual behavior of the encoder and its rate controller. The proposed solution efficiently allocates the available resources, i.e., target bit rate and bitstream buffer space, aiming at maximizing the average scene quality and minimizing quality fluctuations along time and among the various video objects. The results show that this solution outperforms the usual reference solutions, notably those specified in the rate control informative annex of the MPEG-4 visual standard. Paulo J. L. Nunes, Fernando Pereira 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2008 | Error resilient macroblock rate control for H.264/AVC video codingabstractIn this paper, an error resilient rate control scheme for the H.264/AVC standard is proposed. This scheme differs from traditional rate control schemes in that macroblock mode decisions are not made only to minimize their rate-distortion cost, but also take into account that the bitstream will have to be transmitted through an error-prone network. Since channel errors will probably occur, error propagation due to predictive coding should be mitigated by adequate Intra coding refreshes. The proposed scheme works by comparing the rate-distortion cost of coding a macroblock in Intra and Inter modes: if the cost of Intra coding is only slightly larger than the cost of Inter coding, the coding mode is changed to Intra, thus reducing error propagation. Additionally, cyclic Intra refresh is also applied to guarantee that all macroblocks are eventually refreshed. The proposed scheme outperforms the H.264/AVC reference software, for typical test sequences, for error-free transmission and several packet loss rates. Paulo J. L. Nunes, Luís Ducla Soares, Fernando Pereira 0001 |
ICIP | 1 |
| 2007 | Improved Feedback Compensation Mechanisms for Multiple Video Object Encoding Rate ControlabstractThis paper proposes new buffer and video object distortion feedback compensation mechanisms for efficiently dealing with deviations between the ideal and the actual behavior of video scene encoders when jointly encoding multiple arbitrarily shaped video objects in the context of compliant low-delay object-based MPEG-4 video coding. The proposed solution computes target buffer occupancies for each encoding time instant based on the amount and complexity of the video data to encode, and the bit allocation for each encoding time instant is feedback adjusted according to deviations relatively to this ideal behavior. Additionally, each video object bit allocation is also feedback adjusted based on the relative distortion of the various video objects in the scene. The proposed solution outperforms the non-normative MPEG-4 reference rate control algorithm for a wide range of bit rates and spatio-temporal resolutions, for typical test sequences. Paulo J. L. Nunes, Fernando Pereira 0001 |
ICIP (3) | 1 |
| 2002 | Evaluating MPEG-4 video decoding complexity for an alternative video complexity verifier modelabstractMPEG-4 is the first object-based audiovisual coding standard. To control the minimum decoding complexity resources required at the decoder, the MPEG-4 Visual standard defines the so-called video buffering verifier mechanism, which includes three virtual buffer models, among them the video complexity verifier (VCV). This paper proposes an alternative VCV model, based on a set of macroblock (MB) relative decoding complexity weights assigned to the various MB coding types used in MPEG-4 video coding. The new VCV model allows a more efficient use of the available decoding resources by preventing the overevaluation of the decoding complexity of certain MB types and thus making it possible to encode scenes (for the same profile@level decoding resources) which otherwise would be considered too demanding. João Valentim, Paulo J. L. Nunes, Fernando Pereira 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2001 | An alternative complexity model for the MPEG-4 video verifier mechanismabstractMPEG-4 is the first object-based audiovisual coding standard. To control the minimum decoding complexity resources required at the decoder, the MPEG-4 visual standard defines the so-called video complexity verifier (VCV). This paper proposes an alternative VCV model, based on a set of relative macroblock (MB) complexity weights assigned to the various MB coding types used in MPEG-4 video coding. The new VCV model allows a more efficient use of the available decoding resources by preventing the over-evaluation of the decoding complexity of certain MB types and thus making possible to encode scenes (for the same profile@level decoding resources) which otherwise would be considered too demanding. João Valentim, Paulo J. L. Nunes, Fernando Pereira 0001 |
ICIP (1) | 2 |
| 2001 | Scene level rate control algorithm for MPEG-4 video coding
Paulo J. L. Nunes, Fernando Pereira 0001 |
VCIP | 1 |
| 2000 | A contour-based approach to binary shape coding using a multiple grid chain code
Paulo J. L. Nunes, Ferran Marqués, Fernando Pereira 0001, Antoni Gasull |
Signal Process. Image Commun. | 1 |
| 1997 | Multi-grid chain coding of binary shapesabstractThis paper presents a chain code based approach to efficiently code binary shape information of video objects, in the context of object-based video coding. The proposed method tries to meet some of the requirements of the MPEG-4 standard, currently under development, notably efficient coding, and low delay. This approach allows several modes of operation depending on the application requirements, notably lossless, near-lossless, and lossy coding modes. For the lossless case a pure differential chain code method is proposed while for the near-lossless case a multi-grid chain code (MGCC) technique is adopted. Also both INTRA and INTER prediction modes can be used. For the INTER mode, motion compensation is applied without coding the residues. The MGCC is a near-lossless contour coding technique using a contour description based on edges, which combines both contour prediction and contour simplification. Paulo J. L. Nunes, Fernando Pereira 0001, Ferran Marqués |
ICIP (3) | 1 |
| 1995 | Image segmentation towards new image representation methods
Diogo Cortez, Paulo J. L. Nunes, Manuel Menezes de Sequeira, Fernando Pereira 0001 |
Signal Process. Image Commun. | 2 |