Caroline Conti

dblp:10/10699 · DBLP profile ↗
← Back
12ranked-venue papers
7as first author
6since 2021 · last 2025
0000-0002-9197-2627ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 7 first-author · 6 since 2021
YearPublicationVenuePosition
2025 WaveE2VID: Frequency-Aware Event-Based Video Reconstruction
abstract
Event cameras, which detect local brightness changes instead of capturing full-frame images, offer high temporal resolution and low latency. Although existing convolutional neural networks (CNNs) and transformer-based methods for event-based video reconstruction have achieved impressive results, they suffer from high computational costs due to their linear operations. These methods often require 10M-30M parameters and inference times of 30-110 ms per forward pass at a resolution of 640 × 480 on modern GPUs. Furthermore, to reduce computational costs, these methods apply CNN-based downsampling, which leads to the loss of fine details. To address these challenges, we propose an efficient hybrid model, WaveE2VID, which combines the frequency-domain analysis of the wavelet transform with the spatio-temporal context modeling of a deep convolutional recurrent network. Our model achieves 50% faster inference speed and lower GPU memory usage than CNN and transformer-based methods, maintaining reconstruction performance on par with state-of-the-art approaches across benchmark datasets.
Ramna Maqsood, Paulo J. L. Nunes, Caroline Conti, Luís Ducla Soares
ICIP3
2025 Swinscale-LFVS: Parallel Feature Integration for Light Field View Synthesis
abstract
Light Field (LF) view synthesis aims to synthesize a dense set of LF views from a sparse set of input views. Although many recent learning-based methods have shown promising results in this task, they often rely on deep residual networks or on multiple LF representations to extract dense features, without fully exploiting the geometric structure of the LFs. In this paper, we introduce SwinScale-LFVS, a novel framework that combines the strengths of the Swin Transformer and the Multi-Scale Convolutional Network in parallel streams. The first stream uses a Swin Transformer to model local and global features using a geometry-aware Angular Mutual Self Attention (AMSA) network, and the second stream uses multi-scale 3D convolutions to extract dense features and to ensure spatial-angular consistency in synthesized LF views. The outputs from these streams are integrated and processed by an LF View Synthesis (LFVS) network to synthesize high-quality dense LF views. Extensive experiments show that SwinScale-LFVS outperforms existing methods on both real-world and synthetic datasets. The code is publicly available at https://github.com/MSP-IUL/SwinScale-LFVS.
Paulo J. L. Nunes, Caroline Conti, Luís Ducla Soares
ICIP3
2025 LFVS-Mamba: State-Space Model for Light Field View Synthesis
abstract
Light Field View Synthesis (LFVS) methods using Convolutional Neural Networks (CNNs) and Vision Transformers (VTs) have been extensively studied: CNNs excel at learning local spatial features via hierarchical receptive fields but cannot capture long-range global dependencies, while VTs inherently model global context through self-attention at the cost of quadratic computation and memory complexity. To address these issues, we propose LFVS-Mamba, which integrates a State-Space Module (SSM) with a Selective Scanning Mechanism to efficiently capture long-range dependencies. LFVS-Mamba processes 2D slices of the 4D LF to fully exploit spatial context, complementary angular information, and depth cues. The LFVS-Mamba comprises three modules to progressively synthesize dense LFs: (i) Shallow Feature Extraction (SFE), (ii) Spatial-Angular Depth Feature Extraction (SADFE), and (iii) Angular Upsampling (AU). Experimental results on standard LF benchmarks demonstrate that LFVS-Mamba consistently outperforms existing methods.
Paulo J. L. Nunes, Caroline Conti, Luís Ducla Soares
VCIP3
2024 Light Field View Synthesis Using Deformable Convolutional Neural Networks
abstract
Light Field (LF) imaging has emerged as a technology that can simultaneously capture both intensity values and directions of light rays from real-world scenes. Densely sampled LFs are drawing increased attention for their wide application in 3D reconstruction, depth estimation, and digital refocusing. In order to synthesize additional views to obtain a LF with higher angular resolution, many learning-based methods have been proposed. This paper follows a similar approach to Liu et al. [1] but using deformable convolutions to improve the view synthesis performance and depth-wise separable convolutions to reduce the amount of model parameters. The proposed framework consists of two main modules: i) a multi-representation view synthesis module to extract features from different LF representations of the sparse LF, and ii) a geometry-aware refinement module to synthesize a dense LF by exploring the structural characteristics of the corresponding sparse LF. Experimental results over various benchmarks demonstrate the superiority of the proposed method when compared to state-of-the-art ones. The code is available at https://github.com/MSP-IUL/deformable_lfvs.
Paulo J. L. Nunes, Caroline Conti, Luís Ducla Soares
PCS3
2024 Editorial
Caroline Conti, Atanas P. Gotchev, Robert Bregovic, Donald G. Dansereau, Cristian Perra, Toshiaki Fujii
Signal Process. Image Commun.1
2023 Hyperpixels: Flexible 4D Over-Segmentation for Dense and Sparse Light Fields
abstract
4D Light Field (LF) imaging, since it conveys both spatial and angular scene information, can facilitate computer vision tasks and generate immersive experiences for end-users. A key challenge in 4D LF imaging is to flexibly and adaptively represent the included spatio-angular information to facilitate subsequent computer vision applications. Recently, image over-segmentation into homogenous regions with perceptually meaningful information has been exploited to represent 4D LFs. However, existing methods assume densely sampled LFs and do not adequately deal with sparse LFs with large occlusions. Furthermore, the spatio-angular LF cues are not fully exploited in the existing methods. In this paper, the concept of hyperpixels is defined and a flexible, automatic, and adaptive representation for both dense and sparse 4D LFs is proposed. Initially, disparity maps are estimated for all views to enhance over-segmentation accuracy and consistency. Afterwards, a modified weighted K -means clustering using robust spatio-angular features is performed in 4D Euclidean space. Experimental results on several dense and sparse 4D LF datasets show competitive and outperforming performance in terms of over-segmentation accuracy, shape regularity and view consistency against state-of-the-art methods.
Maryam Hamad, Caroline Conti, Paulo J. L. Nunes, Luís Ducla Soares
IEEE Trans. Image Process.2
2018 Light field image coding with jointly estimated self-similarity bi-prediction
Caroline Conti, Paulo J. L. Nunes, Luís Ducla Soares
Signal Process. Image Commun.1
2018 Light Field Coding With Field-of-View Scalability and Exemplar-Based Interlayer Prediction
abstract
Light field imaging based on microlens arrays-a.k.a. holoscopic, plenoptic, and integral imaging-has currently risen up as a feasible and prospective technology for future image and video applications. However, deploying actual light field applications will require identifying more powerful representations and coding solutions that support arising new manipulation and interaction functionalities. In this context, this paper proposes a novel scalable coding solution that supports a new type of scalability, referred to as field-of-view scalability. The proposed scalable coding solution comprises a base layer compliant with the High Efficiency Video Coding (HEVC) standard, complemented by one or more enhancement layers that progressively allow richer versions of the same light field content in terms of content manipulation and interaction possibilities. In addition, to achieve high-compression performance in the enhancement layers, novel exemplar-based interlayer coding tools are also proposed, namely: 1) a direct prediction based on exemplar texture samples from lower layers and 2) an interlayer compensated prediction using a reference picture that is built relying on an exemplar-based algorithm for texture synthesis. Experimental results demonstrate the advantages of the proposed scalable coding solution to cater to users with different preferences/requirements in terms of interaction functionalities, while providing better rate-distortion performance (independently of the optical setup used for acquisition) compared to HEVC and other scalable light field coding solutions in the literature.
Caroline Conti, Luís Ducla Soares, Paulo J. L. Nunes
IEEE Trans. Multim.1
2016 HEVC-based 3D holoscopic video coding using self-similarity compensated prediction
Caroline Conti, Luís Ducla Soares, Paulo J. L. Nunes
Signal Process. Image Commun.1
2013 Inter-Layer Prediction Scheme for Scalable 3-D Holoscopic Video Coding
abstract
Holoscopic imaging has recently become a prospective glassless 3-D technology conquering the attention of researchers seeking more realistic depth-illusion approaches. However, backward compatibility with legacy displays is crucial to progressively introduce this technology into the consumer market and to efficiently deliver 3-D holoscopic content to end-users. Therefore, this letter proposes a new display scalable coding solution for 3-D holoscopic based on an inter-layer prediction scheme that exploits the redundancy between multiview and 3-D holoscopic content representations. Experimental results show that this inter-layer prediction scheme integrated into the High Efficiency Video Coding (HEVC) is advantageous, always outperforming the simulcast approach.
Caroline Conti, Paulo J. L. Nunes, Luís Ducla Soares
IEEE Signal Process. Lett.1
2012 New HEVC prediction modes for 3D holoscopic video coding
abstract
Holoscopic imaging is an advantageous solution for glassless 3D video systems, which promises to revolutionize the 3D market in the near future. Besides freeing the user from wearing any viewing device, it supports full motion parallax, improving this way the users' viewing experience. However, in order to provide 3D holoscopic content with convenient visual quality in terms of resolution and 3D perception, ultra-high resolution acquisition and display devices are required. Consequently, efficient video coding tools to deal with this large amount of data become of paramount importance. The recent standardization project called High Efficiency Video Coding (HEVC) addresses the requirements of high resolution video coding, but does not yet address the specific characteristics of 3D holoscopic content. To remedy this situation, this paper proposes to incorporate new prediction modes in HEVC to explore the particular structure of 3D holoscopic content, in order to further improve the performance of HEVC for this type of content. Experimental results, based on the HEVC test model version 4.0 are presented and clearly show the advantages of using this approach.
Caroline Conti, Paulo J. L. Nunes, Luís Ducla Soares
ICIP1
2011 Spatial prediction based on self-similarity compensation for 3D holoscopic image and video coding
abstract
Holoscopic imaging, also known as integral imaging, provides a solution for glassless 3D, and is promising to change the market for 3D television. To start, this paper briefly describes the general concepts of holoscopic imaging, focusing mainly on the spatial correlations inherent to this new type of content, which appear due to the micro-lens array that is used for both acquisition and display. The micro-images that are formed behind each micro-lens, from which only one pixel is viewed from a given observation point, have a high cross-correlation between them, which can be exploited for coding. A novel scheme for spatial prediction, exploring the particular arrangement of holoscopic images, is proposed. The proposed scheme can be used for both still image coding and intra-coding of video. Experimental results based on an H.264/AVC video codec modified to handle 3D holoscopic images and video are presented, showing the superior performance of this approach.
Caroline Conti, João Lino, Paulo J. L. Nunes, Luís Ducla Soares, Paulo Lobato Correia
ICIP1