Ferran Marqués

dblp:42/2845 · DBLP profile ↗
← Back
93ranked-venue papers
17as first author
5since 2021 · last 2025
0000-0001-8311-1168ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 67 · 15 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 22 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 13 · 3 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2
YearPublicationVenuePosition
2025 BiomSHARP: Biomass Super-Resolution for High Accuracy Prediction
abstract
Accurate estimation of above-ground biomass (AGB) is essential to understanding carbon stocks and flows, monitoring forest health, assessing biodiversity, and tracking ecological disturbances, which together help to inform climate policies. Imminent global satellite biomass missions (such as ESA’s BIOMASS and NASA-ISRO’s NISAR satellites) will offer valuable environmental monitoring, but their low spatial resolution limits their application in detailed local assessments. In this study, we present BiomSHARP (Biomass Super-resolution for High Accuracy Prediction), a deep learning (DL) model that extends the Hierarchical Attention Transformer (HAT) architecture adapting it to enhance coarse-resolution biomass maps by fusing them with high-resolution multispectral data from sensors such as Sentinel-2 or Landsat. BiomSHARP achieves 25-meter biomass predictions—four times the spatial resolution of the input—bridging the gap between global-scale monitoring and local-scale applications. In a first set of experiments, conducted in a local area in Europe, we demonstrate that BiomSHARP outperforms both traditional interpolation methods and state-of-the-art DL interpolation and prediction approaches for high-resolution AGB estimation across all evaluated metrics (MAE, MSE, RMSE, PSNR and SSIM), while using a comparable/lower number of parameters. Moreover, the model exhibits strong global-scale generalization, as demonstrated by its ability to accurately estimate biomass across diverse climatic regions despite being trained on a limited subset of data. Furthermore, the model presents strong temporal generalization, achieving improved performance in estimating AGB from 2020 data even when trained solely on 2010 data. We also analyze the impact of different combinations of spectral bands on biomass estimation, identifying optimal subsets that reduce redundancy and improve computational efficiency. BiomSHARP represents a promising approach to advance global environmental assessments and support improved climate strategies. The code and models are publicly available at https://github.com/laiaalbors/biomsharp.
Laia Albors, Javier Marcello, Ferran Marqués
IEEE Trans. Geosci. Remote. Sens.3
2023 Assessing Tree-Based Phenotype Prediction on the UK Biobank
abstract
Precision medicine relies on the ability to identify associations between genomic data and its phenotypic expression in order to provide personalized predictions. Phenotype prediction using statistical models trained on large-scale genomic and phenotypic data is a critical research area at the intersection of machine learning and genomics. Current genotype-to-phenotype models, such as polygenic risk scores, only account for linear relationships, and the use of nonlinear methods is still partially unexplored. In this work, we evaluate the prediction accuracy and scalability of nine nonlinear decision tree-based algorithms, including ensembling and boosting mechanisms, and compare them to linear prediction models. We assess the prediction performance for 24 anthropometric and disease-related phenotypes present in the UK Biobank. By using random feature selection, we explore how accuracy and computational time vary for each method as a function of the number of genetic variants selected. Our results show that tree-based methods, especially gradientboosted trees, can offer superior predictions with computational times comparable to those of linear methods. Thus, models able to capture nonlinear relationships between genotypes and phenotypes merit consideration for integration in upcoming computational systems for personalized medicine.
Alex Meléndez, Cayetana López, David Bonet, Gerard Sant, Daniel Mas Montserrat, Jordi Abante, Manuel Rivas Pérez, Ferran Marqués, Alexander G. Ioannidis
BIBM8
2023 Assessment of Forest Degradation Using Multitemporal and Multisensor Very High Resolution Satellite Imagery
abstract
The reliable detection of vegetation disease and plant stress are challenges in forest ecosystems. To address this problem, remote sensing existing methods of detection mostly rely on vegetation indices, however, in dense forest, the spectral saturation must be considered to select the most appropriate index. In this work, after a revision of the state of the art, a total of 20 vegetation indices were preliminary selected to perform a thorough statistical analysis with the aim to identify the disease and devitalization phenomena in a complex laurel forest. Multisensor very high resolution imagery, from the same month, with a time difference of a decade have been used. A robust methodology has been implemented to generate accurate vigor maps and to identify the forest areas that have experienced a degradation in plant health after 10 years.
Javier Marcello, Francisco Eugenio, Dionisio Rodríguez-Esparragón, Ferran Marqués
IGARSS4
2022 Multiple Object Tracking from appearance by hierarchically clustering tracklets
Andreu Girbau-Xalabarder, Ferran Marqués, Shin'ichi Satoh 0001
BMVC2
2022 High-Resolution Satellite Bathymetry Mapping: Regression and Machine Learning-Based Approaches
abstract
Remote spectral imaging of coastal areas can provide valuable information for their sustainable management and conservation of their biodiversity. Unfortunately, such areas are very sensitive to changes due to human activity, natural phenomenon, introduction of non-native species, and climate change. Thus, the main objective of this research is the implementation of a robust image processing methodology to produce accurate bathymetry maps in shallow coastal waters using high-resolution multispectral WorldView-2/3 satellite imagery for the monitoring at the maximum spatial and spectral resolutions. Two different island ecosystems have been selected for the assessment, since they stand out for their richness in endemic species and they are more vulnerable to climate change: Cabrera National Park and Maspalomas Natural Protected area, located in the Balearic and Canary Islands, Spain, respectively. In addition, a third example to show the applicability of the mapping methodology to monitor the construction of a new port in Granadilla (Canary Islands) is presented. Contributions of this work focus on improving the preprocessing methodology and, mainly, on the proposal and assessment of new satellite-derived regression and machine learning bathymetric models, which have been validated and compared with respect to measured reference bathymetry. After a thorough analysis of nine techniques, using visual and quantitative statistical parameters, ensemble learning approaches have demonstrated excellent performance, even in challenging scenarios up to 35-m depth, with mean RMSE values around 2 m.
Francisco Eugenio, Javier Marcello, Antonio Mederos-Barrera, Ferran Marqués
IEEE Trans. Geosci. Remote. Sens.4
2019 RVOS: End-To-End Recurrent Network for Video Object Segmentation
abstract
Multiple object video object segmentation is a challenging task, specially for the zero-shot case, when no object mask is given at the initial frame and the model has to find the objects to be segmented along the sequence. In our work, we propose a Recurrent network for multiple object Video Object Segmentation (RVOS) that is fully end-to-end trainable. Our model incorporates recurrence on two different domains: (i) the spatial, which allows to discover the different object instances within a frame, and (ii) the temporal, which allows to keep the coherence of the segmented objects along time. We train RVOS for zero-shot video object segmentation and are the first ones to report quantitative results for DAVIS-2017 and YouTube-VOS benchmarks. Further, we adapt RVOS for one-shot video object segmentation by using the masks obtained in previous time steps as inputs to be processed by the recurrent module. Our model reaches comparable results to state-of-the-art techniques in YouTube-VOS benchmark and outperforms all previous video object segmentation methods not using online learning in the DAVIS-2017 benchmark. Moreover, our model achieves faster inference runtimes than previous methods, reaching 44ms/frame on a P100 GPU.
Carles Ventura, Miriam Bellver, Andreu Girbau-Xalabarder, Amaia Salvador, Ferran Marqués, Xavier Giró-i-Nieto
CVPR5
2019 Bathymetry Mapping using very High Resolution Satellite Multispectral Imagery in Shallow Coastal Waters of Protected Ecosystems
abstract
Remote sensing of coastal areas requires multispectral satellite images with high spatial resolution. In this sense, WorldView-2 is a very high resolution satellite, which provides an advanced multispectral sensor with eight narrow bands, allowing the proliferation of new environmental monitoring and mapping applications in shallow coastal ecosystems. The problem of estimating water depths using a radiative model has yielded good results as it considers the physical phenomena of water absorption-backscattering and the relationship between the albedo of the seafloor and the reflectivity of the shallow waters. The sophisticated model developed and evaluated in this study expands the ratio algorithm model allowing for the increased amount of information provided in WorldView-2 imagery to be included in the retrieval of water depth of shallow coastal waters.
Ferran Marqués, Francisco Eugenio, Monica Alfaro, Javier Marcello
IGARSS1
2019 Multiresolution co-clustering for uncalibrated multiview segmentation
Carles Ventura, David Varas, Verónica Vilaplana, Xavier Giró-i-Nieto, Ferran Marqués
Signal Process. Image Commun.5
2018 Benthic Mapping Using High Resolution Multispectral and Hyperspectral Imagery
abstract
Coastal ecosystems are essential due to their high biodiversity and primary production, however they are extremely complex and with high spatial and temporal variability. Thus, to properly manage them it is necessary a systematic monitoring. Remote sensing can be very useful due to the spatial and spectral improvement of satellites and the availability of airborne or drone hyperspectral sensors. Unfortunately, the mapping of coastal areas is challenging due to the low SNR received at the sensor, as a consequence of the minimum reflectivity of the seafloor and the atmospheric and water column disturbances. In this context, the goal of this work is to obtain a robust classification methodology to generate accurate benthic habitat maps applying object-oriented and pixel-based classification methods in shallow waters using WorldView-2 and AHS (Airborne Hyperspectral Scanner) images. Maspalomas (Gran Canaria, Spain) was studied due to its complexity and the presence of important seagrass meadows.
Javier Marcello, Francisco Eugenio, Ferran Marqués
IGARSS3
2018 3D hierarchical optimization for multi-view depth map coding
Marc Maceira, David Varas, Ramon Morros, Javier Ruiz Hidalgo, Ferran Marqués
Multim. Tools Appl.5
2017 Multiscale Combinatorial Grouping for Image Segmentation and Object Proposal Generation
abstract
We propose a unified approach for bottom-up hierarchical image segmentation and object proposal generation for recognition, called Multiscale Combinatorial Grouping (MCG). For this purpose, we first develop a fast normalized cuts algorithm. We then propose a high-performance hierarchical segmenter that makes effective use of multiscale information. Finally, we propose a grouping strategy that combines our multiscale regions into highly-accurate object proposals by exploring efficiently their combinatorial space. We also present Single-scale Combinatorial Grouping (SCG), a faster version of MCG that produces competitive proposals in under five seconds per image. We conduct an extensive and comprehensive empirical validation on the BSDS500, SegVOC12, SBD, and COCO datasets, showing that MCG produces state-of-the-art contours, hierarchical regions, and object proposals.
Jordi Pont-Tuset, Pablo Andrés Arbeláez, Jonathan T. Barron, Ferran Marqués, Jitendra Malik
IEEE Trans. Pattern Anal. Mach. Intell.4
2016 Bags of Local Convolutional Features for Scalable Instance Search
abstract
This work proposes a simple instance retrieval pipeline based on encoding the convolutional features of CNN using the bag of words aggregation scheme (BoW). Assigning each local array of activations in a convolutional layer to a visual word produces an assignment map, a compact representation that relates regions of an image with a visual word. We use the assignment map for fast spatial reranking, obtaining object localizations that are used for query expansion. We demonstrate the suitability of the BoW representation based on local CNN features for instance retrieval, achieving competitive performance on the Oxford and Paris buildings benchmarks. We show that our proposed system for CNN feature aggregation with BoW outperforms state-of-the-art techniques using sum pooling at a subset of the challenging TRECVid INS benchmark.
Eva Mohedano, Kevin McGuinness, Noel E. O'Connor, Amaia Salvador, Ferran Marqués, Xavier Giró-i-Nieto
ICMR5
2016 Supervised Evaluation of Image Segmentation and Object Proposal Techniques
abstract
This paper tackles the supervised evaluation of image segmentation and object proposal algorithms. It surveys, structures, and deduplicates the measures used to compare both segmentation results and object proposals with a ground truth database; and proposes a new measure: the precision-recall for objects and parts. To compare the quality of these measures, eight state-of-the-art object proposal techniques are analyzed and two quantitative meta-measures involving nine state of the art segmentation methods are presented. The meta-measures consist in assuming some plausible hypotheses about the results and assessing how well each measure reflects these hypotheses. As a conclusion of the performed experiments, this paper proposes the tandem of precision-recall curves for boundaries and for objects-and-parts as the tool of choice for the supervised evaluation of image segmentation. We make the datasets and code of all the measures publicly available.
Jordi Pont-Tuset, Ferran Marqués
IEEE Trans. Pattern Anal. Mach. Intell.2
2015 Multiresolution Hierarchy Co-Clustering for Semantic Segmentation in Sequences with Small Variations
abstract
This paper presents a co-clustering technique that, given a collection of images and their hierarchies, clusters nodes from these hierarchies to obtain a coherent multiresolution representation of the image collection. We formalize the co-clustering as Quadratic Semi-Assignment Problem and solve it with a linear programming relaxation approach that makes effective use of information from hierarchies. Initially, we address the problem of generating an optimal, coherent partition per image and, afterwards, we extend this method to a multiresolution framework. Finally, we particularize this framework to an iterative multiresolution video segmentation algorithm in sequences with small variations. We evaluate the algorithm on the Video Occlusion/Object Boundary Detection Dataset, showing that it produces state-of-the-art results in these scenarios.
David Varas, Monica Alfaro, Ferran Marqués
ICCV3
2015 Improving spatial codification in semantic segmentation
abstract
This paper explores novel approaches for improving the spatial codification for the pooling of local descriptors to solve the semantic segmentation problem. We propose to partition the image into three regions for each object to be described: Figure, Border and Ground. This partition aims at minimizing the influence of the image context on the object description and vice versa by introducing an intermediate zone around the object contour. Furthermore, we also propose a richer visual descriptor of the object by applying a Spatial Pyramid over the Figure region. Two novel Spatial Pyramid configurations are explored: Cartesian-based and crown-based Spatial Pyramids. We test these approaches with state-of-the-art techniques and show that they improve the Figure-Ground based pooling in the Pascal VOC 2011 and 2012 semantic segmentation challenges.
Carles Ventura, Xavier Giró-i-Nieto, Verónica Vilaplana, Kevin McGuinness, Ferran Marqués, Noel E. O'Connor
ICIP5
2015 Precise classification of coastal benthic habitats using high resolution Worldview-2 imagery
abstract
The analysis of the seafloor in shallow waters using remote sensing imagery at very high spatial resolution is a very challenging topic due to the minimum signal level received; the presence of noisy contributions from the atmosphere, solar reflection, foam, turbidity and water column; and the limited spectral information available for the classification at such depths that impedes, for example, the extraction of vegetation indices. In this complex scenario we have developed a mapping methodology that involves the precise application of pre-processing techniques and the use of efficient classification algorithms. In particular, after a detailed assessment, support vector machines achieved the best performance using the appropriate kernel and parameters. Two natural areas located at the Canary Islands (Spain) have been selected for their benthic habitats richness and specially for their preservation of highly protected seagrass regions.
Javier Marcello, Francisco Eugenio, Ferran Marqués, Javier Martín Abasolo
IGARSS3
2014 Multiscale Combinatorial Grouping
abstract
We propose a unified approach for bottom-up hierarchical image segmentation and object candidate generation for recognition, called Multiscale Combinatorial Grouping (MCG). For this purpose, we first develop a fast normalized cuts algorithm. We then propose a high-performance hierarchical segmenter that makes effective use of multiscale information. Finally, we propose a grouping strategy that combines our multiscale regions into highly-accurate object candidates by exploring efficiently their combinatorial space. We conduct extensive experiments on both the BSDS500 and on the PASCAL 2012 segmentation datasets, showing that MCG produces state-of-the-art contours, hierarchical regions and object candidates.
Pablo Andrés Arbeláez, Jordi Pont-Tuset, Jonathan T. Barron, Ferran Marqués, Jitendra Malik
CVPR4
2014 Region-Based Particle Filter for Video Object Segmentation
abstract
We present a video object segmentation approach that extends the particle filter to a region-based image representation. Image partition is considered part of the particle filter measurement, which enriches the available information and leads to a re-formulation of the particle filter. The prediction step uses a co-clustering between the previous image object partition and a partition of the current one, which allows us to tackle the evolution of non-rigid structures. Particles are defined as unions of regions in the current image partition and their propagation is computed through a single co-clustering. The proposed technique is assessed on the SegTrack dataset, leading to satisfactory perceptual results and obtaining very competitive pixel error rates compared with the state-of-the-art methods.
David Varas, Ferran Marqués
CVPR2
2014 Fast generation of LULC maps for temporal studies in North-Western Africa
abstract
This paper provides an objective evaluation of six supervised classification techniques and three state of the art features, with the objective of obtaining a single combination of them that provides both robustness and objective performance improvements. As a conclusion, a simple procedure for obtaining LULC maps with four targeted classes is proposed.
Pol del Aguila Pla, Felipe Calderero, Ferran Marqués, Javier Marcello, Francisco Eugenio
IGARSS3
2014 Analysis of urban and vegetation growing in NW Senegal during the last 25 years using medium resolution imagery
abstract
Land use and land cover information are key information for Governments in developing countries. In this context, remote sensing satellites like Landsat or SPOT can provide valuable data covering several decades. We have developed a methodology to generate land cover maps with the aim to analyze changes in the last 25 years in the NW region of Senegal. In particular, we have applied the radiometric and atmospheric corrections prior to the classification algorithm or to the generation of vegetation indexes and, finally, to analyze the spatial and temporal variability, post-classification change detection techniques have been applied, providing valuable quantitative information.
Javier Marcello, Felipe Calderero, Francisco Eugenio, Ferran Marqués
IGARSS4
2014 Improving retrieval accuracy of Hierarchical Cellular Trees for generic metric spaces
Carles Ventura, Verónica Vilaplana, Xavier Giró-i-Nieto, Ferran Marqués
Multim. Tools Appl.4
2013 Measures and Meta-Measures for the Supervised Evaluation of Image Segmentation
abstract
This paper tackles the supervised evaluation of image segmentation algorithms. First, it surveys and structures the measures used to compare the segmentation results with a ground truth database, and proposes a new measure: the precision-recall for objects and parts. To compare the goodness of these measures, it defines three quantitative meta-measures involving six state of the art segmentation methods. The meta-measures consist in assuming some plausible hypotheses about the results and assessing how well each measure reflects these hypotheses. As a conclusion, this paper proposes the precision-recall curves for boundaries and for objects-and-parts as the tool of choice for the supervised evaluation of image segmentation. We make the datasets and code of all the measures publicly available.
Jordi Pont-Tuset, Ferran Marqués
CVPR2
2013 Real-time user independent hand gesture recognition from time-of-flight camera video using static and dynamic models
Javier Molina, Marcos Escudero-Viñolo, Alessandro Signoriello, Montse Pardàs, Christian Ferran Bennström, Jesús Bescós, Ferran Marqués, José María Martínez Sanchez
Mach. Vis. Appl.7
2012 Supervised Assessment of Segmentation Hierarchies
Jordi Pont-Tuset, Ferran Marqués
ECCV (4)2
2012 Upper-bound assessment of the spatial accuracy of hierarchical region-based image representations
abstract
Hierarchical region-based image representations are versatile tools for segmentation, filtering, object detection, etc. The evaluation of their spatial accuracy has been usually performed assessing the final result of an algorithm based on this representation. Given its wide applicability, however, a direct supervised assessment, independent of any application, would be desirable and fair. A brute-force assessment of all the partitions represented in the hierarchical structure would be a correct approach, but as we prove formally, it is computationally unfeasible. This paper presents an efficient algorithm to find the upper-bound performance of the representation and we show that the previous approximations in the literature can fail at finding this bound.
Jordi Pont-Tuset, Ferran Marqués
ICASSP2
2012 A region-based particle filter for generic object tracking and segmentation
abstract
In this work we present a region-based particle filter for generic object tracking and segmentation. The representation of the object in terms of regions homogeneous in color allows the proposed algorithm to robustly track the object and accurately segment its shape along the sequence. Moreover, this segmentation provides a mechanism to update the target model and allows the tracker to deal with color and shape variations of the object. The performance of the algorithm has been tested using the LabelMe Video public database. The experiments show satisfactory results in both tracking and segmentation of the object without an important increase of the computational time due to an efficient computation of the image partition.
David Varas, Ferran Marqués
ICIP2
2012 Hierarchical Navigation and Visual Search for Video Keyframe Retrieval
Carles Ventura, Manel Martos, Xavier Giró-i-Nieto, Verónica Vilaplana, Ferran Marqués
MMM5
2012 Multiview depth coding based on combined color/depth segmentation
Javier Ruiz Hidalgo, Ramon Morros, P. Aflaki, Felipe Calderero, Ferran Marqués
J. Vis. Commun. Image Represent.5
2012 Multispectral Cooperative Partition Sequence Fusion for Joint Classification and Hierarchical Segmentation
abstract
In this letter, a region-based fusion methodology is presented for joint classification and hierarchical segmentation of specific ground cover classes from high-spatial-resolution remote sensing images. Multispectral information is fused at the partition level using nonlinear techniques, which allows the different relevance of the various bands to be fully exploited. A hierarchical segmentation is performed for each individual band, and the ensuing segmentation results are fused in an iterative and cooperative way. At each iteration, a consensus partition is obtained based on information theory and is combined with a specific ground cover classification. Here, the proposed approach is applied to the extraction and segmentation of vegetation areas. The result is a hierarchy of partitions with the most relevant information of the vegetation areas at different levels of resolution. This system has been tested for vegetation analysis in high-spatial-resolution images from the QuickBird and GeoEye satellites.
Felipe Calderero, Francisco Eugenio, Javier Marcello, Ferran Marqués
IEEE Geosci. Remote. Sens. Lett.4
2011 Diversity ranking for video retrieval from a broadcaster archive
abstract
Video retrieval through text queries is a very common practice in broadcaster archives. The query keywords are compared to the metadata labels that documentalists have previously associated to the video assets. This paper focuses on a ranking strategy to obtain more relevant keyframes among the top hits of the results ranked lists but, at the same time, keeping a diversity of video assets. Previous solutions based on a random walk over a visual similarity graph have been modified to increase the asset diversity by filtering the edges between keyframes depending on their asset. The random walk algorithm is applied separately for ever visual feature to avoid any normalization issue between visual similarity metrics. Finally, this work evaluates performance with two separate metrics: the relevance is measured by the Average Precision and the diversity is assessed by the Average Diversity, a new metric presented in this work.
Xavier Giró-i-Nieto, Monica Alfaro, Ferran Marqués
ICMR3
2010 Region merging parameter dependency as information diversity to create sparse hierarchies of partitions
abstract
Region merging techniques usually include parameters that may be used to optimize or adapt the algorithm to a specific image type. Although, an appropriate tuning may provide a significant improvement, it also introduces a severe performance dependency on the parameter setting. The goal of this work is to transform the parameter dependency into an increase of accuracy and stability of the segmentation results. The idea is to use different parameter settings as specific type of diversity in an information fusion process based on a cooperative region merging approach. The potential of this parameter removal strategy is objectively evaluated on a set of state-of-the-art information theoretical region merging techniques for the removal of parameters: (i) in the region model, and (ii) in the merging order.
Felipe Calderero, Ferran Marqués
ICIP2
2010 Contour detection using Binary Partition Trees
abstract
Contour detection is a hard, challenging, and of paramount importance problem in image processing. State-of-the-art algorithms are approaching human performance but usually entail complex and tailored image models and arduous training. Binary Partition Tree is a hierarchical region-based image model that has been proven to have a wide range of applications in image filtering, information retrieval, object detection, etc. In this paper we propose a contour detection technique based on this versatile image model, extracting the contour information available in the tree yet outperforming one of the most widely-used contour detector.
Jordi Pont-Tuset, Ferran Marqués
ICIP2
2010 Object detection and segmentation on a hierarchical region-based image representation
abstract
In this paper we present a general framework for object detection and segmentation. Using a bottom-up unsupervised merging algorithm, a region-based hierarchy that represents the image at different resolution levels is created. Next, top-down, object class knowledge is used to select and combine regions from the hierarchy, in order to define the exact object shape. We illustrate the usefulness of the approach with four different object classes: sky, caption text, traffic signs and faces.
Verónica Vilaplana, Ferran Marqués, Miriam Leon, Antoni Gasull
ICIP2
2010 GAT: a Graphical Annotation Tool for semantic regions
Xavier Giró-i-Nieto, Neus Camps, Ferran Marqués
Multim. Tools Appl.3
2010 Region Merging Techniques Using Information Theory Statistical Measures
abstract
The purpose of the current work is to propose, under a statistical framework, a family of unsupervised region merging techniques providing a set of the most relevant region-based explanations of an image at different levels of analysis. These techniques are characterized by general and nonparametric region models, with neither color nor texture homogeneity assumptions, and a set of innovative merging criteria, based on information theory statistical measures. The scale consistency of the partitions is assured through i) a size regularization term into the merging criteria and a classical merging order, or ii) using a novel scale-based merging order to avoid the region size homogeneity imposed by the use of a size regularization term. Moreover, a partition significance index is defined to automatically determine the subset of most representative partitions from the created hierarchy. Most significant automatically extracted partitions show the ability to represent the semantic content of the image from a human point of view. Finally, a complete and exhaustive evaluation of the proposed techniques is performed, using not only different databases for the two main addressed problems (object-oriented segmentation of generic images and texture image segmentation), but also specific evaluation features in each case: under- and oversegmentation error, and a large set of region-based, pixel-based and error consistency indicators, respectively. Results are promising, outperforming in most indicators both object-oriented and texture state-of-the-art segmentation techniques.
Felipe Calderero, Ferran Marqués
IEEE Trans. Image Process.2
2009 Hierarchical fusion of color and depth information at partition level by cooperative region merging
abstract
A high level scheme for information fusion to create hierarchical region-based image representations based on a region merging process is presented. The strategy is based on an iterative evolution where the different merging criteria work independently and cooperate at the partition level to obtain a further consensus that increases the reliability of the resulting partitions. This cooperative scheme is applied to the creation of hierarchical region-based representations of the image based on color and depth information. The proposed technique is compared with approaches using only one source of information or linear combinations of both, in datasets with ground truth as well as estimated disparity information.
Felipe Calderero, Ferran Marqués
ICASSP2
2009 Performance evaluation of probability density estimators for unsupervised information theoretical region merging
abstract
Information theoretical region merging techniques have been shown to provide a state-of-the-art unified solution for natural and texture image segmentation. Here, we study how the segmentation results can be further improved by a more accurate estimation of the statistical model characterizing the regions. Concretely, we explore four density estimators that can be used for pdf or joint pdf estimation. The first three are based on different quantization strategies: a general uniform quantization, an MDL-based uniform quantization, and a data-dependent partitioning and estimation. The fourth strategy is based on a computationally efficient kernel-based estimator (averaged shifted histogram). Finally, all estimators are objectively evaluated using a database with available ground truth partitions.
Felipe Calderero, Ferran Marqués, Antonio Ortega
ICIP2
2009 Caption text extraction for indexing purposes using a hierarchical region-based image model
abstract
This paper presents a technique for detecting caption text for indexing purposes. This technique is to be included in a generic indexing system dealing with other semantic concepts. The various object detection algorithms are required to share a common image description which, in our case, is a hierarchical region-based image model. Caption text objects are detected combining texture and geometric features, which are estimated using wavelet analysis and taking advantage of the region-based image model, respectively. Analysis of the region hierarchy provides the final caption text objects.
Miriam Leon, Verónica Vilaplana, Antoni Gasull, Ferran Marqués
ICIP4
2009 Hierarchical Segmentation of Vegetation Areas in High Spatial Resolution Images by Fusion of Multispectral Information
abstract
A new region-based methodology for the automated extraction and hierarchical segmentation of vegetation areas into high spatial resolution images is proposed. This approach is based on the iterative and cooperative fusion of the independent segmentation results of equal or different resolution spectral bands, combined with an unsupervised classification into vegetation and no-vegetation regions. The result is a hierarchy of partitions with most relevant information at different levels of resolution of the vegetation areas. In addition, the high flexibility of the scheme allows different configurations depending on the final purpose. For instance, considering the size of the vegetation areas into the hierarchy, or prioritizing the information into the high resolution panchromatic band to improve the accuracy of both vegetation extraction and segmentation. This general tool for vegetation analysis is tested into high spatial resolution images from IKONOS and QuickBird satellites.
Felipe Calderero, Ferran Marqués, Javier Marcello, Francisco Eugenio
IGARSS (4)2
2009 Cloud Motion Estimation in SEVIRI Image Sequences
abstract
Determination of atmospheric dynamic characteristics from remote sensing imagery is fundamental in weather and climate studies. The SEVIRI radiometer, on board the MSG, with its 12 bands and 15 minutes sensing capability provides an important amount of information for cloud tracking. In this work, we have first conducted a detailed evaluation of twelve region matching techniques in order to select those providing the best results. For this performance evaluation, databases of synthetic and real sequences have been used. Next, the best metrics have been incorporated in a new methodology that includes a preliminary stage that segments cloudy structures to initialize the optimum motion estimation parameters (template size and search window dimensions). Also a study region mask is generated to disable the application of the motion estimation algorithm in unreliable areas, thus, eliminating erroneous vectors and decreasing the computation times.
Javier Marcello, Francisco Eugenio, Ferran Marqués
IGARSS (3)3
2009 Trajectory Tree as an Object-Oriented Hierarchical Representation for Video
abstract
This paper presents the trajectory tree as a hierarchical region-based representation for video sequences. Motion, as well as spatial features from multiple frames are used to generate a set of temporal regions structured within a hierarchy of scale and motion coherency. The resulting representations offer a global description of the entire video sequence and enhance semantic analysis potential. A multiscale segmentation strategy is proposed whereby region-merging criteria of progressively greater complexity are used to define partition layers of increasing aptitude for object detection. A novel data structure, called the trajectory adjacency graph, is defined for the long-term analysis of partition sequences. Furthermore, mechanisms for assessing connectivity, verifying temporal continuity, and proposing merging operations based on color, affine, and translational motion homogeneity characteristics over the entire sequence are also introduced. Finally, as demonstrated through experimental results, the trajectory tree offers a concise yet versatile support for video object segmentation, description and retrieval tasks.
Camilo C. Dorea, Montse Pardàs, Ferran Marqués
IEEE Trans. Circuits Syst. Video Technol.3
2008 General region merging approaches based on information theory statistical measures
abstract
This work presents a new statistical approach to region merging where regions are modeled as arbitrary discrete distributions, directly estimated from the pixel values. Under this framework, two region merging criteria are obtained from two different perspectives, leading to information theory statistical measures: the Kullback-Leibler divergence and the Bhattacharyya coefficient. The developed methods are size-dependent, which assures the size consistency of the partitions but reduces their size resolution. Thus, a size-independent extension of the previous methods, combined with a modified merging order, is also proposed. Additionally, an automatic criterion to select the most statistically significant partitions from the whole merging sequence is presented. Finally, all methods are evaluated and compared with other state-of-the-art region merging techniques.
Felipe Calderero, Ferran Marqués
ICIP2
2008 Region-based mean shift tracking: Application to face tracking
abstract
We present a new technique for object tracking that is an extension of the mean shift tracking algorithm. The proposed technique relies on a segmentation of the area under analysis into a set of color-homogenous regions. The use of regions allows a robust estimation of the likelihood distributions that form the object and background models, as well as a precise shape definition of the object being tracked. Thanks to this accurate object definition, the object model can be updated through the tracking process, handling variations in the object representation. These concepts have been tested in the case of tracking human faces.
Verónica Vilaplana, Ferran Marqués
ICIP2
2008 3D posture estimation using geodesic distance maps
Pedro Correa, Ferran Marqués, Xavier Marichal, Benoît Macq
Multim. Tools Appl.2
2008 Motion Estimation Techniques to Automatically Track Oceanographic Thermal Structures in Multisensor Image Sequences
abstract
The ocean involves a complex set of physical, chemical, biological, and geological processes, interacting with each other to influence our climate and natural environment. One of the most important disciplines in oceanography is the study of the ocean dynamics and, particularly, the ocean surface circulation. One can estimate this by the automated tracking of thermal infrared features in pairs of sequential satellite imagery. In this context, an extensive analysis of different motion estimation techniques has been performed by employing databases with synthetic sequences, real sequences, andinsitumeasurements. Four region- based metrics and two differential algorithms are proposed to estimate surface currents in multitemporal and multisensor AVHRR and MODIS image sequences. Once the appropriate motion estimation techniques have been selected, a new methodology to compute ocean currents is proposed. It includes a preliminary step to precisely segment the oceanographic structures and a second step to track its motion using additional modules (initialization, preprocessing, and postprocessing) to increase effectiveness. The information provided by the segmentation step reduces computing times, initializes the motion estimation parameters with appropriate values, and increases the overall performance. In summary, this two-stage approach combines image processing tools and physical oceanography knowledge to achieve a good ocean current estimation.
Javier Marcello, Francisco Eugenio, Ferran Marqués, Alonso Hernandez-Guerra, Antoni Gasull
IEEE Trans. Geosci. Remote. Sens.3
2008 Binary Partition Trees for Object Detection
abstract
This paper discusses the use of Binary Partition Trees (BPTs) for object detection. BPTs are hierarchical region-based representations of images. They define a reduced set of regions that covers the image support and that spans various levels of resolution. They are attractive for object detection as they tremendously reduce the search space. In this paper, several issues related to the use of BPT for object detection are studied. Concerning the tree construction, we analyze the compromise between computational complexity reduction and accuracy. This will lead us to define two parts in the BPT: one providing accuracy and one representing the search space for the object detection task. Then we analyze and objectively compare various similarity measures for the tree construction. We conclude that different similarity criteria should be used for the part providing accuracy in the BPT and for the part defining the search space and specific criteria are proposed for each case. Then we discuss the object detection strategy based on BPT. The notion of node extension is proposed and discussed. Finally, several object detection examples illustrating the generality of the approach and its efficiency are reported.
Verónica Vilaplana, Ferran Marqués, Philippe Salembier
IEEE Trans. Image Process.2
2007 Multiple View Region Matching as a Lagrangian Optimization Problem
abstract
A method to establish correspondences between regions belonging to independent segmentations of multiple views of a scene is presented. The trade-off between color similarity and projective similarity of the matching regions is formulated in terms of a constrained optimization, analogous to a rate-distortion budget-constrained allocation problem, and solved using Lagrangian optimization techniques.
Felipe Calderero, Ferran Marqués, Antonio Ortega
ICASSP (1)2
2007 Hierarchical Partition-Based Representations for Image Sequences using Trajectory Merging Criteria
abstract
This paper describes a hierarchical analysis framework for image sequences. Region merging schemes traditionally used in the construction of partition hierarchies are extended to multiple frames using trajectory merging criteria. The merging criteria assess homogeneity among features throughout the entire sequence to recursively create partitions in the spatio-temporal domain. We propose similarity measures using long-term affine and translational motion features. Furthermore, the analysis of connectivity relations and the algorithm implementation over trajectory adjacency graphs allow the generation of partition sets containing temporally consistent objects characterized by coherent motion. Lastly, we introduce the novel trajectory tree as a single, hierarchical representation of the partitions generated for the complete sequence. Experimental results are provided, illustrating the usefulness of the approach.
Camilo C. Dorea, Montse Pardàs, Ferran Marqués
ICASSP (1)3
2007 On Building a Hierarchical Region-Based Representation for Generic Image Analysis
abstract
This paper studies the procedure to create a hierarchical region-based image representation aiming at generic image analysis. This study is carried out in the context of bottom-up segmentation algorithms and, specifically, using the Binary Partition Tree implementation. The different steps necessary to create a hierarchical region-based representation are analyzed; namely, (i) the creation of the initial partition in the hierarchy, which is split into the definition of the initial merging criterion and the proposal of a stopping criterion, and (ii) the merging criteria used to produce the different regions in the final hierarchical representation. For both steps, the proposed approach is assessed and compared with previous existing ones over a large data set using well-established partition-based metrics.
Verónica Vilaplana, Ferran Marqués
ICIP (4)2
2007 Methodology for the estimation of ocean surface currents using region matching and differential algorithms
abstract
The ocean involves a complex set of physical, chemical, biological and geological processes, interacting each other to influence our climate and natural environment. One of the most important disciplines in oceanography is the study of the ocean dynamics. Particularly, ocean surface circulation can be recovered by the automated tracking of thermal features (coastal upwellings, filaments and eddies) in pairs of sequential satellite imagery. This paper presents a new methodology for the automatic estimation of ocean surface currents using AVHRR and MODIS imagery. The technique is based on a two-step process: precise structure detection followed by the structure tracking. The precise detection methodology is, as well, composed by two modules. The first one with the goal to obtain a coarse structure segmentation and the second achieving the maximum detail of the structure. The tracking methodology estimates motion fields using region matching and differential techniques with additional modules to guarantee the optimum performance. Twelve matching metrics and four differential algorithms were implemented and tested. Information from the previous structure segmentation stage is of fundamental importance to initialize the optimum motion estimation parameters and to generate the corresponding study area for each sequence. The proposed methodology, combining the segmentation and tracking steps, has been extensively tested and it has demonstrated an adequate performance in the estimation of flow fields in multitemporal and multisensorial sequences of the NW African coast.
Javier Marcello, Francisco Eugenio, Ferran Marqués
IGARSS3
2007 Performance of region-based matching techniques to compute the ocean surface motion
abstract
The study of the ocean circulation is the central core of all dynamical oceanography. The routine derivation of sea surface temperature or infrared brightness temperatures has been used to estimate the surface circulation by calculating the motion of the thermal features (coastal upwellings, filaments and eddies) in successive images. To that respect, a number of authors have developed different methodologies to recover the motion field, but the most straightforward methods match patterns (points, borders or regions) in all possible subwindows of one image with those in the next image. The maximization of the normalized cross- correlation coefficient, known as the Maximum Cross-Correlation (MCC) technique, is the most popular region-based matching metric applied to compute ocean circulation. In this paper a careful analysis of different region matching techniques has been conducted and the performance achieved for each approach is presented. The assessment methodology uses a database of synthetic sequences, real sequences and in-situ speed measurements. After the qualitative and quantitative analysis, we can conclude that the best performance is achieved by ZSAD, ZSSD, NZSSD and NCC metrics. These metrics achieve, when applied to synthetic sequences, mean angular errors around 30deg and magnitude errors around 30% for the worst case. In general the flow field recovered by the 4 previous metrics, perfectly models the motion of the structures in real sequences. Finally, results obtained with comparison with ground-truth data suggest an underestimation in the computed velocity between 35%-45% but with a higher angular accuracy, achieving global errors around 30deg-50deg. To conclude, it is important to emphasize that the prevalent MCC method provides acceptable results but with more errors when compared with the four previous metrics over the three databases.
Javier Marcello, Francisco Eugenio, Ferran Marqués
IGARSS3
2007 Prefetching and Caching Strategies for Remote and Interactive Browsing of JPEG2000 Images
abstract
This paper considers the issues of scheduling and caching JPEG2000 data in client/server interactive browsing applications, under memory and channel bandwidth constraints. It analyzes how the conveyed data have to be selected at the server and managed within the client cache so as to maximize the reactivity of the browsing application. Formally, to render the dynamic nature of the browsing session, we assume the existence of a reaction model that defines when the user launches a novel command as a function of the image quality displayed at the client. As a main outcome, our work demonstrates that, due to the latency inherent to client/server exchanges, a priori expectation about future navigation commands may help to improve the overall reactivity of the system. In our study, the browsing session is defined by the evolution of a rectangular window of interest (WoI) along the time. At any given time, the WoI defines the position and the resolution of the image data to display at the client. The expectation about future navigation commands is then formalized based on a stochastic navigation model, which defines the probability that a given WoI is requested next, knowing previous WoI requests. Based on that knowledge, several scheduling scenarios are considered. The first scenario is conventional and transmits all the data corresponding to the current WoI before prefetching the most promising data outside the current WoI. Alternative scenarios are then proposed to anticipate prefetching, by scheduling data expected to be requested in the future before all the current WoI data have been sent out. Our results demonstrate that, for predictable navigation commands, anticipated prefetching improves the overall reactivity of the system by up to 30% compared to the conventional scheduling approach. They also reveal that an accurate knowledge of the reaction model is not required to get these significant improvements.
Antonin Descampe, Christophe De Vleeschouwer, Marcela Iregui, Benoît Macq, Ferran Marqués
IEEE Trans. Image Process.5
2007 Bayesian Approach for Morphology-Based 2-D Human Motion Capture
abstract
This paper presents a novel technique for 2D human motion capture using a single non calibrated camera. The user's five extremities (head, hands and feet) are extracted, labeled and tracked after silhouette segmentation. As they are the minimal number of points that can be used in order to enable whole body gestural interaction, we henceforth refer to these features as crucial points. The crucial point candidates are defined as the local maxima of the geodesic distance with respect to the center of gravity of the actor region that lie on the silhouette boundary. In order to disambiguate the selected crucial points into head, left and right foot and left and right hand classes, we propose a Bayesian framework that combines a MAP approach weighted by a prior model and the intensities of the tracked crucial points. Due to its low computational complexity, the system can run at real-time paces on standard personal computers, with an average error rate range between 2% and 7% in realistic situations, depending on the context and segmentation quality.
P. C. Correa Hernandez, Jacek Czyz, Ferran Marqués, Toshiyuki Umeda, Xavier Marichal, Benoît Macq
IEEE Trans. Multim.3
2006 Pre-Fetching Strategies for Remote and Interactive Browsing of JPEG2000 Images
abstract
This paper considers the remote interactive browsing of large JPEG2000 images. In contrast with previous contributions, we focus on the dynamic nature of the system. Practically, we study the conditions under which a priori knowledge about the user behavior may help to improve the browsing system reactivity, when combined with appropriate pre-fetching mechanisms. In particular, our simulations show that, due to the latency inherent to client/server exchanges, a benefit can be drawn by scheduling future window of interest (WoI) data before all current WoI data have been sent to the client. They also reveal that an accurate knowledge of the user behavior is not necessary to get important improvements over conventional scheduling approaches.
Antonin Descampe, Christophe De Vleeschouwer, Marcela Iregui, Benoît Macq, Ferran Marqués
ICIP5
2006 Generation of Long-Term Color and Motion Coherent Partitions
abstract
This paper describes a technique for generating partition sequences of regions presenting long-term homogeneity in color and motion coherency in terms of affine models. The technique is based on region merging schemes compatible with hierarchical representation frameworks and can be divided into two stages: partition tracking and partition sequence analysis. Partition tracking is a recursive algorithm whereby regions are constructed according to short-term spatio-temporal features, namely color and motion. Partition sequence analysis proposes the trajectory adjacency graph (TAG) to exploit the long-term connectivity relations of tracked regions. A novel trajectory merging strategy using color homogeneity criteria over multiple frames is introduced. Algorithm performance is assessed and comparisons to other proposals are drawn by means of established evaluation metrics.
Camilo C. Dorea, Montse Pardàs, Ferran Marqués
ICIP3
2006 Face Recognition using Groups of Images in Smart Room Scenarios
abstract
In this paper, we present a technique for face recognition in smart environments. The technique takes advantage of the continuous monitoring of the scenario and combines the information of several images to perform the recognition. Appearance based face recognition techniques are used given that unobtrusive systems are required in this type of applications and that the scenario does not ensure high quality images. Models for the users are created on-the-fly and subsequent face images of the same individual are gathered into groups. Images within a group are jointly compared to the models for identification and verification purposes. Reliable face images are used to update the user's models. The proposed technique is assessed with the BANCA and CHIL databases.
Verónica Vilaplana, Claudi Martinez, Javier Cruz, Ferran Marqués
ICIP4
2005 Silhouette-based probabilistic 2D human motion estimation for real-time applications
abstract
This paper presents a novel technique for 2D human motion estimation using a single non calibrated camera. The user's five crucial human features (head, hands and feet) are extracted, labeled and tracked, after silhouette segmentation. The crucial points candidates are defined as the local maxima of the geodesic distance with respect to the center of gravity of the actor region (silhouette) following the silhouette boundary. Selected crucial points are then classified as head, hands or feet using a probabilistic approach weighted by a prior human model. The system can run at 50 Hz paces on standard personal computers.
Pedro Correa, Jacek Czyz, Toshiyuki Umeda, Ferran Marqués, Xavier Marichal, Benoît Macq
ICIP (3)4
2005 A motion-based binary partition tree approach to video object segmentation
abstract
This paper describes an approach for generating binary partition tree [P. Salembier and L. Garrido, 2000] representations and video object segmentations using a novel region merging strategy based on motion similarity measures of multiple frames of an image sequence. The system operates over color-homogeneous regions, tracked across frames of a shot, representing an over-segmentation of the objects. A long-term motion similarity measure is introduced for region merging, offering accurate segmentation of objects and extending temporal consistency between the tracked partitions to hierarchical representations of every frame within the shot. Experimental results are presented, illustrating the usefulness of the approach.
Camilo C. Dorea, Montse Pardàs, Ferran Marqués
ICIP (2)3
2005 Detection of semantic objects using description graphs
abstract
This paper presents a technique to detect instances of classes (objects) according to their semantic definition in the form of a description graph. Classes are defined as combinations of instances of lower level semantic classes and allow the definition of a semantic tree that organizes classes in semantic levels. At the bottom level of the semantic tree, classes are defined by a perceptual model containing a list of low-level descriptors. The proposed detection algorithm follows a bottom-up/top-down approach, building semantic trees on a region-based representation of the media. The flexibility of the approach is assessed on different examples of planar objects, such as frontal faces, groups of islands, flags and traffic signs.
Xavier Giró-i-Nieto, Ferran Marqués
ICIP (1)2
2005 Automatic tool for the precise detection of upwelling and filaments in remote sensing imagery
abstract
The upward movement of cool and nutrient-rich waters toward the surface leads to horizontal alterations in the distribution of the physical, chemical, and biological properties. Remote sensing is being extensively applied to detect such coastal upwellings; however, the enormous amount of data daily generated obliges to develop automatic detection and prediction tools. The problem of identifying oceanographic mesoscale structures has been studied using a variety of image processing techniques; however, the outstanding difficulties encountered in the traditional approaches are the presence of noise, the fact that gradients are weak, the strong morphological variation, and the absence of a valid analytical model for the structures. In this context, the proposed automatic upwelling extraction methodology overcomes the preceding detection inconveniences and achieves a highly accurate structure extraction. This automatic technique is based on a coarse-segmentation methodology followed by a fine-detail growing process. The complete system has been validated over a database of 378 multisensorial images of years 2000 to 2003, and it has been applied to the detection and feature extraction of coastal upwellings and filaments in three areas with different characteristics, such as the Canary Islands, Cape Ghir, and the Alboran Sea, using imagery from the Advanced Very High Resolution Radiometer 2 and 3 sensors, the Sea-viewing Wide Field-of-view Sensor, and the Moderate Resolution Imaging Spectroradiometer sensor, demonstrating its effectiveness and robustness in a wide variety of climate conditions.
Javier Marcello, Ferran Marqués, Francisco Eugenio
IEEE Trans. Geosci. Remote. Sens.2
2004 Enhanced audio data hiding synchronization using non linear filters
abstract
The paper addresses the problem of synchronization in the context of audio data hiding. For real time transmission purposes, the data decoding process has to deal with synchronization issues. The paper proposes an new synchronization scheme that optimizes the performances of systems which are based on spread spectrum synchronization by the use of mathematical morphological tools that present good performance for peak detection. A brief theoretical presentation of the top-hat filter is recalled and the enhanced system is derived from the analysis of the advantages and disadvantages. The final scheme provides a robust synchronization system and is compared with the classical solutions.
Alejandro LoboGuerrero, Ferran Marqués, Patrick Bas, Joel Lienard
ICASSP (2)2
2004 Object recognition based on binary partition trees
abstract
This paper presents an object recognition method that exploits the representation of the images obtained by means of a binary partition tree (BPT). The shape matching technique in which it is based was first presented in F. Marques et al., (2002). This method compares a transformed version of an object shape model (reference contour) to the contours of a partition of the image. The comparison is based on a distance map that measures the Euclidean distance between any points in the image to the partition contours. In F. Marques et al., (2002), this algorithm was applied using a colour-based segmentation of the image and a full-search was performed to find the best match between the searched object and the contours of this segmentation. Here, the information of the binary partition tree is used both to obtain the segmentation and to guide and reduce the search for the optimum match between the shape and the objects of the image.
Oreste Salerno, Montse Pardàs, Verónica Vilaplana, Ferran Marqués
ICIP4
2004 An automated multisensor satellite imagery registration technique based on the optimization of contour features
abstract
Spatial registration of multidate or multisensor images is required for many applications in remote sensing. Automatic image registration, which has been extensively studied in other areas of image processing, is still a complex problem in the framework of remote sensing. This work explores an alternative strategy for a fully automatic and operational registration system capable of registering multitemporal and multisensor remote sensing satellite images with high accuracy and avoiding the use of ground control points, exploiting the maximum reliable information in both images (coastlines not occluded by clouds). The automatic feature-based approach is summarized as follows: (i) reference image coastline extraction; (ii) sensed image gradient energy map estimation and (iii) contour matching, mapping function estimation and transformation of the sensed images. Several experimental results for single sensor imagery (AVHRR/3) and multisensor imagery (AVHRR/3-SeaWiFS-MODIS-ATSR) from different viewpoints and dates have verified the robustness and accuracy of the proposed automatic registration algorithm, demonstrating its capability of registering satellite images of coastal areas within one pixel.
Francisco Eugenio, Javier Marcello, Ferran Marqués
IGARSS3
2004 Precise upwelling and filaments automatic extraction from multisensorial imagery
abstract
The upward movement of cool and nutrient-rich waters towards the surface leads to horizontal alterations in the distribution of physical, chemical and biological properties. Remote sensing is being extensively applied to detect such coastal upwellings; however, the enormous amount of data daily generated obliges to develop automatic detection and prediction tools. The problem of identifying oceanographic mesoscale structures has been studied using a variety of image processing techniques, however, the outstanding difficulties encountered in the traditional approaches are the presence of noise, mainly due to the clouds and other atmospheric phenomena; the fact that gradients are weak and provide excess of information; the strong morphological variation that impedes an accurate geometric representation and the absence of a valid analytical model for the structures. In this context, the proposed automatic upwelling extraction methodology overcomes the preceding detection inconveniences and achieves a highly accurate structure detection and identification. This automatic technique has been applied to the detection and feature extraction of coastal upwellings and filaments in the northwest African coast, the Alboran Sea and Cape Ghir using imagery from the AVHRR/2&3, SeaWiFS and MODIS sensors. The system has proven to be very effective and robust in a wide variety of climate conditions.
Javier Marcello, Francisco Eugenio, Ferran Marqués
IGARSS3
2003 Automatic structures detection and spatial registration using multisensor satellite imagery
abstract
Mesoscale processes such as upwellings, eddies, or thermal fronts are very energetic and their knowledge is very important not only to study oceanic circulation but also areas of applications that include acoustic propagation anomalies, fisheries management and exploitation, coastal monitoring and offshore or ocean oil detection and exploitation. A variety of techniques and algorithms have been developed to detect such structures. The foremost difficulties encountered in the preceding approaches are the presence of noise, mainly due to clouds and other atmospheric phenomena. In this context, the proposed methodology, due to its region-based nature, overcomes the edge detection inconveniences and obtains the proper structure identification. Moreover and in order to perform an exhaustive analysis of the structure dynamics, it is necessary to compare image sequences. In this context, the use of spatial registration techniques is necessary to achieve that pixels in different images correspond to the same geographic region. An automatic contour based approach for high accuracy registration of multisensor and multitemporal remote sensing images is presented. It avoids the use of ground control points, while exploiting the maximum reliable information in both images. These automatic tools, that combine structures detection techniques and multitemporal and multisensoral registration, have been applied to AVHRR, SeaWiFS and MODIS images of the Canary Island and Alboran Sea areas and have demonstrated that it can be a fundamental tool to validate marine and coastal dynamic studies using remote sensing data.
Francisco Eugenio, Eduardo Rovaris, Javier Marcello, Ferran Marqués
IGARSS4
2003 Automatic satellite image georeferencing using a contour-matching approach
abstract
Multitemporal and multisatellite studies or comparisons between satellite data and local ground measurements require nowadays precise and automatic geometric correction of satellite images. This paper presents a fully automatic geometric correction system capable of georeferencing satellite images with high accuracy. An orbital prediction model, which provides initial earth locations, is combined with the proposed automatic contour-matching technique. This combination allows correcting the low-frequency error component, mainly due to timing and orbital model errors, as well as the high-frequency error component, due to variations in the spacecraft's attitude. The approach aims at exploiting the maximum reliable information in the image to guide the matching algorithm. The contour-matching process has three main steps: 1) estimation of the gradient energy map (edges) and detection of the cloudless (reliable) areas; 2) initialization of the contours positions; 3) estimation of the transformation parameters (affine model) using a contour optimization approach. Three different robust and automatic algorithms are proposed for optimization, and their main features are discussed. Finally, the performance of the three proposed algorithms is assessed using a new error estimation technique applied to Advanced Very High Resolution Radiometer (AVHRR), Sea-viewing Wide Field of view Sensor (SeaWiFS), and multisensor AVHRR-SeaWiFS imagery.
Francisco Eugenio, Ferran Marqués
IEEE Trans. Geosci. Remote. Sens.2
2002 Object matching based on partition information
abstract
This paper presents a new technique for object matching that exploits the information about transitions in the image obtained by means of a segmentation approach. Object matching is performed by comparing a transformed version of an object shape model (reference contour) to the contours in the image partition. The comparison is based on a distance map that measures the Euclidean distance between any point in the image to the partition contours. Examples using parametric and non-parametric reference contours are provided to assess the quality of the proposed technique.
Ferran Marqués, Montse Pardàs, Ramon Morros
ICIP (2)1
2002 A contour-based approach to automatic and accurate registration of multitemporal and multisensor satellite imagery
abstract
An automatic approach for high accuracy registration of multisensor and multitemporal remote sensing images is presented. It avoids the use of ground control points, while exploiting the maximum reliable information in both images. Features to be used for image registration are those contours in both images that have been classified as coastline (reliable information). The automatic contour-based approach is summarized by the following steps: (i) reference image coastline extraction; (ii) sensed image gradient energy map estimation; (iii) contour matching, mapping function estimation and transformation of the sensed images. The algorithm proposed is automatic and of significant value in an operational context. Several experimental results for single sensor imagery (AVHRR) from different viewpoints and dates as well as multisensor imagery (AVHRR-SeaWiFS) have verified the robustness and accuracy of the proposed automatic registration algorithm, demonstrating its capability of registering satellite images of coastal areas within one pixel.
Francisco Eugenio, Ferran Marqués, Javier Marcello
IGARSS2
2002 Automatic feature extraction from multisensorial oceanographic imagery
abstract
The problem of identifying mesoscale structures has been studied using a variety of image processing techniques, mainly, texture analysis, edge detection, mathematical morphology, neural networks and wavelet transform. The foremost difficulties encountered in the preceding approaches are the presence of noise, mainly due to clouds and other atmospheric phenomena; the fact that gradients are weak and provide excess of information; the strong morphological variation that impedes an accurate geometric representation and the absence of a valid analytical model for the structures. In this context, the proposed methodology, due to its region-based nature, overcomes the edge detection inconveniences and obtains the proper structure identification. This automatic technique has been applied to the detection and feature extraction of coastal upwellings and filaments in the northwest African coast and the Alboran Sea using imagery from the AVHRR/2&3, SeaWiFS and MODIS sensors. The system has proven to be very effective and robust in a wide variety of climate conditions.
Javier Marcello, Ferran Marqués, Francisco Eugenio
IGARSS2
2002 Face segmentation and tracking based on connected operators and partition projection
Ferran Marqués, Verónica Vilaplana
Pattern Recognit.1
2002 Image processing for 3D imaging
Ferran Marqués, Fabian Lavagetto, Michael G. Strintzis
Signal Process. Image Commun.1
2001 Pixel and sub-pixel accuracy in satellite image georeferencing using an automatic contour matching approach
abstract
This paper presents a technique for a fully automatic and operational geometric correction system capable of georeferencing satellite images with high accuracy. A simple Keplerian orbital satellite model is considered and mean orbital elements are given as input from ephemeris data. To correct the systematic errors caused by these simplifications, nonzero values for the spacecraft roll, pitch and yaw and failures in the satellite internal clock, an automatic global contour matching approach is proposed. It has three main steps: (i) estimation of the gradient energy map (edges) and detection of the cloudless (reliable) areas; (ii) initialization of the contour positions; (iii) obtaining the transformation parameters (affine model) by means of a global contour optimization approach. Three different algorithms are proposed for optimization. The performance of the overall technique is assessed using AVHRR and multisensor AVHRR-SeaWiFS imagery.
Francisco Eugenio, Ferran Marqués, Javier Marcello
ICIP (1)2
2000 Partition-Based Image Representation as Basis for User-Assisted Segmentation
abstract
This paper discusses the usefulness of a partition-based image representation in the context of user-assisted segmentation, and compares it with the two other most common image representations in this framework: pixel-based and transition-based. Partition-based image representations allow user-interaction at the region level, which is a very natural and friendly manner to interact with images. The paper describes various region-based implementations of the most common types of user interaction; namely, initial object selection, object refinement, imposing object characteristics, selection of specific objects, selection of objects that are similar to a previously selected one. In all cases, partition-based image representations, and their associated region-based tools for user interaction, show up to be very suitable for interactive image segmentation.
Ferran Marqués, Beatriz Marcotegui, M. Francisca Zanoguera, Paulo Lobato Correia, Roland Mech, Michael Wollborn
ICIP1
2000 A Morphological Approach for Segmentation and Tracking of Human Face
abstract
A new technique for segmenting and tracking human faces in video sequences is presented. The technique relies on morphological tools such as using connected operators to extract the connected component that more likely belongs to a face, and partition projection to track this component through the sequence. A binary partition tree (BPT) is used to implement the connected operator. The BPT is constructed based on the chrominance criteria and its nodes are analyzed so that the selected node maximizes an estimation of the likelihood of being part of a face. The tracking is performed using a partition projection approach. Images are divided into face and non-face parts, which are tracked through the sequence. The technique has been successfully assessed using several test sequences from the MPEG-4 (raw format) and the MPEG-7 databases (MPEG-1 format).
Ferran Marqués, Verónica Vilaplana
ICPR1
2000 A contour-based approach to binary shape coding using a multiple grid chain code
Paulo J. L. Nunes, Ferran Marqués, Fernando Pereira 0001, Antoni Gasull
Signal Process. Image Commun.2
1999 A Video Object Generation Tool Allowing Friendly User Interaction
abstract
In this paper we describe an interactive video object segmentation tool developed in the framework of the ACTS-AC098 MOMUSYS project. The Video Object Generator with User Environment (VOGUE) combines three different sets of automatic and semi-automatic-tool (spatial segmentation, object tracking and temporal segmentation) with general purpose tools for user interaction. The result is an integrated environment allowing the user-assisted segmentation of any sort of video sequences in a friendly and efficient manner.
Beatriz Marcotegui, Paulo Lobato Correia, Ferran Marqués, Roland Mech, Ricardo Rosa, Michael Wollborn, M. Francisca Zanoguera
ICIP (2)3
1999 Human Face Segmentation and Tracking Using Connected Operators and Partition Projection
abstract
A new technique for segmenting and tracking human faces in video sequences is presented. The algorithm uses a connected operator to extract the connected component that more likely belongs to a face. Such a connected operator is implemented by means of a binary partition tree. A set of connected regions (a node in the tree) is selected maximizing an estimation of the likelihood of being part of a face. Faces are tracked through the sequence based on the partition projection approach. A face and a non-face core component are obtained in the current image by projecting the previous partition. The technique has been successfully assessed using several test sequences from the MPEG-4 database (raw format) as well as from the MPEG-7 database (MPEG-1 format).
Ferran Marqués, Verónica Vilaplana, Anabel Buxes
ICIP (3)1
1999 A Proposal for Dependent Optimization in Scalable Region-Based Coding Systems
abstract
We address in this paper the problem of optimal coding in the framework of region-based video coding systems, with a special stress on content-based functionalities. We present a coding system that can provide scaled layers (using PSNR or temporal content-based scalability) such that each one has an optimal partition with optimal bit allocation among the resulting regions. This coding system is based on a dependent optimization algorithm that can provide joint optimality for a group of layers or a group of frames.
Ramon Morros, Ferran Marqués
ICIP (4)2
1999 Region-based representations of image and video: segmentation tools for multimedia services
abstract
This paper discusses region-based representations of image and video that are useful for multimedia services such as those supported by the MPEG-4 and MPEG-7 standards. Classical tools related to the generation of the region-based representations are discussed. After a description of the main processing steps and the corresponding choices in terms of feature spaces, decision spaces, and decision algorithms, the state of the art in segmentation is reviewed. Mainly tools useful in the context of the MPEG-4 and MPEG-7 standards are discussed. The review is structured around the strategies used by the algorithms (transition based or homogeneity based) and the decision spaces (spatial, spatio-temporal, and temporal). The second part of this paper proposes a partition tree representation of images and introduces a processing strategy that involves a similarity estimation step followed by a partition creation step. This strategy tries to find a compromise between what can be done in a systematic and universal way and what has to be application dependent. It is shown in particular how a single partition tree created with an extremely simple similarity feature can support a large number of segmentation applications: spatial segmentation, motion estimation, region-based coding, semantic object extraction, and region-based retrieval.
Philippe Salembier, Ferran Marqués
IEEE Trans. Circuits Syst. Video Technol.2
1998 Tracking of Generic Objects for Video Object Generation
Ferran Marqués, Joan Llach
ICIP (3)1
1998 Prediction of image partitions using Fourier descriptors: application to segmentation-based coding schemes
abstract
This paper presents a prediction technique for partition sequences. It uses a region-by-region approach that consists of four steps: region parameterization, region prediction, region ordering, and partition creation. The time evolution of each region is divided into two types: regular motion and shape deformation. Both types of evolution are parameterized by means of the Fourier descriptors and they are separately predicted in the Fourier domain. The final predicted partition is built from the ordered combination of the predicted regions, using morphological tools. With this prediction technique, two different applications are addressed in the context of segmentation-based coding approaches. Noncausal partition prediction is applied to partition interpolation, and examples using complete partitions are presented. In turn, causal partition prediction is applied to partition extrapolation for coding purposes, and examples using complete partitions as well as sequences of binary images--shape information in video object planes (VOPs)--are presented.
Ferran Marqués, Bernat Llorens, Antoni Gasull
IEEE Trans. Image Process.1
1997 Scalable Segmentation-Based Coding of Video Sequences Addressing Content-Based Functionalities
abstract
In this paper, we address video scalability in the framework of a region-based coding system, allowing content-based functionalities. The proposed algorithm can construct scaled layers from a video sequence, each one with either fixed bit-rate or fixed quality, allowing content based manipulation. Two modes of operation have been defined: a supervised mode, that allows the user to select the objects to be coded in the enhancement layer, and an unsupervised mode, where this selection is done by the algorithm itself.
Ramon Morros, Ferran Marqués
ICIP (2)2
1997 Multi-grid chain coding of binary shapes
abstract
This paper presents a chain code based approach to efficiently code binary shape information of video objects, in the context of object-based video coding. The proposed method tries to meet some of the requirements of the MPEG-4 standard, currently under development, notably efficient coding, and low delay. This approach allows several modes of operation depending on the application requirements, notably lossless, near-lossless, and lossy coding modes. For the lossless case a pure differential chain code method is proposed while for the near-lossless case a multi-grid chain code (MGCC) technique is adopted. Also both INTRA and INTER prediction modes can be used. For the INTER mode, motion compensation is applied without coding the residues. The MGCC is a near-lossless contour coding technique using a contour description based on edges, which combines both contour prediction and contour simplification.
Paulo J. L. Nunes, Fernando Pereira 0001, Ferran Marqués
ICIP (3)3
1997 Segmentation-based video coding system allowing the manipulation of objects
abstract
This paper presents a generic video coding algorithm allowing the content-based manipulation of objects. This manipulation is possible thanks to the definition of a spatiotemporal segmentation of the sequences. The coding strategy relies on a joint optimization in the rate-distortion sense of the partition definition and of the coding techniques to be used within each region. This optimization creates the link between the analysis and synthesis parts of the coder. The analysis defines the time evolution of the partition, as well as the elimination or the appearance of regions that are homogeneous either spatially or in motion. The coding of the texture as well as of the partition relies on region-based motion compensation techniques. The algorithm offers a good compromise between the ability to track and manipulate objects and the coding efficiency.
Philippe Salembier, Ferran Marqués, Montse Pardàs, Ramon Morros, Isabelle Corset, Sylvie Jeannin, Lionel Bouchard, Fernand Meyer, Beatriz Marcotegui
IEEE Trans. Circuits Syst. Video Technol.2
1996 Tracking areas of interest for content-based functionalities in segmentation-based video coding
abstract
This paper presents a technique for tracking areas of interest in the framework of segmentation-based video coding. The technique is independent of the type of segmentation technique used in the coding approach. Therefore, it can also be used in block-based coding schemes. The algorithm relies on a double segmentation of the image based on morphological tools. This double segmentation permits to obtain the position and shape of the previous area of interest in the current image. In order to demonstrate the potentialities of this algorithm, it is applied in a specific coding scheme so that content-based selective coding is allowed.
Ferran Marqués, Beatriz Marcotegui, Fernand Meyer
ICASSP1
1996 Partition tree for a segmentation-based video coding system
abstract
In this paper we describe a very low bit rate video coding system that belongs to the class of object based coding systems. The system takes advantage of homogeneity in both gray level and motion, trying to use the most adequate criterion for defining a partition in order to obtain the best compromise between cost and quality. In order to do so, multiple partitions are generated for every image in a hierarchical way, from the coarsest one to the finest one. Then, a decision step will chose which parts of the image need to be coded with regions from one partition or from another. The aim of the paper is to study the features that this multiple partition must have, and propose a way to create it.
Montse Pardàs, Philippe Salembier, Ferran Marqués, Ramon Morros
ICASSP3
1996 Partition coding using multi-grid chain code and motion compensation
abstract
In this paper, a lossy partition coding technique is presented which leads to decoded partitions with unnoticeable losses. It uses a hexagonal grid for contour representation and it is based on the concept of multi-grid chain code. This coding technique can be used in intra-frame mode or, in combination with partition prediction techniques, in interframe mode. In intra-frame mode, it leads to an average saving of 25% of the coding cost with respect to chain code techniques. The savings on inter-frame mode depend on the motion prediction approach. Results of coding binary shapes (concept of video object plane) as well as partitions with an arbitrary set of regions are presented.
Ferran Marqués, Antoni Gasull
ICIP (2)1
1996 A segmentation-based coding system allowing manipulation of objects (SESAME)
abstract
We present a coding scheme that achieves, for each image in the sequence, the best segmentation in terms of rate-distortion theory. It is obtained from a set of initial regions and a set of available coding techniques. The segmentation combines spatial and motion criteria. It selects at each area of the image the most adequate criterion for defining a partition in order to obtain the best compromise between cost and quality. In addition, the proposed scheme is very suitable for addressing content-based functionalities.
Ferran Marqués, Philippe Salembier, Montse Pardàs, Ramon Morros, Isabelle Corset, Sylvie Jeannin, Beatriz Marcotegui, Fernand Meyer
ICIP (3)1
1995 Interpolation and extrapolation of image partitions using Fourier descriptors: application to segmentation-based coding schemes
abstract
This paper presents an interpolation/extrapolation technique for sequence partitions. It consists of four steps: region parametrization, region interpolation, region ordering and partition creation. The evolution of each region is divided into two types: regular motion and random deformations. Both types of evolution are parametrized by means of the Fourier descriptors of the regions and they are separately interpolated in the Fourier domain. The final interpolated partition is built from the ordered combination of the interpolated regions, using morphological tools.
Ferran Marqués, Bernat Llorens, Antoni Gasull
ICIP (3)1
1994 Recursive image sequence segmentation by hierarchical models
abstract
This paper addresses the problem of image sequence segmentation. A technique using a sequence model based on compound random fields is presented. This technique is recursive in the sense that frames are processed in the same cadency as they are produced. New regions appearing in the sequence are detected by a morphological procedure.
Ferran Marqués, Victor Vera, Antoni Gasull
ICPR (1)1
1993 Unsupervised image segmentation controlled by morphological contrast extraction
Ferran Marqués, Jordi Cunillera, Antoni Gasull
ICASSP (5)1
1992 Hierarchical segmentation using compound Gauss-Markov random fields
abstract
The authors discuss an original approach for segmenting still images. In this approach, the image is initially decomposed in several levels of different resolution. The decomposition that has been chosen is a Gaussian pyramid. At each level of the pyramid, the image is modeled by a compound Gauss-Markov random field and the segmentation is obtained by using a maximum a posteriori criterion. The segmentation is carried out first at the top level of the pyramid. Once a level (l) has been segmented, this segmentation is projected onto the following level below it (l-1). The process is iterated until the segmentation at the bottom level (0) is performed.>
Ferran Marqués, Jordi Cunillera, Antoni Gasull
ICASSP1
1991 Coding-oriented segmentation based on Gibbs-Markov random fields and human visual system knowledge
abstract
A new segmentation algorithm for still black and white images is introduced. This algorithm forms the basis of a region-oriented sequence coding technique, currently under development. The algorithm models the human mechanism of selecting regions both by their interior characteristics and their boundaries. This is carried out in two different stages: with a preprocessing that takes into account only gray level information, and with a stochastic model for segmented images that uses both region interior and boundary information. In the stochastic model, the gray level information within the regions is modeled by stationary Gaussian processes, and the boundary information by a Gibbs-Markov random field (GMRF). The segmentation is carried out by finding the most likely realization of the joint process (maximum a posteriori criterion), given the preprocessed image. For decreasing the computational load while avoiding local maxima in the probability function, suboptimal versions of the algorithm are proposed.>
Ferran Marqués, Antoni Gasull, Todd R. Reed, Murat Kunt
ICASSP1