Daniel Berjón

dblp:84/31 · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
8since 2021 · last 2026
0000-0003-0584-7166ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 2 · 2 first-authorHuman-computer interaction and ubiquitous computing · 2 · 2 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 Viewpoint-invariant soccer pitch registration using geometric and learned features
abstract
Automatic registration of broadcast soccer images to a standardized field model enables advanced analytics, augmented reality overlays, and precise player tracking. We propose a fully automatic, viewpoint-independent homography estimation pipeline fusing three complementary geometric cues: white field markings (lines and elliptical arcs), grass-band delimitations, and a binary playing-field mask. Detected primitives are first richly labeled — classifying lines as longitudinal or transversal, characterizing grass-tone transitions, and encoding four-quadrant intersection patterns — to reduce correspondence ambiguity. We then generate and prune candidate subsets of primitives, establish plausible matches to model elements via intersection-pattern rules and projective cross-ratio invariants, and systematically evaluate homography hypotheses using bidirectional mask-projection accuracies and mean reprojection error. An experimental evaluation on the LaSoDa benchmark demonstrates that the proposed method achieves highly accurate registrations with ground-truth primitives and robust performance in the fully automatic end-to-end pipeline. Furthermore, comparative experiments with recent state-of-the-art approaches confirm improved precision and robustness across diverse broadcast scenarios.
Carlos Cuevas, Daniel Berjón, Narciso García
J. Vis. Commun. Image Represent.2
2026 Quality assessment of 3D reconstructed meshes: Bridging objective metrics, subjective perception, and behavioral cues
abstract
Assessing the quality of 3D reconstructed models remains a key challenge in multimedia applications, especially in the context of cultural heritage, where visual fidelity and perceptual realism are equally crucial. This study investigates how reconstruction parameters, as well as existing objective quality metrics, align with human perception. In addition, we analyze how perceived quality and user interaction are related. A dataset of 3D models was generated by varying the number of input images, mesh complexity, and texture resolution. Results from a subjective study show that texture resolution significantly affects perceived quality, whereas variations in number of images and mesh complexity have a limited impact. Furthermore, interaction behavior was found to vary with perceived quality, with participants spending more time and exploring larger viewing angles for models receiving higher scores. These findings highlight the need for perceptually grounded, interaction-aware evaluation methodologies and provide guidelines for future perceptual optimization of 3D reconstruction pipelines.
Anna Ferrarotti, Isabel Rodríguez, Javier Usón, Sara Baldoni, David Barbero García, Daniel Berjón, Francisco Morán, Narciso García, Federica Battisti, Jesús Gutiérrez 0001, Marco Carli, Julián Cabrera
Signal Process. Image Commun.6
2025 Analysis of Objective 3D Mesh Quality Metrics for Cultural Heritage
abstract
Extended reality technologies are increasingly used in cultural heritage for preserving and accessing sites and artworks, where 3D model acquisition and rendering are key. Despite progress in reconstruction methodologies, a standardized approach to quality assessment is still missing. This study aims to evaluate objective quality metrics —both image-based and model-based, Full Reference and No Reference— applied to 3D models generated using the Structure from Motion algorithm. By varying parameters such as the number of images, number of triangles, and texture resolution, we examine the impact of these factors on metric outcomes, aiming to assess their reliability in cultural heritage applications.
Anna Ferrarotti, Isabel Rodríguez, Javier Usón, Sara Baldoni, Jesús Gutiérrez 0001, Daniel Berjón, Francisco Morán, Federica Battisti, Narciso García, Marco Carli, Julián Cabrera
QoMEX6
2024 Real-Time Free Viewpoint Video for Immersive Videoconferencing
abstract
In this work, we propose a demo of an immersive videoconference system using Free Viewpoint Video (FVV) technology. It makes use of the FVV Live system, which covers the entire FVV pipeline (capture, view rendering, and visualization) while working in real-time. The FVV Live system consists of nine cameras that capture an environment and a view renderer that uses the information from the cameras to generate a synthetic view at an arbitrary point.It is designed as a hybrid demo. While the capture and rendering processes take place at our premises, FVV Live can be visualized through devices connected to the Internet.The system allows immersive navigation of a virtual scene with 6 degrees of freedom, and interaction with live-captured avatars integrated in such scene. For this purpose, it uses WebRTC connections to update the position of the virtual camera and to receive the FVV Live view encoded as a video.Additionally, the user will be recorded by a simple camera and microphone setup, and the generated streams will be transmitted to our premises through the same WebRTC server. This way, people being recorded by FVV Live will be able to see and hear the user, enabling bidirectional communication.
Javier Usón, Victoria Muñoz, Carlos Cortés 0001, Daniel Berjón, Francisco Morán, César Díaz, Jesús Gutiérrez 0001, Fernando Jaureguizar, Narciso García, Julián Cabrera
QoMEX4
2024 Automatic highlight detection in videos of martial arts tricking
abstract
Abstract We propose a novel strategy for the automatic detection of highlight events in user-generated tricking videos, to the best of our knowledge, the first one specifically tailored for this complex sport. Most current methods for related sports leverage high-level semantics such as predefined camera angles or common editing practices, or rely on depth cameras to achieve automatic detection. However, our approach only relies on the contents (themselves) in the frames of a given video, and consists in a four stage pipeline. The first stage identifies foreground key points of interest along with an estimation of their motion in the video frames. In the second stage, these points are grouped into regions of interest based on their proximity and motion. Their behavior over time is evaluated in the third stage to generate an attention map indicating the regions participating in the most relevant events. The fourth and final stage provides the extracted video sequences where highlights have been identified. Experimental results attest to the effectiveness of our approach, which shows high recall and precision values at frame level, with detections that fit well the ground truth events.
Marcos Rodrigo, Carlos Cuevas, Daniel Berjón, Narciso García
Multim. Tools Appl.3
2023 Soccer line mark segmentation and classification with stochastic watershed transform
abstract
Augmented reality applications are beginning to change the way sports are broadcast, providing richer experiences and valuable insights to fans. The first step of augmented reality systems is camera calibration, possibly based on detecting the line markings of the playing field. Most existing proposals for line detection rely on edge detection and Hough transform, but radial distortion and extraneous edges cause inaccurate or spurious detections of line markings. We propose a novel strategy to automatically and accurately segment and classify line markings. First, line points are segmented thanks to a stochastic watershed transform that is robust to radial distortions, since it makes no assumptions about line straightness, and is unaffected by the presence of players or the ball. The line points are then linked to primitive structures (straight lines and ellipses) thanks to a very efficient procedure that makes no assumptions about the number of primitives that appear in each image. The strategy has been tested on a new and public database composed by 60 annotated images from matches in five stadiums. The results obtained have proven that the proposed strategy is more robust and accurate than existing approaches, achieving successful line mark detection even under challenging conditions.
Daniel Berjón, Carlos Cuevas, Narciso García
Signal Process. Image Commun.1
2022 Grass band detection in soccer images for improved image registration
abstract
The registration of images of soccer matches is a key stage in many computer vision applications. Until now, this task has been typically carried out from key points obtained from the white line marks drawn on the field of play, but in many cases this does not yield enough keypoints for a robust registration. This article proposes a strategy to detect the borders between the grass bands of the field of play and therefore makes it possible to locate many more key points that will allow to carry out a subsequent registration of the images. First, a preprocessing is applied to obtain a grayscale image in which the grass bands are easily distinguishable, and also to obtain a binary mask of the entire field of play that determines the area of interest. Then, a local analysis is carried out to detect most of the borders between grass bands. Finally, a global analysis based on the intersections between lines is applied to group the detected borders and rule out false detections. The strategy has been evaluated on two databases composed of hundreds of annotated images from matches in several stadiums with different characteristics and light conditions. The results obtained have shown that most of the lines delimiting the grass bands are found successfully, while the number of false detections is very small.
Carlos Cuevas, Daniel Berjón, Narciso García
Signal Process. Image Commun.2
2022 FVV Live: A Real-Time Free-Viewpoint Video System With Consumer Electronics Hardware
abstract
FVV Live is a novel end-to-end free-viewpoint video system, designed for real-time operation, using consumer-grade cameras and hardware, which enables low deployment costs and easy installation for immersive event-broadcasting or videoconferencing. All the blocks of the system have been designed to maximize perceptual video quality, overcoming the limitations imposed by hardware and network, which impact directly the accuracy of depth data and thus the quality of virtual view synthesis. Therefore, it does not sacrifice perceptual video quality with respect to high-end counterparts. The results presented in this paper correspond to an implementation with nine stereo-based depth cameras. However, the design of the acquisition block of FVV Live allows scalability for an arbitrary number of cameras. In addition, FVV Live presents low motion-to-photon and end-to-end delays, which enables a responsive free-viewpoint navigation and bilateral immersive communications. Moreover, the visual quality of FVV Live has been assessed through subjective assessment with satisfactory results, and additional comparative tests show that it is preferred over state-of-the-art DIBR alternatives.
Pablo Carballeira, Carlos Carmona, César Díaz, Daniel Berjón, Daniel Corregidor, Julián Cabrera, Francisco Morán, Carmen Doblado, Sergio Arnaldo, María del Mar Martín, Narciso García
IEEE Trans. Multim.4
2018 Real-time nonparametric background subtraction with tracking-based foreground update
Daniel Berjón, Carlos Cuevas, Francisco Morán, Narciso García
Pattern Recognit.1
2017 Detection of Stationary Foreground Objects Using Multiple Nonparametric Background-Foreground Models on a Finite State Machine
abstract
There is a huge proliferation of surveillance systems that require strategies for detecting different kinds of stationary foreground objects (e.g., unattended packages or illegally parked vehicles). As these strategies must be able to detect foreground objects remaining static in crowd scenarios, regardless of how long they have not been moving, several algorithms for detecting different kinds of such foreground objects have been developed over the last decades. This paper presents an efficient and high-quality strategy to detect stationary foreground objects, which is able to detect not only completely static objects but also partially static ones. Three parallel nonparametric detectors with different absorption rates are used to detect currently moving foreground objects, short-term stationary foreground objects, and long-term stationary foreground objects. The results of the detectors are fed into a novel finite state machine that classifies the pixels among background, moving foreground objects, stationary foreground objects, occluded stationary foreground objects, and uncovered background. Results show that the proposed detection strategy is not only able to achieve high quality in several challenging situations but it also improves upon previous strategies.
Carlos Cuevas, Daniel Berjón, Narciso García
IEEE Trans. Image Process.3
2016 Optimal Piecewise Linear Function Approximation for GPU-Based Applications
abstract
Many computer vision and human-computer interaction applications developed in recent years need evaluating complex and continuous mathematical functions as an essential step toward proper operation. However, rigorous evaluation of these kind of functions often implies a very high computational cost, unacceptable in real-time applications. To alleviate this problem, functions are commonly approximated by simpler piecewise-polynomial representations. Following this idea, we propose a novel, efficient, and practical technique to evaluate complex and continuous functions using a nearly optimal design of two types of piecewise linear approximations in the case of a large budget of evaluation subintervals. To this end, we develop a thorough error analysis that yields asymptotically tight bounds to accurately quantify the approximation performance of both representations. It provides an improvement upon previous error estimates and allows the user to control the tradeoff between the approximation error and the number of evaluation subintervals. To guarantee real-time operation, the method is suitable for, but not limited to, an efficient implementation in modern graphics processing units, where it outperforms previous alternative approaches by exploiting the fixed-function interpolation routines present in their texture units. The proposed technique is a perfect match for any application requiring the evaluation of continuous functions; we have measured in detail its quality and efficiency on several functions, and, in particular, the Gaussian function because it is extensively used in many areas of computer vision and cybernetics, and it is expensive to evaluate.
Daniel Berjón, Guillermo Gallego 0002, Carlos Cuevas, Francisco Morán, Narciso García
IEEE Trans. Cybern.1
2015 Seamless, Static Multi-Texturing of 3D Meshes
abstract
Abstract In the context of 3D reconstruction, we present a static multi‐texturing system yielding a seamless texture atlas calculated by combining the colour information from several photos from the same subject covering most of its surface. These pictures can be provided by shooting just one camera several times when reconstructing a static object, or a set of synchronized cameras, when dealing with a human or any other moving object. We suppress the colour seams due to image misalignments and irregular lighting conditions that multi‐texturing approaches typically suffer from, while minimizing the blurring effect introduced by colour blending techniques. Our system is robust enough to compensate for the almost inevitable inaccuracies of 3D meshes obtained with visual hull–based techniques: errors in silhouette segmentation, inherently bad handling of concavities, etc.
Rafael Pagés, Daniel Berjón, Francisco Morán, Narciso García
Comput. Graph. Forum2
2013 Objective and subjective evaluation of static 3D mesh compression
Daniel Berjón, Francisco Morán, Shankar Manjunatha
Signal Process. Image Commun.1
2013 Automatic system for virtual human reconstruction with 3D mesh multi-texturing and facial enhancement
Rafael Pagés, Daniel Berjón, Francisco Morán
Signal Process. Image Commun.2
2011 Evaluation of backward mapping DIBR for FVV applications
abstract
In this paper, we explore the challenges posed by wide baseline camera configurations for depth-image-based rendering, which should provide greater freedom for choosing the virtual viewpoint in a Free Viewpoint Video context, compared with the usual camera configurations intended for use in 3DTV settings. We implement a backward mapping approach with a custom filtering scheme based on median filters. Whilst the results back our initial assumption that this camera configuration provides good mobility, we show that the usual encoding for depth information referred to a global reference system is wrong and reference systems local to each camera should be used instead.
Daniel Berjón, Alexander Sorkine-Hornung, Francisco Morán, Aljoscha Smolic
ICME1
2006 Schedulability analysis of AR-TP, a Ravenscar compliant communication protocol for high-integrity distributed systems
abstract
A new token-passing algorithm called AR-TP for avoiding the non-determinism of some networking technologies is presented. This protocol allows the schedulability analysis of the network, enabling the use of standard Ethernet hardware for hard real-time behavior while adding congestion management. It is specially designed for high-integrity distributed hard real-time systems, being fully compliant with the Ravenscar profile.
Santiago Urueña, Juan Zamorano, Daniel Berjón, José Antonio Pulido, Juan Antonio de la Puente
IPDPS3