EDBT 2026 Demo / reviewers in the wild / expert
François Pitié
dblp:62/1773
· DBLP profile ↗
27ranked-venue papers
6as first author
11since 2021 · last 2025
0000-0003-4599-0549ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3Human-computer interaction and ubiquitous computing · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LiteVPNet: A Lightweight Network for Video Encoding Control in Quality-Critical Applications
Vibhoothi Vibhoothi, François Pitié, Anil C. Kokaram |
PCS | 2 |
| 2024 | Lightweight Video Denoising Using a Classic Bayesian BackboneabstractIn recent years, state-of-the-art image and video denoising networks have become increasingly large, requiring millions of trainable parameters to achieve best-in-class performance. Improved denoising quality has come at the cost of denoising speed, where modern transformer networks are far slower to run than smaller denoising networks such as FastDVDnet and classic Bayesian denoisers such as the Wiener filter.In this paper, we implement a hybrid Wiener filter which leverages small ancillary networks to increase the original denoiser performance, while retaining fast denoising speeds. These networks are used to refine the Wiener coring estimate, optimise windowing functions and estimate the unknown noise profile. Using these methods, we outperform several popular denoisers and remain within 0.2 dB, on average, of the popular VRT transformer. Our method was found to be over x10 faster than the transformer method, with a far lower parameter cost. Clement Bled, François Pitié |
ICME | 2 |
| 2023 | Pushing the Limits of the Wiener Filter in Image DenoisingabstractAs modern image denoiser networks have grown in size, their reported performance in popular real noise benchmarks such as DND and SIDD have now long outperformed classic non-deep learning denoisers such as Wiener and Wavelet-based methods. In this paper, we propose to revisit the Wiener filter and re-assess its potential performance. We show that carefully considering the implementation of the Wiener filter can yield similar performance to popular networks such as DnCNN. Clement Bled, François Pitié |
ICIP | 2 |
| 2023 | Subjective Assessment of the Impact of a Content Adaptive Optimiser for Compressing 4K HDR Content With AV1abstractSince 2015 video dimensionality has expanded to higher spatial and temporal resolutions and a wider colour gamut. This High Dynamic Range (HDR) content has gained traction in the consumer space as it delivers an enhanced quality of experience. At the same time, the complexity of codecs is growing. This has driven the development of tools for content-adaptive optimisation that achieve optimal rate-distortion performance for HDR video at 4K resolution. While improvements of just a few percentage points in BD-Rate (1-5%) are significant for the streaming media industry, the impact on subjective quality has been less studied especially for HDR/AV1. In this paper, we conduct a subjective quality assessment (42 subjects) of 4K HDR content with a per-clip optimisation strategy. We correlate these subjective scores with existing popular objective metrics used in standard development and show that some perceptual metrics correlate surprisingly well even though they are not tuned for HDR. We find that the DSQCS protocol is too insensitive to categorically compare the methods but the data allows us to make recommendations about the use of experts vs non-experts in HDR studies, and explain the subjective impact of film grain in HDR content under compression. Vibhoothi, Angeliki V. Katsenou, François Pitié, Katarina Domijan, Anil C. Kokaram |
ICIP | 3 |
| 2023 | Comparison of HDR quality metrics in Per-Clip Lagrangian multiplier optimisation with AV1abstractThe complexity of modern codecs along with the increased need of delivering high-quality videos at low bitrates has reinforced the idea of a per-clip tailoring of parameters for optimised rate-distortion performance. While the objective quality metrics used for Standard Dynamic Range (SDR) videos have been well studied, the transitioning of consumer displays to support High Dynamic Range (HDR) videos, poses a new challenge to rate-distortion optimisation. In this paper, we review the popular HDR metrics DeltaE100 (DE100), PSNRL100, wPSNR, and HDR-VQM. We measure the impact of employing these metrics in per-clip direct search optimisation of the rate-distortion Lagrange multiplier in AV1. We report, on 35 HDR videos, average Bjontegaard Delta Rate (BD-Rate) gains of 4.675%, 2.226%, and 7.253% in terms of DE100, PSNRL100, and HDR-VQM. We also show that the inclusion of chroma in the quality metrics has a significant impact on optimisation, which can only be partially addressed by the use of chroma offsets. Vibhoothi, François Pitié, Angeliki V. Katsenou, Yeping Su, Balu Adsumilli, Anil C. Kokaram |
ICME | 2 |
| 2023 | Recommendations for Verifying HDR Subjective Testing WorkflowsabstractOver the past few years, there has been an increase in the demand and availability of High Dynamic Range (HDR) displays and content. To ensure the production of high-quality materials, human evaluation is required. However, ascertaining whether the full playback pipeline is indeed HDR-compliant can be challenging. In this paper, we present a set of recommendations for conformance testing to validate various aspects of the testing workflow, including playback, displays, brightness, colours, and viewing environment. We assessed the effectiveness of HDR conversion techniques used in current standards development (3GPP) for making source materials. Additionally, we evaluate HDR display technologies, including OLED and LCD, using both consumer television and a reference monitor. Vibhoothi, Angeliki V. Katsenou, John Squires, François Pitié, Anil C. Kokaram |
QoMEX | 4 |
| 2022 | Frame-Type Sensitive RDO Control for Content-Adaptive EncodingabstractVideo transcoding is an increasingly important application in the streaming media industry. It has become important to investigate the optimisation of transcoder parameters for a single clip simply because of the immense number of playbacks for popular clips. In this paper, we explore the use of a canned optimiser to estimate the optimal Rate-Distortion (RD) tradeoff achievable for a particular clip. We show that by adjusting the Lagrange multiplier in RD optimisation on keyframes alone we can achieve more than 10× the previous BD-Rate gains possible without affecting quality for any operating point. Vibhoothi, François Pitié, Anil C. Kokaram |
ICIP | 2 |
| 2021 | Investigating Automated Mechanisms for Multi-Modal Prediction of User Online-Video Commenting BehaviourabstractOnline-video commenting is attracting increasing attention among young people, particularly in the form of "danmu" comments in Asia. These provide a channel for engagement enabling users to share time-synchronous comments on videos with other viewers. Danmu form community discussions of video content and frequently provoke extensive further contributions. The motivation to add danmu comments at specific points in videos are not obvious. In this paper, we explore the potential for predicting user online-video comment distributions using multi-modal signals from the video content stream. To address this task we integrate multiple sources of information, including video frame content, the audio signal, as well as video subtitles, in an end-to-end neural framework. Specifically, text, visual and audio input are encoded respectively and then a transformer framework is used to learn and combine attention aware representation of three modalities. We evaluate the system using retrieval-based evaluation metrics, including mean average precision (mAP) and normalized discounted cumulative gain (NDCG). We conduct experiments on an expanded publicly available danmu commenting dataset. Our model significantly outperformed an LSTM multi-modal baseline method. Hao Wu 0110, François Pitié, Gareth J. F. Jones |
CBMI | 2 |
| 2021 | CNN-Based Video Codec Classifier For Multimedia ForensicsabstractIn video forensics, identification of codec type is complicated by a lack of standards compliant compressed bitstreams. Previous work is unable to identify codec types without actually decoding the file successfully. This paper presents a CNN classifier derived from the AlexNet architecture that can detect codec types without decoding the bitstream. It is based on classification of the raw bitstream data itself without decoding. The algorithm is tested on real data in a video forensics setting as well as user generated content supplied in the YouTube test set. Our results show better than 96.73% accuracy with over 43 combinations of codec/containers and, at least, 88.59% accuracy at 20% data corruption across both test sets. Rodrigo Pessoa, Anil C. Kokaram, François Pitié, Mark Sugrue |
ICIP | 3 |
| 2021 | Knowing Where and What to Write in Automated Live Video Comments: A Unified Multi-Task ApproachabstractLive video comments, or “danmu”, are an emerging social feature on Asian online video platforms. These time-synchronous comments are overlaid on the video playback and uniquely enrich the viewing experience, engaging hundreds of millions of users in rich community discussions. The presence of danmu comments has become a determining factor for video popularity. Recent work has proposed a model to automatically generate comments, but very little work has so far considered the problem of where to insert the comments in the video timeline. In this work, we propose to address both the what and where of automatic danmu generation, by jointly predicting the danmu comment content to be generated, as well as its optimal insertion point in the video timeline. Our model exploits the video visual content, subtitles, audio signals, and any existing surrounding comments, in one unified architecture and can handle scenarios where the videos are already heavily commented or when the video has no comments yet. Experiments show that our proposed unified framework is in general observed to outperform state-of-the-art comment generation methods. Hao Wu 0110, Gareth J. F. Jones, François Pitié |
ICMI | 3 |
| 2021 | Near Optimal Per-Clip Lagrangian Multiplier Prediction in HEVCabstractThe majority of internet traffic is video content. This drives the demand for video compression to deliver high quality video at low target bitrates. Optimising the parameters of a video codec for a specific video clip (per-clip optimisation) has been shown to yield significant bitrate savings. In previous work we have shown that per-clip optimisation of the Lagrangian multiplier leads to up to 24% BD-Rate improvement. A key component of these algorithms is modeling the R-D characteristic across the appropriate bitrate range. This is computationally heavy as it usually involves repeated video encodes of the high resolution material at different parameter settings. This work focuses on reducing this computational load by deploying a NN operating on lower bandwidth features. Our system achieves BD-Rate improvement in approximately 90% of a large corpus with comparable results to previous work in direct optimisation. Daniel Ringis, François Pitié, Anil C. Kokaram |
PCS | 2 |
| 2020 | Interactive Training And Architecture For Deep Object SelectionabstractInteractive object cutout tools are the cornerstone of the image editing workflow. Algorithms that can reduce the number of interactions are clearly valuable. Recent deep-learning based interactive segmentation algorithms are capable of rough binary selections with a handful of clicks, yet, they tend to plateau once this rough selection has been reached. In this work, we interpret this plateau as an inability of the algorithm to precisely leverage each user interaction.We introduce a novel interactive architecture and a training scheme that are both tailored to better exploit the user input at higher numbers of clicks. Comprehensive experiments support our approach, and our network achieves state of the art performance. Marco Forte, Brian L. Price, Scott Cohen, Ning Xu 0007, François Pitié |
ICME | 5 |
| 2020 | An Advert Creation System for 3D Product Placements
Ivan Bacher, Hossein Javidnia, Soumyabrata Dev, Rahul Agrahari, Murhaf Hossari, Matthew Nicholson, Clare Conran, Peng Song 0027, David Corrigan, François Pitié |
ECML/PKDD (4) | 11 |
| 2020 | Advances in colour transferabstractColour grading is an essential step in movie post‐production, which is done in the industry by experienced artists on expensive edit hardware and software suites. This paper presents a review of the advances made to automate this process. The review looks in particular at how the state‐of‐the‐art in optimal transport and deep learning has advanced some of the fundamental problems of colour transfer, and how far are we still from being able to automatically grade images. François Pitié |
IET Comput. Vis. | 1 |
| 2019 | The ALOS Dataset for Advert Localization in Outdoor ScenesabstractThe rapid increase in the number of online videos provides the marketing and advertising agents ample opportunities to reach out to their audience. One of the most widely used strategies is product placement, or embedded marketing, wherein new advertisements are integrated seamlessly into existing advertisements in videos. Such strategies involve accurately localizing the position of the advert in the image frame, either manually in the video editing phase, or by using machine learning frameworks. However, these machine learning techniques and deep neural networks need a massive amount of data for training. In this paper, we propose and release the first large-scale dataset of advertisement billboards, captured in outdoor scenes. We also benchmark several state-of-the-art semantic segmentation algorithms on our proposed dataset. Soumyabrata Dev, Murhaf Hossari, Matthew Nicholson, Killian McCabe, Atul Nautiyal, Clare Conran, Wei Xu 0022, François Pitié |
QoMEX | 9 |
| 2018 | Temporal Consistency for Still Image Based Defocus Blur Estimation MethodsabstractMany Defocus blur estimation methods have been proposed in recent years but, when applied to video sequences in a frame-by-frame manner, they typically exhibit temporal inconsistencies or flickering. This paper presents a temporal coherence scheme that can be coupled to any existing defocus blur estimation for still images, aiming to produce spatiotemporally coherent defocus blur map videos. The proposed method is based on the design of a Kalman Filter which is applied on a patch level. Experimental results show that the proposed method can smooth out undesirable temporal fluctuations whilst still being able to preserve the abrupt local appearance changes due to motion, occlusions or dis-occlusions. Ali Karaali, Cláudio R. Jung, François Pitié |
ICIP | 3 |
| 2018 | An Advert Creation System for Next-Gen Publicity
Atul Nautiyal, Killian McCabe, Murhaf Hossari, Soumyabrata Dev, Matthew Nicholson, Clare Conran, Declan McKibben, Wei Xu 0022, François Pitié |
ECML/PKDD (3) | 10 |
| 2016 | An alternative matting LaplacianabstractCutting out and object and estimate its transparency mask is a key task in many applications. We take on the work on closed-form matting by Levin et al.[1], that is used at the core of many matting techniques, and propose an alternative formulation that offers more flexible controls over the matting priors. We also show that this new approach is efficient at upscaling transparency maps from coarse estimates. François Pitié |
ICIP | 1 |
| 2013 | A Non-parametric Framework for Document Bleed-through RemovalabstractThis paper presents recent work on a new framework for non-blind document bleed-through removal. The framework includes image preprocessing to remove local intensity variations, pixel region classification based on a segmentation of the joint recto-verso intensity histogram and connected component analysis on the subsequent image labelling. Finally restoration of the degraded regions is performed using exemplar-based image in painting. The proposed method is evaluated visually and numerically on a freely available database of 25 scanned manuscript image pairs with ground truth, and is shown to outperform recent non-blind bleed-through removal techniques. Róisín Rowley-Brooke, François Pitié, Anil C. Kokaram |
CVPR | 2 |
| 2012 | A Ground Truth Bleed-Through Document Image Database
Róisín Rowley-Brooke, François Pitié, Anil C. Kokaram |
TPDL | 2 |
| 2011 | Reflection detection in image sequencesabstractReflections in image sequences consist of several layers superimposed over each other. This phenomenon causes many image processing techniques to fail as they assume the presence of only one layer at each examined site e.g. motion estimation and object recognition. This work presents an automated technique for detecting reflections in image sequences by analyzing motion trajectories of feature points. It models reflection as regions containing two different layers moving over each other. We present a strong detector based on combining a set of weak detectors. We use novel priors, generate sparse and dense detection maps and our results show high detection rate with rejection to pathological motion and occlusion. Mohamed A. Elgharib, François Pitié, Anil C. Kokaram |
CVPR | 2 |
| 2010 | Matting with a depth mapabstractDepth maps are becoming a readily available commodity of the stereo pipeline. We propose to make use of this new free information to improve a key step of postproduction that is matting. We extend the work of Levin et al on closed form matting to introduce two new depth-aware techniques. First we explore how depth can be used as an extra channel in the matting process. Then we see how depth can be used as a diffusion guide for matting. Our results show that both techniques can reduce the amount of time needed to pull a matte. François Pitié, Anil C. Kokaram |
ICIP | 1 |
| 2009 | Extraction of non-binary blotch mattesabstractAutomated blotch removal is important in film restoration and typically involves a detection/interpolation step. Current algorithms model the corruption as a binary mixture between the original, clean images and an opaque (dirt) field. This typically causes incomplete blotch removal that manifests as blotch haloes in reconstruction. This paper proposes a new approach by modeling the corruption as a continuous mixture between the two components and generating a solution using a Bayesian framework. We use novel priors, propose a computationally efficient scheme for implementation and our results show more complete blotch reconstruction. Mohamed A. Elgharib, François Pitié, Anil C. Kokaram |
ICIP | 2 |
| 2007 | Automated colour grading using colour distribution transfer
François Pitié, Anil C. Kokaram, Rozenn Dahyot |
Comput. Vis. Image Underst. | 1 |
| 2005 | N-Dimensional Probablility Density Function Transfer and its Application to Colour TransferabstractThis article proposes an original method to estimate a continuous transformation that maps one N-dimensional distribution to another. The method is iterative, non-linear, and is shown to converge. Only 1D marginal distribution is used in the estimation process, hence involving low computation costs. As an illustration this mapping is applied to color transfer between two images of different contents. The paper also serves as a central focal point for collecting together the research activity in this area and relating it to the important problem of automated color grading François Pitié, Anil C. Kokaram, Rozenn Dahyot |
ICCV | 1 |
| 2005 | Off-line multiple object tracking using candidate selection and the Viterbi algorithmabstractThis paper presents a probabilistic framework for off-line multiple object tracking. At each timestep, a small set of deterministic candidates is generated which is guaranteed to contain the correct solution. Tracking an object within video then becomes possible using the Viterbi algorithm. In contrast with particle filter methods where candidates are numerous and random, the proposed algorithm involves a few candidates and results in a deterministic solution. Moreover, we consider here off-line applications where past and future information is exploited. This paper shows that, although basic and very simple, this candidate selection allows the solution of many tracking problems in different real-world applications and offers a good alternative to particle filter methods for off-line applications. François Pitié, Sid-Ahmed Berrani, Anil C. Kokaram, Rozenn Dahyot |
ICIP (3) | 1 |
| 2004 | Gradient based dominant motion estimation with integral projections for real time video stabilisationabstractThis paper presents a new expression of the relationship between integral projections and motion in an image pair. The resulting new multiresolution gradient based approach is used to estimate dominant motion in image sequences degraded by random shake. The paper also describes an implementation using the GPU as a coprocessor for the CPU that allows, for the first time, real time video stabilisation in software on broadcast standard definition television images. Andrew Crawford, Hugh Denman, Francis Kelly, François Pitié, Anil C. Kokaram |
ICIP | 4 |