VLDB 2026 Research / reviewers in the wild / expert
Damien Kelly
dblp:76/9877
· DBLP profile ↗
9ranked-venue papers
1as first author
2since 2021 · last 2021
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
4 papers |
Image and video processing · 60% Computational photography and imaging · 16% Audio and music processing · 14% | |
| Artificial intelligence
1 paper |
Deep learning architectures and training · 100% |
Topics — the 15 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Image and video processing
image enhancement |
0.9 | 2 | 2021 | Better Compression With Deep Pre-Editing · IEEE Trans. Image Process. 2021 Handheld multi-frame super-resolution · ACM Trans. Graph. 2019 |
Image and video processing › image restoration
compression artifact removal |
0.5 | 1 | 2021 | Better Compression With Deep Pre-Editing · IEEE Trans. Image Process. 2021 |
Image and video processing › image restoration › image deblurring
defocus deblurring |
0.5 | 1 | 2021 | Learning to Reduce Defocus Blur by Realistically Modeling Dual-Pixel Data · ICCV 2021 |
Computational photography and imaging
dual-pixel imaging |
0.5 | 1 | 2021 | Learning to Reduce Defocus Blur by Realistically Modeling Dual-Pixel Data · ICCV 2021 |
Image and video coding
image compression |
0.5 | 1 | 2021 | Better Compression With Deep Pre-Editing · IEEE Trans. Image Process. 2021 |
Image and video processing › image restoration
image deblurring |
0.5 | 1 | 2021 | Learning to Reduce Defocus Blur by Realistically Modeling Dual-Pixel Data · ICCV 2021 |
Image and video processing
image restoration |
0.5 | 1 | 2021 | Better Compression With Deep Pre-Editing · IEEE Trans. Image Process. 2021 |
Computational photography and imaging › image acquisition
burst photography |
0.4 | 1 | 2019 | Handheld multi-frame super-resolution · ACM Trans. Graph. 2019 |
Image and video processing › super-resolution
multi-frame super-resolution |
0.4 | 1 | 2019 | Handheld multi-frame super-resolution · ACM Trans. Graph. 2019 |
Audio and music processing
audio coding |
0.2 | 1 | 2014 | Perceived Audio Quality for Streaming Stereo Music · ACM Multimedia 2014 |
Audio and music processing
audio quality assessment |
0.2 | 1 | 2014 | Perceived Audio Quality for Streaming Stereo Music · ACM Multimedia 2014 |
Audio and music processing › audio coding
lossy audio compression |
0.2 | 1 | 2014 | Perceived Audio Quality for Streaming Stereo Music · ACM Multimedia 2014 |
Audio and music processing › audio quality assessment
perceived audio quality |
0.2 | 1 | 2014 | Perceived Audio Quality for Streaming Stereo Music · ACM Multimedia 2014 |
Machine learning › Deep learning architectures and training › recurrent neural network
recurrent convolutional network |
0.1 | 1 | 2021 | Learning to Reduce Defocus Blur by Realistically Modeling Dual-Pixel Data · ICCV 2021 |
Multimedia systems and quality of experience › multimedia streaming
music streaming |
0.1 | 1 | 2014 | Perceived Audio Quality for Streaming Stereo Music · ACM Multimedia 2014 |
Methods — techniques the papers use, named apart from their topics
synthetic data generation · 1.0recurrent convolutional network · 1.0optimization · 0.5no-reference image quality assessment · 0.5convolutional neural network · 0.5multi-frame alignment · 0.4CFA raw merging · 0.4MUSHRA · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Learning to Reduce Defocus Blur by Realistically Modeling Dual-Pixel DataabstractRecent work has shown impressive results on data-driven defocus deblurring using the two-image views available on modern dual-pixel (DP) sensors. One significant challenge in this line of research is access to DP data. Despite many cameras having DP sensors, only a limited number provide access to the low-level DP sensor images. In addition, capturing training data for defocus deblurring involves a time-consuming and tedious setup requiring the camera’s aperture to be adjusted. Some cameras with DP sensors (e.g., smartphones) do not have adjustable apertures, further limiting the ability to produce the necessary training data. We address the data capture bottleneck by proposing a procedure to generate realistic DP data synthetically. Our synthesis approach mimics the optical image formation found on DP sensors and can be applied to virtual scenes rendered with standard computer software. Leveraging these realistic synthetic DP images, we introduce a recurrent convolutional network (RCN) architecture that improves deblurring results and is suitable for use with single-frame and multi-frame data (e.g., video) captured by DP sensors. Finally, we show that our synthetic DP data is useful for training DNN models targeting video deblurring applications where access to DP data remains challenging. Abdullah Abuolaim, Mauricio Delbracio, Damien Kelly, Michael S. Brown, Peyman Milanfar |
ICCV | 3 |
| 2021 | Better Compression With Deep Pre-EditingabstractCould we compress images via standard codecs while avoiding visible artifacts? The answer is obvious - this is doable as long as the bit budget is generous enough. What if the allocated bit-rate for compression is insufficient? Then unfortunately, artifacts are a fact of life. Many attempts were made over the years to fight this phenomenon, with various degrees of success. In this work we aim to break the unholy connection between bit-rate and image quality, and propose a way to circumvent compression artifacts by pre-editing the incoming image and modifying its content to fit the given bits. We design this editing operation as a learned convolutional neural network, and formulate an optimization problem for its training. Our loss takes into account a proximity between the original image and the edited one, a bit-budget penalty over the proposed image, and a no-reference image quality measure for forcing the outcome to be visually pleasing. The proposed approach is demonstrated on the popular JPEG compression, showing savings in bits and/or improvements in visual quality, obtained with intricate editing effects. Hossein Talebi Esfandarani, Damien Kelly, Xiyang Luo, Ignacio Garcia-Dorado, Feng Yang 0008, Peyman Milanfar, Michael Elad |
IEEE Trans. Image Process. | 2 |
| 2020 | Motion-based frame interpolation for film and television effectsabstractFrame interpolation is the process of synthesising a new frame in‐between existing frames in an image sequence. It has emerged as a key algorithmic module in motion picture effects. In the context of this special issue, this study provides a review of the technology used to create in‐between frames and presents a Bayesian framework that generalises frame interpolation algorithms using the concept of motion interpolation. Unlike existing literature in this area, the authors also compare performance using the top industrial toolkits used in the post production industry. They find that all successful techniques employ motion‐based interpolation, and the commercial version of the Bayesian approach performs best. Another goal of this study is to compare the performance gains with recent convolutional neural network (CNN) algorithms against the traditional explicit model‐based approaches. They find that CNNs do not clearly outperform the explicit motion‐based techniques, and require significant compute resources, but provide complementary improvements in certain types of sequences. Anil C. Kokaram, Davinder Singh, Damien Kelly, Bill Collis, Kim Libreri |
IET Comput. Vis. | 4 |
| 2019 | Handheld multi-frame super-resolutionabstractCompared to DSLR cameras, smartphone cameras have smaller sensors, which limits their spatial resolution; smaller apertures, which limits their light gathering ability; and smaller pixels, which reduces their signal-to-noise ratio. The use of color filter arrays (CFAs) requires demosaicing, which further degrades resolution. In this paper, we supplant the use of traditional demosaicing in single-frame and burst photography pipelines with a multiframe super-resolution algorithm that creates a complete RGB image directly from a burst of CFA raw images. We harness natural hand tremor, typical in handheld photography, to acquire a burst of raw frames with small offsets. These frames are then aligned and merged to form a single image with red, green, and blue values at every pixel site. This approach, which includes no explicit demosaicing step, serves to both increase image resolution and boost signal to noise ratio. Our algorithm is robust to challenging scene conditions: local motion, occlusion, or scene changes. It runs at 100 milliseconds per 12-megapixel RAW input burst frame on mass-produced mobile phones. Specifically, the algorithm is the basis of the Super-Res Zoom feature, as well as the default merge method in Night Sight mode (whether zooming or not) on Google's flagship phone. Bartlomiej Wronski, Ignacio Garcia-Dorado, Manfred Ernst, Damien Kelly, Michael Krainin, Chia-Kai Liang, Marc Levoy, Peyman Milanfar |
ACM Trans. Graph. | 4 |
| 2016 | Bitrate classification of twice-encoded audio using objective quality featuresabstractWhen a user uploads audio files to a music streaming service, these files are subsequently re-encoded to lower bitrates to target different devices, e.g. low bitrate for mobile. To save time and bandwidth uploading files, some users encode their original files using a lossy codec. The metadata for these files cannot always be trusted as users might have encoded their files more than once. Determining the lowest bitrate of the files allows the streaming service to skip the process of encoding the files to bitrates higher than that of the uploaded files, saving on processing and storage space. This paper presents a model that uses quality predictions from ViSQOLAudio, a full reference objective audio quality metric, as features in combination with a multi-class support vector machine classifier. An experiment on twice-encoded files found that low bitrate codecs could be classified using audio quality features. The experiment also provides insights into the implications of multiple transcodes from a quality perspective. Colm Sloan, Naomi Harte, Damien Kelly, Anil C. Kokaram, Andrew Hines |
QoMEX | 3 |
| 2014 | Temporal synchronization of multiple audio signalsabstractGiven the proliferation of consumer media recording devices, events often give rise to a large number of recordings. These recordings are taken from different spatial positions and do not have reliable timestamp information. In this paper, we present two robust graph-based approaches for synchronizing multiple audio signals. The graphs are constructed atop the over-determined system resulting from pairwise signal comparison using cross-correlation of audio features. The first approach uses a Minimum Spanning Tree (MST) technique, while the second uses Belief Propagation (BP) to solve the system. Both approaches can provide excellent solutions and robustness to pairwise outliers, however the MST approach is much less complex than BP. In addition, an experimental comparison of audio features-based synchronization shows that spectral flatness outperforms the zero-crossing rate and signal energy. Julius Kammerl, Neil Birkbeck, Sasi Inguva, Damien Kelly, Andrew J. Crawford, Hugh Denman, Anil C. Kokaram, Caroline Pantofaru |
ICASSP | 4 |
| 2014 | Perceived Audio Quality for Streaming Stereo MusicabstractUsers of audio-visual streaming services expect an ever increasing quality of experience. Channel bandwidth remains a bottleneck commonly addressed with lossy compression schemes for both the video and audio streams. Anecdotal evidence suggests a strongly perceived link between bit rate and quality. This paper presents three audio quality listening experiments using the ITU MUSHRA methodology to assess a number of audio codecs typically used by streaming services. They were assessed for a range of bit rates using three presentation modes: consumer and studio quality headphones and loudspeakers. Our results indicate that with consumer quality headphones, listeners were not differentiating between codecs with bit rates greater than 48 kb/s (p>=0.228). For studio quality headphones and loudspeakers aac-lc at 128 kb/s and higher was differentiated over other codecs (p<=0.001). The results provide insights into quality of experience that will guide future development of objective audio quality metrics. Andrew Hines, Eoin Gillen, Damien Kelly, Jan Skoglund, Anil C. Kokaram, Naomi Harte |
ACM Multimedia | 3 |
| 2012 | Measuring noise correlation for improved video denoisingabstractThe vast majority of previous work in noise reduction for visual media has assumed uncorrelated, white, noise sources. In practice this is almost always violated by real media. Film grain noise is never white, and this paper highlights that the same applies to almost all consumer video content. We therefore present an algorithm for measuring the spatial and temporal spectral density of noise in archived video content, be it consumer digital camera or film orginated. As an example of how this information can be used for video denoising, the spectral density is then used for spatio-temporal noise reduction in the Fourier frequency domain. Results show improved performance for noise reduction in an easily pipelined system. Anil C. Kokaram, Damien Kelly, Hugh Denman, Andrew Crawford |
ICIP | 2 |
| 2011 | Voxel-based Viterbi Active Speaker Tracking (V-VAST) with best view selection for video lecture post-productionabstractAn automated system is presented for reducing a multi-view lecture recording into a single view video containing a best view summary of active speakers. The system uses skin color detection and voxel-based analysis in locating likely speaker locations. Using time-delay estimates from multiple micro phones, speech activity is analyzed for each speaker position. The Viterbi algorithm is then used to estimate a track of the active speaker which maximizes the observed speech activity. This novel approach is termed Voxel-based Viterbi Active Speaker Tracking (V-VAST) and is shown to track speakers with an accuracy of 0.23m. Using the tracking information, the system then extracts from the available camera views the most frontal face view of the active speaker to display. Damien Kelly, Anil C. Kokaram, Francis M. Boland |
ICASSP | 1 |