Kurt Debattista

dblp:29/5944 · DBLP profile ↗
← Back
79ranked-venue papers
8as first author
28since 2021 · last 2026
0000-0003-2982-5199ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 56 · 6 first-author · 12 since 2021Artificial intelligence and machine learning · 14 · 13 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Security and privacy · 2Systems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Rethinking probabilistic learning for counterfactual low-light image enhancement in robust engineering vision systems
abstract
• 1: Formulates low-light enhancement as a counterfactual intervention on a Retinex SCM. • 2: Introduces a physics-guided structural causal model for illumination–reflectance reasoning. • 3: Achieves superior perceptual fidelity and robustness on multiple benchmark datasets. Many Architecture, Engineering, and Construction (AEC) operations must operate safely under visually compromised conditions, dimensional dimension, including night-time navigation/driving, tunnel or cave exploration, and hazardous or confined sites observation. In such conditions, low-light image enhancement is expected to be more than simply image brightening but also to maintain mission-critical structures and be visually interpretable for deployment in safety-critical conditions. Nonetheless, current deep LIE techniques conceptually overlook image processing, physical priors and causal mechanisms during improvement, contributing to constrained ruggedness and generalisability. This paper presents a Learning Probabilistic Low-light Image Enhancement (LPIE) network that explicitly integrates causal factors into the LPIE and allows for probabilistic counterfactual reasoning. LPIE casts illumination enhancement in the light of an intervention and unrolls it into a structured causal model of image formation. A normalizing-flow module performs invertible and physically coherent illumination processes, while an uncertainty-aware Transformer recovers reflectance in spatially varied and high-textured regions. Consequently, a complete probabilistic refinement further optimizes all elements to optimize the distribution of physical variables, resulting in a natural, fine-detailed photo and supporting across-used scenarios. Experiments on public benchmarks and AEC-relevant datasets show that LPIE achieves state-of-the-art performance and clearly surpasses recent LIE methods on the widely used benchmark dataset, paving the way for more interpretable, robust and trustworthy automated perception systems.
Zixiang Wei, Kurt Debattista, Valentina Donzella
Knowl. Based Syst.3
2026 Zero-Shot infrared-guided HDR video deflickering
Jingchao Peng, Thomas Bashford-Rogers, Francesco Banterle, Haitao Zhao 0002, Kurt Debattista
Pattern Recognit.5
2026 CapHDR2IR: Caption-Driven Transfer From Visible Light to Infrared Domain
Jingchao Peng, Thomas Bashford-Rogers, Haitao Zhao 0002, Aru Ranjan Singh, Abhishek Goswami, Kurt Debattista
IEEE Trans. Multim.7
2026 Image Based Whole Sky Cloud Volume Generation
abstract
Accurate illumination is crucial for many imaging and vision applications, and skies are the dominant source of lighting in many scenes. Most existing work for representing sky illumination has focused on clear skies or more recently generative approaches for synthesizing clouds. However, these are very limited in that they assume distant illumination and do not capture the 3D properties of clouds. This paper presents a novel and principled approach to extract 3D whole-sky volumetric representations of clouds which can be used for imaging applications. Our approach extracts clouds from a single fisheye capture of the sky via an iterative optimization process. We achieve this by exploiting the physical properties of light scattering in clouds and use these to drive a domain-specific light transport simulation algorithm to render the images required for optimization. Results for this method provide high accuracy when re-rendering with our reconstructed clouds compared to real captures, and also enable novel uses of environment maps such as inclusion of captured clouds in renderings, cloud shadows, and more accurate aerial perspective and lighting.
Pinar Satilmis, Kurt Debattista, Thomas Bashford-Rogers
IEEE Trans. Vis. Comput. Graph.2
2025 Lightweight RAW Object Detection for Automated Driving
abstract
Modern sensor and image acquisition technologies have enabled Automated Vehicles (AVs) to leverage image data for enhanced on-board visual perception. To achieve this, AVs deploy Computer Vision (CV) deep learning tools for safe and accurate object detection. These systems typically rely on Image Signal Processing (ISP) pipelines to convert RAW sensor data into human-perceivable RGB images, as state-of-the-art deep learning models are optimised for such inputs. However, in safety-critical real-time applications such as AVs, CV systems must not only be reliable and accurate but also fast. In this work, we demonstrate that a lightweight ISP comprising simple processing steps can effectively operate on RAW image data while maintaining strong object detection performance across diverse scenes. Our experiments show that our proposed simple ISP maintains a balanced trade-off between accuracy and latency.
Georgia Souvalioti, Abhishek Goswami, Aru Ranjan Singh, Kurt Debattista, Valentina Donzella
IV4
2025 Driver Expectations for Automated Vehicle Driving Styles in Mixed-Traffic Interactions
abstract
As highly automated vehicles (AVs) are deployed in various countries, mixed-autonomy traffic will become common and persist for the foreseeable future. In such environments, AVs must operate in ways which are both predictable and acceptable to human drivers, particularly in complex intersections where negotiation is crucial. However, how human drivers expect AVs to interact with them, particularly in scenarios where the right-of-way is ambiguous, remains unclear. In this research, we conducted a simulation-based video survey of UK drivers (N = 87), investigating their perception of aggressive and defensive AV driving styles under unclear right-of-way scenarios. The analysis indicates that a defensive driving style is generally preferred by human drivers, while an aggressive style can also be acceptable at lower-speed interaction zones. These findings provide empirical evidence for algorithm engineers seeking to design motion control and negotiation strategies that align with human expectations in mixed traffic.
Roger Woodman, Zhizhuo Su, Kurt Debattista
IV4
2025 Cross-Modality Distillation for Multi-Modal Tracking
abstract
Contemporary multi-modal trackers achieve strong performance by leveraging complex backbones and fusion strategies, but this comes at the cost of computational efficiency, limiting their deployment in resource-constrained settings. On the other hand, compact multi-modal trackers are more efficient but often suffer from reduced performance due to limited feature representation. To mitigate the performance gap between compact and more complex trackers, we introduce a cross-modality distillation framework. This framework includes a complementarity-aware mask autoencoder designed to enhance cross-modal interactions by selectively masking patches within a modality, thereby forcing the model to learn more robust multi-modal representations. Additionally, we present a specific-common feature distillation module that transfers both modality-specific and shared information from a more powerful model's backbone to the compact model. Moreover, we develop a multi-path selection distillation module to guide a simple fusion module in learning more accurate multi-modal information from a sophisticated fusion mechanism using multiple paths. Extensive experiments on six multi-modal tracking benchmarks demonstrate that the proposed tracker, despite being lightweight, outperforms most state-of-the-art methods, highlighting its effectiveness. Notably, our tiny variant achieves a PR score of 67.5% on LasHeR, a PR score of 58.5% on DepthTrack, and a PR score of 73.1% on VisEvent with only 6.5 M parameters, while operating at 126 FPS on an NVIDIA 2080Ti GPU.
Tianlu Zhang, Qiang Zhang 0020, Kurt Debattista, Jungong Han
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Robustness of Panoptic Segmentation for Degraded Automotive Cameras Data
abstract
Precise situational awareness is essential for the safe deployment of artificial intelligence in real-world applications, particularly in assisted and automated driving (AAD) systems. Among perception techniques, panoptic segmentation is a promising technique to identify and categorise objects, impending hazards, and drivable space at a pixel level. While panoptic quality might be affected by automotive camera data quality, a comprehensive understanding and modelling of their relationship remains underexplored. Motivated by such a need, this work proposes a unifying pipeline to evaluate the robustness of panoptic segmentation models for automotive cameras, correlating it with 8 traditional image quality metrics (IQA). The proposed pipeline begins by generating a novel degraded dataset, D-Cityscapes+, featuring 19 realistic automotive degradation types at varying severity levels, including novel models for darkness and snowfall conditions with veiling effect. Evaluations on 14 state-of-the-art segmentation model backbones yielded key insights: 1) large-particle degradations (e.g., lens droplets, heavy snow) severely degrade segmentation performance, increasing uncertainty and edge-concentrated segmentation errors; 2) Transformer-based models outperform CNN models under adverse conditions; however, longer processing time, a higher number of parameters, and computational cost are limiting their real-world deployment; 3) frequency-based IQA metrics, such as CW-SSIM, strongly correlate with segmentation performance, serving as reliable predictive tools. 4) visual enhancements via restoration do not coherently benefit downstream segmentation tasks, underscoring the need for perception-specific restoration techniques. The benchmark and code:https://github.com/Warwick-Jocelyn/BRPS.
Daniel Gummadi, Mehrdad Dianati, Kurt Debattista, Valentina Donzella
IEEE Trans Autom. Sci. Eng.5
2025 Addressing Inconsistent Labeling With Cross Image Matching for Scribble-Based Medical Image Segmentation
abstract
In recent years, there has been a notable surge in the adoption of weakly-supervised learning for medical image segmentation, utilizing scribble annotation as a means to potentially reduce annotation costs. However, the inherent characteristics of scribble labeling, marked by incompleteness, subjectivity, and a lack of standardization, introduce inconsistencies into the annotations. These inconsistencies become significant challenges for the network's learning process, ultimately affecting the performance of segmentation. To address this challenge, we propose creating a reference set to guide pixel-level feature matching, constructed from class-specific tokens and pixel-level features extracted from variously images. Serving as a repository showcasing diverse pixel styles and classes, the reference set becomes the cornerstone for a pixel-level feature matching strategy. This strategy enables the effective comparison of unlabeled pixels, offering guidance, particularly in learning scenarios characterized by inconsistent and incomplete scribbles. The proposed strategy incorporates smoothing and regression techniques to align pixel-level features across different images. By leveraging the diversity of pixel sources, our matching approach enhances the network's ability to learn consistent patterns from the reference set. This, in turn, mitigates the impact of inconsistent and incomplete labeling, resulting in improved segmentation outcomes. Extensive experiments conducted on three publicly available datasets demonstrate the superiority of our approach over state-of-the-art methods in terms of segmentation accuracy and stability. The code will be made publicly available at https://github.com/jingkunchen/scribble-medical-segmentation.
Jingkun Chen, Wenjian Huang 0001, Jianguo Zhang 0001, Kurt Debattista, Jungong Han
IEEE Trans. Image Process.4
2025 Counterfactual Explainer for Deep Reinforcement Learning Models using Policy Distillation
abstract
Deep Reinforcement Learning (DRL) has demonstrated promising capability in solving complex control problems. However, DRL applications in safety-critical systems are hindered by the inherent lack of robust validation techniques to assure their performance in such applications. One of the key requirements of the verification process is the development of effective techniques to explain the system functionality, providing why the system produces specific results in given circumstances. Recently, interpretation methods based on the Counterfactual (CF) explanation approach have been proposed to address the problem of explanation in DRLs. This article proposes a novel CF explainer to interpret the decisions made by a black-box DRL. To evaluate the efficacy of the proposed explanation framework, we carried out several experiments in the domains of Automated Driving Systems (ADSs) and the Atari Pong game. Our analysis demonstrates that the proposed framework generates plausible and meaningful explanations for various decisions made by deep underlying DRLs. Additionally, we discuss the practical implications of our approach for various automotive stakeholders, illustrating its potential real-world impact. Source codes are available at https://github.com/Amir-Samadi/Counterfactual-Explanation .
Amir Samadi, Konstantinos Koufos, Kurt Debattista, Mehrdad Dianati
ACM Trans. Intell. Syst. Technol.3
2025 Region-Object Relation-Aware Dense Captioning via Transformer
abstract
Dense captioning provides detailed captions of complex visual scenes. While a number of successes have been achieved in recent years, there are still two broad limitations: 1) most existing methods adopt an encoder-decoder framework, where the contextual information is sequentially encoded using long short-term memory (LSTM). However, the forget gate mechanism of LSTM makes it vulnerable when dealing with a long sequence and 2) the vast majority of prior arts consider regions of interests (RoIs) equally important, thus failing to focus on more informative regions. The consequence is that the generated captions cannot highlight important contents of the image, which does not seem natural. To overcome these limitations, in this article, we propose a novel end-to-end transformer-based dense image captioning architecture, termed the transformer-based dense captioner (TDC). TDC learns the mapping between images and their dense captions via a transformer, prioritizing more informative regions. To this end, we present a novel unit, named region-object correlation score unit (ROCSU), to measure the importance of each region, where the relationships between detected objects and the region, alongside the confidence scores of detected objects within the region, are taken into account. Extensive experimental results and ablation studies on the standard dense-captioning datasets demonstrate the superiority of the proposed method to the state-of-the-art methods.
Jungong Han, Demetris Marnerides, Kurt Debattista
IEEE Trans. Neural Networks Learn. Syst.4
2024 The Significance of Interaction in Determining Learning Outcomes in Serious Games
Khairul Fadhli Mohammad, Alan Chalmers, Kurt Debattista
CGI (2)3
2024 Pseudo-labelling Should Be Aware of Disguising Channel Activations
Changrui Chen, Kurt Debattista, Jungong Han
ECCV (63)2
2024 Exploring Generative AI for Sim2Real in Driving Data Synthesis
abstract
Datasets are essential for training and testing vehicle perception algorithms. However, the collection and annotation of real-world images is time-consuming and expensive. Driving simulators offer a solution by automatically generating various driving scenarios with corresponding annotations, but the simulation-to-reality (Sim2Real) domain gap remains a challenge. While most of the Generative Artificial Intelligence (AI) follows the de facto Generative Adversarial Nets (GANs)-based methods, the recent emerging diffusion probabilistic models have not been fully explored in mitigating Sim2Real challenges for driving data synthesis. To explore the performance, this paper applied three different generative AI methods to leverage semantic label maps from a driving simulator as a bridge for the creation of realistic datasets. A comparative analysis of these methods is presented from the perspective of image quality and perception. New synthetic datasets, which include driving images and auto-generated high-quality annotations, are produced with low costs and high scene variability. The experimental results show that although GAN-based methods are adept at generating high-quality images when provided with manually annotated labels, ControlNet produces synthetic datasets with fewer artefacts and more structural fidelity when using simulator-generated labels. This suggests that the diffusion-based approach may provide improved stability and an alternative method for addressing Sim2Real challenges. These insights contribute to the intelligent vehicle community’s understanding of the potential for diffusion models to mitigate the Sim2Real gap.
Thomas Bashford-Rogers, Valentina Donzella, Kurt Debattista
IV5
2024 Revisiting motion information for RGB-Event tracking with MOT philosophy
abstract
RGB-Event single object tracking (SOT) aims to leverage the merits of RGB and event data to achieve higher performance. However, existing frameworks focus on exploring complementary appearance information within multi-modal data, and struggle to address the association problem of targets and distractors in the temporal domain using motion information from the event stream. In this paper, we introduce the Multi-Object Tracking (MOT) philosophy into RGB-E SOT to keep track of targets as well as distractors by using both RGB and event data, thereby improving the robustness of the tracker. Specifically, an appearance model is employed to predict the initial candidates. Subsequently, the initially predicted tracking results, in combination with the RGB-E features, are encoded into appearance and motion embeddings, respectively. Furthermore, a Spatial-Temporal Transformer Encoder is proposed to model the spatial-temporal relationships and learn discriminative features for each candidate through guidance of the appearance-motion embeddings. Simultaneously, a Dual-Branch Transformer Decoder is designed to adopt such motion and appearance information for candidate matching, thus distinguishing between targets and distractors. The proposed method is evaluated on multiple benchmark datasets and achieves state-of-the-art performance on all the datasets tested.
Tianlu Zhang, Kurt Debattista, Qiang Zhang 0020, Guiguang Ding, Jungong Han
NeurIPS2
2024 Virtual Category Learning: A Semi-Supervised Learning Method for Dense Prediction With Extremely Limited Labels
abstract
Due to the costliness of labelled data in real-world applications, semi-supervised learning, underpinned by pseudo labelling, is an appealing solution. However, handling confusing samples is nontrivial: discarding valuable confusing samples would compromise the model generalisation while using them for training would exacerbate the issue of confirmation bias caused by the resulting inevitable mislabelling. To solve this problem, this paper proposes to use confusing samples proactively without label correction. Specifically, a Virtual Category (VC) is assigned to each confusing sample in such a way that it can safely contribute to the model optimisation even without a concrete label. This provides an upper bound for inter-class information sharing capacity, which eventually leads to a better embedding space. Extensive experiments on two mainstream dense prediction tasks - semantic segmentation and object detection, demonstrate that the proposed VC learning significantly surpasses the state-of-the-art, especially when only very few labels are available. Our intriguing findings highlight the usage of VC learning in dense vision tasks.
Changrui Chen, Jungong Han, Kurt Debattista
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Dynamic contrastive learning guided by class confidence and confusion degree for medical image segmentation
Jingkun Chen, Changrui Chen, Wenjian Huang 0001, Jianguo Zhang 0001, Kurt Debattista, Jungong Han
Pattern Recognit.5
2024 Generating Synthetic Training Images to Detect Split Defects in Stamped Components
abstract
Detecting rare and costly defects, such as necks and splits in sheet metal stamping, remains challenging for deep learning models due to low failure rates entailing few available samples to train on. Synthetic images provide a simulated alternative; however, the two main current approaches have limitations for generating split defect images. Image synthesis-based models generate implausible training data, while physics-based models are computationally expensive and lack the diversity required. To address this, we present a novel method combining the advantages of physics-based simulation with synthetic-based defect generation. The method first generates deformed 3-D geometry through finite element simulation with plausible split locations determined using a forming limit curve. Subsequently, the fine details of captured real splits are mapped to the identified locations to generate realistic defect features. Our results show that training a deep neural network with the addition of synthetic images improves the performance significantly.
Aru Ranjan Singh, Thomas Bashford-Rogers, Sumit Hazra 0002, Kurt Debattista
IEEE Trans. Ind. Informatics4
2024 DCMSTRD: End-to-end Dense Captioning via Multi-Scale Transformer Decoding
abstract
Dense captioning creates diverse Region of Interests (RoIs) descriptions for complex visual scenes. While promising results have been obtained, several issues persist. In particular: 1) it is hard to find the optimal parameters for artificially designed modules (e.g., non-maximum suppression (NMS)) causing redundancies and fewer interactions to benefit the two sub-tasks of RoI detection and RoI captioning; 2) the absence of a multi-scale decoder in current methods hinders the acquisition of scale-invariant features, thus leading to poor performance. To tackle these limitations, we bypass the artificially designed modules and present an end-to-end dense captioning framework via multi-scale transformer decoding (DCMSTRD). DCMSTRD solves dense captioning by set matching and prediction instead. To further enhance the discriminative quality of the multi-scale representations during caption generation, we introduce a multi-scale module, termed multi-scale language decoder (MSLD). Our proposed method tested on standard datasets achieves a mean Average Precision (mAP) of 16.7% on the challenging VG-COCO dataset, demonstrating its effectiveness against the current methods.
Jungong Han, Kurt Debattista, Yanwei Pang
IEEE Trans. Multim.3
2024 Self-supervised High Dynamic Range Imaging: What Can Be Learned from a Single 8-bit Video?
abstract
Recently, Deep Learning-based methods for inverse tone mapping standard dynamic range (SDR) images to obtain high dynamic range (HDR) images have become very popular. These methods manage to fill over-exposed areas convincingly both in terms of details and dynamic range. To be effective, deep learning-based methods need to learn from large datasets and transfer this knowledge to the network weights. In this work, we tackle this problem from a completely different perspective. What can we learn from a single SDR 8-bit video? With the presented self-supervised approach, we show that, in many cases, a single SDR video is sufficient to generate an HDR video of the same quality or better than other state-of-the-art methods.
Francesco Banterle, Demetris Marnerides, Thomas Bashford-Rogers, Kurt Debattista
ACM Trans. Graph.4
2023 Accelerating Stereo Image Simulation for Automotive Applications Using Neural Stereo Super Resolution
abstract
Camera image simulation is integral to the virtual validation of autonomous vehicles and robots that use visual perception to understand their environment. It also has applications in creating image datasets for training learning-based vision models. As camera image simulation takes into account a wide variety of external and internal parameters, achieving a high-fidelity simulation is a computationally expensive process. Recently, several neural network-based techniques have been proposed to reduce the computational complexity of image rendering, a critical element of the camera simulation pipeline. However, the existing methods are tailored for monocular camera images and are not optimised for stereo images, which are widely used in autonomous driving applications. To address this, we propose a technique based on Stereo Super Resolution (SSR) to speed up the simulation of stereo images. The proposed method first simulates stereo images at a lower resolution, then super-resolves them to their original resolution using our introduced SSR model, ETSSR. We evaluated the performance of our technique using the CARLA driving simulator and created our own synthetic dataset for training ETSSR. The evaluations indicate that our approach can speed up stereo image simulation by a factor of up to 2.57 over various resolutions. Moreover, it shows that our ETSSR achieves on-par or superior performance compared to the state-of-the-art models, using significantly fewer parameters and FLOPs. We have made our source code and dataset available athttps://github.com/hamedhaghighi/ETSSR.
Hamed Haghighi, Mehrdad Dianati, Valentina Donzella, Kurt Debattista
IEEE Trans. Intell. Transp. Syst.4
2023 Semi-Supervised Unpaired Medical Image Segmentation Through Task-Affinity Consistency
abstract
Deep learning-based semi-supervised learning (SSL) algorithms are promising in reducing the cost of manual annotation of clinicians by using unlabelled data, when developing medical image segmentation tools. However, to date, most existing semi-supervised learning (SSL) algorithms treat the labelled images and unlabelled images separately and ignore the explicit connection between them; this disregards essential shared information and thus hinders further performance improvements. To mine the shared information between the labelled and unlabelled images, we introduce a class-specific representation extraction approach, in which a task-affinity module is specifically designed for representation extraction. We further cast the representation into two different views of feature maps; one is focusing on low-level context, while the other concentrates on structural information. The two views of feature maps are incorporated into the task-affinity module, which then extracts the class-specific representations to aid the knowledge transfer from the labelled images to the unlabelled images. In particular, a task-affinity consistency loss between the labelled images and unlabelled images based on the multi-scale class-specific representations is formulated, leading to a significant performance improvement. Experimental results on three datasets show that our method consistently outperforms existing state-of-the-art methods. Our findings highlight the potential of consistency between class-specific knowledge for semi-supervised medical image segmentation. The code and models are to be made publicly available at https://github.com/jingkunchen/TAC.
Jingkun Chen, Jianguo Zhang 0001, Kurt Debattista, Jungong Han
IEEE Trans. Medical Imaging3
2023 Textual Context-Aware Dense Captioning With Diverse Words
abstract
Dense captioning generates more detailed spoken descriptions for complex visual scenes. Despite several promising leads, existing methods still have two broad limitations: 1) The vast majority of prior arts only consider visual contextual clues during captioning but ignore potentially important textual context; 2) current imbalanced learning mechanisms limit the diversity of vocabulary learned from the dictionary, thus giving rise to low language-learning efficiency. To alleviate these gaps, in this paper, we propose an end-to-end enhanced dense captioning architecture, namely Enhanced Transformer Dense Captioner (ETDC), which obtains textual context from surrounding regions and dynamically diversifies the vocabulary bank during captioning. Concretely, we first propose the Textual Context Module (TCM), which is integrated into each self-attention layer of the Transformer decoder, to capture the surrounding textual context. Moreover, we take full advantage of the class information of object context and propose a Dynamic Vocabulary Frequency Histogram (DVFH) re-sampling strategy during training to balance words with different frequencies. The proposed method is tested on the standard dense captioning datasets and surpasses the state-of-the-art methods in terms of mean Average Precision (mAP).
Jungong Han, Kurt Debattista, Yanwei Pang
IEEE Trans. Multim.3
2023 A genetic algorithm for backlight dimming for HDR displays
Lvyin Duan, Kurt Debattista, Guanghui Yue 0001, Demetris Marnerides, Alan Chalmers
Vis. Comput.2
2022 Semi-supervised Object Detection via VC Learning
Changrui Chen, Kurt Debattista, Jungong Han
ECCV (31)2
2022 Recognising place under distinct weather variability, a comparison between end-to-end and metric learning approaches
abstract
Autonomous driving requires robust and accurate real time localisation information to navigate and perform trajectory planning. Although Global Navigation Satellite Systems (GNSS) are most frequently used in this application, they are unreliable within urban environments because of multipath and non-line-of-sight errors. Alternative solutions exist that exploit rich visual content from images that can be corresponded with a stored representation, such as a map, to determine the vehicles location. However, one major cause of reduced location accuracy are variations in environmental conditions between the images captured and those stored in the representation. We tackle this issue directly by collecting a simulated and real-world dataset captured over a single route under multiple environmental conditions. We demonstrate the effectiveness of an end-to-end approach in recognising place and by extension determining vehicle location.
Stephane Role, Demetris Marnerides, Kurt Debattista, Stefano Cavazzi, Mehrdad Dianati
IV3
2022 Ensemble Metropolis Light Transport
abstract
This article proposes a Markov Chain Monte Carlo ( MCMC ) rendering algorithm based on a family of guided transition kernels. The kernels exploit properties of ensembles of light transport paths, which are distributed according to the lighting in the scene, and utilize this information to make informed decisions for guiding local path sampling. Critically, our approach does not require caching distributions in world space, saving time and memory, yet it is able to make guided sampling decisions based on whole paths. We show how this can be implemented efficiently by organizing the paths in each ensemble and designing transition kernels for MCMC rendering based on a carefully chosen subset of paths from the ensemble. This algorithm is easy to parallelize and leads to improvements in variance when rendering a variety of scenes.
Thomas Bashford-Rogers, Luís Paulo Santos, Demetris Marnerides, Kurt Debattista
ACM Trans. Graph.4
2021 Optical effects on HDR calibration via a multiple exposure noise-based workflow
abstract
Abstract High dynamic range (HDR) technology allows more of the lighting in a specific scene to be captured at a set point in time, and thus is capable of delivering an overall view of the scene that more closely correlates with our visual experience in the real world, compared to standard, or low dynamic range (LDR) technology. Although HDR capabilities of single exposure capture systems are improving, the traditional method for creating HDR images still includes combing a number of different exposures, captured with an LDR system, into a single HDR image. Several use cases requiring absolute calibration of the resulting HDR luminance map have been undertaken, but none of these have provided a detailed analysis of the optical effects of glare on the results. We develop a calibrated HDR radiance map, including methodical linearization of captured image data, while characterizing the limitations due to the effects of optical glare. A purposely designed controlled test scene is used to challenge the calibrated reconstruction efforts, including low luminance levels, spatial inclusion of lens vignette over the full imaged area, and optical glare. Results demonstrate that even with careful processing and recombination of the LDR data, radiometric accuracy is limited as a result of glare. The proposed approach performs better than calibration methods in commercially available HDR recombination software.
Brian Karr, Kurt Debattista, Alan Chalmers
Vis. Comput.2
2020 A technology-aided multi-modal training approach to assist abdominal palpation training and its assessment in medical education
Ali Asadipour 0001, Kurt Debattista, Vinod Patel, Alan Chalmers
Int. J. Hum. Comput. Stud.2
2020 Per-pixel classification of clouds from whole sky HDR images
Pinar Satilmis, Thomas Bashford-Rogers, Alan Chalmers, Kurt Debattista
Signal Process. Image Commun.4
2019 Displaying detail in bright environments: A 10, 000 nit display and its evaluation
Jonathan Hatchett, Domenico Toffoli, Miguel Melo, Maximino Bessa, Kurt Debattista, Alan Chalmers
Signal Process. Image Commun.5
2019 Uniform Color Space-Based High Dynamic Range Video Compression
abstract
Recently, there has been a significant progress in the research and development of the high dynamic range (HDR) video technology and the state-of-the-art video pipelines are able to offer a higher bit depth support to capture, store, encode, and display HDR video content. In this paper, we introduce a novel HDR video compression algorithm, which uses a perceptually uniform color opponent space, a novel perceptual transfer function to encode the dynamic range of the scene, and a novel error minimization scheme for accurate chroma reproduction. The proposed algorithm was objectively and subjectively evaluated against four state-of-the-art algorithms. The objective evaluation was conducted across a set of 39 HDR video sequences, using the latest x265 10-bit video codec along with several perceptual and structural quality assessment metrics at 11 different quality levels. Furthermore, a rating-based subjective evaluation (n=40) was conducted with six sequences at two different output bitrates. Results suggest that the proposed algorithm exhibits the lowest coding error amongst the five algorithms evaluated. Additionally, the rate-distortion characteristics suggest that the proposed algorithm outperforms the existing state-of-the-art at bitrates ≥ 0.4 bits/pixel.
Ratnajit Mukherjee, Kurt Debattista, Thomas Bashford-Rogers, Maximino Bessa, Alan Chalmers
IEEE Trans. Circuits Syst. Video Technol.2
2019 Audio-Visual-Olfactory Resource Allocation for Tri-modal Virtual Environments
abstract
Virtual Environments (VEs) provide the opportunity to simulate a wide range of applications, from training to entertainment, in a safe and controlled manner. For applications which require realistic representations of real world environments, the VEs need to provide multiple, physically accurate sensory stimuli. However, simulating all the senses that comprise the human sensory system (HSS) is a task that requires significant computational resources. Since it is intractable to deliver all senses at the highest quality, we propose a resource distribution scheme in order to achieve an optimal perceptual experience within the given computational budgets. This paper investigates resource balancing for multi-modal scenarios composed of aural, visual and olfactory stimuli. Three experimental studies were conducted. The first experiment identified perceptual boundaries for olfactory computation. In the second experiment, participants ( N=25) were asked, across a fixed number of budgets ( M=5), to identify what they perceived to be the best visual, acoustic and olfactory stimulus quality for a given computational budget. Results demonstrate that participants tend to prioritize visual quality compared to other sensory stimuli. However, as the budget size is increased, users prefer a balanced distribution of resources with an increased preference for having smell impulses in the VE. Based on the collected data, a quality prediction model is proposed and its accuracy is validated against previously unused budgets and an untested scenario in a third and final experiment.
Efstratios Doukakis, Kurt Debattista, Thomas Bashford-Rogers, Amar Dhokia, Ali Asadipour 0001, Alan Chalmers, Carlo Harvey
IEEE Trans. Vis. Comput. Graph.2
2019 An asynchronous method for cloud-based rendering
Keith Bugeja, Kurt Debattista, Sandro Spina
Vis. Comput.2
2018 Application-Specific Tone Mapping Via Genetic Programming
abstract
Abstract High dynamic range (HDR) imagery permits the manipulation of real‐world data distinct from the limitations of the traditional, low dynamic range (LDR), content. The process of retargeting HDR content to traditional LDR imagery via tone mapping operators (TMOs) is useful for visualizing HDR content on traditional displays, supporting backwards‐compatible HDR compression and, more recently, is being frequently used for input into a wide variety of computer vision applications. This work presents the automatic generation of TMOs for specific applications via the evolutionary computing method of genetic programming (GP). A straightforward, generic GP method that generates TMOs for a given fitness function and HDR content is presented. Its efficacy is demonstrated in the context of three applications: Visualization of HDR content on LDR displays, feature mapping and compression. For these applications, results show good performance for the generated TMOs when compared to traditional methods. Furthermore, they demonstrate that the method is generalizable and could be used across various applications that require TMOs but for which dedicated successful TMOs have not yet been discovered.
Kurt Debattista
Comput. Graph. Forum1
2018 Frame Rate vs Resolution: A Subjective Evaluation of Spatiotemporal Perceived Quality Under Varying Computational Budgets
abstract
Abstract Maximizing performance for rendered content requires making compromises on quality parameters depending on the computational resources available . Yet, it is currently unclear which parameters best maximize perceived quality. This work investigates perceived quality across computational budgets for the primary spatiotemporal parameters of resolution and frame rate. Three experiments are conducted. Experiment 1 (n = 26) shows that participants prefer fixed frame rates of 60 frames per second (fps) at lower resolutions over 30 fps at higher resolutions. Experiment 2 (n = 24) explores the relationship further with more budgets and quality settings and again finds 60 fps is generally preferred even when more resources are available. Experiment 3 (n = 25) permits the use of adaptive frame rates, and analyses the resource allocation across seven budgets. Results show that while participants allocate more resources to frame rate at lower budgets the situation reverses once higher budgets are available and a frame rate of around 40 fps is achieved. In the overall, the results demonstrate a complex relationship between frame rate and resolution's effects on perceived quality. This relationship can be harnessed, via the results and models presented, to obtain more cost‐effective virtual experiences.
Kurt Debattista, Keith Bugeja, Sandro Spina, Thomas Bashford-Rogers, Vedad Hulusic
Comput. Graph. Forum1
2018 Audiovisual Resource Allocation for Bimodal Virtual Environments
abstract
Abstract Fidelity is of key importance if virtual environments are to be used as authentic representations of real environments. However, simulating the multitude of senses that comprise the human sensory system is computationally challenging. With limited computational resources, it is essential to distribute these carefully in order to simulate the most ideal perceptual experience. This paper investigates this balance of resources across multiple scenarios where combined audiovisual stimulation is delivered to the user. A subjective experiment was undertaken where participants (N=35) allocated five fixed resource budgets across graphics and acoustic stimuli. In the experiment, increasing the quality of one of the stimuli decreased the quality of the other. Findings demonstrate that participants allocate more resources to graphics; however, as the computational budget is increased, an approximately balanced distribution of resources is preferred between graphics and acoustics. Based on the results, an audiovisual quality prediction model is proposed and successfully validated against previously untested budgets and an untested scenario.
Efstratios Doukakis, Kurt Debattista, Carlo Harvey, Thomas Bashford-Rogers, Alan Chalmers
Comput. Graph. Forum2
2018 Olfaction and Selective Rendering
abstract
Abstract Accurate simulation of all the senses in virtual environments is a computationally expensive task. Visual saliency models have been used to improve computational performance for rendered content, but this is insufficient for multi‐modal environments. This paper considers cross‐modal perception and, in particular, if and how olfaction affects visual attention. Two experiments are presented in this paper. Firstly, eye tracking is gathered from a number of participants to gain an impression about where and how they view virtual objects when smell is introduced compared to an odourless condition. Based on the results of this experiment a new type of saliency map in a selective‐rendering pipeline is presented. A second experiment validates this approach, and demonstrates that participants rank images as better quality, when compared to a reference, for the same rendering budget.
Carlo Harvey, Thomas Bashford-Rogers, Kurt Debattista, Efstratios Doukakis, Alan Chalmers
Comput. Graph. Forum3
2018 ExpandNet: A Deep Convolutional Neural Network for High Dynamic Range Expansion from Low Dynamic Range Content
abstract
Abstract High dynamic range (HDR) imaging provides the capability of handling real world lighting as opposed to the traditional low dynamic range (LDR) which struggles to accurately represent images with higher dynamic range. However, most imaging content is still available only in LDR. This paper presents a method for generating HDR content from LDR content based on deep Convolutional Neural Networks (CNNs) termed ExpandNet. ExpandNet accepts LDR images as input and generates images with an expanded range in an end‐to‐end fashion. The model attempts to reconstruct missing information that was lost from the original signal due to quantization, clipping, tone mapping or gamma correction. The added information is reconstructed from learned features, as the network is trained in a supervised fashion using a dataset of HDR images. The approach is fully automatic and data driven; it does not require any heuristics or human expertise. ExpandNet uses a multiscale architecture which avoids the use of upsampling layers to improve image quality. The method performs well compared to expansion/inverse tone mapping operators quantitatively on multiple metrics, even for badly exposed inputs.
Demetris Marnerides, Thomas Bashford-Rogers, Jonathan Hatchett, Kurt Debattista
Comput. Graph. Forum4
2018 Evaluating practitioner cyber-security attack graph configuration preferences
Harjinder Singh Lallie 0001, Kurt Debattista, Jay Bal
Comput. Secur.2
2018 Subjective Evaluation of High-Fidelity Virtual Environments for Driving Simulations
abstract
Virtual environments (VEs) grant the ability to experience real-world scenarios, such as driving, in a virtual, safe, and reproducible context. However, in order to achieve their full potential, the fidelity of the VE must provide confidence that it replicates the perception of the real-world experience. The computational cost of simulating real-world visuals accurately means that compromises to the fidelity of the visuals must be made. In this paper, a subjective evaluation of driving in a VE at different quality settings is presented. Participants (n = 44) were driven around in the real world and in a purposely built representative VE and the fidelity of the graphics and overall experience at low-, medium-, and high-visual settings were analyzed. Low quality corresponds to the illumination in many current traditional simulators, medium to a higher quality using accurate shadows and reflections, and high to the quality experienced in modern movies and simulations that require hours of computation. Results demonstrate that graphics quality affects the perceived fidelity of the visuals and the overall experience. When judging the overall experience, participants could tell the difference between the lower quality graphics and the rest but did not significantly discriminate between the medium and higher graphical settings. This indicates that future driving simulators should improve the quality, but once the equivalent of the presented medium quality is reached, they may not need to do so significantly.
Kurt Debattista, Thomas Bashford-Rogers, Carlo Harvey, Brian Waterfield, Alan Chalmers
IEEE Trans. Hum. Mach. Syst.1
2018 An Empirical Evaluation of the Effectiveness of Attack Graphs and Fault Trees in Cyber-Attack Perception
abstract
Perceiving and understanding cyber-attacks can be a difficult task. This problem is widely recognized and welldocumented, and more effective techniques are needed to aid cyber-attack perception. Attack modeling techniques (AMTs), such as attack graphs and fault trees, are useful visual aids that can aid cyber-attack perception; however, there is little empirical or comparative research which evaluates the effectiveness of these methods. This paper reports the results of an empirical evaluation between an adapted attack graph method and the fault tree standard to determine which of the two methods is more effective in aiding cyber-attack perception. An empirical evaluation (n = 63) was conducted through a 3 × 2 × 2 factorial design. Participants from computer-science and non-computerscience backgrounds were divided into an adapted attack graph and fault tree group and then asked to complete three tests which tested the ability to recall, comprehend, and apply the AMT. A mean assessment score (mas) was calculated for each test. The results show that the adapted attack graph method is more effective at aiding cyber-attack perception when compared with the fault tree method (p <; 0.01). Participants that have a computer science background outperformed other participants when using both methods (p <; 0.05). These results indicate that the adapted attack graph method can be an effective tool for aiding cyber-attack perception amongst experts. This paper underlines the need for further comparisons in a broader range of settings involving additional techniques, and makes several suggestions for further work.
Harjinder Singh Lallie 0001, Kurt Debattista, Jay Bal
IEEE Trans. Inf. Forensics Secur.2
2018 An evaluation of power transfer functions for HDR video compression
abstract
High dynamic range (HDR) imaging enables the full range of light in a scene to be captured, transmitted and displayed. However, uncompressed 32-bit HDR is four times larger than traditional low dynamic range (LDR) imagery. If HDR is to fulfil its potential for use in live broadcasts and interactive remote gaming, fast, efficient compression is necessary for HDR video to be manageable on existing communications infrastructure. A number of methods have been put forward for HDR video compression. However, these can be relatively complex and frequently require the use of multiple video streams. In this paper, we propose the use of a straightforward Power Transfer Function (PTF) as a practical, computationally fast, HDR video compression solution. The use of PTF is presented and evaluated against four other HDR video compression methods. An objective evaluation shows that PTF exhibits improved quality at a range of bit-rates and, due to its straightforward nature, is highly suited for real-time HDR video applications.
Jonathan Hatchett, Kurt Debattista, Ratnajit Mukherjee, Thomas Bashford-Rogers, Alan Chalmers
Vis. Comput.2
2017 Multi-Modal Perception for Selective Rendering
abstract
Abstract A major challenge in generating high‐fidelity virtual environments (VEs) is to be able to provide realism at interactive rates. The high‐fidelity simulation of light and sound is still unachievable in real time as such physical accuracy is very computationally demanding. Only recently has visual perception been used in high‐fidelity rendering to improve performance by a series of novel exploitations; to render parts of the scene that are not currently being attended to by the viewer at a much lower quality without the difference being perceived. This paper investigates the effect spatialized directional sound has on the visual attention of a user towards rendered images. These perceptual artefacts are utilized in selective rendering pipelines via the use of multi‐modal maps. The multi‐modal maps are tested through psychophysical experiments to examine their applicability to selective rendering algorithms, with a series of fixed cost rendering functions, and are found to perform significantly better than only using image saliency maps that are naively applied to multi‐modal VEs.
Carlo Harvey, Kurt Debattista, Thomas Bashford-Rogers, Alan Chalmers
Comput. Graph. Forum2
2017 A Subjective Evaluation of Texture Synthesis Methods
abstract
This paper presents the results of a user study which quantifies the relative and absolute quality of example-based texture synthesis algorithms. In order to allow such evaluation, a list of texture properties is compiled, and a minimal representative set of textures is selected to cover these. Six texture synthesis methods are compared against each other and a reference on a selection of twelve textures by non-expert participants (N = 67). Results demonstrate certain algorithms successfully solve the problem of texture synthesis for certain textures, but there are no satisfactory results for other types of texture properties. The presented textures and results make it possible for future work to be subjectively compared, thus facilitating the development of future texture synthesis methods.
Martin Kolár, Kurt Debattista, Alan Chalmers
Comput. Graph. Forum2
2017 Context-aware HDR video distribution for mobile devices
Miguel Melo, Luís Barbosa, Maximino Bessa, Kurt Debattista, Alan Chalmers
Multim. Tools Appl.4
2017 HDR video past, present and future: A perspective
abstract
High dynamic range (HDR) video has emerged from research labs around the world and entered the realm of consumer electronics. The dynamic range that a human can see in a scene with minimal eye adaption (approximately 1,000,000:1) is vastly greater than traditional imaging technology which can only capture about 8 f -stops (256:1). HDR technology, on the other hand, has the potential to capture the full range of light in a scene; even more than a human eye can see. This paper examines the field of HDR video from capture to display: past, present and future. In particular the paper looks beyond the current marketing hype around HDR to show how HDR video in the future can and, indeed, should bring about a step change in imaging, analogous to the change from black and white to colour.
Alan Chalmers, Kurt Debattista
Signal Process. Image Commun.2
2017 A model of perceived dynamic range for HDR images
Vedad Hulusic, Kurt Debattista, Giuseppe Valenzise, Frédéric Dufaux
Signal Process. Image Commun.2
2017 Visuohaptic augmented feedback for enhancing motor skills acquisition
abstract
Serious games are accepted as an effective approach to deliver augmented feedback in motor (re-)learning processes. The multi-modal nature of the conventional computer games (e.g. audiovisual representation) plus the ability to interact via haptic-enabled inputs provides a more immersive experience. Thus, particular disciplines such as medical education in which frequent hands on rehearsals play a key role in learning core motor skills (e.g. physical palpations) may benefit from this technique. Challenges such as the impracticality of verbalising palpation experience by tutors and ethical considerations may prevent the medical students from correctly learning core palpation skills. This work presents a new data glove, built from off-the-shelf components which captures pressure sensitivity designed to provide feedback for palpation tasks. In this work the data glove is used to control a serious game adapted from the infinite runner genre to improve motor skill acquisition. A comparative evaluation on usability and effectiveness of the method using multimodal visualisations, as part of a larger study to enhance pressure sensitivity, is presented. Thirty participants divided into a game-playing group ( $$n=15$$ n = 15 ) and a control group ( $$n=15$$ n = 15 ) were invited to perform a simple palpation task. The game-playing group significantly outperformed the control group in which abstract visualisation of force was provided to the users in a blind-folded transfer test. The game-based training approach was positively described by the game-playing group as enjoyable and engaging.
Ali Asadipour 0001, Kurt Debattista, Alan Chalmers
Vis. Comput.2
2017 Preface to the special issue on VS-Games 2015
Kurt Debattista, Fotis Liarokapis
Vis. Comput.1
2016 Perceived dynamic range of HDR images
abstract
Although high dynamic range (HDR) imaging has gained great popularity and acceptance in both the scientific and commercial domains, the relationship between perceptually accurate, content-independent dynamic range and objective measures has not been fully explored. In this paper, a new methodology for perceived dynamic range evaluation of complex stimuli in HDR conditions is proposed. A subjective study with 20 participants was conducted and correlations between mean opinion scores (MOS) and three image features were analyzed. Strong Spearman correlations between MOS and objective DR measure and between MOS and image key were found. An exploratory analysis reveals that additional image characteristics should be considered when modeling perceptually-based dynamic range metrics. Finally, one of the outcomes of the study is the perceptually annotated HDR image dataset with MOS values, that can be used for HDR imaging algorithms and metric validation, content selection and analysis of aesthetic image attributes.
Vedad Hulusic, Giuseppe Valenzise, Edoardo Provenzi, Kurt Debattista, Frédéric Dufaux
QoMEX4
2016 Objective and subjective evaluation of High Dynamic Range video compression
Ratnajit Mukherjee, Kurt Debattista, Thomas Bashford-Rogers, Peter Vangorp, Rafal Mantiuk, Maximino Bessa, Brian Waterfield, Alan Chalmers
Signal Process. Image Commun.2
2016 Repeatable texture sampling with interchangeable patches
Martin Kolár, Alan Chalmers, Kurt Debattista
Vis. Comput.3
2016 A study on user preference of high dynamic range over low dynamic range video
abstract
The increased interest in high dynamic range (HDR) video over existing low dynamic range (LDR) video during the past decade or so was primarily due to its inherent capability to capture, store and display the full range of real-world lighting visible to the human eye with increased precision. This has led to an inherent assumption that HDR video would be preferable by the end-user over LDR video due to the more immersive and realistic visual experience provided by HDR. This assumption has led to a considerable body of research into efficient capture, processing, storage and display of HDR video. Although this is beneficial for scientific research and industrial purposes, very little research has been conducted to test the veracity of this assumption. In this paper, we conduct two subjective studies by means of a ranking and a rating-based experiment where 60 participants in total, 30 in each experiment, were tasked to rank and rate several reference HDR video scenes along with three mapped LDR versions of each scene on an HDR display, in order of their viewing preference. Results suggest that given the option, end-users prefer the HDR representation of the scene over its LDR counterpart.
Ratnajit Mukherjee, Kurt Debattista, Thomas Bashford-Rogers, Brian Waterfield, Alan Chalmers
Vis. Comput.2
2015 Evaluation of Tone-Mapping Operators for HDR Video Under Different Ambient Luminance Levels
abstract
Abstract Since high dynamic range (HDR) displays are not yet widely available, there is still a need to perform a dynamic range reduction of HDR content to reproduce it properly on standard dynamic range (SDR) displays. The most common techniques for performing this reduction are termed tone‐mapping operators (TMOs). Although mobile devices are becoming widespread, methods for displaying HDR content on these SDR screens are still very much in their infancy. While several studies have been conducted to evaluate TMOs, few have been done with a goal of testing small screen displays (SSDs), common on mobile devices. This paper presents an evaluation of six state‐of‐the‐art HDR video TMOs. The experiments considered three different levels of ambient luminance under which 180 participants were asked to rank the TMOs for seven tone‐mapped HDR video sequences. A comparison was conducted between tone‐mapped HDR video footage shown on an SSD and on a large screen SDR display using an HDR display as reference. The results show that there are differences between the performance of the TMOs under different ambient lighting levels and the TMOs that perform well on traditional large screen displays also perform well on SSDs at the same given luminance level.
Miguel Melo, Maximino Bessa, Kurt Debattista, Alan Chalmers
Comput. Graph. Forum3
2015 Optimal exposure compression for high dynamic range content
abstract
High dynamic range (HDR) imaging has become one of the foremost imaging methods capable of capturing and displaying the full range of lighting perceived by the human visual system in the real world. A number of HDR compression methods for both images and video have been developed to handle HDR data, but none of them has yet been adopted as the method of choice. In particular, the backwards-compatible methods that always maintain a stream/image that allow part of the content to be viewed on conventional displays make use of tone mapping operators which were developed to view HDR images on traditional displays. There are a large number of tone mappers, none of which is considered the best as the images produced could be deemed subjective. This work presents an alternative to tone mapping-based HDR content compression by identifying a single exposure that can reproduce the most information from the original HDR image. This single exposure can be adapted to fit within the bit depth of any traditional encoder. Any additional information that may be lost is stored as a residual. Results demonstrate quality is maintained as well, and better, than other traditional methods. Furthermore, the presented method is backwards-compatible, straightforward to implement, fast and does not require choosing tone mappers or settings.
Kurt Debattista, Thomas Bashford-Rogers, Elmedin Selmanovic, Ratnajit Mukherjee, Alan Chalmers
Vis. Comput.1
2014 Evaluation of HDR video tone mapping for mobile devices
Miguel Melo, Maximino Bessa, Kurt Debattista, Alan Chalmers
Signal Process. Image Commun.3
2014 Enabling stereoscopic high dynamic range video
Elmedin Selmanovic, Kurt Debattista, Thomas Bashford-Rogers, Alan Chalmers
Signal Process. Image Commun.2
2014 Importance Driven Environment Map Sampling
abstract
In this paper we present an efficient method for supporting image based lighting (IBL) for bidirectional methods. This improves both sampling of the environment, and the detection and sampling of important regions of the scene, such as windows and doors. These parts of the scene often have a small area proportional to that of the entire scene, so paths which pass through them are generated with a low probability. The method proposed in this paper improves sampling efficiency, by taking into account view importance, and modifies the lighting distribution to use light transport information from the camera. This method automatically constructs a sampling distribution in locations which are relevant to the camera position, thereby improving sampling of light paths. This approach can be applied to several bidirectional rendering methods, and results are shown for bidirectional path tracing, metropolis light transport and progressive photon mapping. When compared to other methods, efficiency results demonstrate speed ups of orders of magnitude.
Thomas Bashford-Rogers, Kurt Debattista, Alan Chalmers
IEEE Trans. Vis. Comput. Graph.2
2013 Foreword - High dynamic range imaging
Luís Paulo Santos, Kurt Debattista
Comput. Graph.2
2013 Generating stereoscopic HDR images using HDR-LDR image pairs
abstract
A number of novel imaging technologies have been gaining popularity over the past few years. Foremost among these are stereoscopy and high dynamic range (HDR) Imaging. While a large body of research has looked into each of these imaging technologies independently, very little work has attempted to combine them. This is mostly due to the current limitations in capture and display. In this article, we mitigate problems of capturing Stereoscopic HDR (SHDR) that would potentially require two HDR cameras, by capturing an HDR and LDR pair and using it to generate 3D stereoscopic HDR content. We ran a detailed user study to compare four different methods of generating SHDR content. The methods investigated were the following: two based on expanding the luminance of the LDR image, and two utilizing stereo correspondence methods, which were adapted for our purposes. Results demonstrate that one of the stereo correspondence methods may be considered perceptually indistinguishable from the ground truth (image pair captured using two HDR cameras), while the other methods are all significantly distinct from the ground truth.
Elmedin Selmanovic, Kurt Debattista, Thomas Bashford-Rogers, Alan Chalmers
ACM Trans. Appl. Percept.2
2013 Smoothness perception - Investigation of beat rate effect on frame rate perception
Vedad Hulusic, Kurt Debattista, Alan Chalmers
Vis. Comput.2
2012 A Significance Cache for Accelerating Global Illumination
abstract
Abstract Rendering using physically based methods requires substantial computational resources. Most methods that are physically based use straightforward techniques that may excessively compute certain types of light transport, while ignoring more important ones. Importance sampling is an effective and commonly used technique to reduce variance in such methods. Most current approaches for physically based rendering based on Monte Carlo methods sample the BRDF and cosine term, but are unable to sample the indirect illumination as this is the term that is being computed. Knowledge of the incoming illumination can be especially useful in the case of hard to find light paths, such as caustics or scenes which rely primarily on indirect illumination. To facilitate the determination of such paths, we propose a caching scheme which stores important directions, and is analytically sampled to calculate important paths. Results show an improvement over BRDF sampling and similar illumination importance sampling.
Thomas Bashford-Rogers, Kurt Debattista, Alan Chalmers
Comput. Graph. Forum2
2012 Cultural Heritage Predictive Rendering
abstract
Abstract High‐fidelity rendering can be used to investigate Cultural Heritage (CH) sites in a scientifically rigorous manner. However, a high degree of realism in the reconstruction of a CH site can be misleading insofar as it can be seen to imply a high degree of certainty about the displayed scene—which is frequently not the case, especially when investigating the past. So far, little effort has gone into adapting and formulating a Predictive Rendering pipeline for CH research applications. In this paper, we first discuss the goals and the workflow of CH reconstructions in general, as well as those of traditional Predictive Rendering. Based on this, we then propose a research framework for CH research, which we refer to as ‘Cultural Heritage Predictive Rendering’ (CHPR). This is an extension to Predictive Rendering that introduces a temporal component and addresses uncertainty that is important for the scene’s historical interpretation. To demonstrate these concepts, two example case studies are detailed.
Jassim Happa, Thomas Bashford-Rogers, Alexander Wilkie, Alessandro Artusi, Kurt Debattista, Alan Chalmers
Comput. Graph. Forum5
2012 Acoustic Rendering and Auditory-Visual Cross-Modal Perception and Interaction
abstract
Abstract In recent years research in the three‐dimensional sound generation field has been primarily focussed upon new applications of spatialized sound. In the computer graphics community the use of such techniques is most commonly found being applied to virtual, immersive environments. However, the field is more varied and diverse than this and other research tackles the problem in a more complete, and computationally expensive manner. Furthermore, the simulation of light and sound wave propagation is still unachievable at a physically accurate spatio‐temporal quality in real time. Although the Human Visual System (HVS) and the Human Auditory System (HAS) are exceptionally sophisticated, they also contain certain perceptional and attentional limitations. Researchers, in fields such as psychology, have been investigating these limitations for several years and have come up with findings which may be exploited in other fields. This paper provides a comprehensive overview of the major techniques for generating spatialized sound and, in addition, discusses perceptual and cross‐modal influences to consider. We also describe current limitations and provide an in‐depth look at the emerging topics in the field.
Vedad Hulusic, Carlo Harvey, Kurt Debattista, Nicolas Tsingos, Steve Walker, David M. Howard 0001, Alan Chalmers
Comput. Graph. Forum3
2012 Guest Editor's Introduction: Special Section on the Eurographics Symposium on Parallel Graphics and Visualization (EGPGV)
abstract
The articles in this special section contain selected papers from the Eurographics Symposium on Parallel Graphics and Visualization (EGPGV).
James P. Ahrens, Kurt Debattista
IEEE Trans. Vis. Comput. Graph.2
2011 Maintaining frame rate perception in interactive environments by exploiting audio-visual cross-modal interaction
Vedad Hulusic, Kurt Debattista, Vibhor Aggarwal, Alan Chalmers
Vis. Comput.2
2010 Preface and biographic notes for the special issue on graphics for serious games
Kurt Debattista, Alberto José Proença, Luís Paulo Santos
Comput. Graph.1
2010 Selective rendering for efficient ray traced stereoscopic images
Cheng-Hung Lo, Chih-Hsing Chu, Kurt Debattista, Alan Chalmers
Vis. Comput.3
2009 High Dynamic Range Imaging and Low Dynamic Range Expansion for Generating HDR Content
abstract
Abstract In the last few years, researchers in the field of High Dynamic Range (HDR) Imaging have focused on providing tools for expanding Low Dynamic Range (LDR) content for the generation of HDR images due to the growing popularity of HDR in applications, such as photography and rendering via Image‐Based Lighting, and the imminent arrival of HDR displays to the consumer market. LDR content expansion is required due to the lack of fast and reliable consumer level HDR capture for still images and videos. Furthermore, LDR content expansion, will allow the re‐use of legacy LDR stills, videos and LDR applications created, over the last century and more, to be widely available. The use of certain LDR expansion methods, those that are based on the inversion of Tone Mapping Operators (TMOs), has made it possible to create novel compression algorithms that tackle the problem of the size of HDR content storage, which remains one of the major obstacles to be overcome for the adoption of HDR. These methods are used in conjunction with traditional LDR compression methods and can evolve accordingly. The goal of this report is to provide a comprehensive overview on HDR Imaging, and an in depth review on these emerging topics.
Francesco Banterle, Kurt Debattista, Alessandro Artusi, Sumanta N. Pattanaik, Karol Myszkowski, Patrick Ledda, Alan Chalmers
Comput. Graph. Forum2
2009 A Psychophysical Evaluation of Inverse Tone Mapping Techniques
abstract
Abstract In recent years inverse tone mapping techniques have been proposed for enhancing low‐dynamic range (LDR) content for a high‐dynamic range (HDR) experience on HDR displays, and for image based lighting. In this paper, we present a psychophysical study to evaluate the performance of inverse (reverse) tone mapping algorithms. Some of these techniques are computationally expensive because they need to resolve quantization problems that can occur when expanding an LDR image. Even if they can be implemented efficiently on hardware, the computational cost can still be high. An alternative is to utilize less complex operators; although these may suffer in terms of accuracy. Our study investigates, firstly, if a high level of complexity is needed for inverse tone mapping and, secondly, if a correlation exists between image content and quality. Two main applications have been considered: visualization on an HDR monitor and image‐based lighting.
Francesco Banterle, Patrick Ledda, Kurt Debattista, Marina Bloj, Alessandro Artusi, Alan Chalmers
Comput. Graph. Forum3
2009 2009 Eurographics Symposium on Parallel Graphics and Visualization
abstract
In this paper, we propose an experimental study of an inexpensive off-the-shelf sort-last volume visualization architecture based upon multiple GPUs and a single CPU.We show how to efficiently make use of this architecture to achieve high performance sort-last volume visualization of large datasets.We analyze the bottlenecks of this architecture.We compare this architecture to a classical sort-last visualization system using a cluster of commodity machines interconnected by a gigabit Ethernet network.Based on extensive experiments, we show that this solution competes very well with a mid-sized PC cluster, while it significantly improves performance compared to a single standard PC.
João Luiz Dihl Comba, Daniel Weiskopf, Kurt Debattista
Comput. Graph. Forum3
2009 Instant Caching for Interactive Global Illumination
abstract
Abstract The ability to interactively render dynamic scenes with global illumination is one of the main challenges in computer graphics. The improvement in performance of interactive ray tracing brought about by significant advances in hardware and careful exploitation of coherence has rendered the potential of interactive global illumination a reality. However, the simulation of complex light transport phenomena, such as diffuse interreflections, is still quite costly to compute in real time. In this paper we present a caching scheme, termed Instant Caching, based on a combination of irradiance caching and instant radiosity. By reutilising calculations from neighbouring computations this results in a speedup over previous instant radiosity‐based approaches. Additionally, temporal coherence is exploited by identifying which computations have been invalidated due to geometric transformations and updating only those paths. The exploitation of spatial and temporal coherence allows us to achieve superior frame rates for interactive global illumination within dynamic scenes, without any precomputation or quality loss when compared to previous methods; handling of lighting and material changes are also demonstrated.
Kurt Debattista, Piotr Dubla, Francesco Banterle, Luís Paulo Santos, Alan Chalmers
Comput. Graph. Forum1
2009 Adaptive Interleaved Sampling for Interactive High-Fidelity Rendering
abstract
Abstract Recent advances have made interactive ray tracing (IRT) possible on consumer desktop machines. These advances have brought about the potential for interactive global illumination (IGI) with enhanced realism through physically based lighting. IGI, unlike IRT, has a much higher computational complexity. Furthermore, since non‐primary rays constitute the majority of the computation, the rays are predominantly incoherent, making impractical many of the methods that have made IRT possible. Two methods that have already shown promise in decreasing the computational time of the GI solution are interleaved sampling and adaptive rendering. Interleaved sampling is a generalized sampling scheme that smoothly blends between regular and irregular sampling while maintaining coherence. Adaptive rendering algorithms adjust rendering quality, non‐uniformally, using a guidance scheme. While adaptive rendering has shown to provide speed‐up when used for off‐line rendering it has not been utilized in IRT due to its naturally incoherent nature. In this paper, we combine adaptive rendering and interleaved sampling within a component‐based solution into a new approach we term adaptive interleaved sampling. This allows us to tailor new adaptive heuristics for interleaved sampling of the individual components of the GI solution significantly improving overall performance. We present a novel component‐based IGI framework for which we achieve interactive frame rates for a range of effects such as indirect diffuse lighting, soft shadows and single scatter homogeneous participating media.
Piotr Dubla, Kurt Debattista, Alan Chalmers
Comput. Graph. Forum2
2009 Towards high-fidelity multi-sensory virtual environments
Alan Chalmers, Kurt Debattista, Belma Ramic-Brkic
Vis. Comput.2
2008 A GPU-friendly method for high dynamic range texture compression using inverse tone mapping
Francesco Banterle, Kurt Debattista, Patrick Ledda, Alan Chalmers
Graphics Interface2
2007 A physically-based client-server rendering solution for mobile devices
abstract
Mobile devices, also known as small-form-factor (SFF) devices such as mobile phones, PDAs and ultra mobile PCs have continued to grow in popularity. Improvements in SFF hardware has enabled a range of suitable applications such as gaming, interactive visualisation and mobile mapping. Although high-fidelity graphic systems typically have significant computational requirements, the time taken may be largely resolution dependent. The limited resolution of SFFs indicates such platforms are prime candidates for running high-fidelity graphics.
Matt Aranha, Piotr Dubla, Kurt Debattista, Thomas Bashford-Rogers, Alan Chalmers
MUM3
2007 Parallel selective rendering of high-fidelity virtual environments
Kurt Debattista, Alan Chalmers, Richard Gillibrand, Peter William Longhurst, Georgia Mastoropoulou, Veronica Sundstedt
Parallel Comput.1
2007 A framework for inverse tone mapping
Francesco Banterle, Patrick Ledda, Kurt Debattista, Alan Chalmers, Marina Bloj
Vis. Comput.3