Carlo Harvey

dblp:33/9591 · DBLP profile ↗
← Back
15ranked-venue papers
3as first author
7since 2021 · last 2024
0000-0002-4809-1592ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 3 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 5 · 3 since 2021Artificial intelligence and machine learning · 1
YearPublicationVenuePosition
2024 Using a NEAT approach with curriculums for dynamic content generation in video games
abstract
Abstract This paper presents a novel exploration of the use of an evolving neural network approach to generate dynamic content for video games, specifically for a tower defence game. The objective is to employ the NeuroEvolution of Augmenting Topologies (NEAT) technique to train a NEAT neural network as a wave manager to generate enemy waves that challenge the player’s defences. The approach is extended to incorporate NEAT-generated curriculums for tower deployments to gradually increase the difficulty for the generated enemy waves, allowing the neural network to learn incrementally. The approach dynamically adapts to changes in the player’s skill level, providing a more personalised and engaging gaming experience. The quality of the machine-generated waves is evaluated through a blind A/B test with the Games Experience Questionnaire (GEQ), and results are compared with manually designed human waves. The study finds no discernible difference in the reported player experience between AI and human-designed waves. The approach can significantly reduce the time and resources required to design game content while maintaining the quality of the player experience. The approach has the potential to be applied to a range of video game genres and within the design and development process, providing a more personalised and engaging gaming experience for players.
Daniel Hind, Carlo Harvey
Pers. Ubiquitous Comput.2
2023 Online alignment of human motion using forward plotting-dynamic time warping
abstract
Abstract A number of approaches to online time warping have been proposed based on plotting an alignment path backwards from a selected end‐point. When continually updating a time warp to align a time series with a live data source, such as a stream of motion capture data, the use of backwards plotting makes it hard to maintain a monotonic constraint. To solve this problem, a number of time warping approaches based on forward plotting, referred to as FP‐DTW, are proposed and evaluated by applying them to human motions. We demonstrate that forward plotting for temporal alignment is a viable solution when backwards plotting is otherwise not possible.
Mathew Randall, Carlo Harvey, Ian Williams 0001
Comput. Animat. Virtual Worlds2
2023 Correlation as a measure for alignment and similarity of human motions
abstract
Abstract The ability to measure similarity and alignment of motions is a key tool in motion retrieval and motion editing. Similarity metrics based on distance functions are often utilized when measuring similarity of human motions, however, metrics based on correlation can also potentially useful for measuring similarity and alignment. This paper evaluates the use of correlation as a method of measuring the alignment and similarity of human motion and compares them against more established distance‐based metrics. Three correlation methods and five methods of parameterising rotation are evaluated. The results show that parameterization based on displacement vectors and Kendall Tau rank correlation are optimal for measuring the alignment between two motions. If measuring similarity of motions, however, an approach based on distance metrics for angular or positional distance should be used.
Mathew Randall, Carlo Harvey, Ian Williams 0001
Comput. Animat. Virtual Worlds2
2022 Acoustic Rendering Based on Geometry Reduction and Acoustic Material Classification
abstract
We present work in progress on a pipeline for audio rendering integrating vision-based systems for acoustic material classification. With a marching cubes algorithm, the pipeline estimates a cuboid acoustic volume encapsulating the listener, a sound source, and the surrounding environment. A variable-resolution binary field, samples and simplifies the input scene, and captures the appearance of surfaces to produce a set of image patches. A classifier infers acoustic materials, expressed as frequency-dependent acoustic absorption coefficients, from image patches. The estimated volume and aggregated acoustic materials provide input to the Image Source Model that models reverberation by generating Room Impulse Responses (RIRs). We conduct preliminary tests by applying our pipeline on a set of indoor and outdoor scenes, producing RIRs with inferred acoustic materials, comparing them against RIRs with manually assigned acoustic materials, extracting and evaluating objective metrics, such as reverberation or clarity. Through a learned metric on subjective responses, we compare perceptual aspects of automatically-generated RIRs against manually tagged. Objective and subjective analysis suggests that the pipeline can automate the acoustic material classification process by producing RIRs indistinguishable from manually-tagged counterparts.
Mattia Colombo, Alan Dolhasz, Jason Hockman, Carlo Harvey
CoG4
2021 Psychometric Mapping of Audio Features to Perceived Physical Characteristics of Virtual Objects
abstract
Physically-based sound synthesis can simulate virtual sound sources whose audio features reflect the physical characteristics of corresponding objects displayed in a virtual environment, allowing for real-time generation of content without relying on pre-existing audio samples. This, however, requires efficient control strategies for sound synthesis models that, depending on the nature of the sounding objects, require to be mapped to varying physical characteristics displayed through visual information. In this experiment, participants were asked to adjust a set of sound synthesis parameters based on varying physical characteristics of a virtual bouncing ball: distance, elasticity and radius. Statistical analysis of recorded subject responses shows that object radius influences evaluation of pitch and amplitude for the object's representation. Similarly, distance influences user evaluation of both reverb and amplitude whilst elasticity doesn't influence user evaluation of the feature distributions. This result is consistent across user groups evaluated: audio experts and naïve listeners. Models are produced that encode these observations using linear regression, enabling automatic parameterisation of this feature space for audio synthesis engines.
Mattia Colombo, Alan Dolhasz, Jason Hockman, Carlo Harvey
CoG4
2021 A robot arm digital twin utilising reinforcement learning
Marius Matulis, Carlo Harvey
Comput. Graph.2
2021 A comparison between expert and beginner learning for motor skill development in a virtual reality serious game
abstract
Abstract In order to be used for skill development and skill maintenance, virtual environments require accurate simulation of the physical phenomena involved in the process of the task being trained. The accuracy needs to be conveyed in a multimodal fashion with varying parameterisations still being quantified, and these are a function of task, prior knowledge, sensory efficacy and human perception. Virtual reality (VR) has been integrated from a didactic perspective in many serious games and shown to be effective in the pedological process. This paper interrogates whether didactic processes introduced into a VR serious game, by taking advantage of augmented virtuality to modify game attributes, can be effective for both beginners and experts to a task. The task in question is subjective performance in a clay pigeon shooting simulation. The investigation covers whether modified game attributes influence skill and learning in a complex motor task and also investigates whether this process is applicable to experts as well as beginners to the task. VR offers designers and developers of serious games the ability to provide information in the virtual world in a fashion that is impossible in the real world. This introduces the question of whether this is effective and transfers skill adoption into the real world and also if a-priori knowledge influences the practical nature of this information in the pedagogic process. Analysis is conducted via a between-subjects repeated measure ANOVA using a $$2 \times 2$$ 2 × 2 factorial design to address these questions. The results show that the different training provided affects the performance in this task ( $$N=57$$ N = 57 ). The skill improvement is still evidenced in repeated measures when information and guidance is removed. This effect does not exist under a control condition. Additionally, we separate by an expert and non-expert group to deduce if a-priori knowledge influences the effect of the presented information, it is shown that it does not.
Carlo Harvey, Elmedin Selmanovic, Jake O'Connor, Malek Chahin
Vis. Comput.1
2020 A Computer Vision Inspired Automatic Acoustic Material Tagging System for Virtual Environments
abstract
This paper presents the ongoing work on an approach to material information retrieval in virtual environments (VEs). Our approach uses convolutional neural networks to classify materials by performing semantic segmentation on images captured in the VE. Class maps obtained are then re-projected onto the environment. We use transfer learning and fine-tune a pretrained segmentation model on images captured in our VEs. The geometry and semantic information can then be used to create mappings between objects in the VE and acoustic absorption coefficients. This can then be input for physically-based audio renderers, allowing a significant reduction in manual material tagging.
Mattia Colombo, Alan Dolhasz, Carlo Harvey
CoG3
2020 Learning to Observe: Approximating Human Perceptual Thresholds for Detection of Suprathreshold Image Transformations
abstract
Many tasks in computer vision are often calibrated and evaluated relative to human perception. In this paper, we propose to directly approximate the perceptual function performed by human observers completing a visual detection task. Specifically, we present a novel methodology for learning to detect image transformations visible to human observers through approximating perceptual thresholds. To do this, we carry out a subjective two-alternative forced-choice study to estimate perceptual thresholds of human observers detecting local exposure shifts in images. We then leverage transformation equivariant representation learning to overcome issues of limited perceptual data. This representation is then used to train a dense convolutional classifier capable of detecting local suprathreshold exposure shifts - a distortion common to image composites. In this context, our model can approximate perceptual thresholds with an average error of 0.1148 exposure stops between empirical and predicted thresholds. It can also be trained to detect a range of different local transformations.
Alan Dolhasz, Carlo Harvey, Ian Williams 0001
CVPR2
2019 Audio-Visual-Olfactory Resource Allocation for Tri-modal Virtual Environments
abstract
Virtual Environments (VEs) provide the opportunity to simulate a wide range of applications, from training to entertainment, in a safe and controlled manner. For applications which require realistic representations of real world environments, the VEs need to provide multiple, physically accurate sensory stimuli. However, simulating all the senses that comprise the human sensory system (HSS) is a task that requires significant computational resources. Since it is intractable to deliver all senses at the highest quality, we propose a resource distribution scheme in order to achieve an optimal perceptual experience within the given computational budgets. This paper investigates resource balancing for multi-modal scenarios composed of aural, visual and olfactory stimuli. Three experimental studies were conducted. The first experiment identified perceptual boundaries for olfactory computation. In the second experiment, participants ( N=25) were asked, across a fixed number of budgets ( M=5), to identify what they perceived to be the best visual, acoustic and olfactory stimulus quality for a given computational budget. Results demonstrate that participants tend to prioritize visual quality compared to other sensory stimuli. However, as the budget size is increased, users prefer a balanced distribution of resources with an increased preference for having smell impulses in the VE. Based on the collected data, a quality prediction model is proposed and its accuracy is validated against previously unused budgets and an untested scenario in a third and final experiment.
Efstratios Doukakis, Kurt Debattista, Thomas Bashford-Rogers, Amar Dhokia, Ali Asadipour 0001, Alan Chalmers, Carlo Harvey
IEEE Trans. Vis. Comput. Graph.7
2018 Audiovisual Resource Allocation for Bimodal Virtual Environments
abstract
Abstract Fidelity is of key importance if virtual environments are to be used as authentic representations of real environments. However, simulating the multitude of senses that comprise the human sensory system is computationally challenging. With limited computational resources, it is essential to distribute these carefully in order to simulate the most ideal perceptual experience. This paper investigates this balance of resources across multiple scenarios where combined audiovisual stimulation is delivered to the user. A subjective experiment was undertaken where participants (N=35) allocated five fixed resource budgets across graphics and acoustic stimuli. In the experiment, increasing the quality of one of the stimuli decreased the quality of the other. Findings demonstrate that participants allocate more resources to graphics; however, as the computational budget is increased, an approximately balanced distribution of resources is preferred between graphics and acoustics. Based on the results, an audiovisual quality prediction model is proposed and successfully validated against previously untested budgets and an untested scenario.
Efstratios Doukakis, Kurt Debattista, Carlo Harvey, Thomas Bashford-Rogers, Alan Chalmers
Comput. Graph. Forum3
2018 Olfaction and Selective Rendering
abstract
Abstract Accurate simulation of all the senses in virtual environments is a computationally expensive task. Visual saliency models have been used to improve computational performance for rendered content, but this is insufficient for multi‐modal environments. This paper considers cross‐modal perception and, in particular, if and how olfaction affects visual attention. Two experiments are presented in this paper. Firstly, eye tracking is gathered from a number of participants to gain an impression about where and how they view virtual objects when smell is introduced compared to an odourless condition. Based on the results of this experiment a new type of saliency map in a selective‐rendering pipeline is presented. A second experiment validates this approach, and demonstrates that participants rank images as better quality, when compared to a reference, for the same rendering budget.
Carlo Harvey, Thomas Bashford-Rogers, Kurt Debattista, Efstratios Doukakis, Alan Chalmers
Comput. Graph. Forum1
2018 Subjective Evaluation of High-Fidelity Virtual Environments for Driving Simulations
abstract
Virtual environments (VEs) grant the ability to experience real-world scenarios, such as driving, in a virtual, safe, and reproducible context. However, in order to achieve their full potential, the fidelity of the VE must provide confidence that it replicates the perception of the real-world experience. The computational cost of simulating real-world visuals accurately means that compromises to the fidelity of the visuals must be made. In this paper, a subjective evaluation of driving in a VE at different quality settings is presented. Participants (n = 44) were driven around in the real world and in a purposely built representative VE and the fidelity of the graphics and overall experience at low-, medium-, and high-visual settings were analyzed. Low quality corresponds to the illumination in many current traditional simulators, medium to a higher quality using accurate shadows and reflections, and high to the quality experienced in modern movies and simulations that require hours of computation. Results demonstrate that graphics quality affects the perceived fidelity of the visuals and the overall experience. When judging the overall experience, participants could tell the difference between the lower quality graphics and the rest but did not significantly discriminate between the medium and higher graphical settings. This indicates that future driving simulators should improve the quality, but once the equivalent of the presented medium quality is reached, they may not need to do so significantly.
Kurt Debattista, Thomas Bashford-Rogers, Carlo Harvey, Brian Waterfield, Alan Chalmers
IEEE Trans. Hum. Mach. Syst.3
2017 Multi-Modal Perception for Selective Rendering
abstract
Abstract A major challenge in generating high‐fidelity virtual environments (VEs) is to be able to provide realism at interactive rates. The high‐fidelity simulation of light and sound is still unachievable in real time as such physical accuracy is very computationally demanding. Only recently has visual perception been used in high‐fidelity rendering to improve performance by a series of novel exploitations; to render parts of the scene that are not currently being attended to by the viewer at a much lower quality without the difference being perceived. This paper investigates the effect spatialized directional sound has on the visual attention of a user towards rendered images. These perceptual artefacts are utilized in selective rendering pipelines via the use of multi‐modal maps. The multi‐modal maps are tested through psychophysical experiments to examine their applicability to selective rendering algorithms, with a series of fixed cost rendering functions, and are found to perform significantly better than only using image saliency maps that are naively applied to multi‐modal VEs.
Carlo Harvey, Kurt Debattista, Thomas Bashford-Rogers, Alan Chalmers
Comput. Graph. Forum1
2012 Acoustic Rendering and Auditory-Visual Cross-Modal Perception and Interaction
abstract
Abstract In recent years research in the three‐dimensional sound generation field has been primarily focussed upon new applications of spatialized sound. In the computer graphics community the use of such techniques is most commonly found being applied to virtual, immersive environments. However, the field is more varied and diverse than this and other research tackles the problem in a more complete, and computationally expensive manner. Furthermore, the simulation of light and sound wave propagation is still unachievable at a physically accurate spatio‐temporal quality in real time. Although the Human Visual System (HVS) and the Human Auditory System (HAS) are exceptionally sophisticated, they also contain certain perceptional and attentional limitations. Researchers, in fields such as psychology, have been investigating these limitations for several years and have come up with findings which may be exploited in other fields. This paper provides a comprehensive overview of the major techniques for generating spatialized sound and, in addition, discusses perceptual and cross‐modal influences to consider. We also describe current limitations and provide an in‐depth look at the emerging topics in the field.
Vedad Hulusic, Carlo Harvey, Kurt Debattista, Nicolas Tsingos, Steve Walker, David M. Howard 0001, Alan Chalmers
Comput. Graph. Forum2