George Alex Koulieris

dblp:137/5755 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
7since 2021 · last 2026
0000-0003-1610-6240ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 On the design fundamentals of diffusion models: A survey
abstract
Diffusion models are learning pattern-learning systems to model and sample from data distributions with three functional components namely the forward process, the reverse process, and the sampling process. The components of diffusion models have gained significant attention with many design factors being considered in common practice. Existing reviews have primarily focused on higher-level solutions, covering less on the design fundamentals of components. This study seeks to address this gap by providing a comprehensive and coherent review of seminal designable factors within each functional component of diffusion models. This provides a finer-grained perspective of diffusion models, benefiting future studies in the analysis of individual components, the design factors for different purposes, and the implementation of diffusion models. • A literature review of design fundamentals of diffusion models. • Identified the three functional components of diffusion models. • Examined the key factors and their functionalities of the components. • Discussed the popular designs and seminal works for each factor.
Ziyi Chang, George Alex Koulieris, Hyung Jin Chang, Hubert P. H. Shum
Pattern Recognit.2
2023 An Inverted Pyramid Acceleration Structure Guiding Foveated Sphere Tracing for Implicit Surfaces in VR
Andreas Polychronakis, George Alex Koulieris, Katerina Mania
EGSR (ST)2
2023 The Statistics of Eye Movements and Binocular Disparities during VR Gaming: Implications for Headset Design
abstract
The human visual system evolved in environments with statistical regularities. Binocular vision is adapted to these such that depth perception and eye movements are more precise, faster, and performed comfortably in environments consistent with the regularities. We measured the statistics of eye movements and binocular disparities in virtual-reality (VR) - gaming environments and found that they are quite different from those in the natural environment. Fixation distance and direction are more restricted in VR, and fixation distance is farther. The pattern of disparity across the visual field is less regular in VR and does not conform to a prominent property of naturally occurring disparities. From this we predict that double vision is more likely in VR than in the natural environment. We also determined the optimal screen distance to minimize discomfort due to the vergence-accommodation conflict, and the optimal nasal-temporal positioning of head-mounted display (HMD) screens to maximize binocular field of view. Finally, in a user study we investigated how VR content affects comfort and performance. Content that is more consistent with the statistics of the natural world yields less discomfort than content that is not. Furthermore, consistent content yields slightly better performance than inconsistent content.
Avi M. Aizenman, George Alex Koulieris, Agostino Gibaldi, Vibhor Sehgal, Dennis M. Levi, Martin S. Banks
ACM Trans. Graph.2
2023 OrthopedVR: clinical assessment and pre-operative planning of paediatric patients with lower limb rotational abnormalities in virtual reality
abstract
Abstract Rotational abnormalities in the lower limbs causing patellar mal-tracking negatively affect patients’ lives, particularly young patients (10–17 years old). Recent studies suggest that rotational abnormalities can increase degenerative effects on the joints of the lower limbs. Rotational abnormalities are diagnosed using 2D CT imaging and X-rays, and these data are then used by surgeons to make decisions during an operation. However, 3D representation of data is preferable in the examination of 3D structures, such as bones. This correlates with added benefits for medical judgement, pre-operative planning, and clinical training. Virtual reality can enable the transformation of standard clinical imaging examination methods (CT/MRI) into immersive examinations and pre-operative planning in 3D. We present a VR system (OrthopedVR) which allows orthopaedic surgeons to examine patients’ specific anatomy of the lower limbs in an immersive three-dimensional environment and to simulate the effect of potential surgical interventions such as corrective osteotomies in VR. In OrthopedVR, surgeons can perform corrective incisions and re-align segments into desired rotational angles. From the system evaluation performed by experienced surgeons we found that OrthopedVR provides a better understanding of lower limb alignment and rotational profiles in comparison with isolated 2D CT scans. In addition, it was demonstrated that using VR software improves pre-operative planning, surgical precision and post-operative outcomes for patients. Our study results indicate that our system can become a stepping stone into simulating corrective surgeries of the lower limbs, and suggest future improvements which will help adopt VR surgical planning into the clinical orthopaedic practice.
David Sibrina, Sarath Bethapudi, George Alex Koulieris
Vis. Comput.3
2022 3D Reconstruction of Sculptures from Single Images via Unsupervised Domain Adaptation on Implicit Models
abstract
Acquiring the virtual equivalent of exhibits, such as sculptures, in virtual reality (VR) museums, can be labour-intensive and sometimes infeasible. Deep learning based 3D reconstruction approaches allow us to recover 3D shapes from 2D observations, among which single-view-based approaches can reduce the need for human intervention and specialised equipment in acquiring 3D sculptures for VR museums. However, there exist two challenges when attempting to use the well-researched human reconstruction methods: limited data availability and domain shift. Considering sculptures are usually related to humans, we propose our unsupervised 3D domain adaptation method for adapting a single-view 3D implicit reconstruction model from the source (real-world humans) to the target (sculptures) domain. We have compared the generated shapes with other methods and conducted ablation studies as well as a user study to demonstrate the effectiveness of our adaptation method. We also deploy our results in a VR application.
Ziyi Chang, George Alex Koulieris, Hubert P. H. Shum
VRST2
2021 Emulating Foveated Path Tracing
abstract
At full resolution, path tracing cannot be deployed in real-time based on current graphics hardware due to slow convergence times and noisy outputs, despite recent advances in denoisers. In this work, we develop a perceptual sandbox emulating a foveated path tracer to determine the eccentricity angle thresholds that enable imperceptible foveated path tracing. In a foveated path tracer the number of rays fired can be decreased, and thus performance can be increased. For this study, due to current hardware limitations prohibiting real-time path-tracing for multiple samples-per-pixel, we pre-render image buffers and emulate foveated rendering as a post-process by selectively blending the pre-rendered content, driven by an eye tracker capturing eye motion. We then perform three experiments to estimate conservative thresholds of eccentricity boundaries for which image manipulations are imperceptible. Contrary to our expectation of a single threshold across the three experiments, our results indicated three different average thresholds, one for each experiment. We hypothesise that this is due to the dissimilarity of the methodologies, i.e., A-B testing vs sequential presentation vs custom adjustment of eccentricities affecting the perceptibility of peripheral blur among others. We estimate, for the first time for path tracing, specific thresholds of eccentricity that limit any perceptual repercussions whilst maintaining high performance. We perform an analysis to determine potential computational complexity reductions due to foveation in path tracing. Our analysis shows a significant boost in path-tracing performance (≥ 2x − 3x) using our foveated rendering method as a result of the reduction in the primary rays.
Andreas Polychronakis, George Alex Koulieris, Katerina Mania
MIG2
2021 Eye Tracking Interaction on Unmodified Mobile VR Headsets Using the Selfie Camera
abstract
Input methods for interaction in smartphone-based virtual and mixed reality (VR/MR) are currently based on uncomfortable head tracking controlling a pointer on the screen. User fixations are a fast and natural input method for VR/MR interaction. Previously, eye tracking in mobile VR suffered from low accuracy, long processing time, and the need for hardware add-ons such as anti-reflective lens coating and infrared emitters. We present an innovative mobile VR eye tracking methodology utilizing only the eye images from the front-facing (selfie) camera through the headset’s lens, without any modifications. Our system first enhances the low-contrast, poorly lit eye images by applying a pipeline of customised low-level image enhancements suppressing obtrusive lens reflections. We then propose an iris region-of-interest detection algorithm that is run only once. This increases the iris tracking speed by reducing the iris search space in mobile devices. We iteratively fit a customised geometric model to the iris to refine its coordinates. We display a thin bezel of light at the top edge of the screen for constant illumination. A confidence metric calculates the probability of successful iris detection. Calibration and linear gaze mapping between the estimated iris centroid and physical pixels on the screen results in low latency, real-time iris tracking. A formal study confirmed that our system’s accuracy is similar to eye trackers in commercial VR headsets in the central part of the headset’s field-of-view. In a VR game, gaze-driven user completion time was as fast as with head-tracked interaction, without the need for consecutive head motions. In a VR panorama viewer, users could successfully switch between panoramas using gaze.
Panagiotis Drakopoulos, George Alex Koulieris, Katerina Mania
ACM Trans. Appl. Percept.2
2019 Near-Eye Display and Tracking Technologies for Virtual and Augmented Reality
abstract
Abstract Virtual and augmented reality (VR/AR) are expected to revolutionise entertainment, healthcare, communication and the manufacturing industries among many others. Near‐eye displays are an enabling vessel for VR/AR applications, which have to tackle many challenges related to ergonomics, comfort, visual quality and natural interaction. These challenges are related to the core elements of these near‐eye display hardware and tracking technologies. In this state‐of‐the‐art report, we investigate the background theory of perception and vision as well as the latest advancements in display engineering and tracking technologies. We begin our discussion by describing the basics of light and image formation. Later, we recount principles of visual perception by relating to the human visual system. We provide two structured overviews on state‐of‐the‐art near‐eye display and tracking technologies involved in such near‐eye displays. We conclude by outlining unresolved research questions to inspire the next generation of researchers.
George Alex Koulieris, Kaan Aksit, Michael Stengel, Rafal Mantiuk, Katerina Mania, Christian Richardt
Comput. Graph. Forum1
2019 DiCE: dichoptic contrast enhancement for VR and stereo displays
abstract
In stereoscopic displays, such as those used in VR/AR headsets, our eyes are presented with two different views. The disparity between the views is typically used to convey depth cues, but it could be also used to enhance image appearance. We devise a novel technique that takes advantage of binocular fusion to boost perceived local contrast and visual quality of images. Since the technique is based on fixed tone curves, it has negligible computational cost and it is well suited for real-time applications, such as VR rendering. To control the trade-off between contrast gain and binocular rivalry, we conduct a series of experiments to explain the factors that dominate rivalry perception in a dichoptic presentation where two images of different contrasts are displayed. With this new finding, we can effectively enhance contrast and control rivalry in mono- and stereoscopic images, and in VR rendering, as confirmed in validation experiments.
Fangcheng Zhong, George Alex Koulieris, George Drettakis, Martin S. Banks, Mathieu Chambe, Frédo Durand, Rafal Mantiuk
ACM Trans. Graph.2
2017 Accommodation and comfort in head-mounted displays
abstract
Head-mounted displays (HMDs) often cause discomfort and even nausea. Improving comfort is therefore one of the most significant challenges for the design of such systems. In this paper, we evaluate the effect of different HMD display configurations on discomfort. We do this by designing a device to measure human visual behavior and evaluate viewer comfort. In particular, we focus on one known source of discomfort: the vergence-accommodation (VA) conflict. The VA conflict is the difference between accommodative and vergence response. In HMDs the eyes accommodate to a fixed screen distance while they converge to the simulated distance of the object of interest, requiring the viewer to undo the neural coupling between the two responses. Several methods have been proposed to alleviate the VA conflict, including Depth-of-Field (DoF) rendering, focus-adjustable lenses, and monovision. However, no previous work has investigated whether these solutions actually drive accommodation to the distance of the simulated object. If they did, the VA conflict would disappear, and we expect comfort to improve. We design the first device that allows us to measure accommodation in HMDs, and we use it to obtain accommodation measurements and to conduct a discomfort study. The results of the first experiment demonstrate that only the focus-adjustable-lens design drives accommodation effectively, while other solutions do not drive accommodation to the simulated distance and thus do not resolve the VA conflict. The second experiment measures discomfort. The results validate that the focus-adjustable-lens design improves comfort significantly more than the other solutions.
George Alex Koulieris, Bee Bui, Martin S. Banks, George Drettakis
ACM Trans. Graph.1
2016 Gaze prediction using machine learning for dynamic stereo manipulation in games
abstract
Comfortable, high-quality 3D stereo viewing is becoming a requirement for interactive applications today. Previous research shows that manipulating disparity can alleviate some of the discomfort caused by 3D stereo, but it is best to do this locally, around the object the user is gazing at. The main challenge is thus to develop a gaze predictor in the demanding context of real-time, heavily task-oriented applications such as games. Our key observation is that player actions are highly correlated with the present state of a game, encoded by game variables. Based on this, we train a classifier to learn these correlations using an eye-tracker which provides the ground-truth object being looked at. The classifier is used at runtime to predict object category - and thus gaze - during game play, based on the current state of game variables. We use this prediction to propose a dynamic disparity manipulation method, which provides rich and comfortable depth. We evaluate the quality of our gaze predictor numerically and experimentally, showing that it predicts gaze more accurately than previous approaches. A subjective rating study demonstrates that our localized disparity manipulation is preferred over previous methods.
George Alex Koulieris, George Drettakis, Douglas W. Cunningham, Katerina Mania
VR1
2014 C-LOD: Context-aware Material Level-of-Detail applied to Mobile Graphics
abstract
Abstract Attention‐based Level‐Of–Detail (LOD) managers downgrade the quality of areas that are expected to go unnoticed by an observer to economize on computational resources. The perceptibility of lowered visual fidelity is determined by the accuracy of the attention model that assigns quality levels. Most previous attention based LOD managers do not take into account saliency provoked by context, failing to provide consistently accurate attention predictions. In this work, we extend a recent high level saliency model with four additional components yielding more accurate predictions: an object‐intrinsic factor accounting for canonical form of objects, an object‐context factor for contextual isolation of objects, a feature uniqueness term that accounts for the number of salient features in an image, and a temporal context that generates recurring fixations for objects inconsistent with the context. We conduct a perceptual experiment to acquire the weighting factors to initialize our model. We design C‐LOD, a LOD manager that maintains a constant frame rate on mobile devices by dynamically re‐adjusting material quality on secondary visual features of non‐attended objects. In a proof of concept study we establish that by incorporating C‐LOD, complex effects such as parallax occlusion mapping usually omitted in mobile devices can now be employed, without overloading GPU capability and, at the same time, conserving battery power.
George Alex Koulieris, George Drettakis, Douglas W. Cunningham, Katerina Mania
Comput. Graph. Forum1
2014 An Automated High-Level Saliency Predictor for Smart Game Balancing
abstract
Successfully predicting visual attention can significantly improve many aspects of computer graphics: scene design, interactivity and rendering. Most previous attention models are mainly based on low-level image features, and fail to take into account high-level factors such as scene context, topology, or task. Low-level saliency has previously been combined with task maps, but only for predetermined tasks. Thus, the application of these methods to graphics (e.g., for selective rendering) has not achieved its full potential. In this article, we present the first automated high-level saliency predictor incorporating two hypotheses from perception and cognitive science that can be adapted to different tasks. The first states that a scene is comprised of objects expected to be found in a specific context as well objects out of context which are salient (scene schemata) while the other claims that viewer’s attention is captured by isolated objects (singletons). We propose a new model of attention by extending Eckstein’s Differential Weighting Model. We conducted a formal eye-tracking experiment which confirmed that object saliency guides attention to specific objects in a game scene and determined appropriate parameters for a model. We present a GPU-based system architecture that estimates the probabilities of objects to be attended in real- time. We embedded this tool in a game level editor to automatically adjust game level difficulty based on object saliency, offering a novel way to facilitate game design. We perform a study confirming that game level completion time depends on object topology as predicted by our system.
George Alex Koulieris, George Drettakis, Douglas W. Cunningham, Katerina Mania
ACM Trans. Appl. Percept.1