Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Alexander Fix

dblp:70/10771 · DBLP profile ↗
← Back
15ranked-venue papers
6as first author
5since 2021 · last 2026
0009-0002-6163-2354ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-author · 4 since 2021Artificial intelligence and machine learning · 6 · 6 first-authorHuman-computer interaction and ubiquitous computing · 3 · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
3 papers
Rendering · 50% Virtual and augmented reality · 28% Image and video processing · 13%
Artificial intelligence
6 papers
Probabilistic and Bayesian machine learning · 54% Optimization for machine learning · 21% Segmentation and scene understanding · 11%
Theoretical computer science
3 papers
Mathematical optimization · 89% Graph algorithms and graph theory · 11%

Topics — the 22 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Virtual and augmented reality › immersive display
head-mounted display
0.622018
DeepFocus: learned image synthesis for computational displays · ACM Trans. Graph. 2018
Focal surface displays · ACM Trans. Graph. 2017
Mathematical optimization
discrete optimization
0.532014
A Primal-Dual Algorithm for Higher-Order Multilabel Markov Random Fields · CVPR 2014
Structured Learning of Sum-of-Submodular Higher Order Energy Functions · ICCV 2013
A graph cut algorithm for higher-order Markov Random Fields · ICCV 2011
Rendering › antialiasing › supersampling
neural supersampling
0.412020
Neural supersampling for real-time rendering · ACM Trans. Graph. 2020
Rendering
real-time rendering
0.412020
Neural supersampling for real-time rendering · ACM Trans. Graph. 2020
Image and video processing
super-resolution
0.412020
Neural supersampling for real-time rendering · ACM Trans. Graph. 2020
Rendering › sampling and reconstruction
temporal upsampling
0.412020
Neural supersampling for real-time rendering · ACM Trans. Graph. 2020
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models › markov random field
higher-order markov random fields
0.432015
A Hypergraph-Based Reduction for Higher-Order Binary Markov Random Fields · IEEE Trans. Pattern Anal. Mach. Intell. 2015
A graph cut algorithm for higher-order Markov Random Fields · ICCV 2011
Structured Learning of Sum-of-Submodular Higher Order Energy Functions · ICCV 2013
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models
0.432014
Duality and the Continuous Graphical Model · ECCV (3) 2014
A graph cut algorithm for higher-order Markov Random Fields · ICCV 2011
Structured Learning of Sum-of-Submodular Higher Order Energy Functions · ICCV 2013
Computational photography and imaging
light field display
0.312017
Focal surface displays · ACM Trans. Graph. 2017
Virtual and augmented reality › near-eye display
multifocal display
0.312017
Focal surface displays · ACM Trans. Graph. 2017
Machine learning › Optimization for machine learning
energy minimization
0.212015
A Hypergraph-Based Reduction for Higher-Order Binary Markov Random Fields · IEEE Trans. Pattern Anal. Mach. Intell. 2015
Machine learning › Optimization for machine learning › energy minimization
graph cuts
0.212015
A Hypergraph-Based Reduction for Higher-Order Binary Markov Random Fields · IEEE Trans. Pattern Anal. Mach. Intell. 2015
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
markov random field
0.212015
A Hypergraph-Based Reduction for Higher-Order Binary Markov Random Fields · IEEE Trans. Pattern Anal. Mach. Intell. 2015
Mathematical optimization › discrete optimization
energy minimization
0.212014
A Primal-Dual Algorithm for Higher-Order Multilabel Markov Random Fields · CVPR 2014
Mathematical optimization
primal-dual method
0.212014
A Primal-Dual Algorithm for Higher-Order Multilabel Markov Random Fields · CVPR 2014
Natural language and speech › Information extraction and text analysis › sequence labeling
binary sequence labeling
0.212013
Structured Learning of Sum-of-Submodular Higher Order Energy Functions · ICCV 2013
Computer vision › Segmentation and scene understanding
image segmentation
0.212013
Structured Learning of Sum-of-Submodular Higher Order Energy Functions · ICCV 2013
Mathematical optimization › discrete optimization
submodular function minimization
0.212013
Structured Learning of Sum-of-Submodular Higher Order Energy Functions · ICCV 2013
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
markov random field inference
0.112012
Approximate MRF Inference Using Bounded Treewidth Subgraphs · ECCV (1) 2012
Graph algorithms and graph theory
graph cut
0.112011
A graph cut algorithm for higher-order Markov Random Fields · ICCV 2011
Computer vision › 3D vision
stereo vision
0.112015
A Hypergraph-Based Reduction for Higher-Order Binary Markov Random Fields · IEEE Trans. Pattern Anal. Mach. Intell. 2015
Computer vision › 3D vision › stereo vision
stereo matching
0.112014
A Primal-Dual Algorithm for Higher-Order Multilabel Markov Random Fields · CVPR 2014

Methods — techniques the papers use, named apart from their topics

convolutional neural network · 0.8max-flow · 0.7graph cuts · 0.6pseudo-boolean optimization · 0.5temporal network · 0.4fusion moves · 0.4alpha-expansion · 0.4submodular flow · 0.3structural SVM · 0.3phase-only spatial light modulator · 0.3optimized blending · 0.3focal stack decomposition · 0.3graph cut transformation · 0.2QPBO · 0.2cutting-plane algorithm · 0.2cutting plane algorithm · 0.2submodular minimization · 0.1
YearPublicationVenuePosition
2026 Digitally Prototype Your Eye Tracker: Simulating Hardware Performance using 3D Synthetic Data
abstract
Eye tracking (ET) is a key enabler for Augmented and Virtual Reality (AR/VR). Prototyping new ET hardware requires assessing the impact of hardware choices on eye tracking performance. This task is compounded by the high cost of obtaining data from sufficiently many variations of real hardware, especially for machine learning, which requires large training datasets. We propose a method for end-to-end evaluation of how hardware changes impact machine learning-based ET performance using only synthetic data. We utilize a dataset of real 3D eyes, reconstructed from light dome data using neural radiance fields (NeRF), to synthesize captured eyes from novel viewpoints and camera parameters. Using this framework, we demonstrate that we can predict the relative performance across various hardware configurations, accounting for variations in sensor noise, illumination brightness, and optical blur. We also compare our simulator with the publicly available eye tracking dataset from the Project Aria glasses, demonstrating a strong correlation with real-world performance. Finally, we present a first-of-its-kind analysis in which we vary ET camera positions, evaluating ET performance ranging from on-axis direct views of the eye to peripheral views on the frame. Such an analysis would have previously required manufacturing physical devices to capture evaluation data. In short, our method enables faster prototyping of ET hardware.
Esther Y. H. Lin, Yimin Ding, Jogendra Kundu, Yatong An, Mohamed T. El-Haddad, Alexander Fix
ETRA6
2026 Enhancing Eye Feature Estimation from Event Data Streams through Adaptive Inference State Space Modeling
abstract
Eye feature extraction from event-based data streams can be performed efficiently and with low energy consumption, offering great utility to real-world eye tracking pipelines. However, few eye feature extractors are designed to handle sudden changes in event density caused by the changes between gaze behaviors that vary in their kinematics, leading to degraded prediction performance. In this work, we address this problem by introducing the adaptive inference state space model (AISSM), a novel architecture for feature extraction that is capable of dynamically adjusting the relative weight placed on current versus recent information. This relative weighting is determined via estimates of the signal-to-noise ratio and event density produced by a complementary dynamic confidence network. Lastly, we craft and evaluate a novel learning technique that improves training efficiency. Experimental results demonstrate that the AISSM system outperforms state-of-the-art models for event-based eye feature extraction.
Viet Dung Nguyen, Mobina Ghorbaninejad, Chengyi Ma, Reynold J. Bailey, Gabriel J. Diaz, Alexander Fix, Ryan J. Suess, Alexander Ororbia
ETRA6
2026 Reconstructing Realistic and Relightable Eyes
abstract
Accurately modeling the eye is a challenging task as it exhibits refraction and reflection at the cornea, complex iris texture, and self-occlusion and shadowing due to eyelids and eyelashes. To address these challenges, we present a system for learning a hybrid relightable eye model which can be relit under near-field point lights. Our hybrid model leverages an eyeball mesh for explicitly representing the cornea surface and the reflections on it while learning the geometry and light transport of the periocular region and eye interior implicitly. To account for refraction, we explicitly handle the refraction of camera rays using Snell’s law and predict the refraction of incident light rays using a neural network. Furthermore, we propose an extension of our method which enables us to relight the eye using a fringe projector to simulate structured light. Through experiments, we demonstrate that our method results in higher fidelity rendering under novel viewpoint and lighting conditions, improves learned iris geometry, and more accurately simulates structured light fringe patterns on the eye.
Wesley Khademi, Jogendra Kundu, Yatong An, Alexander Fix, David Colmenares
WACV4
2026 Polarization-Based Eye Tracking with Personalized Siamese Architecture ETRA011
abstract
Head-mounted devices integrated with eye tracking promise a solution for natural human-computer interaction. However, they typically require per-user calibration for optimal performance due to inter-person variability. A differential personalization approach using Siamese architectures learns relative gaze displacements and reconstructs absolute gaze from a small set of calibration frames. In this paper, we benchmark Siamese personalization on polarization-enabled eye tracking. For benchmarking, we use a 338-subject dataset captured with a polarization-sensitive camera and 850 nm illumination. We achieve performance comparable to linear calibration with 10-fold fewer samples. Using polarization inputs for Siamese personalization reduces gaze error by up to 12% compared to near-infrared (NIR)-based inputs. Combining Siamese personalization with linear calibration yields further improvements of up to 13% over a linearly calibrated baseline. These results establish Siamese personalization as a practical approach enabling accurate eye tracking.
Mantas Zurauskas, Alexander Fix, Beyza Kalkanli, Dmitri Model, Tom Bu, Mahsa Shakeri, Dave Stronks
Proc. ACM Hum. Comput. Interact.2
2022 Event-Based Kilohertz Eye Tracking using Coded Differential Lighting
abstract
Pixels in an event camera operate asynchronously and independently, reporting changes in intensity as events - tuples of (x,y) position, polarity s and timestamp t at microsecond resolution. Event cameras operate at low power (≈ 5mW) and respond to changes in the scene with a latency on the order of microseconds. These properties make event cameras an exciting candidate for eye tracking sensors on mobile platforms such as Augmented/Virtual Reality (AR/VR) headsets, since these systems have hard real-time and power constraints. One proven method for eye tracking and gaze estimation is corneal glint detection. We exploit the fact that corneal glint tracking only requires a sparse set of pixels in the image, by making use of the natural sparsity of event cameras, which only detect changes in the scene. To enhance this effect, we design an illumination scheme, Coded Differential Lighting, which enhances specular reflections, suppresses all other events, and solves the light-to-glint correspondence. This is the first purely event-based corneal glint detection and tracking algorithm, which operates on standard hardware at kHz sampling rate.
Timo Stoffregen, Hossein Daraei, Clare Robinson, Alexander Fix
WACV4
2020 Raycast Calibration for Augmented Reality HMDs with Off-Axis Reflective Combiners
abstract
Augmented reality overlays virtual objects on the real world. To do so, the head mounted display (HMD) needs to be calibrated to establish a mapping between 3D points in the real world with 2D pixels on display panels. This distortion is a high-dimensional function that also depends on pupil position and varifocal settings. We present Raycast calibration, an efficient approach to geometrically calibrate AR displays with off-axis reflective combiners. Our approach requires a small amount of data to estimate a compact, physics-based, and ray-traceable model of the HMD optics. We apply this technique to automatically calibrate an AR prototype with display, SLAM and eye-tracker, without user in the loop.
Qi Guo 0009, Huixuan Tang, Aaron Schmitz, Yang Lou, Alexander Fix, Steven Lovegrove, Hauke Strasdat
ICCP6
2020 Neural supersampling for real-time rendering
abstract
Due to higher resolutions and refresh rates, as well as more photorealistic effects, real-time rendering has become increasingly challenging for video games and emerging virtual reality headsets. To meet this demand, modern graphics hardware and game engines often reduce the computational cost by rendering at a lower resolution and then upsampling to the native resolution. Following the recent advances in image and video superresolution in computer vision, we propose a machine learning approach that is specifically tailored for high-quality upsampling of rendered content in real-time applications. The main insight of our work is that in rendered content, the image pixels are point-sampled, but precise temporal dynamics are available. Our method combines this specific information that is typically available in modern renderers (i.e., depth and dense motion vectors) with a novel temporal network design that takes into account such specifics and is aimed at maximizing video quality while delivering real-time performance. By training on a large synthetic dataset rendered from multiple 3D scenes with recorded camera motion, we demonstrate high fidelity and temporally stable results in real-time, even in the highly challenging 4 × 4 upsampling scenario, significantly outperforming existing superresolution and temporal antialiasing work.
Lei Xiao 0014, Salah Nouri, Matthew Chapman, Alexander Fix, Douglas Lanman, Anton Kaplanyan
ACM Trans. Graph.4
2018 DeepFocus: learned image synthesis for computational displays
abstract
Addressing vergence-accommodation conflict in head-mounted displays (HMDs) requires resolving two interrelated problems. First, the hardware must support viewing sharp imagery over the full accommodation range of the user. Second, HMDs should accurately reproduce retinal defocus blur to correctly drive accommodation. A multitude of accommodation-supporting HMDs have been proposed, with three architectures receiving particular attention: varifocal, multifocal, and light field displays. These designs all extend depth of focus, but rely on computationally expensive rendering and optimization algorithms to reproduce accurate defocus blur (often limiting content complexity and interactive applications). To date, no unified framework has been proposed to support driving these emerging HMDs using commodity content. In this paper, we introduce DeepFocus , a generic, end-to-end convolutional neural network designed to efficiently solve the full range of computational tasks for accommodation-supporting HMDs. This network is demonstrated to accurately synthesize defocus blur, focal stacks, multilayer decompositions, and multiview imagery using only commonly available RGB-D images, enabling real-time, near-correct depictions of retinal blur with a broad set of accommodation-supporting HMDs.
Lei Xiao 0014, Anton Kaplanyan, Alexander Fix, Matthew Chapman, Douglas Lanman
ACM Trans. Graph.3
2017 Focal surface displays
abstract
Conventional binocular head-mounted displays (HMDs) vary the stimulus to vergence with the information in the picture, while the stimulus to accommodation remains fixed at the apparent distance of the display, as created by the viewing optics. Sustained vergence-accommodation conflict (VAC) has been associated with visual discomfort, motivating numerous proposals for delivering near-correct accommodation cues. We introduce focal surface displays to meet this challenge, augmenting conventional HMDs with a phase-only spatial light modulator (SLM) placed between the display screen and viewing optics. This SLM acts as a dynamic freeform lens, shaping synthesized focal surfaces to conform to the virtual scene geometry. We introduce a framework to decompose target focal stacks and depth maps into one or more pairs of piecewise smooth focal surfaces and underlying display images. We build on recent developments in "optimized blending" to implement a multifocal display that allows the accurate depiction of occluding, semi-transparent, and reflective objects. Practical benefits over prior accommodation-supporting HMDs are demonstrated using a binocular focal surface display employing a liquid crystal on silicon (LCOS) phase SLM and an organic light-emitting diode (OLED) display.
Nathan Matsuda, Alexander Fix, Douglas Lanman
ACM Trans. Graph.2
2015 A Hypergraph-Based Reduction for Higher-Order Binary Markov Random Fields
abstract
Higher-order Markov Random Fields, which can capture important properties of natural images, have become increasingly important in computer vision. While graph cuts work well for first-order MRF's, until recently they have rarely been effective for higher-order MRF's. Ishikawa's graph cut technique [1], [2] shows great promise for many higher-order MRF's. His method transforms an arbitrary higher-order MRF with binary labels into a first-order one with the same minima. If all the terms are submodular the exact solution can be easily found; otherwise, pseudoboolean optimization techniques can produce an optimal labeling for a subset of the variables. We present a new transformation with better performance than [1], [2], both theoretically and experimentally. While [1], [2] transforms each higher-order term independently, we use the underlying hypergraph structure of the MRF to transform a group of terms at once. For n binary variables, each of which appears in terms with k other variables, at worst we produce n non-submodular terms, while [1], [2] produces O(nk). We identify a local completeness property under which our method perform even better, and show that under certain assumptions several important vision problems (including common variants of fusion moves) have this property. We show experimentally that our method produces smaller weight of non-submodular edges, and that this metric is directly related to the effectiveness of QPBO [3]. Running on the same field of experts dataset used in [1], [2] we optimally label significantly more variables (96 versus 80 percent) and converge more rapidly to a lower energy. Preliminary experiments suggest that some other higher-order MRF's used in stereo [4] and segmentation [5] are also locally complete and would thus benefit from our work.
Alexander Fix, Aritanan Gruber, Endre Boros, Ramin Zabih
IEEE Trans. Pattern Anal. Mach. Intell.1
2014 A Primal-Dual Algorithm for Higher-Order Multilabel Markov Random Fields
abstract
Graph cuts method such as α-expansion [4] and fusion moves [22] have been successful at solving many optimization problems in computer vision. Higher-order Markov Random Fields (MRF's), which are important for numerous applications, have proven to be very difficult, especially for multilabel MRF's (i.e. more than 2 labels). In this paper we propose a new primal-dual energy minimization method for arbitrary higher-order multilabel MRF's. Primal-dual methods provide guaranteed approximation bounds, and can exploit information in the dual variables to improve their efficiency. Our algorithm generalizes the PD3 [19] technique for first-order MRFs, and relies on a variant of max-flow that can exactly optimize certain higher-order binary MRF's [14]. We provide approximation bounds similar to PD3 [19], and the method is fast in practice. It can optimize non-submodular MRF's, and additionally can in- corporate problem-specific knowledge in the form of fusion proposals. We compare experimentally against the existing approaches that can efficiently handle these difficult energy functions [6, 10, 11]. For higher-order denoising and stereo MRF's, we produce lower energy while running significantly faster.
Alexander Fix, Chen Wang 0050, Ramin Zabih
CVPR1
2014 Duality and the Continuous Graphical Model
Alexander Fix, Sameer Agarwal 0001
ECCV (3)1
2013 Structured Learning of Sum-of-Submodular Higher Order Energy Functions
abstract
Sub modular functions can be exactly minimized in polynomial time, and the special case that graph cuts solve with max flow [19] has had significant impact in computer vision [5, 21, 28]. In this paper we address the important class of sum-of-sub modular (SoS) functions [2, 18], which can be efficiently minimized via a variant of max flow called sub modular flow [6]. SoS functions can naturally express higher order priors involving, e.g., local image patches, however, it is difficult to fully exploit their expressive power because they have so many parameters. Rather than trying to formulate existing higher order priors as an SoS function, we take a discriminative learning approach, effectively searching the space of SoS functions for a higher order prior that performs well on our training set. We adopt a structural SVM approach [15, 34] and formulate the training problem in terms of quadratic programming, as a result we can efficiently search the space of SoS priors via an extended cutting-plane algorithm. We also show how the state-of-the-art max flow method for vision problems [11] can be modified to efficiently solve the sub modular flow problem. Experimental comparisons are made against the OpenCV implementation of the Grab Cut interactive segmentation technique [28], which uses hand-tuned parameters instead of machine learning. On a standard dataset [12] our method learns higher order priors with hundreds of parameter values, and produces significantly better segmentations. While our focus is on binary labeling problems, we show that our techniques can be naturally generalized to handle more than two labels.
Alexander Fix, Thorsten Joachims, Sung Min Park 0002, Ramin Zabih
ICCV1
2012 Approximate MRF Inference Using Bounded Treewidth Subgraphs
Alexander Fix, Joyce Chen, Endre Boros, Ramin Zabih
ECCV (1)1
2011 A graph cut algorithm for higher-order Markov Random Fields
abstract
Higher-order Markov Random Fields, which can capture important properties of natural images, have become increasingly important in computer vision. While graph cuts work well for first-order MRF's, until recently they have rarely been effective for higher-order MRF's. Ishikawa's graph cut technique [8, 9] shows great promise for many higher-order MRF's. His method transforms an arbitrary higher-order MRF with binary labels into a first-order one with the same minima. If all the terms are submodular the exact solution can be easily found; otherwise, pseudo-boolean optimization techniques can produce an optimal labeling for a subset of the variables. We present a new transformation with better performance than [8, 9], both theoretically and experimentally. While [8, 9] transforms each higher-order term independently, we transform a group of terms at once. For n binary variables, each of which appears in terms with k other variables, at worst we produce n non-submodular terms, while [8, 9] produces O(nk). We identify a local completeness property that makes our method perform even better, and show that under certain assumptions several important vision problems (including common variants of fusion moves) have this property. Running on the same field of experts dataset used in [8, 9] we optimally label significantly more variables (96% versus 80%) and converge more rapidly to a lower energy. Preliminary experiments suggest that some other higher-order MRF's used in stereo [20] and segmentation [1] are also locally complete and would thus benefit from our work.
Alexander Fix, Aritanan Gruber, Endre Boros, Ramin Zabih
ICCV1