EDBT 2026 Demo / reviewers in the wild / expert
Alexander Fix
dblp:70/10771
· DBLP profile ↗
15ranked-venue papers
6as first author
5since 2021 · last 2026
0009-0002-6163-2354ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-author · 4 since 2021Artificial intelligence and machine learning · 6 · 6 first-authorHuman-computer interaction and ubiquitous computing · 3 · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
3 papers |
Rendering · 50% Virtual and augmented reality · 28% Image and video processing · 13% | |
| Artificial intelligence
6 papers |
Probabilistic and Bayesian machine learning · 54% Optimization for machine learning · 21% Segmentation and scene understanding · 11% | |
| Theoretical computer science
3 papers |
Mathematical optimization · 89% Graph algorithms and graph theory · 11% |
Topics — the 22 heaviest of 24, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Virtual and augmented reality › immersive display
head-mounted display |
0.6 | 2 | 2018 | DeepFocus: learned image synthesis for computational displays · ACM Trans. Graph. 2018 Focal surface displays · ACM Trans. Graph. 2017 |
Mathematical optimization
discrete optimization |
0.5 | 3 | 2014 | A Primal-Dual Algorithm for Higher-Order Multilabel Markov Random Fields · CVPR 2014 Structured Learning of Sum-of-Submodular Higher Order Energy Functions · ICCV 2013 A graph cut algorithm for higher-order Markov Random Fields · ICCV 2011 |
Rendering › antialiasing › supersampling
neural supersampling |
0.4 | 1 | 2020 | Neural supersampling for real-time rendering · ACM Trans. Graph. 2020 |
Rendering
real-time rendering |
0.4 | 1 | 2020 | Neural supersampling for real-time rendering · ACM Trans. Graph. 2020 |
Image and video processing
super-resolution |
0.4 | 1 | 2020 | Neural supersampling for real-time rendering · ACM Trans. Graph. 2020 |
Rendering › sampling and reconstruction
temporal upsampling |
0.4 | 1 | 2020 | Neural supersampling for real-time rendering · ACM Trans. Graph. 2020 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models › markov random field
higher-order markov random fields |
0.4 | 3 | 2015 | A Hypergraph-Based Reduction for Higher-Order Binary Markov Random Fields · IEEE Trans. Pattern Anal. Mach. Intell. 2015 A graph cut algorithm for higher-order Markov Random Fields · ICCV 2011 Structured Learning of Sum-of-Submodular Higher Order Energy Functions · ICCV 2013 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models |
0.4 | 3 | 2014 | Duality and the Continuous Graphical Model · ECCV (3) 2014 A graph cut algorithm for higher-order Markov Random Fields · ICCV 2011 Structured Learning of Sum-of-Submodular Higher Order Energy Functions · ICCV 2013 |
Computational photography and imaging
light field display |
0.3 | 1 | 2017 | Focal surface displays · ACM Trans. Graph. 2017 |
Virtual and augmented reality › near-eye display
multifocal display |
0.3 | 1 | 2017 | Focal surface displays · ACM Trans. Graph. 2017 |
Machine learning › Optimization for machine learning
energy minimization |
0.2 | 1 | 2015 | A Hypergraph-Based Reduction for Higher-Order Binary Markov Random Fields · IEEE Trans. Pattern Anal. Mach. Intell. 2015 |
Machine learning › Optimization for machine learning › energy minimization
graph cuts |
0.2 | 1 | 2015 | A Hypergraph-Based Reduction for Higher-Order Binary Markov Random Fields · IEEE Trans. Pattern Anal. Mach. Intell. 2015 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
markov random field |
0.2 | 1 | 2015 | A Hypergraph-Based Reduction for Higher-Order Binary Markov Random Fields · IEEE Trans. Pattern Anal. Mach. Intell. 2015 |
Mathematical optimization › discrete optimization
energy minimization |
0.2 | 1 | 2014 | A Primal-Dual Algorithm for Higher-Order Multilabel Markov Random Fields · CVPR 2014 |
Mathematical optimization
primal-dual method |
0.2 | 1 | 2014 | A Primal-Dual Algorithm for Higher-Order Multilabel Markov Random Fields · CVPR 2014 |
Natural language and speech › Information extraction and text analysis › sequence labeling
binary sequence labeling |
0.2 | 1 | 2013 | Structured Learning of Sum-of-Submodular Higher Order Energy Functions · ICCV 2013 |
Computer vision › Segmentation and scene understanding
image segmentation |
0.2 | 1 | 2013 | Structured Learning of Sum-of-Submodular Higher Order Energy Functions · ICCV 2013 |
Mathematical optimization › discrete optimization
submodular function minimization |
0.2 | 1 | 2013 | Structured Learning of Sum-of-Submodular Higher Order Energy Functions · ICCV 2013 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
markov random field inference |
0.1 | 1 | 2012 | Approximate MRF Inference Using Bounded Treewidth Subgraphs · ECCV (1) 2012 |
Graph algorithms and graph theory
graph cut |
0.1 | 1 | 2011 | A graph cut algorithm for higher-order Markov Random Fields · ICCV 2011 |
Computer vision › 3D vision
stereo vision |
0.1 | 1 | 2015 | A Hypergraph-Based Reduction for Higher-Order Binary Markov Random Fields · IEEE Trans. Pattern Anal. Mach. Intell. 2015 |
Computer vision › 3D vision › stereo vision
stereo matching |
0.1 | 1 | 2014 | A Primal-Dual Algorithm for Higher-Order Multilabel Markov Random Fields · CVPR 2014 |
Methods — techniques the papers use, named apart from their topics
convolutional neural network · 0.8max-flow · 0.7graph cuts · 0.6pseudo-boolean optimization · 0.5temporal network · 0.4fusion moves · 0.4alpha-expansion · 0.4submodular flow · 0.3structural SVM · 0.3phase-only spatial light modulator · 0.3optimized blending · 0.3focal stack decomposition · 0.3graph cut transformation · 0.2QPBO · 0.2cutting-plane algorithm · 0.2cutting plane algorithm · 0.2submodular minimization · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Digitally Prototype Your Eye Tracker: Simulating Hardware Performance using 3D Synthetic DataabstractEye tracking (ET) is a key enabler for Augmented and Virtual Reality (AR/VR). Prototyping new ET hardware requires assessing the impact of hardware choices on eye tracking performance. This task is compounded by the high cost of obtaining data from sufficiently many variations of real hardware, especially for machine learning, which requires large training datasets. We propose a method for end-to-end evaluation of how hardware changes impact machine learning-based ET performance using only synthetic data. We utilize a dataset of real 3D eyes, reconstructed from light dome data using neural radiance fields (NeRF), to synthesize captured eyes from novel viewpoints and camera parameters. Using this framework, we demonstrate that we can predict the relative performance across various hardware configurations, accounting for variations in sensor noise, illumination brightness, and optical blur. We also compare our simulator with the publicly available eye tracking dataset from the Project Aria glasses, demonstrating a strong correlation with real-world performance. Finally, we present a first-of-its-kind analysis in which we vary ET camera positions, evaluating ET performance ranging from on-axis direct views of the eye to peripheral views on the frame. Such an analysis would have previously required manufacturing physical devices to capture evaluation data. In short, our method enables faster prototyping of ET hardware. Esther Y. H. Lin, Yimin Ding, Jogendra Kundu, Yatong An, Mohamed T. El-Haddad, Alexander Fix |
ETRA | 6 |
| 2026 | Enhancing Eye Feature Estimation from Event Data Streams through Adaptive Inference State Space ModelingabstractEye feature extraction from event-based data streams can be performed efficiently and with low energy consumption, offering great utility to real-world eye tracking pipelines. However, few eye feature extractors are designed to handle sudden changes in event density caused by the changes between gaze behaviors that vary in their kinematics, leading to degraded prediction performance. In this work, we address this problem by introducing the adaptive inference state space model (AISSM), a novel architecture for feature extraction that is capable of dynamically adjusting the relative weight placed on current versus recent information. This relative weighting is determined via estimates of the signal-to-noise ratio and event density produced by a complementary dynamic confidence network. Lastly, we craft and evaluate a novel learning technique that improves training efficiency. Experimental results demonstrate that the AISSM system outperforms state-of-the-art models for event-based eye feature extraction. Viet Dung Nguyen, Mobina Ghorbaninejad, Chengyi Ma, Reynold J. Bailey, Gabriel J. Diaz, Alexander Fix, Ryan J. Suess, Alexander Ororbia |
ETRA | 6 |
| 2026 | Reconstructing Realistic and Relightable EyesabstractAccurately modeling the eye is a challenging task as it exhibits refraction and reflection at the cornea, complex iris texture, and self-occlusion and shadowing due to eyelids and eyelashes. To address these challenges, we present a system for learning a hybrid relightable eye model which can be relit under near-field point lights. Our hybrid model leverages an eyeball mesh for explicitly representing the cornea surface and the reflections on it while learning the geometry and light transport of the periocular region and eye interior implicitly. To account for refraction, we explicitly handle the refraction of camera rays using Snell’s law and predict the refraction of incident light rays using a neural network. Furthermore, we propose an extension of our method which enables us to relight the eye using a fringe projector to simulate structured light. Through experiments, we demonstrate that our method results in higher fidelity rendering under novel viewpoint and lighting conditions, improves learned iris geometry, and more accurately simulates structured light fringe patterns on the eye. Wesley Khademi, Jogendra Kundu, Yatong An, Alexander Fix, David Colmenares |
WACV | 4 |
| 2026 | Polarization-Based Eye Tracking with Personalized Siamese Architecture ETRA011abstractHead-mounted devices integrated with eye tracking promise a solution for natural human-computer interaction. However, they typically require per-user calibration for optimal performance due to inter-person variability. A differential personalization approach using Siamese architectures learns relative gaze displacements and reconstructs absolute gaze from a small set of calibration frames. In this paper, we benchmark Siamese personalization on polarization-enabled eye tracking. For benchmarking, we use a 338-subject dataset captured with a polarization-sensitive camera and 850 nm illumination. We achieve performance comparable to linear calibration with 10-fold fewer samples. Using polarization inputs for Siamese personalization reduces gaze error by up to 12% compared to near-infrared (NIR)-based inputs. Combining Siamese personalization with linear calibration yields further improvements of up to 13% over a linearly calibrated baseline. These results establish Siamese personalization as a practical approach enabling accurate eye tracking. Mantas Zurauskas, Alexander Fix, Beyza Kalkanli, Dmitri Model, Tom Bu, Mahsa Shakeri, Dave Stronks |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2022 | Event-Based Kilohertz Eye Tracking using Coded Differential LightingabstractPixels in an event camera operate asynchronously and independently, reporting changes in intensity as events - tuples of (x,y) position, polarity s and timestamp t at microsecond resolution. Event cameras operate at low power (≈ 5mW) and respond to changes in the scene with a latency on the order of microseconds. These properties make event cameras an exciting candidate for eye tracking sensors on mobile platforms such as Augmented/Virtual Reality (AR/VR) headsets, since these systems have hard real-time and power constraints. One proven method for eye tracking and gaze estimation is corneal glint detection. We exploit the fact that corneal glint tracking only requires a sparse set of pixels in the image, by making use of the natural sparsity of event cameras, which only detect changes in the scene. To enhance this effect, we design an illumination scheme, Coded Differential Lighting, which enhances specular reflections, suppresses all other events, and solves the light-to-glint correspondence. This is the first purely event-based corneal glint detection and tracking algorithm, which operates on standard hardware at kHz sampling rate. Timo Stoffregen, Hossein Daraei, Clare Robinson, Alexander Fix |
WACV | 4 |
| 2020 | Raycast Calibration for Augmented Reality HMDs with Off-Axis Reflective CombinersabstractAugmented reality overlays virtual objects on the real world. To do so, the head mounted display (HMD) needs to be calibrated to establish a mapping between 3D points in the real world with 2D pixels on display panels. This distortion is a high-dimensional function that also depends on pupil position and varifocal settings. We present Raycast calibration, an efficient approach to geometrically calibrate AR displays with off-axis reflective combiners. Our approach requires a small amount of data to estimate a compact, physics-based, and ray-traceable model of the HMD optics. We apply this technique to automatically calibrate an AR prototype with display, SLAM and eye-tracker, without user in the loop. Qi Guo 0009, Huixuan Tang, Aaron Schmitz, Yang Lou, Alexander Fix, Steven Lovegrove, Hauke Strasdat |
ICCP | 6 |
| 2020 | Neural supersampling for real-time renderingabstractDue to higher resolutions and refresh rates, as well as more photorealistic effects, real-time rendering has become increasingly challenging for video games and emerging virtual reality headsets. To meet this demand, modern graphics hardware and game engines often reduce the computational cost by rendering at a lower resolution and then upsampling to the native resolution. Following the recent advances in image and video superresolution in computer vision, we propose a machine learning approach that is specifically tailored for high-quality upsampling of rendered content in real-time applications. The main insight of our work is that in rendered content, the image pixels are point-sampled, but precise temporal dynamics are available. Our method combines this specific information that is typically available in modern renderers (i.e., depth and dense motion vectors) with a novel temporal network design that takes into account such specifics and is aimed at maximizing video quality while delivering real-time performance. By training on a large synthetic dataset rendered from multiple 3D scenes with recorded camera motion, we demonstrate high fidelity and temporally stable results in real-time, even in the highly challenging 4 × 4 upsampling scenario, significantly outperforming existing superresolution and temporal antialiasing work. Lei Xiao 0014, Salah Nouri, Matthew Chapman, Alexander Fix, Douglas Lanman, Anton Kaplanyan |
ACM Trans. Graph. | 4 |
| 2018 | DeepFocus: learned image synthesis for computational displaysabstractAddressing vergence-accommodation conflict in head-mounted displays (HMDs) requires resolving two interrelated problems. First, the hardware must support viewing sharp imagery over the full accommodation range of the user. Second, HMDs should accurately reproduce retinal defocus blur to correctly drive accommodation. A multitude of accommodation-supporting HMDs have been proposed, with three architectures receiving particular attention: varifocal, multifocal, and light field displays. These designs all extend depth of focus, but rely on computationally expensive rendering and optimization algorithms to reproduce accurate defocus blur (often limiting content complexity and interactive applications). To date, no unified framework has been proposed to support driving these emerging HMDs using commodity content. In this paper, we introduce DeepFocus , a generic, end-to-end convolutional neural network designed to efficiently solve the full range of computational tasks for accommodation-supporting HMDs. This network is demonstrated to accurately synthesize defocus blur, focal stacks, multilayer decompositions, and multiview imagery using only commonly available RGB-D images, enabling real-time, near-correct depictions of retinal blur with a broad set of accommodation-supporting HMDs. Lei Xiao 0014, Anton Kaplanyan, Alexander Fix, Matthew Chapman, Douglas Lanman |
ACM Trans. Graph. | 3 |
| 2017 | Focal surface displaysabstractConventional binocular head-mounted displays (HMDs) vary the stimulus to vergence with the information in the picture, while the stimulus to accommodation remains fixed at the apparent distance of the display, as created by the viewing optics. Sustained vergence-accommodation conflict (VAC) has been associated with visual discomfort, motivating numerous proposals for delivering near-correct accommodation cues. We introduce focal surface displays to meet this challenge, augmenting conventional HMDs with a phase-only spatial light modulator (SLM) placed between the display screen and viewing optics. This SLM acts as a dynamic freeform lens, shaping synthesized focal surfaces to conform to the virtual scene geometry. We introduce a framework to decompose target focal stacks and depth maps into one or more pairs of piecewise smooth focal surfaces and underlying display images. We build on recent developments in "optimized blending" to implement a multifocal display that allows the accurate depiction of occluding, semi-transparent, and reflective objects. Practical benefits over prior accommodation-supporting HMDs are demonstrated using a binocular focal surface display employing a liquid crystal on silicon (LCOS) phase SLM and an organic light-emitting diode (OLED) display. Nathan Matsuda, Alexander Fix, Douglas Lanman |
ACM Trans. Graph. | 2 |
| 2015 | A Hypergraph-Based Reduction for Higher-Order Binary Markov Random FieldsabstractHigher-order Markov Random Fields, which can capture important properties of natural images, have become increasingly important in computer vision. While graph cuts work well for first-order MRF's, until recently they have rarely been effective for higher-order MRF's. Ishikawa's graph cut technique [1], [2] shows great promise for many higher-order MRF's. His method transforms an arbitrary higher-order MRF with binary labels into a first-order one with the same minima. If all the terms are submodular the exact solution can be easily found; otherwise, pseudoboolean optimization techniques can produce an optimal labeling for a subset of the variables. We present a new transformation with better performance than [1], [2], both theoretically and experimentally. While [1], [2] transforms each higher-order term independently, we use the underlying hypergraph structure of the MRF to transform a group of terms at once. For n binary variables, each of which appears in terms with k other variables, at worst we produce n non-submodular terms, while [1], [2] produces O(nk). We identify a local completeness property under which our method perform even better, and show that under certain assumptions several important vision problems (including common variants of fusion moves) have this property. We show experimentally that our method produces smaller weight of non-submodular edges, and that this metric is directly related to the effectiveness of QPBO [3]. Running on the same field of experts dataset used in [1], [2] we optimally label significantly more variables (96 versus 80 percent) and converge more rapidly to a lower energy. Preliminary experiments suggest that some other higher-order MRF's used in stereo [4] and segmentation [5] are also locally complete and would thus benefit from our work. Alexander Fix, Aritanan Gruber, Endre Boros, Ramin Zabih |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2014 | A Primal-Dual Algorithm for Higher-Order Multilabel Markov Random FieldsabstractGraph cuts method such as α-expansion [4] and fusion moves [22] have been successful at solving many optimization problems in computer vision. Higher-order Markov Random Fields (MRF's), which are important for numerous applications, have proven to be very difficult, especially for multilabel MRF's (i.e. more than 2 labels). In this paper we propose a new primal-dual energy minimization method for arbitrary higher-order multilabel MRF's. Primal-dual methods provide guaranteed approximation bounds, and can exploit information in the dual variables to improve their efficiency. Our algorithm generalizes the PD3 [19] technique for first-order MRFs, and relies on a variant of max-flow that can exactly optimize certain higher-order binary MRF's [14]. We provide approximation bounds similar to PD3 [19], and the method is fast in practice. It can optimize non-submodular MRF's, and additionally can in- corporate problem-specific knowledge in the form of fusion proposals. We compare experimentally against the existing approaches that can efficiently handle these difficult energy functions [6, 10, 11]. For higher-order denoising and stereo MRF's, we produce lower energy while running significantly faster. Alexander Fix, Chen Wang 0050, Ramin Zabih |
CVPR | 1 |
| 2014 | Duality and the Continuous Graphical Model
Alexander Fix, Sameer Agarwal 0001 |
ECCV (3) | 1 |
| 2013 | Structured Learning of Sum-of-Submodular Higher Order Energy FunctionsabstractSub modular functions can be exactly minimized in polynomial time, and the special case that graph cuts solve with max flow [19] has had significant impact in computer vision [5, 21, 28]. In this paper we address the important class of sum-of-sub modular (SoS) functions [2, 18], which can be efficiently minimized via a variant of max flow called sub modular flow [6]. SoS functions can naturally express higher order priors involving, e.g., local image patches, however, it is difficult to fully exploit their expressive power because they have so many parameters. Rather than trying to formulate existing higher order priors as an SoS function, we take a discriminative learning approach, effectively searching the space of SoS functions for a higher order prior that performs well on our training set. We adopt a structural SVM approach [15, 34] and formulate the training problem in terms of quadratic programming, as a result we can efficiently search the space of SoS priors via an extended cutting-plane algorithm. We also show how the state-of-the-art max flow method for vision problems [11] can be modified to efficiently solve the sub modular flow problem. Experimental comparisons are made against the OpenCV implementation of the Grab Cut interactive segmentation technique [28], which uses hand-tuned parameters instead of machine learning. On a standard dataset [12] our method learns higher order priors with hundreds of parameter values, and produces significantly better segmentations. While our focus is on binary labeling problems, we show that our techniques can be naturally generalized to handle more than two labels. Alexander Fix, Thorsten Joachims, Sung Min Park 0002, Ramin Zabih |
ICCV | 1 |
| 2012 | Approximate MRF Inference Using Bounded Treewidth Subgraphs
Alexander Fix, Joyce Chen, Endre Boros, Ramin Zabih |
ECCV (1) | 1 |
| 2011 | A graph cut algorithm for higher-order Markov Random FieldsabstractHigher-order Markov Random Fields, which can capture important properties of natural images, have become increasingly important in computer vision. While graph cuts work well for first-order MRF's, until recently they have rarely been effective for higher-order MRF's. Ishikawa's graph cut technique [8, 9] shows great promise for many higher-order MRF's. His method transforms an arbitrary higher-order MRF with binary labels into a first-order one with the same minima. If all the terms are submodular the exact solution can be easily found; otherwise, pseudo-boolean optimization techniques can produce an optimal labeling for a subset of the variables. We present a new transformation with better performance than [8, 9], both theoretically and experimentally. While [8, 9] transforms each higher-order term independently, we transform a group of terms at once. For n binary variables, each of which appears in terms with k other variables, at worst we produce n non-submodular terms, while [8, 9] produces O(nk). We identify a local completeness property that makes our method perform even better, and show that under certain assumptions several important vision problems (including common variants of fusion moves) have this property. Running on the same field of experts dataset used in [8, 9] we optimally label significantly more variables (96% versus 80%) and converge more rapidly to a lower energy. Preliminary experiments suggest that some other higher-order MRF's used in stereo [20] and segmentation [1] are also locally complete and would thus benefit from our work. Alexander Fix, Aritanan Gruber, Endre Boros, Ramin Zabih |
ICCV | 1 |