VLDB 2026 Research / reviewers in the wild / expert
Kenneth Chen
dblp:149/5705
· DBLP profile ↗
14ranked-venue papers
6as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 5 · 3 first-author · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GeneVA: A Dataset of Human Annotations for Generative Text to Video ArtifactsabstractVideo generated by the current state-of-the-art generative models contain undesirable artifacts. We introduce GeneVA, the first large-scale dataset of human-annotated artifact bounding boxes in AI-generated videos. The dataset consists of 16,356 AI-generated videos, each labeled by a human annotator with per-frame artifact bounding boxes, their labels and descriptions, and video quality ratings. A custom data collection pipeline was developed in Prolific, and a novel taxonomy for spatio-temporal artifacts present in AI-generated videos was defined. The videos were from the VidProM [41] dataset, with text prompts from this dataset then used to generate an additional subset of videos using Sora. We trained an artifact detector and caption generator using a pre-trained image-based model, and a custom temporal fusion module. The dataset can be found at https://www.immersivecomputinglab.org/publication/geneva. We hope that datasets like GeneVA will encourage improvements in artifact detection in AI-generated video towards applications such as deepfake detection. Jenna Kang, Maria Beatriz Silva, Patsorn Sangkloy, Kenneth Chen, Niall L. Williams, Qi Sun 0003 |
WACV | 4 |
| 2026 | Perceptually Guided 3DGS Streaming and Rendering for Mixed RealityabstractRecent breakthroughs in radiance fields, particularly 3D Gaussian Splatting (3DGS), have unlocked real-time, high-quality rendering of complex environments, enabling a wide range of applications. However, the stringent requirements of mixed reality (MR) rendering, such as rapid refresh rates, high-resolution stereo viewing, and constrained computing budgets, remain out of reach for current 3DGS techniques. Nevertheless, the wide field-of-view design of MR displays, which mimics human vision, presents a unique opportunity to exploit human visual system’s own perceptual limitations to reduce computational overhead while not compromising user-perceived rendering quality.To this end, we propose a perception-guided, continuous level-of-detail (LOD) framework for 3DGS that maximizes perceived quality under given compute resources. We distill a visual quality metric, which encodes the spatial, temporal, and peripheral characteristics of human visual perception, into a lightweight, gaze-contingent model that predicts and adaptively modulates rendering LOD across the visual field based on each region’s contributions to perceptual quality. This budget-driven LOD modulation, guided by both scene content and gaze behavior, enables significant computation reduction with minimal loss in perceived quality. To support low-power, untethered MR setups, we design an edge-cloud collaborative rendering framework to partially offload computation to the cloud, further reducing overhead on the edge MR devices. Objective metrics and MR user study evidence that, compared to vanilla and foveated LOD baselines, our method achieves superior trade-offs between computational efficiency and user-perceived visual quality. Sai Harsha Mupparaju, Kenneth Chen, Jenna Kang, Maito Omori, Kazuyuki Arimatsu, Qi Sun 0003 |
WACV | 3 |
| 2026 | ML-PEA: Machine Learning-Based Perceptual Algorithms for Display Power OptimizationabstractAbstract Image processing techniques can be used to modulate the pixel intensities of an image to reduce the power consumption of the display device. A simple example of this consists of uniformly dimming the entire image. Such algorithms should strive to minimize the impact on image quality while maximizing power savings. Techniques based on heuristics or human perception have been proposed, both for traditional flat panel displays and modern display modalities such as virtual and augmented reality (VR/AR). In this paper, we focus on developing and evaluating display power‐saving techniques that use machine learning (ML) in VR displays. We developed a U‐Net‐based technique paired with perceptual and power optimization loss functions that generates spatially varying dimming maps. These dimming maps are used to modulate input images, per‐pixel, to generate a power‐efficient image. Our pipeline was validated via quantitative analysis using image quality metrics and through a subjective study. Our subjective validation provides results scaled in perceptual just‐objectionable‐difference (JOD) units. This data, when rescaled, allows for comparisons of our technique with recent studies on VR display power optimization. Our results show that participants prefer our technique over a uniform dimming baseline for high target power saving conditions. This model and study serve as a template and baseline for future applications of deep learning to display power optimization. Model training code and data can be found at kenchen10.github.io/projects/mlpea/index.html . Kenneth Chen, Nathan Matsuda, Thomas Wan, Ajit Ninan, Alexandre Chapiro, Qi Sun 0003 |
Comput. Graph. Forum | 1 |
| 2025 | Process Only Where You Look: Hardware and Algorithm Co-optimization for Efficient Gaze-Tracked Foveated Rendering in Virtual RealityabstractVirtual reality (VR) plays a crucial role in advancing immersive, interactive experiences that transform learning, work, and entertainment by enhancing user engagement and expanding possibilities across various fields.Image rendering is one of the most crucial application in VR, as it produces high-quality, realistic visuals that are vital for maintaining immersive user experiences and preventing visual discomfort or motion sickness.However, the cost of image rendering in VR environment is considerable, primarily due to the demands of high-quality visual experiences from users.This challenge is even greater in real-time applications, where maintaining low latency further increases the complexity of the rendering process.On the other hand, VR devices, such as head-mounted displays (HMDs), are intrinsically linked to human behavior, using insights from perception and cognition to enhance user experience.In this work, we aim to reduce the high computational costs of the rendering process in VR by leveraging natural human eye dynamics and focusing on processing only where you look (POLO).This involves co-optimizing AI algorithms with underlying hardware for greater efficiency.We introduce POLONet, an efficient multitask deep learning framework designed to track human eye movements with minimal latency.Integrated with the POLO accelerator as a plug-in for VR HMD SoCs, this approach significantly lowers image rendering costs, achieving up to a 3.9× reduction in end-to-end latency compared to the latest gaze tracking methods. Wenxuan Liu 0006, Kenneth Chen, Qi Sun 0003, Sai Qian Zhang |
ISCA | 3 |
| 2025 | Perceptually-Guided Acoustic "Foveation"abstractRealistic spatial audio rendering improves immersion in virtual environments. However, the computational complexity of acoustic propagation increases linearly with the number of sources. Consequently, real-time accurate acoustic rendering becomes challenging in highly dynamic scenarios such as virtual and augmented reality (VR/AR). Exploiting the fact that human spatial sensitivity of acoustic sources is not equal at azimuth eccentricities in the horizontal plane, we introduce a perceptually-aware acoustic "foveation" guidance model to the audio rendering pipeline, which can integrate audio sources that are not spatially resolvable by human listeners. To this end, we first conduct a series of psychophysical studies to measure the minimum resolvable audible angular distance under various spatial and background conditions. We leverage this data to derive an azimuth-characterized real-time acoustic foveation algorithm. Numerical analysis and subjective user studies in VR environments demonstrate our method’s effectiveness in significantly reducing acoustic rendering workload, without compromising users’ spatial perception of audio sources. We believe that the presented research will motivate future investigation into the new frontier of modeling and leveraging human multimodal perceptual limitations — beyond the extensively studied visual acuity — for designing efficient VR/AR systems. Kenneth Chen, Irán R. Román, Juan Pablo Bello, Qi Sun 0003, Praneeth Chakravarthula |
VR | 2 |
| 2024 | Exploiting Human Color Discrimination for Memory- and Energy-Efficient Image Encoding in Virtual RealityabstractVirtual Reality (VR) has the potential of becoming the next ubiquitous computing platform. Continued progress in the burgeoning field of VR depends critically on an efficient computing substrate. In particular, DRAM access energy is known to contribute to a significant portion of system energy. Today's framebuffer compression system alleviates the DRAM traffic by using a numerically lossless compression algorithm. Being numerically lossless, however, is unnecessary to preserve perceptual quality for humans. This paper proposes a perceptually lossless, but numerically lossy, system to compress DRAM traffic. Our idea builds on top of long-established psychophysical studies that show that humans cannot discriminate colors that are close to each other. The discrimination ability becomes even weaker (i.e., more colors are perceptually indistinguishable) in our peripheral vision. Leveraging the color discrimination (in)ability, we propose an algorithm that adjusts pixel colors to minimize the bit encoding cost without introducing visible artifacts. The algorithm is coupled with lightweight architectural support that, in real-time, reduces the DRAM traffic by 66.9% and outperforms existing framebuffer compression mechanisms by up to 20.4%. Psychophysical studies on human participants show that our system introduce little to no perceptual fidelity degradation. Nisarg Ujjainkar, Ethan Shahan, Kenneth Chen, Budmonde Duinkharjav, Qi Sun 0003, Yuhao Zhu 0001 |
ASPLOS (1) | 3 |
| 2024 | PEA-PODs: Perceptual Evaluation of Algorithms for Power Optimization in XR DisplaysabstractDisplay power consumption is an emerging concern for untethered devices. This goes double for augmented and virtual extended reality (XR) displays, which target high refresh rates and high resolutions while conforming to an ergonomically light form factor. A number of image mapping techniques have been proposed to extend battery usage. However, there is currently no comprehensive quantitative understanding of how the power savings provided by these methods compare to their impact on visual quality. We set out to answer this question. To this end, we present a perceptual evaluation of algorithms (PEA) for power optimization in XR displays (PODs). Consolidating a portfolio of six power-saving display mapping approaches, we begin by performing a large-scale perceptual study to understand the impact of each method on perceived quality in the wild. This results in a unified quality score for each technique, scaled in just-objectionable-difference (JOD) units. In parallel, each technique is analyzed using hardware-accurate power models. The resulting JOD-to-Milliwatt transfer function provides a first-of-its-kind look into tradeoffs offered by display mapping techniques, and can be directly employed to make architectural decisions for power budgets on XR displays. Finally, we leverage our study data and power models to address important display power applications like the choice of display primary, power implications of eye tracking, and more 1 . Kenneth Chen, Thomas Wan, Nathan Matsuda, Ajit Ninan, Alexandre Chapiro, Qi Sun 0003 |
ACM Trans. Graph. | 1 |
| 2023 | Freeform Templates: Combining Freeform Curation with Structured TemplatesabstractOnline whiteboards are becoming a popular way to facilitate collaborative design work, providing a free-form environment to curate ideas. However, as templates are increasingly being used to scaffold contributions from non-experts designers, it is crucial to understand their impact on the creative process. In this paper, we present the results from a study with 114 students in a large introductory design course. Our results confirm prior findings that templates benefit students by providing a starting point, a shared process, and the ability to access their own work from previous steps. While prior research has criticized templates for being too rigid, we discovered that using templates within a free-form environment resulted in visual patterns of free-form curation where concepts were spatially organized, clustered, color-coded, and connected using arrows and lines. We introduce the concept of ‘Free-form Templates’ to illustrate how templates and free-form curation can be synergistic. Stephen MacNeil, Ziheng Huang 0002, Kenneth Chen, Zijian Ding, Alexander Yu, Kendall Nakai, Steven Dow |
Creativity & Cognition | 3 |
| 2023 | Towards Learning and Generating Audience Motion from VideoabstractThere has recently been an explosion of interest in creating large-scale shared virtual spaces for multiplayer content. However, rendering player-controllable avatars in real-time creates latency issues when scaling to thousands of players. We introduce a human audience video dataset to support applications in deep learning-based 2D video audience simulation, bypassing the need for background 3D virtual humans. This dataset consists of YouTube videos that depict audiences with diverse lighting conditions, color, dress, and movement patterns. We describe the dataset statistics, our implicit data collection strategy, and audience video extraction pipeline. We apply deep learning tasks on this data based on video prediction techniques, and propose a novel method for 2D audience simulations. Kenneth Chen, Norman I. Badler |
SCA | 1 |
| 2022 | Color-Perception-Guided Display Power Reduction for Virtual RealityabstractBattery life is an increasingly urgent challenge for today's untethered VR and AR devices. However, the power efficiency of head-mounted displays is naturally at odds with growing computational requirements driven by better resolution, refresh rate, and dynamic ranges, all of which reduce the sustained usage time of untethered AR/VR devices. For instance, the Oculus Quest 2, under a fully-charged battery, can sustain only 2 to 3 hours of operation time. Prior display power reduction techniques mostly target smartphone displays. Directly applying smartphone display power reduction techniques, however, degrades the visual perception in AR/VR with noticeable artifacts. For instance, the "power-saving mode" on smartphones uniformly lowers the pixel luminance across the display and, as a result, presents an overall darkened visual perception to users if directly applied to VR content. Our key insight is that VR display power reduction must be cognizant of the gaze-contingent nature of high field-of-view VR displays. To that end, we present a gaze-contingent system that, without degrading luminance, minimizes the display power consumption while preserving high visual fidelity when users actively view immersive video sequences. This is enabled by constructing 1) a gaze-contingent color discrimination model through psychophysical studies, and 2) a display power model (with respect to pixel color) through real-device measurements. Critically, due to the careful design decisions made in constructing the two models, our algorithm is cast as a constrained optimization problem with a closed-form solution, which can be implemented as a real-time, image-space shader. We evaluate our system using a series of psychophysical studies and large-scale analyses on natural images. Experiment results show that our system reduces the display power by as much as 24% (14% on average) with little to no perceptual fidelity degradation. Budmonde Duinkharjav, Kenneth Chen, Abhishek Tyagi, Yuhao Zhu 0001, Qi Sun 0003 |
ACM Trans. Graph. | 2 |
| 2019 | Towards Design Principles for Fashion in Interactive Emergent Narrative
Kenneth Chen |
ICIDS | 1 |
| 2018 | Towards Design Principles for Humor in Interactive Emergent Narrative
Kenneth Chen, Stefan Rank |
DiGRA Conference | 1 |
| 2017 | Towards more meaningful interactive narrative with intelligent affective charactersabstractInteractive emergent narrative uses affective agents to create stories, but these stories often lack a thematic direction. I propose an approach towards a modular system using reified theme objects which can point a story towards the author's intended direction while maintaining the autonomy of all involved agents. I describe my current progress using an authored story as a test case, leading towards my future work to create theme objects as a computational system. Kenneth Chen |
ACII | 1 |
| 2001 | Stanford SKOLAR, M.D.: A Model for Learner-initiated, Learner-manipulated, In-context Continuing Medical Education
Howard R. Strasberg, Thomas C. Rindfleisch, Kenneth Chen, Jeremy C. Durack, Helen Deng, Jie S. Yan, Daniel Greenwald, Kenneth L. Melmon |
AMIA | 3 |