VLDB 2026 Research / reviewers in the wild / expert
Edward Lu
dblp:17/1945
· DBLP profile ↗
8ranked-venue papers
3as first author
8since 2021 · last 2026
0009-0007-5627-0244ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 6 · 2 first-author · 6 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GenAssist: Interactive Prompt-Driven XR Program GenerationabstractThis paper introduces GenAssist, a system for generating interactive Extended Reality (XR) programs from natural language prompts. Given plain text descriptions of desired programs, our system uses Retrieval-Augmented Generation (RAG) to retrieve related documentation and example code, which is then used to prompt Large Language Models (LLMs) to generate and execute hot-pluggable XR programs in real time. To ensure that the programs are written correctly to the user’s specifications, we add a closed-loop feedback mechanism using virtual cameras in the scene that iteratively refines the system’s output, mimicking the development cycle of human developers that compile and then interactively test programs. GenAssist generates scripts that can not only place multiple primitives and 3D models in various locations in a virtual scene, but it can also animate and enable user interactions with those objects. We show that across a benchmark of 50 diverse XR program prompts, our system achieves high output accuracy and program generation quality. Furthermore, we conduct a user study with 18 participants that demonstrates GenAssist’s effectiveness and usability (NASA TLX = 36.6) for XR program generation. We compare GenAssist to prior systems and show that it is significantly faster (<10 seconds per run) and requires fewer LLM calls. Sruti Srinidhi, Akul Singh, Edward Lu, Anthony Rowe 0001 |
VR | 3 |
| 2025 | An XR Platform that Integrates Large Language Models with the Physical WorldabstractAs Artificial Intelligence (AI) and eXtended Reality (XR) evolve, integrating them effectively remains a challenge. Although multimodal large language models (MLLMs) offer powerful reasoning over text and images, they lack an inherent understanding of 3D space. Additionally, XR headsets are resource-constrained and cannot run these models locally. To address this gap, we introduce XaiR, a system that integrates MLLMs with XR to enable AI-driven spatial reasoning and interaction. XaiR employs a client-server architecture in which an XR headset (client) captures spatial data, generates 2D snapshots of the 3D environment, and renders augmented reality (AR) content, while a remote server runs multiple parallel MLLMs to generate contextually aware responses. Our demo showcases an XR cognitive assistant application that guides a user through a series of instructions. Deployed on a mobile AR headset, our system dynamically interprets user actions, tracks task progress in real time, and provides textual feedback and AR-guided assistance. Sruti Srinidhi, Edward Lu, Akul Singh, Saisha Kartik, Audi Lin, Tarana Laroia, Anthony Rowe 0001 |
SenSys | 2 |
| 2025 | QUASAR: Quad-based Adaptive Streaming And RenderingabstractAs AR/VR systems evolve to demand increasingly powerful GPUs, physically separating compute from display hardware emerges as a natural approach to enable a lightweight, comfortable form factor. Unfortunately, splitting the system into a client-server architecture leads to challenges in transporting graphical data. Simply streaming rendered images over a network suffers in terms of latency and reliability, especially given variable bandwidth. Although image-based reprojection techniques can help, they often do not support full motion parallax or disocclusion events. Instead, scene geometry can be streamed to the client, allowing local rendering of novel views. Traditionally, this has required a prohibitively large amount of interconnect bandwidth, excluding the use of practical networks. This paper presents a new quad-based geometry streaming approach that is designed with compression and the ability to adjust Quality-of-Experience (QoE) in response to target network bandwidths. Our approach advances previous work by introducing a more compact data structure and a temporal compression technique that reduces data transfer overhead by up to 15×, reducing bandwidth usage to as low as 100 Mbps. We optimized our design for hardware video codec compatibility and support an adaptive data streaming strategy that prioritizes transmitting only the most relevant geometry updates. Our approach achieves image quality comparable to, and in many cases exceeds, state-of-the-art techniques while requiring only a fraction of the bandwidth, enabling real-time geometry streaming on commodity headsets over WiFi. Edward Lu, Anthony Rowe 0001 |
ACM Trans. Graph. | 1 |
| 2024 | XaiR: An XR Platform that Integrates Large Language Models with the Physical WorldabstractThis paper discusses the integration of Multimodal Large Language Models (MLLMs) with Extended Reality (XR) headsets, focusing on enhancing machine understanding of physical spaces. By combining the contextual capabilities of MLLMs with the sensory inputs from XR, there is potential for more intuitive spatial interactions. However, the integration faces challenges due to the inherent limitations of MLLMs in processing 3D inputs and their significant resource demands for XR headsets. We introduce XaiR, a platform that facilitates integrating MLLMs with XR applications. XaiR uses a split architecture that offloads complex MLLM operations to a server while handling 3D world processing on the headset. This setup manages multiple input modalities, parallel models, and links them with real-time pose data, improving AR content placement in physical scenes. We tested XaiR’s effectiveness with a “cognitive assistant” application that guides users through tasks like making coffee or assembling furniture. Results from a 15-participant study shows over 90% accuracy in task guidance and 85% accuracy in AR content anchoring. Additionally, we evaluate MLLMs against human operators for cognitive assistant tasks which provides insights into the quality of the captured data as well as the current gap in performance for cognitive assistant tasks. Sruti Srinidhi, Edward Lu, Anthony Rowe 0001 |
ISMAR | 2 |
| 2023 | RenderFusion: Balancing Local and Remote Rendering for Interactive 3D ScenesabstractMany modern-day XR devices (e.g. mobile headsets, phones, etc.) lack the computing resources required to render complex 3D scenes in real-time. Typically, to render a high-resolution scene on a lightweight XR device, 3D designers arduously decimate and fine-tune the objects. As an alternative, remote rendering systems can utilize powerful nearby servers to stream rendering results to a client. While this is a promising solution, it can introduce a variety of latency and reliability issues, especially under variable network conditions. In this paper, we present a distributed rendering system that combines both remote rendering and on-device, “local” rendering to add robustness to network fluctuations and device workloads. To maximize user QoE, our approach dynamically swaps an object’s rendering medium, adjusting for client workload, low frame rates, and several perceptual characteristics. To model these characteristics, we perform a study under simulated conditions to measure how users perceive latency and complexity differences between objects in a scene. Using the results of the study, we then provide an algorithm for choosing the optimal object rendering medium, based on rendering complexity as well as network and latency models, ensuring that a target frame rate will be met. Finally, we evaluate this algorithm on a prototype implementation that can provide cross-platform split rendering using web technologies. Edward Lu, Sagar Bharadwaj, Mallesham Dasari, Connor Smith, Srinivasan Seshan, Anthony Rowe 0001 |
ISMAR | 1 |
| 2023 | Scaling VR Video ConferencingabstractVirtual Reality (VR) telepresence platforms are being challenged to support live performances, sporting events, and conferences with thousands of users across seamless virtual worlds. Current systems have struggled to meet these demands which has led to high-profile performance events with groups of users isolated in parallel sessions. The core difference in scaling VR environments compared to classic 2D video content delivery comes from the dynamic peer-to-peer spatial dependence on communication. Users have many pair-wise interactions that grow and shrink as they explore spaces. In this paper, we discuss the challenges of VR scaling and present an architecture that supports hundreds of users with spatial audio and video in a single virtual environment. We leverage the property of spatial locality with two key optimizations: (1) a Quality of Service (QoS) scheme to prioritize audio and video traffic based on users' locality, and (2) a resource manager that allocates client connections across multiple servers based on user proximity within the virtual world. Through real-world deployments and extensive evaluations under real and simulated environments, we demonstrate the scalability of our platform while showing improved QoS compared with existing approaches. Mallesham Dasari, Edward Lu, Michael W. Farb, Nuno Pereira 0001, Ivan Liang, Anthony Rowe 0001 |
VR | 2 |
| 2021 | ARENA: The Augmented Reality Edge Networking ArchitectureabstractMany have predicted the future of the Web to be the integration of Web content with the real-world through technologies such as Augmented Reality (AR). This has led to the rise of Extended Reality (XR) Web Browsers used to shorten the long AR application development and deployment cycle of native applications especially across different platforms. As XR Browsers mature, we face new challenges related to collaborative and multi-user applications that span users, devices, and machines. These collaborative XR applications require: (1) networking support for scaling to many users, (2) mechanisms for content access control and application isolation, and (3) the ability to host application logic near clients or data sources to reduce application latency. In this paper, we present the design and evaluation of the AR Edge Networking Architecture (ARENA) which is a platform that simplifies building and hosting collaborative XR applications on WebXR capable browsers. ARENA provides a number of critical components including: a hierarchical geospatial directory service that connects users to nearby servers and content, a token-based authentication system for controlling user access to content, and an application/service runtime supervisor that can dispatch programs across any network connected device. All of the content within ARENA exists as endpoints in a PubSub scene graph model that is synchronized across all users. We evaluate ARENA in terms of client performance as well as benchmark end-to-end response-time as load on the system scales. We show the ability to horizontally scale the system to Internet-scale with scenes containing hundreds of users and latencies on the order of tens of milliseconds. Finally, we highlight projects built using ARENA and showcase how our approach dramatically simplifies collaborative multi-user XR development compared to monolithic approaches. Nuno Pereira 0001, Anthony Rowe 0001, Michael W. Farb, Ivan Liang, Edward Lu, Eric Riebling |
ISMAR | 5 |
| 2021 | FLASH: Video-Embeddable AR Anchors for Live EventsabstractPublic spaces like concert stadiums and sporting arenas are ideal venues for AR content delivery to crowds of mobile phone users. Unfortunately, these environments tend to be some of the most challenging in terms of lighting and dynamic staging for vision-based relocalization. In this paper, we introduce FLASH1, a system for delivering AR content within challenging lighting environments that uses active tags (i.e., blinking) with detectable features from passive tags (quads) for marking regions of interest and determining pose. This combination allows the tags to be detectable from long distances with significantly less computational overhead per frame, making it possible to embed tags in existing video displays like large jumbotrons. To aid in pose acquisition, we implement a gravity-assisted pose solver that removes the ambiguous solutions that are often encountered when trying to localize using standard passive tags. We show that our technique outperforms similarly sized passive tags in terms of range by 20-30% and is fast enough to run at 30 FPS even within a mobile web browser on a smartphone. Edward Lu, John Miller 0002, Nuno Pereira 0001, Anthony Rowe 0001 |
ISMAR | 1 |