VLDB 2026 Research / reviewers in the wild / expert
Ben Lundell
dblp:126/6862 · also Benjamin E. Lundell, Benjamin Lundell
· DBLP profile ↗
7ranked-venue papers
0as first author
7since 2021 · last 2026
0009-0003-8349-5773ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Schema-Guided Scene-Graph Reasoning Based on Multi-Agent Large Language Model SystemabstractScene graphs have emerged as a structured and serializable environment representation for grounded spatial reasoning with Large Language Models (LLMs). In this work, we propose SG2, an iterative Schema-Guided Scene-Graph reasoning framework based on multi-agent LLMs. The agents are grouped into two modules: a (1) Reasoner module for abstract task planning and graph information queries generation, and a (2) Retriever module for extracting corresponding graph information based on code-writing following the queries. Two modules collaborate iteratively, enabling sequential reasoning and adaptive attention to graph information. The scene graph schema, prompted to both modules, serves to not only streamline both reasoning and retrieval process, but also guide the cooperation between two modules. This eliminates the need to prompt LLMs with full graph data, reducing the chance of hallucination due to irrelevant information. Through experiments in multiple simulation environments, we show that our framework surpasses existing LLM-based approaches and baseline single-agent, tool-based Reason-while-Retrieve strategy in numerical Q&A and planning tasks. Yiye Chen, Harpreet Sawhney, Nicholas Gyde, Yanan Jian, Jack Saunders, Patricio A. Vela, Ben Lundell |
AAAI | 7 |
| 2025 | How Do I Do That? Synthesizing 3D Hand Motion and Contacts for Everyday InteractionsabstractWe tackle the novel problem of predicting 3D hand motion and contact maps (or Interaction Trajectories) given a single RGB view, action text, and a 3D contact point on the object as input. Our approach consists of (1) Interaction Codebook: a VQVAE model to learn a latent codebook of hand poses and contact points, effectively tokenizing interaction trajectories, (2) Interaction Predictor: a transformer-decoder module to predict the interaction trajectory from test time inputs by using an indexer module to retrieve a latent affordance from the learned codebook. To train our model, we develop a data engine that extracts 3D hand poses and contact trajectories from the diverse HoloAssist dataset. We evaluate our model on a benchmark that is 2.5-10× larger than existing works, in terms of diversity of objects and interactions observed, and test for generalization of the model across object categories, action categories, tasks, and scenes. Experimental results show the effectiveness of our approach over transformer & diffusion baselines across all settings. Ben Lundell, Dmitry Andreychuk, David Forsyth, Saurabh Gupta 0001, Harpreet Sawhney |
CVPR | 2 |
| 2025 | GASP: Gaussian Avatars with Synthetic PriorsabstractGaussian Splatting has changed the game for real-time photo-realistic rendering. One of the most popular applications of Gaussian Splatting is to create animatable avatars, known as Gaussian Avatars. Recent works have pushed the boundaries of quality and rendering efficiency but suffer from two main limitations. Either they require expensive multi-camera rigs to produce avatars with free-viewpoint rendering, or they can be trained with a single camera but only rendered at high quality from this fixed viewpoint. An ideal model would be trained using a short monocular video or image from available hardware, such as a webcam, and rendered from any view. To this end, we propose GASP: Gaussian Avatars with Synthetic Priors. To overcome the limitations of existing datasets, we exploit the pixel-perfect nature of synthetic data to train a Gaussian Avatar prior. By fitting this prior model to a single photo or video and fine-tuning it, we get a high-quality Gaussian Avatar, which supports 360° rendering. Our prior is only required for fitting, not inference, enabling real-time applications. Through our method, we obtain high-quality, animatable Avatars from limited data which can be animated and rendered at 70fps on commercial hardware. Jack R. Saunders, Charlie Hewitt, Yanan Jian, Marek Kowalski, Tadas Baltrusaitis, Yiye Chen, Darren Cosker, Virginia Estellers, Nicholas Gyde, Vinay P. Namboodiri, Ben Lundell |
CVPR | 11 |
| 2024 | AR-in-VR simulator: A toolbox for rapid augmented reality simulation and user researchabstractWhile providing exciting opportunities for spatial computing, current generation augmented reality (AR) headsets face visual quality and visual comfort issues. Designing headsets that are both perceptually accurate and comfortable to use for extended periods is difficult as these devices rely on state-of-the-art developments in optics, display technology, spatial computing and graphics. Understanding the requirements and parameters needed to create great AR devices currently requires expensive and time consuming prototype development and user research. Here, we present the AR-in-VR Simulator, a tool to facilitate rapid AR device prototyping and user research of the perceptual requirements needed for AR systems by simulating various display, optics and rendering properties of AR systems and presenting them in a virtual reality (VR) headset. Our simulator is suitable for conducting perceptual research on various potential AR device properties and artefacts, such as FOV size and shape, stereoscopic display configurations, and luminance and color non-uniformity. Simulations can be presented in realistic 3D environments, passthrough VR, or even Gaussian Splat environments allowing for a range of naturalistic data collection possibilities. We further demonstrate the utility of the simulator by using it to extend prior work on the perception of partially overlapping stereoscopic AR displays to a 6DOF VR simulation, capturing more ecologically valid and natural user interaction compared to traditional stereoscope based perceptual studies. Jacob Hadnett-Hunter, Ben Lundell, Ian Ellison-Taylor, Michaela Porubanova, Tapani Alasaarela, Maria Olkkonen |
SAP | 2 |
| 2023 | STEPs: Self-Supervised Key Step Extraction and Localization from Unlabeled Procedural VideosabstractWe address the problem of extracting key steps from un-labeled procedural videos, motivated by the potential of Augmented Reality (AR) headsets to revolutionize job training and performance. We decompose the problem into two steps: representation learning and key steps extraction. We propose a training objective, Bootstrapped Multi-Cue Contrastive (BMC2) loss to learn discriminative representations for various steps without any labels. Different from prior works, we develop techniques to train a light-weight temporal module which uses off-the-shelf features for self supervision. Our approach can seamlessly leverage information from multiple cues like optical flow, depth or gaze to learn discriminative features for key-steps, making it amenable for AR applications. We finally extract key steps via a tunable algorithm that clusters the representations and samples. We show significant improvements over prior works for the task of key step localization and phase classification. Qualitative results demonstrate that the extracted key steps are meaningful and succinctly represent various steps of the procedural tasks. Our code can be found at https://github.com/anshulbshah/STEPs. Anshul Shah 0001, Ben Lundell, Harpreet Sawhney, Rama Chellappa |
ICCV | 2 |
| 2023 | Self-supervised Learning with Local Contrastive Loss for Detection and Semantic SegmentationabstractWe present a self-supervised learning (SSL) method suitable for semi-global tasks such as object detection and semantic segmentation. We enforce local consistency between self-learned features that represent corresponding image locations of transformed versions of the same image, by minimizing a pixel-level local contrastive (LC) loss during training. LC-loss can be added to existing self-supervised learning methods with minimal overhead. We evaluate our SSL approach on two downstream tasks – object detection and semantic segmentation, using COCO, PASCAL VOC, and CityScapes datasets. Our method outperforms the existing state-of-the-art SSL approaches by 1.9% on COCO object detection, 1.4% on PASCAL VOC detection, and 0.6% on CityScapes segmentation. Ashraful Islam, Ben Lundell, Harpreet Sawhney, Sudipta Sinha, Peter Morales, Richard J. Radke |
WACV | 2 |
| 2021 | Full-Body Motion from a Single Head-Mounted Device: Generating SMPL Poses from Partial ObservationsabstractThe increased availability and maturity of head-mounted and wearable devices opens up opportunities for remote communication and collaboration. However, the signal streams provided by these devices (e.g., head pose, hand pose, and gaze direction) do not represent a whole person. One of the main open problems is therefore how to leverage these signals to build faithful representations of the user. In this paper, we propose a method based on variational autoencoders to generate articulated poses of a human skeleton based on noisy streams of head and hand pose. Our approach relies on a model of pose likelihood that is novel and theoretically well-grounded. We demonstrate on publicly available datasets that our method is effective even from very impoverished signals and investigate how pose prediction can be made more accurate and realistic. Andrea Dittadi, Sebastian Dziadzio, Darren Cosker, Ben Lundell, Thomas J. Cashman 0001, Jamie Shotton |
ICCV | 4 |