Andrea Colaco

dblp:117/4791 · also Andrea Colaço · DBLP profile ↗
← Back
15ranked-venue papers
3as first author
9since 2021 · last 2026
0009-0001-6661-2216ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 From Videos to Conversations: Egocentric Instructions for Task Assistance
Lavisha Aggarwal, Vikas Bahirwani, Andrea Colaco
ICPR (10)3
2026 SurfaceXR: Fusing Smartwatch IMUs and Egocentric Hand Pose for Seamless Surface Interactions
abstract
Mid-air gestures in Extended Reality (XR) often lead to fatigue, discomfort and imprecision, limiting their suitability for extended use. Surface-based interactions offer a compelling alternative, providing improved accuracy, speed, and comfort. However, current egocentric vision-based methods struggle with reliable surface inputs due to challenges in hand tracking and surface plane estimation from oblique and occluded viewing angles. To this extent, we introduce SurfaceXR, a novel sensor fusion approach that combines headset based hand tracking with micro-vibration data sampled from commodity smartwatch IMUs to enable precise and robust inputs on everyday surfaces. Our system is designed with flexibility in mind - it can function using only hand tracking, only IMU sensing, or optimally with both modalities combined, and remains robust even without explicit surface calibration. Our key insight is that these modalities are complementary - hand tracking provides 3D positional data of hand joints, whereas IMUs supply high-frequency wrist/hand motion data. Our user study across 21 participants validates SurfaceXR's effectiveness in augmenting surface touch tracking and 8 class hand-surface gesture recognition, demonstrating significant improvements over single-modality approaches. Enabled by SurfaceXR, we demonstrate a series of interactive apps for both AR and VR, ranging from on-surface sketching, text entry and gesture-based navigation.
Vasco Xu, Eric J. Gonzalez, Andrea Colaco, Henry Hoffmann, Mar González-Franco, Karan Ahuja
IEEE Trans. Vis. Comput. Graph.4
2025 Online-EYE: Multimodal Implicit Eye Tracking Calibration for XR
abstract
Unlike other inputs for extended reality (XR) that work out of the box, eye tracking typically requires custom calibration per user or session. We present a multimodal inputs approach for implicit calibration of eye tracker in VR, leveraging UI interaction for continuous, background calibration. Our method analyzes gaze data alongside controller interaction with UI elements, and employing ML techniques it continuously refines the calibration matrix without interrupting users from their current tasks. Potentially eliminating the need for explicit calibration. We demonstrate the accuracy and effectiveness of this implicit approach across various tasks and real time applications achieving comparable eye tracking accuracy to native, explicit calibration. While our evaluation focuses on VR and controller-based interactions, we anticipate the broader applicability of this approach to various XR devices and input modalities.
Baosheng James Hou, Lucy Abramyan, Prasanthi Gurumurthy, Haley Adams, Ivana Tosic Rodgers, Eric J. Gonzalez, Khushman Patel, Andrea Colaco, Ken Pfeuffer, Hans-Werner Gellersen, Karan Ahuja, Mar González-Franco
CHI8
2025 EI-Lite: Electrical Impedance Sensing for Micro-gesture Recognition and Pinch Force Estimation
Junyi Zhu 0001, Tianyu Xu 0008, Emily Guan, JaeYoung Moon, Stiven Morvan, D. Shin, Andrea Colaco, Stefanie Mueller 0001, Karan Ahuja, Yiyue Luo, Ishan Chatterjee
UIST8
2025 EgoTrigger: Toward Audio-Driven Image Capture for Human Memory Enhancement in All-Day Energy-Efficient Smart Glasses
abstract
All-day smart glasses are likely to emerge as platforms capable of continuous contextual sensing, uniquely positioning them for unprecedented assistance in our daily lives. Integrating the multi-modal AI agents required for human memory enhancement while performing continuous sensing, however, presents a major energy efficiency challenge for all-day usage. Achieving this balance requires intelligent, context-aware sensor management. Our approach, EgoTrigger, leverages audio cues from the microphone to selectively activate power-intensive cameras, enabling efficient sensing while preserving substantial utility for human memory enhancement. EgoTrigger uses a lightweight audio model (YAMNet) and a custom classification head to trigger image capture from hand-object interaction (HOI) audio cues, such as the sound of a drawer opening or a medication bottle being opened. In addition to evaluating on the QA-Ego4D dataset, we introduce and evaluate on the Human Memory Enhancement Question-Answer (HME-QA) dataset. Our dataset contains 340 human-annotated first-person QA pairs from full-length Ego4D videos that were curated to ensure that they contained audio, focusing on HOI moments critical for contextual understanding and memory. Our results show EgoTrigger can use 54% fewer frames on average, significantly saving energy in both power-hungry sensing components (e.g., cameras) and downstream operations (e.g., wireless transmission), while achieving comparable performance on datasets for an episodic memory task. We believe this context-aware triggering strategy represents a promising direction for enabling energy-efficient, functional smart glasses capable of all-day use - supporting applications like helping users recall where they placed their keys or information about their routine activities (e.g., taking medications).
Akshay Paruchuri, Sinan Hersek, Lavisha Aggarwal, Xin Liu 0034, Achin Kulshrestha, Andrea Colaco, Henry Fuchs, Ishan Chatterjee
IEEE Trans. Vis. Comput. Graph.7
2024 Diffuse, Attend, and Segment: Unsupervised Zero-Shot Segmentation using Stable Diffusion
abstract
Producing quality segmentation masks for images is a fundamental problem in computer vision. Recent research has explored large-scale supervised training to enable zero- shot transfer segmentation on virtually any image style and unsupervised training to enable segmentation without dense annotations. However, constructing a model capable of segmenting anything in a zero-shot manner without any anno-tations is still challenging. In this paper, we propose to uti-lize the self-attention layers in stable diffusion models to achieve this goal because the pre-trained stable diffusion model has learned inherent concepts of objects within its attention layers. Specifically, we introduce a simple yet ef-fective iterative merging process based on measuring KL divergence among attention maps to merge them into valid segmentation masks. The proposed method does not re-quire any training or language dependency to extract qual-ity segmentation for any images. On COCO-Stuff-27, our method surpasses the prior unsupervised zero-shot trans-fer SOTA method by an absolute 26% in pixel accuracy and 17% in mean IoU. The project page is at https://sites.google.com/view/diffseg/home.11Georgia Institute of Technology
Junjiao Tian, Lavisha Aggarwal, Andrea Colaco, Zsolt Kira, Mar González-Franco
CVPR3
2024 Geometry Fidelity for Spherical Images
Anders Christensen, Nooshin Mojab, Khushman Patel, Karan Ahuja, Zeynep Akata, Ole Winther, Mar González-Franco, Andrea Colaco
ECCV (80)8
2024 Augmented Object Intelligence with XR-Objects
abstract
Seamless integration of physical objects as interactive digital entities remains a challenge for spatial computing. This paper explores Augmented Object Intelligence (AOI) in the context of XR, an interaction paradigm that aims to blur the lines between digital and physical by equipping real-world objects with the ability to interact as if they were digital, where every object has the potential to serve as a portal to digital functionalities. Our approach utilizes real-time object segmentation and classification, combined with the power of Multimodal Large Language Models (MLLMs), to facilitate these interactions without the need for object pre-registration. We implement the AOI concept in the form of XR-Objects, an open-source prototype system that provides a platform for users to engage with their physical environment in contextually relevant ways using object-based context menus. This system enables analog objects to not only convey information but also to initiate digital actions, such as querying for details or executing tasks. Our contributions are threefold: (1) we define the AOI concept and detail its advantages over traditional AI assistants, (2) detail the XR-Objects system’s open-source design and implementation, and (3) show its versatility through various use cases and a user study.
Mustafa Doga Dogan, Eric J. Gonzalez, Karan Ahuja, Ruofei Du, Andrea Colaco, Johnny Lee, Mar González-Franco, David Kim 0002
UIST5
2023 HIME: Efficient Headshot Image Super-Resolution with Multiple Exemplars
abstract
A promising direction for recovering the lost information in low-resolution headshot images is utilizing a set of high-resolution exemplars from the same identity. Complementary images in the reference set can improve the generated headshot quality across many different views and poses. However, it is challenging to make the best use of multiple exemplars: the quality and alignment of each exemplar cannot be guaranteed. Using low-quality and mismatched images as references will impair the output results. To overcome these issues, we propose the Headshot Image Super-Resolution with Multiple Exemplars network (HIME) method. Compared with previous methods, our network can effectively handle the misalignment between the input and the reference without requiring facial priors and learn the aggregated reference set representation in an end-to-end manner. Furthermore, to reconstruct more detailed facial features, we propose a correlation loss that provides a rich representation of the local texture in a controllable spatial range. Experimental results demonstrate that the proposed framework not only has significantly fewer computation cost than recent exemplar-guided methods but also achieves better qualitative and quantitative performance.
Xiaoyu Xiang, Jon Morton, Fitsum A. Reda, Lucas D. Young, Federico Perazzi, Amit Kumar 0013, Andrea Colaco, Jan P. Allebach
WACV8
2013 Phase unwrapping and denoising for time-of-flight imaging using generalized approximate message passing
abstract
We present a new method for simultaneously denoising and unwrapping phase in multi-frequency homodyne time-of-flight ranging for the formation of accurate depth maps despite low SNR of raw measurements. This is achieved with a new generalized approximate message passing (GAMP) algorithm for minimum mean-squared error estimation of the phase. A detailed, physically-accurate acquisition model is central in achieving high accuracy, and the use of the GAMP methodology allows low computational complexity despite dense dependencies and the nonlinearity and non-Gaussianity of the acquisition model. Numerical simulations demonstrate that our integrated approach performs better than separate unwrapping followed by denoising. This performance translates to lowering the optical power consumption of time-of-flight cameras for a fixed acquisition quality.
Jonathan Mei, Ahmed Kirmani, Andrea Colaco, Vivek K. Goyal
ICIP3
2013 Mime: compact, low power 3D gesture sensing for interaction with head mounted displays
abstract
We present Mime, a compact, low-power 3D sensor for unencumbered free-form, single-handed gestural interaction with head-mounted displays (HMDs). Mime introduces a real-time signal processing framework that combines a novel three-pixel time-of-flight (TOF) module with a standard RGB camera. The TOF module achieves accurate 3D hand localization and tracking, and it thus enables motion-controlled gestures. The joint processing of 3D information with RGB image data enables finer, shape-based gestural interaction.
Andrea Colaco, Ahmed Kirmani, Hye Soo Yang, Nan-Wei Gong, Chris Schmandt, Vivek K. Goyal
UIST1
2012 Compressive depth map acquisition using a single photon-counting detector: Parametric signal processing meets sparsity
abstract
Active range acquisition systems such as light detection and ranging (LIDAR) and time-of-flight (TOF) cameras achieve high depth resolution but suffer from poor spatial resolution. In this paper we introduce a new range acquisition architecture that does not rely on scene raster scanning as in LIDAR or on a two-dimensional array of sensors as used in TOF cameras. Instead, we achieve spatial resolution through patterned sensing of the scene using a digital micromirror device (DMD) array. Our depth map reconstruction uses parametric signal modeling to recover the set of distinct depth ranges present in the scene. Then, using a convex program that exploits the sparsity of the Laplacian of the depth map, we recover the spatial content at the estimated depth ranges. In our experiments we acquired 64×64-pixel depth maps of fronto-parallel scenes at ranges up to 2.1 M using a pulsed laser, a DMD array and a single photon-counting detector. We also demonstrated imaging in the presence of unknown partially-transmissive occluders. The prototype and results provide promising directions for non-scanning, low-complexity range acquisition devices for various computer vision applications.
Andrea Colaco, Ahmed Kirmani, Gregory A. Howland, John C. Howell, Vivek K. Goyal
CVPR1
2012 CoDAC: A compressive depth acquisition camera framework
abstract
Light detection and ranging (LIDAR) systems use time of flight (TOF) in combination with raster scanning of the scene to form depth maps, and TOF cameras instead make TOF measurements in parallel by using an array of sensors. Here we present a framework for depth map acquisition using neither raster scanning by the illumination source nor an array of sensors. Our architecture uses a spatial light modulator (SLM) to spatially pattern a temporally-modulated light source. Then, measurements from a single omnidirectional sensor provide adequate information for depth map estimation at a resolution equal that of the SLM. Proof-of-concept experiments have verified the validity of our modeling and algorithms.
Ahmed Kirmani, Andrea Colaco, Franco N. C. Wong, Vivek K. Goyal
ICASSP2
2011 Back Talk: An auditory environment for sociable television viewing
abstract
Video content is being consumed in a host of new ways-viewers are no longer restricted to same-time or same-place viewing. However, the experience of watching with a group is inherently social and often desirable despite the physical distribution of group members. This paper introduces Back Talk, a system designed to create a sociable television watching experience. We enhance TV viewing with an auditory environment around a listener. We have explored and leveraged the richness of audio to convey presence of remote viewers. We have developed a novel framework for capturing and translating engagement of an individual into a set of audio cues that are played spatially around a listener. Such auditory enhancements can augment video content consumption in the future.
Andrea Colaco, Ig-Jae Kim, Chris Schmandt
CCNC1
2010 My Second Bike: A TV-Enabled Social and Interactive Riding Experience
abstract
In this paper, we propose a novel concept for a social TV application targeting the demographic of viewers enjoying live sports events, such as road bicycle racing. We intend to enhance the viewing experiences of spectators with sensor-fitted bikes tied to an interactive biking environment on television. The system enables a new form of personalized, physical, and virtual-reality interaction between viewers and a TV program, as well as interactions within or between communities of friends. We also describe a prototype we have implemented to demonstrate the feasibility of our idea. The prototype, my second bike, uses a 3D mirrored world environment (Google Earth) to visually represent participating spectators, competing athletes and outdoor bikers. We contend that the system has the potential to attract and support a large user base on account of its scalability, ease of deployment and ability to promote audience participation in live sports events on TV.
Jaewoo Chung, Kuang Xu, Andrea Colaco, Chris Schmandt, Victor O. K. Li
CCNC3