VLDB 2026 Research / reviewers in the wild / expert
Christian Vogler
dblp:57/2818
· DBLP profile ↗
28ranked-venue papers
6as first author
9since 2021 · last 2026
0000-0003-2590-6880ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 11 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 9 · 5 first-authorApplied, interdisciplinary, general and emerging computing · 7 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Deaf and Hard of Hearing Access to Intelligent Personal Assistants: Comparison of Voice-Based Options with an LLM-Powered Touch InterfaceabstractWe investigate intelligent personal assistants (IPAs) accessibility for deaf and hard of hearing (DHH) people who can use their voice in everyday communication. The inability of IPAs to understand diverse accents including deaf speech renders them largely inaccessible to non-signing and speaking DHH individuals. Using an Echo Show, we compared the usability of natural language input via two spoken English methods against that of a large language model (LLM)-assisted touch interface in a mixed-methods study. The two spoken English methods consisted of Alexa’s built-in automatic speech recognition and a Wizard-of-Oz setting with a trained facilitator re-speaking commands. The touch method was navigated through an LLM-powered ‘task prompter,’ which integrated the user’s history and smart environment to suggest contextually-appropriate commands. Quantitative results showed no significant differences across both spoken English conditions vs LLM-assisted touch. Qualitative results showed variability in opinions on the usability of each method. Ultimately, it will be necessary to have robust deaf-accented speech recognized natively by IPAs. Paige S. DeVries, Michaela Okosi, Nora Dunphy, Gidey Gezae, Dante Conway, Abraham Glasser, Raja S. Kushalnagar, Christian Vogler |
CHI | 9 |
| 2026 | Beyond the Touchscreen: Hands-Free Sign Language and Head-Pointing Interfaces for Deaf Interaction with Intelligent AssistantsabstractAbstract Intelligent Personal Assistants (IPAs) are currently limited to mostly voice input by users, which often does not work for Deaf and Hard of Hearing (DHH) users’ accessibility. Touch interfaces are an accessible alternative in principle, and these recently have been combined with large language models (LLMs) for usability enhancements. However, these are not hands-free, and it is not always feasible to walk up to a device and interact with its touchscreen, such as in the kitchen with dirty hands. This paper situates an LLM-powered touch interface against hands-free options. We present a study with 23 DHH participants who tested three potential input methods for interacting with IPAs: American Sign Language (ASL) in a Wizard-of-Oz setting, LLM-assisted touch through a touchscreen, and LLM-assisted touch through hands-free head pointing. ASL and LLM-assisted touch had comparable usability scores, while headpointing scored much worse. Despite comparable usability between ASL and touch, participants were much more enthusiastic about ASL input. This suggests ASL recognition should be the ultimate goal, but because such technology is not yet commercially viable, further research is needed for identifying practical hands-free alternatives to voice interaction with IPAs. Nora Dunphy, Gidey Gezae, Paige S. DeVries, Pranav Pidathala, Abraham Glasser, Raja S. Kushalnagar, Christian Vogler |
ICCHP (1) | 8 |
| 2026 | Accessibility for the Deaf and Hard-of-Hearing Introduction to the Special Thematic Session
Raja S. Kushalnagar, Matjaz Debevc, Christian Vogler |
ICCHP (1) | 3 |
| 2026 | Accessible Deaf and Hard of Hearing Hybrid Events in the 2020sabstractAbstract Hybrid events have become increasingly common, yet supporting accessible participation for deaf and hard of hearing (DHH) audiences remains challenging. Although accessibility practices for in-person and virtual settings are relatively well established, hybrid environments introduce additional coordination demands, particularly in aligning sign language interpretation, captioning, and audiovisual (AV) workflows across modalities. We present a case study of a large DHH-focused hybrid conference with over 300 virtual and approximately 90 in-person and hybrid attendees. Drawing on planning materials, live workflow observations, and post-event reflections, we examine how accessibility was implemented across in-person and remote contexts. Our analysis identifies recurring challenges in interpreter configuration, Q&A management, and AV coordination. We further identify strategies that supported equitable participation, including the use of separate interpreter teams, structured Q&A workflows, and coordinated AV control across environments. Our findings highlight key trade-offs in supporting visual communication across modalities. Michaela Okosi, Joshua Prado, Abraham Glasser, Raja S. Kushalnagar, Christian Vogler |
ICCHP (1) | 5 |
| 2024 | How Users Experience Closed Captions on Live Television: Quality Metrics Remain a ChallengeabstractThis paper presents a mixed methods study on how deaf, hard of hearing and hearing viewers perceive live TV caption quality with captioned video stimuli designed to mirror TV captioning experiences. To assess caption quality, we used four commonly-used quality metrics focusing on accuracy: word error rate, weighted word error rate, automated caption evaluation (ACE), and its successor ACE2. We calculated the correlation between the four quality metrics and viewer ratings for subjective quality and found that the correlation was weak, revealing that other factors besides accuracy affect user ratings. Additionally, even high-quality captions are perceived to have problems, despite controlling for confounding factors. Qualitative analysis of viewer comments revealed three major factors affecting their experience: Errors within captions, difficulty in following captions, and caption appearance. The findings raise questions as to how objective caption quality metrics can be reconciled with the user experience across a diverse spectrum of viewers. Mariana Arroyo Chavez, Molly Feanny, Matthew Seita, Bernard Thompson, Keith Delk, Skyler Officer, Abraham Glasser, Raja S. Kushalnagar, Christian Vogler |
CHI | 9 |
| 2024 | Towards Co-Creating Access and Inclusion: A Group Autoethnography on a Hearing Individual's Journey Towards Effective Communication in Mixed-Hearing Ability Higher Education SettingsabstractWe present a group autoethnography detailing a hearing student’s journey in adopting communication technologies at a mixed-hearing ability summer research camp. Our study focuses on how this student, a research assistant with emerging American Sign Language (ASL) skills, (in)effectively communicates with deaf and hard-of-hearing (DHH) peers and faculty during the ten-week program. The DHH members also reflected on their communication with the hearing student. We depict scenarios and analyze the (in)effectiveness of how emerging technologies like live automatic speech recognition (ASR) and typing are utilized to facilitate communication. We outline communication strategies to engage everyone with diverse signing skills in conversations - directing visual attention, pause-for-attention-and-proceed, and back-channeling via expressive body. These strategies promote inclusive collaboration and leverage technology advancements. Furthermore, we delve into the factors that have motivated individuals to embrace more inclusive communication practices and provide design implications for accessible communication technologies within the mixed-hearing ability context. Si Chen 0006, James M. Waller, Matthew Seita, Christian Vogler, Raja S. Kushalnagar, Qi Wang 0088 |
CHI | 4 |
| 2024 | Assessment of Sign Language-Based versus Touch-Based Input for Deaf Users Interacting with Intelligent Personal AssistantsabstractWith the recent advancements in intelligent personal assistants (IPAs), their popularity is rapidly increasing when it comes to utilizing Automatic Speech Recognition within households. In this study, we used a Wizard-of-Oz methodology to evaluate and compare the usability of American Sign Language (ASL), Tap to Alexa, and smart home apps among 23 deaf participants within a limited-domain smart home environment. Results indicate a slight usability preference for ASL. Linguistic analysis of the participants' signing reveals a diverse range of expressions and vocabulary as they interacted with IPAs in the context of a restricted-domain application. On average, deaf participants exhibited a vocabulary of 47 +/- 17 signs with an additional 10 +/- 7 fingerspelled words, for a total of 246 different signs and 93 different fingerspelled words across all participants. We discuss the implications for the design of limited-vocabulary applications as a stepping-stone toward general-purpose ASL recognition in the future. Nina Tran, Paige S. DeVries, Matthew Seita, Raja S. Kushalnagar, Abraham Glasser, Christian Vogler |
CHI | 6 |
| 2024 | Closed Sign Language Interpreting: A Usability StudyabstractAbstract Closed sign language interpreting makes media accessible to deaf and hard-of-hearing viewers who use sign language as their primary mode of communication. Analogous to subtitles, this feature allows to toggle sign language interpretation on and off, and customize its appearance in conjunction with videos. This paper provides information on designing closed interpreting in a media player through a pair of mixed-method studies. The first study assesses the usability of technical sign language interpreting features, while the second one assesses how users interact with the content. Results indicate above-average usability for the technical features. Additionally, preliminary results suggest that the optimal configuration of the SLI depends on the type of content viewed and that user preferences vary. Overall, the customizability of features and placement will be important in closed-interpreting implementations. Patrick Boudreault, Muhammad Abubakar, Andrew Duran, Bridget Lam, Zehui Liu, Christian Vogler, Raja S. Kushalnagar |
ICCHP (2) | 6 |
| 2024 | Customization of Closed Captions via Large Language ModelsabstractAbstract This study investigates the feasibility of employing artificial intelligence and large language models (LLMs) to customize closed captions/subtitles to match the personal needs of deaf and hard of hearing viewers. Drawing on recorded live TV samples, it compares user ratings of caption quality, speed, and understandability across five experimental conditions: unaltered verbatim captions, slowed-down verbatim captions, moderately and heavily edited captions via ChatGPT, and lightly edited captions by an LLM optimized for TV content by AppTek, LLC. Results across 16 deaf and hard of hearing participants show a significant preference for verbatim captions, both at original speeds and in the slowed-down version, over those edited by ChatGPT. However, a small number of participants also rated AI-edited captions as best. Despite the overall poor showing of AI, the results suggest that LLM-driven customization of captions on a per-user and per-video basis remains an important avenue for future research. Mariana Arroyo Chavez, Bernard Thompson, Molly Feanny, Kafayat Alabi, Lu Ming, Abraham Glasser, Raja S. Kushalnagar, Christian Vogler |
ICCHP (2) | 9 |
| 2020 | Teleconference Accessibility and Guidelines for Deaf and Hard of Hearing UsersabstractIn this experience report, we describe the accessibility challenges that deaf and hard of hearing users face in teleconferences, based on both our first-hand participation in meetings, and as User Interface and Experience experts. Teleconferencing poses new accessibility challenges compared to face-to-face communication because of limited social, emotional, and haptic feedback. Above all, teleconferencing participants and organizers need to be flexible, because deaf or hard of hearing people have diverse communication preferences. We explain what recurring problems users experience, where current teleconferencing software falls short, and how to address these shortcomings. We offer specific recommendations for best practices and the experiential reasons behind them. Raja S. Kushalnagar, Christian Vogler |
ASSETS | 2 |
| 2020 | Readability of Punctuation in Automatic Subtitles
Promiti Datta, Pablo Jakubowicz, Christian Vogler, Raja S. Kushalnagar |
ICCHP (2) | 3 |
| 2020 | Hearing Systems and Accessories for People with Hearing Loss - Introduction to the Special Thematic Session
Matjaz Debevc, Christian Vogler |
ICCHP (2) | 2 |
| 2019 | Sign Language Recognition, Generation, and Translation: An Interdisciplinary PerspectiveabstractDeveloping successful sign language recognition, generation, and translation systems requires expertise in a wide range of fields, including computer vision, computer graphics, natural language processing, human-computer interaction, linguistics, and Deaf culture. Despite the need for deep interdisciplinary knowledge, existing research occurs in separate disciplinary silos, and tackles separate portions of the sign language processing pipeline. This leads to three key questions: 1) What does an interdisciplinary view of the current landscape reveal? 2) What are the biggest challenges facing the field? and 3) What are the calls to action for people working in the field? To help answer these questions, we brought together a diverse group of experts for a two-day workshop. This paper presents the results of that interdisciplinary workshop, providing key background that is often overlooked by computer scientists, a review of the state-of-the-art, a set of pressing challenges, and a call to action for the research community. Danielle Bragg, Oscar Koller, Mary Bellard, Larwan Berke, Patrick Boudreault, Annelies Braffort, Naomi Caselli, Matt Huenerfauth, Hernisa Kacorri, Tessa Verhoef, Christian Vogler, Meredith Ringel Morris |
ASSETS | 11 |
| 2019 | Voice Telephony for Individuals with Hearing Loss: The Effects of Audio Bandwidth, Bit Rate and Packet LossabstractThis paper describes three studies conducted with a total of 114 individuals with hearing loss and 12 hearing controls, with the goal of investigating the impact of audio quality parameters on the accessibility of voice telecommunications. Three categories of parameters are covered: (1) narrowband (NB) versus wideband (WB) audio; (2) encoding audio at varying bit rates, ranging from typical rates used in today's telecom networks to the highest quality supported by these audio codecs; and (3) absence of packet loss to worst-case packet loss in VoIP telephony. With WB audio, individuals with hearing loss exhibit better speech recognition, expend less perceived mental effort, and rate speech quality higher than with NB audio. Bit rate affects speech recognition for NB audio, and speech quality ratings for both NB and WB audio. Packet loss affects all of speech recognition, mental effort, and speech quality ratings. WB versus NB audio also affects hearing individuals, especially under packet loss. Linda Kozma-Spytek, Paula Tucker, Christian Vogler |
ASSETS | 3 |
| 2015 | Head-Mounted Display Visualizations to Support Sound Awareness for the Deaf and Hard of HearingabstractPersons with hearing loss use visual signals such as gestures and lip movement to interpret speech. While hearing aids and cochlear implants can improve sound recognition, they generally do not help the wearer localize sound necessary to leverage these visual cues. In this paper, we design and evaluate visualizations for spatially locating sound on a head-mounted display (HMD). To investigate this design space, we developed eight high-level visual sound feedback dimensions. For each dimension, we created 3-12 example visualizations and evaluated these as a design probe with 24 deaf and hard of hearing participants (Study 1). We then implemented a real-time proof-of-concept HMD prototype and solicited feedback from 4 new participants (Study 2). Study 1 findings reaffirm past work on challenges faced by persons with hearing loss in group conversations, provide support for the general idea of sound awareness visualizations on HMDs, and reveal preferences for specific design options. Although preliminary, Study 2 further contextualizes the design probe and uncovers directions for future work. Dhruv Jain, Leah Findlater, Jamie Gilkeson, Benjamin Holland, Ramani Duraiswami, Dmitry N. Zotkin, Christian Vogler, Jon Froehlich |
CHI | 7 |
| 2013 | Audio-visual speech understanding in simulated telephony applications by individuals with hearing lossabstractWe present a study into the effects of the addition of a video channel, video frame rate, and audio-video synchrony, on the ability of people with hearing loss to understand spoken language during video telephone conversations. Analysis indicates that higher frame rates result in a significant improvement in speech understanding, even when audio and video are not perfectly synchronized. At lower frame rates, audio-video synchrony is critical: if the audio is perceived 100 ms ahead of video, understanding drops significantly; if on the other hand the audio is perceived 100 ms behind video, understanding does not degrade versus perfect audio-video synchrony. These findings are validated in extensive statistical analysis over two within-subjects experiments with 24 and 22 participants, respectively. Linda Kozma-Spytek, Paula Tucker, Christian Vogler |
ASSETS | 3 |
| 2013 | Standardization of real-time text in instant messagingabstractWe demonstrate new standardized ways of how real-time text can be seamlessly integrated into instant messaging environments. Real-time text is text transmitted instantly while it is being typed or created. The recipient can immediately read the sender's text as it is written, without waiting. Mark Rejhon, Christian Vogler, Norman Williams, Gunnar Hellström |
ASSETS | 2 |
| 2013 | Mixed local and remote participation in teleconferences from a deaf and hard of hearing perspectiveabstractIn this experience report we describe the accessibility challenges that deaf and hard of hearing committee members faced while collaborating with a larger group of hearing committee members over a period of 2½ years. We explain what some recurring problems are, how audio-only conferences fall short even when relay services and interpreters are available, and how we devised a videoconferencing setup using FuzeMeeting to minimize the accessibility barriers. We also describe some best practices, as well as lessons learned, and pitfalls to avoid in deploying this type of setup. Christian Vogler, Paula Tucker, Norman Williams |
ASSETS | 1 |
| 2007 | The Best of Both Worlds: Combining 3D Deformable Models with Active Shape ModelsabstractReliable 3D tracking is still a difficult task. Most parametrized 3D deformable models rely on the accurate extraction of image features for updating their parameters, and are prone to failures when the underlying feature distribution assumptions are invalid. Active Shape Models (ASMs), on the other hand, are based on learning, and thus require fewer reliable local image features than parametrized 3D models, but fail easily when they encounter a situation for which they were not trained. In this paper, we develop an integrated framework that combines the strengths of both 3D deformable models and ASMs. The 3D model governs the overall shape, orientation and location, and provides the basis for statistical inference on both the image features and the parameters. The ASMs, in contrast, provide the majority of reliable 2D image features over time, and aid in recovering from drift and total occlusions. The framework dynamically selects among different ASMs to compensate for large viewpoint changes due to head rotations. This integration allows the robust tracking effaces and the estimation of both their rigid and non- rigid motions. We demonstrate the strength of the framework in experiments that include automated 3D model fitting and facial expression tracking for a variety of applications, including sign language. Christian Vogler, Atul Kanaujia, Siome Goldenstein, Dimitris N. Metaxas |
ICCV | 1 |
| 2007 | Outlier rejection in high-dimensional deformable models
Christian Vogler, Siome Goldenstein, Jorge Stolfi, Vladimir Pavlovic 0001, Dimitris N. Metaxas |
Image Vis. Comput. | 1 |
| 2007 | Human gait recognition at sagittal plane
Christian Vogler, Dimitris N. Metaxas |
Image Vis. Comput. | 2 |
| 2005 | Adaptive Deformable Models for Graphics and VisionabstractABSTRACT Deformable models are a powerful tool in both computer graphics and computer vision. The description and implementation of the deformations have to be simultaneously flexible and powerful, otherwise the technique may not satisfy the requirements of all the distinct applications. In this paper, we introduce a new method for the deformable model specification: deformable fields. Deformable fields are conceptually simple, lead to an easy implementation, and are suitable for adaptive models. We apply our new technique to describe an adaptive deformable face, and compare three different adaptation strategies. We show how our technique is suitable to describe different individuals, how to construct a model based on information from a single image, and how it allows the tracking of the deformation parameters over a video sequence. Siome Goldenstein, Christian Vogler, Luiz Velho 0001 |
Comput. Graph. Forum | 2 |
| 2004 | 3D Facial Tracking from Corrupted Movie Sequences
Siome Goldenstein, Christian Vogler, Dimitris N. Metaxas |
CVPR (1) | 2 |
| 2003 | Statistical Cue Integration in DAG Deformable ModelsabstractDeformable models are a useful modeling paradigm in computer vision. A deformable model is a curve, a surface, or a volume, whose shape, position, and orientation are controlled through a set of parameters. They can represent manufactured objects, human faces and skeletons, and even bodies of fluid. With low-level computer vision and image processing techniques, such as optical flow, we extract relevant information from images. Then, we use this information to change the parameters of the model iteratively until we find a good approximation of the object in the images. When we have multiple computer vision algorithms providing distinct sources of information (cues), we have to deal with the difficult problem of combining these, sometimes conflicting contributions in a sensible way. In this paper, we introduce the use of a directed acyclic graph (DAG) to describe the position and Jacobian of each point of deformable models. This representation is dynamic, flexible, and allows computational optimizations that would be difficult to do otherwise. We then describe a new method for statistical cue integration method for tracking deformable models that scales well with the dimension of the parameter space. We use affine forms and affine arithmetic to represent and propagate the cues and their regions of confidence. We show that we can apply the Lindeberg theorem to approximate each cue with a Gaussian distribution, and can use a maximum-likelihood estimator to integrate them. Finally, we demonstrate the technique at work in a 3D deformable face tracking system on monocular image sequences with thousands of frames. Siome Goldenstein, Christian Vogler, Dimitris N. Metaxas |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2001 | Affine Arithmetic Based Estimation of Cue Distributions in Deformable Model TrackingabstractIn this paper we describe a statistical method for the integration of an unlimited number of cues within a deformable model framework. We treat each cue as a random variable, each of which is the sum of a large number of local contributions with unknown probability distribution functions. Under the assumption that these distributions are independent, the overall distributions of the generalized cue forces can be approximated with multidimensional Gaussians, as per the central limit theorem. Estimating the covariance matrix of these Gaussian distributions, however, is difficult, because the probability distributions of the local contributions are unknown. We use affine arithmetic as a novel approach toward overcoming these difficulties. It lets us track and integrate the support of bounded distributions without having to know their actual probability distributions, and without having to make assumptions about their properties. We present a method for converting the resulting affine forms into the estimated Gaussian distributions of the generalized cue forces. This method scales well with the number of cues. We apply a Kalman filter as a maximum likelihood estimator to merge all Gaussian estimates of the cues into a single best fit Gaussian. Its mean is the deterministic result of the algorithm, and its covariance matrix provides a measure of the confidence in the result. We demonstrate in experiments how to apply this framework to improve the results of a face tracking system. Siome Goldenstein, Christian Vogler, Dimitris N. Metaxas |
CVPR (1) | 2 |
| 2001 | A Framework for Recognizing the Simultaneous Aspects of American Sign Language
Christian Vogler, Dimitris N. Metaxas |
Comput. Vis. Image Underst. | 1 |
| 1999 | Parallel Hidden Markov Models for American Sign Language RecognitionabstractThe major challenge that faces American Sign Language (ASL) recognition now is to develop methods that will scale well with increasing vocabulary size. Unlike in spoken languages, phonemes can occur simultaneously in ASL. The number of possible combinations of phonemes after enforcing linguistic constraints is approximately 5.5/spl times/10/sup 8/. Gesture recognition, which is less constrained than ASL recognition, suffers from the same problem. Thus, it is not feasible to train conventional hidden Markov models (HMMs) for large-scab ASL applications. Factorial HMMs and coupled HMMs are two extensions to HMMs that explicitly attempt to model several processes occuring in parallel. Unfortunately, they still require consideration of the combinations at training time. In this paper we present a novel approach to ASL recognition that aspires to being a solution to the scalability problems. It is based on parallel HMMs (PaHMMs), which model the parallel processes independently. Thus, they can also be trained independently, and do not require consideration of the different combinations at training time. We develop the recognition algorithm for PaHMMs and show that it runs in time polynomial in the number of states, and in time linear in the number of parallel processes. We run several experiments with a 22 sign vocabulary and demonstrate that PaHMMs can improve the robustness of HMM-based recognition even on a small scale. Thus, PaHMMs are a very promising general recognition scheme with applications in both gesture and ASL recognition. Christian Vogler, Dimitris N. Metaxas |
ICCV | 1 |
| 1998 | ASL Recognition Based on a Coupling Between HMMs and 3D Motion AnalysisabstractWe present a framework for recognizing isolated and continuous American Sign Language (ASL) sentences from three-dimensional data. The data are obtained by using physics-based three-dimensional tracking methods and then presented as input to Hidden Markov Models (HMMs) for recognition. To improve recognition performance, we model context-dependent HMMs and present a novel method of coupling three-dimensional computer vision methods and HMMs by temporally segmenting the data stream with vision methods. We then use the geometric properties of the segments to constrain the HMM framework for recognition. We show in experiments with a 53 sign vocabulary that three-dimensional features outperform two-dimensional features in recognition performance. Furthermore, we demonstrate that context-dependent modeling and the coupling of vision methods and HMMs improve the accuracy of continuous ASL recognition. Christian Vogler, Dimitris N. Metaxas |
ICCV | 1 |