Cosmin Munteanu

dblp:49/5750 · DBLP profile ↗
← Back
53ranked-venue papers
17as first author
11since 2021 · last 2026
0000-0002-0635-9124ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 40 · 11 first-author · 11 since 2021Artificial intelligence and machine learning · 11 · 6 first-authorGraphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 2Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 From Barriers to Blueprints: A Critical Systematic Review of Older Adults and Social VR
abstract
The dominant narrative in HCI positions older adults as struggling to adopt virtual reality (VR), framing resistance as user deficit. We challenge this through a critical systematic review of 85 papers (2017–2024) using resistant reading, adapted from feminist literary theory. Our analysis surfaces 284 instances documented as “adoption failures” but reinterpreted as design intelligence. We identify a “Vicious Cycle of Deficit-Focused Inquiry” whereby researchers document friction, interpret it as inadequacy, and design “solutions” reproducing deficit assumptions. This cycle operates across four dimensions: Activity (prescribed tasks vs. authentic presence), Embodiment (prosthetic correction vs. embodied dignity), Environment (spatial rescue vs. permeable boundaries), and Accessibility (independence vs. interdependence). We contribute resistant reading as method for meta-analysis, empirical evidence of institutional epistemic injustice in HCI, and four Critical Re-Orientations reframing friction as testimony about systemic inadequacy. Documented resistance constitutes a coherent critique of platforms designed for gaming rather than comfort, agency, and connection.
Sho Conte, Cosmin Munteanu, Aava Sapkota
CHI2
2024 Politics of the Past: Understanding the Role of Memory, Postmemory, and Remembrance in Navigating the History of Migrant Families
abstract
The importance of history as an HCI method has been gaining increasing attention in HCI literature. However, the mainstream historical sources (books, documentaries, etc.) and methods often risk (re)producing western colonial biases potentially providing a narrow one-sided perspective on history and detaching “sanitized facts” from people’s emotional accounts. While oral history and similar alternative methods are often used as a countermeasure, their applicability has remained underexplored in HCI, especially in a sensitive context, such as migration. We build on the rich body of social science work on collective memory to introduce a complementary way of navigating the past of the migrant families, and also reveal the corresponding challenges to advance this literature. Our interview study with 17 migrant families highlights how the politics of remembrance, family dynamics, and postmemory shape the past stories of migrant families. We discuss how these findings inform the HCI literature on migration, design, and postcolonial computing.
Nabila Chowdhury, Natasha Shokri, Cibeles Herrera Valera, Carolina Reyes Marquez, Md. Rashidujjaman Rifat, Marisol Wong-Villacres, Cosmin Munteanu, Negin Dahya, Syed Ishtiaque Ahmed
CHI8
2023 Avoiding mixed messages: research-based fact-checking the media portrayals of voice user interfaces for older adults
abstract
It is often suggested that older adults (those 60 years or older) constitute a viable target market for voice user interfaces (VUIs) and that VUIs can provide many benefits for older adults. The ma...
Jaisie Sin, Cosmin Munteanu, Dongqing Chen, Jalena G. Threatt
Hum. Comput. Interact.2
2023 An Underdeveloped Metaphor: The Mismatched Designs and Motivations of Digital Picture Interactions
abstract
Picture interactions are key to daily and long-term social connections between families and communities, especially through reminiscence. Across the nearly 200-year history of domestic photography, this social reminiscence has been accomplished largely through photo albums. However, in the now common digital setting, albums are pushed aside for the endless film roll metaphor. In this article, we explore this metaphor through the 20-year history of proposed digital picture interactions from Human–Computer Interaction research, and compare this to ongoing interactions with modern picture tools. Through these, we reveal that this prominent design metaphor does not create space for social reminiscence, but does fit the novel use of immediate sharing seen across social networking. Furthermore, the endless and ever-growing nature of digital film rolls are not meaningfully browsable for either intended use. We close by reconnecting to the past works that explore the broad potential of digital pictures.
Benett Axtell, Eleen Gong, Cosmin Munteanu
ACM Trans. Comput. Hum. Interact.3
2022 Design is Worth a Thousand Words: The Effect of Digital Interaction Design on Picture-Prompted Reminiscence
abstract
Interactions with our personal and family pictures are essential to continued social reminiscence, leading to long-term benefits, including reduced social isolation. Previous research has identified how designs of digital picture tools fall short of physical options specifically in terms of reminiscence. However, the relative prompting abilities of different digital interactions, including the types of memories prompted like external facts or person-centred memories, have not yet been explored. To investigate this, we present a controlled study of the memories prompted by three digital picture interactions (slideshow, gallery, and tabletop) on personal touchscreen devices. We find differences in how these tools and the interactions they support prompt reminiscence. In particular, gallery views prompt significantly fewer memories than either the tabletop or slideshow. Slideshows prompt significantly more external, factual memories, but not more person-centred memories, which are key to reminiscence. This has implications for the overall social usability of digital picture interactions.
Benett Axtell, Raheleh Saryazdi, Cosmin Munteanu
CHI3
2022 "Rewind to the Jiggling Meat Part": Understanding Voice Control of Instructional Videos in Everyday Tasks
abstract
Voice interaction has long been envisioned as enabling users to transform physical interaction into hands-free, such as allowing fine-grained control of instructional videos without physically disengaging from the task at hand. While significant engineering advances have brought us closer to this ideal, we do not fully understand the user requirements for voice interactions that should be supported in such contexts. This paper presents an ecologically-valid wizard-of-oz elicitation study exploring realistic user requirements for an ideal instructional video playback control while cooking. Through the analysis of the issued commands and performed actions during this non-linear and complex task, we identify (1) patterns of command formulation, (2) challenges for design, and (3) how task and voice-based commands are interwoven in real-life. We discuss implications for the design and research of voice interactions for navigating instructional videos while performing complex tasks.
Yaxi Zhao, Razan Jaber, Donald McMillan, Cosmin Munteanu
CHI4
2022 Partners in life and online search: Investigating older couples' collaborative information seeking
abstract
Older adults frequently collaborate with their spouses in daily tasks and problem solving. Despite information seeking being an important aspect of collaboration, the information seeking behaviour of older adults and in particular couples remains under investigated. To address this gap, in this paper we present a qualitative investigation of older adults’ collaborative information seeking. Through in-depth interviews and demonstrations of real-life search tasks with eleven older couples, we show that older couples frequently engage in collaborative information seeking in daily tasks, interests, and to satisfy curiosity. Our research suggests that collaborative information seeking is a relationship maintenance behaviour among older couples, and that their long-term relationships may play a key role in how they communicate, make decisions, and develop divide and conquer strategies by taking on various roles during their collaborative information seeking. We also found that older couples construct shared views toward technology adoption and usage despite their individual differences. We include some reflections on the existing collaborative information systems and how they may adapt to fit older couples’ collaborative information seeking.
Winter Wei, Cosmin Munteanu, Martin Halvey
CHIIR2
2022 "With a hint she will remember": Collaborative Storytelling and Culture Sharing between Immigrant Grandparents and Grandchildren Via Magic Thing Designs
abstract
The timeless social activity of passing down oral stories preserves family memory, identity, values, and culture. Existing tools for family memories often take a techno-determinist approach by focusing on the mechanics of connecting families and the resulting documentation, rather than the social process of sharing stories and morals, and largely without considering the specific needs of immigrant families. For immigrant families, cultural exchange, particularly crucial across grandparent and grandchild generations, is threatened by the language and cultural barriers emerging from displacement and migration. As a result, immigrant grandparents and their young grandchildren struggle with fostering social kinship, leading to social disconnect and loss of cultural heritage. In our research, we collaborate with multi-generational and culturally-at-risk immigrant families through Participatory Design activities towards the design of reminiscence tools that support their needs focusing on language and cultural connection. We report on the designs created by families and propose design guidelines supporting cultural resilience, focusing on flexible, visual storytelling.
Amna Liaqat, Benett Axtell, Cosmin Munteanu
Proc. ACM Hum. Comput. Interact.3
2021 Tea, Earl Grey, Hot: Designing Speech Interactions from the Imagined Ideal of Star Trek
abstract
Speech is now common in daily interactions with our devices, thanks to voice user interfaces (VUIs) like Alexa. Despite their seeming ubiquity, designs often do not match users’ expectations. Science fiction, which is known to influence design of new technologies, has included VUIs for decades. Star Trek: The Next Generation is a prime example of how people envisioned ideal VUIs. Understanding how current VUIs live up to Star Trek’s utopian technologies reveals mismatches between current designs and user expectations, as informed by popular fiction. Combining conversational analysis and VUI user analysis, we study voice interactions with the Enterprise’s computer and compare them to current interactions. Independent of futuristic computing power, we find key design-based differences: Star Trek interactions are brief and functional, not conversational, they are highly multimodal and context-driven, and there is often no spoken computer response. From this, we suggest paths to better align VUIs with user expectations.
Benett Axtell, Cosmin Munteanu
CHI2
2021 Digital Design Marginalization: New Perspectives on Designing Inclusive Interfaces
abstract
We conceptualize Digital Design Marginalization (DDM) as the process in which a digital interface design excludes certain users and contributes to marginalization in other areas of their lives. Due to non-inclusive designs, many underrepresented users face barriers in accessing essential services that are moving increasingly, sometimes exclusively, online – services such as personal finance, healthcare, social connectivity, and shopping. This can further perpetuate the “digital divide,” a technology-based form of social inequality that has offline consequences. We introduce the term Marginalizing Design to describe designs that contribute to DDM. In this paper, we focus on the impact of Marginalizing Design on older adults through examples from our research and discussions of services that may have marginalizing designs for older adults. Our aim is to provide a conceptual lens for designers, service providers, and policy makers through which they can use to purposely lessen or avoid digitally marginalizing groups of users.
Jaisie Sin, Rachel L. Franz, Cosmin Munteanu, Bárbara Barbosa Neves
CHI3
2021 Participatory Design for Intergenerational Culture Exchange in Immigrant Families: How Collaborative Narration and Creation Fosters Democratic Engagement
abstract
Language and cultural barriers critically threaten the social relationships between grandparents and grandchildren in immigrant families. Cultural exchange activities, like shared storytelling, can foster these crucial connections. However, existing barriers make these seemingly routine interactions challenging for families to navigate. The resulting intergenerational drift places grandparents at high risk of sustained social isolation from their families. Past works have presented technology-mediated supports for grandparent-grandchild social interactions in non-immigrant families and have found that these interventions do foster stronger connections in both physically close and distant multigenerational families. We explore how to support the specific needs of immigrant families through Magic Thing participatory design workshops with grandchildren and grandparents together in order to reveal the social interactions that would support their cultural exchange. We use the Magic Thing to move the standard dialogic grandparent-grandchild relationship into a trialogic one, creating space for comfortable social connection and storytelling through the shared creation of the design. We find that technology-mediated support of intergenerational immigrant cultural exchange must be designed for this trialogic process, consider the role of expressing values as a form of meta-commentary on a story, and shift the perspective on existing "barriers" to consider how they might foster further engagement.
Amna Liaqat, Benett Axtell, Cosmin Munteanu
Proc. ACM Hum. Comput. Interact.3
2020 Designing Voice Interfaces: Back to the (Curriculum) Basics
abstract
Voice user interfaces (VUIs) are rapidly increasing in popularity in the consumer space. This leads to a concurrent explosion of available applications for such devices, with many industries rushing to offer voice interactions for their products. This pressure is then transferred to interface designers; however, a large majority of designers have been only trained to handle the usability challenges specific to Graphical User Interfaces (GUIs). Since VUIs differ significantly in design and usability from GUIs, we investigate in this paper the extent to which current educational resources prepare designers to handle the specific challenges of VUI design. For this, we conducted a preliminary scoping scan and syllabi meta review of HCI curricula at more than twenty top international HCI departments, revealing that the current offering of VUI design training within HCI education is rather limited. Based on this, we advocate for the updating of HCI curricula to incorporate VUI design, and for the development of VUI-specific pedagogical artifacts to be included in new curricula.
Christine Murad, Cosmin Munteanu
CHI2
2020 FAB: The French Absolute Beginner Corpus for Pronunciation Training
abstract
We introduce the French Absolute Beginner (FAB) speech corpus. The corpus is intended for the development and study of Computer-Assisted Pronunciation Training (CAPT) tools for absolute beginner learners. Data were recorded during two experiments focusing on using a CAPT system in paired role-play tasks. The setting grants FAB three distinguishing features from other non-native corpora: the experimental setting is ecologically valid, closing the gap between training and deployment; it features a label set based on teacher feedback, allowing for context-sensitive CAPT; and data have been primarily collected from absolute beginners, a group often ignored. Participants did not read prompts, but instead recalled and modified dialogues that were modelled in videos. Unable to distinguish modelled words solely from viewing videos, speakers often uttered unintelligible or out-of-L2 words. The corpus is split into three partitions: one from an experiment with minimal feedback; another with explicit, word-level feedback; and a third with supplementary read-and-record data. A subset of words in the first partition has been labelled as more or less native, with inter-annotator agreement reported. In the explicit feedback partition, labels are derived from the experiment’s online feedback. The FAB corpus is scheduled to be made freely available by the end of 2020.
Sean Robertson, Cosmin Munteanu, Gerald Penn
LREC2
2020 An empirically grounded sociotechnical perspective on designing virtual agents for older adults
abstract
Autonomous, intelligent virtual agents (IVAs) are increasingly used commercially in essential information spaces such as healthcare. Existing IVA research has focused on microscale interaction patterns, for example those related to the usability of artificial intelligence systems. However, the sociotechnical patterns of users’ information practices and their relationship with the design and adoption of IVAs have been largely understudied, especially when it comes to older adults, who stand to benefit greatly from IVAs. Yet, exposing such patterns may more meaningfully relate sociotechnical considerations to users’ perceptions and attitudes toward the adoption of emerging technologies such as IVAs. We explore here the feasibility of information models in informing our understanding of how older adults may use and perceive an IVA. To do this, we relate the insights and findings from a case study of health information IVAs to the six stages of the information search process model (ISP). By doing this, we uncover sociotechnical issues pertinent to each stage of the ISP which help to better contextualize (older) users’ interaction with intelligent interfaces such as IVAs. Through this, we argue for the potential of information models to inform the design of interactive user interfaces from a sociotechnical approach.
Jaisie Sin, Cosmin Munteanu
Hum. Comput. Interact.2
2020 Leveraging Peer Support for Mature Immigrants Learning to Write in Informal Contexts
abstract
For adult newcomers to countries such as Canada, learning language is more than an academic task. Language proficiency is their gateway to long-term economic and social stability, but limited access to resources contributes to systemic inequities which disproportionately place immigrants at socioeconomic disadvantages. Many new immigrants rely heavily on informal peer-networks to pursue avenues of success within an unfamiliar and inadequate system. To explore how we could leverage such a peer-based approach to meet their needs for feedback and support when learning to write in English, we deployed a peer-based writing app with 16 participants. Post-deployment focus groups and analysis of writing artifacts reveal that the design of writing support tools should present transparent feedback from both peers and automated sources, foster community through semi-structured discussions, incorporate guided review, and scaffold affective development. We discuss how incorporating these elements into the design of community learning platforms can address the language literacy needs of diverse immigrant learners and foster more positive experiences for newcomers as they negotiate their evolving identities.
Amna Liaqat, Cosmin Munteanu
Proc. ACM Hum. Comput. Interact.2
2019 What Makes a Good Conversation?: Challenges in Designing Truly Conversational Agents
abstract
Conversational agents promise conversational interaction but fail to deliver. Efforts often emulate functional rules from human speech, without considering key characteristics that conversation must encapsulate. Given its potential in supporting long-term human-agent relationships, it is paramount that HCI focuses efforts on delivering this promise. We aim to understand what people value in conversation and how this should manifest in agents. Findings from a series of semi-structured interviews show people make a clear dichotomy between social and functional roles of conversation, emphasising the long-term dynamics of bond and trust along with the importance of context and relationship stage in the types of conversations they have. People fundamentally questioned the need for bond and common ground in agent communication, shifting to more utilitarian definitions of conversational qualities. Drawing on these findings we discuss key challenges for conversational agent design, most notably the need to redefine the design parameters for conversational agent interaction.
Leigh Clark, Nadia Pantidi, Orla Cooney, Philip R. Doyle, Diego Garaialde, Justin Edwards, Brendan Spillane, Emer Gilmartin, Christine Murad, Cosmin Munteanu, Vincent P. Wade, Benjamin R. Cowan
CHI10
2019 Mature ELLs' Perceptions Towards Automated and Peer Writing Feedback
Amna Liaqat, Gökçe Akçayir, Carrie Demmans Epp, Cosmin Munteanu
EC-TEL4
2019 Help!: I'm Stuck, and there's no F1 Key on My Tablet!
abstract
Older adults are often considered to be less frequent adopters of new technologies, in part due to increased efforts required to learn new interaction paradigms, especially if these need to overcome long-established mental models of technology use. Many current interfaces such as mobile devices often do not incorporate elements that align with older adults' models of use: explicit help menus, user manuals, navigation affordances. The lack of such reassuring elements may cause anxiety to those trying to learn interaction paradigms that are new to them. This paper details the help and support paradigms behind the design of a contextual support interface for tablet devices and describes the results of a usability evaluation with older adult participants.
Sho Conte, Cosmin Munteanu
MobileHCI2
2019 Engaging Seniors through Automatically-Generated Photo Digests from their Families' Social Media
abstract
Seniors are increasingly using the Internet. However, their adoption of available services such as social media is often restricted by their limited experience with new technologies. At the same time, there is significant interest in designing communication applications, especially mobile, that improve seniors' social connectedness. These are mostly implemented as dedicated social networking tools for seniors and their families. A barrier to the full adoption of such tools is the requirement for younger family members to actively manage a platform parallel to the social media tools they already use (e.g., Facebook). We propose PhotoDigest -- a user-centred application that allows seniors to passively engage in their families' social media activities. PhotoDigest automatically harvests families' Facebook photo posts and delivers them to seniors as weekly digests. We conducted a preliminary deployment study and show that PhotoDigest is easily adopted by seniors, does not interfere with younger generations' life routines, and enhances the entire family's social connectedness.
Yichen Dang, Cosmin Munteanu, Carrie Demmans Epp
MobileHCI2
2019 Effects of WER on ASR Correction Interfaces for Mobile Text Entry
abstract
Speech is increasingly being used as a method for text entry, especially on commercial mobile devices such as smartphones. While automatic speech recognition has seen great advances, factors like acoustic noise, differences in language or accents can affect the accuracy of speech dictation for mobile text entry. There has been some research on interfaces that enable users to intervene in the process, by correcting speech recognition errors. However, there is currently little research that investigates the effect of Automatic Speech Recognition (ASR) metrics, such as word error rate, on human performance and usability of speech recognition correction interfaces for mobile devices. This research explores how word error rates affect the usability and usefulness of touch-based speech recognition correction interfaces in the context of mobile device text entry.
Christine Murad, Cosmin Munteanu, Wolfgang Stuerzlinger
MobileHCI2
2019 An Information Behaviour-Based Approach to Virtual Doctor Design
abstract
Information behaviour models have been used extensively to explain people's interactions with information, such as in information search and user behaviour in libraries. However, we do not yet know the connection between components of information models and the interface design of digital systems, particularly when these are designed to support marginalized users such as older adults (OAs). Yet, this connection may relate to users' perceptions and subsequent adoption of emerging technologies, such as the autonomous virtual agents (VAs) functioning as advice-dispensing chatbots (increasingly present on mobile devices). We explore here the feasibility of information models in informing our understanding of how OAs may use and perceive a VA. For this, we use the information search process (ISP) model to explain the results of a case study with health information VAs and speculate on the implications of the ISP on the design of mobile-based VAs, chatbots, and voice-based interfaces.
Jaisie Sin, Cosmin Munteanu
MobileHCI2
2019 The State of Speech in HCI: Trends, Themes and Challenges
abstract
Abstract Speech interfaces are growing in popularity. Through a review of 99 research papers this work maps the trends, themes, findings and methods of empirical research on speech interfaces in the field of human–computer interaction (HCI). We find that studies are usability/theory-focused or explore wider system experiences, evaluating Wizard of Oz, prototypes or developed systems. Measuring task and interaction was common, as was using self-report questionnaires to measure concepts like usability and user attitudes. A thematic analysis of the research found that speech HCI work focuses on nine key topics: system speech production, design insight, modality comparison, experiences with interactive voice response systems, assistive technology and accessibility, user speech production, using speech technology for development, peoples’ experiences with intelligent personal assistants and how user memory affects speech interface interaction. From these insights we identify gaps and challenges in speech research, notably taking into account technological advancements, the need to develop theories of speech interface interaction, grow critical mass in this domain, increase design work and expand research from single to multiple user interaction contexts so as to reflect current use contexts. We also highlight the need to improve measure reliability, validity and consistency, in the wild deployment and reduce barriers to building fully functional speech interfaces for research. RESEARCH HIGHLIGHTS Most papers focused on usability/theory-based or wider system experience research with a focus on Wizard of Oz and developed systems Questionnaires on usability and user attitudes often used but few were reliable or validated Thematic analysis showed nine primary research topics Challenges identified in theoretical approaches and design guidelines, engaging with technological advances, multiple user and in the wild contexts, critical research mass and barriers to building speech interfaces
Leigh Clark, Philip R. Doyle, Diego Garaialde, Emer Gilmartin, Stephan Schlögl, Jens Edlund, Matthew P. Aylett, João P. Cabral, Cosmin Munteanu, Justin Edwards, Benjamin R. Cowan
Interact. Comput.9
2018 Designing Pronunciation Learning Tools: The Case for Interactivity against Over-Engineering
abstract
Paired role-play is a common collaborative activity in language learning classrooms, adding meaning and cultural context to the learning process. This is complemented by teachers' immediate and explicit feedback. Interactive tools that provide explicit feedback during collaborative learning are scarce, however. More commonly, supporting dialogue practice takes the form of computer-aided single-student read-and-record activities. This limitation is partly due to the complexity of processing language learners' speech in unconstrained tasks. In this paper, we assess the value of pronunciation error detection algorithms within a realistic, software-aided, paired role-playing task with beginning learners of French. We found that students' pronunciations improve regardless of the type of error detector employed -- even for those using simple heuristics. We suggest that speech technologies for language learning have been too focused on engineering goals. Instead, new interactive designs supporting collaboration may be used to overcome engineering limitations and properly support students' engagement.
Sean Robertson, Cosmin Munteanu, Gerald Penn
CHI2
2018 Understanding Older Users' Acceptance of Wearable Interfaces for Sensor-based Fall Risk Assessment
abstract
Algorithms processing data from wearable sensors promise to more accurately predict risks of falling -- a significant concern for older adults. Substantial engineering work is dedicated to increasing the prediction accuracy of these algorithms; yet fewer efforts are dedicated to better engaging users through interactive visualizations in decision-making using these data. We present an investigation of the acceptance of a sensor-based fall risk assessment wearable device. A participatory design was employed to develop a mobile interface providing visualizations of sensor data and algorithmic assessments of fall risks. We then investigated the acceptance of this interface and its potential to motivate behavioural changes through a field deployment, which suggested that the interface and its belt-mounted wearable sensors are perceived as usable. We also found that providing contextual information for fall risk estimation combined with relevant practical fall prevention instructions may facilitate the acceptance of such technologies, potentially leading to behaviour change.
Alan Yusheng Wu, Cosmin Munteanu
CHI2
2018 Session details: Session H3: Edgy, Airy Interaction
Cosmin Munteanu
Graphics Interface1
2018 Touch-Supported Voice Recording to Facilitate Forced Alignment of Text and Speech in an E-Reading Interface
abstract
Reading a book together with a family member who has impaired vision or other difficulties reading is an important social bonding activity. However, for the person being read to, there is little support in making these experiences repeatable. While audio can easily be recorded, synchronizing it with the text for later playback requires the use of forced alignment algorithms, which do not perform well on amateur read-aloud speech. We propose a human-in-the-loop approach to augmenting such algorithms, in the form of touch metaphors during collocated read-aloud sessions using tablet e-readers. The metaphor is implemented as a finger-follows-text tracker. We explore how this could better handle the variability of amateur reading, which poses accuracy challenges for existing forced alignment techniques. Data collected from users reading aloud as assisted by touch metaphors show increases in the accuracy of forced alignment algorithms and reveal opportunities for how to better support reading aloud.
Benett Axtell, Cosmin Munteanu, Carrie Demmans Epp, Yomna Aly, Frank Rudzicz
IUI2
2018 Towards a writing analytics framework for adult english language learners
abstract
Improving the written literacy of newcomers to English-speaking countries can lead to better education, employment, or social integration opportunities. However, this remains a challenge in traditional classrooms where providing frequent, timely, and personalized feedback is not always possible. Analytics can scaffold the writing development of English Language Learners (ELLs) by providing such feedback. To design these analytics, we conducted a field study analyzing essay samples from immigrant adult ELLs (a group often overlooked in writing analytics research) and identifying their epistemic beliefs and learning motivations. We identified common themes across individual learner differences and patterns of errors in the writing samples. The study revealed strong associations between epistemic writing beliefs and learning strategies. The results are used to develop guidelines for designing writing analytics for adult ELLs, and to propose ideas for analytics that scaffold writing development for this group.
Amna Liaqat, Cosmin Munteanu
LAK2
2017 Using frame of mind: documenting reminiscence through unstructured digital picture interaction
abstract
Mobile technologies have made family photo collections extremely portable. People can now carry all their pictures with them wherever they go and show them to others in any setting with smartphones or tablets. However, current options for portable photo viewing are not intended for in-person sharing and reminiscence. Frame of Mind presents a new way to interact with digital pictures on a touch screen that encourages storytelling through its free-flowing interaction using the metaphor of looking at pictures on a table top. This allows family reminiscence to be lightweight, portable, and more accessible by supporting photo viewing on tablets that can have access to complete picture collections. So Frame of Mind moves towards digital tools that support our current photo viewing and sharing activities.
Benett Axtell, Cosmin Munteanu
MobileHCI2
2017 Finger tracking: facilitating non-commercial content production for mobile e-reading applications
abstract
Limited literacy and visual impairment reduce the ability of many to read on their own. Current e-reader solutions rely on either unnatural synthetic voices or professionally produced audio e-books. Neither provide the same enjoyment as having a family member read to a user, especially when the user requires assistive reading (following printed text while listening to it being read). Unfortunately, the support for non-commercial production of such e-books is limited and requires significant effort. We evaluate a novel, assistive mobile interaction technique that facilitates the recording of audio e-books and their synchronization with the read text. We show that a technique based on a finger tracking metaphor provides optimal support with respect to reading speed. These human-in-the-loop, adaptive techniques can now be used to reduce the content-creation burden that is associated with supporting those who cannot read on their own.
Carrie Demmans Epp, Cosmin Munteanu, Benett Axtell, Keerthika Ravinthiran, Yomna Aly, Elman Mansimov
MobileHCI2
2017 Speech and Hands-free interaction: myths, challenges, and opportunities
abstract
HCI research has for long been dedicated to better and more naturally facilitating information transfer between humans and machines. Unfortunately, humans' most natural form of communication, speech, is also one of the most difficult modalities to be understood by machines - despite, and perhaps, because it is the highest-bandwidth communication channel we possess. While significant research efforts, from engineering, to linguistic, and to cognitive sciences, have been spent on improving machines' ability to understand speech, the MobileHCI community (and the HCI field at large) has been relatively timid in embracing this modality as a central focus of research. This can be attributed in part to the unexpected variations in error rates when processing speech, in contrast with often-unfounded claims of success from industry, but also to the intrinsic difficulty of designing and especially evaluating speech and natural language interfaces. As such, the development of interactive speech-based systems is mostly driven by engineering efforts to improve such systems with respect to largely arbitrary performance metrics. Such developments have often been void of any user-centered design principles or consideration for usability or usefulness.
Cosmin Munteanu, Gerald Penn
MobileHCI1
2016 Pronunciation Error Detection for New Language Learners
Sean Robertson, Cosmin Munteanu, Gerald Penn
INTERSPEECH2
2015 Situational Ethics: Re-thinking Approaches to Formal Ethics Requirements for Human-Computer Interaction
abstract
Most Human-Computer Interaction (HCI) researchers are accustomed to the process of formal ethics review for their evaluation or field trial protocol. Although this process varies by country, the underlying principles are universal. While this process is often a formality, for field research or lab-based studies with vulnerable users, formal ethics requirements can be challenging to navigate -- a common occurrence in the social sciences; yet, in many cases, foreign to HCI researchers. Nevertheless, with the increase in new areas of research such as mobile technologies for marginalized populations or assistive technologies, this is a current reality. In this paper we present our experiences and challenges in conducting several studies that evaluate interactive systems in difficult settings, from the perspective of the ethics process. Based on these, we draft recommendations for mitigating the effect of such challenges to the ethical conduct of research. We then issue a call for interaction researchers, together with policy makers, to refine existing ethics guidelines and protocols in order to more accurately capture the particularities of such field-based evaluations, qualitative studies, challenging lab-based evaluations, and ethnographic observations.
Cosmin Munteanu, Heather Molyneaux, Wendy Moncur, Mario Romero, Susan O'Donnell, John Vines
CHI1
2015 "My Hand Doesn't Listen to Me!": Adoption and Evaluation of a Communication Technology for the 'Oldest Old'
abstract
Adoption and use of novel technology by the institutionalized 'oldest old' (80+) is understudied. This population is the fastest growing demographic group in developed countries, providing design opportunities and challenges for HCI. Since the recruitment of oldest old people is challenging, research tends to focus on older adults (65+) and their use of and attitudes towards existing communication technologies, or on their caregivers and social ties. Our study deployed a novel communication appliance among five frail oldest old people living in a long-term care facility, which included field observations and usability and accessibility tests. Our findings suggest factors that facilitate and hinder the adoption of communication technologies, such as social, attitudinal, digital literacy, physical, and usability. We also discuss issues that arise in studying technology adoption by the oldest old, including usability and accessibility testing, and suggest solutions that may be helpful to HCI researchers working with this population.
Bárbara Barbosa Neves, Rachel L. Franz, Cosmin Munteanu, Ronald Baecker, Mags Ngo
CHI3
2015 Speech-based Interaction: Myths, Challenges, and Opportunities
abstract
HCI research has for long been dedicated to better and more naturally facilitating information transfer between humans and machines. Unfortunately, humans' most natural form of communication, speech, is also one of the most difficult modalities to be understood by machines -- despite, and perhaps, because it is the highest-bandwidth communication channel we possess. While significant research efforts, from engineering, to linguistic, and to cognitive sciences, have been spent on improving machines' ability to understand speech, the HCI community has been relatively timid in embracing this modality as a central focus of research. This can be attributed in part to the relatively discouraging levels of accuracy in understanding speech, in contrast with often-unfounded claims of success from industry, but also to the intrinsic difficulty of designing and especially evaluating speech and natural language interfaces.
Cosmin Munteanu, Gerald Penn
IUI1
2014 Speech-based interaction: myths, challenges, and opportunities
abstract
Human-Computer Interaction (HCI) research has for long been dedicated to better and more naturally facilitating information transfer between humans and machines. Unfortunately, humans' most natural form of communication, speech, is also one of the most difficult modalities to be understood by machines. This is largely due to speech being the highest-bandwidth communication channel we possess. As such, significant research efforts, from engineering, to linguistic, and to cognitive sciences, have been spent during the past several decades on improving machines' ability to understand speech. Yet, the MobileHCI community (and HCI in general) has been relatively timid in embracing this modality as a central focus of research. This can be attributed in part to the relatively discouraging levels of accuracy in understanding speech, in contrast with often-unfounded claims of success from industry, but also to the intrinsic difficulty of designing and especially evaluating speech and natural language interfaces.
Cosmin Munteanu, Gerald Penn
Mobile HCI1
2014 Hidden in plain sight: low-literacy adults in a developed country overcoming social and educational challenges through mobile learning support tools
Cosmin Munteanu, Heather Molyneaux, Julie Maitland, Daniel McDonald, Rock Leung, Hélène Fournier, Joanna Lumsden
Pers. Ubiquitous Comput.1
2013 Automatic human utility evaluation of ASR systems: does WER really predict performance?
abstract
International audience
Benoît Favre, Kyla Cheung, Siavash Kazemian, Adam Lee, Yang Liu 0004, Cosmin Munteanu, Ani Nenkova, Dennis Ochei, Gerald Penn, Stephen Tratz, Clare R. Voss, Frauke Zeller
INTERSPEECH6
2013 An accessible, large-print, listening and talking e-book to support families reading together
abstract
Reading is an activity that is not only informative or pleasurable, but can have significant social benefits. Especially in a family setting, it is part of the interaction between children and their parents, it helps create a bond between children and their grandparents, and even bring adults and their older parents closer. However, with families increasingly living or spending time in different locations or managing busy schedules that afford very little time together, the social opportunities enabled by reading are often lost. Furthermore, reading can be a challenge for older adults or for those with impaired eyesight. To address these problems, we are proposing ALLT -- an Accessible, Large-Print, Listening and Talking e-book. ALLT is a tablet-based e-reading application that enhances the capabilities of e-book readers through customizable and intelligent accessibility features. It provides support for asynchronous "reading together" by synchronizing the audio recording of one user with the text that is later read by another user. This addresses the needs of a variety of users, from visually impaired adults reading together with a loved one, to children being able to replay an interactive story previously read together with their grandparents. In this demo paper we present ALLT's features and detail how they support asynchronously reading together.
Abbas Attarwala, Cosmin Munteanu, Ronald Baecker
Mobile HCI2
2012 Ecological validity and the evaluation of speech summarization quality
abstract
There is little evidence of widespread adoption of speech summarization systems. This may be due in part to the fact that the natural language heuristics used to generate summaries are often optimized with respect to a class of evaluation measures that, while computationally and experimentally inexpensive, rely on subjectively selected gold standards against which automatically generated summaries are scored. This evaluation protocol does not take into account the usefulness of a summary in assisting the listener in achieving his or her goal. In this paper we study how current measures and methods for evaluating summarization systems compare to human-centric evaluation criteria. For this, we have designed and conducted an ecologically valid evaluation that determines the value of a summary when embedded in a task, rather than how closely a summary resembles a gold standard. The results of our evaluation demonstrate that in the domain of lecture summarization, the well-known baseline of maximal marginal relevance [1] is statistically significantly worse than human-generated extractive summaries, and even worse than having no summary at all in a simple quiz-taking task. Priming seems to have no statistically significant effect on the usefulness of the human summaries. This is interesting because priming had been proposed as a technique for increasing kappa scores and/or maintaining goal orientation among summary authors. In addition, our results suggest that ROUGE scores, regardless of whether they are derived from numerically-ranked reference data or ecologically valid human-extracted summaries, may not always be reliable as inexpensive proxies for task-embedded evaluations. In fact, under some conditions, relying exclusively on ROUGE may lead to scoring human-generated summaries very favourably even when a task-embedded score calls their usefulness into question relative to using no summaries at all.
Anthony McCallum, Gerald Penn, Cosmin Munteanu, Xiaodan Zhu 0001
SLT3
2011 "Showing off" your mobile device: adult literacy learning in the classroom and beyond
abstract
For a very large number of adults, tasks such as reading. understanding, and using everyday items are a challenge. Although many community-based organizations offer resources and support for adults with limited literacy skills. current programs have difficulty reaching and retaining those that would benefit most. In this paper we present the findings of an exploratory study aimed at investigating how a technological solution that addresses these challenges is received and adopted by adult learners. For this, we have developed a mobile application to support literacy programs and to assist low-literacy adults in today's information-centric society. ALEX© (Adult Literacy support application for Experiential learning) is a mobile language assistant that is designed to be used both in the classroom and in daily life in order to help low-literacy adults become increasingly literate and independent. Through a long-term study with adult learners we show that such a solution complements literacy programs by increasing users' motivation and interest in learning, and raising their confidence levels both in their education pursuits and in facing the challenges of their daily lives.
Cosmin Munteanu, Heather Molyneaux, Daniel McDonald, Joanna Lumsden, Rock Leung, Hélène Fournier, Julie Maitland
Mobile HCI1
2010 I-smooth for improved minimum classification error training
abstract
Increasing the generalization capability of Discriminative Training (DT) of Hidden Markov Models (HMM) has recently gained an increased interest within the speech recognition field. In particular, achieving such increases with only minor modifications to the existing DT method is of significant practical importance. In this paper, we propose a solution for increasing the generalization capability of a widely-used training method - the Minimum Classification Error (MCE) training of HMM - with limited changes to its original framework. For this, we define boundary data - obtained by applying a large steep parameter, and confusion data - obtained by applying a small steep parameter on the training samples, and then do a soft interpolation between these according to the number points of occupancies of boundary data and the number points ratio between the boundary and the confusion occupancies. The final HMM parameters are then tuned in the same manner as in MCE by using the interpolated boundary data. We show that the proposed method achieves lower error rates than a standard HMM training framework on a phoneme classification task for the TIMIT speech corpus.
Haozheng Li, Cosmin Munteanu
ICASSP2
2010 ALEX: mobile language assistant for low-literacy adults
abstract
Basic literacy skills are fundamental building blocks of education, yet for a very large number of adults tasks such as understanding and using everyday items is a challenge. While research, industry, and policy-making is looking at improving access to textual information for low-literacy adults, the literacy-based demands of today's society are continually increasing. Although many community-based organizations offer resources and support to adults with limited literacy skills, current programs have difficulties reaching and retaining those that would benefit most from them. To address these challenges, the National Research Council of Canada is proposing a technological solution to support literacy programs and to assist low-literacy adults in today's information-centric society: ALEX© - Adult Literacy support application for EXperiential learning. ALEX© has been created together with low-literacy adults, following guidelines for inclusive design of mobile assistive tools. It is a mobile language assistant that is designed to be used both in the classroom and in daily life, in order to help low-literacy adults become increasingly literate and independent.
Cosmin Munteanu, Joanna Lumsden, Hélène Fournier, Rock Leung, Danny D'Amours, Daniel McDonald, Julie Maitland
Mobile HCI1
2009 Improving Automatic Speech Recognition for Lectures through Transformation-based Rules Learned from Minimal Data
Cosmin Munteanu, Gerald Penn, Xiaodan Zhu 0001
ACL/IJCNLP1
2008 Collaborative editing for improved usefulness and usability of transcript-enhanced webcasts
abstract
One challenge in facilitating skimming or browsing through archives of on-line recordings of webcast lectures is the lack of text transcripts of the recorded lecture. Ideally, transcripts would be obtainable through Automatic Speech Recognition (ASR). However, current ASR systems can only deliver, in realistic lecture conditions, a Word Error Rate of around 45% -- above the accepted threshold of 25%. In this paper, we present the iterative design of a webcast extension that engages users to collaborate in a wiki-like manner on editing the ASR-produced imperfect transcripts, and show that this is a feasible solution for improving the quality of lecture transcripts. We also present the findings of a field study carried out in a real lecture environment investigating how students use and edit the transcripts.
Cosmin Munteanu, Ronald Baecker, Gerald Penn
CHI1
2008 Using latent Dirichlet allocation to incorporate domain knowledge for topic transition detection
abstract
This paper studies automatic detection of topic transitions for recorded presentations. This can be achieved by matching slide content with presentation transcripts directly with some similarity metrics. Such literal matching, however, misses domain-specific knowledge and is sensitive to speech recognition errors. In this paper, we incorporate relevant written materials, e.g., textbooks for lectures, which convey semantic relationships, in particular domain-specific relationships, between words. To this end, we train latent Dirichlet allocation (LDA) models on these materials and measure the similarity between slides and transcripts in the acquired hidden-topic space. This similarity is then combined with literal matchings. Experiments show that the proposed approach reduces the errors in slide transition detection by 17-41 % on manual transcripts and 27-37% on automatic transcripts. Index Terms: slides transition detection, boundary detection. 1.
Xiaodan Zhu 0001, Xuming He 0001, Cosmin Munteanu, Gerald Penn
INTERSPEECH3
2007 Web-based language modelling for automatic lecture transcription
abstract
Universities have long relied on written text to share knowledge. As more lectures are made available on-line, these must be accompanied by textual transcripts in order to provide the same access to information as textbooks. While Automatic Speech Recognition (ASR) is a cost-effective method to deliver transcriptions, its accuracy for lectures is not yet satisfactory. One approach for improving lecture ASR is to build smaller, topic-dependent Language Models (LMs) and combine them (through LM interpolation or hypothesis space combination) with general-purpose, large-vocabulary LMs. In this paper, we propose a simple solution for lecture ASR with similar or better Word Error Rate reductions (as well as topic-specific keyword identification accuracies) than combination-based approaches. Our method eliminates the need for two types of LMs by exploiting the lecture slides to collect a web corpus appropriate for modelling both the conversational and the topic-specific styles of lectures. Index Terms: speech recognition, language modelling, corpus building, topic dependent, lecture transcription.
Cosmin Munteanu, Gerald Penn, Ronald Baecker
INTERSPEECH1
2006 The effect of speech recognition accuracy rates on the usefulness and usability of webcast archives
abstract
The widespread availability of broadband connections has led to an increase in the use of Internet broadcasting (webcasting). Most webcasts are archived and accessed numerous times retrospectively. In the absence of transcripts of what was said, users have difficulty searching and scanning for specific topics. This research investigates user needs for transcription accuracy in webcast archives, and measures how the quality of transcripts affects user performance in a question-answering task, and how quality affects overall user experience. We tested 48 subjects in a within-subjects design under 4 conditions: perfect transcripts, transcripts with 25% Word Error Rate (WER), transcripts with 45% WER, and no transcript. Our data reveals that speech recognition accuracy linearly influences both user performance and experience, shows that transcripts with 45% WER are unsatisfactory, and suggests that transcripts having a WER of 25% or less would be useful and usable in webcast archives.
Cosmin Munteanu, Ronald Baecker, Gerald Penn, Elaine Toms, David James
CHI1
2006 Automatic speech recognition for webcasts: how good is good enough and what to do when it isn't
abstract
The increased availability of broadband connections has recently led to an increase in the use of Internet broadcasting (webcasting). Most webcasts are archived and accessed numerous times retrospectively. One challenge to skimming and browsing through such archives is the lack of text transcripts of the webcast's audio channel. This paper describes a procedure for prototyping an Automatic Speech Recognition (ASR) system that generates realistic transcripts of any desired Word Error Rate (WER), thus overcoming the drawbacks of both prototype-based and Wizard of Oz simulations. We used such a system in a user study showing that transcripts with WERs less than 25% are acceptable for use in webcast archives. As current ASR systems can only deliver, in realistic conditions, Word Error Rates (WERs) of around 45%, we also describe a solution for reducing the WER of such transcripts by engaging users to collaborate in a wiki fashion on editing the imperfect transcripts obtained through ASR.
Cosmin Munteanu, Gerald Penn, Ronald Baecker, Yuecheng Zhang
ICMI1
2006 Measuring the acceptable word error rate of machine-generated webcast transcripts
abstract
The increased availability of broadband connections has recently led to an increase in the use of Internet broadcasting (webcasting). Most webcasts are archived and accessed numerous times retrospectively. One of the hurdles users face when browsing and skimming through archives is the lack of text transcripts of the audio channel of the webcast archive. In this paper, we proposed a procedure for prototyping an Automatic Speech Recognition (ASR) system that generates realistic transcripts of any desired Word Error Rate (WER), thus overcoming the drawbacks of both prototypebased and Wizard of Oz simulations. We used such a system in a study where human subjects perform question-answering tasks using archives of webcast lectures, and showed that their performance and perception of transcript quality is linearly affected by WER, and that transcripts of WER equal or less than 25 % would be acceptable for use in webcast archives.
Cosmin Munteanu, Gerald Penn, Ronald Baecker, Elaine Toms, David James
INTERSPEECH1
2004 Optimizing Typed Feature Structure Grammar Parsing through Non-Statistical Indexing
abstract
This paper introduces an indexing method based on static analysis of grammar rules and type signatures for typed feature structure grammars (TFSGs). The static analysis tries to predict at compile-time which feature paths will cause unification failure during parsing at run-time. To support the static analysis, we introduce a new classification of the instances of variables used in TFSGs, based on what type of structure sharing they create. The indexing actions that can be performed during parsing are also enumerated. Non-statistical indexing has the advantage of not requiring training, and, as the evaluation using large-scale HPSGs demonstrates, the improvements are comparable with those of statistical optimizations. Such statistical optimizations rely on data collected during training, and their performance does not always compensate for the training costs.
Cosmin Munteanu, Gerald Penn
ACL1
2003 A Tabulation-Based Parsing Method that Reduces Copying
abstract
This paper presents a new bottom-up chart parsing algorithm for Prolog along with a compilation procedure that reduces the amount of copying at run-time to a constant number (2) per edge. It has applications to unification-based grammars with very large partially ordered categories, in which copying is expensive, and can facilitate the use of more sophisticated indexing strategies for retrieving such categories that may otherwise be overwhelmed by the cost of such copying. It also provides a new perspective on "quick-checking" and related heuristics, which seems to confirm that forcing an early failure (as opposed to seeking an early guarantee of success) is in fact the best approach to use. A preliminary empirical evaluation of its performance is also provided.
Gerald Penn, Cosmin Munteanu
ACL2
2003 Indexing methods for efficient parsing
Cosmin Munteanu
HLT-NAACL1
2000 MDWOZ: A Wizard of Oz Environment for Dialog Systems Development
Cosmin Munteanu, Marian Boldea
LREC1