EDBT 2026 Demo / reviewers in the wild / expert
Julia Hirschberg
dblp:h/JuliaHirschberg
· DBLP profile ↗
180ranked-venue papers
21as first author
28since 2021 · last 2026
0000-0003-0689-7616ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 151 · 18 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 113 · 14 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Disentangling Annotator Skill from Verifier Strictness in Cross-Verified Dialogue AnnotationabstractSome dialogue corpus projects use a verify-after-annotation workflow: a second team member reviews a submitted file and records corrections. The resulting correction count mixes two signals, annotator accuracy and verifier strictness. We separate these signals for RASwDA, an audio-anchored re-alignment of 1,045 Switchboard file sides (105,005 corrections across 977 change logs) produced by five team members in 2024-2025. Verification is crossed: each of three identified annotators was checked by four different verifiers, with overlap in both directions. A cross-classified mixed-effects model on per-file corrections-per-interval assigns 10.8% of the variance to annotator identity, while the verifier random effect collapses to zero (singular fit, stable across seven leave-one-out and response-choice refits). Thus, for this boundary-realignment task, we find no detectable verifier identity effect once annotator identity, batch, and file length are controlled. Boundary placement, not label selection, accounts for 58.5% of corrections corpus-wide. This helps explain why verifier-specific strictness has little room to appear: timestamp adjustments are anchored in the audio, while DA-label changes account for only 3.0% of corrections. We release the action-typed change logs so other projects can run the same annotator-verifier decomposition on their own verification data. Zihao Tao, John A Prado, Ignazio Steven LaManna, Ryan Puterbaugh, Mim Datta, Julia Hirschberg |
SIGDIAL | 6 |
| 2025 | Pragmatics in the Era of Large Language Models: A Survey on Datasets, Evaluation, Opportunities and ChallengesabstractUnderstanding pragmatics—the use of language in context—is crucial for developing NLP systems capable of interpreting nuanced language use. Despite recent advances in language technologies, including large language models, evaluating their ability to handle pragmatic phenomena such as implicatures and references remains challenging. To advance pragmatic abilities in models, it is essential to understand current evaluation trends and identify existing limitations. In this survey, we provide a comprehensive review of resources designed for evaluating pragmatic capabilities in NLP, categorizing datasets by the pragmatic phenomena they address. We analyze task designs, data collection methods, evaluation approaches, and their relevance to real-world applications. By examining these resources in the context of modern language models, we highlight emerging trends, challenges, and gaps in existing benchmarks. Our survey aims to clarify the landscape of pragmatic evaluation and guide the development of more comprehensive and targeted benchmarks, ultimately contributing to more nuanced and context-aware NLP models. Bolei Ma, Wei Zhou 0067, Ziwei Gong, Yang Janet Liu, Katja Jasinskaja, Annemarie Friedrich, Julia Hirschberg, Frauke Kreuter, Barbara Plank |
ACL (1) | 8 |
| 2025 | NovAScore: A New Automated Metric for Evaluating Document Level NoveltyabstractThe rapid expansion of online content has intensified the issue of information redundancy, underscoring the need for solutions that can identify genuinely new information. Despite this challenge, the research community has seen a decline in focus on novelty detection, particularly with the rise of large language models (LLMs). Additionally, previous approaches have relied heavily on human annotation, which is time-consuming, costly, and particularly challenging when annotators must compare a target document against a vast number of historical documents. In this work, we introduce NovAScore (Novelty Evaluation in Atomicity Score), an automated metric for evaluating document-level novelty. NovAScore aggregates the novelty and salience scores of atomic information, providing high interpretability and a detailed analysis of a document’s novelty. With its dynamic weight adjustment scheme, NovAScore offers enhanced flexibility and an additional dimension to assess both the novelty level and the importance of information within a document. Our experiments show that NovAScore strongly correlates with human judgments of novelty, achieving a 0.626 Point-Biserial correlation on the TAP-DLND 1.0 dataset and a 0.920 Pearson correlation on an internal human-annotated dataset. Lin Ai, Ziwei Gong, Harshsaiprasad Deshpande, Alexander Johnson, Emmy Phung, Ahmad Emami, Julia Hirschberg |
COLING | 7 |
| 2025 | PropaInsight: Toward Deeper Understanding of Propaganda in Terms of Techniques, Appeals, and IntentabstractPropaganda plays a critical role in shaping public opinion and fueling disinformation. While existing research primarily focuses on identifying propaganda techniques, it lacks the ability to capture the broader motives and the impacts of such content. To address these challenges, we introduce PropaInsight, a conceptual framework grounded in foundational social science research, which systematically dissects propaganda into techniques, arousal appeals, and underlying intent. PropaInsight offers a more granular understanding of how propaganda operates across different contexts. Additionally, we present PropaGaze, a novel dataset that combines human-annotated data with high-quality synthetic data generated through a meticulously designed pipeline. Our experiments show that off-the-shelf LLMs struggle with propaganda analysis, but PropaGaze significantly improves performance. Fine-tuned Llama-7B-Chat achieves 203.4% higher text span IoU in technique identification and 66.2% higher BertScore in appeal analysis compared to 1-shot GPT-4-Turbo. Moreover, PropaGaze complements limited human-annotated data in data-sparse and cross-domain scenarios, demonstrating its potential for comprehensive and generalizable propaganda analysis. Jiateng Liu, Lin Ai, Zizhou Liu, Payam Karisani, Zheng Hui, Yi R. Fung 0001, Preslav Nakov, Julia Hirschberg, Heng Ji 0001 |
COLING | 8 |
| 2025 | Discourse-Driven Code-Switching: Analyzing the Role of Content and Communicative Function in Spanish-English Bilingual SpeechabstractCode-switching (CSW) is commonly observed among bilingual speakers, and is motivated by various paralinguistic, syntactic, and morphological aspects of conversation.We build on prior work by asking: how do discourse-level aspects of dialogue -i.e. the content and function of speech -influence patterns of CSW?To answer this, we analyze the named entities and dialogue acts present in a Spanish-English spontaneous speech corpus, and build a predictive model of CSW based on our statistical findings.We show that discourse content and function interact with patterns of CSW to varying degrees, with a stronger influence from function overall.Our work is the first to take a discourse-sensitive approach to understanding the pragmatic and referential cues of bilingual speech and has potential applications in improving the prediction, recognition, and synthesis of code-switched speech that is grounded in authentic aspects of multilingual discourse. Debasmita Bhattacharya, Juan Junco, Divya Tadimeti, Julia Hirschberg |
EMNLP | 4 |
| 2025 | Does Context Matter? A Prosodic Comparison of English and Spanish in Monolingual and Multilingual Discourse SettingsabstractDifferent languages are known to have typical and distinctive prosodic profiles.However, the majority of work on prosody across languages has been restricted to monolingual discourse contexts.We build on prior studies by asking: how does the nature of the discourse context influence variations in the prosody of monolingual speech?To answer this question, we compare the prosody of spontaneous, conversational monolingual English and Spanish both in monolingual and in multilingual speech settings.For both languages, we find that monolingual speech produced in a monolingual context is prosodically different from that produced in a multilingual context, with more marked differences having increased proximity to multilingual discourse.Our work is the first to incorporate multilingual discourse contexts into the study of native-level monolingual prosody, and has potential downstream applications for the recognition and synthesis of multilingual speech. Debasmita Bhattacharya, David Sasu, Michela Marchini, Natalie Schluter, Julia Hirschberg |
EMNLP | 5 |
| 2025 | Read to Hear: A Zero-Shot Pronunciation Assessment Using Textual Descriptions and LLMsabstractAutomatic pronunciation assessment is typically performed by acoustic models trained on audio-score pairs.Although effective, these systems provide only numerical scores, without the information needed to help learners understand their errors.Meanwhile, large language models (LLMs) have proven effective in supporting language learning, but their potential for assessing pronunciation remains unexplored.In this work, we introduce TextPA, a zero-shot, Textual description-based Pronunciation Assessment approach.TextPA utilizes human-readable representations of speech signals, which are fed into an LLM to assess pronunciation accuracy and fluency, while also providing reasoning behind the assigned scores.Finally, a phoneme sequence match scoring method is used to refine the accuracy scores.Our work highlights a previously overlooked direction for pronunciation assessment.Instead of relying on supervised training with audioscore examples, we exploit the rich pronunciation knowledge embedded in written text.Experimental results show that our approach is both cost-efficient and competitive in performance.Furthermore, TextPA significantly improves the performance of conventional audioscore-trained models on out-of-domain data by offering a complementary perspective. Yuwen Chen 0006, Melody Ma, Julia Hirschberg |
EMNLP | 3 |
| 2025 | From Context to Code-switching: Examining the Interplay of Language Proficiency and Multilingualism in Speech
Debasmita Bhattacharya, Aanya Tolat, Julia Hirschberg |
INTERSPEECH | 3 |
| 2025 | Comparison-Based Automatic Evaluation for Meeting Summarization
Ziwei Gong, Lin Ai, Harsh Deshpande, Alexander Johnson, Emmy Phung, Zehui Wu, Ahmad Emami, Julia Hirschberg |
INTERSPEECH | 8 |
| 2025 | Learning More with Less: Self-Supervised Approaches forLow-Resource Speech Emotion Recognition
Ziwei Gong, Pengyuan Shi, Kaan Donbekci, Lin Ai, Run Chen, David Sasu, Zehui Wu, Julia Hirschberg |
INTERSPEECH | 8 |
| 2025 | PAPILLON: Privacy Preservation from Internet-based and Local Language Model EnsemblesabstractSiyan Li, Vethavikashini Chithrra Raghuram, Omar Khattab, Julia Hirschberg, Zhou Yu. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Siyan Li, Vethavikashini Chithrra Raghuram, Omar Khattab, Julia Hirschberg |
NAACL (Long Papers) | 4 |
| 2025 | Beyond electronic health record data: leveraging natural language processing and machine learning to uncover cognitive insights from patient-nurse verbal communicationsabstractBACKGROUND: Mild cognitive impairment and early-stage dementia significantly impact healthcare utilization and costs, yet more than half of affected patients remain underdiagnosed. This study leverages audio-recorded patient-nurse verbal communication in home healthcare settings to develop an artificial intelligence-based screening tool for early detection of cognitive decline. OBJECTIVE: To develop a speech processing algorithm using routine patient-nurse verbal communication and evaluate its performance when combined with electronic health record (EHR) data in detecting early signs of cognitive decline. METHOD: We analyzed 125 audio-recorded patient-nurse verbal communication for 47 patients from a major home healthcare agency in New York City. Out of 47 patients, 19 experienced symptoms associated with the onset of cognitive decline. A natural language processing algorithm was developed to extract domain-specific linguistic and interaction features from these recordings. The algorithm's performance was compared against EHR-based screening methods. Both standalone and combined data approaches were assessed using F1-score and area under the curve (AUC) metrics. RESULTS: The initial model using only patient-nurse verbal communication achieved an F1-score of 85 and an AUC of 86.47. The model based on EHR data achieved an F1-score of 75.56 and an AUC of 79. Combining patient-nurse verbal communication with EHR data yielded the highest performance, with an F1-score of 88.89 and an AUC of 90.23. Key linguistic indicators of cognitive decline included reduced linguistic diversity, grammatical challenges, repetition, and altered speech patterns. Incorporating audio data significantly enhanced the risk prediction models for hospitalization and emergency department visits. DISCUSSION: Routine verbal communication between patients and nurses contains critical linguistic and interactional indicators for identifying cognitive impairment. Integrating audio-recorded patient-nurse communication with EHR data provides a more comprehensive and accurate method for early detection of cognitive decline, potentially improving patient outcomes through timely interventions. This combined approach could revolutionize cognitive impairment screening in home healthcare settings. Maryam Zolnoori, Ali Zolnour, Sasha Vergez, Sridevi Sridharan, Ian Spens, Maxim Topaz, James Noble 0003, Suzanne Bakken, Julia Hirschberg, Kathryn H. Bowles, Nicole Onorato, Margaret V. McDonald |
J. Am. Medical Informatics Assoc. | 9 |
| 2024 | Enhancing Pre-Trained Generative Language Models with Question Attended Span Extraction on Machine Reading ComprehensionabstractMachine Reading Comprehension (MRC) poses a significant challenge in the field of Natural Language Processing (NLP).While mainstream MRC methods predominantly leverage extractive strategies using encoder-only models such as BERT, generative approaches face the issue of out-of-control generation -a critical problem where answers generated are often incorrect, irrelevant, or unfaithful to the source text.To address these limitations in generative models for extractive MRC, we introduce the Question-Attended Span Extraction (QASE) module.Integrated during the finetuning phase of pre-trained generative language models (PLMs), QASE significantly enhances their performance, allowing them to surpass the extractive capabilities of advanced Large Language Models (LLMs) such as GPT-4 in few-shot settings.Notably, these gains in performance do not come with an increase in computational demands.The efficacy of the QASE module has been rigorously tested across various datasets, consistently achieving or even surpassing state-of-the-art (SOTA) results, thereby bridging the gap between generative and extractive models in extractive MRC tasks.Our code is available at this GitHub repository. Lin Ai, Zheng Hui, Zizhou Liu, Julia Hirschberg |
EMNLP | 4 |
| 2024 | Defending Against Social Engineering Attacks in the Age of LLMsabstractLin Ai, Tharindu Sandaruwan Kumarage, Amrita Bhattacharjee, Zizhou Liu, Zheng Hui, Michael S. Davinroy, James Cook, Laura Cassani, Kirill Trapeznikov, Matthias Kirchner, Arslan Basharat, Anthony Hoogs, Joshua Garland, Huan Liu, Julia Hirschberg. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Lin Ai, Tharindu Kumarage, Amrita Bhattacharjee, Zizhou Liu, Zheng Hui, Michael Davinroy, James Cook, Laura Cassani, Kirill Trapeznikov, Matthias Kirchner, Arslan Basharat, Anthony Hoogs, Joshua Garland, Huan Liu 0001, Julia Hirschberg |
EMNLP | 15 |
| 2024 | EmoKnob: Enhance Voice Cloning with Fine-Grained Emotion ControlabstractWhile recent advances in Text-to-Speech (TTS) technology produce natural and expressive speech, they lack the option for users to select emotion and control intensity.We propose EmoKnob, a framework that allows fine-grained emotion control in speech synthesis with fewshot demonstrative samples of arbitrary emotion.Our framework leverages the expressive speaker representation space made possible by recent advances in foundation voice cloning models.Based on the few-shot capability of our emotion control framework, we propose two methods to apply emotion control on emotions described by open-ended text, enabling an intuitive interface for controlling a diverse array of nuanced emotions.To facilitate a more systematic emotional speech synthesis field, we introduce a set of evaluation metrics designed to rigorously assess the faithfulness and recognizability of emotion control frameworks.Through objective and subjective evaluations, we show that our emotion control framework effectively embeds emotions into speech and surpasses emotion expressiveness of commercial TTS services.1 Run Chen, Julia Hirschberg |
EMNLP | 3 |
| 2024 | TweetIntent@Crisis: A Dataset Revealing Narratives of Both Sides in the Russia-Ukraine CrisisabstractThis paper introduces TweetIntent@Crisis, a novel Twitter dataset centered on the Russia-Ukraine crisis. Comprising over 17K tweets from government-affiliated accounts of both nations, the dataset is meticulously annotated to identify underlying intents and detailed intent-related information. Our analysis demonstrates the dataset's capability in revealing fine-grained intents and nuanced narratives within the tweets from both parties involved in the crisis. We aim for TweetIntent@Crisis to provide the research community with a valuable tool for understanding and analyzing granular media narratives and their impact in this geopolitical conflict. Lin Ai, Sameer Gupta, Shreya Oak, Zheng Hui, Zizhou Liu, Julia Hirschberg |
ICWSM | 6 |
| 2024 | Switching Tongues, Sharing Hearts: Identifying the Relationship between Empathy and Code-switching in Speech
Debasmita Bhattacharya, Eleanor Lin, Run Chen, Julia Hirschberg |
INTERSPEECH | 4 |
| 2024 | Detecting Empathy in Speech
Run Chen, Anushka Kulkarni, Eleanor Lin, Linda Pang, Divya Tadimeti, Jun Shin, Julia Hirschberg |
INTERSPEECH | 8 |
| 2024 | MultiPA: A Multi-task Speech Pronunciation Assessment Model for Open Response Scenarios
Yuwen Chen 0006, Julia Hirschberg |
INTERSPEECH | 3 |
| 2024 | Measuring Entrainment in Spontaneous Code-switched SpeechabstractDebasmita Bhattacharya, Siying Ding, Alayna Nguyen, Julia Hirschberg. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Debasmita Bhattacharya, Siying Ding, Alayna Nguyen, Julia Hirschberg |
NAACL-HLT | 4 |
| 2024 | Multimodal Multi-loss Fusion Network for Sentiment AnalysisabstractZehui Wu, Ziwei Gong, Jaywon Koo, Julia Hirschberg. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Zehui Wu, Ziwei Gong, Jaywon Koo, Julia Hirschberg |
NAACL-HLT | 4 |
| 2023 | Identifying Entrainment in Task-Oriented ConversationsabstractHuman interlocutors adapt their behavior to each other in a conversation through entrainment. While entrainment has been found in long chit-chat conversations, much less research has been conducted on task-oriented dialogs. In this paper, we investigate short task-oriented Wizard-of-Oz conversations for acoustic-prosodic and lexical entrainment. We conduct significance tests that reveal changes in speech pitch and frequent words as important indicators of entrainment. Our findings will guide user-entraining dialog systems to improve the quality of conversations. Run Chen, Seokhwan Kim, Alexandros Papangelis, Julia Hirschberg, Yang Liu 0004, Dilek Hakkani-Tür |
ICASSP | 4 |
| 2023 | Capturing Formality in Speech Across Domains and LanguagesabstractThe linguistic notion of formality is one dimension of stylistic variation in human communication. A universal characteristic of language production, formality has surface-level realizations in written and spoken language. In this work, we explore ways of measuring the formality of such realizations in multilingual speech corpora across a wide range of domains. We compare measures of formality, contrasting textual and acoustic-prosodic metrics. We believe that a combination of these should correlate well with downstream applications. Our findings include: an indication that certain prosodic variables might play a stronger role than others; no correlation between prosodic and textual measures; limited evidence for anticipated inter-domain trends, but some evidence of consistency of measures between languages. We conclude that non-lexical indicators of formality in speech may be more subtle than our initial expectations, motivating further work on reliably encoding spoken formality. Debasmita Bhattacharya, Jie Chi, Julia Hirschberg, Peter Bell 0001 |
INTERSPEECH | 3 |
| 2023 | Investigating prosodic entrainment from global conversations to local turns and tones in Mandarin conversations
Zhihua Xia, Julia Hirschberg, Rivka Levitan |
Speech Commun. | 2 |
| 2021 | Identifying the Popularity and Persuasiveness of Right- and Left-Leaning Group Videos on Social MediaabstractWe have collected over 30,000 right- and left-leaning groups’ videos from YouTube, Bitchute, 4Chan and Vimeo to identify aspects of their content and presentation which make these videos more popular and also potentially more persuasive. To date we have collected videos for and against Antifa and other anti-Fascist groups, Black Lives Matter, Proud Boys, Oath Keepers and QAnon and manually labelled subsets for style, stance toward the group, persuasiveness, techniques used and other features. We have also extracted video features including titles, descriptions, time of upload, captions and ASR transcripts, topic categories, and users’ likes, dislikes, comments, and views. We are currently using these to automatically identify information such as the stance of the video (for or against a group), changes in popularity and in the sentiment of viewers toward the videos over time, correlating these changes with major events. We are also extracting text and audio features from videos and their comments to develop multimodal Machine Learning models for use in identifying different types of videos (e.g. pro- and anti- a group, extremely popular or unpopular) and eventually to use in identifying new radical groups and tracking their success. We will also be crowdsourcing surveys of subsets of these videos to understand how persons with different demographics and personality types perceive and are potentially influenced by different groups and different types of videos. Lin Ai, Anika Kathuria, Subhadarshi Panda, Arushi Sahai, Yuwen Yu, Sarah Ita Levitan, Julia Hirschberg |
IEEE BigData | 7 |
| 2021 | "Talk to me with left, right, and angles": Lexical entrainment in spoken Hebrew dialogueabstractAndreas Weise, Vered Silber-Varod, Anat Lerner, Julia Hirschberg, Rivka Levitan. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Andreas Weise, Vered Silber-Varod, Anat Lerner, Julia Hirschberg, Rivka Levitan |
EACL | 4 |
| 2021 | CHoRaL: Collecting Humor Reaction Labels from Millions of Social Media UsersabstractHumor detection has gained attention in recent years due to the desire to understand usergenerated content with figurative language.However, substantial individual and cultural differences in humor perception make it very difficult to collect a large-scale humor dataset with reliable humor labels.We propose CHoRaL, a framework to generate perceived humor labels on Facebook posts, using the naturally available user reactions to these posts with no manual annotation needed.CHoRaL provides both binary labels and continuous scores of humor and non-humor.We present the largest dataset to date with labeled humor on 785K posts related to COVID-19.Additionally, we analyze the expression of COVIDrelated humor in social media by extracting lexico-semantic and affective features from the posts, and build humor detection models with performance similar to humans.CHoRaL enables the development of large-scale humor detection models on any topic and opens a new path to the study of humor on social media. Zixiaofan Yang, Shayan Hooshmand, Julia Hirschberg |
EMNLP (1) | 3 |
| 2021 | Acoustic-Prosodic, Lexical and Demographic Cues to Persuasiveness in Competitive Debate Speeches
Ralph Vente, David Lupea, Sarah Ita Levitan, Julia Hirschberg |
Interspeech | 5 |
| 2020 | LieCatcher: Game Framework for Collecting Human Judgments of Deceptive SpeechabstractHumans are notoriously poor at detecting deception --- most are worse than chance. To address this issue we have developed LieCatcher, a single-player web-based Game With A Purpose (GWAP) that allows players to assess their lie detection skills while providing human judgments of deceptive speech. Players listen to audio recordings drawn from a corpus of deceptive and non-deceptive interview dialogues, and guess if the speaker is lying or telling the truth. They are awarded points for correct guesses and at the end of the game they receive a score summarizing their performance at lie detection. We present the game design and implementation, and describe a crowdsourcing experiment conducted to study perceived deception. Sarah Ita Levitan, Xinyue Tan, Julia Hirschberg |
ICMI | 3 |
| 2020 | Multimodal Deception Detection Using Automatically Extracted Acoustic, Visual, and Lexical Features
Sarah Ita Levitan, Julia Hirschberg |
INTERSPEECH | 3 |
| 2020 | An empirical study of the effect of acoustic-prosodic entrainment on the perceived trustworthiness of conversational avatars
Ramiro H. Gálvez, Agustín Gravano, Stefan Benus, Rivka Levitan, Marián Trnka, Julia Hirschberg |
Speech Commun. | 6 |
| 2020 | Acoustic-Prosodic and Lexical Cues to Deception and Trust: Deciphering How People Detect LiesabstractHumans rarely perform better than chance at lie detection. To better understand human perception of deception, we created a game framework, LieCatcher, to collect ratings of perceived deception using a large corpus of deceptive and truthful interviews. We analyzed the acoustic-prosodic and linguistic characteristics of language trusted and mistrusted by raters and compared these to characteristics of actual truthful and deceptive language to understand how perception aligns with reality. With this data we built classifiers to automatically distinguish trusted from mistrusted speech, achieving an F1 of 66.1%. We next evaluated whether the strategies raters said they used to discriminate between truthful and deceptive responses were in fact useful. Our results show that, although several prosodic and lexical features were consistently perceived as trustworthy, they were not reliable cues. Also, the strategies that judges reported using in deception detection were not helpful for the task. Our work sheds light on the nature of trusted language and provides insight into the challenging problem of human deception detection. Xi Leslie Chen, Sarah Ita Levitan, Michelle Levine, Marko Mandic, Julia Hirschberg |
Trans. Assoc. Comput. Linguistics | 5 |
| 2019 | Sincerity in Acted Speech: Presenting the Sincere Apology Corpus and ResultsabstractThe ability to discern an individual's level of sincerity varies from person to person and across cultures.Sincerity is typically a key indication of personality traits such as trustworthiness, and portraying sincerity can be integral to an abundance of scenarios, e. g. , when apologising.Speech signals are one important factor when discerning sincerity and, with more modern interactions occurring remotely, automatic approaches for the recognition of sincerity from speech are beneficial during both interpersonal and professional scenarios.In this study we present details of the Sincere Apology Corpus (SINA-C ).Annotated by 22 individuals for their perception of sincerity, SINA-C is an English acted-speech corpus of 32 speakers, apologising in multiple ways.To provide an updated baseline for the corpus, various machine learning experiments are conducted.Finding that extracting deep data-representations (utilising the DEEP SPECTRUM toolkit) from the speech signals is best suited.Classification results on the binary (sincere / not sincere) task are at best 79.2 % Unweighted Average Recall and for regression, in regards to the degree of sincerity, a Root Mean Square Error of 0.395 from the standardised range [-1.51; 1.72] is obtained. Alice Baird, Eduardo Coutinho, Julia Hirschberg, Björn W. Schuller |
INTERSPEECH | 3 |
| 2019 | Improving Code-Switched Language Modeling Performance Using Cognate Features
Victor Soto, Julia Hirschberg |
INTERSPEECH | 2 |
| 2019 | Linguistically-Informed Training of Acoustic Word Embeddings for Low-Resource Languages
Zixiaofan Yang, Julia Hirschberg |
INTERSPEECH | 2 |
| 2019 | Predicting Humor by Learning from Time-Aligned Comments
Zixiaofan Yang, Bingyan Hu, Julia Hirschberg |
INTERSPEECH | 3 |
| 2019 | Individual differences in acoustic-prosodic entrainment in spoken dialogue
Andreas Weise, Sarah Ita Levitan, Julia Hirschberg, Rivka Levitan |
Speech Commun. | 3 |
| 2018 | Deep Personality Recognition for Deception Detection
Guozhen An, Sarah Ita Levitan, Julia Hirschberg, Rivka Levitan |
INTERSPEECH | 3 |
| 2018 | A Comparison of Speaker-based and Utterance-based Data Selection for Text-to-Speech Synthesis
Kai-Zhan Lee, Erica Cooper, Julia Hirschberg |
INTERSPEECH | 3 |
| 2018 | Acoustic-Prosodic Indicators of Deception and Trust in Interview Dialogues
Sarah Ita Levitan, Angel Maredia, Julia Hirschberg |
INTERSPEECH | 3 |
| 2018 | The Role of Cognate Words, POS Tags and Entrainment in Code-Switching
Victor Soto, Nishmar Cestero, Julia Hirschberg |
INTERSPEECH | 3 |
| 2018 | Predicting Arousal and Valence from Waveforms and Spectrograms Using Deep Neural Networks
Zixiaofan Yang, Julia Hirschberg |
INTERSPEECH | 2 |
| 2018 | Collecting Code-Switched Data from Social Media
Gideon Mendels, Victor Soto, Aaron Jaech, Julia Hirschberg |
LREC | 4 |
| 2018 | Evaluating the WordsEye Text-to-Scene System: Imaginative and Realistic Sentences
Morgan Ulinski, Bob Coyne, Julia Hirschberg |
LREC | 3 |
| 2018 | Linguistic Cues to Deception and Perceived Deception in Interview DialoguesabstractSarah Ita Levitan, Angel Maredia, Julia Hirschberg. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Sarah Ita Levitan, Angel Maredia, Julia Hirschberg |
NAACL-HLT | 3 |
| 2017 | Utterance Selection for Optimizing Intelligibility of TTS Voices Trained on ASR Data
Erica Cooper, Alison Chang, Yocheved Levitan, Julia Hirschberg |
INTERSPEECH | 5 |
| 2017 | Hybrid Acoustic-Lexical Deep Learning Approach for Deception Detection
Gideon Mendels, Sarah Ita Levitan, Kai-Zhan Lee, Julia Hirschberg |
INTERSPEECH | 4 |
| 2017 | Crowdsourcing Universal Part-of-Speech Tags for Code-SwitchingabstractCode-switching is the phenomenon by which bilingual speakers switch between multiple languages during communication.The importance of developing language technologies for codeswitching data is immense, given the large populations that routinely code-switch.High-quality linguistic annotations are extremely valuable for any NLP task, and performance is often limited by the amount of high-quality labeled data.However, little such data exists for code-switching.In this paper, we describe crowd-sourcing universal part-of-speech tags for the Miami Bangor Corpus of Spanish-English code-switched speech.We split the annotation task into three subtasks: one in which a subset of tokens are labeled automatically, one in which questions are specifically designed to disambiguate a subset of high frequency words, and a more general cascaded approach for the remaining data in which questions are displayed to the worker following a decision tree structure.Each subtask is extended and adapted for a multilingual setting and the universal tagset.The quality of the annotation process is measured using hidden check questions annotated with gold labels.The overall agreement between gold standard labels and the majority vote is between 0.95 and 0.96 for just three labels and the average recall across part-of-speech tags is between 0.87 and 0.99, depending on the task. Victor Soto, Julia Hirschberg |
INTERSPEECH | 2 |
| 2016 | Incrementally Learning a Dependency Parser to Support Language Documentation in Field LinguisticsabstractWe present experiments in incrementally learning a dependency parser. The parser will be used in the WordsEye Linguistics Tools (WELT) (Ulinski et al., 2014) which supports field linguists documenting a language’s syntax and semantics. Our goal is to make syntactic annotation faster for field linguists. We have created a new parallel corpus of descriptions of spatial relations and motion events, based on pictures and video clips used by field linguists for elicitation of language from native speaker informants. We collected descriptions for each picture and video from native speakers in English, Spanish, German, and Egyptian Arabic. We compare the performance of MSTParser (McDonald et al., 2006) and MaltParser (Nivre et al., 2006) when trained on small amounts of this data. We find that MaltParser achieves the best performance. We also present the results of experiments using the parser to assist with annotation. We find that even when the parser is trained on a single sentence from the corpus, annotation time significantly decreases. Morgan Ulinski, Julia Hirschberg, Owen Rambow |
COLING | 2 |
| 2016 | Testing the consistency assumption: Pronunciation variant forced alignment in read and spontaneous speech synthesisabstractForced alignment for speech synthesis traditionally aligns a phoneme sequence predetermined by the front-end text processing system. This sequence is not altered during alignment, i.e., it is forced, despite possibly being faulty. The consistency assumption is the assumption that these mistakes do not degrade models, as long as the mistakes are consistent across training and synthesis. We present evidence that in the alignment of both standard read prompts and spontaneous speech this phoneme sequence is often wrong, and that this is likely to have a negative impact on acoustic models. A lattice-based forced alignment system allowing for pronunciation variation is implemented, resulting in improved phoneme identity accuracy for both types of speech. A perceptual evaluation of HMM-based voices showed that spontaneous models trained on this improved alignment also improved standard synthesis, despite breaking the consistency assumption. Rasmus Dall, Sandrine Brognaux, Korin Richmond, Cassia Valentini-Botinhao, Gustav Eje Henter, Julia Hirschberg, Junichi Yamagishi, Simon King 0001 |
ICASSP | 6 |
| 2016 | Automatically Classifying Self-Rated Personality Scores from Speech
Guozhen An, Sarah Ita Levitan, Rivka Levitan, Andrew Rosenberg, Michelle Levine, Julia Hirschberg |
INTERSPEECH | 6 |
| 2016 | Data Selection and Adaptation for Naturalness in HMM-Based Speech Synthesis
Erica Cooper, Alison Chang, Yocheved Levitan, Julia Hirschberg |
INTERSPEECH | 4 |
| 2016 | Computational Approaches to Linguistic Code Switching
Mona T. Diab, Pascale Fung, Julia Hirschberg, Thamar Solorio |
INTERSPEECH | 3 |
| 2016 | Combining Acoustic-Prosodic, Lexical, and Phonotactic Features for Automatic Deception Detection
Sarah Ita Levitan, Guozhen An, Rivka Levitan, Andrew Rosenberg, Julia Hirschberg |
INTERSPEECH | 6 |
| 2016 | Implementing Acoustic-Prosodic Entrainment in a Conversational Avatar
Rivka Levitan, Stefan Benus, Ramiro H. Gálvez, Agustín Gravano, Florencia Savoretti, Marián Trnka, Andreas Weise, Julia Hirschberg |
INTERSPEECH | 8 |
| 2016 | The INTERSPEECH 2016 Computational Paralinguistics Challenge: Deception, Sincerity & Native LanguageabstractThe INTERSPEECH 2016 Computational Paralinguistics Challenge addresses three different problems for the first time in research competition under well-defined conditions: classification of deceptive vs. non-deceptive speech, the estimation of the degree of sincerity, and the identification of the native language out of eleven L1 classes of English L2 speakers.In this paper, we describe these sub-challenges, their conditions, the baseline feature extraction and classifiers, and the resulting baselines, as provided to the participants. Björn W. Schuller, Stefan Steidl, Anton Batliner, Julia Hirschberg, Judee K. Burgoon, Alice Baird, Aaron C. Elkins, Yue Zhang 0014, Eduardo Coutinho, Keelan Evanini |
INTERSPEECH | 4 |
| 2016 | The Deception Sub-Challenge: The Data
Björn W. Schuller, Stefan Steidl, Anton Batliner, Julia Hirschberg, Judee K. Burgoon, Alice Baird, Aaron C. Elkins, Yue Zhang 0014, Eduardo Coutinho, Keelan Evanini |
INTERSPEECH | 4 |
| 2016 | The Sincerity Sub-Challenge: The Data
Björn W. Schuller, Stefan Steidl, Anton Batliner, Julia Hirschberg, Judee K. Burgoon, Alice Baird, Aaron C. Elkins, Yue Zhang 0014, Eduardo Coutinho, Keelan Evanini |
INTERSPEECH | 4 |
| 2016 | The Native Language Sub-Challenge: The Data
Björn W. Schuller, Stefan Steidl, Anton Batliner, Julia Hirschberg, Judee K. Burgoon, Alice Baird, Aaron C. Elkins, Yue Zhang 0014, Eduardo Coutinho, Keelan Evanini |
INTERSPEECH | 4 |
| 2016 | The INTERSPEECH 2016 Computational Paralinguistics Challenge: A Summary of Results
Björn W. Schuller, Stefan Steidl, Anton Batliner, Julia Hirschberg, Judee K. Burgoon, Alice Baird, Aaron C. Elkins, Yue Zhang 0014, Eduardo Coutinho, Keelan Evanini |
INTERSPEECH | 4 |
| 2016 | Discussion
Björn W. Schuller, Stefan Steidl, Anton Batliner, Julia Hirschberg, Judee K. Burgoon, Alice Baird, Aaron C. Elkins, Yue Zhang 0014, Eduardo Coutinho, Keelan Evanini |
INTERSPEECH | 4 |
| 2015 | Backward mimicry and forward influence in prosodic contour choice in standard American EnglishabstractEntrainment is the tendency of speakers engaged in conversation to align different aspects of their communicative behavior. In this study we explore in more detail a measure of prosodic entrainment defined in previous work, which uses a discrete parametrization of intonational contours defined by the ToBI conventions for prosodic description. We divide this measure into two asymmetric variants: backward mimicry (in which a speaker uses a contour used previously by the interlocutor) and forward influence (in which a speaker’s contour appears later in the speech of the interlocutor). This distinction sheds new light on significant correlations with a number of social variables related to the level of engagement of speakers in a corpus of task-oriented dialogues in Standard American English. Agustín Gravano, Stefan Benus, Rivka Levitan, Julia Hirschberg |
INTERSPEECH | 4 |
| 2015 | Improving speech recognition and keyword search for low resource languages using web dataabstractWe describe the use of text data scraped from the web to augment language models for Automatic Speech Recognition and Keyword Search for Low Resource Languages.We scrape text from multiple genres including blogs, online news, translated TED talks, and subtitles.Using linearly interpolated language models, we find that blogs and movie subtitles are more relevant for language modeling of conversational telephone speech and obtain large reductions in out-of-vocabulary keywords.Furthermore, we show that the web data can improve Term Error Rate Performance by 3.8% absolute and Maximum Term-Weighted Value in Keyword Search by 0.0076-0.1059absolute points.Much of the gain comes from the reduction of out-of-vocabulary items. Gideon Mendels, Erica Cooper, Victor Soto, Julia Hirschberg, Mark J. F. Gales, Kate M. Knill, Anton Ragni |
INTERSPEECH | 4 |
| 2015 | Acoustic-prosodic entrainment in Slovak, Spanish, English and Chinese: A cross-linguistic comparisonabstractIt is well established that speakers of Standard American English entrain, or become more similar to each other as they speak, in acoustic-prosodic features of their speech as well as other behaviors.Entrainment in other languages is less well understood.This work uses a variety of metrics to measure acoustic-prosodic entrainment in four comparable corpora of task-oriented conversational speech in Slovak, Spanish, English and Chinese.We report the results of these experiments and describe trends and patterns that can be observed from comparing acoustic-prosodic entrainment in these four languages.We find evidence of a variety of forms of entrainment across all the languages studied, with some evidence of individual differences as well within the languages. Rivka Levitan, Stefan Benus, Agustín Gravano, Julia Hirschberg |
SIGDIAL Conference | 4 |
| 2014 | Rescoring Confusion Networks for Keyword SearchabstractWe introduce a two-stage cascaded scheme to rescore Confusion Networks (CNs) for Keyword Search in the context of Low-Resource Languages. In the first stage we rescore the CN to improve the error rate of the 1-best hypothesis using a large number of lexical, phonetic, false alarms and structural features. Using a rank learning Support Vector Machine classifier, we obtain WER gains between 0.54% and 2.84% on Cantonese, Tagalog, Turkish, Pashto and Vietnamese. In the second stage we generate keyword hits from the rescored CN and use logistic regression to detect true hits and false alarms. We compare these to hits generated from the unrescored CN and obtain gains between 0.45% and 0.9% on the MTWV metric by using the mentioned features and including acoustic and prosodic features on Tagalog, Turkish and Pashto. Victor Soto, Erica Cooper, Lidia Mangu, Andrew Rosenberg, Julia Hirschberg |
ICASSP | 5 |
| 2014 | Strategies for rescoring keyword search results using word-burst and acoustic featuresabstractThe identification of keyword queries in speech data from lowresources languages poses a challenge for current methods as speech recognition algorithms lack sufficient training data to produce high accuracy transcript. To compensate for these shortcomings, we extract signals from the data that are useful in keyword identification but are not being used by the speech recognizer. These signals take multiple forms — word burstiness, rescored confusion network posteriors and acoustic/prosodic qualities. The former denotes the tendency for keywords to occur in bursts within a conversational topic. We employ three different strategies to exploit this information: 1) a four-way classification of keyword hypotheses that targets low-scoring correct hits and high-scoring false alarms, 2) ranking algorithms, and 3) a direct adjustment of keyword hit scores based on hypothesized repetition. We find that interpolating the results of these three strategies in an ensemble provides a reliable way to improve the results of keyword search. Justin Richards, Victor Soto, Julia Hirschberg, Andrew Rosenberg |
INTERSPEECH | 4 |
| 2014 | A comparison of multiple methods for rescoring keyword search lists for low resource languagesabstractWe review the performance of a new two-stage cascaded machine learning approach for rescoring keyword search output for low resource languages. In the first stage Confusion Networks (CNs) are rescored for improved Automatic Speech Recognition (ASR) by reranking the arcs of each confusion bin. In the second stage we generate keyword search hypotheses from the rescored ASR output and rescore them using logistic regression classifiers to detect true hits and false alarms. We compare the performance of our system with state of the art rescoring techniques, including probability of false alarm normalization, exponential normalization, rank-normalized posterior scores and sum-to-one normalization and show promising results. Experimental validation is performed using the Term Weighted Value (TWV) metric on four corpora from the IARPA-Babel program for keyword search on low resource languages, including Assamese, Bengali, Lao and Zulu. Victor Soto, Lidia Mangu, Andrew Rosenberg, Julia Hirschberg |
INTERSPEECH | 4 |
| 2014 | Teenage and adult speech in school context: building and processing a corpus of European Portuguese
Ana Isabel Mata, Helena Moniz, Fernando Batista, Julia Hirschberg |
LREC | 4 |
| 2014 | Detecting Inappropriate Clarification Requests in Spoken Dialogue SystemsabstractAlex Liu, Rose Sloan, Mei-Vern Then, Svetlana Stoyanchev, Julia Hirschberg, Elizabeth Shriberg. Proceedings of the 15th Annual Meeting of the Special Interest Group on Discourse and Dialogue (SIGDIAL). 2014. Rose Sloan, Mei-Vern Then, Svetlana Stoyanchev, Julia Hirschberg, Elizabeth Shriberg |
SIGDIAL Conference | 5 |
| 2014 | Three ToBI-based measures of prosodic entrainment and their correlations with speaker engagementabstractEntrainment is the propensity of conversational partners to align different aspects of their communicative behavior. In this study we present three novel measures of prosodic entrainment based on intonational contours as defined by the ToBI conventions for prosodic description. Each of these measures estimates the similarity of contours used by speakers in different ways: by means of the perplexity of n-gram models, the Levenshtein distance, and the Kullback-Leibler divergence measure. We report significant correlations between each of these measures and manual annotations of a number of social variables related to the level of engagement of speakers, in a corpus of task-oriented dialogues in Standard American English. Agustín Gravano, Stefan Benus, Rivka Levitan, Julia Hirschberg |
SLT | 4 |
| 2014 | Entrainment, dominance and alliance in supreme court hearings
Stefan Benus, Agustín Gravano, Rivka Levitan, Sarah Ita Levitan, Laura Willson, Julia Hirschberg |
Knowl. Based Syst. | 6 |
| 2013 | "Can you give me another word for hyperbaric?": Improving speech translation using targeted clarification questionsabstractWe present a novel approach for improving communication success between users of speech-to-speech translation systems by automatically detecting errors in the output of automatic speech recognition (ASR) and statistical machine translation (SMT) systems. Our approach initiates system-driven targeted clarification about errorful regions in user input and repairs them given user responses. Our system has been evaluated by unbiased subjects in live mode, and results show improved success of communication between users of the system. Necip Fazil Ayan, Arindam Mandal, Michael W. Frandsen, Jing Zheng 0001, Peter Blasco, Andreas Kathol, Frédéric Béchet, Benoît Favre, Alex Marin, Tom Kwiatkowski, Mari Ostendorf, Luke Zettlemoyer, Philipp Salletmayr, Julia Hirschberg, Svetlana Stoyanchev |
ICASSP | 14 |
| 2013 | Cross-language phrase boundary detectionabstractWe describe models of prosodic phrasing trained on multiple languages to identify boundaries in an unseen language. Our goal is to create models from High Resource languages, in which hand-annotated prosodic phrase boundaries are available, to use in identifying boundaries in a Low Resource language, with little or no training material. We train models on American English, Italian, Mandarin, and German and test on each of these languages. We find that, while pause is the most important feature for phrase boundary prediction in all languages examined, the role of pause in boundary identification varies by annotator and the relative importance of other features varies significantly by language. We also find that different acoustic correlates of prosodic boundaries characterize different languages. In some, the relative importance of features is silence > pitch > intensity > duration, while for other languages intensity is more important than pitch. These differences do not appear to be attributable to language family, since, e.g. English and German display different patterns. Victor Soto, Erica Cooper, Andrew Rosenberg, Julia Hirschberg |
ICASSP | 4 |
| 2013 | Exploring Features For Localized Detection of Speech Recognition Errors
Eli Pincus, Svetlana Stoyanchev, Julia Hirschberg |
SIGDIAL Conference | 3 |
| 2013 | Modelling Human Clarification Strategies
Svetlana Stoyanchev, Julia Hirschberg |
SIGDIAL Conference | 3 |
| 2013 | Automatic detection of speaker state: Lexical, prosodic, and phonetic approaches to level-of-interest and intoxication classification
William Yang Wang, Fadi Biadsy, Andrew Rosenberg, Julia Hirschberg |
Comput. Speech Lang. | 4 |
| 2012 | A Corpus-Based Study of Interruptions in Spoken DialogueabstractWe examine interruptions in a corpus of spontaneous taskoriented dialogue. We present evidence that interruptions occur at particular places in conversation. They are likely to occur during or after speech with certain acoustic/prosodic properties. We also examine the speech of interruptions themselves and find a number of significant differences between interrupting and non-interrupting turns. Index Terms: interruption, turn-taking, prosody, dialogue. Agustín Gravano, Julia Hirschberg |
INTERSPEECH | 2 |
| 2012 | Acoustic-Prosodic Entrainment and Social Behavior
Rivka Levitan, Agustín Gravano, Laura Willson, Stefan Benus, Julia Hirschberg, Ani Nenkova |
HLT-NAACL | 5 |
| 2012 | Localized detection of speech recognition errorsabstractWe address the problem of localized error detection in Automatic Speech Recognition (ASR) output. Localized error detection seeks to identify which particular words in a user's utterance have been misrecognized. Identifying misrecognized words permits one to create targeted clarification strategies for spoken dialogue systems, allowing the system to ask clarification questions targeting the particular type of misrecognition, in contrast to the “please repeat/rephrase” strategies used in most current dialogue systems. We present results of machine learning experiments using ASR confidence scores together with prosodic and syntactic features to predict whether 1) an utterance contains an error, and 2) whether a word in a misrecognized utterance is misrecognized. We show that by adding syntactic features to the ASR features when predicting misrecognized utterances the F-measure improves by 13.3% compared to using ASR features alone. By adding syntactic and prosodic features when predicting misrecognized words F-measure improves by 40%. Svetlana Stoyanchev, Philipp Salletmayr, Julia Hirschberg |
SLT | 4 |
| 2012 | Affirmative Cue Words in Task-Oriented DialogueabstractWe present a series of studies of affirmative cue words—a family of cue words such as “okay” or “alright” that speakers use frequently in conversation. These words pose a challenge for spoken dialogue systems because of their ambiguity: They may be used for agreeing with what the interlocutor has said, indicating continued attention, or for cueing the start of a new topic, among other meanings. We describe differences in the acoustic/prosodic realization of such functions in a corpus of spontaneous, task-oriented dialogues in Standard American English. These results are important both for interpretation and for production in spoken language applications. We also assess the predictive power of computational methods for the automatic disambiguation of these words. We find that contextual information and final intonation figure as the most salient cues to automatic disambiguation. Agustín Gravano, Julia Hirschberg, Stefan Benus |
Comput. Linguistics | 2 |
| 2011 | Dialect and Accent Recognition Using Phonetic-Segmentation SupervectorsabstractWe describe a new approach to automatic dialect and accent recognition which exceeds state-of-the-art performance in three recognition tasks.This approach improves the accuracy and substantially lower the time complexity of our earlier phoneticbased kernel approach for dialect recognition.In contrast to state-of-the-art acoustic-based systems, our approach employs phone labels and segmentation to constrain the acoustic models.Given a speaker's utterance, we first obtain phone hypotheses using a phone recognizer and then extract GMM-supervectors for each phone type, effectively summarizing the speaker's phonetic characteristics in a single vector of phone-type supervectors.Using these vectors, we design a kernel function that computes the phonetic similarities between pairs of utterances to train SVM classifiers to identify dialects.Comparing this approach to the state-of-the-art, we obtain a 12.9% relative improvement in EER on Arabic dialects, and a 17.9% relative improvement for American vs. Indian English dialects.We also see a 53.5% relative improvement over a GMM-UBM on American Southern vs. Non-Southern English. Fadi Biadsy, Julia Hirschberg, Daniel P. W. Ellis |
INTERSPEECH | 2 |
| 2011 | Intoxication Detection Using Phonetic, Phonotactic and Prosodic Cues
Fadi Biadsy, William Yang Wang, Andrew Rosenberg, Julia Hirschberg |
INTERSPEECH | 4 |
| 2011 | Acoustic and Prosodic Correlates of Social BehaviorabstractWe describe acoustic/prosodic and lexical correlates of social variables annotated on a large corpus of task-oriented spontaneous speech.We employ Amazon Mechanical Turk to label the corpus with a large number of social behaviors, examining results of three of these here.We find significant differences between male and female speakers for perceptions of attempts to be liked, likeability, speech planning, that also differ depending upon the gender of their conversational partners. Agustín Gravano, Rivka Levitan, Laura Willson, Stefan Benus, Julia Hirschberg, Ani Nenkova |
INTERSPEECH | 5 |
| 2011 | Accounting for Prosodic Information to Improve ASR-Based Topic Tracking for TV Broadcast NewsabstractThe increasing quantity of video material available on line requires improved methods to help users navigate such data, among which are topic tracking techniques. The goal of this paper is to show that prosodic information can improve an ASR based topic tracking system for French TV Broadcast News. To this end, two kinds of prosodic information--extracted with and without a learning phase--are integrated in the system. This integration shows significant improvements in the F1-measure, by 13 and 8 points for the two techniques compared with the baseline system. Camille Guinaudeau, Julia Hirschberg |
INTERSPEECH | 2 |
| 2011 | Speaking More Like You: Entrainment in Conversational Speech
Julia Hirschberg |
INTERSPEECH | 1 |
| 2011 | Measuring Acoustic-Prosodic Entrainment with Respect to Multiple Levels and DimensionsabstractIn conversation, speakers become more like each other in various dimensions.This phenomenon, commonly called entrainment, coordination, or alignment, is widely believed to be crucial to the success and naturalness of human interactions.We investigate entrainment in four acoustic and prosodic dimensions.We explore whether speakers coordinate with each other in these dimensions over the conversation as a whole as well as on a turn-by-turn basis and in both relative and absolute terms, and whether this coordination improves over the course of the conversation. Rivka Levitan, Julia Hirschberg |
INTERSPEECH | 2 |
| 2011 | Detecting Levels of Interest from Spoken Dialog with Multistream Prediction Feedback and Similarity Based Hierarchical Fusion Learning
William Yang Wang, Julia Hirschberg |
SIGDIAL Conference | 2 |
| 2011 | Turn-taking cues in task-oriented dialogue
Agustín Gravano, Julia Hirschberg |
Comput. Speech Lang. | 2 |
| 2010 | Dialect recognition using a phone-GMM-supervector-based SVM kernelabstractIn this paper, we introduce a new approach to dialect recognition which relies on the hypothesis that certain phones are realized differently across dialects. Given a speaker’s utterance, we first obtain the most likely phone sequence using a phone recognizer. We then extract GMM Supervectors for each phone instance. Using these vectors, we design a kernel function that computes the similarities of phones between pairs of utterances. We employ this kernel to train SVM classifiers that estimate posterior probabilities, used during recognition. Testing our approach on four Arabic dialects from 30s cuts, we compare our performance to five approaches: PRLM; GMM-UBM; our own improved version of GMM-UBM which employs fMLLR adaptation; our recent discriminative phonotactic approach; and a state-of-the-art system: SDC-based GMM-UBM discriminatively trained. Our kernel-based technique outperforms all these previous approaches; the overall EER of our system is 4.9%. Fadi Biadsy, Julia Hirschberg, Michael Collins 0001 |
INTERSPEECH | 2 |
| 2010 | Pitch similarity in the vicinity of backchannelsabstractDynamic modeling of spoken dialogue seeks to capture how interlocutors change their speech over the course of a conversation. Much work has focused on how speakers adapt or entrain to different aspects of one another’s speaking style. In this paper we focus on local aspects of this adaptation. We investigate the relationship between backchannels and the interlocutor utterances that precede them with respect to pitch. We demonstrate that the pitch of backchannels is more similar to the immediately preceding utterance than non-backchannels. This inter-speaker pitch relationship captures the same distinctions as more cumbersome intra-speaker relations, and supports the intuition that, in terms of pitch, such similarity may be one of the mechanisms by which backchannels are rendered ’unobtrusive’. Mattias Heldner, Jens Edlund, Julia Hirschberg |
INTERSPEECH | 3 |
| 2010 | Sparse representations for text categorizationabstractSparse representations (SRs) are often used to characterize a test signal using few support training examples, and allow the number of supports to be adapted to the specific signal being categorized. Given the good performance of SRs compared to other classifiers for both image classification and phonetic clas-sification, in this paper, we extended the use of SRs for text classification, a method which has thus far not been explored for this domain. Specifically, we demonstrate how sparse repre-sentations can be used for text classification and how their per-formance varies with the vocabulary size of the documents. In addition, we also show that this method offers promising results over the Naive Bayes (NB) classifier, a standard baseline classi-fier used for text categorization, thus introducing an alternative class of methods for text categorization. 1. Tara N. Sainath, Sameer Maskey, Dimitri Kanevsky, Bhuvana Ramabhadran, David Nahamoo, Julia Hirschberg |
INTERSPEECH | 6 |
| 2010 | Frame Semantics in Text-to-Scene Generation
Bob Coyne, Owen Rambow, Julia Hirschberg, Richard Sproat |
KES (4) | 3 |
| 2009 | Using prosody and phonotactics in Arabic dialect identificationabstractWhile Modern Standard Arabic is the formal spoken and written language of the Arab world, dialects are the major communication mode for everyday life; identifying a speaker’s dialect is thus critical to speech processing tasks such as automatic speech recognition, as well as speaker identification. We examine the role of prosodic features (intonation and rhythm) across four Arabic dialects: Gulf, Iraqi, Levantine, and Egyptian, for the purpose of automatic dialect identification. We show that prosodic features can significantly improve identification, over a purely phonotactic-based approach, with an identification accuracy of 86.33 % for 2m utterances. 1. Fadi Biadsy, Julia Hirschberg |
INTERSPEECH | 2 |
| 2009 | Cross-cultural perception of discourse phenomenaabstractWe discuss perception studies of two low level indicators of discourse phenomena by Swedish, Japanese, and Chinese native speakers. Subjects were asked to identify upcoming prosodic boundaries and disfluencies in Swedish spontaneous speech. We hypothesize that speakers of prosodically unrelated languages should be less able to predict upcoming phrase boundaries but potentially better able to identify disfluencies, since indicators of disfluency are more likely to depend upon lexical, as well as acoustic information. However, surprisingly, we found that both phenomena were fairly well recognized by native and non-native speakers, with, however, some possible interference from word tones for the Chinese subjects. Rolf Carlson, Julia Hirschberg |
INTERSPEECH | 2 |
| 2009 | Pause and gap length in face-to-face interactionabstractIt has long been noted that conversational partners tend to exhibit increasingly similar pitch, intensity, and timing behavior over the course of a conversation. However, the metrics developed to measure this similarity to date have generally failed to capture the dynamic temporal aspects of this process. In this paper, we propose new approaches to measuring interlocutor similarity in spoken dialogue. define similarity in terms of convergence and synchrony and propose approaches to capture these, illustrating our techniques on gap and pause production in Swedish spontaneous dialogues. Jens Edlund, Mattias Heldner, Julia Hirschberg |
INTERSPEECH | 3 |
| 2009 | Backchannel-inviting cues in task-oriented dialogueabstractWe examine BACKCHANNEL-INVITING CUES — distinct prosodic, acoustic and lexical events in the speaker’s speech that tend to precede a short response produced by the interlocutor to convey continued attention — in the Columbia Games Corpus, a large corpus of task-oriented dialogues. We show that the likelihood of occurrence of a backchannel increases quadratically with the number of cues conjointly displayed by the speaker. Our results are important for improving the coordination of conversational turns in interactive voice-response systems, so that systems can produce backchannels in appropriate places, and so that they can elicit backchannels from users in expected places. Agustín Gravano, Julia Hirschberg |
INTERSPEECH | 2 |
| 2009 | Improving the Arabic Pronunciation Dictionary for Phone and Word Recognition with Linguistically-Based Pronunciation Rules
Fadi Biadsy, Nizar Habash, Julia Hirschberg |
HLT-NAACL | 3 |
| 2009 | Turn-Yielding Cues in Task-Oriented Dialogue
Agustín Gravano, Julia Hirschberg |
SIGDIAL Conference | 2 |
| 2009 | Charisma perception from text and speech
Andrew Rosenberg, Julia Hirschberg |
Speech Commun. | 2 |
| 2008 | An Unsupervised Approach to Biography Production Using Wikipedia
Fadi Biadsy, Julia Hirschberg, Elena Filatova |
ACL | 2 |
| 2008 | Intonational phrases for speech summarizationabstractExtractive speech summarization approaches select relevant segments of spoken documents and concatenate them to generate a summary. The extraction unit chosen, whether a sentence, syntactic constituent, or other segment, has a significant impact on the overall quality and fluency of the summary. Even though sentences tend to be the choice of most the extractive speech summarizers, in this paper, we present the results of an empirical study indicating that intonational phrases are better units of extraction for summarization. Our study compared four types of input segmentation: sentences, two pause-based segmentation, and intonational phrases (IP). We found that IPs are the best candidates for extractive summarization, improving over the second highest-performing approach, sentence-based summarization, by 8.2% F-measure. Sameer Maskey, Andrew Rosenberg, Julia Hirschberg |
INTERSPEECH | 3 |
| 2007 | On the role of context and prosody in the interpretation of 'okay'
Agustín Gravano, Stefan Benus, Héctor Chávez, Julia Hirschberg, Lauren Wilcox |
ACL | 4 |
| 2007 | V-Measure: A Conditional Entropy-Based External Cluster Evaluation Measure
Andrew Rosenberg, Julia Hirschberg |
EMNLP-CoNLL | 2 |
| 2007 | Prosody, emotions, and... 'whatever'abstractWe examine the role of prosody in cueing a scale of negative meanings associated with the use of whatever. The analysis of a corpus of elicited examples shows that the more negative the token, the more likely it is to have an additional pitch accent, extended duration, and expanded pitch range on the first syllable. These findings are analyzed as a link between pragmatic meaning and the strength of the prosodic boundary between the first two syllables (what#ever). The results of perception experiments show that the prosody of whatever itself is a systematic cue for the degree of negative connotation associated with the utterance in which whatever occurs. Potential applications of this result for spoken dialogue systems and synthesis of emotional speech are discussed. Stefan Benus, Agustín Gravano, Julia Hirschberg |
INTERSPEECH | 3 |
| 2007 | Comparing american and palestinian perceptions of charisma using acoustic-prosodic and lexical analysisabstractCharisma, the ability to lead by virtue of personality alone, is difficult to define but relatively easy to identify. However, cultural factors clearly affect perceptions of charisma. In this paper we compare results from parallel perception studies investigating charismatic speech in Palestinian Arabic and American English. We examine acoustic/prosodic and lexical correlates of charisma ratings to determine how the two cultures differ with respect to their views of charismatic speech. Fadi Biadsy, Julia Hirschberg, Andrew Rosenberg, Wisam Dakka |
INTERSPEECH | 2 |
| 2007 | Detecting deception using critical segmentsabstractWe present an investigation of segments that map to GLOBAL LIES, that is, the intent to deceive with respect to salient topics of the discourse. We propose that identifying the truth or falsity of these CRITICAL SEGMENTS may be important in determining a speaker’s veracity over the larger topic of discourse. Further, answers to key questions, which can be identified a priori, may represent emotional and cognitive HOT SPOTS, analogous to those observed by psychologists who study gestural and facial cues to deception. We present results of experiments that use two different definitions of CRITICAL SEGMENTS and employ machine learning techniques that compensate for imbalances in the dataset. Using this approach, we achieve a performance gain of 23.8% relative to chance, in contrast with human performance on a similar task, which averages substantially below chance. We discuss the features used by the models, and consider how these findings can influence future research. Frank Enos, Elizabeth Shriberg, Martin Graciarena, Julia Hirschberg, Andreas Stolcke |
INTERSPEECH | 4 |
| 2007 | Classification of discourse functions of affirmative words in spoken dialogueabstractWe present results of a series of machine learning experiments that address the classification of the discourse function of single affirmative cue words such as alright, okay and mm-hm in a spoken dialogue corpus. We suggest that a simple discourse/sentential distinction is not sufficient for such words and propose two additional classification sub-tasks: identifying (a) whether such words convey acknowledgment or agreement, and (b) whether they cue the beginning or end of a discourse segment. We also study the classification of each individual word into its most common discourse functions. We show that models based on contextual features extracted from the time-aligned transcripts approach the error rate of trained human aligners. Agustín Gravano, Stefan Benus, Julia Hirschberg, Shira Mitchell, Ilia Vovsha |
INTERSPEECH | 3 |
| 2007 | Detecting pitch accent using pitch-corrected energy-based predictorsabstractPrevious work has shown that the energy components of frequency subbands with a variety of frequencies and bandwidths predict pitch accent with various degrees of accuracy, and produce correct predictions for distinct subsets of data points. In this paper, we describe a series of experiments exploring techniques to leverage the predictive power of these energy components by including pitch and duration features – other known correlates to pitch accent. We perform these experiments on Standard American English read, spontaneous and broadcast news speech, each corpus containing at least four speakers. Using an approach by which we correct energy-based predictions using pitch and duration information prior to using a majority voting classifier, we were able to detect pitch accent in read, spontaneous and broadcast news speech at 84.0%, 88.3 % and 88.5 % accuracy, respectively. Human performance at pitch accent detection is generally taken to be between 85 % and 90%. Andrew Rosenberg, Julia Hirschberg |
INTERSPEECH | 2 |
| 2007 | Varying input segmentation for story boundary detection in English, Arabic and Mandarin broadcast newsabstractStory segmentation of news broadcasts has been shown to improve the accuracy of the subsequent processes such as question answering and information retrieval. In previous work, a decision tree trained on automatically extracted lexical and acoustic features was trained to predict story boundaries, using hypothesized sentence boundaries to define potential story boundaries. In this paper, we empirically evaluate several alternatives to choice of segmentation on three languages: English, Mandarin and Arabic. Our results suggest that the best performance can be achieved by using 250ms pause-based segmentation or sentence boundaries determined using a very low confidence score threshold. Andrew Rosenberg, Mehrbod Sharifi, Julia Hirschberg |
INTERSPEECH | 3 |
| 2007 | Accessing speech data using strategic fixation
Steve Whittaker 0001, Julia Hirschberg |
Comput. Speech Lang. | 2 |
| 2006 | Combining Prosodic Lexical and Cepstral Systems for Deceptive Speech DetectionabstractWe report on machine learning experiments to distinguish deceptive from nondeceptive speech in the Columbia-SRI-Colorado (CSC) corpus. Specifically, we propose a system combination approach using different models and features for deception detection. Scores from an SVM system based on prosodic/lexical features are combined with scores from a Gaussian mixture model system based on acoustic features, resulting in improved accuracy over the individual systems. Finally, we compare results from the prosodic-only SVM system using features derived either from recognized words or from human transcriptions. Martin Graciarena, Elizabeth Shriberg, Andreas Stolcke, Frank Enos, Julia Hirschberg, Sachin S. Kajarekar |
ICASSP (1) | 5 |
| 2006 | Personality factors in human deception detection: comparing human to machine performanceabstractPrevious studies of human performance in deception detection have found that humans generally are quite poor at this task, comparing unfavorably even to the performance of automated procedures.However, different scenarios and speakers may be harder or easier to judge.In this paper we compare human to machine performance detecting deception on a single corpus, the Columbia-SRI-Colorado Corpus of deceptive speech.On average, our human judges scored worse than chance -and worse than current best machine learning performance on this corpus.However, not all judges scored poorly.Based on personality tests given before the task, we find that several personality factors appear to correlate with the ability of a judge to detect deception in speech. Frank Enos, Stefan Benus, Robin L. Cautin, Martin Graciarena, Julia Hirschberg, Elizabeth Shriberg |
INTERSPEECH | 5 |
| 2006 | Effect of genre, speaker, and word class on the realization of given and new informationabstractThere is much evidence in the literature that speakers tend to deaccent discourse-given entities, while accenting new ones.However, speakers do not always follow this simple strategy and the causes for such variation are not yet well understood.In this paper, we describe several new forms of variability in the relationship between given/new information and accenting behavior, variation due to individual differences and to word class.We present results indicating that different speakers have different strategies for making new words prominent.We analyze two word-classes -nouns and verbs -in a corpus of spontaneous and read direction-giving monologues, and show that speakers use different combinations of pitch, intensity and inter-word pauses to distinguish between given and new information.Most interestingly, we find that in both genres all speakers tend to produce given verbs with higher intensity than new verbs. Agustín Gravano, Julia Hirschberg |
INTERSPEECH | 2 |
| 2006 | Detecting question-bearing turns in spoken tutorial dialoguesabstractCurrent speech-enabled Intelligent Tutoring Systems do not model student question behavior the way human tutors do, despite evidence indicating the importance of doing so.Our study examined a corpus of spoken tutorial dialogues collected for development of ITSpoke, an Intelligent Tutoring Spoken Dialogue System.The authors extracted prosodic, lexical, syntactic, and student and task dependent information from student turns.Results of running 5-fold cross validation machine learning experiments using AdaBoosted C4.5 decision trees show prediction of student question-bearing turns at a rate of 79.7%.The most useful features were prosodic, especially the pitch slope of the last 200 milliseconds of the student turn.Student pre-test score was the most-used feature.Findings indicate that using turn-based units is acceptable for incorporating question detection capability into practical Intelligent Tutoring Systems. Jackson Liscombe, Jennifer J. Venditti, Julia Hirschberg |
INTERSPEECH | 3 |
| 2006 | Soundbite detection in broadcast news domainabstractIn this paper, we present results of a study designed to identify SOUNDBITES in Broadcast News. We describe a Conditional Random Field-based model for the detection of these included speech segments uttered by individuals who are interviewed or who are the subject of a news story. Our goal is to identify direct quotations in spoken corpora which can be directly attributable to particular individuals, as well as to associate these soundbites with their speakers. We frame soundbite detection as a binary classification problem in which each turn is categorized either as a soundbite or not. We use lexical, acoustic/prosodic and structural features on a turn level to train a CRF. We performed a 10-fold cross validation experiment in which we obtained an accuracy of 67.4 % and an F-measure of 0.566 which is 20.9 % and 38.6 % higher than a chance baseline. Index Terms: soundbite detection, speaker roles, speech summarization, information extraction. Sameer Maskey, Julia Hirschberg |
INTERSPEECH | 2 |
| 2006 | On the correlation between energy and pitch accent in read English speechabstractIn this paper, we describe a set of experiments that examine the correlation between energy and pitch accent.We tested the discriminative power of the energy component of frequency subbands with a variety of frequencies and bandwidths on read speech spoken by four native speakers of Standard American English, using an analysis by classification approach.We found that the frequency region most robust to speaker differences is between 2 and 20 bark.Across all speakers, using only energy features we were able to predict pitch accent in read speech with accuracy of 81.9%. Andrew Rosenberg, Julia Hirschberg |
INTERSPEECH | 2 |
| 2006 | Text-independent cross-language voice conversionabstractSo far, cross-language voice conversion requires at least one bilingual speaker and parallel speech data to perform the training. This paper shows how these obstacles can be overcome by means of a recently presented text-independent training method based on unit selection. The new method is evaluated in the framework of the European speech-to-speech translation project TC-Star and achieves a performance similar to that of text-dependent intralingual voice conversion. David Suendermann-Oeft, Harald Höge, Antonio Bonafonte, Hermann Ney, Julia Hirschberg |
INTERSPEECH | 5 |
| 2006 | Intonational cues to student questions in tutoring dialogsabstractSuccessful Intelligent Tutoring Systems (ITSs) must be able to recognize when their students are asking a question.They must identify question form as well as function in order to respond appropriately.Our study examines whether intonational features, specifically, F0 height and rise range, are useful cues to student question type in a corpus of 643 American English questions.Results show a quantitative effect of both form and function.In addition, among clarification-seeking questions, we observed differences based on the type of clarification being sought. 1 Jennifer J. Venditti, Julia Hirschberg, Jackson Liscombe |
INTERSPEECH | 2 |
| 2006 | Summarizing Speech Without Text Using Hidden Markov Models
Sameer Maskey, Julia Hirschberg |
HLT-NAACL | 2 |
| 2006 | Story Segmentation of Broadcast News in English, Mandarin and Arabic
Andrew Rosenberg, Julia Hirschberg |
HLT-NAACL | 2 |
| 2006 | Characterizing and Predicting Corrections in Spoken Dialogue SystemsabstractThis article focuses on the analysis and prediction of corrections, defined as turns where a user tries to correct a prior error made by a spoken dialogue system. We describe our labeling procedure of various corrections types and statistical analyses of their features in a corpus collected from a train information spoken dialogue system. We then present results of machine-learning experiments designed to identify user corrections of speech recognition errors. We investigate the predictive power of features automatically computable from the prosody of the turn, the speech recognition process, experimental conditions, and the dialogue history. Our best-performing features reduce classification error from baselines of 25.70–28.99% to 15.72%. Diane J. Litman, Julia Hirschberg, Marc Swerts |
Comput. Linguistics | 2 |
| 2005 | From text to speech summarizationabstractIn this paper, we present approaches used in text summarization, showing how they can be adapted for speech summarization and where they fall short. Informal style and apparent lack of structure in speech mean that the typical approaches used for text summarization must be extended for use with speech. We illustrate how features derived from speech can help determine summary content within two ongoing summarization projects at Columbia University. Kathy McKeown, Julia Hirschberg, Michel Galley, Sameer Maskey |
ICASSP (5) | 2 |
| 2005 | Distinguishing deceptive from non-deceptive speechabstractTo date, studies of deceptive speech have largely been confined to descriptive studies and observations from subjects, researchers, or practitioners, with few empirical studies of the specific lexical or acoustic/prosodic features which may characterize deceptive speech.We present results from a study seeking to distinguish deceptive from non-deceptive speech using machine learning techniques on features extracted from a large corpus of deceptive and non-deceptive speech.This corpus employs an interview paradigm that includes subject reports of truth vs. lie at multiple temporal scales.We present current results comparing the performance of acoustic/prosodic, lexical, and speaker-dependent features and discuss future research directions. Julia Hirschberg, Stefan Benus, Jason M. Brenier, Frank Enos, Sarah Friedman, Sarah Gilman, Cynthia Girand, Martin Graciarena, Andreas Kathol, Laura A. Michaelis, Bryan L. Pellom, Elizabeth Shriberg, Andreas Stolcke |
INTERSPEECH | 1 |
| 2005 | Detecting certainness in spoken tutorial dialoguesabstractWhat role does affect play in spoken tutorial systems and is it automatically detectable? We investigated the classification of student certainness in a corpus collected for ITSPOKE, a speech-enabled Intelligent Tutorial System (ITS). Our study suggests that tutors respond to indications of student uncertainty differently from student certainty. Results of machine learning experiments indicate that acoustic-prosodic features can distinguish student certainness from other student states. A combination of acoustic-prosodic features extracted at two levels of intonational analysis --- breath groups and turns --- achieves 76.42% classification accuracy, a 15.8% relative improvement over baseline performance. Our results suggest that student certainness can be automatically detected and utilized to create better spoke dialog ITSs. Jackson Liscombe, Julia Hirschberg, Jennifer J. Venditti |
INTERSPEECH | 2 |
| 2005 | Comparing lexical, acoustic/prosodic, structural and discourse features for speech summarizationabstractWe present results of an empirical study of the usefulness of different types of features in selecting extractive summaries of news broadcasts for our Broadcast News Summarization System. We evaluate lexical, prosodic, structural and discourse features as predictors of those news segments which should be included in a summary. We show that a summarization system that uses a combination of these feature sets produces the most accurate summaries, and that a combination of acoustic /prosodic and structural features are enough to build a `good' summarizer when speech transcription is not available. Sameer Maskey, Julia Hirschberg |
INTERSPEECH | 2 |
| 2005 | Acoustic/prosodic and lexical correlates of charismatic speechabstractCharisma, the ability to command authority on the basis of personal qualities, is more difcult to dene than to identify. How do charismatic leaders such as Fidel Castro or Pope John Paul II attract and retain their followers? We present results of an analysis of subjective ratings of charisma from a corpus of American political speech. We identify the associations be- tween charisma ratings and ratings of other personal attributes. We also examine acoustic/prosodic and lexical features of this speech and correlate these with charisma ratings. Andrew Rosenberg, Julia Hirschberg |
INTERSPEECH | 2 |
| 2005 | Do summaries help?abstractWe describe a task-based evaluation to determine whether multi-document summaries measurably improve user performance whe using online news browsing systems for directed research. We evaluated the multi-document summaries generated by Newsblaster, a robust news browsing system that clusters online news articles and summarizes multiple articles on each event. Four groups of subjects were asked to perform the same time-restricted fact-gathering tasks, reading news under different conditions: no summaries at all, single sentence summaries drawn from one of the articles, Newsblaster multi-document summaries, and human summaries. Our results show that, in comparison to source documents only, the quality of reports assembled using Newsblaster summaries was significantly better and user satisfaction was higher with both Newsblaster and human summaries. Kathy McKeown, Rebecca J. Passonneau, David K. Elson, Ani Nenkova, Julia Hirschberg |
SIGIR | 5 |
| 2005 | Error handling in spoken dialogue systems
Rolf Carlson, Julia Hirschberg, Marc Swerts |
Speech Commun. | 2 |
| 2005 | Cues to upcoming Swedish prosodic boundaries: Subjective judgment studies and acoustic correlates
Rolf Carlson, Julia Hirschberg, Marc Swerts |
Speech Commun. | 2 |
| 2004 | Identifying Agreement and Disagreement in Conversational Speech: Use of Bayesian Networks to Model Pragmatic DependenciesabstractWe describe a statistical approach for modeling agreements and disagreements in conversational interaction. Our approach first identifies adjacency pairs using maximum entropy ranking based on a set of lexical, durational, and structural features that look both forward and backward in the discourse. We then classify utterances as agreement or disagreement using these adjacency pairs and features that represent various pragmatic influences of previous agreement or disagreement on the current utterance. Our approach achieves 86.9% accuracy, a 4.9% increase over previous work. Michel Galley, Kathy McKeown, Julia Hirschberg, Elizabeth Shriberg |
ACL | 3 |
| 2004 | Prosodic and other cues to speech recognition failures
Julia Hirschberg, Diane J. Litman, Marc Swerts |
Speech Commun. | 1 |
| 2004 | Introduction to the Special Issue on Spontaneous Speech Processing
Sadaoki Furui, Mary E. Beckman, Julia Hirschberg, Shuichi Itahashi, Tatsuya Kawahara, Satoshi Nakamura 0001, Shri Narayanan |
IEEE Trans. Speech Audio Process. | 3 |
| 2003 | Classifying subject ratings of emotional speech using acoustic featuresabstractThis paper presents results from a study examining emotional speech using acoustic features and their use in automatic machine learning classification.In addition, we propose a classification scheme for the labeling of emotions on continuous scales.Our findings support those of previous research as well as indicate possible future directions utilizing spectral tilt and pitch contour to distinguish emotions in the valence dimension. Emotion Recognition Survey: Sound File 1 of 47not at all a little somewhat quite extremely How frustrated does this person sound?How confident does this person sound?How interested does this person sound?How sad does this person sound?How happy does this person sound?How friendly does this person sound?How angry does this person sound?How anxious does this person sound?How bored does this person sound?How encouraging does this person sound? Jackson Liscombe, Jennifer J. Venditti, Julia Hirschberg |
INTERSPEECH | 3 |
| 2003 | Automatic summarization of broadcast news using structural featuresabstractWe present a method of summarizing brodcast news that is not affected by word errors in the transcript of broadcast news. We built a graphical model to represent the probability distribution and dependencies among the structural features. We trained the model by filling the probability table by multinomial counts on the training sentences of summary. Then we ranked the new test segments of broadcast news and extracted the highest ranked ones as a summary. 1. Sameer Maskey, Julia Hirschberg |
INTERSPEECH | 2 |
| 2002 | SCANMail: a voicemail interface that makes speech browsable, readable and searchableabstractIncreasing amounts of public, corporate, and private speech data are now available on-line. These are limited in their usefulness, however, by the lack of tools to permit their browsing and search. The goal of our research is to provide tools to overcome the inherent difficulties of speech access, by supporting visual scanning, search, and information extraction. We describe a novel principle for the design of UIs to speech data: What You See Is Almost What You Hear (WYSIAWYH). In WYSIAWYH, automatic speech recognition (ASR) generates a transcript of the speech data. The transcript is then used as a visual analogue to that underlying data. A graphical user interface allows users to visually scan, read, annotate and search these transcripts. Users can also use the transcript to access and play specific regions of the underlying message. We first summarize previous studies of voicemail usage that motivated the WYSIAWYH principle, and describe a voicemail UI, SCANMail, that embodies WYSIAWYH. We report on a laboratory experiment and a two-month field trial evaluation. SCANMail outperformed a state of the art voicemail system on core voicemail tasks. This was attributable to SCANMail's support for visual scanning, search and information extraction. While the ASR transcripts contain errors, they nevertheless improve the efficiency of voicemail processing. Transcripts either provide enough information for users to extract key points or to navigate to important regions of the underlying speech, which they can then play directly Steve Whittaker 0001, Julia Hirschberg, Brian Amento, Litza A. Stark, Michiel Bacchiani, Philip L. Isenhour, Larry Stead, Gary Zamchick, Aaron E. Rosenberg |
CHI | 2 |
| 2002 | Exploring features from natural language generation for prosody modeling
Shimei Pan, Kathy McKeown, Julia Hirschberg |
Comput. Speech Lang. | 3 |
| 2002 | Communication and prosody: Functional aspects of prosody
Julia Hirschberg |
Speech Commun. | 1 |
| 2001 | Predicting User Reactions to System ErrorabstractThis paper focuses on the analysis and prediction of so-called aware sites, defined as turns where a user of a spoken dialogue system first becomes aware that the system has made a speech recognition error. We describe statistical comparisons of features of these aware sites in a train timetable spoken dialogue corpus, which reveal significant prosodic differences between such turns, compared with turns that 'correct' speech recognition errors as well as with 'normal' turns that are neither aware sites nor corrections. We then present machine learning results in which we show how prosodic features in combination with other automatically available features can predict whether or not a user turn was a normal turn, a correction, and/or an aware site. Diane J. Litman, Julia Hirschberg, Marc Swerts |
ACL | 2 |
| 2001 | SCANMail: browsing and searching speech data by contentabstractIncreasing amounts of public, corporate, and private audio data are available for use, but limited in usefulness by the lack of tools to permit their browsing and search. In this paper, we describe SCANMail, a system that employs automatic speech recognition, information retrieval, information extraction, and human computer interaction technology to permit users to browse and search their voicemail messages by content through a graphical user interface interface. The SCANMail client also provides note-taking capabilities as well as browsing and querying features. A CallerId server also proposes caller names from existing caller acoustic models and is trained from user feedback. An Email server sends the original message plus its transcription to a mailing address specified in the user's profile. 1. Julia Hirschberg, Michiel Bacchiani, Donald Hindle, Philip L. Isenhour, Aaron E. Rosenberg, Litza A. Stark, Larry Stead, Steve Whittaker 0001, Gary Zamchick |
INTERSPEECH | 1 |
| 2001 | Learning prosodic features using a tree representationabstractWe describe experiments designed to learn associations between two types of intonational features, pitch accent and phrasing, from a tree-based corpus annotated with various intonational and syntactic features, for a concept-to-speech system. We show that using novel tree-based features improves the quality of boundary prediction over using only the linear orderbased features normally used in text-to-speech. Julia Hirschberg, Owen Rambow |
INTERSPEECH | 1 |
| 2001 | Semantic abnormality and its realization in spoken languageabstractIn this paper we investigate the relationship between various lexical and prosodic features and semantic abnormality, the occurrence of unusual or unexpected events, in generating speech for MAGIC, which employs a Concept-to-Speech system to generate post-operative reports for patients who have undergone bypass surgery. Using the speech corpus collected for this application, we conducted empirical analysis to systematically discover significantly correlated prosodic and lexical features. The automatically learned abnormality model not only can be used in building comprehensive prosody prediction systems for Concept-to-Speech generation, but also help identify unusual information during speech analysis and understanding. Shimei Pan, Kathy McKeown, Julia Hirschberg |
INTERSPEECH | 3 |
| 2001 | Caller identification for the SCANMail voicemail browserabstractSCANMail is a prototype system developed at AT&T Labs for the purpose of providing useful tools for managing and searching through voicemail messages. Content is extracted from voicemail messages using various speech and text processing tools. One such content category is the identity of the message caller. This paper describes CallerID, the server tool attached to SCANMail for the purpose of providing caller labels for voicemail messages. CallerID make use of text independent speaker recognition techniques. Two kinds of requests are handled by the CallerID server. A request triggered by the arrival of a new voicemail message results in the processing of the message to score it against the models of callers assigned to the user (recipient) in order to propose the identity of the caller. A second request is initiated by a user who provides a caller label for a message he/she has reviewed. CallerID processes the message and uses it to train or adapt a speaker model for the caller whose label is provided. The paper describes in detail the CallerID functions and provides some results of performance evaluations of the caller identification capability. Aaron E. Rosenberg, Julia Hirschberg, Michiel Bacchiani, Sarangarajan Parthasarathy, Philip L. Isenhour, Larry Stead |
INTERSPEECH | 2 |
| 2001 | Identifying User Corrections Automatically in Spoken Dialogue Systems
Julia Hirschberg, Diane J. Litman, Marc Swerts |
NAACL | 1 |
| 2001 | Automatic ToBI prediction and alignment to speed manual labeling of prosody
Ann K. Syrdal, Julia Hirschberg, Julie Tevis McGory, Mary E. Beckman |
Speech Commun. | 2 |
| 2001 | The character, value, and management of personal paper archivesabstractWe explored general issues concerning personal information management by investigating the characteristics of office workers' paper-based information, in an industrial research environment. we examined the reasons people collect paper, types of data they collect, problems encountered in handling paper, and strategies used for processing it. We tested three specific hypotheses in the course of an office move. The greater availability of public digital data along with changes in people's jobs or interests should lead to wholescale discarding of paper data, while preparing for the move. Instead we found workers kept large, highly valued papar archives. We also expected that the major part of people's personal archives would be unique documents. However, only 49% of people's archives were unique documents, the remainder being copies of publicly available data and unread information, and we explore reasons for this. We examined the effects of paper-processing strategies on archive structure. We discovered different paper-processing strategies ( filing and piling )that were relatively independent of job type. We predicated that filers' attempted to evaluate and catergorize incoming documents would produce smaller archives that were accessed frequently. Contrary to our predictions, filers amassed more information, and accessed it less frequently than pilers. We argue that filers may engage in premature filing : to clear their workspace, they archives information that later turns out to be of low value. Given the effort involved in organzing data, they are also loath to discard filed information, even when its value is uncertain. We discuss the implications of this research for digital personal information management. Steve Whittaker 0001, Julia Hirschberg |
ACM Trans. Comput. Hum. Interact. | 2 |
| 2000 | Modeling Local Context for Pitch Accent PredictionabstractPitch accent placement is a major topic in intonational phonology research and its application to speech synthesis. What factors influence whether or not a word is made intonationally prominent or not is an open question. In this paper, we investigate how one aspect of a word's local context --- its collocation with neighboring words --- influences whether it is accented or not. Results of experiments on two transcribed speech corpora in a medical domain show that such collocation information is a useful predictor of pitch accent placement. Shimei Pan, Julia Hirschberg |
ACL | 2 |
| 2000 | Jotmail: a voicemail interface that enables you to see what was saidabstractVoicemail is a pervasive, but under-researched tool for workplace communication. Despite potential advantages of voicemail over email, current phone-based voicemail UIs are highly problematic for users. We present a novel, Web-based, voicemail interface, Jotmail. The design was based on data from several studies of voicemail tasks and user strategies. The GUI has two main elements: (a) personal annotations that serve as a visual analogue to underlying speech; (b) automatically derived message header information. We evaluated Jotmail in an 8-week field trial, where people used it as their only means for accessing voicemail. Jotmail was successful in supporting most key voicemail tasks, although users' electronic annotation and archiving behaviors were different from our initial predictions. Our results argue for the utility of a combination of annotation based indexing and automatically derived information, as a general technique for accessing speech archives. Steve Whittaker 0001, Julia Hirschberg, Urs Muller |
CHI | 3 |
| 2000 | Improving intonational phrasing with syntactic informationabstractThe prediction of intonational phrase boundaries from raw text is an important step for a text-to-speech system: locating where to place short pauses enables more natural sounding speech, that can be more easily understood. We improved upon earlier work [Hirschberg and Prieto, 1996] by adding syntactic information gained from a high-accuracy parser [Collins, 1999]. We report significant improvement using various experimental setups. We also show that our improved method comes close to interannotator agreement. Philipp Koehn, Steven P. Abney, Julia Hirschberg, Michael Collins 0001 |
ICASSP | 3 |
| 2000 | Generalizing prosodic prediction of speech recognition errorsabstractSince users of spoken dialogue systems have difficulty correcting system misconceptions, it is important for automatic speech recognition (ASR) systems to know when their best hypothesis is incorrect. We compare results of previous experiments which showed that prosody improves the detection of ASR errors to experiments with a new system and new domain, the W99 conference registration system. Our new results again show that prosodic features can improve prediction of ASR misrecognitions over the use of other standard techniques for ASR rejection. Julia Hirschberg, Diane J. Litman, Marc Swerts |
INTERSPEECH | 1 |
| 2000 | Foldering voicemail messages by caller using text independent speaker recognitionabstractThe ability to automatically scan voicemail messages for content and caller identity cues would be a useful service. This paper describes a system which automatically files voicemail messages into caller folders using text independent speaker recognition techniques. Callers are represented by Gaussian mixture models (GMM’s). The speech for an incoming message is processed and scored against caller models created for a subscriber. A message whose matching score exceeds a threshold is filed in the matching caller folder; otherwise it is tagged as “unknown”. The subscriber has the ability to listen to an “unknown” message and file it in the proper folder, if it exists, or create a new folder, if it does not. Such subscriber labelled messages are used to train and adapt caller models. The system has been evaluated on a database of voicemail messages collected at AT&T Labs. A set of 20 callers from this database is designated as “ingroup”. Each of these callers has recorded at least 20 messages totalling 10 or more minutes in duration. A distinct set of 220 messages, each from a dierent caller, are designated as “outgroup”. Representative performance figures with threshold parameters set to ensure that outgroup acceptance is low compared with ingroup rejection are the following. The average ingroup message rejection rate is 11.0% and the average ingroup message confusion rate (matching the wrong caller) is 1.0%, while the average outgroup message accept rate is 2.7%. Aaron E. Rosenberg, Sarangarajan Parthasarathy, Julia Hirschberg, Steve Whittaker 0001 |
INTERSPEECH | 3 |
| 2000 | ASR satisficing: the effects of ASR accuracy on speech retrievalabstractWe examine how differences in the accuracy of Automatic Speech Recognition transcripts affect users' ability to use these in tasks requiring the retrieval of speech "documents". We compare performance measures, processing strategies, and preference data for subjects using transcripts and speech data to perform a series of relevance judgment and summary tasks on transcripts with different levels of accuracy. Results show effects for transcript quality on solution accuracy, time to solution, amount of speech played for the task, likelihood of subjects abandoning use of a transcript, and subject perceptions of task difficulty, transcript utility, readability, and comprehensibility. 1. INTRODUCTION Research on Automatic Speech Recognition (ASR) generally assumes perfect transcription accuracy to be its holy grail. The more accurately a system can transcribe an utterance, the better that system performs (modulo time factors) Recently this assumption is being challenged, however. Different... Litza A. Stark, Steve Whittaker 0001, Julia Hirschberg |
INTERSPEECH | 3 |
| 2000 | Corrections in spoken dialogue systemsabstractThis study analyzes user corrections of system errors in the TOOT spoken dialogue system. We find that corrections differ from noncorrections prosodically, in ways consistent with hyperarticulated speech, although many corrections are not hyperarticulated. Yet both are misrecognized more frequently than non-corrections --- though no more likely to be rejected by the system. Corrections more distant from the error they correct tend to exhibit greater prosodic differences, and also to be recognized more poorly. System dialogue strategy affects users' choice of correction type, suggesting that strategy-specific methods of detecting or coaching users on corrections may be useful. Strategies that produce longer tasks but fewer misrecognitions and subsequent corrections are preferred by users. 1. INTRODUCTION Since spoken dialogue systems often make mistakes in recognizing user input, accurate methods of detecting and correcting system errors are essential to supporting successful interact... Marc Swerts, Diane J. Litman, Julia Hirschberg |
INTERSPEECH | 3 |
| 1999 | SCAN: Designing and Evaluating User Interfaces to Support Retrieval From Speech ArchivesabstractPrevious examinations of search in textual archives have assumed that users first retrieve a ranked set of documents relevant to their query, and then visually scan through these documents, to identify the information they seek.While document scanning is possible in text, it is much more laborious in speech archives, due to the inherently serial nature of speech.Yet, in developing tools for speech access, little attention has so far been paid to users' problems in scanning and extracting information from within "speech documents".We demonstrate the extent of these problems in two user studies.We show that users experience severe problems with local navigation in extracting relevant information from within "speech documents".Based on these results, we propose a new user interface (UI) design paradigm: What You See Is (Almost) What You Hear, (WYSIAWYH) -a multimodal method for accessing speech archives.This paradigm presents a visual analogue to the underlying speech, enabling visual scanning for effective local navigation.We empirically evaluate a UI based on this paradigm.We compare our WYSIAWYH UI with a visual "tape recorder", in relevance ranking, fact-finding, and summarization tasks involving broadcast news data.Our findings indicate that an interface supporting local navigation multimodally helps relevance ranking and fact-finding, but not summarization.We analyze the reasons for system success and identify outstanding research issues in UI design for speech archives. Steve Whittaker 0001, Julia Hirschberg, Donald Hindle, Fernando Pereira 0003, Amit Singhal 0001 |
SIGIR | 2 |
| 1998 | SCAN - speech content based audio navigator: a system overview
Donald Hindle, Julia Hirschberg, Ivan Magrin-Chagnolleau, Christine H. Nakatani, Fernando Pereira 0003, Amit Singhal 0001, Steve Whittaker 0001 |
ICSLP | 3 |
| 1998 | Acoustic indicators of topic segmentationabstractThe segmentation of text and speech into topics and subtopics is an important step in document interpretation. For text, formatting information, such as headings and paragraphing, is available to aid in this endeavor, although this information is by no means su cient. For speech, the task is even more di cult. We present results of the application of machine learning techniques to the automatic identi cation of intonational phrases beginning and ending 'topics ' determined independently by annotators for two corpora | the Boston Directions Corpus and the Broadcast News (HUB-4) DARPA/NIST database. 1. Julia Hirschberg, Christine H. Nakatani |
ICSLP | 1 |
| 1998 | Now you hear it, now you don't: empirical studies of audio browsing behavior behavior
Christine H. Nakatani, Steve Whittaker 0001, Julia Hirschberg |
ICSLP | 3 |
| 1998 | What you see is (almost) what you hear: design principles for user interfaces for accessing speech archivesabstractDespite the recent growth and potential utility of speech archives, we currently lack tools for effective archival access. Previous research on search of textual archives has assumed that the system goal should be to retrieve sets of relevant documents, leaving users to visually scan through those documents to identify relevant information. However, in previous work we show that in accessing real speech archives, it is insufficient to only retrieve "document" sets [9,10]. Users experience huge problems of local navigation in attempting to extract relevant information from within speech "documents". These studies also show that users address these problems by taking handwritten notes. These notes detail both the content of the speech and serve as indices to help access relevant regions of the archive. From these studies we derive a new principle for the design of speech access systems: What You See Is (Almost) What You Hear. We present a new user interface to a broadcast news archive, d... Steve Whittaker 0001, Julia Hirschberg, Christine H. Nakatani |
ICSLP | 3 |
| 1998 | Speech Research: Near and Not-so-near Results and What They Might Mean for IUI (Panel)abstractNo abstract available. Candace L. Sidner, Alex Acero, Janet E. Cahn, Julia Hirschberg, Salim Roukos |
IUI | 4 |
| 1996 | A Prosodic Analysis of Discourse Segments in Direction-Giving MonologuesabstractThis paper reports on corpus-based research into the relationship between intonational variation and discourse structure. We examine the effects of speaking style (read versus spontaneous) and of discourse segmentation method (text-alone versus text-and-speech) on the nature of this relationship. We also compare the acoustic-prosodic features of initial, medial, and final utterances in a discourse segment. Julia Hirschberg, Christine H. Nakatani |
ACL | 1 |
| 1996 | Training intonational phrasing rules automatically for English and Spanish text-to-speech
Julia Hirschberg, Pilar Prieto |
Speech Commun. | 1 |
| 1994 | Evaluation of prosodic transcription labeling reliability in the tobi framework
John F. Pitrelli, Mary E. Beckman, Julia Hirschberg |
ICSLP | 3 |
| 1994 | Segmental effects on timing and height of pitch contours
Jan P. H. van Santen, Julia Hirschberg |
ICSLP | 2 |
| 1993 | A Speech-First Model for Repair Detection and CorrectionabstractInterpreting fully natural speech is an important goal for spoken language understanding systems. However, while corpus studies have shown that about 10% of spontaneous utterances contain self-corrections, or REPAIRS, little is known about the extent to which cues in the speech signal may facilitate repair processing. We identify several cues based on acoustic and prosodic analysis of repairs in a corpus of spontaneous speech, and propose methods for exploiting these cues to detect and correct repairs. We test our acoustic-prosodic cues with other lexical cues to repair identification and find that precision rates of 89--93% and recall of 78--83% can be achieved, depending upon the cues employed, from a prosodically labeled corpus. Christine H. Nakatani, Julia Hirschberg |
ACL | 2 |
| 1993 | A speech-first model for repair identification in spoken language systems
Julia Hirschberg, Christine H. Nakatani |
EUROSPEECH | 1 |
| 1993 | Deaccentuation and persistence of grammatical function and surface position
Julia Hirschberg, Jacques M. B. Terken |
EUROSPEECH | 1 |
| 1993 | Pitch Accent in Context: Predicting Intonational Prominence from Text
Julia Hirschberg |
Artif. Intell. | 1 |
| 1993 | Empirical Studies on the Disambiguation of Cue Phrases
Julia Hirschberg, Diane J. Litman |
Comput. Linguistics | 1 |
| 1992 | Some intonational characteristics of discourse structureabstractThis paper reports on a study of the relationship between acoustic-prosodic variation and discourse structure, as determined from an independent model of discourse. We present results of two pilot studies. Our corpus consisted of three AP news stories recorded by a professional speaker. Discourse structure was labeled by subjects either from text alone or from text (with all orthographic markings except sentence-final punctuation removed) and speech, following Grosz & Sidner 1986; average inter-labeler agreement for structural elements varied from 74.3%-95.1%, depending upon feature. These elements of global structure, together with elements of local structure such as parentheticals and attributive tags, were correlated with variation in into- This research was partly supported by NSF grant #IRI-9009018. national and acoustic features such as pitch range, contour, timing, and amplitude. We found statistically significant associations between aspects of pitch range, amplitude, and t... Barbara J. Grosz, Julia Hirschberg |
ICSLP | 2 |
| 1992 | TOBI: a standard for labeling English prosody
Kim E. A. Silverman, Mary E. Beckman, John F. Pitrelli, Mari Ostendorf, Colin W. Wightman, Patti Price, Janet B. Pierrehumbert, Julia Hirschberg |
ICSLP | 8 |
| 1992 | A corpus-based synthesizer
Richard Sproat, Julia Hirschberg, David Yarowsky |
ICSLP | 2 |
| 1991 | Predicting Intonational Phrasing from TextabstractDetermining the relationship between the intonational characteristics of an utterance and other features inferable from its text is important both for speech recognition and for speech synthesis. This work investigates the use of text analysis in predicting the location of intonational phrase boundaries in natural speech, through analyzing 298 utterances from the DARPA Air Travel Information Service database. For statistical modeling, we employ Classification and Regression Tree (CART) techniques. We achieve success rates of just over 90%, representing a major improvement over other attempts at boundary prediction from unrestricted text. Michelle Q. Wang, Julia Hirschberg |
ACL | 2 |
| 1991 | Using text analysis to predict intonational boundaries
Julia Hirschberg |
EUROSPEECH | 1 |
| 1990 | Accent and Discourse Context: Assigning Pitch Accent in Synthetic Speech
Julia Hirschberg |
AAAI | 1 |
| 1990 | Disambiguating Cue Phrases in Text and Speech
Diane J. Litman, Julia Hirschberg |
COLING | 2 |
| 1988 | Assigning Intonational Features in Synthesized Spoken DirectionsabstractSpeakers convey much of the information hearers use to interpret discourse by varying prosodic features such as PHRASING, PITCH ACCENT placement, TUNE, and PITCH RANGE. The ability to emulate such variation is crucial to effective (synthetic) speech generation. While text-to-speech synthesis must rely primarily upon structural information to determine appropriate intonational features, speech synthesized from an abstract representation of the message to be conveyed may employ much richer sources. The implementation of an intonation assignment component for Direction Assistance, a program which generates spoken directions, provides a first approximation of how recent models of discourse structure can be used to control intonational variation in ways that build upon recent research in intonational meaning. The implementation further suggests ways in which these discourse models might be augmented to permit the assignment of appropriate intonational features. James Raymond Davis, Julia Hirschberg |
ACL | 2 |
| 1987 | Now let's Talk about Now; Identifying Cue Phrases IntonationallyabstractCue phrases are words and phrases such as now and by the way which may be used to convey explicit information about the structure of a discourse.However, while cue phrases may convey discourse structure, each may also be used to different effect.The question of how speakers and hearers distinguish between such uses of cue phrases has not been addressed in discourse studies to date.Based on a study of now in natural recorded discourse, we propose that cue and non-cue usage can be distinguished intonationally, on the basis of phrasing and accent. Julia Hirschberg, Diane J. Litman |
ACL | 1 |
| 1987 | Intonation and the Intentional Structure of Discourse
Julia Hirschberg, Diane J. Litman, Janet B. Pierrehumbert, G. Ward |
IJCAI | 1 |
| 1986 | The intonational Structuring of DiscourseabstractWe propose a mapping between prosodic phenomena and semantico-pragmatic effects based upon the hypothesis that intonation conveys information about the intentional as well as the attentional structure of discourse. In particular, we discuss how variations in pitch range and choice of accent and tune can help to convey such information as: discourse segmentation and topic structure, appropriate choice of referent, the distinction between 'given' and 'new' information, conceptual contrast or parallelism between mentioned items, and subordination relationships between propositions salient in the discourse. Our goals for this research are practical as well as theoretical. In particular, we are investigating the problem of intonational assignment in synthetic speech. Julia Hirschberg, Janet B. Pierrehumbert |
ACL | 1 |
| 1984 | Toward a Redefinition of Yes/No QuestionsabstractWhile both theoretical and empirical studies of question-answering have revealed the inadequacy of traditional definitions of yes-no questions (YNQs), little progress has been made toward a more satisfactory redefinition. This paper reviews the limitations of several proposed revisions. It proposes a new definition of YNQs based upon research on a type of conversational implicature, termed here scalar implicature, that helps define appropriate responses to YNQs. By representing YNQs as scalar queries it is possible to support a wider variety of system and user responses in a principled way. Julia Hirschberg |
COLING | 1 |
| 1982 | User Participation in the Reasoning Processes of Expert Systems
Martha E. Pollack, Julia Hirschberg, Bonnie L. Webber |
AAAI | 2 |