Gerald Penn

dblp:37/1531 · DBLP profile ↗
← Back
77ranked-venue papers
14as first author
13since 2021 · last 2026
0000-0003-3553-8305ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 54 · 11 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 11 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-authorTheory of computation · 2 · 2 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2026 Beyond Step Pruning: Information Theory Based Step-level Optimization for Self-Refining Large Language Models
abstract
Large language models (LLMs) have shown impressive capabilities in natural language tasks, yet they continue to struggle with multi-step mathematical reasoning, where correctness depends on a precise chain of intermediate steps. Preference optimization methods such as Direct Preference Optimization (DPO) have improved answer-level alignment, but they often overlook the reasoning process itself, providing little supervision over intermediate steps that are critical for complex problem-solving. Existing fine-grained approaches typically rely on strong annotators or reward models to assess the quality of individual steps. However, reward models are vulnerable to reward hacking. To address this, we propose ISLA, a reward-model-free framework that constructs step-level preference data directly from SFT gold traces. ISLA also introduces a self-improving pruning mechanism that identifies informative steps based on two signals: their marginal contribution to final accuracy (relative accuracy) and the model’s uncertainty, inspired by the concept of information gain. Empirically, ISLA achieves better performance than DPO while using only 12% of the training tokens, demonstrating that careful step-level selection can significantly improve both reasoning accuracy and training efficiency.
Jinman Zhao, Erxue Min, Ziheng Li 0003, Zexu Sun, Hengyi Cai, Shuaiqiang Wang, Xu Chen 0017, Gerald Penn
AAAI9
2025 Making the Write Connections: Linking Writing Support Tools with Writer Needs
abstract
This work sheds light on whether and how creative writers' needs are met by existing research and commercial writing support tools (WST). We conducted a need finding study to gain insight into the writers' process during creative writing through a qualitative analysis of the response from an online questionnaire and Reddit discussions on r/Writing. Using a systematic analysis of 115 tools and 67 research papers, we map out the landscape of how digital tools facilitate the writing process. Our triangulation of data reveals that research predominantly focuses on the writing activity and overlooks pre-writing activities and the importance of visualization. We distill 10 key takeaways to inform future research on WST and point to opportunities surrounding underexplored areas. Our work offers a holistic and up-to-date account of how tools have transformed the writing process, guiding the design of future tools that address writers' evolving and unmet needs.
Damien Masson, Young-Ho Kim, Gerald Penn, Fanny Chevalier
CHI4
2025 Inside-Outside Algorithm for Probabilistic Product-Free Lambek Categorial Grammar
abstract
The inside-outside algorithm is widely utilized in statistical models related to context-free grammars. It plays a key role in the EM estimation of probabilistic context-free grammars. In this work, we introduce an inside-outside algorithm for Probabilistic Lambek Categorical Grammar (PLCG)
Jinman Zhao, Gerald Penn
COLING2
2025 Sheaf Discovery with Joint Computation Graph Pruning and Flexible Granularity
abstract
In this paper, we introduce DiscoGP, a novel framework for extracting self-contained modular units, or sheaves, within neural language models (LMs).Sheaves extend the concept of functional circuits, a unit widely explored in interpretability research, by considering not only subsets of edges in an LM's computation graph but also the model's weight parameters.Our framework identifies sheaves through a gradient-based pruning algorithm that operates on both of these in such a way that reduces the original LM to a sparse skeleton that preserves certain core capabilities.Experimental results demonstrate that, across a range of linguistic and reasoning tasks, DiscoGP extracts sheaves that preserve 93-100% of a model's performance on the identified task while comprising only 1-7% of the original weights and connections.Furthermore, our analysis reveals that, compared to previously identified LM circuits, the sheaves discovered by DiscoGP exhibit superior modularity and functional fidelity.Extending our method to the neuron level also unveils novel insights into the inner workings of LLMs. 1 * Equal contribution. 1 The code and results of DiscoGP are available online: https://github.com/frankniujc/disco_gp.
Jingcheng Niu, Zining Zhu 0001, Gerald Penn
EMNLP5
2025 Tiny Budgets, Big Gains: Parameter Placement Strategy in Parameter Super-Efficient Fine-Tuning
abstract
In this work, we propose FoRA-UA, a novel method that, using only 1-5% of the standard LoRA's parameters, achieves state-ofthe-art performance across a wide range of tasks.Specifically, we explore scenarios with extremely limited parameter budgets and derive two key insights: (1) fix-sized sparse frequency representations approximate small matrices more accurately; and (2) with a fixed number of trainable parameters, introducing a smaller intermediate representation to approximate larger matrices results in lower construction error.These findings form the foundation of our FoRA-UA method.By inserting a small intermediate parameter set, we achieve greater model compression without sacrificing performance.We evaluate FoRA-UA across diverse tasks, including natural language understanding (NLU), natural language generation (NLG), instruction tuning, and image classification, demonstrating strong generalisation and robustness under extreme compression. 1
Jinman Zhao, Jiaru Li, Jingcheng Niu, Yulan Hu, Erxue Min, Gerald Penn
EMNLP7
2025 An Ecologically Valid Approach to Evaluating Online Gatekeepers
abstract
CAPTCHAs are commonly used as gatekeepers to protect online services from automated bots. Previous CAPTCHA evaluations have solely focussed on whether CAPTCHA challenges are unbreakable by bots but solvable by humans and have ignored a fundamental question: what effect do these gatekeepers have on their human users? Our study proposes a more realistic approach that complements existing CAPTCHA evaluations by measuring CAPTCHA’s effect on the main task that the user wants to complete. Through our experimental setup, we show that failing to correctly answer CAPTCHAs can turn away a significant percentage of legitimate, human users. We also show that these failures have temporary knock-on effects on the quality of tasks that users later perform. Methodologically, our study also reveals limitations of current approaches to batch evaluations of CAPTCHAs that do not accurately capture the effects that ordering and variance in difficulty have on users’ CAPTCHA-solving performance.
Sajad Shirali-Shahreza, Gerald Penn
Int. J. Hum. Comput. Interact.2
2024 LCGbank: A Corpus of Syntactic Analyses Based on Proof Nets
abstract
In syntactic parsing, proof nets are graphical structures that have the advantageous property of invariance to spurious ambiguities. Semantically-equivalent derivations correspond to a single proof net. Recent years have seen fresh interest in statistical syntactic parsing with proof nets, including the development of methods based on neural networks. However, training of statistical parsers requires corpora that provide ground-truth syntactic analyses. Unfortunately, there has been a paucity of corpora in formalisms for which proof nets are applicable, such as Lambek categorial grammar (LCG), a formalism related to combinatory categorial grammar (CCG). To address this, we leverage CCGbank and the relationship between LCG and CCG to develop LCGbank, an English-language corpus of syntactic analyses based on LCG proof nets. In contrast to CCGbank, LCGbank eschews type-changing and uses only categorial rules; the syntactic analyses thus provide fully compositional semantics, exploiting the transparency between syntax and semantics that so characterizes categorial grammars.
Aditya Bhargava, Timothy A. D. Fowler, Gerald Penn
LREC/COLING3
2024 A Generative Model for Lambek Categorial Sequents
abstract
In this work, we introduce a generative model, PLC+, for generating Lambek Categorial Grammar(LCG) sequents. We also introduce a simple method to numerically estimate the model’s parameters from an annotated corpus. Then we compare our model with probabilistic context-free grammars (PCFGs) and show that PLC+ simultaneously assigns a higher probability to a common corpus, and has greater coverage.
Jinman Zhao, Gerald Penn
LREC/COLING2
2024 What does the Knowledge Neuron Thesis Have to do with Knowledge?
abstract
We reassess the Knowledge Neuron (KN) Thesis: an interpretation of the mechanism underlying the ability of large language models to recall facts from a training corpus. This nascent thesis proposes that facts are recalled from the training corpus through the MLP weights in a manner resembling key-value memory, implying in effect that "knowledge" is stored in the network. Furthermore, by modifying the MLP modules, one can control the language model's generation of factual information. The plausibility of the KN thesis has been demonstrated by the success of KN-inspired model editing methods (Dai et al., 2022; Meng et al., 2022). We find that this thesis is, at best, an oversimplification. Not only have we found that we can edit the expression of certain linguistic phenomena using the same model editing methods but, through a more comprehensive evaluation, we have found that the KN thesis does not adequately explain the process of factual expression. While it is possible to argue that the MLP weights store complex patterns that are interpretable both syntactically and semantically, these patterns do not constitute "knowledge." To gain a more comprehensive understanding of the knowledge representation process, we must look beyond the MLP weights and explore recent models' complex layer structures and attention mechanisms.
Jingcheng Niu, Zining Zhu 0001, Gerald Penn
ICLR4
2024 Quantifying the Role of Textual Predictability in Automatic Speech Recognition
abstract
A long-standing question in automatic speech recognition research is how to attribute errors to the ability of a model to model the acoustics, versus its ability to leverage higher-order context (lexicon, morphology, syntax, semantics). We validate a novel approach which models error rates as a function of relative textual predictability, and yields a single number, $k$, which measures the effect of textual predictability on the recognizer. We use this method to demonstrate that a Wav2Vec 2.0-based model makes greater stronger use of textual context than a hybrid ASR model, in spite of not using an explicit language model, and also use it to shed light on recent results demonstrating poor performance of standard ASR systems on African-American English. We demonstrate that these mostly represent failures of acoustic--phonetic modelling. We show how this approach can be used straightforwardly in diagnosing and improving ASR.
Sean Robertson, Gerald Penn, Ewan Dunbar
INTERSPEECH2
2022 Does BERT Rediscover a Classical NLP Pipeline?
abstract
Does BERT store surface knowledge in its bottom layers, syntactic knowledge in its middle layers, and semantic knowledge in its upper layers? In re-examining Jawahar et al. (2019) and Tenney et al.’s (2019a) probes into the structure of BERT, we have found that the pipeline-like separation that they asserted lacks conclusive empirical support. BERT’s structure is, however, linguistically founded, although perhaps in a way that is more nuanced than can be explained by layers alone. We introduce a novel probe, called GridLoc, through which we can also take into account token positions, training rounds, and random seeds. Using GridLoc, we are able to detect other, stronger regularities that suggest that pseudo-cognitive appeals to layer depth may not be the preferable mode of explanation for BERT’s inner workings.
Jingcheng Niu, Gerald Penn
COLING3
2021 Reanalyzing the Most Probable Sentence Problem: A Case Study in Explicating the Role of Entropy in Algorithmic Complexity
abstract
When working with problems in natural language processing, we can find ourselves in situations where the traditional measurements of descriptive complexity are ineffective at describing the behaviour of our algorithms. It is easy to see why — the models we use are often general frameworks into which difficult-to-define tasks can be embedded. These frameworks can have more power than we typically use, and so complexity measures such as worst-case running time can drastically overestimate the cost of running our algorithms. In particular, they can make an apparently tractable problem seem NP-complete. Using empirical studies to evaluate performance is a necessary but incomplete method of dealing with this mismatch, since these studies no longer act as a guarantee of good performance. In this paper we use statistical measures such as entropy to give an updated analysis of the complexity of the NP-complete Most Probable Sentence problem for pCFGs, which can then be applied to word sense disambiguation and inference tasks. We can bound both the running time and the error in a simple search algorithm, allowing for a much faster search than the NP-completeness of this problem would suggest.
Eric Corlett, Gerald Penn
EACL2
2021 The Chinese Remainder Theorem for Compact, Task-Precise, Efficient and Secure Word Embeddings
abstract
The growing availability of powerful mobile devices and other edge devices, together with increasing regulatory and security concerns about the exchange of personal information across networks of these devices has challenged the Computational Linguistics community to develop methods that are at once fast, space-efficient, accurate and amenable to secure encoding schemes such as homomorphic encryption.Inspired by recent work that restricts floating point precision to speed up neural network training in hardware-based SIMD, we have developed a method for compressing word vector embeddings into integers using the Chinese Reminder Theorem that speeds up addition by up to 48.27% and at the same time compresses GloVe word embedding libraries by up to 25.86%.We explore the practicality of this simple approach by investigating the trade-off between precision and performance in two NLP tasks: compositional semantic relatedness and opinion target sentiment classification.We find that in both tasks, lowering floating point number precision results in negligible changes to performance.
Patricia Thaine, Gerald Penn
EACL2
2020 Temporal Histories of Epidemic Events (THEE): A Case Study in Temporal Annotation for Public Health
abstract
We present a new temporal annotation standard, THEE-TimeML, and a corpus TheeBank enabling precise temporal information extraction (TIE) for event-based surveillance (EBS) systems in the public health domain. Current EBS must estimate the occurrence time of each event based on coarse document metadata such as document publication time. Because of the complicated language and narration style of news articles, estimated case outbreak times are often inaccurate or even erroneous. Thus, it is necessary to create annotation standards and corpora to facilitate the development of TIE systems in the public health domain to address this problem. We will discuss the adaptations that have proved necessary for this domain as we present THEE-TimeML and TheeBank. Finally, we document the corpus annotation process, and demonstrate the immediate benefit to public health applications brought by the annotations.
Jingcheng Niu, Victoria Ng, Gerald Penn, Erin E. Rees
LREC3
2020 FAB: The French Absolute Beginner Corpus for Pronunciation Training
abstract
We introduce the French Absolute Beginner (FAB) speech corpus. The corpus is intended for the development and study of Computer-Assisted Pronunciation Training (CAPT) tools for absolute beginner learners. Data were recorded during two experiments focusing on using a CAPT system in paired role-play tasks. The setting grants FAB three distinguishing features from other non-native corpora: the experimental setting is ecologically valid, closing the gap between training and deployment; it features a label set based on teacher feedback, allowing for context-sensitive CAPT; and data have been primarily collected from absolute beginners, a group often ignored. Participants did not read prompts, but instead recalled and modified dialogues that were modelled in videos. Unable to distinguish modelled words solely from viewing videos, speakers often uttered unintelligible or out-of-L2 words. The corpus is split into three partitions: one from an experiment with minimal feedback; another with explicit, word-level feedback; and a third with supplementary read-and-record data. A subset of words in the first partition has been labelled as more or less native, with inter-annotator agreement reported. In the explicit feedback partition, labels are derived from the experiment’s online feedback. The FAB corpus is scheduled to be made freely available by the end of 2020.
Sean Robertson, Cosmin Munteanu, Gerald Penn
LREC3
2019 Rationally Reappraising ATIS-based Dialogue Systems
abstract
The Air Travel Information Service (ATIS) corpus has been the most common benchmark for evaluating Spoken Language Understanding (SLU) tasks for more than three decades since it was released.Recent state-of-the-art neural models have obtained F1-scores near 98% on the task of slot filling.We developed a rule-based grammar for the ATIS domain that achieves a 95.82% F1-score on our evaluation set.In the process, we furthermore discovered numerous shortcomings in the ATIS corpus annotation, which we have fixed.This paper presents a detailed account of these shortcomings, our proposed repairs, our rulebased grammar and the neural slot-filling architectures associated with ATIS.We also rationally reappraise the motivations for choosing a neural architecture in view of this account.Fixing the annotation errors results in a relative error reduction of between 19.4 and 52% across all architectures.We nevertheless argue that neural models must play a different role in ATIS dialogues because of the latter's lack of variety.
Jingcheng Niu, Gerald Penn
ACL (1)2
2019 Extracting Mel-Frequency and Bark-Frequency Cepstral Coefficients from Encrypted Signals
Patricia Thaine, Gerald Penn
INTERSPEECH2
2018 Designing Pronunciation Learning Tools: The Case for Interactivity against Over-Engineering
abstract
Paired role-play is a common collaborative activity in language learning classrooms, adding meaning and cultural context to the learning process. This is complemented by teachers' immediate and explicit feedback. Interactive tools that provide explicit feedback during collaborative learning are scarce, however. More commonly, supporting dialogue practice takes the form of computer-aided single-student read-and-record activities. This limitation is partly due to the complexity of processing language learners' speech in unconstrained tasks. In this paper, we assess the value of pronunciation error detection algorithms within a realistic, software-aided, paired role-playing task with beginning learners of French. We found that students' pronunciations improve regardless of the type of error detector employed -- even for those using simple heuristics. We suggest that speech technologies for language learning have been too focused on engineering goals. Instead, new interactive designs supporting collaboration may be used to overcome engineering limitations and properly support students' engagement.
Sean Robertson, Cosmin Munteanu, Gerald Penn
CHI3
2018 Can Deep Learning Compensate for a Shallow Evaluation?
abstract
The last ten years have witnessed an enormous increase in the application of "deep learning" methods to both spoken and textual natural language processing. Have they helped? With respect to some well-defined tasks such as language modelling and acoustic modelling, the answer is most certainly affirmative, but those are mere components of the real applications that are driving the increasing interest in our field. In many of these real applications, the answer is surprisingly that we cannot be certain because of the shambolic evaluation standards that have been commonplace --- long before the deep learning renaissance --- in the communities that specialized in advancing them.
Gerald Penn
DocEng1
2018 MOS Naturalness and the Quest for Human-Like Speech
abstract
This paper reconsiders the use of MOS naturalness as an instrument for measuring the quality (vs. intelligibility) of speech. We reconsider an earlier proposed alternative, the paired comparison or “AB” test, and present new empirical evidence that this is indeed a better method for evaluating TTS quality. Using this, we evaluate three older TTS systems along with a recent deep-learning approach against native North-American and Indian speech and show that, in fact, TTS had already crossed the threshold of human-like speech synthesis some time ago. This suggests that a systematic reappraisal of the concept of abstract “naturalness” of speech is in order.
Sajad Shirali-Shahreza, Gerald Penn
SLT2
2017 Speech and Hands-free interaction: myths, challenges, and opportunities
abstract
HCI research has for long been dedicated to better and more naturally facilitating information transfer between humans and machines. Unfortunately, humans' most natural form of communication, speech, is also one of the most difficult modalities to be understood by machines - despite, and perhaps, because it is the highest-bandwidth communication channel we possess. While significant research efforts, from engineering, to linguistic, and to cognitive sciences, have been spent on improving machines' ability to understand speech, the MobileHCI community (and the HCI field at large) has been relatively timid in embracing this modality as a central focus of research. This can be attributed in part to the unexpected variations in error rates when processing speech, in contrast with often-unfounded claims of success from industry, but also to the intrinsic difficulty of designing and especially evaluating speech and natural language interfaces. As such, the development of interactive speech-based systems is mostly driven by engineering efforts to improve such systems with respect to largely arbitrary performance metrics. Such developments have often been void of any user-centered design principles or consideration for usability or usefulness.
Cosmin Munteanu, Gerald Penn
MobileHCI2
2016 Evaluating Sentiment Analysis in the Context of Securities Trading
abstract
There are numerous studies suggesting that published news stories have an important effect on the direction of the stock market, its volatility, the volume of trades, and the value of individual stocks mentioned in the news.There is even some published research suggesting that automated sentiment analysis of news documents, quarterly reports, blogs and/or twitter data can be productively used as part of a trading strategy.This paper presents just such a family of trading strategies, and then uses this application to re-examine some of the tacit assumptions behind how sentiment analyzers are generally evaluated, in spite of the contexts of their application.This discrepancy comes at a cost.
Siavash Kazemian, Shunan Zhao, Gerald Penn
ACL (1)3
2016 Pronunciation Error Detection for New Language Learners
Sean Robertson, Cosmin Munteanu, Gerald Penn
INTERSPEECH3
2015 Deep bi-directional recurrent networks over spectral windows
abstract
Long short-term memory (LSTM) acoustic models have recently achieved state-of-the-art results on speech recognition tasks. As a type of recurrent neural network, LSTMs potentially have the ability to model long-span phenomena relating the spectral input to linguistic units. However, it has not been clear whether their observed performance is actually due to this capability, or instead if it is due to a better modeling of short term dynamics through the recurrence. In this paper. we answer this question by applying a windowed (truncated) LSTM to conversational speech transcription, and find that a limited context is adequate, and that it is not necessaary to scan the entire utterance. The sliding window approach allows not only incremental (online) recognition with a bidirectional model, but also frame-wise randomization (as opposed to utterance randomization), which results in faster convergence. On the SWBD/Fisher corpus, applying bidirectional LSTM RNNs to spectral windows of about 0.5s improves WER on the Hub5'00 benchmark set by 16% relative compared to our best sequence-trained DNN. On an extended 3850h training set that that also includes lectures, the relative gain becomes 28% (Hub5'00 WER 9.2%). In-house conversational data improves by 12 to 17% relative.
Abdel-rahman Mohamed, Frank Seide, Dong Yu 0001, Jasha Droppo, Andreas Stolcke, Geoffrey Zweig, Gerald Penn
ASRU7
2015 Speech-based Interaction: Myths, Challenges, and Opportunities
abstract
HCI research has for long been dedicated to better and more naturally facilitating information transfer between humans and machines. Unfortunately, humans' most natural form of communication, speech, is also one of the most difficult modalities to be understood by machines -- despite, and perhaps, because it is the highest-bandwidth communication channel we possess. While significant research efforts, from engineering, to linguistic, and to cognitive sciences, have been spent on improving machines' ability to understand speech, the HCI community has been relatively timid in embracing this modality as a central focus of research. This can be attributed in part to the relatively discouraging levels of accuracy in understanding speech, in contrast with often-unfounded claims of success from industry, but also to the intrinsic difficulty of designing and especially evaluating speech and natural language interfaces.
Cosmin Munteanu, Gerald Penn
IUI2
2014 Unsupervised Sentence Enhancement for Automatic Summarization
abstract
We present sentence enhancement as a novel technique for text-to-text genera-tion in abstractive summarization. Com-pared to extraction or previous approaches to sentence fusion, sentence enhancement increases the range of possible summary sentences by allowing the combination of dependency subtrees from any sentence from the source text. Our experiments in-dicate that our approach yields summary sentences that are competitive with a sen-tence fusion baseline in terms of con-tent quality, but better in terms of gram-maticality, and that the benefit of sen-tence enhancement relies crucially on an event coreference resolution algorithm us-ing distributional semantics. We also consider how text-to-text generation ap-proaches to summarization can be ex-tended beyond the source text by exam-ining how human summary writers incor-porate source-text-external elements into their summary sentences. 1
Jackie Chi Kit Cheung, Gerald Penn
EMNLP2
2014 Speech-based interaction: myths, challenges, and opportunities
abstract
Human-Computer Interaction (HCI) research has for long been dedicated to better and more naturally facilitating information transfer between humans and machines. Unfortunately, humans' most natural form of communication, speech, is also one of the most difficult modalities to be understood by machines. This is largely due to speech being the highest-bandwidth communication channel we possess. As such, significant research efforts, from engineering, to linguistic, and to cognitive sciences, have been spent during the past several decades on improving machines' ability to understand speech. Yet, the MobileHCI community (and HCI in general) has been relatively timid in embracing this modality as a central focus of research. This can be attributed in part to the relatively discouraging levels of accuracy in understanding speech, in contrast with often-unfounded claims of success from industry, but also to the intrinsic difficulty of designing and especially evaluating speech and natural language interfaces.
Cosmin Munteanu, Gerald Penn
Mobile HCI2
2014 Convolutional Neural Networks for Speech Recognition
abstract
Recently, the hybrid deep neural network (DNN)-hidden Markov model (HMM) has been shown to significantly improve speech recognition performance over the conventional Gaussian mixture model (GMM)-HMM. The performance improvement is partially attributed to the ability of the DNN to model complex correlations in speech features. In this paper, we show that further error rate reduction can be obtained by using convolutional neural networks (CNNs). We first present a concise description of the basic CNN and explain how it can be used for speech recognition. We further propose a limited-weight-sharing scheme that can better model speech features. The special structure such as local connectivity, weight sharing, and pooling in CNNs exhibits some degree of invariance to small shifts of speech features along the frequency axis, which is important to deal with speaker and environment variations. Experimental results show that CNNs reduce the error rate by 6%-10% compared with DNNs on the TIMIT phone recognition and the voice search large vocabulary speech recognition tasks.
Ossama Abdel-Hamid, Abdel-rahman Mohamed, Hui Jiang 0001, Li Deng 0001, Gerald Penn, Dong Yu 0001
IEEE ACM Trans. Audio Speech Lang. Process.5
2013 Probabilistic Domain Modelling With Contextualized Distributional Semantic Vectors
Jackie Chi Kit Cheung, Gerald Penn
ACL (1)2
2013 Towards Robust Abstractive Multi-Document Summarization: A Caseframe Analysis of Centrality and Domain
Jackie Chi Kit Cheung, Gerald Penn
ACL (1)2
2013 SeeSay and HearSay CAPTCHA for mobile interaction
abstract
Speech certainly has advantages as an input modality for smartphone applications, especially in scenarios where using touch or keyboard entry is difficult, on increasingly miniaturized devices where useable keyboards are difficult to accommodate, or in scenarios where only small amounts of text need to be input, such as when entering SMS texts or responding to a CAPTCHA challenge. In this paper, we propose two new alternative ways to design CAPTCHAs in which the user says the answer instead of typing it with (a) output stimuli provided visually (SeeSay) or (b) auditorily (HearSay). Our user study results show that SeeSay CAPTCHA requires less time to be solved and users prefer it over current text-based CAPTCHA methods.
Sajad Shirali-Shahreza, Gerald Penn, Ravin Balakrishnan, Yashar Ganjali
CHI2
2013 Automatic human utility evaluation of ASR systems: does WER really predict performance?
abstract
International audience
Benoît Favre, Kyla Cheung, Siavash Kazemian, Adam Lee, Yang Liu 0004, Cosmin Munteanu, Ani Nenkova, Dennis Ochei, Gerald Penn, Stephen Tratz, Clare R. Voss, Frauke Zeller
INTERSPEECH9
2013 A Graph-Partitioning Framework for Aligning Hierarchical Topic Structures to Presentations
abstract
This paper studies the problem of imposing an existing hierarchical semantic structure onto a corresponding spoken document in which the structures are embedded, with the goal of indexing such documents for easier access. We propose a graph-partitioning framework to solve a semantic tree-to-string alignment problem through optimizing a normalized-cut criterion. We present models with different modeling capabilities and time complexities in this framework and provide experimental evidence of their performance. We relate graph partitioning to conventional dynamic time warping (DTW) as it applies to this problem, and show that the proposed framework can naturally include topic segmentation to accommodate cohesion constraints.
Xiaodan Zhu 0001, Colin Cherry, Gerald Penn
IEEE Trans. Speech Audio Process.3
2012 Flexible Structural Analysis of Near-Meet-Semilattices for Typed Unification-Based Grammar Design
Rouzbeh Farahmand, Gerald Penn
COLING2
2012 Evaluating Distributional Models of Semantics for Syntactically Invariant Inference
Jackie Chi Kit Cheung, Gerald Penn
EACL2
2012 Unsupervised Detection of Downward-Entailing Operators By Maximizing Classification Certainty
Jackie Chi Kit Cheung, Gerald Penn
EACL2
2012 Applying Convolutional Neural Networks concepts to hybrid NN-HMM model for speech recognition
abstract
Convolutional Neural Networks (CNN) have showed success in achieving translation invariance for many image processing tasks. The success is largely attributed to the use of local filtering and max-pooling in the CNN architecture. In this paper, we propose to apply CNN to speech recognition within the framework of hybrid NN-HMM model. We propose to use local filtering and max-pooling in frequency domain to normalize speaker variance to achieve higher multi-speaker speech recognition performance. In our method, a pair of local filtering layer and max-pooling layer is added at the lowest end of neural network (NN) to normalize spectral variations of speech signals. In our experiments, the proposed CNN architecture is evaluated in a speaker independent speech recognition task using the standard TIMIT data sets. Experimental results show that the proposed CNN method can achieve over 10% relative error reduction in the core TIMIT test sets when comparing with a regular NN using the same number of hidden layers and weights. Our results also show that the best result of the proposed CNN model is better than previously published results on the same TIMIT test sets that use a pre-trained deep NN model.
Ossama Abdel-Hamid, Abdel-rahman Mohamed, Hui Jiang 0001, Gerald Penn
ICASSP4
2012 Understanding how Deep Belief Networks perform acoustic modelling
abstract
Deep Belief Networks (DBNs) are a very competitive alternative to Gaussian mixture models for relating states of a hidden Markov model to frames of coefficients derived from the acoustic input. They are competitive for three reasons: DBNs can be fine-tuned as neural networks; DBNs have many non-linear hidden layers; and DBNs are generatively pre-trained. This paper illustrates how each of these three aspects contributes to the DBN's good recognition performance using both phone recognition performance on the TIMIT corpus and a dimensionally reduced visualization of the relationships between the feature vectors learned by the DBNs that preserves the similarity structure of the feature vectors at multiple scales. The same two methods are also used to investigate the most suitable type of input representation for a DBN.
Abdel-rahman Mohamed, Geoffrey E. Hinton, Gerald Penn
ICASSP3
2012 Ecological validity and the evaluation of speech summarization quality
abstract
There is little evidence of widespread adoption of speech summarization systems. This may be due in part to the fact that the natural language heuristics used to generate summaries are often optimized with respect to a class of evaluation measures that, while computationally and experimentally inexpensive, rely on subjectively selected gold standards against which automatically generated summaries are scored. This evaluation protocol does not take into account the usefulness of a summary in assisting the listener in achieving his or her goal. In this paper we study how current measures and methods for evaluating summarization systems compare to human-centric evaluation criteria. For this, we have designed and conducted an ecologically valid evaluation that determines the value of a summary when embedded in a task, rather than how closely a summary resembles a gold standard. The results of our evaluation demonstrate that in the domain of lecture summarization, the well-known baseline of maximal marginal relevance [1] is statistically significantly worse than human-generated extractive summaries, and even worse than having no summary at all in a simple quiz-taking task. Priming seems to have no statistically significant effect on the usefulness of the human summaries. This is interesting because priming had been proposed as a technique for increasing kappa scores and/or maintaining goal orientation among summary authors. In addition, our results suggest that ROUGE scores, regardless of whether they are derived from numerically-ranked reference data or ecologically valid human-extracted summaries, may not always be reliable as inexpensive proxies for task-embedded evaluations. In fact, under some conditions, relying exclusively on ROUGE may lead to scoring human-generated summaries very favourably even when a task-embedded score calls their usefulness into question relative to using no summaries at all.
Anthony McCallum, Gerald Penn, Cosmin Munteanu, Xiaodan Zhu 0001
SLT2
2012 Realistic answer verification: An analysis of user errors in a sentence-repetition task
abstract
Speech authentication protocols should have a challenge/response feature to be protected against replay attacks. As a result, they need to verify whether the user responded to an interactive prompt. However, it is usually assumed that the user will provide their answer perfectly. In this paper, we report on an ecologically valid user study that we conducted to test this assumption. Our results show that 40% of user answers are imperfect, even in a task as simple as sentence repetition. Error analysis reveals that 60% of the imperfect answers contain small errors that should be deemed acceptable, which increases the total acceptance rate of this task to 84%. We also tested a forced alignment algorithm as a means of verifying answers automatically.
Sajad Shirali-Shahreza, Gerald Penn
SLT2
2011 Indexing Spoken Documents with Hierarchical Semantic Structures: Semantic Tree-to-string Alignment Models
Xiaodan Zhu 0001, Colin Cherry, Gerald Penn
IJCNLP3
2010 Entity-Based Local Coherence Modelling Using Topological Fields
Jackie Chi Kit Cheung, Gerald Penn
ACL2
2010 An Exact A* Method for Deciphering Letter-Substitution Ciphers
Eric Corlett, Gerald Penn
ACL2
2010 Accurate Context-Free Parsing with Combinatory Categorial Grammar
Timothy A. D. Fowler, Gerald Penn
ACL2
2010 A Generalized-Zero-Preserving Method for Compact Encoding of Concept Lattices
Matthew Skala, Victoria Krakovna, János Kramár, Gerald Penn
ACL4
2010 Utilizing Extra-Sentential Context for Parsing
Jackie Chi Kit Cheung, Gerald Penn
EMNLP2
2009 Topological Field Parsing of German
Jackie Chi Kit Cheung, Gerald Penn
ACL/IJCNLP2
2009 Improving Automatic Speech Recognition for Lectures through Transformation-based Rules Learned from Minimal Data
Cosmin Munteanu, Gerald Penn, Xiaodan Zhu 0001
ACL/IJCNLP2
2009 Summarizing multiple spoken documents: finding evidence from untranscribed audio
Xiaodan Zhu 0001, Gerald Penn, Frank Rudzicz
ACL/IJCNLP2
2009 DocuBurst: Visualizing Document Content using Language Structure
abstract
Abstract Textual data is at the forefront of information management problems today. One response has been the development of visualizations of text data. These visualizations, commonly based on simple attributes such as relative word frequency, have become increasingly popular tools. We extend this direction, presenting the first visualization of document content which combines word frequency with the human‐created structure in lexical databases to create a visualization that also reflects semantic content. DocuBurst is a radial, space‐filling layout of hyponymy (the IS‐A relation), overlaid with occurrence counts of words in a document of interest to provide visual summaries at varying levels of granularity. Interactive document analysis is supported with geometric and semantic zoom, selectable focus on individual words, and linked access to source text.
Christopher Collins 0001, Sheelagh Carpendale, Gerald Penn
Comput. Graph. Forum3
2009 Bubble Sets: Revealing Set Relations with Isocontours over Existing Visualizations
abstract
While many data sets contain multiple relationships, depicting more than one data relationship within a single visualization is challenging. We introduce Bubble Sets as a visualization technique for data that has both a primary data relation with a semantically significant spatial organization and a significant set membership relation in which members of the same set are not necessarily adjacent in the primary layout. In order to maintain the spatial rights of the primary data relation, we avoid layout adjustment techniques that improve set cluster continuity and density. Instead, we use a continuous, possibly concave, isocontour to delineate set membership, without disrupting the primary layout. Optimizations minimize cluster overlap and provide for calculation of the isocontours at interactive speeds. Case studies show how this technique can be used to indicate multiple sets on a variety of common visualizations.
Christopher Collins 0001, Gerald Penn, Sheelagh Carpendale
IEEE Trans. Vis. Comput. Graph.2
2008 A Critical Reassessment of Evaluation Baselines for Speech Summarization
Gerald Penn, Xiaodan Zhu 0001
ACL1
2008 Collaborative editing for improved usefulness and usability of transcript-enhanced webcasts
abstract
One challenge in facilitating skimming or browsing through archives of on-line recordings of webcast lectures is the lack of text transcripts of the recorded lecture. Ideally, transcripts would be obtainable through Automatic Speech Recognition (ASR). However, current ASR systems can only deliver, in realistic lecture conditions, a Word Error Rate of around 45% -- above the accepted threshold of 25%. In this paper, we present the iterative design of a webcast extension that engages users to collaborate in a wiki-like manner on editing the ASR-produced imperfect transcripts, and show that this is a feasible solution for improving the quality of lecture transcripts. We also present the findings of a field study carried out in a real lecture environment investigating how students use and edit the transcripts.
Cosmin Munteanu, Ronald Baecker, Gerald Penn
CHI3
2008 Using latent Dirichlet allocation to incorporate domain knowledge for topic transition detection
abstract
This paper studies automatic detection of topic transitions for recorded presentations. This can be achieved by matching slide content with presentation transcripts directly with some similarity metrics. Such literal matching, however, misses domain-specific knowledge and is sensitive to speech recognition errors. In this paper, we incorporate relevant written materials, e.g., textbooks for lectures, which convey semantic relationships, in particular domain-specific relationships, between words. To this end, we train latent Dirichlet allocation (LDA) models on these materials and measure the similarity between slides and transcripts in the acquired hidden-topic space. This similarity is then combined with literal matchings. Experiments show that the proposed approach reduces the errors in slide transition detection by 17-41 % on manual transcripts and 27-37% on automatic transcripts. Index Terms: slides transition detection, boundary detection. 1.
Xiaodan Zhu 0001, Xuming He 0001, Cosmin Munteanu, Gerald Penn
INTERSPEECH4
2008 Identifying salient utterances of online spoken documents using descriptive hypertext
abstract
The Internet has become an important supply channel of spoken documents. Efficient ways of navigating their content are highly desirable. This paper aims to identify the most salient utterances from online spoken documents using relevant hypertext that encapsulates key information. Experimental results show that hypertext features are helpful when properly utilized and if the bit rates used to compress the spoken documents are reasonable.
Xiaodan Zhu 0001, Siavash Kazemian, Gerald Penn
SLT3
2007 Web-based language modelling for automatic lecture transcription
abstract
Universities have long relied on written text to share knowledge. As more lectures are made available on-line, these must be accompanied by textual transcripts in order to provide the same access to information as textbooks. While Automatic Speech Recognition (ASR) is a cost-effective method to deliver transcriptions, its accuracy for lectures is not yet satisfactory. One approach for improving lecture ASR is to build smaller, topic-dependent Language Models (LMs) and combine them (through LM interpolation or hypothesis space combination) with general-purpose, large-vocabulary LMs. In this paper, we propose a simple solution for lecture ASR with similar or better Word Error Rate reductions (as well as topic-specific keyword identification accuracies) than combination-based approaches. Our method eliminates the need for two types of LMs by exploiting the lecture slides to collect a web corpus appropriate for modelling both the conversational and the topic-specific styles of lectures. Index Terms: speech recognition, language modelling, corpus building, topic dependent, lecture transcription.
Cosmin Munteanu, Gerald Penn, Ronald Baecker
INTERSPEECH2
2007 Visualization of Uncertainty in Lattices to Support Decision-Making
abstract
Lattice graphs are used as underlying data structures in many statistical processing systems, including natural language processing. Lattices compactly represent multiple possible outputs and are usually hidden from users. We present a novel visualization intended to reveal the uncertainty and variability inherent in statistically-derived lattice structures. Applications such as machine translation and automated speech recognition typically present users with a best-guess about the appropriate output, with apparent complete confidence. Through case studies we show how our visualization uses a hybrid layout along with varying transparency, colour, and size to reveal the lattice structure, expose the inherent uncertainty in statistical processing, and help users make better-informed decisions about statistically-derived outputs.
Christopher Collins 0001, Sheelagh Carpendale, Gerald Penn
EuroVis3
2006 The effect of speech recognition accuracy rates on the usefulness and usability of webcast archives
abstract
The widespread availability of broadband connections has led to an increase in the use of Internet broadcasting (webcasting). Most webcasts are archived and accessed numerous times retrospectively. In the absence of transcripts of what was said, users have difficulty searching and scanning for specific topics. This research investigates user needs for transcription accuracy in webcast archives, and measures how the quality of transcripts affects user performance in a question-answering task, and how quality affects overall user experience. We tested 48 subjects in a within-subjects design under 4 conditions: perfect transcripts, transcripts with 25% Word Error Rate (WER), transcripts with 45% WER, and no transcript. Our data reveals that speech recognition accuracy linearly influences both user performance and experience, shows that transcripts with 45% WER are unsatisfactory, and suggests that transcripts having a WER of 25% or less would be useful and usable in webcast archives.
Cosmin Munteanu, Ronald Baecker, Gerald Penn, Elaine Toms, David James
CHI3
2006 Utterance-Level Extractive Summarization of Open-Domain Spontaneous Conversations with Rich Features
abstract
To identify important utterances from open-domain spontaneous conversations, previous work has focused on using textual features that are extracted from transcripts, e.g., word frequencies and noun senses. In this paper, we summarize spontaneous conversations with features of a wide variety that have not been explored before. Experiments show that the use of speech-related features improves summarization performance. In addition, the effectiveness of individual features is examined and compared
Xiaodan Zhu 0001, Gerald Penn
ICME2
2006 Automatic speech recognition for webcasts: how good is good enough and what to do when it isn't
abstract
The increased availability of broadband connections has recently led to an increase in the use of Internet broadcasting (webcasting). Most webcasts are archived and accessed numerous times retrospectively. One challenge to skimming and browsing through such archives is the lack of text transcripts of the webcast's audio channel. This paper describes a procedure for prototyping an Automatic Speech Recognition (ASR) system that generates realistic transcripts of any desired Word Error Rate (WER), thus overcoming the drawbacks of both prototype-based and Wizard of Oz simulations. We used such a system in a user study showing that transcripts with WERs less than 25% are acceptable for use in webcast archives. As current ASR systems can only deliver, in realistic conditions, Word Error Rates (WERs) of around 45%, we also describe a solution for reducing the WER of such transcripts by engaging users to collaborate in a wiki fashion on editing the imperfect transcripts obtained through ASR.
Cosmin Munteanu, Gerald Penn, Ronald Baecker, Yuecheng Zhang
ICMI2
2006 Measuring the acceptable word error rate of machine-generated webcast transcripts
abstract
The increased availability of broadband connections has recently led to an increase in the use of Internet broadcasting (webcasting). Most webcasts are archived and accessed numerous times retrospectively. One of the hurdles users face when browsing and skimming through archives is the lack of text transcripts of the audio channel of the webcast archive. In this paper, we proposed a procedure for prototyping an Automatic Speech Recognition (ASR) system that generates realistic transcripts of any desired Word Error Rate (WER), thus overcoming the drawbacks of both prototypebased and Wizard of Oz simulations. We used such a system in a study where human subjects perform question-answering tasks using archives of webcast lectures, and showed that their performance and perception of transcript quality is linearly affected by WER, and that transcripts of WER equal or less than 25 % would be acceptable for use in webcast archives.
Cosmin Munteanu, Gerald Penn, Ronald Baecker, Elaine Toms, David James
INTERSPEECH2
2006 Summarization of spontaneous conversations
abstract
Spontaneous conversations are an integral element in many CSCW environments. Although speech is often regarded as the most natural and effective way of communication between human beings, speech data are not efficient for quick review. One solution to help people access speech data efficiently in CSCW environments is to conduct speech summarization. Up till now, most speech summarization research has focused on broadcast news; nevertheless summarizing spontaneous conversations is more valuable for CSCW. The task is also more challenging, for example, spontaneous conversations often contain more speech disfluencies, which need to be coped with properly; they are also more vulnerable to speech recognition errors. This demonstration is built to show the prototype of our summarization system. Compared with previous work, our summarizer addresses the problem further in several important respects. First, the system summarizes spontaneous conversations with a wide variety of information/features that have not been explored before, which improve summarization performance according to our experiments. Second, our summarizer handles speech disfluencies, which in all previous work was either not explicitly handled or removed as noise.
Xiaodan Zhu 0001, Gerald Penn
INTERSPEECH2
2006 Quantitative Methods for Classifying Writing Systems
Gerald Penn, Travis Choma
HLT-NAACL1
2006 Comparing the roles of textual, acoustic and spoken-language features on spontaneous-conversation summarization
Xiaodan Zhu 0001, Gerald Penn
HLT-NAACL2
2006 Efficient transitive closure of sparse matrices over closed semirings
Gerald Penn
Theor. Comput. Sci.1
2004 Head-Driven Parsing for Word Lattices
abstract
We present the first application of the head-driven statistical parsing model of Collins (1999) as a simultaneous language model and parser for large-vocabulary speech recognition. The model is adapted to an online left to right chart-parser for word lattices, integrating acoustic, n-gram, and parser probabilities. The parser uses structural and lexical dependencies not considered by n-gram models, conditioning recognition on more linguistically-grounded relationships. Experiments on the Wall Street Journal treebank and lattice corpora show word error rates competitive with the standard n-gram language model while extracting additional structural information useful for speech understanding.
Christopher Collins 0001, Bob Carpenter, Gerald Penn
ACL3
2004 Optimizing Typed Feature Structure Grammar Parsing through Non-Statistical Indexing
abstract
This paper introduces an indexing method based on static analysis of grammar rules and type signatures for typed feature structure grammars (TFSGs). The static analysis tries to predict at compile-time which feature paths will cause unification failure during parsing at run-time. To support the static analysis, we introduce a new classification of the instances of variables used in TFSGs, based on what type of structure sharing they create. The indexing actions that can be performed during parsing are also enumerated. Non-statistical indexing has the advantage of not requiring training, and, as the evaluation using large-scale HPSGs demonstrates, the improvements are comparable with those of statistical optimizations. Such statistical optimizations rely on data collected during training, and their performance does not always compensate for the training costs.
Cosmin Munteanu, Gerald Penn
ACL2
2004 Balancing Clarity and Efficiency in Typed Feature Logic Through Delaying
abstract
The purpose of this paper is to re-examine the balance between clarity and efficiency in HPSG design, with particular reference to the design decisions made in the English Resource Grammar (LinGO, 1999, ERG). It is argued that a simple generalization of the conventional delay statements used in logic programming is sufficient to restore much of the functionality and concomitant benefit that the ERG elected to forego, with an acceptable although still perceptible computational cost.
Gerald Penn
ACL1
2003 A Tabulation-Based Parsing Method that Reduces Copying
abstract
This paper presents a new bottom-up chart parsing algorithm for Prolog along with a compilation procedure that reduces the amount of copying at run-time to a constant number (2) per edge. It has applications to unification-based grammars with very large partially ordered categories, in which copying is expensive, and can facilitate the use of more sophisticated indexing strategies for retrieving such categories that may otherwise be overwhelmed by the cost of such copying. It also provides a new perspective on "quick-checking" and related heuristics, which seems to confirm that forcing an early failure (as opposed to seeking an early guarantee of success) is in fact the best approach to use. A preliminary empirical evaluation of its performance is also provided.
Gerald Penn, Cosmin Munteanu
ACL1
2003 AVM Description Compilation using Types as Modes
Gerald Penn
EACL1
2003 Topological Parsing
Gerald Penn, Mohammad Haji-Abdolhosseini
EACL1
2003 Implementing Typed Feature Structure Grammars By Ann Copestake
Gerald Penn
Comput. Linguistics1
2002 Generalized Encoding of Description Spaces and its Application to Typed Feature Structures
abstract
This paper presents a new formalization of a unification-or join-preserving encoding of partially ordered sets that more essentially captures what it means for an encoding to preserve joins, generalizing the standard definition in AI research.It then shows that every statically typable ontology in the logic of typed feature structures can be encoded in a data structure of fixed size without the need for resizing or additional union-find operations.This is important for any grammar implementation or development system based on typed feature structures, as it significantly reduces the overhead of memory management and reference-pointer-chasing during unification.
Gerald Penn
ACL1
2001 Tractability and Structural Closures in Attribute Logic Type Signatures
abstract
This paper considers three assumptions conventionally made about signatures in typed feature logic that are in potential disagreement with current practice among grammar developers and linguists working within feature-based frameworks such as HPSG: meet-semi-latticehood, unique feature introduction, and the absence of subtype covering. It also discusses the conditions under which each of these can be tractably restored in realistic grammar signatures where they do not already exist.
Gerald Penn
ACL1
2001 Flexible Web Document Analysis for Delivery to Narrow-Bandwidth Devices
abstract
We propose a set of baseline heuristics for identifying genuinely tabular information and news links in HTML documents. A prototype implementation of these heuristics is described for delivering content from news providers' home pages to a narrow-bandwidth device such as a portable digital assistant or cellular phone display. Its evaluation on 75 Web sites is provided, along with a discussion of topics for future research.
Gerald Penn, Jianying Hu, Hengbin Luo, Ryan T. McDonald
ICDAR1
1999 An Optimized Prolog Encoding of Typed Feature Structures
Gerald Penn
ICLP1
1999 ALE for speech: a translation prototype
abstract
This paper discusses the use of a secondary likelihood classifier scheme for improving speaker recognition performance. The system models the output likelihoods of a typical Gaussian Mixture Model system across multiple speakers. The Output Probability Distributions (OPD) of the primary classifiers contain information on inter-speaker relationships, and are modelled by secondary classifiers to improve recognition accuracies. A comparison of the OPD system with the traditional likelihood ratio and maximum likelihood scoring schemes for verification and identification is performed. Fusion of traditional measures with OPDs is shown to enhance overall recognition performance.
Gerald Penn, Bob Carpenter
EUROSPEECH1