VLDB 2026 Research / reviewers in the wild / expert
Kenneth Church 0001
dblp:c/KennethWardChurch · also Ken Ward Church, Kenneth Ward Church
· DBLP profile ↗
126ranked-venue papers
57as first author
34since 2021 · last 2026
0000-0001-8378-6069ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 104 · 51 first-author · 29 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 6 first-author · 8 since 2021Databases, data management, data science and information retrieval · 12 · 3 first-author · 5 since 2021Theory of computation · 3Systems, architecture and hardware · 1Computer networks · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Scholar API: Search and Recommendations for Academic SearchabstractThere has been considerable work on Academic Search. Academic Search is widely used; in addition to recommending papers to read, authors need to find papers they should cite, and program committees and funding agencies need to assign submissions to reviewers that are well-informed and sympathetic to the topic area. High-quality recommendations impact reviews, publication quality, and move a field into new directions. Kenneth Church 0001, Omar Alonso |
WSDM | 1 |
| 2025 | Is Peer-Reviewing Worth the Effort?abstractHow effective is peer-reviewing in identifying important papers? We treat this question as a forecasting task. Can we predict which papers will be highly cited in the future based on venue and “early returns” (citations soon after publication)? We show early returns are more predictive than venue. Finally, we end with a constructive suggestion to simplify reviewing. Kenneth Church 0001, Raman Chandrasekar, John E. Ortega, Ibrahim Said Ahmad |
COLING | 1 |
| 2025 | Rapid Prototyping for AI-Based Applications: A Hands-on Tutorial for Connecting the Dots
Omar Alonso, Kenneth Church 0001 |
ECIR (5) | 2 |
| 2024 | No Culture Left Behind: ArtELingo-28, a Benchmark of WikiArt with Captions in 28 LanguagesabstractYoussef Mohamed, Runjia Li, Ibrahim Said Ahmad, Kilichbek Haydarov, Philip Torr, Kenneth Church, Mohamed Elhoseiny. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Youssef Mohamed, Runjia Li, Ibrahim Said Ahmad, Kilichbek Haydarov, Philip Torr 0001, Kenneth Church 0001, Mohamed Elhoseiny 0001 |
EMNLP | 6 |
| 2024 | Some Useful Things to Know When Combining IR and NLP: The Easy, the Hard and the UglyabstractDeep nets such as GPT are at the core of the current advances in many systems and applications. Things are moving fast; techniques become obsolete quickly (within weeks). How can we take advantage of new discoveries and incorporate them into our existing work? Are new developments radical improvements, or incremental repetitions of established concepts, or combinations of both? Omar Alonso, Kenneth Church 0001 |
WSDM | 2 |
| 2024 | Emerging trends: When can users trust GPT, and when should they intervene?abstractAbstract Usage of large language models and chat bots will almost surely continue to grow, since they are so easy to use, and so (incredibly) credible. I would be more comfortable with this reality if we encouraged more evaluations with humans-in-the-loop to come up with a better characterization of when the machine can be trusted and when humans should intervene. This article will describe a homework assignment, where I asked my students to use tools such as chat bots and web search to write a number of essays. Even after considerable discussion in class on hallucinations, many of the essays were full of misinformation that should have been fact-checked. Apparently, it is easier to believe ChatGPT than to be skeptical. Fact-checking and web search are too much trouble. Kenneth Church 0001 |
Nat. Lang. Eng. | 1 |
| 2024 | Emerging trends: evaluating general purpose foundation modelsabstractAbstract We suggest that foundation models are general purpose solutions similar to general purpose programmable microprocessors, where fine-tuning and prompt-engineering are analogous to coding for microprocessors. Evaluating general purpose solutions is not like hypothesis testing. We want to know how well the machine will perform on an unknown program with unknown inputs for unknown users with unknown budgets and unknown utility functions. This paper is based on an invited talk by John Mashey, “Lessons from SPEC,” at an ACL-2021 workshop on benchmarking. Mashey started by describing Standard Performance Evaluation Corporation (SPEC), a benchmark that has had more impact than benchmarks in our field because SPEC addresses an import commercial question: which CPU should I buy? In addition, SPEC can be interpreted to show that CPUs are 50,000 faster than they were 40 years ago. It is remarkable that we can make such statements without specifying the program, users, task, dataset, etc. It would be desirable to make quantitative statements about improvements of general purpose foundation models over years/decades without specifying tasks, datasets, use cases, etc. Kenneth Church 0001, Omar Alonso |
Nat. Lang. Eng. | 1 |
| 2024 | Emerging trends: a gentle introduction to RAGabstractAbstract Retrieval-augmented generation (RAG) adds a simple but powerful feature to chatbots, the ability to upload files just-in-time. Chatbots are trained on large quantities of public data. The ability to upload files just-in-time makes it possible to reduce hallucinations by filling in gaps in the knowledge base that go beyond the public training data such as private data and recent events. For example, in a customer service scenario, with RAG, we can upload your private bill and then the bot can discuss questions about your bill as opposed to generic FAQ questions about bills in general. This tutorial will show how to upload files and generate responses to prompts; see https://github.com/kwchurch/RAG for multiple solutions based on tools from OpenAI, LangChain, HuggingFace transformers and VecML. Kenneth Church 0001, Jiameng Sun, Richard Yue, Peter Vickers, Walid S. Saba, Raman Chandrasekar |
Nat. Lang. Eng. | 1 |
| 2023 | Some Useful Things to Know When Combining IR and NLP: the Easy, the Hard and the UglyabstractDeep nets such as GPT are at the core of the current advances in many systems and applications. Things are moving very fast, and it appears that techniques are out of date within weeks. How can we take advantage of new discoveries and incorporate them into our existing work? Are these radical new developments, repetitions of older concepts, or both? Omar Alonso, Kenneth Church 0001 |
CIKM | 2 |
| 2023 | An Example of (Too Much) Hyper-Parameter Tuning In Suicide Ideation DetectionabstractThis work starts with the TWISCO baseline, a benchmark of suicide-related content from Twitter. We find that hyper-parameter tuning can improve this baseline by 9%. We examined 576 combinations of hyper-parameters: learning rate, batch size, epochs and date range of training data. Reasonable settings of learning rate and batch size produce better results than poor settings. Date range is less conclusive. Balancing the date range of the training data to match the benchmark ought to improve performance, but the differences are relatively small. Optimal settings of learning rate and batch size are much better than poor settings, but optimal settings of date range are not that different from poor settings of date range. Finally, we end with concerns about reproducibility. Of the 576 experiments, 10% produced F1 performance above baseline. It is common practice in the literature to run many experiments and report the best, but doing so may be risky, especially given the sensitive nature of Suicide Ideation Detection. Annika Marie Schoene, John E. Ortega, Silvio Amir, Kenneth Church 0001 |
ICWSM | 4 |
| 2023 | Improved Contextualized Speech Representations for Tonal Analysis
Jiahong Yuan, Xingyu Cai, Kenneth Church 0001 |
INTERSPEECH | 3 |
| 2023 | Emerging trends: Risks 3.0 and proliferation of spyware to 50,000 cell phonesabstractAbstract Our last emerging trend article introduced Risks 1.0 (fairness and bias) and Risks 2.0 (addictive, dangerous, deadly, and insanely profitable). This article introduces Risks 3.0 (spyware and cyber weapons). Risks 3.0 are less profitable, but more destructive. We will summarize two recent books, Pegasus: How a Spy in Your Pocket Threatens the End of Privacy, Dignity, and Democracy and This is How They Tell Me the World Ends: The Cyberweapons Arms Race. The first book starts with a leak of 50,000 phone numbers, targeted by spyware named Pegasus. Pegasus uses a zero-click exploit to obtain root access to your phone, taking control of the microphone, camera, GPS, text messages, etc. The list of 50,000 numbers includes journalists, politicians, and academics, as well as their friends and family. Some of these people have been murdered. The second book describes the history of cyber weapons such as Stuxnet, which is described as crossing the Rubicon. In the short term, it sets back Iran’s nuclear program for less than the cost of conventional weapons, but it did not take long for Iran to build the fourth-biggest cyber army in the world. As spyware continues to proliferate, we envision a future dystopia where everyone spies on everyone. Nothing will be safe from hacking: not your identity, or your secrets, or your passwords, or your bank accounts. When the endpoints (phones) have been compromised, technologies such as end-to-end encryption and multi-factor authentication offer a false sense of security; encryption and authentication are as pointless as closing the proverbial barn door after the fact. To address Risks 3.0, journalists are using the tools of their trade to raise awareness in the court of public opinion. We should do what we can to support them. This paper is a small step in that direction. Kenneth Church 0001, Raman Chandrasekar |
Nat. Lang. Eng. | 1 |
| 2023 | Emerging trends: Unfair, biased, addictive, dangerous, deadly, and insanely profitableabstractAbstract There has been considerable work recently in the natural language community and elsewhere on Responsible AI. Much of this work focuses on fairness and biases (henceforth Risks 1.0), following the 2016 best seller:Weapons of Math Destruction. Two books published in 2022, The Chaos MachineandLike, Comment, Subscribe, raise additional risks to public health/safety/security such as genocide, insurrection, polarized politics, vaccinations (henceforth, Risks 2.0). These books suggest that the use of machine learning to maximize engagement in social media has created a Frankenstein Monster that is exploiting human weaknesses with persuasive technology, the illusory truth effect, Pavlovian conditioning, and Skinner’s intermittent variable reinforcement. Just as we cannot expect tobacco companies to sell fewer cigarettes and prioritize public health ahead of profits, so too, it may be asking too much of companies (and countries) to stop trafficking in misinformation given that it is so effective and so insanely profitable (at least in the short term). Eventually, we believe the current chaos will end, like the lawlessness in Wild West, because chaos is bad for business. As computer scientists, this paper will summarize criticisms from other fields and focus on implications for computer science; we will not attempt to contribute to those other fields. There is quite a bit of work in computer science on these risks, especially on Risks 1.0 (bias and fairness), but more work is needed, especially on Risks 2.0 (addictive, dangerous, and deadly). Kenneth Church 0001, Annika Marie Schoene, John E. Ortega, Raman Chandrasekar, Valia Kordoni |
Nat. Lang. Eng. | 1 |
| 2023 | Emerging trends: Smooth-talking machinesabstractAbstract Large language models (LLMs) have achieved amazing successes. They have done well on standardized tests in medicine and the law. That said, the bar has been raised so high that it could take decades to make good on expectations. To buy time for this long-term research program, the field needs to identify some good short-term applications for smooth-talking machines that are more fluent than trustworthy. Kenneth Church 0001, Richard Yue |
Nat. Lang. Eng. | 1 |
| 2022 | ArtELingo: A Million Emotion Annotations of WikiArt with Emphasis on Diversity over Language and CultureabstractYoussef Mohamed, Mohamed Abdelfattah, Shyma Alhuwaider, Feifan Li, Xiangliang Zhang, Kenneth Church, Mohamed Elhoseiny. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Youssef Mohamed, Mohamed Abdelfattah, Shyma Alhuwaider, Xiangliang Zhang 0001, Kenneth Church 0001, Mohamed Elhoseiny 0001 |
EMNLP | 6 |
| 2022 | W-CTC: a Connectionist Temporal Classification Loss with Wild Cards
Xingyu Cai, Jiahong Yuan, Yuchen Bian, Guangxu Xun, Jiaji Huang, Kenneth Church 0001 |
ICLR | 6 |
| 2022 | Training on Lexical ResourcesabstractWe propose using lexical resources (thesaurus, VAD) to fine-tune pretrained deep nets such as BERT and ERNIE. Then at inference time, these nets can be used to distinguish synonyms from antonyms, as well as VAD distances. The inference method can be applied to words as well as texts such as multiword expressions (MWEs), out of vocabulary words (OOVs), morphological variants and more. Code and data are posted on https://github.com/kwchurch/syn_ant. Kenneth Church 0001, Xingyu Cai, Yuchen Bian |
LREC | 1 |
| 2022 | Emerging trends: Deep nets thrive on scaleabstractAbstract Deep nets are becoming larger and larger in practice, with no respect for (non)-factors that ought to limit growth including the so-called curse of dimensionality (CoD). Donoho suggested that dimensionality can be a blessing as well as a curse. Current practice in industry is well ahead of theory, but there are some recent theoretical results from Weinan E’s group suggesting that errors may be independent of dimensions $d$ . Current practice suggests an even stronger conjecture: deep nets are not merely immune to CoD, but actually, deep nets thrive on scale. Kenneth Church 0001 |
Nat. Lang. Eng. | 1 |
| 2022 | Emerging trends: General fine-tuning (gft)abstractAbstract This paper describes gft (general fine-tuning), a little language for deep nets, introduced at an ACL-2022 tutorial. gft makes deep nets accessible to a broad audience including non-programmers. It is standard practice in many fields to use statistics packages such as R. One should not need to know how to program in order to fit a regression or classification model and to use the model to make predictions for novel inputs. With gft, fine-tuning and inference are similar to fit and predict in regression and classification. gft demystifies deep nets; no one would suggest that regression-like methods are “intelligent.” Kenneth Church 0001, Xingyu Cai, Yibiao Ying, Guangxu Xun, Yuchen Bian |
Nat. Lang. Eng. | 1 |
| 2022 | Emerging Trends: SOTA-ChasingabstractAbstract Many papers are chasing state-of-the-art (SOTA) numbers, and more will do so in the future. SOTA-chasing comes with many costs. SOTA-chasing squeezes out more promising opportunities such as coopetition and interdisciplinary collaboration. In addition, there is a risk that too much SOTA-chasing could lead to claims of superhuman performance, unrealistic expectations, and the next AI winter. Two root causes for SOTA-chasing will be discussed: (1) lack of leadership and (2) iffy reviewing processes. SOTA-chasing may be similar to the replication crisis in the scientific literature. The replication crisis is yet another example, like evaluation, of over-confidence in accepted practices and the scientific method, even when such practices lead to absurd consequences. Kenneth Church 0001, Valia Kordoni |
Nat. Lang. Eng. | 1 |
| 2021 | Decoupling Recognition and Transcription in Mandarin ASRabstractMuch of the recent literature on automatic speech recognition (ASR) is taking an end-to-end approach. Unlike English where the writing system is closely related to sound, Chinese characters (Hanzi) represent meaning, not sound. We propose factoring audio → Hanzi into two sub-tasks: (1) audio → Pinyin and (2) Pinyin → Hanzi, where Pinyin is a system of phonetic transcription of standard Chinese. Factoring the audio → Hanzi task in this way achieves 3.9% CER (character error rate) on the Aishell-1 corpus, the best result reported on this dataset so far. Jiahong Yuan, Xingyu Cai, Dongji Gao, Renjie Zheng, Liang Huang 0001, Kenneth Church 0001 |
ASRU | 6 |
| 2021 | Data Collection vs. Knowledge Graph Completion: What is Needed to Improve Coverage?abstractThis survey/position paper discusses ways to improve coverage of resources such as Word-Net.Rapp estimated correlations, ρ, between corpus statistics and psycholinguistic norms.ρ improves with quantity (corpus size) and quality (balance).1M words are enough for simple estimates (unigram frequencies), but at least 100M are required for pairs of words (word associations, edges).Knowledge Graph Completion (KGC) attempts to learn missing links in WN18.Unfortunately, WN18 is flawed with information leaking from train to test.More seriously, WN18 is based on SemCor (just 200k words) and dated (collected in 1960s).KGC cannot learn anything that happened since the 1960s, or associations requiring 100M words. Kenneth Church 0001, Yuchen Bian |
EMNLP (1) | 1 |
| 2021 | Large Margin Training Improves Language Models for ASRabstractLanguage models (LM) have been widely deployed in modern ASR systems. The LM is often trained by minimizing its perplexity on speech transcript. However, few studies try to discriminate a "gold" reference against inferior hypotheses. In this work, we propose a large margin language model (LMLM). LMLM is a general framework that enforces an LM to assign a higher score to the "gold" reference, and a lower one to the inferior hypothesis. The general framework is applied to three pretrained LM architectures: left-to-right LSTM, transformer encoder, and transformer decoder. Results show that LMLM can significantly outperform traditional LMs that are trained by minimizing perplexity. Especially for challenging noisy cases. Finally, among the three architectures, transformer encoder achieves the best performance. Jilin Wang, Jiaji Huang, Kenneth Church 0001 |
ICASSP | 3 |
| 2021 | Speaking Rate and Tonal Realization in Mandarin Chinese: What Can We Learn From Large Speech Corpora?abstractTwo Mandarin speech corpora were used to investigate tonal realization in terms of duration and pitch. The data consist of nearly 1000 hours of speech from more than 1600 speakers. The two corpora, both developed for ASR, differ in speaking rate by approximately 25%. This provides an opportunity to examine the influence of speaking rate on the realization of tones in natural speech. Our analysis found two differences for slower speaking rates: (1) lower "static" tones and (2) more change for "dynamic" tones. Tone 1 was higher and Tone 3 was lower on the first syllable of disyllabic words, suggesting a metrical structure of left-prominence. On the other hand, however, the second syllable was longer, and the slope of Tone 2 and Tone 4 was higher on the second syllable in one of the corpora, both of which suggest right-prominence. We also found a shift from right-prominence to left-prominence, with respect to the realization of the "dynamic" tones, when the speaking rate became slower. Our study demonstrated that both phrasing and metrical structure play an important role in tonal realization. Jiahong Yuan, Kenneth Church 0001 |
ICASSP | 2 |
| 2021 | Pause-Encoded Language Models for Recognition of Alzheimer's Disease and EmotionabstractWe propose enhancing Transformer language models (BERT, RoBERTa) to take advantage of pauses. Pauses play an important role in speech. In previous work we developed a method to encode pauses in transcripts for recognition of Alzheimer's disease. In this study, we extend this idea to language models. We re-train BERT and RoBERTa using a large collection of pause-encoded transcripts, and conduct fine- tuning for two downstream tasks, recognition of Alzheimer's disease and emotion. Pause-encoded language models outperform text-only language models on these tasks. Pause augmentation by duration perturbation for training is shown to improve pause-encoded language models. Jiahong Yuan, Xingyu Cai, Kenneth Church 0001 |
ICASSP | 3 |
| 2021 | Exploring Long Tail Visual Relationship Recognition with Large VocabularyabstractSeveral approaches have been proposed in recent literature to alleviate the long-tail problem, mainly in object classification tasks. In this paper, we make the first largescale study concerning the task of Long-Tail Visual Relationship Recognition (LTVRR). LTVRR aims at improving the learning of structured visual relationships that come from the long-tail (e.g., "rabbit grazing on grass"). In this setup, the subject, relation, and object classes each follow a long-tail distribution. To begin our study and make a future benchmark for the community, we introduce two LTVRR-related benchmarks, dubbed VG8K-LT and GQA-LT, built upon the widely used Visual Genome and GQA datasets. We use these benchmarks to study the performance of several state-of-the-art long-tail models on the LTVRR setup. Lastly, we propose a visiolinguistic hubless (VilHub) loss and a Mixup augmentation technique adapted to LTVRR setup, dubbed as RelMix. Both VilHub and RelMix can be easily integrated on top of existing models and despite being simple, our results show that they can remarkably improve the performance, especially on tail classes. Benchmarks, code, and models have been made available at: https://github.com/Vision-CAIR/LTVRR. Sherif Abdelkarim, Aniket Agarwal, Panos Achlioptas, Jun Chen 0021, Jiaji Huang, Boyang Li 0001, Kenneth Church 0001, Mohamed Elhoseiny 0001 |
ICCV | 7 |
| 2021 | Isotropy in the Contextual Embedding Space: Clusters and Manifolds
Xingyu Cai, Jiaji Huang, Yuchen Bian, Kenneth Church 0001 |
ICLR | 4 |
| 2021 | Speech Emotion Recognition with Multi-Task Learning
Xingyu Cai, Jiahong Yuan, Renjie Zheng, Liang Huang 0001, Kenneth Church 0001 |
Interspeech | 5 |
| 2021 | The Third DIHARD Diarization ChallengeabstractDIHARD III was the third in a series of speaker diarization challenges intended to improve the robustness of diarization systems to variability in recording equipment, noise conditions, and conversational domain. Speaker diarization was evaluated under two speech activity conditions (diarization from a reference speech activity vs. diarization from scratch) and 11 diverse domains. The domains span a range of recording conditions and interaction types, including read audio-books, meeting speech, clinical interviews, web videos, and, for the first time, conversational telephone speech. A total of 30 organizations (forming 21teams) from industry and academia submitted 499 valid system outputs. The evaluation results indicate that speaker diarization has improved markedly since DIHARD I, particularly for two-party interactions, but that for many domains (e.g., web video) the problem remains far from solved. Neville Ryant, Prachi Singh, Venkat Krishnamohan, Rajat Varma, Kenneth Church 0001, Christopher Cieri, Jun Du 0002, Sriram Ganapathy, Mark Y. Liberman |
Interspeech | 5 |
| 2021 | On Attention Redundancy: A Comprehensive StudyabstractYuchen Bian, Jiaji Huang, Xingyu Cai, Jiahong Yuan, Kenneth Church. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Yuchen Bian, Jiaji Huang, Xingyu Cai, Jiahong Yuan, Kenneth Church 0001 |
NAACL-HLT | 5 |
| 2021 | Exploiting a Zoo of Checkpoints for Unseen TasksabstractThere are so many models in the literature that it is difficult for practitioners to decide which combinations are likely to be effective for a new task. This paper attempts to address this question by capturing relationships among checkpoints published on the web. We model the space of tasks as a Gaussian process. The covariance can be estimated from checkpoints and unlabeled probing data. With the Gaussian process, we can identify representative checkpoints by a maximum mutual information criterion. This objective is submodular. A greedy method identifies representatives that are likely to "cover'' the task space. These representatives generalize to new tasks with superior performance. Empirical evidence is provided for applications from both computational linguistics as well as computer vision. Jiaji Huang, Kenneth Church 0001 |
NeurIPS | 3 |
| 2021 | Emerging trends: A gentle introduction to fine-tuningabstractAbstract The previous Emerging Trends article (Churchet al., 2021.Natural Language Engineering27(5), 631–645.) introduced deep nets to poets. Poets is an imperfect metaphor, intended as a gesture toward inclusion. The future for deep nets will benefit by reaching out to a broad audience of potential users, including people with little or no programming skills, and little interest in training models. That paper focused on inference, the use of pre-trained models, as is, without fine-tuning. The goal of this paper is to make fine-tuning more accessible to a broader audience. Since fine-tuning is more challenging than inference, the examples in this paper will require modest programming skills, as well as access to a GPU. Fine-tuning starts with a general purpose base (foundation) model and uses a small training set of labeled data to produce a model for a specific downstream application. There are many examples of fine-tuning in natural language processing (question answering (SQuAD) and GLUE benchmark), as well as vision and speech. Kenneth Church 0001, Yanjun Ma |
Nat. Lang. Eng. | 1 |
| 2021 | Emerging trends: Ethics, intimidation, and the Cold WarabstractAbstract There are well-meaning efforts to address ethics that will likely make the world a better place, but care needs to be taken to avoid repeating mistakes of the past. In particular, ACL has recently introduced a new process where there are special reviews of some papers for ethics. We would be more comfortable with the new ethics process if there were more checks and balances, due process and transparency. Otherwise, there is a risk that the process could intimidate authors in ways that are not that dissimilar from the ways that academics were intimidated during the Cold War on both sides of the Iron Curtain. Kenneth Church 0001, Valia Kordoni |
Nat. Lang. Eng. | 1 |
| 2021 | Emerging trends: Deep nets for poetsabstractAbstract Deep nets have done well with early adopters, but the future will soon depend on crossing the chasm. The goal of this paper is to make deep nets more accessible to a broader audience including people with little or no programming skills, and people with little interest in training new models. A github is provided with simple implementations of image classification, optical character recognition, sentiment analysis, named entity recognition, question answering (QA/SQuAD), machine translation, speech to text (SST), and speech recognition (STT). The emphasis is on instant gratification. Non-programmers should be able to install these programs and use them in 15 minutes or less (per program). Programs are short (10–100 lines each) and readable by users with modest programming skills. Much of the complexity is hidden behind abstractions such as pipelines and auto classes, and pretrained models and datasets provided by hubs: PaddleHub, PaddleNLP, HuggingFaceHub, and Fairseq. Hubs have different priorities than research. Research is training models from corpora and fine-tuning them for tasks. Users are already overwhelmed with an embarrassment of riches (13k models and 1k datasets). Do they want more? We believe the broader market is more interested in inference (how to run pretrained models on novel inputs) and less interested in training (how to create even more models). Kenneth Church 0001, Xiaopeng Yuan, Zewu Wu, Yehua Yang |
Nat. Lang. Eng. | 1 |
| 2020 | Improving Bilingual Lexicon Induction for Low Frequency WordsabstractThis paper designs a Monolingual Lexicon Induction task and observes that two factors accompany the degraded accuracy of bilingual lexicon induction for rare words.First, a diminishing margin between similarities in low frequency regime, and secondly, exacerbated hubness at low frequency.Based on the observation, we further propose two methods to address these two factors, respectively.The larger issue is hubness.Addressing that improves induction accuracy significantly, especially for low-frequency words. Jiaji Huang, Xingyu Cai, Kenneth Church 0001 |
EMNLP (1) | 3 |
| 2020 | Compositional Language Continual Learning
Yuanpeng Li 0001, Kenneth Church 0001, Mohamed Elhoseiny 0001 |
ICLR | 3 |
| 2020 | Disfluencies and Fine-Tuning Pre-Trained Language Models for Detection of Alzheimer's Disease
Jiahong Yuan, Yuchen Bian, Xingyu Cai, Jiaji Huang, Kenneth Church 0001 |
INTERSPEECH | 6 |
| 2020 | Emerging trends: Reviewing the reviewers (again)abstractAbstract The ACL-2019 Business meeting ended with a discussion of reviewing. Conferences are experiencing a success catastrophe. They are becoming bigger and bigger, which is not only a sign of success but also a challenge (for reviewing and more). Various proposals for reducing submissions were discussed at the Business meeting. IMHO, the problem is not so much too many submissions, but rather, random reviewing. We cannot afford to do reviewing as badly as we do (because that leads to even more submissions). Negative feedback loops are effective. The reviewing process will improve over time if reviewers teach authors how to write better submissions, and authors teach reviewers how to write more constructive reviews. If you have received a not-ok (unhelpful/offensive) review, please help program committees improve by sharing your not-ok reviews on social media. Kenneth Church 0001 |
Nat. Lang. Eng. | 1 |
| 2020 | Emerging trends: Subwords, seriously?abstractAbstract Subwords have become very popular, but the BERTa and ERNIEb tokenizers often produce surprising results. Byte pair encoding (BPE) trains a dictionary with a simple information theoretic criterion that sidesteps the need for special treatment of unknown words. BPE is more about training (populating a dictionary of word pieces) than inference (parsing an unknown word into word pieces). The parse at inference time can be ambiguous. Which parse should we use? For example, “electroneutral” can be parsed as electron-eu-tral or electro-neutral, and “bidirectional” can be parsed as bid-ire-ction-al and bi-directional. BERT and ERNIE tend to favor the parse with more word pieces. We propose minimizing the number of word pieces. To justify our proposal, a number of criteria will be considered: sound, meaning, etc. The prefix, bi-, has the desired vowel (unlike bid) and the desired meaning (bi is Latin for two, unlike bid, which is Germanic for offer). Kenneth Church 0001 |
Nat. Lang. Eng. | 1 |
| 2020 | Benchmarks and goalsabstractAbstract Benchmarks can be a useful step toward the goals of the field (when the benchmark is on the critical path), as demonstrated by the GLUE benchmark, and deep nets such as BERT and ERNIE. The case for other benchmarks such as MUSE and WN18RR is less well established. Hopefully, these benchmarks are on a critical path toward progress on bilingual lexicon induction (BLI) and knowledge graph completion (KGC). Many KGC algorithms have been proposed such as Trans[DEHRM], but it remains to be seen how this work improves WordNet coverage. Given how much work is based on these benchmarks, the literature should have more to say than it does about the connection between benchmarks and goals. Is optimizing P@10 on WN18RR likely to produce more complete knowledge graphs? Is MUSE likely to improve Machine Translation? Kenneth Church 0001 |
Nat. Lang. Eng. | 1 |
| 2019 | Hubless Nearest Neighbor Search for Bilingual Lexicon InductionabstractBilingual Lexicon Induction (BLI) is the task of translating words from corpora in two languages.Recent advances in BLI work by aligning the two word embedding spaces.Following that, a key step is to retrieve the nearest neighbor (NN) in the target space given the source word.However, a phenomenon called hubness often degrades the accuracy of NN.Hubness appears as some data points, called hubs, being extra-ordinarily close to many of the other data points.Reducing hubness is necessary for retrieval tasks.One successful example is Inverted SoFtmax (ISF), recently proposed to improve NN.This work proposes a new method, Hubless Nearest Neighbor (HNN), to mitigate hubness.HNN differs from NN by imposing an additional equal preference assumption.Moreover, the HNN formulation explains why ISF works as well as it does.Empirical results demonstrate that HNN outperforms NN, ISF and other state-ofthe-art.For reproducibility and follow-ups, we have published all code 1 . Jiaji Huang, Kenneth Church 0001 |
ACL (1) | 3 |
| 2019 | The Second DIHARD Diarization Challenge: Dataset, Task, and BaselinesabstractThis paper introduces the second DIHARD challenge, the second in a series of speaker diarization challenges intended to improve the robustness of diarization systems to variation in recording equipment, noise conditions, and conversational domain. The challenge comprises four tracks evaluating diarization performance under two input conditions (single channel vs. multi-channel) and two segmentation conditions (diarization from a reference speech segmentation vs. diarization from scratch). In order to prevent participants from overtuning to a particular combination of recording conditions and conversational domain, recordings are drawn from a variety of sources ranging from read audiobooks to meeting speech, to child language acquisition recordings, to dinner parties, to web video. We describe the task and metrics, challenge design, datasets, and baseline systems for speech enhancement, speech activity detection, and diarization. Neville Ryant, Kenneth Church 0001, Christopher Cieri, Alejandrina Cristià, Jun Du 0002, Sriram Ganapathy, Mark Y. Liberman |
INTERSPEECH | 2 |
| 2019 | Language Modeling at ScaleabstractWe show how Zipf's Law can be used to scale up language modeling (LM) to take advantage of more training data and more GPUs. LM plays a key role in many important natural language applications such as speech recognition and machine translation. Scaling up LM is important since it is widely accepted by the community that there is no data like more data. Eventually, we would like to train on terabytes (TBs) of text (trillions of words). Modern training methods are far from this goal, because of various bottlenecks, especially memory (within GPUs) and communication (across GPUs). This paper shows how Zipf's Law can address these bottlenecks by grouping parameters for common words and character sequences, because U ≪ N, where U is the number of unique words (types) and N is the size of the training set (tokens). For a local batch size K with G GPUs and a D-dimension embedding matrix, we reduce the original per-GPU memory and communication asymptotic complexity from Θ(GKD) to Θ(GK + UD). Empirically, we find U ∝ (GK)^0.64 on four publicly available large datasets. When we scale up the number of GPUs to 64, a factor of 8, training time speeds up by factors up to 6.7× (for character LMs) and 6.3× (for word LMs) with negligible loss of accuracy. Our weak scaling on 192 GPUs on the Tieba dataset shows a 35% improvement in LM prediction accuracy by training on 93 GB of data (2.5× larger than publicly available SOTA dataset), but taking only 1.25× increase in training time, compared to 3 GB of the same dataset running on 6 GPUs. Md. Mostofa Ali Patwary, Milind Chabbi, Heewoo Jun, Jiaji Huang, Gregory Frederick Diamos, Kenneth Church 0001 |
IPDPS | 6 |
| 2019 | GANs vs. good enoughabstractAbstract General Adversarial Networks are hot. Given Murphy’s Law, it is prudent to be paranoid. Best not to design for the average case. There is a long tradition of designing for the hundred-year flood (and five 9s reliability). What is good enough? Historically, the market hasn’t been willing to pay for five 9s. Hard to justify upfront costs for future benefits that will only payoff under unlikely scenarios, and might not work when needed. If the market isn’t willing to pay for five 9s, can we afford to design for the worst case? Kenneth Church 0001 |
Nat. Lang. Eng. | 1 |
| 2019 | A survey of 25 years of evaluationabstractAbstract Evaluation was not a thing when the first author was a graduate student in the late 1970s. There was an Artificial Intelligence (AI) boom then, but that boom was quickly followed by a bust and a long AI Winter. Charles Wayne restarted funding in the mid-1980s by emphasizing evaluation. No other sort of program could have been funded at the time, at least in America. His program was so successful that these days, shared tasks and leaderboards have become common place in speech and language (and Vision and Machine Learning). It is hard to remember that evaluation was a tough sell 25 years ago. That said, we may be a bit too satisfied with current state of the art. This paper will survey considerations from other fields such as reliability and validity from psychology and generalization from systems. There has been a trend for publications to report better and better numbers, but what do these numbers mean? Sometimes the numbers are too good to be true, and sometimes the truth is better than the numbers. It is one thing for an evaluation to fail to find a difference between man and machine, and quite another thing to pass the Turing Test. As Feynman said, “the first principle is that you must not fool yourself–and you are the easiest person to fool.” Kenneth Church 0001, Joel Hestness |
Nat. Lang. Eng. | 1 |
| 2018 | Enhancement and Analysis of Conversational Speech: JSALT 2017abstractAutomatic speech recognition is more and more widely and effectively used. Nevertheless, in some automatic speech analysis tasks the state of the art is surprisingly poor. One of these is “diarization”, the task of determining who spoke when. Diarization is key to processing meeting audio and clinical interviews, extended recordings such as police body cam or child language acquisition data, and any other speech data involving multiple speakers whose voices are not cleanly separated into individual channels. Overlapping speech, environmental noise and suboptimal recording techniques make the problem harder. During the JSALT Summer Workshop at CMU in 2017, an international team of researchers worked on several aspects of this problem, including calibration of the state of the art, detection of overlaps, enhancement of noisy recordings, and classification of shorter speech segments. This paper sketches the workshop's results, and announces plans for a “Diarization Challenge” to encourage further progress. Neville Ryant, Elika Bergelson, Kenneth Church 0001, Alejandrina Cristià, Jun Du 0002, Sriram Ganapathy, Sanjeev Khudanpur, Diana Kowalski, Mahesh Krishnamoorthy, Rajat Kulshreshta, Mark Y. Liberman, Yu-Ding Lu, Matthew Maciejewski, Florian Metze, Ján Profant, Lei Sun 0010, Yu Tsao 0001 |
ICASSP | 3 |
| 2018 | Emerging trends: A tribute to Charles WayneabstractAbstract Charles Wayne restarted funding in speech and language in the mid-1980s after a funding winter brought on by Pierce’s glamour-and-deceit criticisms in the ALPAC report and ‘Whither Speech Recognition’. Wayne introduced a new glamour-and-deceit-proof idea, an emphasis on evaluation. No other sort of program could have been funded at the time, at least in America. One could argue that Wayne has been so successful that the program no longer needs him to continue on. These days, shared tasks and leaderboards have become common place in speech and language (and vision and machine learning) research. That said, I am concerned that the community may not appreciate what it has got until it’s gone. Wayne has been doing much more than merely running competitions, but he did what he did in such a subtle Columbo-like way. Going forward, government funding is being eclipsed by consumer markets. Those of us with research to sell need to find more and more ways to be relevant to potential sponsors given this new world order. Kenneth Church 0001 |
Nat. Lang. Eng. | 1 |
| 2018 | Emerging trends: Artificial Intelligence, China and my new job at BaiduabstractAbstract It is amazing that I have written for more than a year on Emerging Trends without mentioning China’s investments in Artificial Intelligence (AI). Now that I have moved to Baidu, I would like to take this opportunity to share some of my personal observations with what’s happening in China over the past 25 years. The top universities in China have always been very good, but they are better today than they were 25 years ago, and they are on a trajectory to become the biggest and the best in the world. China is investing big time in what we do, both in the private sector and the public sector. Kai-Fu Lee is bullish on his investments in AI and China. There is a bold government plan for AI with specific milestones for parity with the West in 2020, major breakthroughs by 2025 and the envy of the world by 2030. Kenneth Church 0001 |
Nat. Lang. Eng. | 1 |
| 2018 | Emerging trends: APIs for speech and machine translation and moreabstractAbstract Lots of companies are offering lots of APIs. Reviews are not always as constructive as they could be. Some reviews encourage unproductive work on checkbox features that no one wants. It makes no sense to do the wrong thing badly. Constructive reviews should help focus priorities on what matters. Users care more about a great box opening experience than small improvements in word error rate and BLEU, popular metrics for speech and translation. In 15 minutes or less, can we teach potential users something new (and fun) that does something useful, such as how to translate PowerPoint between English and Chinese, preserving many of the features that are important to PowerPoint such as graphics and animations? See Appendix for the solution. Kenneth Church 0001 |
Nat. Lang. Eng. | 1 |
| 2017 | Speaker diarization: A perspective on challenges and opportunities from theory to practiceabstractThis paper discusses some challenges and opportunities in developing a speaker diarization system for operation on real world call center telephony data. We contrast some of the differences between a standard data set akin to NIST evaluations and those found in call centers. In exploring these differences we discovered vulnerabilities and proposed changes to address them. In moving from theory into practice we introduce two tasks in which speaker diarization and recognition can be leveraged. First, we show that speaker diarization and recognition systems can be integrated to find the common speaker (the call center agent) across multiple calls and consequently their role. Furthermore, once the role is determined the corresponding speech recognition output can be analyzed to determine the type of support call. Kenneth Church 0001, Weizhong Zhu, Josef Vopicka, Jason W. Pelecanos, Dimitrios Dimitriadis, Petr Fousek |
ICASSP | 1 |
| 2017 | Symbol Sequence Search from Telephone Conversation
Masayuki Suzuki, Gakuto Kurata, Abhinav Sethy, Bhuvana Ramabhadran, Kenneth Church 0001, Mark Drake |
INTERSPEECH | 5 |
| 2017 | Word2VecabstractAbstract My last column ended with some comments about Kuhn and word2vec. Word2vec has racked up plenty of citations because it satisifies both of Kuhn’s conditions for emerging trends: (1) a few initial (promising, if not convincing) successes that motivate early adopters (students) to do more, as well as (2) leaving plenty of room for early adopters to contribute and benefit by doing so. The fact that Google has so much to say on ‘How does word2vec work’ makes it clear that the definitive answer to that question has yet to be written. It also helps citation counts to distribute code and data to make it that much easier for the next generation to take advantage of the opportunities (and cite your work in the process). Kenneth Church 0001 |
Nat. Lang. Eng. | 1 |
| 2017 | Emerging trends: I did it, I did it, I did it, but. . abstractAbstract There has been a trend for publications to report better and better numbers, but less and less insight. The literature is turning into a giant leaderboard, where publication depends on numbers and little else (such as insight and explanation). It is considered a feature that machine learning has become so powerful (and so opaque) that it is no longer necessary (or even relevant) to talk about how it works. Insight is not only not required any more, but perhaps, insight is no longer even considered desirable. Transparency is good and opacity is bad. A recent best seller, Weapons of Math Destruction, is concerned that big data (and WMDs) increase inequality and threaten democracy largely because of opacity. Algorithms are being used to make lots of important decisions like who gets a loan and who goes to jail. If we tell the machine to maximize an objective function like making money, it will do exactly that, for better and for worse. Who is responsible for the consequences? Does it make it ok for machines to do bad things if no one knows what’s happening and why, including those of us who created the machines? Kenneth Church 0001 |
Nat. Lang. Eng. | 1 |
| 2017 | Emerging trends: InflationabstractAbstract Our field has enjoyed amazing growth over the years. Is this a good thing or a bad thing, or just a thing? Good: Growth sounds good. It is hard to imagine a politician arguing against jobs. There are more people working in the field than ever before, and they are publishing more and more, and creating more and more value. What could be wrong with that? Bad: Whatever you measure you get. We are all under too much pressure to publish too much too quickly. Students are graduating these days with more publications than what used to be expected for tenure. So many people are publishing so much that no one has time to think great thoughts, or take time to learn about things that may not be directly relevant to the next publication. Neutral: Inflation is a fact of life. There are long-term macro trends on publication rates that are beyond our control. These trends hold over tens and hundreds of years, and will continue over the foreseeable future. Kenneth Church 0001 |
Nat. Lang. Eng. | 1 |
| 2016 | The next generationabstractAbstract I’m sure you want me to tell you about the next new emerging trend, but I’m not going to do that. It is much easier to suggest where trends come from (the next generation), and how to distinguish passing fads (bubbles) from emerging trends. Young people are often the early adopters, the first to see what is about to happen, but most people don’t see what’s coming until well after the fact. Those with the most to lose (the establishment) tend to be the most resistant to change. Kenneth Church 0001 |
Nat. Lang. Eng. | 1 |
| 2014 | TALIP Perspectives, Guest Editorial Commentary: What Counts (and What Ought to Count)?abstracteditorial Free Access Share on TALIP Perspectives, Guest Editorial Commentary: What Counts (and What Ought to Count)? Author: Kenneth Church IBM IBMView Profile Authors Info & Claims ACM Transactions on Asian Language Information ProcessingVolume 13Issue 1February 2014 Article No.: 5pp 1–5https://doi.org/10.1145/2559789Published:01 February 2014Publication History 1citation286DownloadsMetricsTotal Citations1Total Downloads286Last 12 Months26Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my Alerts New Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Kenneth Church 0001 |
ACM Trans. Asian Lang. Inf. Process. | 1 |
| 2013 | Intent focused summarization of caller-agent conversationsabstractIn this paper, we propose a conditional random field (CRF) based to identify segments within call center conversations that convey caller intent. A distinguishing aspect of our approach is the use of context information of the intent bearing segments to predict the presence or absence of intents within various segments. The context is represented through a set of phrase features that are frequently present in and around the intent bearing segments. These phrases, identified in a data-driven manner, are used along with conventional word features in a CRF based sequence labeling framework to assign intent/non-intent labels to each utterance in a conversation. Another distinguishing aspect of our approach is that instead of using 1-best label alignment, we extract N-best label alignments at the output of CRF and combine evidences from them to rank the utterances according to their intent bearing potential, so that top ranked utterances can be chosen as the intent summary. To demonstrate the effectiveness of our approach and to evaluate the influence of automatic speech recognition (ASR) errors we evaluated our approach using manually transcribed and ASR transcribed conversations. Experimental results show improved summarization accuracy using our approach. Specifically, in 92% of the manually transcribed conversations accurate summaries of just one utterance length can be extracted using the proposed approach. Shajith Ikbal, Ashish Verma 0001, Prasanta Ghosh, Kenneth Church 0001, Jeffrey Marcus |
ICASSP | 4 |
| 2013 | A summary of the 2012 JHU CLSP workshop on zero resource speech technologies and models of early language acquisitionabstractWe summarize the accomplishments of a multi-disciplinary workshop exploring the computational and scientific issues surrounding zero resource (unsupervised) speech technologies and related models of early language acquisition. Centered around the tasks of phonetic and lexical discovery, we consider unified evaluation metrics, present two new approaches for improving speaker independence in the absence of supervision, and evaluate the application of Bayesian word segmentation algorithms to automatic subword unit tokenizations. Finally, we present two strategies for integrating zero resource techniques into supervised settings, demonstrating the potential of unsupervised methods to improve mainstream technologies. Aren Jansen, Emmanuel Dupoux, Sharon Goldwater, Mark Johnson 0001, Sanjeev Khudanpur, Kenneth Church 0001, Naomi Feldman, Hynek Hermansky, Florian Metze, Richard C. Rose, Mike Seltzer, Pascal Clark, Ian McGraw, Balakrishnan Varadarajan, Erin D. Bennett, Benjamin Börschinger, Justin T. Chiu, Ewan Dunbar, Abdellah Fourtassi, David F. Harwath, Chia-ying Lee, Keith D. Levin, Atta Norouzian, Vijayaditya Peddinti, Rachael Richardson, Thomas Schatz, Samuel Thomas 0001 |
ICASSP | 6 |
| 2013 | Deep neural network features and semi-supervised training for low resource speech recognitionabstractWe propose a new technique for training deep neural networks (DNNs) as data-driven feature front-ends for large vocabulary continuous speech recognition (LVCSR) in low resource settings. To circumvent the lack of sufficient training data for acoustic modeling in these scenarios, we use transcribed multilingual data and semi-supervised training to build the proposed feature front-ends. In our experiments, the proposed features provide an absolute improvement of 16% in a low-resource LVCSR setting with only one hour of in-domain training data. While close to three-fourths of these gains come from DNN-based features, the remaining are from semi-supervised training. Samuel Thomas 0001, Michael L. Seltzer, Kenneth Church 0001, Hynek Hermansky |
ICASSP | 3 |
| 2013 | Approximate inference: A sampling based modeling technique to capture complex dependencies in a language model
Anoop Deoras, Tomás Mikolov, Stefan Kombrink, Kenneth Church 0001 |
Speech Commun. | 4 |
| 2012 | Inverting the Point Process Model for Fast Phonetic Keyword SearchabstractNormally, we represent speech as a long sequence of frames and model the keyword with a relatively small set of parameters, commonly with a hidden Markov model (HMM). However, since the input speech is much longer than the keyword, suppose instead that we represent the speech as a relatively sparse set of impulses (roughly one per phoneme) and model the keyword as a filter-bank where each filter’s impulse response relates to the likelihood of a phone at a given position within a word. Evaluating keyword detections can then be seen as a convolution of an impulse train with an array of filters. This view enables huge speedups; runtime no longer depends on the frame rate and is instead linear in the number of events (impulses). We apply this intuition to redesign the runtime engine behind the point process model for keyword spotting. We demonstrate impressive real-time speedups (500,000x faster than real-time) with minimal loss in search accuracy. Keith Kintzley, Aren Jansen, Kenneth Church 0001, Hynek Hermansky |
INTERSPEECH | 3 |
| 2011 | Using Large Monolingual and Bilingual Corpora to Improve Coordination Disambiguation
Shane Bergsma, David Yarowsky, Kenneth Church 0001 |
ACL | 3 |
| 2011 | Bootstrapping a spoken language identification system using unsupervised integrated sensing and processing decision treesabstractIn many inference and learning tasks, collecting large amounts of labeled training data is time consuming and expensive, and oftentimes impractical. Thus, being able to efficiently use small amounts of labeled data with an abundance of unlabeled data-the topic of semi-supervised learning (SSL) [1]-has garnered much attention. In this paper, we look at the problem of choosing these small amounts of labeled data, the first step in a bootstrapping paradigm. Contrary to traditional active learning where an initial trained model is employed to select the unlabeled data points which would be most informative if labeled, our selection has to be done in an unsupervised way, as we do not even have labeled data to train an initial model. We propose using unsupervised clustering algorithms, in particular integrated sensing and processing decision trees (ISPDTs) [2], to select small amounts of data to label and subsequently use in SSL (e.g. transductive SVMs). In a language identification task on the CallFriend1 and 2003 NIST Language Recognition Evaluation corpora [3], we demonstrate that the proposed method results in significantly improved performance over random selection of equivalently sized training data. Damianos Karakos, Glen A. Coppersmith, Kenneth Church 0001, Sabato Marco Siniscalchi |
ASRU | 4 |
| 2011 | Estimating document frequencies in a speech corpusabstractInverse Document Frequency (IDF) is an important quantity in many applications, including Information Retrieval. IDF is defined in terms of document frequency, df (w), the number of documents that mention w at least once. This quantity is relatively easy to compute over textual documents, but spoken documents are more challenging. This paper considers two baselines: (1) an estimate based on the 1-best ASR output and (2) an estimate based on expected term frequencies computed from the lattice. We improve over these baselines by taking advantage of repetition. Whatever the document is about is likely to be repeated, unlike ASR errors, which tend to be more random (Poisson). In addition, we find it helpful to consider an ensemble of language models. There is an opportunity for the ensemble to reduce noise, assuming that the errors across language models are relatively uncorrelated. The opportunity for improvement is larger when WER is high. This paper considers a pairing task application that could benefit from improved estimates of df. The pairing task inputs conversational sides from the English Fisher corpus and outputs estimates of which sides were from the same conversation. Better estimates of df lead to better performance on this task. Damianos Karakos, Mark Dredze, Kenneth Church 0001, Aren Jansen, Sanjeev Khudanpur |
ASRU | 3 |
| 2011 | A Fast Re-scoring Strategy to Capture Long-Distance Dependencies
Anoop Deoras, Tomás Mikolov, Kenneth Church 0001 |
EMNLP | 3 |
| 2011 | Towards Unsupervised Training of Speaker Independent Acoustic ModelsabstractCan we automatically discover speaker independent phonemelike subword units with zero resources in a surprise language? There have been a number of recent efforts to automatically discover repeated spoken terms without a recognizer. This paper investigates the feasibility of using these results as constraints for unsupervised acoustic model training. We start with a relatively small set of word types, as well as their locations in the speech. The training process assumes that repetitions of the same (unknown) word share the same (unknown) sequence of subword units. For each word type, we train a whole-word hidden Markov model with Gaussian mixture observation densities and collapse correlated states across the word types using spectral clustering. We find that the resulting state clusters align reasonably well along phonetic lines. In evaluating cross-speaker word similarity, the proposed techniques outperform both raw acoustic features and language-mismatched acoustic models. Aren Jansen, Kenneth Church 0001 |
INTERSPEECH | 2 |
| 2010 | Using Web-scale N-grams to Improve Base NP Parsing Performance
Emily Pitler, Shane Bergsma, Dekang Lin, Kenneth Church 0001 |
COLING | 4 |
| 2010 | NLP on Spoken Documents Without ASR
Mark Dredze, Aren Jansen, Glen A. Coppersmith, Kenneth Church 0001 |
EMNLP | 4 |
| 2010 | Towards spoken term discovery at scale with zero resourcesabstractThe spoken term discovery task takes speech as input and identifies terms of possible interest. The challenge is to perform this task efficiently on large amounts of speech with zero resources (no training data and no dictionaries), where we must fall back to more basic properties of language. We find that long (∼ 1 s) repetitions tend to be contentful phrases (e.g. University of Pennsylvania) and propose an algorithm to search for these long repetitions without first recognizing the speech. To address efficiency concerns, we take advantage of (i) sparse feature representations and (ii) inherent low occurrence frequency of long content terms to achieve orders-of-magnitude speedup relative to the prior art. We frame our evaluation in the context of spoken document information retrieval, and demonstrate our method’s competence at identifying repeated terms in conversational telephone speech. Index Terms: spoken term discovery, zero resource speech recognition, dotplots Aren Jansen, Kenneth Church 0001, Hynek Hermansky |
INTERSPEECH | 2 |
| 2010 | New Tools for Web-Scale N-grams
Dekang Lin, Kenneth Church 0001, Heng Ji 0001, Satoshi Sekine, David Yarowsky, Shane Bergsma, Kailash Patil, Emily Pitler, Rachel Lathbury, Vikram Rao, Kapil Dalwani, Sushant Narsale |
LREC | 2 |
| 2009 | Has Computational Linguistics Become More Applied?
Kenneth Church 0001 |
CICLing | 1 |
| 2009 | Substring Statistics
Kyoji Umemura, Kenneth Church 0001 |
CICLing | 2 |
| 2009 | Using Word-Sense Disambiguation Methods to Classify Web Queries by Intent
Emily Pitler, Kenneth Church 0001 |
EMNLP | 2 |
| 2009 | A Data Structure for Sponsored SearchabstractInverted files have been very successful for document retrieval, but sponsored search is different. Inverted files are designed to find documents that match the query (all the terms in the query need to be in the document, but not vice versa). For sponsored search, ads are associated with bids. When a user issues a search query, bids are typically matched to the query using broad-match semantics: all the terms in the bid need to be in the query (but not vice versa). This means that the roles of the query and the bid/document are reversed in sponsored search, in turn making standard retrieval techniques based on inverted indexes ill-suited for sponsored search. This paper proposes novel index structures and query processing algorithms for sponsored search. We evaluate these structures using a real corpus of 180 million advertisements. Arnd Christian König, Kenneth Church 0001, Martin Markov |
ICDE | 2 |
| 2008 | Query suggestion using hitting timeabstractGenerating alternative queries, also known as query suggestion, has long been proved useful to help a user explore and express his information need. In many scenarios, such suggestions can be generated from a large scale graph of queries and other accessory information, such as the clickthrough. However, how to generate suggestions while ensuring their semantic consistency with the original query remains a challenging problem. Qiaozhu Mei, Dengyong Zhou, Kenneth Church 0001 |
CIKM | 3 |
| 2008 | On Delivering Embarrassingly Distributed Cloud Services
Kenneth Church 0001, Albert G. Greenberg, James R. Hamilton |
HotNets | 1 |
| 2008 | One sketch for all: Theory and Application of Conditional Random SamplingabstractConditional Random Sampling (CRS) was originally proposed for efficiently computing pairwise ($l_2$, $l_1$) distances, in static, large-scale, and sparse data sets such as text and Web data. It was previously presented using a heuristic argument. This study extends CRS to handle dynamic or streaming data, which much better reflect the real-world situation than assuming static data. Compared with other known sketching algorithms for dimension reductions such as stable random projections, CRS exhibits a significant advantage in that it is ``one-sketch-for-all.'' In particular, we demonstrate that CRS can be applied to efficiently compute the $l_p$ distance and the Hilbertian metrics, both are popular in machine learning. Although a fully rigorous analysis of CRS is difficult, we prove that, with a simple modification, CRS is rigorous at least for an important application of computing Hamming norms. A generic estimator and an approximate variance formula are provided and tested on various applications, for computing Hamming norms, Hamming distances, and $\chi^2$ distances. Ping Li 0001, Kenneth Church 0001, Trevor J. Hastie |
NIPS | 2 |
| 2008 | Entropy of search logs: how hard is search? with personalization? with backoff?abstractHow many pages are there on the Web? 5B? 20B? More? Less? Big bets on clusters in the clouds could be wiped out if a small cache of a few million urls could capture much of the value. Language modeling techniques are applied to MSN's search logs to estimate entropy. The perplexity is surprisingly small: millions, not billions. Qiaozhu Mei, Kenneth Church 0001 |
WSDM | 2 |
| 2007 | Nonlinear Estimators and Tail Bounds for Dimension Reduction in l 1 Using Cauchy Random Projections
Ping Li 0001, Trevor J. Hastie, Kenneth Church 0001 |
COLT | 3 |
| 2007 | Compressing Trigram Language Models With Golomb Coding
Kenneth Church 0001, Ted Hart, Jianfeng Gao 0001 |
EMNLP-CoNLL | 1 |
| 2007 | Heavy-tailed distributions and multi-keyword queriesabstractIntersecting inverted indexes is a fundamental operation for many applications in information retrieval and databases. Efficient indexing for this operation is known to be a hard problem for arbitrary data distributions. However, text corpora used in Information Retrieval applications often have convenient power-law constraints (also known as Zipf’s Law and long tails) that allow us to materialize carefully chosen combinations of multi-keyword indexes, which significantly improve worst-case performance without requiring excessive storage. These multi-keyword indexes limit the number of postings accessed when computing arbitrary index intersections. Our evaluation on an e-commerce collection of 20 million products shows that the indexes of up to four arbitrary keywords can be intersected while accessing less than 20 % of the postings in the largest single-keyword index. Surajit Chaudhuri, Kenneth Church 0001, Arnd Christian König, Liying Sui |
SIGIR | 2 |
| 2007 | The wild thing goes localabstractSuppose you are on a mobile device with no keyboard (e.g., a cell phone) and you want to perform a "near me" search. Where is the nearest pizza? How do you enter queries quickly? T9? The Wild Thing encourages users to enter patterns with implicit and explicit wild cards (regular expressions). The search engine uses Microsoft Local Live logs to find the most likely queries for a particular location. For example, 7#6 is short-hand for the regular expression: /^[PQRS].*[ ][MNO].*/, which matches "post office" in many places (but "Space Needle" in Seattle). Some queries are more local than others. Pizza is likely everywhere, whereas "Boeing Company," is very likely in Seattle and Chicago, moderately likely nearby, and somewhat likely elsewhere. Smoothing is important. Not every query is observed everywhere. Kenneth Church 0001, Bo Thiesson |
SIGIR | 1 |
| 2007 | A Sketch Algorithm for Estimating Two-Way and Multi-Way AssociationsabstractWe should not have to look at the entire corpus (e.g., the Web) to know if two (or more) words are strongly associated or not. One can often obtain estimates of associations from a small sample. We develop a sketch-based algorithm that constructs a contingency table for a sample. One can estimate the contingency table for the entire population using straightforward scaling. However, one can do better by taking advantage of the margins (also known as document frequencies). The proposed method cuts the errors roughly in half over Broder's sketches. Ping Li 0001, Kenneth Church 0001 |
Comput. Linguistics | 2 |
| 2007 | Nonlinear Estimators and Tail Bounds for Dimension Reduction in l1 Using Cauchy Random Projections
Ping Li 0001, Trevor J. Hastie, Kenneth Church 0001 |
J. Mach. Learn. Res. | 3 |
| 2006 | Improving Random Projections Using Marginal Information
Ping Li 0001, Trevor J. Hastie, Kenneth Church 0001 |
COLT | 3 |
| 2006 | Very sparse random projectionsabstractThere has been considerable interest in random projections, an approximate algorithm for estimating distances between pairs of points in a high-dimensional vector space. Let A in Rn x D be our n points in D dimensions. The method multiplies A by a random matrix R in RD x k, reducing the D dimensions down to just k for speeding up the computation. R typically consists of entries of standard normal N(0,1). It is well known that random projections preserve pairwise distances (in the expectation). Achlioptas proposed sparse random projections by replacing the N(0,1) entries in R with entries in -1,0,1 with probabilities 1/6, 2/3, 1/6, achieving a threefold speedup in processing time.We recommend using R of entries in -1,0,1 with probabilities 1/2√D, 1-1√D, 1/2√D for achieving a significant √D-fold speedup, with little loss in accuracy. Ping Li 0001, Trevor J. Hastie, Kenneth Church 0001 |
KDD | 3 |
| 2006 | Conditional Random Sampling: A Sketch-based Sampling Technique for Sparse DataabstractWe1 develop Conditional Random Sampling (CRS), a technique particularly suit- able for sparse data. In large-scale applications, the data are often highly sparse. CRS combines sketching and sampling in that it converts sketches of the data into conditional random samples online in the estimation stage, with the sample size determined retrospectively. This paper focuses on approximating pairwise l2 and l1 distances and comparing CRS with random projections. For boolean (0/1) data, CRS is provably better than random projections. We show using real-world data that CRS often outperforms random projections. This technique can be applied in learning, data mining, information retrieval, and database query optimizations. Ping Li 0001, Kenneth Church 0001, Trevor J. Hastie |
NIPS | 2 |
| 2005 | The Wild Thing
Kenneth Church 0001, Bo Thiesson |
ACL | 1 |
| 2005 | Reviewing the ReviewersabstractRecall is More Subtle than PrecisionI just returned from the Association for Computational Linguistics' 43rd Annual Meeting (ACL-2005).The acceptance rate was 18%.Is this a good thing or a bad thing?When the acceptance rate is low, precision tends to be high.The audience can judge precision for itself.If the presentations are good, everyone knows it.And if they aren't, they know that as well.ACL-2005 had great precision.Recall is more subtle.When there is an issue with recall, it isn't immediately obvious to everyone.If you listen closely, you'll hear some grumbling in the halls.And then the rejects start to appear elsewhere.ACL's low recall has been great for other conferences.The best of the rejects are very good, better than most of the accepted papers, and often strong contenders for the best paper award at EMNLP.I used to be surprised by the quality of these rejects, but after seeing so many great rejects over so many years, I am no longer surprised by anything.The practice of setting EMNLP's submission date immediately after ACL's notification date is a not-so-subtle hint: Please do something about the low recall.When you read some of the ACL reviews for these top EMNLP papers, you realize what is happening.ACL reviewing is paying too much attention to abstentions (and objections from people outside the area).If a reviewer isn't qualified to say anything on a particular topic, that's okay.An abstention shouldn't kill a paper.Controversial papers are great; boring unobjectionable incremental papers are not.The only bad paper is a paper without an advocate.A paper with a single advocate should trump a paper with lots of seconds, but no advocates.Don't average votes.The key votes are the advocates.Negative votes matter only if they convince the advocates to change their votes.Recall is a problem for many conferences, not just ACL; SIGIR, for example, rejected the classic paper on page rank, a hugely successful paper in terms of citations, perhaps more successful than anything SIGIR ever published. Kenneth Church 0001 |
Comput. Linguistics | 1 |
| 2004 | The submatrices character count problem: an efficient solution using separable values
Amihood Amir, Kenneth Church 0001, Emanuel Dar |
Inf. Comput. | 2 |
| 2003 | Speech and language processing: where have we been and where are we going?abstractCan we use the past to predict the future? Moore’s Law is a great example: performance doubles and prices halve approximately every 18 months. This trend has held up well to the test of time and is expected to continue for some time. Similar arguments can be found in speech demonstrating consistent progress over decades. Unfortunately, there are also cases where history repeats itself, as well as major dislocations, fundamental changes that invalidate fundamental assumptions. What will happen, for example, when petabytes become a commodity? Can demand keep up with supply? How much text and speech would it take to match this supply? Priorities will change. Search will become more important than coding and dictation. Kenneth Church 0001 |
INTERSPEECH | 1 |
| 2002 | NLP Found Helpful (at least for one Text Categorization Task)abstractAttempts to use natural language processing (NLP) for text categorization and information retrieval (IR) have had mixed results. Nevertheless, there is a strong intuition that NLP is important, at least for some tasks. In this paper, we discuss a task involving captioned images for which the subject and the predicate are critical. The usefulness of NLP for this task is established in two ways. In addition to the standard method of introducing a new system and comparing its performance with others in the literature, we also present evidence from experiments with human subjects showing that NLP generally improves speed and accuracy. Carl L. Sable, Kathy McKeown, Kenneth Church 0001 |
EMNLP | 3 |
| 2002 | Separable attributes: a technique for solving the sub matrices character count problem
Amihood Amir, Kenneth Church 0001, Emanuel Dar |
SODA | 2 |
| 2002 | Dedication to William A. GaleabstractGale, as he liked to be called by his friends and family, had extremely broad interests, both professionally and otherwise. His professional career at Bell Labs included radio astronomy, economics, statistics and computational linguistics. Kenneth Church 0001 |
Nat. Lang. Eng. | 1 |
| 2001 | Using Bins to Empirically Estimate Term Weights for Text Categorization
Carl L. Sable, Kenneth Church 0001 |
EMNLP | 2 |
| 2001 | Using Suffix Arrays to Compute Term Frequency and Document Frequency for All Substrings in a CorpusabstractBigrams and trigrams are commonly used in statistical natural language processing; this paper will describe techniques for working with much longer n-grams. Suffix arrays (Manber and Myers 1990) were first introduced to compute the frequency and location of a substring (n-gram) in a sequence (corpus) of length N. To compute frequencies over all N(N+1)/2 substrings in a corpus, the substrings are grouped into a manageable number of equivalence classes. In this way, a prohibitive computation over substrings is reduced to a manageable computation over classes. This paper presents both the algorithms and the code that were used to compute term frequency (tf) and document frequency (df) for all n-grams in two large corpora, an English corpus of 50 million words of Wall Street Journal and a Japanese corpus of 216 million characters of Mainichi Shimbun. The second half of the paper uses these frequencies to find “interesting” substrings. Lexicographers have been interested in n-grams with high mutual information (MI) where the joint term frequency is higher than what would be expected by chance, assuming that the parts of the n-gram combine independently. Residual inverse document frequency (RIDF) compares document frequency to another model of chance where terms with a particular term frequency are distributed randomly throughout the collection. MI tends to pick out phrases with noncompositional semantics (which often violate the independence assumption) whereas RIDF tends to highlight technical terminology, names, and good keywords for information retrieval (which tend to exhibit nonrandom distributions over documents). The combination of both MI and RIDF is better than either by itself in a Japanese word extraction task. Mikio Yamamoto, Kenneth Church 0001 |
Comput. Linguistics | 2 |
| 2000 | Empirical Estimates of Adaptation: The chance of Two Noriegas is closer to p/2 than p2
Kenneth Church 0001 |
COLING | 1 |
| 2000 | Empirical Term Weighting and Expansion FrequencyabstractWe propose an empirical method for estimating term weights directly from relevance judgments, avoiding various standard but potentially trouble-some assumptions. It is common to assume, for example, that weights vary with term frequency (tf) and inverse document frequency (idf) in a particular way, e.g., tf .idf, but the fact that there are so many variants of this formula in the literature suggests that there remains considerable uncertainty about these assumptions. Our method is similar to the Berkeley regression method where labeled relevance judgments are fit as a linear combination of (transforms of) tf, idf, etc. Training methods not only improve performance, but also extend naturally to include additional factors such as burstiness and query expansion. The proposed histogram-based training method provides a simple way to model complicated interactions among factors such as tf, idf, burstiness and expansion frequency (a generalization of query expansion). The correct handling of expanded term is realized based on statistical information. Expansion frequency dramatically improves performance from a level comparable to BKJJBIDS, Berkeley's entry in the Japanese NACSIS NTCIR-1 evaluation for short queries, to the level of JCB1, the top system in the evaluation. JCB1 uses sophisticated (and proprietary) natural language processing techniques developed by Just System, a leader in the Japanese word-processing industry. We are encouraged that the proposed method, which is simple to understand and replicate, can reach this level of performance. Kyoji Umemura, Kenneth Church 0001 |
EMNLP | 2 |
| 2000 | Engineering the compression of massive tables: an experimental approach
Adam L. Buchsbaum, Donald F. Caldwell, Kenneth Church 0001, Glenn S. Fowler, S. Muthukrishnan 0001 |
SODA | 3 |
| 1999 | What's Happened Since the First SIGDAT Meeting?
Kenneth Church 0001 |
EMNLP | 1 |
| 1997 | Termight: Coordinating Humans and Machines in Bilingual Terminology Acquisition
Ido Dagan, Kenneth Church 0001 |
Mach. Transl. | 2 |
| 1997 | Preface
Pierre Isabelle, Kenneth Church 0001 |
Mach. Transl. | 2 |
| 1995 | One Term Or Two?abstractArticle One term or two? Share on Author: Kenneth Ward Church Kenneth Ward Church, AT&T Bell Laboratories, Murray Hill, NJ Kenneth Ward Church, AT&T Bell Laboratories, Murray Hill, NJView Profile Authors Info & Claims SIGIR '95: Proceedings of the 18th annual international ACM SIGIR conference on Research and development in information retrievalJuly 1995 Pages 310–318https://doi.org/10.1145/215206.215376Online:01 July 1995Publication History 19citation501DownloadsMetricsTotal Citations19Total Downloads501Last 12 Months3Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Kenneth Church 0001 |
SIGIR | 1 |
| 1995 | Poisson mixturesabstractAbstract Shannon (1948) showed that a wide range of practical problems can be reduced to the problem of estimating probability distributions of words and ngrams in text. It has become standard practice in text compression, speech recognition, information retrieval and many other applications of Shannon's theory to introduce a “bag-of-words” assumption. But obviously, word rates vary from genre to genre, author to author, topic to topic, document to document, section to section, and paragraph to paragraph. The proposed Poisson mixture captures much of this heterogeneous structure by allowing the Poisson parameter θ to vary over documents subject to a density function φ. φ is intended to capture dependencies on hidden variables such genre, author, topic, etc. (The Negative Binomial is a well-known special case where φ is a Г distribution.) Poisson mixtures fit the data better than standard Poissons, producing more accurate estimates of the variance over documents (σ 2 ), entropy (H), inverse document frequency (IDF), and adaptation (Pr( x ≥ 2/ x ≥ 1)). Kenneth Church 0001, William A. Gale |
Nat. Lang. Eng. | 1 |
| 1994 | Fax: An Alternative to SGML
Kenneth Church 0001, William A. Gale, Jonathan Helfman, David D. Lewis |
COLING | 1 |
| 1994 | K-vec: A New Approach for Aligning Parallel Texts
Pascale Fung, Kenneth Church 0001 |
COLING | 2 |
| 1994 | Using OCR and equalization to downsample documentsabstractDocuments need to be sampled at different rates for different output devices: 300-600 dpi for laser printers, 100-200 dpi for fax, and 75-100 dpi for bitmap terminals. To output a high resolution document on a low resolution device, it may be necessary to introduce downsampling. Standard signal processing techniques such as linear filtering and decimation don't work very well at low resolutions. Better results are obtained by a nonlinear filtering technique we introduce in this paper, called nonlinear document equalization. Even better results are obtained by taking advantage of fonts designed specifically for bitmap terminals and other low resolution devices. However, character-level information is required to make use of fonts. This information is not always available; OCR is not 100% accurate. We propose a hybrid approach: downsample by font substitution when possible, and decimate when necessary. Unfortunately, the result tends to look like a "ransom note". Equalization is used to blend the two cases together so that gaps in the OCR analysis become almost unnoticeable. Oscar E. Agazzi, Kenneth Church 0001, William A. Gale |
ICPR (2) | 2 |
| 1993 | Char_align: A Program for Aligning Parallel Texts at the Character LevelabstractThere have been a number of recent papers on aligning parallel texts at the sentence level, e.g., Brown et al (1991), Gale and Church (to appear), Isabelle (1992), Kay and Rösenschein (to appear), Simard et al (1992), Warwick-Armstrong and Russell (1990). On clean inputs, such as the Canadian Hansards, these methods have been very successful (at least 96% correct by sentence). Unfortunately, if the input is noisy (due to OCR and/or unknown markup conventions), then these methods tend to break down because the noise can make it difficult to find paragraph boundaries, let alone sentences. This paper describes a new program, char_align, that aligns texts at the character level rather than at the sentence/paragraph level, based on the cognate approach proposed by Simard et al. Kenneth Church 0001 |
ACL | 1 |
| 1993 | Introduction to the Special Issue on Computational Linguistics Using Large Corpora
Kenneth Church 0001, Robert L. Mercer |
Comput. Linguistics | 1 |
| 1993 | A Program for Aligning Sentences in Bilingual Corpora
William A. Gale, Kenneth Church 0001 |
Comput. Linguistics | 2 |
| 1993 | Good applications for crummy machine translation
Kenneth Church 0001, Eduard H. Hovy |
Mach. Transl. | 1 |
| 1992 | Estimating Upper and Lower Bounds on the Performance of Word-Sense Disambiguation ProgramsabstractWe have recently reported on two new word-sense disambiguation systems, one trained on bilingual material (the Canadian Hansards) and the other trained on monolingual material (Roget's Thesaurus and Grolier's Encyclopedia). After using both the monolingual and bilingual classifiers for a few months, we have convinced ourselves that the performance is remarkably good. Nevertheless, we would really like to be able to make a stronger statement, and therefore, we decided to try to develop some more objective evaluation measures. Although there has been a fair amount of literature on sense-disambiguation, the literature does not offer much guidance in how we might establish the success or failure of a proposed solution such as the two systems mentioned in the previous paragraph. Many papers avoid quantitative evaluations altogether, because it is so difficult to come up with credible estimates of performance.This paper will attempt to establish upper and lower bounds on the level of performance that can be expected in an evaluation. An estimate of the lower bound of 75% (averaged over ambiguous types) is obtained by measuring the performance produced by a baseline system that ignores context and simply assigns the most likely sense in all cases. An estimate of the upper bound is obtained by assuming that our ability to measure performance is largely limited by our ability obtain reliable judgments from human informants. Not surprisingly, the upper bound is very dependent on the instructions given to the judges. Jorgensen, for example, suspected that lexicographers tend to depend too much on judgments by a single informant and found considerable variation over judgments (only 68% agreement), as she had suspected. In our own experiments, we have set out to find word-sense disambiguation tasks where the judges can agree often enough so that we could show that they were outperforming the baseline system. Under quite different conditions, we have found 96.8% agreement over judges. William A. Gale, Kenneth Church 0001, David Yarowsky |
ACL | 2 |
| 1991 | A Program for Aligning Sentences in Bilingual CorporaabstractResearchers in both machine translation (e.g., Brown et al., 1990) and bilingual lexicography (e.g., Klavans and Tzoukermann, 1990) have recently become interested in studying parallel texts, texts such as the Canadian Hansards (parliamentary proceedings) which are available in multiple languages (French and English). This paper describes a method for aligning sentences in these parallel texts, based on a simple statistical model of character lengths. The method was developed and tested on a small trilingual sample of Swiss economic reports. A much larger sample of 90 million words of Canadian Hansards has been aligned and donated to the ACL/DCI. William A. Gale, Kenneth Church 0001 |
ACL | 2 |
| 1990 | A Spelling Correction Program Based on a Noisy Channel Model
Mark D. Kernighan, Kenneth Church 0001, William A. Gale |
COLING | 2 |
| 1990 | Word Association Norms, Mutual Information, and Lexicography
Kenneth Church 0001, Patrick Hanks |
Comput. Linguistics | 1 |
| 1989 | Word Association Norms, Mutual Information and LexicographyabstractThe term word assaciation is used in a very particular sense in the psycholinguistic literature.(Generally speaking, subjects respond quicker than normal to the word "nurse" if it follows a highly associated word such as "doctor.")We wilt extend the term to provide the basis for a statistical description of a variety of interesting linguistic phenomena, ranging from semantic relations of the doctor/nurse type (content word/content word) to lexico-syntactic co-occurrence constraints between verbs and prepositions (content word/function word).This paper will propose a new objective measure based on the information theoretic notion of mutual information, for estimating word association norms from computer readable corpora.(The standard method of obtaining word association norms, testing a few thousand subjects on a few hundred words, is both costly and unreliable.)The , proposed measure, the association ratio, estimates word association norms directly from computer readable corpora, waki,~g it possible to estimate norms for tens of thousands of words. I. Meaning and AssociationIt is common practice in linguistics to classify words not only on the basis of their meanings but also on the basis of their co-occurrence with other words.Running through the whole Firthian tradition, for example, is the theme that "You shall know a word by the company it keeps" (Firth, 1957)."On the one hand, bank ¢o.occors with words and expression such u money, nmu.loan, account, ~m.c~z~c.o~.ctal, manager, robbery, vaults, wortln# in a, lu action, Fb~Nadonal. of F.ngland, and so forth.On the other hand, we find bank m-occorring with r~r.~bn, boa:.am (end of course West and Sou~, which have tcqu/red special meanings of their own), on top of the, and of the Rhine. Kenneth Church 0001, Patrick Hanks |
ACL | 1 |
| 1989 | A stochastic parts program and noun phrase parser for unrestricted textabstractA program that tags each word in an input sentence with the most likely part of speech has been written. The program uses a linear-time dynamic programming algorithm to find an assignment of parts of speech to words that optimizes the product of (a) lexical probabilities (probability of observing part of speech i given word i) and (b) contextual probabilities (probability of observing part of speech i given n following parts of speech). Program performance is encouraging; a 400-word sample is presented and is judged to be 99.5% correct.> Kenneth Church 0001 |
ICASSP | 1 |
| 1988 | Complexity, two-level morphology and Finnish
Kimmo Koskenniemi, Kenneth Church 0001 |
COLING | 2 |
| 1986 | Morphoogicai Decomposition and Stress Assignment for Speech SynthesisabstractArticle Free AccessMorphological decomposition and stress assignment for speech synthesis Share on Author: Kenneth Church Bell Laboratories, Murray Hill, N.J. Bell Laboratories, Murray Hill, N.J.View Profile Authors Info & Claims ACL '86: Proceedings of the 24th annual meeting on Association for Computational LinguisticsJuly 1986 Pages 156–164https://doi.org/10.3115/981131.981154Online:10 July 1986Publication History 0citation264DownloadsMetricsTotal Citations0Total Downloads264Last 12 Months55Last 6 weeks6 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Kenneth Church 0001 |
ACL | 1 |
| 1986 | Stress assignment in letter to sound rules for speech synthesisabstractThis paper will discuss how to determine word stress from spelling. Stress assignment is a well-established weak point for many speech synthesizers because stress dependencies cannot be determined locally, by looking through a five or six character window, as many speech synthesizers do. Examples such as degráde / dègradátion and télegraph / telégraphy demonstrate that stress dependencies can span over two and three syllables. Etymology is an even more striking problem for local determination of stress. Consider cálculi / tortóni which should have the same stress pattern since they have the same sequence of stops, liquids and vowels. However, tortóni violates the English main stress rule (which is derived from Latin) and takes penultimate stress like most other Italian loan words: Aldrighetti, Angeletti, Bellotti Iannucci, Italiano, Lombardino, Marconi, Morillo, Olivetti. Stressing Italian names with the Latin pattern yields amusing results as will be demonstrated. Kenneth Church 0001 |
ICASSP | 1 |
| 1985 | Stress Assignment in Letter to Sound Rules for Speech SynthesisabstractThis paper will discuss how to determine word stress from spelling. Stress assignment is a well-established weak point for many speech synthesizers because stress dependencies cannot be determined locally. It is impossible to determine the stress of a word by looking through a five or six character window, as many speech synthesizers do. Well-known examples such as degráde / dègradátion and télegraph / telégraphy demonstrate that stress dependencies can span over two and three syllables. This paper will present a principled framework for dealing with these long distance dependencies. Stress assignment will be formulated in terms of Waltz' style constraint propagation with four sources of constraints: (1) syllable weight, (2) part of speech, (3) morphology and (4) etymology. Syllable weight is perhaps the most interesting, and will be the main focus of this paper. Most of what follows has been implemented. Kenneth Church 0001 |
ACL | 1 |
| 1983 | A Finite-State Parser for Use in Speech RecognitionabstractThis paper is divided into two parts. The first section motivates the application of finite-state parsing techniques at the phonetic level in order to exploit certain classes of contextual constraints. In the second section, the parsing framework is extended in order to account for 'feature spreading' (e.g., agreement and co-articulation) in a natural way. Kenneth Church 0001 |
ACL | 1 |
| 1983 | Allophonic and Phonotactic Constraints Are Useful
Kenneth Church 0001 |
IJCAI | 1 |
| 1982 | Coping with Syntactic Ambiguity or How to Put the Block in the Box on the Table
Kenneth Church 0001, Ramesh S. Patil |
Am. J. Comput. Linguistics | 1 |
| 1980 | On Parsing Strategies and ClosureabstractThis paper proposes a welcom hypothesis: a computationally simple deviceis sufficient for processing natural language. Traditionally it has been argued that processing natural language syntax requires very powerful machinery. Many engineers have come to this rather grim conclusion; almost all working parsers are actually Turing Machines (TM). For example, Woods believed that a parser should have TM complexity and specifically designed his Augmented Transition Networks (ATNs) to be Turing Equivalent.(1) "It is well known (cf. [Chomsky64]) that the strict context-free grammar model is not an adequate mechanism for characterizing the subtleties of natural languages." [Woods70]If the problem is really as hard as it appears, then the only solution is to grin and bear it. Our own position is that parsing acceptable sentences is simpler because there are constraints on human performance that drastically reduce the computational complexity. Although Woods correctly observes that competence models are very complex, this observation may not apply directly to a performance problem such as parsing.The claim is that performance limitations actually reduce parsing complexity. This suggests two interesting questions: (a) How is the performance model constrained so as to reduce its complexity, and (b) How can the constrained performance model naturally approximate competence idealizations? Kenneth Church 0001 |
ACL | 1 |
| 1979 | Co-ordinate Square: Solution to Many Chess Pawn Endgames
Kenneth Church 0001 |
IJCAI | 1 |