Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Martin Jansche

dblp:07/816 · DBLP profile ↗
← Back
25ranked-venue papers
9as first author
0since 2021 · last 2020
0000-0003-0484-4879ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 8 first-authorGraphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-authorDatabases, data management, data science and information retrieval · 2Human-computer interaction and ubiquitous computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Information extraction and text analysis · 46% Probabilistic and Bayesian machine learning · 24% Speech recognition and synthesis · 15%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 9 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › document retrieval › spoken document retrieval
spoken term detection
0.112009
Web derived pronunciations for spoken term detection · SIGIR 2009
Natural language and speech › Information extraction and text analysis › sequence labeling
binary sequence labeling
0.112007
A Maximum Expected Utility Framework for Binary Sequence Labeling · ACL 2007
Natural language and speech › Information extraction and text analysis
sequence labeling
0.112007
A Maximum Expected Utility Framework for Binary Sequence Labeling · ACL 2007
Machine learning › Kernel, tree and ensemble methods › support vector machine
support vector regression
0.112007
A Support Vector Approach to Censored Targets · ICDM 2007
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
survival analysis
0.112007
A Support Vector Approach to Censored Targets · ICDM 2007
Machine learning › Probabilistic and Bayesian machine learning
count data modeling
0.012003
Parametric Models of Linguistic Count Data · ACL 2003
Natural language and speech › Information extraction and text analysis
text classification
0.012003
Parametric Models of Linguistic Count Data · ACL 2003
Natural language and speech › Speech recognition and synthesis
automatic speech recognition
0.012002
Information Extraction from Voicemail Transcripts · EMNLP 2002
Natural language and speech › Speech recognition and synthesis
spoken language understanding
0.012002
Information Extraction from Voicemail Transcripts · EMNLP 2002

Methods — techniques the papers use, named apart from their topics

maximum expected utility framework · 0.1web mining · 0.1IPA normalization · 0.1support vector regression · 0.1convex optimization · 0.1gamma-poisson mixture · 0.0beta-binomial mixture · 0.0named entity recognition · 0.0
YearPublicationVenuePosition
2020 Open-source Multi-speaker Speech Corpora for Building Gujarati, Kannada, Malayalam, Marathi, Tamil and Telugu Speech Synthesis Systems
abstract
We present free high quality multi-speaker speech corpora for Gujarati, Kannada, Malayalam, Marathi, Tamil and Telugu, which are six of the twenty two official languages of India spoken by 374 million native speakers. The datasets are primarily intended for use in text-to-speech (TTS) applications, such as constructing multilingual voices or being used for speaker or language adaptation. Most of the corpora (apart from Marathi, which is a female-only database) consist of at least 2,000 recorded lines from female and male native speakers of the language. We present the methodological details behind corpora acquisition, which can be scaled to acquiring data for other languages of interest. We describe the experiments in building a multilingual text-to-speech model that is constructed by combining our corpora. Our results indicate that using these corpora results in good quality voices, with Mean Opinion Scores (MOS) > 3.6, for all the languages tested. We believe that these resources, released with an open-source license, and the described methodology will help in the progress of speech applications for the languages described and aid corpora development for other, smaller, languages of India and beyond.
Shan-Hui Cathy Chu, Oddur Kjartansson, Clara Rivera, Anna Katanova, Alexander Gutkin, Isin Demirsahin, Cibu Johny, Martin Jansche, Supheakmungkol Sarin, Knot Pipatsrisawat
LREC9
2020 Burmese Speech Corpus, Finite-State Text Normalization and Pronunciation Grammars with an Application to Text-to-Speech
abstract
This paper introduces an open-source crowd-sourced multi-speaker speech corpus along with the comprehensive set of finite-state transducer (FST) grammars for performing text normalization for the Burmese (Myanmar) language. We also introduce the open-source finite-state grammars for performing grapheme-to-phoneme (G2P) conversion for Burmese. These three components are necessary (but not sufficient) for building a high-quality text-to-speech (TTS) system for Burmese, a tonal Southeast Asian language from the Sino-Tibetan family which presents several linguistic challenges. We describe the corpus acquisition process and provide the details of our finite state-based approach to Burmese text normalization and G2P. Our experiments involve building a multi-speaker TTS system based on long short term memory (LSTM) recurrent neural network (RNN) models, which were previously shown to perform well for other languages in a low-resource setting. Our results indicate that the data and grammars that we are announcing are sufficient to build reasonably high-quality models comparable to other systems. We hope these resources will facilitate speech and language research on the Burmese language, which is considered by many to be low-resource due to the limited availability of free linguistic data.
Yin May Oo, Theeraphol Wattanavekin, Chenfang Li, Pasindu De Silva, Supheakmungkol Sarin, Knot Pipatsrisawat, Martin Jansche, Oddur Kjartansson, Alexander Gutkin
LREC7
2019 Sampling from Stochastic Finite Automata with Applications to CTC Decoding
abstract
Stochastic finite automata arise naturally in many language and speech processing tasks. They include stochastic acceptors, which represent certain probability distributions over random strings. We consider the problem of efficient sampling: drawing random string variates from the probability distribution represented by stochastic automata and transformations of those. We show that path-sampling is effective and can be efficient if the epsilon-graph of a finite automaton is acyclic. We provide an algorithm that ensures this by conflating epsilon-cycles within strongly connected components. Sampling is also effective in the presence of non-injective transformations of strings. We illustrate this in the context of decoding for Connectionist Temporal Classification (CTC), where the predictive probabilities yield auxiliary sequences which are transformed into shorter labeling strings. We can sample efficiently from the transformed labeling distribution and use this in two different strategies for finding the most probable CTC labeling.
Martin Jansche, Alexander Gutkin
INTERSPEECH1
2019 Cross-Lingual Consistency of Phonological Features: An Empirical Study
Cibu Johny, Alexander Gutkin, Martin Jansche
INTERSPEECH3
2018 FonBund: A Library for Combining Cross-lingual Phonological Segment Data
Alexander Gutkin, Martin Jansche, Tatiana Merkulova
LREC2
2018 Building Open Javanese and Sundanese Corpora for Multilingual Text-to-Speech
Jaka Aris Eko Wibawa, Supheakmungkol Sarin, Chenfang Li, Knot Pipatsrisawat, Keshan Sodimana, Oddur Kjartansson, Alexander Gutkin, Martin Jansche, Linne Ha
LREC8
2017 Rapid Development of TTS Corpora for Four South African Languages
abstract
This paper describes the development of text-to-speech corpora \nfor four South African languages. The approach followed investigated \nthe possibility of using low-cost methods including informal \nrecording environments and untrained volunteer speakers. \nThis objective and the additional future goal of expanding \nthe corpus to increase coverage of South Africa’s 11 official \nlanguages necessitated experimenting with multi-speaker and \ncode-switched data. The process and relevant observations are \ndetailed throughout. The latest version of the corpora are available \nfor download under an open-source license and will likely \nsee further development and refinement in future. \nIndex Terms: text-to-speech corpus, under-resourced languages
Daniel R. van Niekerk, Charl Johannes van Heerden, Marelie H. Davel, Neil Kleynhans, Oddur Kjartansson, Martin Jansche, Linne Ha
INTERSPEECH6
2016 TTS for Low Resource Languages: A Bangla Synthesizer
Alexander Gutkin, Linne Ha, Martin Jansche, Knot Pipatsrisawat, Richard Sproat
LREC3
2014 Computer-Aided Quality Assurance of an Icelandic Pronunciation Dictionary
Martin Jansche
LREC1
2012 Google's cross-dialect Arabic voice search
abstract
We present a large scale effort to build a commercial Automatic Speech Recognition (ASR) product for Arabic. Our goal is to support voice search, dictation, and voice control for the general Arabic-speaking public, including support for multiple Arabic dialects. We describe our ASR system design and compare recognizers for five Arabic dialects, with the potential to reach more than 125 million people in Egypt, Jordan, Lebanon, Saudi Arabia, and the United Arab Emirates (UAE). We compare systems built on diacritized vs. non-diacritized text. We also conduct cross-dialect experiments, where we train on one dialect and test on the others. Our average word error rate (WER) is 24.8% for voice search.
Fadi Biadsy, Pedro J. Moreno 0001, Martin Jansche
ICASSP3
2011 A Web-Based Tool for Developing Multilingual Pronunciation Lexicons
Samantha Ainsley, Linne Ha, Martin Jansche, Ara Kim, Masayuki Nanzawa
INTERSPEECH3
2011 Deploying Google Search by Voice in Cantonese
abstract
We describe our efforts in deploying Google search by voice for Cantonese, a southern Chinese dialect widely spoken in and around Hong Kong and Guangzhou. We collected audio data from local Cantonese speakers in Hong Kong and Guangzhou by using our DataHound smartphone application. This data was used to create appropriate acoustic models. Language models were trained on anonymized query logs from Google Web Search for Hong Kong. Because users in Hong Kong frequently mix English and Cantonese in their queries, we designed our system from the ground up to handle both languages. We report on experiments with different techniques for mapping the phoneme inventories for both languages into a common space. Based on extensive experiments we report word error rates and web scores for both Hong Kong and Guangzhou data. Cantonese Google search by voice was launched in December 2010. Index Terms: voice search, Cantonese speech recognition, multilingual speech recognition
Yun-Hsuan Sung, Martin Jansche, Pedro J. Moreno 0001
INTERSPEECH2
2010 Reading difficulty in adults with intellectual disabilities: analysis with a hierarchical latent trait model
abstract
In prior work, adults with intellectual disabilities answered comprehension questions after reading texts. We apply a latent trait model to this data to infer the intrinsic difficulty of texts for the participant group. We then analyze the correlation between grade levels predicted by an automatic readability assessment tool and the inferred text difficulty.
Martin Jansche, Lijun Feng, Matt Huenerfauth
ASSETS1
2010 Search by voice in Mandarin Chinese
abstract
In this paper we describe our efforts to build a Mandarin Chinese voice search system. We describe our strategies for data collection, language, lexicon and acoustic modeling, as well as issues related to text normalization that are an integral part of building voice search systems. We show excellent performance on typical spoken search queries under a variety of accents and acoustic conditions. The system has been in operation since October 2009 and has received very positive user reviews.
Jiulong Shan, Genqing Wu, Zhihong Hu, Xiliu Tang, Martin Jansche, Pedro J. Moreno 0001
INTERSPEECH5
2009 WEB-derived pronunciations
abstract
Pronunciation information is available in large quantities on the Web, in the form of IPA and ad-hoc transcriptions. We describe techniques for extracting candidate pronunciations from Web pages and associating them with orthographic words, filtering out poorly extracted pronunciations, normalizing IPA pronunciations to better conform to a common transcription standard, and generating phonemic from ad-hoc transcriptions. We show improvements on a letter-to-phoneme task when using web-derived vs. Pronlex pronunciations.
Arnab Ghoshal, Martin Jansche, Sanjeev Khudanpur, Michael Riley 0001, Morgan Ulinski
ICASSP2
2009 Restoring punctuation and capitalization in transcribed speech
abstract
Adding punctuation and capitalization greatly improves the readability of automatic speech transcripts. We discuss an approach for performing both tasks in a single pass using a purely text-based n-gram language model. We study the effect on performance of varying the n-gram order (from n = 3 to n = 6) and the amount of training data (from 58 million to 55 billion tokens). Our results show that using larger training data sets consistently improves performance, while increasing the n-gram order does not help nearly as much.
Agustín Gravano, Martin Jansche, Michiel Bacchiani
ICASSP2
2009 Web derived pronunciations for spoken term detection
abstract
Indexing and retrieval of speech content in various forms such as broadcast news, customer care data and on-line media has gained a lot of interest for a wide range of applications, from customer analytics to on-line media search. For most retrieval applications, the speech content is typically first converted to a lexical or phonetic representation using automatic speech recognition (ASR). The first step in searching through indexes built on these representations is the generation of pronunciations for named entities and foreign language query terms. This paper summarizes the results of the work conducted during the 2008 JHU Summer Workshop by the Multilingual Spoken Term Detection team, on mining the web for pronunciations and analyzing their impact on spoken term detection. We will first present methods to use the vast amount of pronunciation information available on the Web, in the form of IPA and ad-hoc transcriptions. We describe techniques for extracting candidate pronunciations from Web pages and associating them with orthographic words, filtering out poorly extracted pronunciations, normalizing IPA pronunciations to better conform to a common transcription standard, and generating phonemic representations from ad-hoc transcriptions. We then present an analysis of the effectiveness of using these pronunciations to represent Out-Of-Vocabulary (OOV) query terms on the performance of a spoken term detection (STD) system. We will provide comparisons of Web pronunciations against automated techniques for pronunciation generation as well as pronunciations generated by human experts. Our results cover a range of speech indexes based on lattices, confusion networks and one-best transcriptions at both word and word fragments levels.
Dogan Can, Erica Cooper, Arnab Ghoshal, Martin Jansche, Sanjeev Khudanpur, Bhuvana Ramabhadran, Michael Riley 0001, Murat Saraclar, Abhinav Sethy, Morgan Ulinski, Christopher M. White
SIGIR4
2007 A Maximum Expected Utility Framework for Binary Sequence Labeling
Martin Jansche
ACL1
2007 A Support Vector Approach to Censored Targets
abstract
Censored targets, such as the time to events in survival analysis, can generally be represented by intervals on the real line. In this paper, we propose a novel support vector technique (named SVCR) for regression on censored targets. SVCR inherits the strengths of support vector methods, such as a globally optimal solution by convex programming, fast training speed and strong generalization capacity. In contrast to ranking approaches to survival analysis, our approach is able not only to achieve superior ordering performance, but also to predict the survival time very well. Experiments show a significant performance improvement when the majority of the training data is censored. Experimental results on several survival analysis datasets demonstrate that SVCR is very competitive against classical survival analysis models.
Pannagadatta K. Shivaswamy, Martin Jansche
ICDM3
2003 Parametric Models of Linguistic Count Data
abstract
It is well known that occurrence counts of words in documents are often modeled poorly by standard distributions like the binomial or Poisson. Observed counts vary more than simple models predict, prompting the use of overdispersed models like Gamma-Poisson or Beta-binomial mixtures as robust alternatives. Another deficiency of standard models is due to the fact that most words never occur in a given document, resulting in large amounts of zero counts. We propose using zero-inflated models for dealing with this, and evaluate competing models on a Naive Bayes text classification task. Simple zero-inflated models can account for practically relevant variation, and can be easier to work with than overdispersed models.
Martin Jansche
ACL1
2002 Named Entity Extraction with Conditional Markov Models and Classifiers
Martin Jansche
CoNLL1
2002 Information Extraction from Voicemail Transcripts
abstract
Voicemail is not like email. Even such basic information as the name of the caller/sender or a phone number for returning calls is not represented explicitly and must be obtained from message transcripts or other sources. We discuss techniques for doing this and the challenges these tasks present.
Martin Jansche, Steven Abney
EMNLP1
2001 Information extraction via heuristics for a movie showtime query system
abstract
Semantic interpretation for limited-domain spoken dialogue systems often amounts to extracting information from utterances. For a system that provides movie showtime information, queries are classified along four dimensions: question type, and movie titles, towns and theaters that were mentioned. Simple heuristics suffice for constructing highly accurate classifiers for the latter three attributes; classifiers for the question type attribute are induced from data using features tailored to spoken language phenomena. Since separate classifiers are used for the four attributes, which are not independent, certain errors can be detected and corrected, thus increasing robustness. 1.
Martin Jansche
INTERSPEECH1
2001 Re-Engineering Letter-to-Sound Rules
Martin Jansche
NAACL1
1998 Abductive Reasoning For Syntactic Realization
Ralf Klabunde, Martin Jansche
INLG2