Kenji Sagae

dblp:18/141 · DBLP profile ↗
← Back
45ranked-venue papers
8as first author
4since 2021 · last 2023
0000-0003-3371-0618ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 35 · 7 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Human-computer interaction and ubiquitous computing · 3Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Information extraction and text analysis · 41% Transfer learning and domain adaptation · 13% Language models and text generation · 13%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Bioinformatics and computational biology · 78% Computational social science and digital humanities · 22%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 20 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Transfer learning and domain adaptation
cross-lingual transfer
0.512021
Language Embeddings for Typology and Cross-lingual Transfer Learning · ACL/IJCNLP (1) 2021
Natural language and speech › Language models and text generation › text representation
language embedding
0.512021
Language Embeddings for Typology and Cross-lingual Transfer Learning · ACL/IJCNLP (1) 2021
Natural language and speech › Information extraction and text analysis
linguistic typology
0.512021
Language Embeddings for Typology and Cross-lingual Transfer Learning · ACL/IJCNLP (1) 2021
Natural language and speech › Information extraction and text analysis
syntactic parsing
0.352008
Task-oriented Evaluation of Syntactic Parsers and Their Representations · ACL 2008
HPSG Parsing with Shallow Dependency Constraints · ACL 2007
A Fast, Accurate Deterministic Parser for Chinese · ACL 2006
Natural language and speech › Information extraction and text analysis › syntactic parsing
dependency parsing
0.122007
Dependency Parsing and Domain Adaptation with LR Models and Parser Ensembles · EMNLP-CoNLL 2007
HPSG Parsing with Shallow Dependency Constraints · ACL 2007
Knowledge, reasoning and agents › Knowledge representation and reasoning › commonsense reasoning
causal commonsense reasoning
0.112011
Commonsense Causal Reasoning Using Millions of Personal Stories · AAAI 2011
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning
0.112011
Commonsense Causal Reasoning Using Millions of Personal Stories · AAAI 2011
Compilers and program optimization › parsing
incremental parsing
0.112010
Dynamic Programming for Linear-Time Incremental Parsing · ACL 2010
Compilers and program optimization
parsing
0.112010
Dynamic Programming for Linear-Time Incremental Parsing · ACL 2010
Algorithms and data structures
dynamic programming
0.112010
Dynamic Programming for Linear-Time Incremental Parsing · ACL 2010
Bioinformatics and computational biology
biomedical text mining
0.112009
Evaluating contributions of natural language parsers to protein-protein interaction extraction · Bioinform. 2009
Bioinformatics and computational biology › biomedical text mining › relation extraction
protein-protein interaction extraction
0.112009
Evaluating contributions of natural language parsers to protein-protein interaction extraction · Bioinform. 2009
Natural language and speech › Information extraction and text analysis › syntactic parsing
parser evaluation
0.112008
Task-oriented Evaluation of Syntactic Parsers and Their Representations · ACL 2008
Natural language and speech › Information extraction and text analysis › syntactic parsing › unification-based parsing
HPSG parsing
0.112007
HPSG Parsing with Shallow Dependency Constraints · ACL 2007
Natural language and speech › Information extraction and text analysis › syntactic parsing
parser adaptation
0.112007
Dependency Parsing and Domain Adaptation with LR Models and Parser Ensembles · EMNLP-CoNLL 2007
Natural language and speech › Information extraction and text analysis › syntactic parsing
constituency parsing
0.112006
A Fast, Accurate Deterministic Parser for Chinese · ACL 2006
Natural language and speech › Information extraction and text analysis › syntactic parsing
deterministic parsing
0.112006
A Fast, Accurate Deterministic Parser for Chinese · ACL 2006
Natural language and speech › Information extraction and text analysis › syntactic parsing › transition-based parsing
shift-reduce parsing
0.112006
A Best-First Probabilistic Shift-Reduce Parser · ACL 2006
Natural language and speech › Information extraction and text analysis › syntactic parsing
statistical parsing
0.112006
A Best-First Probabilistic Shift-Reduce Parser · ACL 2006
Computational social science and digital humanities › psycholinguistics
child language acquisition
0.112005
Automatic Measurement of Syntactic Development in Child Language · ACL 2005

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 0.5language embedding · 0.5dynamic programming · 0.2pointwise mutual information · 0.1information retrieval · 0.1natural language parsing · 0.1comparative evaluation · 0.1task-oriented evaluation · 0.1shallow dependency constraints · 0.1parser ensembles · 0.1LR models · 0.1shift-reduce parsing · 0.1best-first search · 0.1statistical parsing · 0.1memory-based learning · 0.1
YearPublicationVenuePosition
2023 Double PP Constituent Ordering Preferences in English Early Child Language
Zoey Liu, Lauren E. Namdar, Stefanie Wulff, Kenji Sagae
CogSci4
2021 Language Embeddings for Typology and Cross-lingual Transfer Learning
abstract
Dian Yu, Taiqi He, Kenji Sagae. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Dian Yu 0002, Taiqi He, Kenji Sagae
ACL/IJCNLP (1)3
2021 Automatically Exposing Problems with Neural Dialog Models
abstract
Neural dialog models are known to suffer from problems such as generating unsafe and inconsistent responses.Even though these problems are crucial and prevalent, they are mostly manually identified by model designers through interactions.Recently, some research instructs crowdworkers to goad the bots into triggering such problems.However, humans leverage superficial clues such as hate speech, while leaving systematic problems undercover.In this paper, we propose two methods including reinforcement learning to automatically trigger a dialog model into generating problematic responses.We show the effect of our methods in exposing safety and contradiction issues with state-of-the-art dialog models.
Dian Yu 0002, Kenji Sagae
EMNLP (1)2
2021 Beyond NVD: Cybersecurity meets the Semantic Web
abstract
Cybersecurity experts rely on the knowledge stored in databases like the NVD to do their work, but these are not the only sources of information about threats and vulnerabilities. Much of that information flows through social media channels. In this paper we argue that security experts and general users alike can benefit from the technologies of the Semantic Web, merging heterogeneous sources of knowledge in an ontological representation. We present a system that has an ontology of vulnerabilities at its core, but that is enhanced with NLP tools to identify cybersecurity-related information in social media and to launch queries over heterogeneous data sources. The transformative power of Semantic Web technologies for cybersecurity, which has been proven in the biomedical field, is evaluated and discussed.
Raúl Aranovich, Muting Wu, Dian Yu 0002, Katya Katsy, Benyamin Ahmadnia, Matthew Bishop, Vladimir Filkov, Kenji Sagae
NSPW8
2020 Developing NLP Tools with a New Corpus of Learner Spanish
abstract
The development of effective NLP tools for the L2 classroom depends largely on the availability of large annotated corpora of language learner text. While annotated learner corpora of English are widely available, large learner corpora of Spanish are less common. Those Spanish corpora that are available do not contain the annotations needed to facilitate the development of tools beneficial to language learners, such as grammatical error correction. As a result, the field has seen little research in NLP tools designed to benefit Spanish language learners and teachers. We introduce COWS-L2H, a freely available corpus of Spanish learner data which includes error annotations and parallel corrected text to help researchers better understand L2 development, to examine teaching practices empirically, and to develop NLP tools to better serve the Spanish teaching community. We demonstrate the utility of this corpus by developing a neural-network based grammatical error correction system for Spanish learner writing.
Sam Davidson, Aaron Yamada, Paloma Fernandez Mira, Agustina Carando, Claudia H. Sánchez Gutiérrez, Kenji Sagae
LREC6
2019 Studying the difference between natural and programming language corpora
Casey Casalnuovo, Kenji Sagae, Premkumar T. Devanbu
Empir. Softw. Eng.2
2018 Language in Context: Incorporating Demographic Embeddings into Language Understanding
Justin Garten, Brendan Kennedy 0001, Joe Hoover, Kenji Sagae, Morteza Dehghani
CogSci4
2016 Supertagging With LSTMs
abstract
Ashish Vaswani, Yonatan Bisk, Kenji Sagae, Ryan Musa. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Ashish Vaswani, Yonatan Bisk, Kenji Sagae, Ryan Musa
HLT-NAACL3
2016 Efficient Structured Inference for Transition-Based Parsing with Neural Networks and Error States
abstract
Transition-based approaches based on local classification are attractive for dependency parsing due to their simplicity and speed, despite producing results slightly below the state-of-the-art. In this paper, we propose a new approach for approximate structured inference for transition-based parsing that produces scores suitable for global scoring using local models. This is accomplished with the introduction of error states in local training, which add information about incorrect derivation paths typically left out completely in locally-trained models. Using neural networks for our local classifiers, our approach achieves 93.61% accuracy for transition-based dependency parsing in English.
Ashish Vaswani, Kenji Sagae
Trans. Assoc. Comput. Linguistics2
2016 Multimodal Analysis and Prediction of Persuasiveness in Online Social Multimedia
Sunghyun Park 0001, Han Suk Shim, Moitreya Chatterjee, Kenji Sagae, Louis-Philippe Morency
ACM Trans. Interact. Intell. Syst.4
2015 Acoustic and para-verbal indicators of persuasiveness in social multimedia
abstract
Persuasive communication and interaction play an important and pervasive role in many aspects of our lives. With the rapid growth of social multimedia websites such as YouTube, it has become more important and useful to understand persuasiveness in the context of online social multimedia content. In this paper, we present our results of conducting various analyses of persuasiveness in speech with our multimedia corpus of 1,000 movie review videos obtained from ExpoTV.com, a popular social multimedia website. Our experiments firstly show that a speaker's level of persuasiveness can be predicted from acoustic characteristics and para-verbal cues related to speech fluency. Secondly, we show that taking acoustic cues in different time periods of a movie review can improve the performance of predicting a speaker's level of persuasiveness. Lastly, we show that a speaker's positive or negative attitude toward a topic influences the prediction performance as well.
Han Suk Shim, Sunghyun Park 0001, Moitreya Chatterjee, Stefan Scherer, Kenji Sagae, Louis-Philippe Morency
ICASSP5
2014 Data-driven Measurement of Child Language Development with Simple Syntactic Templates
Shannon Lubetich, Kenji Sagae
COLING2
2014 Computational Analysis of Persuasiveness in Social Multimedia: A Novel Dataset and Multimodal Prediction Approach
abstract
Our lives are heavily influenced by persuasive communication, and it is essential in almost any types of social interactions from business negotiation to conversation with our friends and family. With the rapid growth of social multimedia websites, it is becoming ever more important and useful to understand persuasiveness in the context of social multimedia content online. In this paper, we introduce our newly created multimedia corpus of 1,000 movie review videos obtained from a social multimedia website called ExpoTV.com, which will be made freely available to the research community. Our research results presented here revolve around the following 3 main research hypotheses. Firstly, we show that computational descriptors derived from verbal and nonverbal behavior can be predictive of persuasiveness. We further show that combining descriptors from multiple communication modalities (audio, text and visual) improve the prediction performance compared to using those from single modality alone. Secondly, we investigate if having prior knowledge of a speaker expressing a positive or negative opinion helps better predict the speaker's persuasiveness. Lastly, we show that it is possible to make comparable prediction of persuasiveness by only looking at thin slices (shorter time windows) of a speaker's behavior.
Sunghyun Park 0001, Han Suk Shim, Moitreya Chatterjee, Kenji Sagae, Louis-Philippe Morency
ICMI4
2014 Towards Learning Nonverbal Identities from the Web: Automatically Identifying Visually Accentuated Words
Amir Zadeh 0001, Kenji Sagae, Louis-Philippe Morency
IVA2
2014 Improving Classification-Based Natural Language Understanding with Non-Expert Annotation
abstract
Although data-driven techniques are com-monly used for Natural Language Under-standing in dialogue systems, their effi-cacy is often hampered by the lack of ap-propriate annotated training data in suffi-cient amounts. We present an approach for rapid and cost-effective annotation of training data for classification-based lan-guage understanding in conversational di-alogue systems. Experiments using a web-accessible conversational character that in-teracts with a varied user population show that a dramatic improvement in natural language understanding and a substantial reduction in expert annotation effort can be achieved by leveraging non-expert an-notation.
Fabrizio Morbini, Eric Forbell, Kenji Sagae
SIGDIAL Conference3
2013 Who is persuasive?: the role of perceived personality and communication modality in social multimedia
abstract
Persuasive communication is part of everyone's daily life. With the emergence of social websites like YouTube, Facebook and Twitter, persuasive communication is now seen online on a daily basis. This paper explores the effect of multi-modality and perceived personality on persuasiveness of social multimedia content. The experiments are performed over a large corpus of movie review clips from Youtube which is presented to online annotators in three different modalities: only text, only audio and video. The annotators evaluated the persuasiveness of each review across different modalities and judged the personality of the speaker. Our detailed analysis confirmed several research hypotheses designed to study the relationships between persuasion, perceived personality and communicative channel, namely modality. Three hypotheses are designed: the first hypothesis studies the effect of communication modality on persuasion, the second hypothesis examines the correlation between persuasion and personality perception and finally the third hypothesis, derived from the first two hypotheses explores how communication modality influence the personality perception.
Gelareh Mohammadi, Sunghyun Park 0001, Kenji Sagae, Alessandro Vinciarelli, Louis-Philippe Morency
ICMI3
2013 Roundtable: An Online Framework for Building Web-based Conversational Agents
Eric Forbell, Nicolai Kalisch, Fabrizio Morbini, Kelly Christoffersen, Kenji Sagae, David R. Traum, Albert A. Rizzo
SIGDIAL Conference5
2013 Which ASR should I choose for my dialogue system?
Fabrizio Morbini, Kartik Audhkhasi, Kenji Sagae, Ron Artstein, Dogan Can, Panayiotis G. Georgiou, Shri Narayanan, Anton Leuski, David R. Traum
SIGDIAL Conference3
2012 Semi-supervised discriminative language modeling for Turkish ASR
abstract
We present our work on semi-supervised learning of discriminative language models where the negative examples for sentences in a text corpus are generated using confusion models for Turkish at various granularities, specifically, word, sub-word, syllable and phone levels. We experiment with different language models and various sampling strategies to select competing hypotheses for training with a variant of the perceptron algorithm. We find that morph-based confusion models with a sample selection strategy aiming to match the error distribution of the baseline ASR system gives the best performance. We also observe that substituting half of the supervised training examples with those obtained in a semi-supervised manner gives similar results.
Arda Çelebi, Hasim Sak, Erinç Dikici, Murat Saraclar, Maider Lehr, Emily Tucker Prud'hommeaux, Puyang Xu, Nathan Glenn, Damianos Karakos, Sanjeev Khudanpur, Brian Roark, Kenji Sagae, Izhak Shafran, Dan Bikel, Chris Callison-Burch, Yuan Cao 0007, Keith B. Hall, Eva Hasler, Philipp Koehn, Adam Lopez, Matt Post, Darcey Riley
ICASSP12
2012 Hallucinated n-best lists for discriminative language modeling
abstract
This paper investigates semi-supervised methods for discriminative language modeling, whereby n-best lists are “hallucinated” for given reference text and are then used for training n-gram language models using the perceptron algorithm. We perform controlled experiments on a very strong baseline English CTS system, comparing three methods for simulating ASR output, and compare the results with training with “real” n-best list output from the baseline recognizer. We find that methods based on extracting phrasal cohorts - similar to methods from machine translation for extracting phrase tables - yielded the largest gains of our three methods, achieving over half of the WER reduction of the fully supervised methods.
Kenji Sagae, Maider Lehr, Emily Tucker Prud'hommeaux, Puyang Xu, Nathan Glenn, Damianos Karakos, Sanjeev Khudanpur, Brian Roark, Murat Saraclar, Izhak Shafran, Dan Bikel, Chris Callison-Burch, Yuan Cao 0007, Keith B. Hall, Eva Hasler, Philipp Koehn, Adam Lopez, Matt Post, Darcey Riley
ICASSP1
2012 Continuous space discriminative language modeling
abstract
Discriminative language modeling is a structured classification problem. Log-linear models have been previously used to address this problem. In this paper, the standard dot-product feature representation used in log-linear models is replaced by a non-linear function parameterized by a neural network. Embeddings are learned for each word and features are extracted automatically through the use of convolutional layers. Experimental results show that as a stand-alone model the continuous space model yields significantly lower word error rate (1% absolute), while having a much more compact parameterization (60%-90% smaller). If the baseline scores are combined, our approach performs equally well.
Puyang Xu, Sanjeev Khudanpur, Maider Lehr, Emily Tucker Prud'hommeaux, Nathan Glenn, Damianos Karakos, Brian Roark, Kenji Sagae, Murat Saraclar, Izhak Shafran, Dan Bikel, Chris Callison-Burch, Yuan Cao 0007, Keith B. Hall, Eva Hasler, Philipp Koehn, Adam Lopez, Matt Post, Darcey Riley
ICASSP8
2012 Deriving conversation-based features from unlabeled speech for discriminative language modeling
Damianos Karakos, Brian Roark, Izhak Shafran, Kenji Sagae, Maider Lehr, Emily Tucker Prud'hommeaux, Puyang Xu, Nathan Glenn, Sanjeev Khudanpur, Murat Saraclar, Dan Bikel, Mark Dredze, Chris Callison-Burch, Yuan Cao 0007, Keith B. Hall, Eva Hasler, Philipp Koehn, Adam Lopez, Matt Post, Darcey Riley
INTERSPEECH4
2012 Practical Evaluation of Human and Synthesized Speech for Virtual Human Dialogue Systems
Kallirroi Georgila, Alan W. Black, Kenji Sagae, David R. Traum
LREC3
2012 A Mixed-Initiative Conversational Dialogue System for Healthcare
Fabrizio Morbini, Eric Forbell, David DeVault, Kenji Sagae, David R. Traum, Albert A. Rizzo
SIGDIAL Conference4
2012 A reranking approach for recognition and classification of speech input in conversational dialogue systems
abstract
We address the challenge of interpreting spoken input in a conversational dialogue system with an approach that aims to exploit the close relationship between the tasks of speech recognition and language understanding through joint modeling of these two tasks. Instead of using a standard pipeline approach where the output of a speech recognizer is the input of a language understanding module, we merge multiple speech recognition and utterance classification hypotheses into one list to be processed by a joint reranking model. We obtain substantially improved performance in language understanding in experiments with thousands of user utterances collected from a deployed spoken dialogue system.
Fabrizio Morbini, Kartik Audhkhasi, Ron Artstein, Maarten Van Segbroeck, Kenji Sagae, Panayiotis G. Georgiou, David R. Traum, Shri Narayanan
SLT5
2011 Commonsense Causal Reasoning Using Millions of Personal Stories
abstract
The personal stories that people write in their Internet weblogs include a substantial amount of information about the causal relationships between everyday events. In this paper we describe our efforts to use millions of these stories for automated commonsense causal reasoning. Casting the commonsense causal reasoning problem as a Choice of Plausible Alternatives, we describe four experiments that compare various statistical and information retrieval approaches to exploit causal information in story corpora. The top performing system in these experiments uses a simple co-occurrence statistic between words in the causal antecedent and consequent, calculated as the Pointwise Mutual Information between words in a corpus of millions of personal stories.
Andrew S. Gordon, Cosmin Adrian Bejan, Kenji Sagae
AAAI3
2011 Analyzing Conservative and Liberal Blogs Related to the Construction of the 'Ground Zero Mosque'
Morteza Dehghani, Jonathan Gratch, Sonya Sachdeva, Kenji Sagae
CogSci4
2011 An Evaluation of Alternative Strategies for Implementing Dialogue Policies Using Statistical Classification and Hand-Authored Rules
David DeVault, Anton Leuski, Kenji Sagae
IJCNLP3
2011 Detecting the Status of a Predictive Incremental Speech Understanding Model for Real-Time Decision-Making in a Spoken Dialogue System
abstract
We explore the potential for a responsive spoken dialogue system to use the real-time status of an incremental speech understanding model to guide its incremental decision-making about how to respond to a user utterance that is still in progress. Spoken dialogue systems have a range of potentially useful realtime response options as a user is speaking, such as providing acknowledgments or backchannels, interrupting the user to ask a clarification question or to initiate the system’s response, or even completing the user’s utterance at appropriate moments. However, implementing such incremental response capabilities seems to require that a system be able to assess its own level of understanding incrementally, so that an appropriate response can be selected at each moment. In this paper, we use a datadriven classification approach to explore the trade-offs that a virtual human dialogue system faces in reliably identifying how its understanding is progressing during a user utterance.
David DeVault, Kenji Sagae, David R. Traum
INTERSPEECH2
2011 Toward Learning and Evaluation of Dialogue Policies with Text Examples
David DeVault, Anton Leuski, Kenji Sagae
SIGDIAL Conference3
2010 Dynamic Programming for Linear-Time Incremental Parsing
Liang Huang 0001, Kenji Sagae
ACL2
2010 Latent Mixture of Discriminative Experts for Multimodal Prediction Modeling
Derya Ozkan, Kenji Sagae, Louis-Philippe Morency
COLING2
2010 Practical Evaluation of Speech Recognizers for Virtual Human Dialogue Systems
Xuchen Yao, Pravin Bhutada, Kallirroi Georgila, Kenji Sagae, Ron Artstein, David R. Traum
LREC4
2009 Can I Finish? Learning When to Respond to Incremental Interpretation Results in Interactive Dialogue
David DeVault, Kenji Sagae, David R. Traum
SIGDIAL Conference2
2009 Evaluating contributions of natural language parsers to protein-protein interaction extraction
abstract
MOTIVATION: While text mining technologies for biomedical research have gained popularity as a way to take advantage of the explosive growth of information in text form in biomedical papers, selecting appropriate natural language processing (NLP) tools is still difficult for researchers who are not familiar with recent advances in NLP. This article provides a comparative evaluation of several state-of-the-art natural language parsers, focusing on the task of extracting protein-protein interaction (PPI) from biomedical papers. We measure how each parser, and its output representation, contributes to accuracy improvement when the parser is used as a component in a PPI system. RESULTS: All the parsers attained improvements in accuracy of PPI extraction. The levels of accuracy obtained with these different parsers vary slightly, while differences in parsing speed are larger. The best accuracy in this work was obtained when we combined Miyao and Tsujii's Enju parser and Charniak and Johnson's reranking parser, and the accuracy is better than the state-of-the-art results on the same data. AVAILABILITY: The PPI extraction system used in this work (AkanePPI) is available online at http://www-tsujii.is.s.u-tokyo.ac.jp/downloads/downloads.cgi. The evaluated parsers are also available online from each developer's site.
Yusuke Miyao, Kenji Sagae, Rune Sætre, Takuya Matsuzaki, Jun'ichi Tsujii
Bioinform.2
2008 Task-oriented Evaluation of Syntactic Parsers and Their Representations
Yusuke Miyao, Rune Sætre, Kenji Sagae, Takuya Matsuzaki, Jun'ichi Tsujii
ACL3
2008 Shift-Reduce Dependency DAG Parsing
Kenji Sagae, Jun'ichi Tsujii
COLING1
2008 GENIA-GR: a Grammatical Relation Corpus for Parser Evaluation in the Biomedical Domain
Yuka Tateisi, Yusuke Miyao, Kenji Sagae, Jun'ichi Tsujii
LREC3
2007 HPSG Parsing with Shallow Dependency Constraints
Kenji Sagae, Yusuke Miyao, Jun'ichi Tsujii
ACL1
2007 Dependency Parsing and Domain Adaptation with LR Models and Parser Ensembles
Kenji Sagae, Jun'ichi Tsujii
EMNLP-CoNLL1
2006 A Best-First Probabilistic Shift-Reduce Parser
Kenji Sagae, Alon Lavie
ACL1
2006 A Fast, Accurate Deterministic Parser for Chinese
abstract
We present a novel classifier-based deterministic parser for Chinese constituency parsing. Our parser computes parse trees from bottom up in one pass, and uses classifiers to make shift-reduce decisions. Trained and evaluated on the standard training and test sets, our best model (using stacked classifiers) runs in linear time and has labeled precision and recall above 88% using gold-standard part-of-speech tags, surpassing the best published results. Our SVM parser is 2-13 times faster than state-of-the-art parsers, while producing more accurate results. Our Maxent and DTree parsers run at speeds 40-270 times faster than state-of-the-art parsers, but with 5-6% losses in accuracy.
Mengqiu Wang, Kenji Sagae, Teruko Mitamura
ACL2
2006 Parser Combination by Reparsing
Kenji Sagae, Alon Lavie
HLT-NAACL1
2005 Automatic Measurement of Syntactic Development in Child Language
abstract
To facilitate the use of syntactic information in the study of child language acquisition, a coding scheme for Grammatical Relations (GRs) in transcripts of parent-child dialogs has been proposed by Sagae, MacWhinney and Lavie (2004). We discuss the use of current NLP techniques to produce the GRs in this annotation scheme. By using a statistical parser (Charniak, 2000) and memory-based learning tools for classification (Daelemans et al., 2004), we obtain high precision and recall of several GRs. We demonstrate the usefulness of this approach by performing automatic measurements of syntactic development with the Index of Productive Syntax (Scarborough, 1990) at similar levels to what child language researchers compute manually.
Kenji Sagae, Alon Lavie, Brian MacWhinney
ACL1
2004 Adding Syntactic Annotations to Transcripts of Parent-Child Dialogs
Kenji Sagae, Brian MacWhinney, Alon Lavie
LREC1