VLDB 2026 Research / reviewers in the wild / expert
Kenji Sagae
dblp:18/141
· DBLP profile ↗
45ranked-venue papers
8as first author
4since 2021 · last 2023
0000-0003-3371-0618ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 35 · 7 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Human-computer interaction and ubiquitous computing · 3Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Information extraction and text analysis · 41% Transfer learning and domain adaptation · 13% Language models and text generation · 13% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 78% Computational social science and digital humanities · 22% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% |
Topics — the 20 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Transfer learning and domain adaptation
cross-lingual transfer |
0.5 | 1 | 2021 | Language Embeddings for Typology and Cross-lingual Transfer Learning · ACL/IJCNLP (1) 2021 |
Natural language and speech › Language models and text generation › text representation
language embedding |
0.5 | 1 | 2021 | Language Embeddings for Typology and Cross-lingual Transfer Learning · ACL/IJCNLP (1) 2021 |
Natural language and speech › Information extraction and text analysis
linguistic typology |
0.5 | 1 | 2021 | Language Embeddings for Typology and Cross-lingual Transfer Learning · ACL/IJCNLP (1) 2021 |
Natural language and speech › Information extraction and text analysis
syntactic parsing |
0.3 | 5 | 2008 | Task-oriented Evaluation of Syntactic Parsers and Their Representations · ACL 2008 HPSG Parsing with Shallow Dependency Constraints · ACL 2007 A Fast, Accurate Deterministic Parser for Chinese · ACL 2006 |
Natural language and speech › Information extraction and text analysis › syntactic parsing
dependency parsing |
0.1 | 2 | 2007 | Dependency Parsing and Domain Adaptation with LR Models and Parser Ensembles · EMNLP-CoNLL 2007 HPSG Parsing with Shallow Dependency Constraints · ACL 2007 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › commonsense reasoning
causal commonsense reasoning |
0.1 | 1 | 2011 | Commonsense Causal Reasoning Using Millions of Personal Stories · AAAI 2011 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning |
0.1 | 1 | 2011 | Commonsense Causal Reasoning Using Millions of Personal Stories · AAAI 2011 |
Compilers and program optimization › parsing
incremental parsing |
0.1 | 1 | 2010 | Dynamic Programming for Linear-Time Incremental Parsing · ACL 2010 |
Compilers and program optimization
parsing |
0.1 | 1 | 2010 | Dynamic Programming for Linear-Time Incremental Parsing · ACL 2010 |
Algorithms and data structures
dynamic programming |
0.1 | 1 | 2010 | Dynamic Programming for Linear-Time Incremental Parsing · ACL 2010 |
Bioinformatics and computational biology
biomedical text mining |
0.1 | 1 | 2009 | Evaluating contributions of natural language parsers to protein-protein interaction extraction · Bioinform. 2009 |
Bioinformatics and computational biology › biomedical text mining › relation extraction
protein-protein interaction extraction |
0.1 | 1 | 2009 | Evaluating contributions of natural language parsers to protein-protein interaction extraction · Bioinform. 2009 |
Natural language and speech › Information extraction and text analysis › syntactic parsing
parser evaluation |
0.1 | 1 | 2008 | Task-oriented Evaluation of Syntactic Parsers and Their Representations · ACL 2008 |
Natural language and speech › Information extraction and text analysis › syntactic parsing › unification-based parsing
HPSG parsing |
0.1 | 1 | 2007 | HPSG Parsing with Shallow Dependency Constraints · ACL 2007 |
Natural language and speech › Information extraction and text analysis › syntactic parsing
parser adaptation |
0.1 | 1 | 2007 | Dependency Parsing and Domain Adaptation with LR Models and Parser Ensembles · EMNLP-CoNLL 2007 |
Natural language and speech › Information extraction and text analysis › syntactic parsing
constituency parsing |
0.1 | 1 | 2006 | A Fast, Accurate Deterministic Parser for Chinese · ACL 2006 |
Natural language and speech › Information extraction and text analysis › syntactic parsing
deterministic parsing |
0.1 | 1 | 2006 | A Fast, Accurate Deterministic Parser for Chinese · ACL 2006 |
Natural language and speech › Information extraction and text analysis › syntactic parsing › transition-based parsing
shift-reduce parsing |
0.1 | 1 | 2006 | A Best-First Probabilistic Shift-Reduce Parser · ACL 2006 |
Natural language and speech › Information extraction and text analysis › syntactic parsing
statistical parsing |
0.1 | 1 | 2006 | A Best-First Probabilistic Shift-Reduce Parser · ACL 2006 |
Computational social science and digital humanities › psycholinguistics
child language acquisition |
0.1 | 1 | 2005 | Automatic Measurement of Syntactic Development in Child Language · ACL 2005 |
Methods — techniques the papers use, named apart from their topics
reinforcement learning · 0.5language embedding · 0.5dynamic programming · 0.2pointwise mutual information · 0.1information retrieval · 0.1natural language parsing · 0.1comparative evaluation · 0.1task-oriented evaluation · 0.1shallow dependency constraints · 0.1parser ensembles · 0.1LR models · 0.1shift-reduce parsing · 0.1best-first search · 0.1statistical parsing · 0.1memory-based learning · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Double PP Constituent Ordering Preferences in English Early Child Language
Zoey Liu, Lauren E. Namdar, Stefanie Wulff, Kenji Sagae |
CogSci | 4 |
| 2021 | Language Embeddings for Typology and Cross-lingual Transfer LearningabstractDian Yu, Taiqi He, Kenji Sagae. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Dian Yu 0002, Taiqi He, Kenji Sagae |
ACL/IJCNLP (1) | 3 |
| 2021 | Automatically Exposing Problems with Neural Dialog ModelsabstractNeural dialog models are known to suffer from problems such as generating unsafe and inconsistent responses.Even though these problems are crucial and prevalent, they are mostly manually identified by model designers through interactions.Recently, some research instructs crowdworkers to goad the bots into triggering such problems.However, humans leverage superficial clues such as hate speech, while leaving systematic problems undercover.In this paper, we propose two methods including reinforcement learning to automatically trigger a dialog model into generating problematic responses.We show the effect of our methods in exposing safety and contradiction issues with state-of-the-art dialog models. Dian Yu 0002, Kenji Sagae |
EMNLP (1) | 2 |
| 2021 | Beyond NVD: Cybersecurity meets the Semantic WebabstractCybersecurity experts rely on the knowledge stored in databases like the NVD to do their work, but these are not the only sources of information about threats and vulnerabilities. Much of that information flows through social media channels. In this paper we argue that security experts and general users alike can benefit from the technologies of the Semantic Web, merging heterogeneous sources of knowledge in an ontological representation. We present a system that has an ontology of vulnerabilities at its core, but that is enhanced with NLP tools to identify cybersecurity-related information in social media and to launch queries over heterogeneous data sources. The transformative power of Semantic Web technologies for cybersecurity, which has been proven in the biomedical field, is evaluated and discussed. Raúl Aranovich, Muting Wu, Dian Yu 0002, Katya Katsy, Benyamin Ahmadnia, Matthew Bishop, Vladimir Filkov, Kenji Sagae |
NSPW | 8 |
| 2020 | Developing NLP Tools with a New Corpus of Learner SpanishabstractThe development of effective NLP tools for the L2 classroom depends largely on the availability of large annotated corpora of language learner text. While annotated learner corpora of English are widely available, large learner corpora of Spanish are less common. Those Spanish corpora that are available do not contain the annotations needed to facilitate the development of tools beneficial to language learners, such as grammatical error correction. As a result, the field has seen little research in NLP tools designed to benefit Spanish language learners and teachers. We introduce COWS-L2H, a freely available corpus of Spanish learner data which includes error annotations and parallel corrected text to help researchers better understand L2 development, to examine teaching practices empirically, and to develop NLP tools to better serve the Spanish teaching community. We demonstrate the utility of this corpus by developing a neural-network based grammatical error correction system for Spanish learner writing. Sam Davidson, Aaron Yamada, Paloma Fernandez Mira, Agustina Carando, Claudia H. Sánchez Gutiérrez, Kenji Sagae |
LREC | 6 |
| 2019 | Studying the difference between natural and programming language corpora
Casey Casalnuovo, Kenji Sagae, Premkumar T. Devanbu |
Empir. Softw. Eng. | 2 |
| 2018 | Language in Context: Incorporating Demographic Embeddings into Language Understanding
Justin Garten, Brendan Kennedy 0001, Joe Hoover, Kenji Sagae, Morteza Dehghani |
CogSci | 4 |
| 2016 | Supertagging With LSTMsabstractAshish Vaswani, Yonatan Bisk, Kenji Sagae, Ryan Musa. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Ashish Vaswani, Yonatan Bisk, Kenji Sagae, Ryan Musa |
HLT-NAACL | 3 |
| 2016 | Efficient Structured Inference for Transition-Based Parsing with Neural Networks and Error StatesabstractTransition-based approaches based on local classification are attractive for dependency parsing due to their simplicity and speed, despite producing results slightly below the state-of-the-art. In this paper, we propose a new approach for approximate structured inference for transition-based parsing that produces scores suitable for global scoring using local models. This is accomplished with the introduction of error states in local training, which add information about incorrect derivation paths typically left out completely in locally-trained models. Using neural networks for our local classifiers, our approach achieves 93.61% accuracy for transition-based dependency parsing in English. Ashish Vaswani, Kenji Sagae |
Trans. Assoc. Comput. Linguistics | 2 |
| 2016 | Multimodal Analysis and Prediction of Persuasiveness in Online Social Multimedia
Sunghyun Park 0001, Han Suk Shim, Moitreya Chatterjee, Kenji Sagae, Louis-Philippe Morency |
ACM Trans. Interact. Intell. Syst. | 4 |
| 2015 | Acoustic and para-verbal indicators of persuasiveness in social multimediaabstractPersuasive communication and interaction play an important and pervasive role in many aspects of our lives. With the rapid growth of social multimedia websites such as YouTube, it has become more important and useful to understand persuasiveness in the context of online social multimedia content. In this paper, we present our results of conducting various analyses of persuasiveness in speech with our multimedia corpus of 1,000 movie review videos obtained from ExpoTV.com, a popular social multimedia website. Our experiments firstly show that a speaker's level of persuasiveness can be predicted from acoustic characteristics and para-verbal cues related to speech fluency. Secondly, we show that taking acoustic cues in different time periods of a movie review can improve the performance of predicting a speaker's level of persuasiveness. Lastly, we show that a speaker's positive or negative attitude toward a topic influences the prediction performance as well. Han Suk Shim, Sunghyun Park 0001, Moitreya Chatterjee, Stefan Scherer, Kenji Sagae, Louis-Philippe Morency |
ICASSP | 5 |
| 2014 | Data-driven Measurement of Child Language Development with Simple Syntactic Templates
Shannon Lubetich, Kenji Sagae |
COLING | 2 |
| 2014 | Computational Analysis of Persuasiveness in Social Multimedia: A Novel Dataset and Multimodal Prediction ApproachabstractOur lives are heavily influenced by persuasive communication, and it is essential in almost any types of social interactions from business negotiation to conversation with our friends and family. With the rapid growth of social multimedia websites, it is becoming ever more important and useful to understand persuasiveness in the context of social multimedia content online. In this paper, we introduce our newly created multimedia corpus of 1,000 movie review videos obtained from a social multimedia website called ExpoTV.com, which will be made freely available to the research community. Our research results presented here revolve around the following 3 main research hypotheses. Firstly, we show that computational descriptors derived from verbal and nonverbal behavior can be predictive of persuasiveness. We further show that combining descriptors from multiple communication modalities (audio, text and visual) improve the prediction performance compared to using those from single modality alone. Secondly, we investigate if having prior knowledge of a speaker expressing a positive or negative opinion helps better predict the speaker's persuasiveness. Lastly, we show that it is possible to make comparable prediction of persuasiveness by only looking at thin slices (shorter time windows) of a speaker's behavior. Sunghyun Park 0001, Han Suk Shim, Moitreya Chatterjee, Kenji Sagae, Louis-Philippe Morency |
ICMI | 4 |
| 2014 | Towards Learning Nonverbal Identities from the Web: Automatically Identifying Visually Accentuated Words
Amir Zadeh 0001, Kenji Sagae, Louis-Philippe Morency |
IVA | 2 |
| 2014 | Improving Classification-Based Natural Language Understanding with Non-Expert AnnotationabstractAlthough data-driven techniques are com-monly used for Natural Language Under-standing in dialogue systems, their effi-cacy is often hampered by the lack of ap-propriate annotated training data in suffi-cient amounts. We present an approach for rapid and cost-effective annotation of training data for classification-based lan-guage understanding in conversational di-alogue systems. Experiments using a web-accessible conversational character that in-teracts with a varied user population show that a dramatic improvement in natural language understanding and a substantial reduction in expert annotation effort can be achieved by leveraging non-expert an-notation. Fabrizio Morbini, Eric Forbell, Kenji Sagae |
SIGDIAL Conference | 3 |
| 2013 | Who is persuasive?: the role of perceived personality and communication modality in social multimediaabstractPersuasive communication is part of everyone's daily life. With the emergence of social websites like YouTube, Facebook and Twitter, persuasive communication is now seen online on a daily basis. This paper explores the effect of multi-modality and perceived personality on persuasiveness of social multimedia content. The experiments are performed over a large corpus of movie review clips from Youtube which is presented to online annotators in three different modalities: only text, only audio and video. The annotators evaluated the persuasiveness of each review across different modalities and judged the personality of the speaker. Our detailed analysis confirmed several research hypotheses designed to study the relationships between persuasion, perceived personality and communicative channel, namely modality. Three hypotheses are designed: the first hypothesis studies the effect of communication modality on persuasion, the second hypothesis examines the correlation between persuasion and personality perception and finally the third hypothesis, derived from the first two hypotheses explores how communication modality influence the personality perception. Gelareh Mohammadi, Sunghyun Park 0001, Kenji Sagae, Alessandro Vinciarelli, Louis-Philippe Morency |
ICMI | 3 |
| 2013 | Roundtable: An Online Framework for Building Web-based Conversational Agents
Eric Forbell, Nicolai Kalisch, Fabrizio Morbini, Kelly Christoffersen, Kenji Sagae, David R. Traum, Albert A. Rizzo |
SIGDIAL Conference | 5 |
| 2013 | Which ASR should I choose for my dialogue system?
Fabrizio Morbini, Kartik Audhkhasi, Kenji Sagae, Ron Artstein, Dogan Can, Panayiotis G. Georgiou, Shri Narayanan, Anton Leuski, David R. Traum |
SIGDIAL Conference | 3 |
| 2012 | Semi-supervised discriminative language modeling for Turkish ASRabstractWe present our work on semi-supervised learning of discriminative language models where the negative examples for sentences in a text corpus are generated using confusion models for Turkish at various granularities, specifically, word, sub-word, syllable and phone levels. We experiment with different language models and various sampling strategies to select competing hypotheses for training with a variant of the perceptron algorithm. We find that morph-based confusion models with a sample selection strategy aiming to match the error distribution of the baseline ASR system gives the best performance. We also observe that substituting half of the supervised training examples with those obtained in a semi-supervised manner gives similar results. Arda Çelebi, Hasim Sak, Erinç Dikici, Murat Saraclar, Maider Lehr, Emily Tucker Prud'hommeaux, Puyang Xu, Nathan Glenn, Damianos Karakos, Sanjeev Khudanpur, Brian Roark, Kenji Sagae, Izhak Shafran, Dan Bikel, Chris Callison-Burch, Yuan Cao 0007, Keith B. Hall, Eva Hasler, Philipp Koehn, Adam Lopez, Matt Post, Darcey Riley |
ICASSP | 12 |
| 2012 | Hallucinated n-best lists for discriminative language modelingabstractThis paper investigates semi-supervised methods for discriminative language modeling, whereby n-best lists are “hallucinated” for given reference text and are then used for training n-gram language models using the perceptron algorithm. We perform controlled experiments on a very strong baseline English CTS system, comparing three methods for simulating ASR output, and compare the results with training with “real” n-best list output from the baseline recognizer. We find that methods based on extracting phrasal cohorts - similar to methods from machine translation for extracting phrase tables - yielded the largest gains of our three methods, achieving over half of the WER reduction of the fully supervised methods. Kenji Sagae, Maider Lehr, Emily Tucker Prud'hommeaux, Puyang Xu, Nathan Glenn, Damianos Karakos, Sanjeev Khudanpur, Brian Roark, Murat Saraclar, Izhak Shafran, Dan Bikel, Chris Callison-Burch, Yuan Cao 0007, Keith B. Hall, Eva Hasler, Philipp Koehn, Adam Lopez, Matt Post, Darcey Riley |
ICASSP | 1 |
| 2012 | Continuous space discriminative language modelingabstractDiscriminative language modeling is a structured classification problem. Log-linear models have been previously used to address this problem. In this paper, the standard dot-product feature representation used in log-linear models is replaced by a non-linear function parameterized by a neural network. Embeddings are learned for each word and features are extracted automatically through the use of convolutional layers. Experimental results show that as a stand-alone model the continuous space model yields significantly lower word error rate (1% absolute), while having a much more compact parameterization (60%-90% smaller). If the baseline scores are combined, our approach performs equally well. Puyang Xu, Sanjeev Khudanpur, Maider Lehr, Emily Tucker Prud'hommeaux, Nathan Glenn, Damianos Karakos, Brian Roark, Kenji Sagae, Murat Saraclar, Izhak Shafran, Dan Bikel, Chris Callison-Burch, Yuan Cao 0007, Keith B. Hall, Eva Hasler, Philipp Koehn, Adam Lopez, Matt Post, Darcey Riley |
ICASSP | 8 |
| 2012 | Deriving conversation-based features from unlabeled speech for discriminative language modeling
Damianos Karakos, Brian Roark, Izhak Shafran, Kenji Sagae, Maider Lehr, Emily Tucker Prud'hommeaux, Puyang Xu, Nathan Glenn, Sanjeev Khudanpur, Murat Saraclar, Dan Bikel, Mark Dredze, Chris Callison-Burch, Yuan Cao 0007, Keith B. Hall, Eva Hasler, Philipp Koehn, Adam Lopez, Matt Post, Darcey Riley |
INTERSPEECH | 4 |
| 2012 | Practical Evaluation of Human and Synthesized Speech for Virtual Human Dialogue Systems
Kallirroi Georgila, Alan W. Black, Kenji Sagae, David R. Traum |
LREC | 3 |
| 2012 | A Mixed-Initiative Conversational Dialogue System for Healthcare
Fabrizio Morbini, Eric Forbell, David DeVault, Kenji Sagae, David R. Traum, Albert A. Rizzo |
SIGDIAL Conference | 4 |
| 2012 | A reranking approach for recognition and classification of speech input in conversational dialogue systemsabstractWe address the challenge of interpreting spoken input in a conversational dialogue system with an approach that aims to exploit the close relationship between the tasks of speech recognition and language understanding through joint modeling of these two tasks. Instead of using a standard pipeline approach where the output of a speech recognizer is the input of a language understanding module, we merge multiple speech recognition and utterance classification hypotheses into one list to be processed by a joint reranking model. We obtain substantially improved performance in language understanding in experiments with thousands of user utterances collected from a deployed spoken dialogue system. Fabrizio Morbini, Kartik Audhkhasi, Ron Artstein, Maarten Van Segbroeck, Kenji Sagae, Panayiotis G. Georgiou, David R. Traum, Shri Narayanan |
SLT | 5 |
| 2011 | Commonsense Causal Reasoning Using Millions of Personal StoriesabstractThe personal stories that people write in their Internet weblogs include a substantial amount of information about the causal relationships between everyday events. In this paper we describe our efforts to use millions of these stories for automated commonsense causal reasoning. Casting the commonsense causal reasoning problem as a Choice of Plausible Alternatives, we describe four experiments that compare various statistical and information retrieval approaches to exploit causal information in story corpora. The top performing system in these experiments uses a simple co-occurrence statistic between words in the causal antecedent and consequent, calculated as the Pointwise Mutual Information between words in a corpus of millions of personal stories. Andrew S. Gordon, Cosmin Adrian Bejan, Kenji Sagae |
AAAI | 3 |
| 2011 | Analyzing Conservative and Liberal Blogs Related to the Construction of the 'Ground Zero Mosque'
Morteza Dehghani, Jonathan Gratch, Sonya Sachdeva, Kenji Sagae |
CogSci | 4 |
| 2011 | An Evaluation of Alternative Strategies for Implementing Dialogue Policies Using Statistical Classification and Hand-Authored Rules
David DeVault, Anton Leuski, Kenji Sagae |
IJCNLP | 3 |
| 2011 | Detecting the Status of a Predictive Incremental Speech Understanding Model for Real-Time Decision-Making in a Spoken Dialogue SystemabstractWe explore the potential for a responsive spoken dialogue system to use the real-time status of an incremental speech understanding model to guide its incremental decision-making about how to respond to a user utterance that is still in progress. Spoken dialogue systems have a range of potentially useful realtime response options as a user is speaking, such as providing acknowledgments or backchannels, interrupting the user to ask a clarification question or to initiate the system’s response, or even completing the user’s utterance at appropriate moments. However, implementing such incremental response capabilities seems to require that a system be able to assess its own level of understanding incrementally, so that an appropriate response can be selected at each moment. In this paper, we use a datadriven classification approach to explore the trade-offs that a virtual human dialogue system faces in reliably identifying how its understanding is progressing during a user utterance. David DeVault, Kenji Sagae, David R. Traum |
INTERSPEECH | 2 |
| 2011 | Toward Learning and Evaluation of Dialogue Policies with Text Examples
David DeVault, Anton Leuski, Kenji Sagae |
SIGDIAL Conference | 3 |
| 2010 | Dynamic Programming for Linear-Time Incremental Parsing
Liang Huang 0001, Kenji Sagae |
ACL | 2 |
| 2010 | Latent Mixture of Discriminative Experts for Multimodal Prediction Modeling
Derya Ozkan, Kenji Sagae, Louis-Philippe Morency |
COLING | 2 |
| 2010 | Practical Evaluation of Speech Recognizers for Virtual Human Dialogue Systems
Xuchen Yao, Pravin Bhutada, Kallirroi Georgila, Kenji Sagae, Ron Artstein, David R. Traum |
LREC | 4 |
| 2009 | Can I Finish? Learning When to Respond to Incremental Interpretation Results in Interactive Dialogue
David DeVault, Kenji Sagae, David R. Traum |
SIGDIAL Conference | 2 |
| 2009 | Evaluating contributions of natural language parsers to protein-protein interaction extractionabstractMOTIVATION: While text mining technologies for biomedical research have gained popularity as a way to take advantage of the explosive growth of information in text form in biomedical papers, selecting appropriate natural language processing (NLP) tools is still difficult for researchers who are not familiar with recent advances in NLP. This article provides a comparative evaluation of several state-of-the-art natural language parsers, focusing on the task of extracting protein-protein interaction (PPI) from biomedical papers. We measure how each parser, and its output representation, contributes to accuracy improvement when the parser is used as a component in a PPI system. RESULTS: All the parsers attained improvements in accuracy of PPI extraction. The levels of accuracy obtained with these different parsers vary slightly, while differences in parsing speed are larger. The best accuracy in this work was obtained when we combined Miyao and Tsujii's Enju parser and Charniak and Johnson's reranking parser, and the accuracy is better than the state-of-the-art results on the same data. AVAILABILITY: The PPI extraction system used in this work (AkanePPI) is available online at http://www-tsujii.is.s.u-tokyo.ac.jp/downloads/downloads.cgi. The evaluated parsers are also available online from each developer's site. Yusuke Miyao, Kenji Sagae, Rune Sætre, Takuya Matsuzaki, Jun'ichi Tsujii |
Bioinform. | 2 |
| 2008 | Task-oriented Evaluation of Syntactic Parsers and Their Representations
Yusuke Miyao, Rune Sætre, Kenji Sagae, Takuya Matsuzaki, Jun'ichi Tsujii |
ACL | 3 |
| 2008 | Shift-Reduce Dependency DAG Parsing
Kenji Sagae, Jun'ichi Tsujii |
COLING | 1 |
| 2008 | GENIA-GR: a Grammatical Relation Corpus for Parser Evaluation in the Biomedical Domain
Yuka Tateisi, Yusuke Miyao, Kenji Sagae, Jun'ichi Tsujii |
LREC | 3 |
| 2007 | HPSG Parsing with Shallow Dependency Constraints
Kenji Sagae, Yusuke Miyao, Jun'ichi Tsujii |
ACL | 1 |
| 2007 | Dependency Parsing and Domain Adaptation with LR Models and Parser Ensembles
Kenji Sagae, Jun'ichi Tsujii |
EMNLP-CoNLL | 1 |
| 2006 | A Best-First Probabilistic Shift-Reduce Parser
Kenji Sagae, Alon Lavie |
ACL | 1 |
| 2006 | A Fast, Accurate Deterministic Parser for ChineseabstractWe present a novel classifier-based deterministic parser for Chinese constituency parsing. Our parser computes parse trees from bottom up in one pass, and uses classifiers to make shift-reduce decisions. Trained and evaluated on the standard training and test sets, our best model (using stacked classifiers) runs in linear time and has labeled precision and recall above 88% using gold-standard part-of-speech tags, surpassing the best published results. Our SVM parser is 2-13 times faster than state-of-the-art parsers, while producing more accurate results. Our Maxent and DTree parsers run at speeds 40-270 times faster than state-of-the-art parsers, but with 5-6% losses in accuracy. Mengqiu Wang, Kenji Sagae, Teruko Mitamura |
ACL | 2 |
| 2006 | Parser Combination by Reparsing
Kenji Sagae, Alon Lavie |
HLT-NAACL | 1 |
| 2005 | Automatic Measurement of Syntactic Development in Child LanguageabstractTo facilitate the use of syntactic information in the study of child language acquisition, a coding scheme for Grammatical Relations (GRs) in transcripts of parent-child dialogs has been proposed by Sagae, MacWhinney and Lavie (2004). We discuss the use of current NLP techniques to produce the GRs in this annotation scheme. By using a statistical parser (Charniak, 2000) and memory-based learning tools for classification (Daelemans et al., 2004), we obtain high precision and recall of several GRs. We demonstrate the usefulness of this approach by performing automatic measurements of syntactic development with the Index of Productive Syntax (Scarborough, 1990) at similar levels to what child language researchers compute manually. Kenji Sagae, Alon Lavie, Brian MacWhinney |
ACL | 1 |
| 2004 | Adding Syntactic Annotations to Transcripts of Parent-Child Dialogs
Kenji Sagae, Brian MacWhinney, Alon Lavie |
LREC | 1 |