VLDB 2026 Research / reviewers in the wild / expert
Yusuke Miyao
dblp:34/467
· DBLP profile ↗
133ranked-venue papers
15as first author
44since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 127 · 14 first-author · 41 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The Imperfective Paradox in Large Language ModelsabstractDo Large Language Models (LLMs) genuinely grasp the compositional semantics of events, or do they rely on surface-level probabilistic heuristics?We investigate the Imperfective Paradox, a logical phenomenon where the past progressive aspect entails event realization for activities (e.g., running → ran) but not for accomplishments (e.g., building ↛ built).We introduce IMPERFECTIVENLI, a diagnostic dataset designed to probe this distinction across diverse semantic classes.Evaluating state-ofthe-art open-weight models, we uncover a pervasive Teleological Bias: models systematically hallucinate completion for goal-oriented events, even overriding explicit textual cancellation.Prompting interventions partially reduce this bias but trigger a calibration crisis, causing models to incorrectly reject valid entailments for atelic verbs.Representational analyses further show that while internal embeddings often distinguish progressive from simple past forms, inference decisions are dominated by strong priors about goal attainment.Taken together, our findings indicate that these current openweight LLMs operate as predictive narrative engines rather than faithful logical reasoners, and that resolving aspectual inference requires moving beyond prompting toward structurally grounded alignment.1 * Work done while visiting The University of Tokyo. 1 The data and code are available at: https://github. com/boleima/ImperfectiveParadox. Class Examples (Premise → Hypothesis)Entail?Telic P: The carpenter was building a gazebo.✗ No H: The carpenter built a gazebo. Bolei Ma, Yusuke Miyao |
ACL (1) | 2 |
| 2026 | Multi-Agent Debate for Machine Translation: A Case Study on English-Japanese TranslationabstractAs machine translation increasingly requires deeper contextual, linguistic, and cultural understanding, multi-agent collaboration has emerged as a promising approach. Multi-agent debate (MAD) frameworks, in which multiple agents deliberate to produce a final output, have shown strong performance on objective tasks, but remain underexplored in translation, where multiple valid renderings often exist. We adapt three MAD frameworks for English-Japanese translation and evaluate them against strong generative baselines, reasoning-capable LLMs, and a prompt-based self-reflection baseline. Across general-domain and culturally grounded datasets, the Society of Mind (SoM) variant yields the strongest results in the English-to-Japanese direction, showing that zero-shot translations leave substantial room for improvement through structured deliberation. Yet the gains of debate are front-loaded: later rounds do not reliably improve quality and often reintroduce translation errors. Diagnostic and error-span analyses show that hand-designed debate protocols tend to over-revise already strong translations, leading to semantic drift and process-induced degradation. These findings highlight both the promise and the limitations of agentic translation, and suggest that effective debate-based systems require mechanisms for preserving strong intermediate outputs. Zhan Shen, Jason Naradowsky, Yusuke Miyao |
EAMT (1) | 4 |
| 2026 | JBE-QA: Japanese Bar Exam QA Dataset for Assessing Legal Domain KnowledgeabstractWe introduce JBE-QA, a Japanese Bar Exam Question-Answering dataset to evaluate large language models' legal knowledge. Derived from the multiple-choice (tanto-shiki) section of the Japanese bar exam (2015-2024), JBE-QA provides the first comprehensive benchmark for Japanese legal-domain evaluation of LLMs. It covers the Civil Code, the Penal Code, and the Constitution, extending beyond the Civil Code focus of prior Japanese resources. Each question is decomposed into independent true/false judgments with structured contextual fields. The dataset contains 3,464 items with balanced labels. We evaluate 26 LLMs, including proprietary, open-weight, Japanese-specialised, and reasoning models. Our results show that proprietary models with reasoning enabled perform best, and the Constitution questions are generally easier than the Civil Code or the Penal Code questions. Zhihan Cao, Fumihito Nishino, Hiroaki Yamada 0002, Ha Thanh Nguyen, Yusuke Miyao, Ken Satoh |
LREC | 5 |
| 2026 | Evaluating Social Intelligence in LLMs via Japanese Honorifics in Email Generation: A Social Semiotic System Perspective
Muxuan Liu, Tatsuya Ishigaki, Yusuke Miyao, Hiroya Takamura, Ichiro Kobayashi 0001 |
LREC | 3 |
| 2025 | A Statistical and Multi-Perspective Revisiting of the Membership Inference Attack in Large Language ModelsabstractThe lack of data transparency in Large Language Models (LLMs) has highlighted the importance of Membership Inference Attack (MIA), which differentiates trained (member) and untrained (non-member) data. Though it shows success in previous studies, recent research reported a near-random performance in different settings, highlighting a significant performance inconsistency. We assume that a single setting doesn’t represent the distribution of the vast corpora, causing members and non-members with different distributions to be sampled and causing inconsistency. In this study, instead of a single setting, we statistically revisit MIA methods from various settings with thousands of experiments for each MIA method, along with study in text feature, embedding, threshold decision, and decoding dynamics of members and non-members. We found that (1) MIA performance improves with model size and varies with domains, while most methods do not statistically outperform baselines, (2) Though MIA performance is generally low, a notable amount of differentiable member and non-member outliers exists and vary across MIA methods, (3) Deciding a threshold to separate members and non-members is an overlooked challenge, (4) Text dissimilarity and long text benefit MIA performance, (5) Differentiable or not is reflected in the LLM embedding, (6) Member and non-members show different decoding dynamics. Namgi Han, Yusuke Miyao |
ACL (1) | 3 |
| 2025 | Do Self-Supervised Speech Models Exhibit the Critical Period Effects in Language Acquisition?abstractThis paper investigates whether the Critical Period (CP) effects in human language acquisition are observed in self-supervised speech models (S3Ms). CP effects refer to greater difficulty in acquiring a second language (L2) with delayed L2 exposure onset, and greater retention of their first language (L1) with delayed L1 exposure offset. While previous work has studied these effects using textual language models, their presence in speech models remains underexplored despite the central role of spoken language in human language acquisition. We train S3Ms with varying L2 training onsets and L1 training offsets on child-directed speech and evaluate their phone discrimination performance. We find that S3Ms do not exhibit clear evidence of either CP effects in terms of phonological acquisition. Notably, models with delayed L2 exposure onset tend to perform better on L2 and delayed L1 exposure offset leads to L1 forgetting. Yurie Koga, Shunsuke Kando, Yusuke Miyao |
ASRU | 3 |
| 2025 | GADFA: Generator-Assisted Decision-Focused Approach for Opinion Expressing Timing IdentificationabstractThe advancement of text generation models has granted us the capability to produce coherent and convincing text on demand. Yet, in real-life circumstances, individuals do not continuously generate text or voice their opinions. For instance, consumers pen product reviews after weighing the merits and demerits of a product, and professional analysts issue reports following significant news releases. In essence, opinion expression is typically prompted by particular reasons or signals. Despite long-standing developments in opinion mining, the appropriate timing for expressing an opinion remains largely unexplored. To address this deficit, our study introduces an innovative task - the identification of news-triggered opinion expressing timing. We ground this task in the actions of professional stock analysts and develop a novel dataset for investigation. Our Generator-Assisted Decision-Focused Approach (GADFA) is decision-focused, leveraging text generation models to steer the classification model, thus enhancing overall performance. Our experimental findings demonstrate that the text generated by our model contributes fresh insights from various angles, effectively aiding in identifying the optimal timing for opinion expression. Chung-Chi Chen 0001, Hiroya Takamura, Ichiro Kobayashi 0001, Yusuke Miyao, Hsin-Hsi Chen |
COLING | 4 |
| 2025 | Evaluating Local LLMs on Japanese National University Entrance Examination Dataset in Comparison with Student Performance
Kyosuke Takami, Satoshi Sekine, Yusuke Miyao |
EDM | 3 |
| 2025 | Massive Supervised Fine-tuning Experiments Reveal How Data, Layer, and Training Factors Shape LLM Alignment QualityabstractSupervised fine-tuning (SFT) is a critical step in aligning large language models (LLMs) with human instructions and values, yet many aspects of SFT remain poorly understood.We trained a wide range of base models on a variety of datasets including code generation, mathematical reasoning, and general-domain tasks, resulting in 1,000+ SFT models under controlled conditions.We then identified the dataset properties that matter most and examined the layer-wise modifications introduced by SFT.Our findings reveal that some training-task synergies persist across all models while others vary substantially, emphasizing the importance of model-specific strategies.Moreover, we demonstrate that perplexity consistently predicts SFT effectiveness, often surpassing superficial similarity between the training data and the benchmark, and that mid-layer weight changes correlate most strongly with performance gains.We release these 1,000+ SFT models and benchmark results to accelerate further research.All resources are available at https://github. Yuto Harada, Yusuke Yamauchi, Yusuke Oda, Yohei Oseki, Yusuke Miyao, Yu Takagi |
EMNLP | 5 |
| 2025 | Improving Unsupervised Constituency Parsing via Maximizing Semantic InformationabstractUnsupervised constituency parsers organize phrases within a sentence into a tree-shaped syntactic constituent structure that reflects the organization of sentence semantics.
However, the traditional objective of maximizing sentence log-likelihood (LL) does not explicitly account for the close relationship between the constituent structure and the semantics, resulting in a weak correlation between LL values and parsing accuracy.
In this paper, we introduce a novel objective that trains parsers by maximizing SemInfo, the semantic information encoded in constituent structures.
We introduce a bag-of-substrings model to represent the semantics and estimate the SemInfo value using the probability-weighted information metric.
We apply the SemInfo maximization objective to training Probabilistic Context-Free Grammar (PCFG) parsers and develop a Tree Conditional Random Field (TreeCRF)-based model to facilitate the training.
Experiments show that SemInfo correlates more strongly with parsing accuracy than LL, establishing SemInfo as a better unsupervised parsing objective.
As a result, our algorithm significantly improves parsing accuracy by an average of 7.85 sentence-F1 scores across five PCFG variants and in four languages, achieving state-of-the-art level results in three of the four languages. Xiangheng He, Yusuke Miyao, Danushka Bollegala |
ICLR | 3 |
| 2025 | Evaluating LLMs' Ability to Understand Numerical Time Series for Text GenerationabstractData-to-text generation tasks often involve processing numerical time-series as input such as financial statistics or meteorological data. Although large language models (LLMs) are a powerful approach to data-to-text, we still lack a comprehensive understanding of how well they actually understand time-series data. We therefore introduce a benchmark with 18 evaluation tasks to assess LLMs’ abilities of interpreting numerical time-series, which are categorized into: 1) event detection—identifying maxima and minima; 2) computation—averaging and summation; 3) pairwise comparison—comparing values over time; and 4) inference—imputation and forecasting. Our experiments reveal five key findings: 1) even state-of-the-art LLMs struggle with complex multi-step reasoning; 2) tasks that require extracting values or performing computations within a specified range of the time-series significantly reduce accuracy; 3) instruction tuning offers inconsistent improvements for numerical interpretation; 4) reasoning-based models outperform standard LLMs in complex numerical tasks; and 5) LLMs perform interpolation better than forecasting. These results establish a clear baseline and serve as a wake-up call for anyone aiming to blend fluent language with trustworthy numeric precision in time-series scenarios. Mizuki Arai, Tatsuya Ishigaki, Masayuki Kawarada, Yusuke Miyao, Hiroya Takamura, Ichiro Kobayashi 0001 |
INLG | 4 |
| 2025 | Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models
Shunsuke Kando, Yusuke Miyao, Shinnosuke Takamichi |
INTERSPEECH | 2 |
| 2025 | A Simple-Yet-Effective Data Augmentation Method for Speaker Identification in Novels
Wenjie Zhong, Jason Naradowsky, Yusuke Miyao |
INTERSPEECH | 3 |
| 2024 | VAD Emotion Control in Visual Art Captioning via Disentangled Multimodal RepresentationabstractArt evokes distinct affective responses, leading to the generation of emotional verbal expressions in response to visual stimuli like visual art images. While previous research has addressed controlling linguistic impressions in terms of basic emotion categories, less attention has been given to their continuous, dimensional nature. This paper aims to continuously modulate emotions across valence, arousal, and dominance (VAD) dimensions by employing multimodal representation learning (MMRL) and crossmodal inference. We utilize a multimodal variational autoencoder for MMRL, encoding visual and textual stimuli into a continuous joint vector representation. We then auxiliarily tune this representation to explicitly include the VAD dimensions in a disentangled manner. This enables nuanced emotional control in image captioning, where visual inputs are converted into vectors and then into text. Trained on the ArtEmis dataset, which includes emotion-evoking captions for visual art, as well as additional datasets annotated with VAD scores for either text or visual art, our model demonstrates that manipulating VAD intensities in the vector representation results in continuous caption variations, reflecting the intended emotional changes. Ryo Ueda, Hiromi Narimatsu, Yusuke Miyao, Shiro Kumano |
ACII | 3 |
| 2024 | Evaluating Intention Detection Capability of Large Language Models in Persuasive DialoguesabstractWe investigate intention detection in persuasive multi-turn dialogues employing the largest available Large Language Models (LLMs).Much of the prior research measures the intention detection capability of machine learning models without considering the conversational history.To evaluate LLMs' intention detection capability in conversation, we modified the existing datasets of persuasive conversation and created datasets using a multiplechoice paradigm.1 It is crucial to consider others' perspectives through their utterances when engaging in a persuasive conversation, especially when making a request or reply that is inconvenient for others.This feature makes the persuasive dialogue suitable for the dataset of measuring intention detection capability.We incorporate the concept of face acts, which categorize how utterances affect mental states.This approach enables us to measure intention detection capability by focusing on crucial intentions and to conduct comprehensible analysis according to intention types. Hiromasa Sakurai, Yusuke Miyao |
ACL (1) | 2 |
| 2024 | Professionalism-Aware Pre-Finetuning for Profitability RankingabstractOpinion mining, specifically in the investment sector, has experienced a significant increase in interest over recent years. This paper presents a novel approach to overcome current limitations in assessing and ranking investor opinions based on profitability. The study introduces a pre-finetuning scheme to improve language models' capacity to distinguish professionalism, thus enabling ranking of all available opinions. Furthermore, the paper evaluates ranking results using traditional metrics and suggests the use of a pairwise setting for better performances over a regression setting. Lastly, our method is shown to be effective across various investor opinion tasks, encompassing both professional and amateur investors. The results indicate that this approach significantly enhances the efficiency and accuracy of opinion mining in the investment sector. Chung-Chi Chen 0001, Hiroya Takamura, Ichiro Kobayashi 0001, Yusuke Miyao |
CIKM | 4 |
| 2024 | Emergent Communication with Stack-Based Agents
Daichi Kato, Ryo Ueda, Jason Naradowsky, Yusuke Miyao |
CogSci | 4 |
| 2024 | What Is Needed for Intra-document Disambiguation of Math Identifiers?abstractIn automated scientific document analysis, accurately interpreting math formulae is imperative alongside comprehending natural language. Ambiguity in math identifiers within a single document poses significant challenges to understanding math formulae. While disambiguating math identifiers across documents has seen some progress, resolving ambiguity within a document remains inadequately researched due to complexity and insufficient datasets. The level of difficulty and information required to accomplish this task was uncertain. This study aims to determine which information is necessary for the intra-document disambiguation of math identifiers. Our findings indicate that the position data and local formula structure surrounding the identifiers, including modifiers, are particularly critical. For our study, we expanded a dataset for formula grounding and doubled its size to include annotations for 27,655 math identifier occurrences. We have created a multi-layer perceptron model that performs similarly to humans, with an 85% accuracy and a kappa value of 0.73, outperforming rule-based baselines. We trained and evaluated the model with papers in natural language processing (NLP). Our findings were also confirmed valid in fields other than NLP by applying the trained models to papers from various fields. These results will aid in improving mathematical language processing, such as mathematical information retrieval. Takuto Asakura, Yusuke Miyao |
LREC/COLING | 2 |
| 2024 | Integrating Headedness Information into an Auto-generated Multilingual CCGbank for Improved Semantic InterpretationabstractPreviously, we introduced a method to generate a multilingual Combinatory Categorial Grammar (CCG) treebank by converting from the Universal Dependencies (UD). However, the method only produces bare CCG derivations without any accompanying semantic representations, which makes it difficult to obtain satisfactory analyses for constructions that involve non-local dependencies, such as control/raising or relative clauses, and limits the general applicability of the treebank. In this work, we present an algorithm that adds semantic representations to existing CCG derivations, in the form of predicate-argument structures. Through hand-crafted rules, we enhance each CCG category with headedness information, with which both local and non-local dependencies can be properly projected. This information is extracted from various sources, including UD, Enhanced UD, and proposition banks. Evaluation of our projected dependencies on the English PropBank and the Universal PropBank 2.0 shows that they can capture most of the semantic dependencies in the target corpora. Further error analysis measures the effectiveness of our algorithm for each language tested, and reveals several issues with the previous method and source data. Tu-Anh Tran, Yusuke Miyao |
LREC/COLING | 2 |
| 2024 | Who Said What: Formalization and Benchmarks for the Task of Quote AttributionabstractThe task of quote attribution seeks to pair textual utterances with the name of their speakers. Despite continuing research efforts on the task, models are rarely evaluated systematically against previous models in comparable settings on the same datasets. This has resulted in a poor understanding of the relative strengths and weaknesses of various approaches. In this work we formalize the task of quote attribution, and in doing so, establish a basis of comparison across existing models. We present an exhaustive benchmark of known models, including natural extensions to larger LLM base models, on all available datasets in both English and Chinese. Our benchmarking results reveal that the CEQA model attains state-of-the-art performance among all supervised methods, and ChatGPT, operating in a four-shot setting, demonstrates performance on par with or surpassing that of supervised methods on some datasets. Detailed error analysis identify several key factors contributing to prediction errors. Wenjie Zhong, Jason Naradowsky, Hiroya Takamura, Ichiro Kobayashi 0001, Yusuke Miyao |
LREC/COLING | 5 |
| 2024 | A Multi-Perspective Analysis of Memorization in Large Language ModelsabstractLarge Language Models (LLMs) can generate the same sequences contained in the pre-train corpora, known as memorization. Previous research studied it at a macro level, leaving micro yet important questions under-explored, e.g., what makes sentences memorized, the dynamics when generating memorized sequence, its connection to unmemorized sequence, and its predictability. We answer the above questions by analyzing the relationship of memorization with outputs from LLM, namely, embeddings, probability distributions, and generated tokens. A memorization score is calculated as the overlap between generated tokens and actual continuations when the LLM is prompted with a context sequence from the pre-train corpora. Our findings reveal: (1) The inter-correlation between memorized/unmemorized sentences, model size, continuation size, and context size, as well as the transition dynamics between sentences of different memorization scores, (2) A sudden drop and increase in the frequency of input tokens when generating memorized/unmemorized sequences (boundary effect), (3) Cluster of sentences with different memorization scores in the embedding space, (4) An inverse boundary effect in the entropy of probability distributions for generated memorized/unmemorized sequences, (5) The predictability of memorization is related to model size and continuation length. In addition, we show a Transformer model trained by the hidden states of LLM can predict unmemorized tokens. Bowen Chen 0001, Namgi Han, Yusuke Miyao |
EMNLP | 3 |
| 2024 | Forecasting Implicit Emotions Elicited in ConversationsabstractThis paper aims to forecast the implicit emotion elicited in the dialogue partner by a textual input utterance.Forecasting the interlocutor's emotion is beneficial for natural language generation in dialogue systems to avoid generating utterances that make the users uncomfortable.Previous studies forecast the emotion conveyed in the interlocutor's response, assuming it will explicitly reflect their elicited emotion.However, true emotions are not always expressed verbally.We propose a new task to directly forecast the implicit emotion elicited by an input utterance, which does not rely on this assumption.We compare this task with related ones to investigate the impact of dialogue history and one's own utterance on predicting explicit and implicit emotions.Our result highlights the importance of dialogue history for predicting implicit emotions.It also reveals that, unlike explicit emotions, implicit emotions show limited improvement in predictive performance with one's own utterance, and that they are more difficult to predict than explicit emotions.We find that even a large language model (LLM) struggles to forecast implicit emotions accurately. Yurie Koga, Shunsuke Kando, Yusuke Miyao |
INLG | 3 |
| 2024 | Leveraging Plug-and-Play Models for Rhetorical Structure Control in Text GenerationabstractWe propose a method that extends a BARTbased language generator using the plug-andplay language model to control the rhetorical structure of generated text.Our approach considers rhetorical relations between clauses and generates sentences that reflect this structure using plug-and-play language models.We evaluated our method using the Newsela corpus, which consists of texts at various levels of English proficiency.Our experiments demonstrated that our method outperforms the vanilla BART in terms of the correctness of output discourse and rhetorical structures.In existing methods, the rhetorical structure tends to deteriorate when compared to the baseline, the vanilla BART, as measured by n-gram overlap metrics such as BLEU.However, our proposed method does not exhibit this significant deterioration, demonstrating its advantage. Yuka Yokogawa, Tatsuya Ishigaki, Hiroya Takamura, Yusuke Miyao, Ichiro Kobayashi 0001 |
INLG | 4 |
| 2024 | Textless Dependency Parsing by Labeled Sequence Prediction
Shunsuke Kando, Yusuke Miyao, Jason Naradowsky, Shinnosuke Takamichi |
INTERSPEECH | 2 |
| 2024 | Language Model Based Unsupervised Dependency Parsing with Conditional Mutual Information and Grammatical ConstraintsabstractJunjie Chen, Xiangheng He, Yusuke Miyao. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Xiangheng He, Yusuke Miyao |
NAACL-HLT | 3 |
| 2024 | Evaluating LlaMA-2's Adaptation to Social Context in Japanese Emails via Fine-Tuning
Muxuan Liu, Tatsuya Ishigaki, Yusuke Miyao, Hiroya Takamura, Ichiro Kobayashi 0001 |
PACLIC | 3 |
| 2024 | Self-Emotion Blended Dialogue Generation in Social Simulation AgentsabstractWhen engaging in conversations, dialogue agents in a virtual simulation environment may exhibit their own emotional states that are unrelated to the immediate conversational context, a phenomenon known as self-emotion.This study explores how such self-emotion affects the agents' behaviors in dialogue strategies and decision-making within a large language model (LLM)-driven simulation framework.In a dialogue strategy prediction experiment, we analyze the dialogue strategy choices employed by agents both with and without self-emotion, comparing them to those of humans.The results show that incorporating self-emotion helps agents exhibit more human-like dialogue strategies.In an independent experiment comparing the performance of models fine-tuned on GPT-4 generated dialogue datasets, we demonstrate that self-emotion can lead to better overall naturalness and humanness.Finally, in a virtual simulation environment where agents have discussions on multiple topics, we show that self-emotion of agents can significantly influence the decision-making process of the agents, leading to approximately a 50% change in decisions. Jason Naradowsky, Yusuke Miyao |
SIGDIAL | 3 |
| 2024 | Travel Agency Task Dialogue Corpus: A Multimodal Dataset with Age-Diverse SpeakersabstractWhen individuals communicate, they use different vocabularies, speaking speeds, facial expressions, and gestural languages, depending on those with whom they are speaking. This study focuses on the age of the speaker as a factor that affects the style of communication. We collected a multimodal dialogue corpus with various speaker ages. We used travel as the topic, as it interests people of all ages, and we set up a task based on a tourism consultation between an operator and a customer at a travel agency. This article presents the details of the dialogue task, collection procedures and annotations, and analysis of the characteristics of the dialogues and facial expressions, focusing on the age of the speakers. The results of the analysis suggest that the adult speakers have more independent opinions, the older speakers express their opinions more frequently than other age groups, and those in the operator role smile more frequently at minors. Michimasa Inaba, Yuya Chiba, Zhiyang Qi, Ryuichiro Higashinaka, Kazunori Komatani, Yusuke Miyao, Takayuki Nagai |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 6 |
| 2023 | Emotion-Controllable Impression Utterance Generation for Visual ArtabstractThe degree of subjectivity and the type of emotions when people express their impressions of an object depend on various factors, including their affective states, psychological traits, goals, and social norms. However, most previous efforts in generating people’s impressions of objects have focused on extreme cases, i.e., fully objective or fully emotional, as represented by the MS-COCO and ArtEmis datasets, respectively. We propose an emotional impression generation method that continuously controls the degree of subjectivity versus objectivity and the type of emotion in a unified manner. In the framework of ConCap, which allows controlling text style by modifying an auxiliary input called prefix, we propose to use two types of prefixes to jointly control the degree of subjectivity and the type of emotion: subjective-style prefix and categorical-emotional-style prefix. The subjective-style prefixes are holistic and attempt to describe the entire emotion felt by a group of people, although different individuals may feel different emotions. The categorical-emotional-style prefix is more individual-oriented and tries to focus on a single or few specific emotions. An experiment using both objectively and subjectively descriptive datasets shows qualitatively and quantitatively that the proposed method provides good gradual control of impression expressions of visual art in terms of both subjectivity and objectivity as well as emotion categories. Ryo Ueda, Hiromi Narimatsu, Yusuke Miyao, Shiro Kumano |
ACII | 3 |
| 2023 | Tree-shape Uncertainty for Analyzing the Inherent Branching Bias of Unsupervised Parsing ModelsabstractThis paper presents the formalization of treeshape uncertainty that enables us to analyze the inherent branching bias of unsupervised parsing models using raw texts alone.Previous work analyzed the branching bias of unsupervised parsing models by comparing the outputs of trained parsers with gold syntactic trees.However, such approaches do not consider the fact that texts can be generated by different grammars with different syntactic trees, possibly failing to clearly separate the inherent bias of the model and the bias in train data learned by the model.To this end, we formulate tree-shape uncertainty and derive sufficient conditions that can be used for creating texts that are expected to contain no biased information on branching.In the experiment, we show that training parsers on such unbiased texts can effectively detect the branching bias of existing unsupervised parsing models.Such bias may depend only on the algorithm, or it may depend on seemingly unrelated dataset statistics such as sequence length and vocabulary size. Taiga Ishii, Yusuke Miyao |
CoNLL | 2 |
| 2023 | Fiction-Writing Mode: An Effective Control for Human-Machine Collaborative WritingabstractWe explore the idea of incorporating concepts from writing skills curricula into humanmachine collaborative writing scenarios, focusing on adding writing modes as a control for text generation models.Using crowd-sourced workers, we annotate a corpus of narrative text paragraphs with writing mode labels.Classifiers trained on this data achieve an average accuracy of ∼ 87% on held-out data.We finetune a set of large language models to condition on writing mode labels, and show that the generated text is recognized as belonging to the specified mode with high accuracy.To study the ability of writing modes to provide fine-grained control over generated text, we devise a novel turn-based text reconstruction game to evaluate the difference between the generated text and the author's intention.We show that authors prefer text suggestions made by writing mode-controlled models on average 61.1% of the time, with satisfaction scores 0.5 higher on a 5-point ordinal scale.When evaluated by humans, stories generated via collaboration with writing mode-controlled models achieve high similarity with the professionally written target story.We conclude by identifying the most common mistakes found in the generated stories.The datasets and codes are available at the Github 1 . Wenjie Zhong, Jason Naradowsky, Hiroya Takamura, Ichiro Kobayashi 0001, Yusuke Miyao |
EACL | 5 |
| 2023 | On the Word Boundaries of Emergent Languages Based on Harris's Articulation Scheme
Ryo Ueda, Taiga Ishii, Yusuke Miyao |
ICLR | 3 |
| 2023 | Constructing a Japanese Business Email Corpus Based on Social Situations
Muxuan Liu, Tatsuya Ishigaki, Yusuke Miyao, Hiroya Takamura, Ichiro Kobayashi 0001 |
PACLIC | 3 |
| 2023 | Comprehensive Evaluation of Translation Error Correction Models
Masatoshi Otake, Yusuke Miyao |
PACLIC | 2 |
| 2022 | Modeling Syntactic-Semantic Dependency Correlations in Semantic Role Labeling Using Mixture ModelsabstractIn this paper, we propose a mixture modelbased end-to-end method to model the syntactic-semantic dependency correlation in Semantic Role Labeling (SRL).Semantic dependencies in SRL are modeled as a distribution over semantic dependency labels conditioned on a predicate and an argument word.The semantic label distribution varies depending on Shortest Syntactic Dependency Path (SSDP) hop patterns.We target the variation of semantic label distributions using a mixture model, separately estimating semantic label distributions for different hop patterns and probabilistically clustering hop patterns with similar semantic label distributions.Experiments show that the proposed method successfully learns a cluster assignment reflecting the variation of semantic label distributions.Modeling the variation improves performance in predicting short distance semantic dependencies, in addition to the improvement on long distance semantic dependencies that previous syntax-aware methods have achieved.The proposed method achieves a small but statistically significant improvement over baseline methods in English, German, and Spanish and obtains competitive performance with state-of-the-art methods in English. 1 Xiangheng He, Yusuke Miyao |
ACL (1) | 3 |
| 2022 | StoryER: Automatic Story Evaluation via Ranking, Rating and ReasoningabstractExisting automatic story evaluation methods place a premium on story lexical level coherence, deviating from human preference.We go beyond this limitation by considering a novel Story Evaluation method that mimics human preference when judging a story, namely StoryER, which consists of three sub-tasks: Ranking, Rating and Reasoning.Given either a machine-generated or a human-written story, StoryER requires the machine to output 1) a preference score that corresponds to human preference, 2) specific ratings and their corresponding confidences and 3) comments for various aspects (e.g., opening, character-shaping).To support these tasks, we introduce a wellannotated dataset comprising (i) 100k ranked story pairs; and (ii) a set of 46k ratings and comments on various aspects of the story.We finetune Longformer-Encoder-Decoder (LED) on the collected dataset, with the encoder responsible for preference score and aspect prediction and the decoder for comment generation.Our comprehensive experiments result in a competitive benchmark for each task, showing the high correlation to human preference.In addition, we have witnessed the joint learning of the preference scores, the aspect ratings, and the comments brings gain in each single task.Our dataset and benchmarks are publicly available to advance the research of story evaluation tasks. 1 Hong Chen 0017, Duc Minh Vo, Hiroya Takamura, Yusuke Miyao, Hideki Nakayama |
EMNLP | 4 |
| 2022 | Open-domain Video Commentary GenerationabstractEdison Marrese-Taylor, Yumi Hamazono, Tatsuya Ishigaki, Goran Topić, Yusuke Miyao, Ichiro Kobayashi, Hiroya Takamura. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Edison Marrese-Taylor, Yumi Hamazono, Tatsuya Ishigaki, Goran Topic, Yusuke Miyao, Ichiro Kobayashi 0001, Hiroya Takamura |
EMNLP | 5 |
| 2022 | Building Dataset for Grounding of Formulae - Annotating Coreference Relations Among Math IdentifiersabstractGrounding the meaning of each symbol in math formulae is important for automated understanding of scientific documents. Generally speaking, the meanings of math symbols are not necessarily constant, and the same symbol is used in multiple meanings. Therefore, coreference relations between symbols need to be identified for grounding, and the task has aspects of both description alignment and coreference analysis. In this study, we annotated 15 papers selected from arXiv.org with the grounding information. In total, 12,352 occurrences of math identifiers in these papers were annotated, and all coreference relations between them were made explicit in each paper. The constructed dataset shows that regardless of the ambiguity of symbols in math formulae, coreference relations can be labeled with a high inter-annotator agreement. The constructed dataset enables us to achieve automation of formula grounding, and in turn, make deeper use of the knowledge in scientific documents using techniques such as math information extraction. The built grounding dataset is available at https://sigmathling.kwarc.info/resources/grounding- dataset/. Takuto Asakura, Yusuke Miyao, Akiko Aizawa |
LREC | 2 |
| 2022 | Collection and Analysis of Travel Agency Task Dialogues with Age-Diverse SpeakersabstractWhen individuals communicate with each other, they use different vocabulary, speaking speed, facial expressions, and body language depending on the people they talk to. This paper focuses on the speaker’s age as a factor that affects the change in communication. We collected a multimodal dialogue corpus with a wide range of speaker ages. As a dialogue task, we focus on travel, which interests people of all ages, and we set up a task based on a tourism consultation between an operator and a customer at a travel agency. This paper provides details of the dialogue task, the collection procedure and annotations, and the analysis on the characteristics of the dialogues and facial expressions focusing on the age of the speakers. Results of the analysis suggest that the adult speakers have more independent opinions, the older speakers more frequently express their opinions frequently compared with other age groups, and the operators expressed a smile more frequently to the minor speakers. Michimasa Inaba, Yuya Chiba, Ryuichiro Higashinaka, Kazunori Komatani, Yusuke Miyao, Takayuki Nagai |
LREC | 5 |
| 2022 | Development of a Multilingual CCG Treebank via Universal Dependencies ConversionabstractThis paper introduces an algorithm to convert Universal Dependencies (UD) treebanks to Combinatory Categorial Grammar (CCG) treebanks. As CCG encodes almost all grammatical information into the lexicon, obtaining a high-quality CCG derivation from a dependency tree is a challenging task. Our algorithm relies on hand-crafted rules to assign categories to constituents, and a non-statistical parser to derive full CCG parses given the assigned categories. To evaluate our converted treebanks, we perform lexical, sentential, and syntactic rule coverage analysis, as well as CCG parsing experiments. Finally, we discuss how our method handles complex constructions, and propose possible future extensions. Tu-Anh Tran, Yusuke Miyao |
LREC | 2 |
| 2021 | Generating Racing Game Commentary from Vision, Language, and Structured DataabstractWe propose the task of automatically generating commentaries for races in a motor racing game, from vision, structured numerical, and textual data.Commentaries provide information to support spectators in understanding events in races.Commentary generation models need to interpret the race situation and generate the correct content at the right moment.We divide the task into two subtasks: utterance timing identification and utterance generation.Because existing datasets do not have such alignments of data in multiple modalities, this setting has not been explored in depth.In this study, we introduce a new large-scale dataset that contains aligned video data, structured numerical data, and transcribed commentaries that consist of 129,226 utterances in 1,389 races in a game.Our analysis reveals that the characteristics of commentaries change depending on time and viewpoints.Our experiments on the subtasks show that it is still challenging for a state-of-the-art vision encoder to capture useful information from videos to generate accurate commentaries.We make the dataset and baseline implementation publicly available for further research.1 Tatsuya Ishigaki, Goran Topic, Yumi Hamazono, Hiroshi Noji, Ichiro Kobayashi 0001, Yusuke Miyao, Hiroya Takamura |
INLG | 6 |
| 2021 | Unpredictable Attributes in Market Comment Generation
Yumi Hamazono, Tatsuya Ishigaki, Yusuke Miyao, Hiroya Takamura, Ichiro Kobayashi 0001 |
PACLIC | 3 |
| 2021 | Talking with the Theorem Prover to Interactively Solve Natural Language Inference
Atsushi Sumita, Yusuke Miyao, Koji Mineshima |
PACLIC | 2 |
| 2021 | Controlling contents in data-to-document generation with human-designed topic labels
Kasumi Aoki, Akira Miyazawa, Tatsuya Ishigaki, Tatsuya Aoki, Hiroshi Noji, Keiichi Goshima, Hiroya Takamura, Yusuke Miyao, Ichiro Kobayashi 0001 |
Comput. Speech Lang. | 8 |
| 2020 | An empirical analysis of existing systems and datasets toward general simple question answeringabstractIn this paper, we evaluate the progress of our field toward solving simple factoid questions over a knowledge base, a practically important problem in natural language interface to database. As in other natural language understanding tasks, a common practice for this task is to train and evaluate a model on a single dataset, and recent studies suggest that SimpleQuestions, the most popular and largest dataset, is nearly solved under this setting. However, this common setting does not evaluate the robustness of the systems outside of the distribution of the used training data. We rigorously evaluate such robustness of existing systems using different datasets. Our analysis, including shifting of training and test datasets and training on a union of the datasets, suggests that our progress in solving SimpleQuestions dataset does not indicate the success of more general simple question answering. We discuss a possible future direction toward this goal. Namgi Han, Goran Topic, Hiroshi Noji, Hiroya Takamura, Yusuke Miyao |
COLING | 5 |
| 2020 | Learning with Contrastive Examples for Data-to-Text GenerationabstractYui Uehara, Tatsuya Ishigaki, Kasumi Aoki, Hiroshi Noji, Keiichi Goshima, Ichiro Kobayashi, Hiroya Takamura, Yusuke Miyao. Proceedings of the 28th International Conference on Computational Linguistics. 2020. Yui Uehara, Tatsuya Ishigaki, Kasumi Aoki, Hiroshi Noji, Keiichi Goshima, Ichiro Kobayashi 0001, Hiroya Takamura, Yusuke Miyao |
COLING | 8 |
| 2020 | Market Comment Generation from Data with Noisy AlignmentsabstractEnd-to-end models on data-to-text learn the mapping of data and text from the aligned pairs in the dataset.However, these alignments are not always obtained reliably, especially for the time-series data, for which real time comments are given to some situation and there might be a delay in the comment delivery time compared to the actual event time.To handle this issue of possible noisy alignments in the dataset, we propose a neural network model with multitimestep data and a copy mechanism, which allows the models to learn the correspondences between data and text from the dataset with noisier alignments.We focus on generating market comments in Japanese that are delivered each time an event occurs in the market.The core idea of our approach is to utilize multitimestep data, which is not only the latest market price data when the comment is delivered, but also the data obtained at several timesteps earlier.On top of this, we employ a copy mechanism that is suitable for referring to the content of data records in the market price data.We confirm the superiority of our proposal by two evaluation metrics and show the accuracy improvement of the sentence generation using the time series data by our proposed method. Yumi Hamazono, Yui Uehara, Hiroshi Noji, Yusuke Miyao, Hiroya Takamura, Ichiro Kobayashi 0001 |
INLG | 4 |
| 2020 | Analyzing Word Embedding Through Structural Equation ModelingabstractMany researchers have tried to predict the accuracies of extrinsic evaluation by using intrinsic evaluation to evaluate word embedding. The relationship between intrinsic and extrinsic evaluation, however, has only been studied with simple correlation analysis, which has difficulty capturing complex cause-effect relationships and integrating external factors such as the hyperparameters of word embedding. To tackle this problem, we employ partial least squares path modeling (PLS-PM), a method of structural equation modeling developed for causal analysis. We propose a causal diagram consisting of the evaluation results on the BATS, VecEval, and SentEval datasets, with a causal hypothesis that linguistic knowledge encoded in word embedding contributes to solving downstream tasks. Our PLS-PM models are estimated with 600 word embeddings, and we prove the existence of causal relations between linguistic knowledge evaluated on BATS and the accuracies of downstream tasks evaluated on VecEval and SentEval in our PLS-PM models. Moreover, we show that the PLS-PM models are useful for analyzing the effect of hyperparameters, including the training algorithm, corpus, dimension, and context window, and for validating the effectiveness of intrinsic evaluation. Namgi Han, Katsuhiko Hayashi 0001, Yusuke Miyao |
LREC | 3 |
| 2019 | Learning to Select, Track, and Generate for Data-to-TextabstractHayate Iso, Yui Uehara, Tatsuya Ishigaki, Hiroshi Noji, Eiji Aramaki, Ichiro Kobayashi, Yusuke Miyao, Naoaki Okazaki, Hiroya Takamura. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019. Hayate Iso, Yui Uehara, Tatsuya Ishigaki, Hiroshi Noji, Eiji Aramaki, Ichiro Kobayashi 0001, Yusuke Miyao, Naoaki Okazaki, Hiroya Takamura |
ACL (1) | 7 |
| 2019 | Controlling Contents in Data-to-Document Generation with Human-Designed Topic LabelsabstractKasumi Aoki, Akira Miyazawa, Tatsuya Ishigaki, Tatsuya Aoki, Hiroshi Noji, Keiichi Goshima, Ichiro Kobayashi, Hiroya Takamura, Yusuke Miyao. Proceedings of the 12th International Conference on Natural Language Generation. 2019. Kasumi Aoki, Akira Miyazawa, Tatsuya Ishigaki, Tatsuya Aoki, Hiroshi Noji, Keiichi Goshima, Ichiro Kobayashi 0001, Hiroya Takamura, Yusuke Miyao |
INLG | 9 |
| 2018 | A Simple Method to Remove Reviews against Guideline for Online Review ServicesabstractReviews written by customers on the Web can influence many people when they decide what to do. Offensive or irrelevant reviews are often posted to review services (e.g. TripAdvisor1and Glassdoor2), and they can make people displeased and ruin services' reputation. To avoid this, review service providers issue guidelines that define what are inappropriate reviews and employ human workers to manually remove reviews violating guideline. Such manual operations incur high costs and human filtering results may vary; then, automatic filtering is desirable. Unfortunately, although several filtering methods are available, their accuracy and efficiency are still not enough to work well on actual review services because of their costs, complexities and reviews' noisiness. In this paper, we introduce a simple, accurate, and efficient method that detects whether a review violates guidelines or not, using logistic regression models with features based on word n-gram. We show through experiments on real review data that the method works well under practical and difficult situations. The method can be applied to any review services, and can eliminate the costs of manual operations. Yasutaka Shindoh, Atsunori Kanemura, Yusuke Miyao |
IEEE BigData | 3 |
| 2018 | An Empirical Investigation of Error Types in Vietnamese ParsingabstractSyntactic parsing plays a crucial role in improving the quality of natural language processing tasks. Although there have been several research projects on syntactic parsing in Vietnamese, the parsing quality has been far inferior than those reported in major languages, such as English and Chinese. In this work, we evaluated representative constituency parsing models on a Vietnamese Treebank to look for the most suitable parsing method for Vietnamese. We then combined the advantages of automatic and manual analysis to investigate errors produced by the experimented parsers and find the reasons for them. Our analysis focused on three possible sources of parsing errors, namely limited training data, part-of-speech (POS) tagging errors, and ambiguous constructions. As a result, we found that the last two sources, which frequently appear in Vietnamese text, significantly attributed to the poor performance of Vietnamese parsing. Quy Nguyen, Yusuke Miyao, Hiroshi Noji, Nhung Nguyen |
COLING | 2 |
| 2018 | Generating Market Comments Referring to External ResourcesabstractTatsuya Aoki, Akira Miyazawa, Tatsuya Ishigaki, Keiichi Goshima, Kasumi Aoki, Ichiro Kobayashi, Hiroya Takamura, Yusuke Miyao. Proceedings of the 11th International Conference on Natural Language Generation. 2018. Tatsuya Aoki, Akira Miyazawa, Tatsuya Ishigaki, Keiichi Goshima, Kasumi Aoki, Ichiro Kobayashi 0001, Hiroya Takamura, Yusuke Miyao |
INLG | 8 |
| 2018 | Universal Dependencies Version 2 for Japanese
Masayuki Asahara, Hiroshi Kanayama, Takaaki Tanaka, Yusuke Miyao, Sumire Uematsu, Shinsuke Mori, Yuji Matsumoto 0001, Mai Omura, Yugo Murawaki |
LREC | 4 |
| 2018 | Universal Dependencies for Amharic
Binyam Ephrem Seyoum, Yusuke Miyao, Baye Yimam Mekonnen |
LREC | 2 |
| 2018 | Inducing Temporal Relations from Time Anchor AnnotationabstractRecognizing temporal relations among events and time expressions has been an essential but challenging task in natural language processing.Conventional annotation of judging temporal relations puts a heavy load on annotators.In reality, the existing annotated corpora include annotations on only "salient" event pairs, or on pairs in a fixed window of sentences.In this paper, we propose a new approach to obtain temporal relations from absolute time value (a.k.a.time anchors), which is suitable for texts containing rich temporal information such as news articles.We start from time anchors for events and time expressions, and temporal relation annotations are induced automatically by computing relative order of two time anchors.This proposal shows several advantages over the current methods for temporal relation annotation: it requires less annotation effort, can induce inter-sentence relations easily, and increases informativeness of temporal relations.We compare the empirical statistics and automatic recognition results with our data against a previous temporal relation corpus.We also reveal that our data contributes to a significant improvement of the downstream time anchor prediction task, demonstrating 14.1 point increase in overall accuracy. Fei Cheng 0002, Yusuke Miyao |
NAACL-HLT | 2 |
| 2018 | Do systems pass university entrance exams?
Álvaro Rodrigo, Anselmo Peñas, Yusuke Miyao, Noriko Kando |
Inf. Process. Manag. | 3 |
| 2017 | Learning to Generate Market Comments from Stock PricesabstractSoichiro Murakami, Akihiko Watanabe, Akira Miyazawa, Keiichi Goshima, Toshihiko Yanase, Hiroya Takamura, Yusuke Miyao. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2017. Soichiro Murakami, Akihiko Watanabe, Akira Miyazawa, Keiichi Goshima, Toshihiko Yanase, Hiroya Takamura, Yusuke Miyao |
ACL (1) | 7 |
| 2017 | On-demand Injection of Lexical Knowledge for Recognising Textual EntailmentabstractPascual Martínez-Gómez, Koji Mineshima, Yusuke Miyao, Daisuke Bekki. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017. Pascual Martínez-Gómez, Koji Mineshima, Yusuke Miyao, Daisuke Bekki |
EACL (1) | 3 |
| 2017 | Finding Prototypes of Answers for Improving Answer Sentence SelectionabstractAnswer sentence selection has been widely adopted recently for benchmarking techniques in Question Answering. Previous proposals for the task are essentially general solutions taking the form of neural networks that measure semantic similarity. In contrast, the present paper describes a simple technique to take advantage of such general-purpose tools for dealing with questions and answer sentences without changing the base system. The technique involves replacing wh-words in input questions with a word denoting the prototype of all answers. These transformed questions are passed as input to an existing neural network built for measuring semantic similarity. This technique is evaluated on two different neural network architectures over two datasets: TrecQA and WikiQA. Results of our experiments show improvement in overall accuracy across most question types we are interested in: `who', `when' and `where'-type questions. Wai Lok Tam, Namgi Han, Juan Ignacio Navarro-Horñiacek, Yusuke Miyao |
IJCAI | 4 |
| 2017 | MANet: A Modal Attention Network for Describing VideosabstractExploiting multimodal features has become a standard approach towards many video applications, including the video captioning task. One problem with the existing work is that it models the relevance of each type of features evenly, which neutralizes the impact of each individual modality to the word to be generated. In this paper, we propose a novel Modal Attention Network (MANet) to account for this issue. Our MANet extends the standard encoder-decoder network by adapting the attention mechanism to video modalities. As a result, MANet emphasizes the impact of each modality with respect to the word to be generated. Experimental results show that our MANet effectively utilizes multimodal features to generate better video descriptions. Especially, our MANet system was ranked among the top three systems at the 2nd Video to Language Challenge in both automatic metrics and human evaluations. Sang Phan Le, Yusuke Miyao, Shin'ichi Satoh 0001 |
ACM Multimedia | 2 |
| 2016 | Generating Video Description using Sequence-to-sequence Model with Temporal AttentionabstractAutomatic video description generation has recently been getting attention after rapid advancement in image caption generation. Automatically generating description for a video is more challenging than for an image due to its temporal dynamics of frames. Most of the work relied on Recurrent Neural Network (RNN) and recently attentional mechanisms have also been applied to make the model learn to focus on some frames of the video while generating each word in a describing sentence. In this paper, we focus on a sequence-to-sequence approach with temporal attention mechanism. We analyze and compare the results from different attention model configuration. By applying the temporal attention mechanism to the system, we can achieve a METEOR score of 0.310 on Microsoft Video Description dataset, which outperformed the state-of-the-art system so far. Natsuda Laokulrat, Sang Phan Le, Noriki Nishida, Raphael Shu, Yo Ehara, Naoaki Okazaki, Yusuke Miyao, Hideki Nakayama |
COLING | 7 |
| 2016 | Video Event Detection by Exploiting Word Dependencies from Image CaptionsabstractVideo event detection is a challenging problem in information and multimedia retrieval. Different from single action detection, event detection requires a richer level of semantic information from video. In order to overcome this challenge, existing solutions often represent videos using high level features such as concepts. However, concept-based representation can be confusing because it does not encode the relationship between concepts. This issue can be addressed by exploiting the co-occurrences of the concepts, however, it often leads to a very huge number of possible combinations. In this paper, we propose a new approach to obtain the relationship between concepts by exploiting the syntactic dependencies between words in the image captions. The main advantage of this approach is that it significantly reduces the number of informative combinations between concepts. We conduct extensive experiments to analyze the effectiveness of using the new dependency representation for event detection on two large-scale TRECVID Multimedia Event Detection 2013 and 2014 datasets. Experimental results show that i) Dependency features are more discriminative than concept-based features. ii) Dependency features can be combined with our current event detection system to further improve the performance. For instance, the relative improvement can be as far as 8.6% on the MEDTEST14 10Ex setting. Sang Phan Le, Yusuke Miyao, Duy-Dinh Le, Shin'ichi Satoh 0001 |
COLING | 2 |
| 2016 | Rule Extraction for Tree-to-Tree Transducers by Cost Minimization
Pascual Martínez-Gómez, Yusuke Miyao |
EMNLP | 2 |
| 2016 | Building compositional semantics and higher-order inference system for a wide-coverage Japanese CCG parserabstractThis paper presents a system that compositionally maps outputs of a wide-coverage Japanese CCG parser onto semantic representations and performs automated inference in higher-order logic.The system is evaluated on a textual entailment dataset.It is shown that the system solves inference problems that focus on a variety of complex linguistic phenomena, including those that are difficult to represent in the standard first-order logic. Koji Mineshima, Ribeka Tanaka, Pascual Martínez-Gómez, Yusuke Miyao, Daisuke Bekki |
EMNLP | 4 |
| 2016 | Using Left-corner Parsing to Encode Universal Structural Constraints in Grammar InductionabstractCenter-embedding is difficult to process and is known as a rare syntactic construction across languages.In this paper we describe a method to incorporate this assumption into the grammar induction tasks by restricting the search space of a model to trees with limited centerembedding.The key idea is the tabulation of left-corner parsing, which captures the degree of center-embedding of a parse via its stack depth.We apply the technique to learning of famous generative model, the dependency model with valence (Klein and Manning, 2004).Cross-linguistic experiments on Universal Dependencies show that often our method boosts the performance from the baseline, and competes with the current state-ofthe-art model in a number of languages. Hiroshi Noji, Yusuke Miyao, Mark Johnson 0001 |
EMNLP | 2 |
| 2016 | Challenges and Solutions for Consistent Annotation of Vietnamese Treebank
Quy Nguyen, Yusuke Miyao, Ha Le, Ngan L. T. Nguyen |
LREC | 2 |
| 2016 | Towards Comparability of Linguistic Graph Banks for Semantic Parsing
Stephan Oepen, Marco Kuhlmann, Yusuke Miyao, Daniel Zeman, Silvie Cinková, Dan Flickinger, Jan Hajic 0001, Angelina Ivanova, Zdenka Uresová |
LREC | 3 |
| 2016 | Universal Dependencies for Japanese
Takaaki Tanaka, Yusuke Miyao, Masayuki Asahara, Sumire Uematsu, Hiroshi Kanayama, Shinsuke Mori, Yuji Matsumoto 0001 |
LREC | 2 |
| 2016 | Typed Entity and Relation Annotation on Computer Science Papers
Yuka Tateisi, Tomoko Ohta, Sampo Pyysalo, Yusuke Miyao, Akiko Aizawa |
LREC | 4 |
| 2015 | Optimal Shift-Reduce Constituent Parsing with Structured PerceptronabstractLe Quang Thang, Hiroshi Noji, Yusuke Miyao. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Le Quang Thang, Hiroshi Noji, Yusuke Miyao |
ACL (1) | 3 |
| 2015 | Higher-order logical inference with compositional semanticsabstractWe present a higher-order inference system based on a formal compositional semantics and the wide-coverage CCG parser.We develop an improved method to bridge between the parser and semantic composition.The system is evaluated on the FraCaS test suite.In contrast to the widely held view that higher-order logic is unsuitable for efficient logical inferences, the results show that a system based on a reasonably-sized semantic lexicon and a manageable number of non-first-order axioms enables efficient logical inferences, including those concerned with generalized quantifiers and intensional operators, and outperforms the state-of-the-art firstorder inference system. Koji Mineshima, Pascual Martínez-Gómez, Yusuke Miyao, Daisuke Bekki |
EMNLP | 3 |
| 2015 | Paraphrase Detection Based on Identical Phrase and Similar Word Matching
Hoang-Quoc Nguyen-Son, Yusuke Miyao, Isao Echizen |
PACLIC | 2 |
| 2015 | Integrating Multiple Dependency Corpora for Inducing Wide-Coverage Japanese CCG ResourcesabstractA novel method to induce wide-coverage Combinatory Categorial Grammar (CCG) resources for Japanese is proposed in this article. For some languages including English, the availability of large annotated corpora and the development of data-based induction of lexicalized grammar have enabled deep parsing, i.e., parsing based on lexicalized grammars. However, deep parsing for Japanese has not been widely studied. This is mainly because most Japanese syntactic resources are represented in chunk-based dependency structures, while previous methods for inducing grammars are dependent on tree corpora. To translate syntactic information presented in chunk-based dependencies to phrase structures as accurately as possible, integration of annotation from multiple dependency-based corpora is proposed. Our method first integrates dependency structures and predicate-argument information and converts them into phrase structure trees. The trees are then transformed into CCG derivations in a similar way to previously proposed methods. The quality of the conversion is empirically evaluated in terms of the coverage of the obtained CCG lexicon and the accuracy of the parsing with the grammar. While the transforming process used in this study is specialized for Japanese, the framework of our method would be applicable to other languages for which dependency-based analysis has been regarded as more appropriate than phrase structure-based analysis due to morphosyntactic features. Sumire Uematsu, Takuya Matsuzaki, Hiroki Hanaoka, Yusuke Miyao, Hideki Mima |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2014 | Logical Inference on Dependency-based Compositional SemanticsabstractDependency-based Compositional Semantics (DCS) is a framework of natural language semantics with easy-to-process structures as well as strict semantics.In this paper, we equip the DCS framework with logical inference, by defining abstract denotations as an abstraction of the computing process of denotations in original DCS.An inference engine is built to achieve inference on abstract denotations.Furthermore, we propose a way to generate on-the-fly knowledge in logical inference, by combining our framework with the idea of tree transformation.Experiments on FraCaS and PASCAL RTE datasets show promising results. Yusuke Miyao, Takuya Matsuzaki |
ACL (1) | 2 |
| 2014 | Left-corner Transitions on Dependency Parsing
Hiroshi Noji, Yusuke Miyao |
COLING | 2 |
| 2014 | Formalizing Word Sampling for Vocabulary Prediction as Graph-based Active LearningabstractPredicting vocabulary of second language learners is essential to support their language learning; however, because of the large size of language vocabularies, we cannot collect information on the entire vocabulary.For practical measurements, we need to sample a small portion of words from the entire vocabulary and predict the rest of the words.In this study, we propose a novel framework for this sampling method.Current methods rely on simple heuristic techniques involving inflexible manual tuning by educational experts.We formalize these heuristic techniques as a graph-based non-interactive active learning method as applied to a special graph.We show that by extending the graph, we can support additional functionality such as incorporating domain specificity and sampling from multiple corpora.In our experiments, we show that our extended methods outperform other methods in terms of vocabulary prediction accuracy when the number of samples is small. Yo Ehara, Yusuke Miyao, Hidekazu Oiwa, Issei Sato, Hiroshi Nakagawa |
EMNLP | 2 |
| 2014 | Overview of Todai Robot Project and Evaluation Framework of its NLP-based Problem Solving
Akira Fujita, Akihiro Kameda, Ai Kawazoe, Yusuke Miyao |
LREC | 4 |
| 2014 | Annotation of Computer Science Papers for Semantic Relation Extrac-tion
Yuka Tateisi, Yo Shidahara, Yusuke Miyao, Akiko Aizawa |
LREC | 3 |
| 2014 | Encoding Generalized Quantifiers in Dependency-based Compositional Semantics
Yubing Dong, Yusuke Miyao |
PACLIC | 3 |
| 2013 | Integrating Multiple Dependency Corpora for Inducing Wide-coverage Japanese CCG Resources
Sumire Uematsu, Takuya Matsuzaki, Hiroki Hanaoka, Yusuke Miyao, Hideki Mima |
ACL (1) | 4 |
| 2013 | Improvements to the Bayesian Topic N-Gram ModelsabstractOne of the language phenomena that n-gram language model fails to capture is the topic information of a given situation.We advance the previous study of the Bayesian topic language model by Wallach (2006) in two directions: one, investigating new priors to alleviate the sparseness problem caused by dividing all ngrams into exclusive topics, and two, developing a novel Gibbs sampler that enables moving multiple n-grams across different documents to another topic.Our blocked sampler can efficiently search for higher probability space even with higher order n-grams.In terms of modeling assumption, we found it is effective to assign a topic to only some parts of a document. Hiroshi Noji, Daichi Mochihashi, Yusuke Miyao |
EMNLP | 3 |
| 2013 | Statistical Parsing with Probabilistic Symbol-Refined Tree Substitution Grammars
Hiroyuki Shindo, Yusuke Miyao, Akinori Fujino, Masaaki Nagata |
IJCAI | 2 |
| 2013 | Two-Stage Pre-ordering for Japanese-to-English Statistical Machine Translation
Sho Hoshino, Yusuke Miyao, Katsuhito Sudoh, Masaaki Nagata |
IJCNLP | 2 |
| 2013 | University Entrance Examinations as a Benchmark Resource for NLP-based Problem Solving
Yusuke Miyao, Ai Kawazoe |
IJCNLP | 1 |
| 2013 | Alignment-based Annotation of Proofreading Texts toward Professional Writing Assistance
Ngan L. T. Nguyen, Yusuke Miyao |
IJCNLP | 2 |
| 2013 | Effects of Parsing Errors on Pre-Reordering Performance for Chinese-to-Japanese SMT
Pascual Martínez-Gómez, Yusuke Miyao, Katsuhito Sudoh, Masaaki Nagata |
PACLIC | 3 |
| 2012 | Incremental Joint Approach to Word Segmentation, POS Tagging, and Dependency Parsing in Chinese
Jun Hatori, Takuya Matsuzaki, Yusuke Miyao, Jun'ichi Tsujii |
ACL (1) | 3 |
| 2012 | Bayesian Symbol-Refined Tree Substitution Grammars for Syntactic Parsing
Hiroyuki Shindo, Yusuke Miyao, Akinori Fujino, Masaaki Nagata |
ACL (1) | 2 |
| 2012 | Answering Yes/No Questions via Question Inversion
Hiroshi Kanayama, Yusuke Miyao, John Prager |
COLING | 2 |
| 2012 | Framework of Semantic Role Assignment based on Extended Lexical Conceptual Structure: Comparison with VerbNet and FrameNet
Yuichiroh Matsubayashi, Yusuke Miyao, Akiko Aizawa |
EACL | 2 |
| 2012 | Annotating Factive Verbs
Alvin Grissom II, Yusuke Miyao |
LREC | 2 |
| 2012 | Building Japanese Predicate-argument Structure Corpus using Lexical Conceptual Structure
Yuichiroh Matsubayashi, Yusuke Miyao, Akiko Aizawa |
LREC | 2 |
| 2012 | Evaluating Textual Entailment Recognition for University Entrance ExaminationsabstractThe present article addresses an attempt to apply questions in university entrance examinations to the evaluation of textual entailment recognition. Questions in several fields, such as history and politics, primarily test the examinee’s knowledge in the form of choosing true statements from multiple choices. Answering such questions can be regarded as equivalent to finding evidential texts from a textbase such as textbooks and Wikipedia. Therefore, this task can be recast as recognizing textual entailment between a description in a textbase and a statement given in a question. We focused on the National Center Test for University Admission in Japan and converted questions into the evaluation data for textual entailment recognition by using Wikipedia as a textbase. Consequently, it is revealed that nearly half of the questions can be mapped into textual entailment recognition; 941 text pairs were created from 404 questions from six subjects. This data set is provided for a subtask of NTCIR RITE (Recognizing Inference in Text), and 16 systems from six teams used the data set for evaluation. The evaluation results revealed that the best system achieved a correct answer ratio of 56%, which is significantly better than a random choice baseline. Yusuke Miyao, Hideki Shima 0001, Hiroshi Kanayama, Teruko Mitamura |
ACM Trans. Asian Lang. Inf. Process. | 1 |
| 2011 | Learning with Lookahead: Can History-Based Models Rival Globally Optimized Models?
Yoshimasa Tsuruoka, Yusuke Miyao, Jun'ichi Kazama |
CoNLL | 2 |
| 2011 | Training Dependency Parsers from Partially Annotated Corpora
Daniel Flannery, Yusuke Miyao, Graham Neubig, Shinsuke Mori |
IJCNLP | 2 |
| 2011 | Exploring Difficulties in Parsing Imperatives and Questions
Tadayoshi Hara, Takuya Matsuzaki, Yusuke Miyao, Jun'ichi Tsujii |
IJCNLP | 3 |
| 2011 | Incremental Joint POS Tagging and Dependency Parsing in Chinese
Jun Hatori, Takuya Matsuzaki, Yusuke Miyao, Jun'ichi Tsujii |
IJCNLP | 3 |
| 2010 | Entity-Focused Sentence Simplification for Relation Extraction
Makoto Miwa, Rune Sætre, Yusuke Miyao, Jun'ichi Tsujii |
COLING | 3 |
| 2010 | A Modular Architecture for the Wide-Coverage Translation of Natural Language Texts into Predicate Logic Formulas
Yusuke Miyao, Alastair Butler, Kei Yoshimoto, Jun'ichi Tsujii |
PACLIC | 1 |
| 2009 | Descriptive and Empirical Approaches to Capturing Underlying Dependencies among Parsing Errors
Tadayoshi Hara, Yusuke Miyao, Jun'ichi Tsujii |
EMNLP | 2 |
| 2009 | A Rich Feature Vector for Protein-Protein Interaction Extraction from Multiple Corpora
Makoto Miwa, Rune Sætre, Yusuke Miyao, Jun'ichi Tsujii |
EMNLP | 3 |
| 2009 | Supervised Learning of a Probabilistic Lexicon of Verb Semantic Classes
Yusuke Miyao, Jun'ichi Tsujii |
EMNLP | 1 |
| 2009 | Design of Chinese HPSG Framework for Data-Driven Parsing
Xiangli Wang 0001, Shun'ya Iwasawa, Yusuke Miyao, Takuya Matsuzaki, Jun'ichi Tsujii |
PACLIC | 3 |
| 2009 | Evaluating contributions of natural language parsers to protein-protein interaction extractionabstractMOTIVATION: While text mining technologies for biomedical research have gained popularity as a way to take advantage of the explosive growth of information in text form in biomedical papers, selecting appropriate natural language processing (NLP) tools is still difficult for researchers who are not familiar with recent advances in NLP. This article provides a comparative evaluation of several state-of-the-art natural language parsers, focusing on the task of extracting protein-protein interaction (PPI) from biomedical papers. We measure how each parser, and its output representation, contributes to accuracy improvement when the parser is used as a component in a PPI system. RESULTS: All the parsers attained improvements in accuracy of PPI extraction. The levels of accuracy obtained with these different parsers vary slightly, while differences in parsing speed are larger. The best accuracy in this work was obtained when we combined Miyao and Tsujii's Enju parser and Charniak and Johnson's reranking parser, and the accuracy is better than the state-of-the-art results on the same data. AVAILABILITY: The PPI extraction system used in this work (AkanePPI) is available online at http://www-tsujii.is.s.u-tokyo.ac.jp/downloads/downloads.cgi. The evaluated parsers are also available online from each developer's site. Yusuke Miyao, Kenji Sagae, Rune Sætre, Takuya Matsuzaki, Jun'ichi Tsujii |
Bioinform. | 1 |
| 2008 | Task-oriented Evaluation of Syntactic Parsers and Their Representations
Yusuke Miyao, Rune Sætre, Kenji Sagae, Takuya Matsuzaki, Jun'ichi Tsujii |
ACL | 1 |
| 2008 | Towards Data and Goal Oriented Analysis: Tool Inter-operability and Combinatorial Comparison
Yoshinobu Kano, Ngan L. T. Nguyen, Rune Sætre, Kazuhiro Yoshida, Keiichiro Fukamachi, Yusuke Miyao, Yoshimasa Tsuruoka, Sophia Ananiadou, Jun'ichi Tsujii |
IJCNLP | 6 |
| 2008 | GENIA-GR: a Grammatical Relation Corpus for Parser Evaluation in the Biomedical Domain
Yuka Tateisi, Yusuke Miyao, Kenji Sagae, Jun'ichi Tsujii |
LREC | 2 |
| 2008 | Feature Forest Models for Probabilistic HPSG ParsingabstractProbabilistic modeling of lexicalized grammars is difficult because these grammars exploit complicated data structures, such as typed feature structures. This prevents us from applying common methods of probabilistic modeling in which a complete structure is divided into sub-structures under the assumption of statistical independence among sub-structures. For example, part-of-speech tagging of a sentence is decomposed into tagging of each word, and CFG parsing is split into applications of CFG rules. These methods have relied on the structure of the target problem, namely lattices or trees, and cannot be applied to graph structures including typed feature structures. This article proposes the feature forest model as a solution to the problem of probabilistic modeling of complex data structures including typed feature structures. The feature forest model provides a method for probabilistic modeling without the independence assumption when probabilistic events are represented with feature forests. Feature forests are generic data structures that represent ambiguous trees in a packed forest structure. Feature forest models are maximum entropy models defined over feature forests. A dynamic programming algorithm is proposed for maximum entropy estimation without unpacking feature forests. Thus probabilistic modeling of any data structures is possible when they are represented by feature forests. This article also describes methods for representing HPSG syntactic structures and predicate-argument structures with feature forests. Hence, we describe a complete strategy for developing probabilistic models for HPSG parsing. The effectiveness of the proposed methods is empirically evaluated through parsing experiments on the Penn Treebank, and the promise of applicability to parsing of real-world sentences is discussed. Yusuke Miyao, Jun'ichi Tsujii |
Comput. Linguistics | 1 |
| 2007 | HPSG Parsing with Shallow Dependency Constraints
Kenji Sagae, Yusuke Miyao, Jun'ichi Tsujii |
ACL | 2 |
| 2007 | Efficient HPSG Parsing with Supertagging and CFG-Filtering
Takuya Matsuzaki, Yusuke Miyao, Jun'ichi Tsujii |
IJCAI | 2 |
| 2007 | Ambiguous Part-of-Speech Tagging for Improving Accuracy and Domain Portability of Syntactic Parsers
Kazuhiro Yoshida, Yoshimasa Tsuruoka, Yusuke Miyao, Jun'ichi Tsujii |
IJCAI | 3 |
| 2006 | Semantic Retrieval for the Accurate Identification of Relational Concepts in Massive TextbasesabstractThis paper introduces a novel framework for the accurate retrieval of relational concepts from huge texts. Prior to retrieval, all sentences are annotated with predicate argument structures and ontological identifiers by applying a deep parser and a term recognizer. During the run time, user requests are converted into queries of region algebra on these annotations. Structural matching with pre-computed semantic annotations establishes the accurate and efficient retrieval of relational concepts. This framework was applied to a text retrieval system for MEDLINE. Experiments on the retrieval of biomedical correlations revealed that the cost is sufficiently small for real-time applications and that the retrieval precision is significantly improved. Yusuke Miyao, Tomoko Ohta, Katsuya Masuda, Yoshimasa Tsuruoka, Kazuhiro Yoshida, Takashi Ninomiya, Jun'ichi Tsujii |
ACL | 1 |
| 2006 | An Intelligent Search Engine and GUI-based Efficient MEDLINE Search Tool Based on Deep Syntactic ParsingabstractWe present a practical HPSG parser for English, an intelligent search engine to retrieve MEDLINE abstracts that represent biomedical events and an efficient MED-LINE search tool helping users to find information about biomedical entities such as genes, proteins, and the interactions between them. Tomoko Ohta, Yusuke Miyao, Takashi Ninomiya, Yoshimasa Tsuruoka, Akane Yakushiji, Katsuya Masuda, Jumpei Takeuchi, Kazuhiro Yoshida, Tadayoshi Hara, Jin-Dong Kim, Yuka Tateisi, Jun'ichi Tsujii |
ACL | 2 |
| 2006 | Improving the Scalability of Semi-Markov Conditional Random Fields for Named Entity RecognitionabstractThis paper presents techniques to apply semi-CRFs to Named Entity Recognition tasks with a tractable computational cost. Our framework can handle an NER task that has long named entities and many labels which increase the computational cost. To reduce the computational cost, we propose two techniques: the first is the use of feature forests, which enables us to pack feature-equivalent states, and the second is the introduction of a filtering process which significantly reduces the number of candidate states. This framework allows us to use a rich set of features extracted from the chunk-based representation that can capture informative characteristics of entities. We also introduce a simple trick to transfer information about distant entities by embedding label information into non-entity labels. Experimental results show that our model achieves an F-score of 71.48% on the JNLPBA 2004 shared task without using any external resources or post-processing techniques. Daisuke Okanohara, Yusuke Miyao, Yoshimasa Tsuruoka, Jun'ichi Tsujii |
ACL | 2 |
| 2006 | Translating HPSG-Style Outputs of a Robust Parser into Typed Dynamic Logic
Manabu Sato, Daisuke Bekki, Yusuke Miyao, Jun'ichi Tsujii |
ACL | 3 |
| 2006 | Trimming CFG Parse Trees for Sentence Compression Using Machine Learning Approaches
Yuya Unno, Takashi Ninomiya, Yusuke Miyao, Jun'ichi Tsujii |
ACL | 3 |
| 2006 | Extremely Lexicalized Models for Accurate and Fast HPSG Parsing
Takashi Ninomiya, Takuya Matsuzaki, Yoshimasa Tsuruoka, Yusuke Miyao, Jun'ichi Tsujii |
EMNLP | 4 |
| 2006 | Automatic Construction of Predicate-argument Structure Patterns for Biomedical Information Extraction
Akane Yakushiji, Yusuke Miyao, Tomoko Ohta, Yuka Tateisi, Jun'ichi Tsujii |
EMNLP | 2 |
| 2005 | Probabilistic CFG with Latent AnnotationsabstractThis paper defines a generative probabilistic model of parse trees, which we call PCFG-LA. This model is an extension of PCFG in which non-terminal symbols are augmented with latent variables. Fine-grained CFG rules are automatically induced from a parsed corpus by training a PCFG-LA model using an EM-algorithm. Because exact parsing with a PCFG-LA is NP-hard, several approximations are described and empirically compared. In experiments using the Penn WSJ corpus, our automatically trained model gave a performance of 86.6% (F1, sentences ≤ 40 words), which is comparable to that of an unlexicalized PCFG parser created using extensive manual feature selection. Takuya Matsuzaki, Yusuke Miyao, Jun'ichi Tsujii |
ACL | 2 |
| 2005 | Probabilistic Disambiguation Models for Wide-Coverage HPSG ParsingabstractThis paper reports the development of log-linear models for the disambiguation in wide-coverage HPSG parsing. The estimation of log-linear models requires high computational cost, especially with wide-coverage grammars. Using techniques to reduce the estimation cost, we trained the models using 20 sections of Penn Tree-bank. A series of experiments empirically evaluated the estimation techniques, and also examined the performance of the disambiguation models on the parsing of real-world sentences. Yusuke Miyao, Jun'ichi Tsujii |
ACL | 1 |
| 2005 | Adapting a Probabilistic Disambiguation Model of an HPSG Parser to a New Domain
Tadayoshi Hara, Yusuke Miyao, Jun'ichi Tsujii |
IJCNLP | 2 |
| 2004 | Deep Linguistic Analysis for the Accurate Identification of Predicate-Argument Relations
Yusuke Miyao, Jun'ichi Tsujii |
COLING | 1 |
| 2004 | Corpus-Oriented Grammar Development for Acquiring a Head-Driven Phrase Structure Grammar from the Penn Treebank
Yusuke Miyao, Takashi Ninomiya, Jun'ichi Tsujii |
IJCNLP | 1 |
| 2004 | A Persistent Feature-Object Database for Intelligent Text Archive Systems
Takashi Ninomiya, Jun'ichi Tsujii, Yusuke Miyao |
IJCNLP | 3 |
| 2003 | An efficient clustering algorithm for class-based language models
Takuya Matsuzaki, Yusuke Miyao, Jun'ichi Tsujii |
CoNLL | 2 |
| 2003 | A model of syntactic disambiguation based on lexicalized grammars
Yusuke Miyao, Jun'ichi Tsujii |
CoNLL | 1 |
| 2003 | Lexicalized Grammar Acquisition
Yusuke Miyao, Takashi Ninomiya, Jun'ichi Tsujii |
EACL | 1 |
| 2003 | A Robust Retrieval Engine for Proximal and Structural Search
Katsuya Masuda, Takashi Ninomiya, Yusuke Miyao, Tomoko Ohta, Jun'ichi Tsujii |
HLT-NAACL | 3 |
| 2002 | Lenient Default Unification for Robust Processing within Unification Based Grammar Formalisms
Takashi Ninomiya, Yusuke Miyao, Jun'ichi Tsujii |
COLING | 2 |
| 2000 | The LiLFeS Abstract Machine and its evaluation with the LinGO grammar
Yusuke Miyao, Takaki Makino, Kentaro Torisawa, Jun'ichi Tsujii |
Nat. Lang. Eng. | 1 |
| 2000 | An HPSG parser with CFG filtering
Kentaro Torisawa, Kenji Nishida, Yusuke Miyao, Jun'ichi Tsujii |
Nat. Lang. Eng. | 3 |
| 1999 | Packing of Feature Structures for Efficient Unification of Disjunctive Feature StructuresabstractThis paper proposes a method for packing feature structures, which automatically collapses equivalent parts of lexical/phrasal feature structures of HPSG into a single packed feature structure. This method avoids redundant repetition of unification of those parts. Preliminary experiments show that this method can significantly improve a unification speed in parsing. Yusuke Miyao |
ACL | 1 |