Jun Suzuki 0001

dblp:78/6923 · DBLP profile ↗
← Back
72ranked-venue papers
14as first author
24since 2021 · last 2026
0000-0003-2108-1340ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 64 · 13 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Vulnerability of LLM Outputs to Heuristics-Inducing Prompt Structures
abstract
Large Language Models (LLMs) have become indispensable tools in daily life. Although LLM applications have rapidly expanded across various domains, distorted outputs (specifically, bias and hallucination) are unresolved problems that threaten the reliability of LLM-based artificial intelligence (AI) agents. Focusing on internal mechanisms and social biases, prior research has rarely considered the possibility of distortion-induction even from non-malicious, ordinary prompts with specific input patterns. This study empirically investigates whether LLMs exposed to certain prompt structures can induce biases commonly elicited in humans, namely, representativeness heuristics, anchoring heuristics, and framing heuristics. To this end, we constructed a test set that triggers one of the three heuristics and evaluated the outputs of state-of-the-art LLMs. We further examined the effectiveness of prompt engineering and debiasing interventions. The LLMs continued to produce heuristic-derived biased outputs under certain prompt conditions. Anchoring heuristics were observed at rates significantly above chance, whereas the representativeness and framing heuristics depended on the model and prompt structure. Debiasing interventions notably reduced the representativeness heuristics but exerted limited impact on anchoring and framing heuristics. This study highlights the need for enhanced awareness of vulnerabilities in LLM outputs against particular prompts. It also reveals that typical prompt-engineering strategies offer insufficient protection against such prompt structures. These results will contribute to the safe and effective use of LLMs in human–computer interactions and AI deployment.
Toshiki Kuramoto, Ryohei Kamei, Jun Suzuki 0001
IUI3
2025 MQM-Chat: Multidimensional Quality Metrics for Chat Translation
abstract
The complexities of chats, such as the stylized contents specific to source segments and dialogue consistency, pose significant challenges for machine translation. Recognizing the need for a precise evaluation metric to address the issues associated with chat translation, this study introduces Multidimensional Quality Metrics for Chat Translation (MQM-Chat), which encompasses seven error types, including three specifically designed for chat translations: ambiguity and disambiguation, buzzword or loanword issues, and dialogue inconsistency. In this study, human annotations were applied to the translations of chat data generated by five translation models. Based on the error distribution of MQM-Chat and the performance of relabeling errors into chat-specific types, we concluded that MQM-Chat effectively classified the errors while highlighting chat-specific issues explicitly. The results demonstrate that MQM-Chat can qualify both the lexical accuracy and semantical accuracy of translation models in chat translation tasks.
Yunmeng Li, Jun Suzuki 0001, Makoto Morishita, Kaori Abe, Kentaro Inui
COLING2
2025 Evaluating Model Alignment with Human Perception: A Study on Shitsukan in LLMs and LVLMs
abstract
We evaluate the alignment of large language models (LLMs) and large vision-language models (LVLMs) with human perception, focusing on the Japanese concept of shitsukan, which reflects the sensory experience of perceiving objects. We created a dataset of shitsukan terms elicited from individuals in response to object images. With it, we designed benchmark tasks for three dimensions of understanding shitsukan: (1) accurate perception in object images, (2) commonsense knowledge of typical shitsukan terms for objects, and (3) distinction of valid shitsukan terms. Models demonstrated mixed accuracy across benchmark tasks, with limited overlap between model- and human-generated terms. However, manual evaluations revealed that the model-generated terms were still natural to humans. This work identifies gaps in culture-specific understanding and contributes to aligning models with human sensory perception. We publicly release the dataset to encourage further research in this area.
Daiki Shiono, Ana Brassard, Yukiko Ishizuki, Jun Suzuki 0001
COLING4
2025 VDocRAG: Retrieval-Augmented Generation over Visually-Rich Documents
abstract
We aim to develop a retrieval-augmented generation (RAG) framework that answers questions over a corpus of visually- rich documents presented in mixed modalities (e.g., charts, tables) and diverse formats (e.g., PDF, PPTX). In this paper, we introduce a new RAG framework, VDocRAG, which can directly understand varied documents and modalities in a unified image format to prevent missing information that occurs by parsing documents to obtain text. To improve the performance, we propose novel self-supervised pre-training tasks that adapt large vision-language models for retrieval by compressing visual information into dense token representations while aligning them with textual content in documents. Furthermore, we introduce OpenDocVQA, the first unified collection of open-domain document visual question answering datasets, encompassing diverse document types and formats. OpenDocVQA provides a comprehensive resource for training and evaluating retrieval and question answering models on visually-rich documents in an open-domain setting. Experiments show that VDocRAG substantially outperforms conventional text-based RAG and has strong generalization capability, highlighting the potential of an effective RAG paradigm for real-world documents.
Ryota Tanaka, Taichi Iki, Taku Hasegawa, Kyosuke Nishida, Kuniko Saito, Jun Suzuki 0001
CVPR6
2025 Drop-Upcycling: Training Sparse Mixture of Experts with Partial Re-initialization
abstract
The Mixture of Experts (MoE) architecture reduces the training and inference cost significantly compared to a dense model of equivalent capacity. Upcycling is an approach that initializes and trains an MoE model using a pre-trained dense model. While upcycling leads to initial performance gains, the training progresses slower than when trained from scratch, leading to suboptimal performance in the long term. We propose Drop-Upcycling - a method that effectively addresses this problem. Drop-Upcycling combines two seemingly contradictory approaches: utilizing the knowledge of pre-trained dense models while statistically re-initializing some parts of the weights. This approach strategically promotes expert specialization, significantly enhancing the MoE model's efficiency in knowledge acquisition. Extensive large-scale experiments demonstrate that Drop-Upcycling significantly outperforms previous MoE construction methods in the long term, specifically when training on hundreds of billions of tokens or more. As a result, our MoE model with 5.9B active parameters achieves comparable performance to a 13B dense model in the same model family, while requiring approximately 1/4 of the training FLOPs. All experimental resources, including source code, training data, model checkpoints and logs, are publicly available to promote reproducibility and future research on MoE.
Taishi Nakamura, Takuya Akiba, Kazuki Fujii, Yusuke Oda, Rio Yokota, Jun Suzuki 0001
ICLR6
2025 Transformer Key-Value Memories Are Nearly as Interpretable as Sparse Autoencoders
abstract
Recent interpretability work on large language models (LLMs) has been increasingly dominated by a feature-discovery approach with the help of proxy modules. Then, the quality of features learned by, e.g., sparse auto-encoders (SAEs), is evaluated. This paradigm naturally raises a critical question: do such learned features have better properties than those already represented within the original model parameters, and unfortunately, only a few studies have made such comparisons systematically so far. In this work, we revisit the interpretability of feature vectors stored in feed-forward (FF) layers, given the perspective of FF as key-value memories, with modern interpretability benchmarks. Our extensive evaluation revealed that SAE and FFs exhibits a similar range of interpretability, although SAEs displayed an observable but minimal improvement in some aspects. Furthermore, in certain aspects, surprisingly, even vanilla FFs yielded better interpretability than the SAEs, and features discovered in SAEs and FFs diverged. These bring questions about the advantage of SAEs from both perspectives of feature quality and faithfulness, compared to directly interpreting FF feature vectors, and FF key-value parameters serve as a strong baseline in modern interpretability research.
Mengyu Ye, Jun Suzuki 0001, Tatsuro Inaba, Tatsuki Kuribayashi
NeurIPS2
2025 Understanding Cross-Lingual Generalization of English-Centric LLMs: The Role of Representation Similarity and Data Exposure
Suchun Xie, Shota Sasaki, Hwichan Kim, Yunmeng Li, Reina Akama, Jun Suzuki 0001
PRICAI6
2025 How Stylistic Similarity Shapes Preferences in Dialogue Dataset with User and Third Party Evaluations
abstract
Recent advancements in dialogue generation have broadened the scope of human–bot interactions, enabling not only contextually appropriate responses but also the analysis of human affect and sensitivity. While prior work has suggested that stylistic similarity between user and system may enhance user impressions, the distinction between subjective and objective similarity is often overlooked. To investigate this issue, we introduce a novel dataset that includes users’ preferences, subjective stylistic similarity based on users’ own perceptions, and objective stylistic similarity annotated by third party evaluators in open-domain dialogue settings. Analysis using the constructed dataset reveals a strong positive correlation between subjective stylistic similarity and user preference. Furthermore, our analysis suggests an important finding: users’ subjective stylistic similarity differs from third party objective similarity. This underscores the importance of distinguishing between subjective and objective evaluations and understanding the distinct aspects each captures when analyzing the relationship between stylistic similarity and user preferences. The dataset presented in this paper is available online.
Ikumi Numaya, Shoji Moriya, Shiki Sato, Reina Akama, Jun Suzuki 0001
SIGDIAL5
2024 InstructDoc: A Dataset for Zero-Shot Generalization of Visual Document Understanding with Instructions
abstract
We study the problem of completing various visual document understanding (VDU) tasks, e.g., question answering and information extraction, on real-world documents through human-written instructions. To this end, we propose InstructDoc, the first large-scale collection of 30 publicly available VDU datasets, each with diverse instructions in a unified format, which covers a wide range of 12 tasks and includes open document types/formats. Furthermore, to enhance the generalization performance on VDU tasks, we design a new instruction-based document reading and understanding model, InstructDr, that connects document images, image encoders, and large language models (LLMs) through a trainable bridging module. Experiments demonstrate that InstructDr can effectively adapt to new VDU datasets, tasks, and domains via given instructions and outperforms existing multimodal LLMs and ChatGPT without specific training.
Ryota Tanaka, Taichi Iki, Kyosuke Nishida, Kuniko Saito, Jun Suzuki 0001
AAAI5
2023 Refactoring Programs Using Large Language Models with Few-Shot Examples
abstract
A less complex and more straightforward program is a crucial factor that enhances its maintainability and makes writing secure and bug-free programs easier. However, due to its heavy workload and the risks of breaking the working programs, programmers are reluctant to do code refactoring, and thus, it also causes the loss of potential learning experiences. To mitigate this, we demonstrate the application of using a large language model (LLM), GPT-3.5, to suggest less complex versions of the user-written Python program, aiming to encourage users to learn how to write better programs. We propose a method to leverage the prompting with few-shot examples of the LLM by selecting the best-suited code refactoring examples for each target programming problem based on the prior evaluation of prompting with the one-shot example. The quantitative evaluation shows that 95.68% of programs can be refactored by generating 10 candidates each, resulting in a 17.35% reduction in the average cyclomatic complexity and a 25.84% decrease in the average number of lines after filtering only generated programs that are semantically correct. Further-more, the qualitative evaluation shows outstanding capability in code formatting, while unnecessary behaviors such as deleting or translating comments are also observed.
Atsushi Shirafuji, Yusuke Oda, Jun Suzuki 0001, Makoto Morishita, Yutaka Watanobe
APSEC3
2023 A Challenging Multimodal Video Summary: Simultaneously Extracting and Generating Keyframe-Caption Pairs from Video
abstract
This paper proposes a practical multimodal video summarization task setting and a dataset to train and evaluate the task.The target task involves summarizing a given video into a predefined number of keyframe-caption pairs and displaying them in a listable format to grasp the video content quickly.This task aims to extract crucial scenes from the video in the form of images (keyframes) and generate corresponding captions explaining each keyframe's situation.This task is useful as a practical application and presents a highly challenging problem worthy of study.Specifically, achieving simultaneous optimization of the keyframe selection performance and caption quality necessitates careful consideration of the mutual dependence on both preceding and subsequent keyframes and captions.To facilitate subsequent research in this field, we also construct a dataset by expanding upon existing datasets and propose an evaluation framework.Furthermore, we develop two baseline systems and report their respective performance.1 Keyframe 4 Keyframe 3 Keyframe 2 Keyframe 1 Keyframe 4 Keyframe 3 Keyframe 2 Keyframe 1 InstructBLIP (few-shot)
Keito Kudo, Haruki Nagasawa, Jun Suzuki 0001, Nobuyuki Shimizu
EMNLP3
2023 Assessing Step-by-Step Reasoning against Lexical Negation: A Case Study on Syllogism
abstract
Large language models (LLMs) take advantage of step-by-step reasoning instructions, e.g., chain-of-thought (CoT) prompting.Building on this, their ability to perform CoT-style reasoning robustly is of interest from a probing perspective.In this study, we inspect the stepby-step reasoning ability of LLMs with a focus on negation, which is a core linguistic phenomenon that is difficult to process.In particular, we introduce several controlled settings (e.g., reasoning on fictional entities) to evaluate the logical reasoning abilities of the models.We observed that dozens of modern LLMs were not robust against lexical negation (e.g., plausi-ble→implausible) when performing CoT-style reasoning, and the results highlight unique limitations in each LLM family.https://github.com/muyo8692/ stepbystep-reasoning-vs-negation Setting Few-shot exemplars Target example If fails at this setting BASE Is a sentence "A does B" plausible?A is a C player.B happens in C/X.So the answer is yes/no.Is a sentence "D does E" plausible?D is a F player.E happens in F/Y.So the answer is __ CoT-style reasoning fails. FIC Is a sentence "A does B" plausible?A is a C player.B happens in C/X.So the answer is yes/no.Is a sentence "α does β" plausible?α is a γ player.β happens in γ/χ.So the answer is __ Reasoning cannot be abstracted to fictional texts. FICNEG Is a sentence "A does B" implausible?A is a C player.B happens in C/X.So the answer is yes/no.Is a sentence "α does β" implausible?α is a γ player.β happens in γ/χ.So the answer is __ Abstract CoT-style reasoning is only achieved on the affirmative domain. FICNEG-O Is a sentence "A does B" plausible?A is a C player.B happens in C/X.So the answer is yes/no.Is a sentence "α does β" implausible?α is a γ player.β happens in γ/χ.So the answer is __ Model cannot handle domain shift in terms of negation.Table 1: General task format in each setting.Few-shot exemplars are first shown to a model, and then the model answers to the target example given its question and reasoning chain.Symbols, e.g., A and α, are replaced with certain real or fictional entities in the actual input.The REAL setting indicates that the entity choices reflect the factual reality, and FIC.indicates that the entity choices do not reflect factual reality, e.g., Is "Judy Tate was safe at first." plausible?Judy Tate is a turboglide player.Getting out at first happens in turboglide.So the answer is yes.Refer to Appendix A for the exact input. Format: SyllogismWe evaluated the LLMs' ability to judge the validity of particular types of syllogisms.Here, we utilized three settings to ensure the robustness of the results (Section 3); however, we consider the following SPORTS TASK (SP) format as an example to explain the settings.The base format of the syllogism is as follows:Premise1: PERSON is a SPORT player.Premise2: ACTION happens in the SPORT. Conclusion: PERSON does ACTION.Here, the transitivity of reasoning (A>B, B>C, then A>C) is targeted. C.3 Answer Distribution
Mengyu Ye, Tatsuki Kuribayashi, Jun Suzuki 0001, Goro Kobayashi, Hiroaki Funayama
EMNLP3
2023 Use of an AI-powered Rewriting Support Software in Context with Other Tools: A Study of Non-Native English Speakers
abstract
Academic writing in English can be challenging for non-native English speakers (NNESs). AI-powered rewriting tools can potentially improve NNESs’ writing outcomes at a low cost. However, whether and how NNESs make valid assessments of the revisions provided by these algorithmic recommendations remains unclear. We report a study where NNESs leverage an AI-powered rewriting tool, Langsmith, to polish their drafted academic essays. We examined the participants’ interactions with the tool via user studies and interviews. Our data reveal that most participants used Langsmith in combination with other tools, such as machine translation (MT), and those who used MT had different ways of understanding and evaluating Langsmith’s suggestions than those who did not. Based on these findings, we assert that NNESs’ quality assessment in AI-powered rewriting tools is influenced by the simultaneous use of multiple tools, offering valuable insights into the design of future rewriting tools for NNESs.
Takumi Ito, Naomi Yamashita, Tatsuki Kuribayashi, Masatoshi Hidaka, Jun Suzuki 0001, Ge Gao 0001, Jack Jamieson, Kentaro Inui
UIST5
2023 Examining the effect of whitening on static and contextualized word embeddings
abstract
Static word embeddings (SWE) and contextualized word embeddings (CWE) are the foundation of modern natural language processing. However, these embeddings suffer from spatial bias in the form of anisotropy, which has been demonstrated to reduce their performance. A method to alleviate the anisotropy is the “whitening” transformation. Whitening is a standard method in signal processing and other areas, however, its effect on SWE and CWE is not well understood. In this study, we conduct an experiment to elucidate the effect of whitening on SWE and CWE. The results indicate that whitening predominantly removes the word frequency bias in SWE, and biases other than the word frequency bias in CWE.
Shota Sasaki, Benjamin Heinzerling, Jun Suzuki 0001, Kentaro Inui
Inf. Process. Manag.3
2023 Extracting representative subset from extensive text data for training pre-trained language models
abstract
This paper investigates the existence of a representative subset obtained from a large original dataset that can achieve the same performance level obtained using the entire dataset in the context of training neural language models. We employ the likelihood-based scoring method based on two distinct types of pre-trained language models to select a representative subset. We conduct our experiments on widely used 17 natural language processing datasets with 24 evaluation metrics. The experimental results showed that the representative subset obtained using the likelihood difference score can achieve the 90% performance level even when the size of the dataset is reduced to approximately two to three orders of magnitude smaller than the original dataset. We also compare the performance with the models trained with the same amount of subset selected randomly to show the effectiveness of the representative subset.
Jun Suzuki 0001, Heiga Zen, Hideto Kazawa
Inf. Process. Manag.1
2022 Balancing Cost and Quality: An Exploration of Human-in-the-Loop Frameworks for Automated Short Answer Scoring
Hiroaki Funayama, Tasuku Sato, Yuichiroh Matsubayashi, Tomoya Mizumoto, Jun Suzuki 0001, Kentaro Inui
AIED (1)5
2022 Target-Guided Open-Domain Conversation Planning
abstract
Prior studies addressing target-oriented conversational tasks lack a crucial notion that has been intensively studied in the context of goal-oriented artificial intelligence agents, namely, planning. In this study, we propose the task of Target-Guided Open-Domain Conversation Planning (TGCP) task to evaluate whether neural conversational agents have goal-oriented conversation planning abilities. Using the TGCP task, we investigate the conversation planning abilities of existing retrieval models and recent strong generative models. The experimental results reveal the challenges facing current technology.
Yosuke Kishinami, Reina Akama, Shiki Sato, Ryoko Tokuhisa, Jun Suzuki 0001, Kentaro Inui
COLING5
2022 JParaCrawl v3.0: A Large-scale English-Japanese Parallel Corpus
abstract
Most current machine translation models are mainly trained with parallel corpora, and their translation accuracy largely depends on the quality and quantity of the corpora. Although there are billions of parallel sentences for a few language pairs, effectively dealing with most language pairs is difficult due to a lack of publicly available parallel corpora. This paper creates a large parallel corpus for English-Japanese, a language pair for which only limited resources are available, compared to such resource-rich languages as English-German. It introduces a new web-based English-Japanese parallel corpus named JParaCrawl v3.0. Our new corpus contains more than 21 million unique parallel sentence pairs, which is more than twice as many as the previous JParaCrawl v2.0 corpus. Through experiments, we empirically show how our new corpus boosts the accuracy of machine translation models on various domains. The JParaCrawl v3.0 corpus will eventually be publicly available online for research purposes.
Makoto Morishita, Katsuki Chousa, Jun Suzuki 0001, Masaaki Nagata
LREC3
2022 N-best Response-based Analysis of Contradiction-awareness in Neural Response Generation Models
abstract
Avoiding the generation of responses that contradict the preceding context is a significant challenge in dialogue response generation.One feasible method is post-processing, such as filtering out contradicting responses from a resulting n-best response list.In this scenario, the quality of the n-best list considerably affects the occurrence of contradictions because the final response is chosen from this n-best list.This study quantitatively analyzes the contextual contradiction-awareness of neural response generation models using the consistency of the n-best lists.Particularly, we used polar questions as stimulus inputs for concise and quantitative analyses.Our tests illustrate the contradiction-awareness of recent neural response generation models and methodologies, followed by a discussion of their properties and limitations.
Shiki Sato, Reina Akama, Hiroki Ouchi, Ryoko Tokuhisa, Jun Suzuki 0001, Kentaro Inui
SIGDIAL5
2022 Prompt Sensitivity of Language Model for Solving Programming Problems
abstract
A popular language model that can solve introductory programming problems, OpenAI’s Codex, has drawn much attention not only in the natural language processing field but also in the software engineering field. It supports programmers by suggesting the next tokens to write, and it can even generate a whole function definition from a document string. We focus on its capability of automatically solving programming problems through code generation from problem descriptions. We investigate the model’s sensitivity to problem descriptions by formatting and modifying them. The experimental results show that the more explicitly formatted problem description enhances the code generation performance from 30.9% (raw) to 39.9% (formatted). Additionally, we observe that code generation relies on information specified in the problem description, such as variable names and constant values, as anonymizing them reduces the performance significantly. Moreover, statistical biases in code generation are identified, such as the generated programs ignoring the problem modification and answering the exact opposite problem. The changes in accuracy across formats suggest that the model does not correctly understand the natural language explaining the problem specification even if the model could solve the programming problems with high accuracy.
Atsushi Shirafuji, Takumi Ito, Makoto Morishita, Yuki Nakamura, Yusuke Oda, Jun Suzuki 0001, Yutaka Watanobe
SoMeT6
2021 Context-aware Neural Machine Translation with Mini-batch Embedding
abstract
It is crucial to provide an inter-sentence context in Neural Machine Translation (NMT) models for higher-quality translation.With the aim of using a simple approach to incorporate inter-sentence information, we propose minibatch embedding (MBE) as a way to represent the features of sentences in a mini-batch.We construct a mini-batch by choosing sentences from the same document, and thus the MBE is expected to have contextual information across sentences.Here, we incorporate MBE in an NMT model, and our experiments show that the proposed method consistently outperforms the translation capabilities of strong baselines and improves writing style or terminology to fit the document's context. 1
Makoto Morishita, Jun Suzuki 0001, Tomoharu Iwata, Masaaki Nagata
EACL2
2021 SHAPE : Shifted Absolute Position Embedding for Transformers
abstract
Position representation is crucial for building position-aware representations in Transformers.Existing position representations suffer from a lack of generalization to test data with unseen lengths or high computational cost.We investigate shifted absolute position embedding (SHAPE) to address both issues.The basic idea of SHAPE is to achieve shift invariance, which is a key property of recent successful position representations, by randomly shifting absolute positions during training.We demonstrate that SHAPE is empirically comparable to its counterpart while being simpler and faster 1 .
Shun Kiyono, Sosuke Kobayashi, Jun Suzuki 0001, Kentaro Inui
EMNLP (1)3
2021 Instance-Based Neural Dependency Parsing
abstract
Abstract Interpretable rationales for model predictions are crucial in practical applications. We develop neural models that possess an interpretable inference process for dependency parsing. Our models adopt instance-based inference, where dependency edges are extracted and labeled by comparing them to edges in a training set. The training edges are explicitly used for the predictions; thus, it is easy to grasp the contribution of each edge to the predictions. Our experiments show that our instance-based models achieve competitive accuracy with standard neural models and have the reasonable plausibility of instance-based explanations.
Hiroki Ouchi, Jun Suzuki 0001, Sosuke Kobayashi, Sho Yokoi, Tatsuki Kuribayashi, Masashi Yoshikawa, Kentaro Inui
Trans. Assoc. Comput. Linguistics2
2021 Subword-Based Compact Reconstruction for Open-Vocabulary Neural Word Embeddings
abstract
The methodology of neural word embeddings has become an important fundamental resource for tackling many applications in the artificial intelligence (AI) research field. They have successfully been proven to capture high-quality syntactic and semantic relationships in a vector space. Despite their significant impact, neural word embeddings have several disadvantages. In this paper, we focus on two issues regarding well-trained word embeddings: (i) the massive memory requirement and (ii) the inapplicability of out-of-vocabulary (OOV) words. To overcome these two issues, we propose a method of reconstructing pre-trained word embeddings by using subword information that can effectively represent a large number of subword embeddings in a considerably small fixed space while preventing quality degradation from the original word embeddings. The key techniques of our method are twofold: memory-shared embeddings and a variant of the key-value-query self-attention mechanism. Our experiments show that our reconstructed subword-based word embeddings can successfully imitate well-trained word embeddings in a small fixed space while preventing quality degradation across several linguistic benchmark datasets and can simultaneously predict effective embeddings of OOV words. We also demonstrate the effectiveness of our reconstruction method when it is applied to downstream tasks, such as named entity recognition and natural language inference tasks.
Shota Sasaki, Jun Suzuki 0001, Kentaro Inui
IEEE ACM Trans. Audio Speech Lang. Process.2
2020 Encoder-Decoder Models Can Benefit from Pre-trained Masked Language Models in Grammatical Error Correction
abstract
This paper investigates how to effectively incorporate a pre-trained masked language model (MLM), such as BERT, into an encoderdecoder (EncDec) model for grammatical error correction (GEC).The answer to this question is not as straightforward as one might expect because the previous common methods for incorporating a MLM into an EncDec model have potential drawbacks when applied to GEC.For example, the distribution of the inputs to a GEC model can be considerably different (erroneous, clumsy, etc.) from that of the corpora used for pre-training MLMs; however, this issue is not addressed in the previous methods.Our experiments show that our proposed method, where we first fine-tune a MLM with a given GEC corpus and then use the output of the finetuned MLM as additional features in the GEC model, maximizes the benefit of the MLM.The best-performing model achieves state-ofthe-art performances on the BEA-2019 and CoNLL-2014 benchmarks.Our code is publicly available at: https://github.com/ kanekomasahiro/bert-gec.
Masahiro Kaneko, Masato Mita, Shun Kiyono, Jun Suzuki 0001, Kentaro Inui
ACL4
2020 Language Models as an Alternative Evaluator of Word Order Hypotheses: A Case Study in Japanese
abstract
We examine a methodology using neural language models (LMs) for analyzing the word order of language.This LM-based method has the potential to overcome the difficulties existing methods face, such as the propagation of preprocessor errors in count-based methods.In this study, we explore whether the LMbased method is valid for analyzing the word order.As a case study, this study focuses on Japanese due to its complex and flexible word order.To validate the LM-based method, we test (i) parallels between LMs and human word order preference, and (ii) consistency of the results obtained using the LM-based method with previous linguistic studies.Through our experiments, we tentatively conclude that LMs display sufficient word order knowledge for usage as an analysis tool.Finally, using the LMbased method, we demonstrate the relationship between the canonical word order and topicalization, which had yet to be analyzed by largescale experiments.C Data used in Section 5.2, Section 6, and Appendix F
Tatsuki Kuribayashi, Takumi Ito, Jun Suzuki 0001, Kentaro Inui
ACL3
2020 Single Model Ensemble using Pseudo-Tags and Distinct Vectors
abstract
Model ensemble techniques often increase task performance in neural networks; however, they require increased time, memory, and management effort.In this study, we propose a novel method that replicates the effects of a model ensemble with a single model.Our approach creates K-virtual models within a single parameter space using K-distinct pseudotags and K-distinct vectors.Experiments on text classification and sequence labeling tasks on several datasets demonstrate that our method emulates or outperforms a traditional model ensemble with 1/K-times fewer parameters.
Ryosuke Kuwabara, Jun Suzuki 0001, Hideki Nakayama
ACL2
2020 Instance-Based Learning of Span Representations: A Case Study through Named Entity Recognition
abstract
Interpretable rationales for model predictions play a critical role in practical applications.In this study, we develop models possessing interpretable inference process for structured prediction.Specifically, we present a method of instance-based learning that learns similarities between spans.At inference time, each span is assigned a class label based on its similar spans in the training set, where it is easy to understand how much each training instance contributes to the predictions.Through empirical analysis on named entity recognition, we demonstrate that our method enables to build models that have high interpretability without sacrificing performance.
Hiroki Ouchi, Jun Suzuki 0001, Sosuke Kobayashi, Sho Yokoi, Tatsuki Kuribayashi, Ryuto Konno, Kentaro Inui
ACL2
2020 Evaluating Dialogue Generation Systems via Response Selection
abstract
Existing automatic evaluation metrics for open-domain dialogue response generation systems correlate poorly with human evaluation.We focus on evaluating response generation systems via response selection.To evaluate systems properly via response selection, we propose a method to construct response selection test sets with well-chosen false candidates.Specifically, we propose to construct test sets filtering out some types of false candidates: (i) those unrelated to the ground-truth response and (ii) those acceptable as appropriate responses.Through experiments, we demonstrate that evaluating systems via response selection with the test set developed by our method correlates more strongly with human evaluation, compared with widely used automatic evaluation metrics such as BLEU.
Shiki Sato, Reina Akama, Hiroki Ouchi, Jun Suzuki 0001, Kentaro Inui
ACL4
2020 PheMT: A Phenomenon-wise Dataset for Machine Translation Robustness on User-Generated Contents
abstract
Neural Machine Translation (NMT) has shown drastic improvement in its quality when translating clean input, such as text from the news domain.However, existing studies suggest that NMT still struggles with certain kinds of input with considerable noise, such as User-Generated Contents (UGC) on the Internet.To make better use of NMT for cross-cultural communication, one of the most promising directions is to develop a model that correctly handles these expressions.Though its importance has been recognized, it is still not clear as to what creates the great gap in performance between the translation of clean input and that of UGC.To answer the question, we present a new dataset, PheMT, for evaluating the robustness of MT systems against specific linguistic phenomena in Japanese-English translation.Our experiments with the created dataset revealed that not only our in-house models but even widely used off-the-shelf systems are greatly disturbed by the presence of certain phenomena.
Ryo Fujii, Masato Mita, Kaori Abe, Kazuaki Hanawa, Makoto Morishita, Jun Suzuki 0001, Kentaro Inui
COLING6
2020 Filtering Noisy Dialogue Corpora by Connectivity and Content Relatedness
abstract
Large-scale dialogue datasets have recently become available for training neural dialogue agents. However, these datasets have been reported to contain a non-negligible number of unacceptable utterance pairs. In this paper, we propose a method for scoring the quality of utterance pairs in terms of their connectivity and relatedness. The proposed scoring method is designed based on findings widely shared in the dialogue and linguistics research communities. We demonstrate that it has a relatively good correlation with the human judgment of dialogue quality. Furthermore, the method is applied to filter out potentially unacceptable utterance pairs from a large-scale noisy dialogue corpus to ensure its quality. We experimentally confirm that training data filtered by the proposed method improves the quality of neural dialogue agents in response generation.
Reina Akama, Sho Yokoi, Jun Suzuki 0001, Kentaro Inui
EMNLP (1)3
2020 Word Rotator's Distance
abstract
A key principle in assessing textual similarity is measuring the degree of semantic overlap between two texts by considering the word alignment.Such alignment-based approaches are intuitive and interpretable; however, they are empirically inferior to the simple cosine similarity between general-purpose sentence vectors.To address this issue, we focus on and demonstrate the fact that the norm of word vectors is a good proxy for word importance, and their angle is a good proxy for word similarity.Alignment-based approaches do not distinguish them, whereas sentence-vector approaches automatically use the norm as the word importance.Accordingly, we propose a method that first decouples word vectors into their norm and direction, and then computes alignment-based similarity using earth mover's distance (i.e., optimal transport cost), which we refer to as word rotator's distance.Besides, we find how to "grow" the norm and direction of word vectors (vector converter), which is a new systematic approach derived from sentence-vector estimation methods.On several textual similarity datasets, the combination of these simple proposed methods outperformed not only alignment-based approaches but also strong baselines.1
Sho Yokoi, Reina Akama, Jun Suzuki 0001, Kentaro Inui
EMNLP (1)4
2020 JParaCrawl: A Large Scale Web-Based English-Japanese Parallel Corpus
abstract
Recent machine translation algorithms mainly rely on parallel corpora. However, since the availability of parallel corpora remains limited, only some resource-rich language pairs can benefit from them. We constructed a parallel corpus for English-Japanese, for which the amount of publicly available parallel corpora is still limited. We constructed the parallel corpus by broadly crawling the web and automatically aligning parallel sentences. Our collected corpus, called JParaCrawl, amassed over 8.7 million sentence pairs. We show how it includes a broader range of domains and how a neural machine translation model trained with it works as a good pre-trained model for fine-tuning specific domains. The pre-training and fine-tuning approaches achieved or surpassed performance comparable to model training from the initial state and reduced the training time. Additionally, we trained the model with an in-domain dataset and JParaCrawl to show how we achieved the best performance with them. JParaCrawl and the pre-trained models are freely available online for research purposes.
Makoto Morishita, Jun Suzuki 0001, Masaaki Nagata
LREC2
2020 Massive Exploration of Pseudo Data for Grammatical Error Correction
abstract
Collecting a large amount of training data for grammatical error correction (GEC) models has been an ongoing challenge in the field of GEC. Recently, it has become common to use data demanding deep neural models such as an encoder-decoder for GEC; thus, tackling the problem of data collection has become increasingly important. The incorporation of pseudo data in the training of GEC models is one of the main approaches for mitigating the problem of data scarcity. However, a consensus is lacking on experimental configurations, namely, (i) the methods for generating pseudo data, (ii) the seed corpora used as the source of the pseudo data, and (iii) the means of optimizing the model. In this study, these configurations are thoroughly explored through massive amount of experiments, with the aim of providing an improved understanding of pseudo data. Our main experimental finding is that pretraining a model with pseudo data generated by back-translation-based method is the most effective approach. Our findings are supported by the achievement of state-of-the-art performance on multiple benchmark test sets (the CoNLL-2014 test set and the official test set of the BEA-2019 shared task) without requiring any modifications to the model architecture. We also perform an in-depth analysis of our model with respect to the grammatical error type and proficiency level of the text. Finally, we suggest future directions for further improving model performance.
Shun Kiyono, Jun Suzuki 0001, Tomoya Mizumoto, Kentaro Inui
IEEE ACM Trans. Audio Speech Lang. Process.2
2019 Mixture of Expert/Imitator Networks: Scalable Semi-Supervised Learning Framework
abstract
The current success of deep neural networks (DNNs) in an increasingly broad range of tasks involving artificial intelligence strongly depends on the quality and quantity of labeled training data. In general, the scarcity of labeled data, which is often observed in many natural language processing tasks, is one of the most important issues to be addressed. Semisupervised learning (SSL) is a promising approach to overcoming this issue by incorporating a large amount of unlabeled data. In this paper, we propose a novel scalable method of SSL for text classification tasks. The unique property of our method, Mixture of Expert/Imitator Networks, is that imitator networks learn to “imitate” the estimated label distribution of the expert network over the unlabeled data, which potentially contributes a set of features for the classification. Our experiments demonstrate that the proposed method consistently improves the performance of several types of baseline DNNs. We also demonstrate that our method has the more data, better performance property with promising scalability to the amount of unlabeled data.
Shun Kiyono, Jun Suzuki 0001, Kentaro Inui
AAAI2
2019 Character n-Gram Embeddings to Improve RNN Language Models
abstract
This paper proposes a novel Recurrent Neural Network (RNN) language model that takes advantage of character information. We focus on character n-grams based on research in the field of word embedding construction (Wieting et al. 2016). Our proposed method constructs word embeddings from character ngram embeddings and combines them with ordinary word embeddings. We demonstrate that the proposed method achieves the best perplexities on the language modeling datasets: Penn Treebank, WikiText-2, and WikiText-103. Moreover, we conduct experiments on application tasks: machine translation and headline generation. The experimental results indicate that our proposed method also positively affects these tasks
Sho Takase, Jun Suzuki 0001, Masaaki Nagata
AAAI2
2019 An Empirical Study of Span Representations in Argumentation Structure Parsing
abstract
Tatsuki Kuribayashi, Hiroki Ouchi, Naoya Inoue, Paul Reisert, Toshinori Miyoshi, Jun Suzuki, Kentaro Inui. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.
Tatsuki Kuribayashi, Hiroki Ouchi, Naoya Inoue, Paul Reisert, Toshinori Miyoshi, Jun Suzuki 0001, Kentaro Inui
ACL (1)6
2019 Effective Adversarial Regularization for Neural Machine Translation
abstract
A regularization technique based on adversarial perturbation, which was initially developed in the field of image processing, has been successfully applied to text classification tasks and has yielded attractive improvements.We aim to further leverage this promising methodology into more sophisticated and critical neural models in the natural language processing field, i.e., neural machine translation (NMT) models.However, it is not trivial to apply this methodology to such models.Thus, this paper investigates the effectiveness of several possible configurations of applying the adversarial perturbation and reveals that the adversarial regularization technique can significantly and consistently improve the performance of widely used NMT models, such as LSTMbased and Transformer-based models. 1
Motoki Sato, Jun Suzuki 0001, Shun Kiyono
ACL (1)2
2019 An Empirical Study of Incorporating Pseudo Data into Grammatical Error Correction
abstract
Shun Kiyono, Jun Suzuki, Masato Mita, Tomoya Mizumoto, Kentaro Inui. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Shun Kiyono, Jun Suzuki 0001, Masato Mita, Tomoya Mizumoto, Kentaro Inui
EMNLP/IJCNLP (1)2
2019 Transductive Learning of Neural Language Models for Syntactic and Semantic Analysis
abstract
Hiroki Ouchi, Jun Suzuki, Kentaro Inui. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Hiroki Ouchi, Jun Suzuki 0001, Kentaro Inui
EMNLP/IJCNLP (1)2
2019 Select and Attend: Towards Controllable Content Selection in Text Generation
abstract
Xiaoyu Shen, Jun Suzuki, Kentaro Inui, Hui Su, Dietrich Klakow, Satoshi Sekine. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Xiaoyu Shen 0001, Jun Suzuki 0001, Kentaro Inui, Hui Su, Dietrich Klakow, Satoshi Sekine
EMNLP/IJCNLP (1)2
2019 Diamonds in the Rough: Generating Fluent Sentences from Early-Stage Drafts for Academic Writing Assistance
abstract
The writing process consists of several stages such as drafting, revising, editing, and proofreading. Studies on writing assistance, such as grammatical error correction (GEC), have mainly focused on sentence editing and proofreading, where surface-level issues such as typographical errors, spelling errors, or grammatical errors should be corrected. We broaden this focus to include the earlier revising stage, where sentences require adjustment to the information included or major rewriting and propose Sentence-level Revision (SentRev) as a new writing assistance task. Well-performing systems in this task can help inexperienced authors by producing fluent, complete sentences given their rough, incomplete drafts. We build a new freely available crowdsourced evaluation dataset consisting of incomplete sentences authored by non-native writers paired with their final versions extracted from published academic papers for developing and evaluating SentRev models. We also establish baseline performance on SentRev using our newly built evaluation dataset.
Takumi Ito, Tatsuki Kuribayashi, Hayato Kobayashi, Ana Brassard, Masato Hagiwara, Jun Suzuki 0001, Kentaro Inui
INLG6
2018 Improving Neural Machine Translation by Incorporating Hierarchical Subword Features
abstract
This paper focuses on subword-based Neural Machine Translation (NMT). We hypothesize that in the NMT model, the appropriate subword units for the following three modules (layers) can differ: (1) the encoder embedding layer, (2) the decoder embedding layer, and (3) the decoder output layer. We find the subword based on Sennrich et al. (2016) has a feature that a large vocabulary is a superset of a small vocabulary and modify the NMT model enables the incorporation of several different subword units in a single embedding layer. We refer these small subword features as hierarchical subword features. To empirically investigate our assumption, we compare the performance of several different subword units and hierarchical subword features for both the encoder and decoder embedding layers. We confirmed that incorporating hierarchical subword features in the encoder consistently improves BLEU scores on the IWSLT evaluation datasets.
Makoto Morishita, Jun Suzuki 0001, Masaaki Nagata
COLING2
2018 Direct Output Connection for a High-Rank Language Model
abstract
This paper proposes a state-of-the-art recurrent neural network (RNN) language model that combines probability distributions computed not only from a final RNN layer but also from middle layers.Our proposed method raises the expressive power of a language model based on the matrix factorization interpretation of language modeling introduced by Yang et al. (2018).The proposed method improves the current state-of-the-art language model and achieves the best score on the Penn Treebank and WikiText-2, which are the standard benchmark datasets.Moreover, we indicate our proposed method contributes to two application tasks: machine translation and headline generation.
Sho Takase, Jun Suzuki 0001, Masaaki Nagata
EMNLP2
2018 Pointwise HSIC: A Linear-Time Kernelized Co-occurrence Norm for Sparse Linguistic Expressions
abstract
In this paper, we propose a new kernel-based co-occurrence measure that can be applied to sparse linguistic expressions (e.g., sentences) with a very short learning time, as an alternative to pointwise mutual information (PMI).As well as deriving PMI from mutual information, we derive this new measure from the Hilbert-Schmidt independence criterion (HSIC); thus, we call the new measure the pointwise HSIC (PHSIC).PHSIC can be interpreted as a smoothed variant of PMI that allows various similarity metrics (e.g., sentence embeddings) to be plugged in as kernels.Moreover, PHSIC can be estimated by simple and fast (linear in the size of the data) matrix calculations regardless of whether we use linear or nonlinear kernels.Empirically, in a dialogue response selection task, PHSIC is learned thousands of times faster than an RNNbased PMI while outperforming PMI in accuracy.In addition, we also demonstrate that PH-SIC is beneficial as a criterion of a data selection task for machine translation owing to its ability to give high (low) scores to a consistent (inconsistent) pair with other pairs.
Sho Yokoi, Sosuke Kobayashi, Kenji Fukumizu, Jun Suzuki 0001, Kentaro Inui
EMNLP4
2018 Interpretable Adversarial Perturbation in Input Embedding Space for Text
abstract
Following great success in the image processing field, the idea of adversarial training has been applied to tasks in the natural language processing (NLP) field. One promising approach directly applies adversarial training developed in the image processing field to the input word embedding space instead of the discrete input space of texts. However, this approach abandons such interpretability as generating adversarial texts to significantly improve the performance of NLP tasks. This paper restores interpretability to such methods by restricting the directions of perturbations toward the existing words in the input embedding space. As a result, we can straightforwardly reconstruct each input with perturbations to an actual text by considering the perturbations to be the replacement of words in the sentence while maintaining or even improving the task performance.
Motoki Sato, Jun Suzuki 0001, Hiroyuki Shindo, Yuji Matsumoto 0001
IJCAI2
2018 Reducing Odd Generation from Neural Headline Generation
Shun Kiyono, Sho Takase, Jun Suzuki 0001, Naoaki Okazaki, Kentaro Inui, Masaaki Nagata
PACLIC3
2017 Enumeration of Extractive Oracle Summaries
abstract
To analyze the limitations and the future directions of the extractive summarization paradigm, this paper proposes an Integer Linear Programming (ILP) formulation to obtain extractive oracle summaries in terms of ROUGE n .We also propose an algorithm that enumerates all of the oracle summaries for a set of reference summaries to exploit F-measures that evaluate which system summaries contain how many sentences that are extracted as an oracle summary.Our experimental results obtained from Document Understanding Conference (DUC) corpora demonstrated the following: (1) room still exists to improve the performance of extractive summarization; (2) the F-measures derived from the enumerated oracle summaries have significantly stronger correlations with human judgment than those derived from single oracle summaries.
Tsutomu Hirao, Masaaki Nishino, Jun Suzuki 0001, Masaaki Nagata
EACL (1)3
2016 Neural Headline Generation on Abstract Meaning Representation
abstract
Neural network-based encoder-decoder models are among recent attractive methodologies for tackling natural language generation tasks.This paper investigates the usefulness of structural syntactic and semantic information additionally incorporated in a baseline neural attention-based model.We encode results obtained from an abstract meaning representation (AMR) parser using a modified version of Tree-LSTM.Our proposed attention-based AMR encoder-decoder model improves headline generation benchmarks compared with the baseline neural attention-based model.
Sho Takase, Jun Suzuki 0001, Naoaki Okazaki, Tsutomu Hirao, Masaaki Nagata
EMNLP2
2016 Learning Compact Neural Word Embeddings by Parameter Space Sharing
Jun Suzuki 0001, Masaaki Nagata
IJCAI1
2016 Right-truncatable Neural Word Embeddings
abstract
This paper proposes an incremental learning strategy for neural word embedding methods, such as SkipGrams and Global Vectors.Since our method iteratively generates embedding vectors one dimension at a time, obtained vectors equip a unique property.Namely, any right-truncated vector matches the solution of the corresponding lower-dimensional embedding.Therefore, a single embedding vector can manage a wide range of dimensional requirements imposed by many different uses and applications.
Jun Suzuki 0001, Masaaki Nagata
HLT-NAACL1
2015 Summarizing a Document by Trimming the Discourse Tree
abstract
Recent studies on extractive text summarization formulate it as a combinatorial optimization problem, extracting the optimal subset from a set of the textual units that maximizes an objective function without violating the length constraint. Although these methods successfully improve automatic evaluation scores, they do not consider the discourse structure in the source document. Thus, summaries generated by these methods may lack logical coherence. In previous work, we proposed a method that exploits a discourse tree structure to produce coherent summaries. By transforming a traditional discourse tree, namely a rhetorical structure theory-based discourse tree (RST-DT), into a dependency-based discourse tree (DEP-DT), we formulated the summarization procedure as a Tree Knapsack Problem whose tree corresponds to the DEP-DT. This paper extends the work with a detailed discussion of the approach together with a novel efficient dynamic programming algorithm for solving the Tree Knapsack Problem. Experiments show that our method not only achieved the highest score in both automatic and human evaluation, but also obtained good performance in terms of the linguistic qualities of the summaries.
Tsutomu Hirao, Masaaki Nishino, Yasuhisa Yoshida, Jun Suzuki 0001, Norihito Yasuda, Masaaki Nagata
IEEE ACM Trans. Audio Speech Lang. Process.4
2014 Fused Feature Representation Discovery for High-Dimensional and Sparse Data
Jun Suzuki 0001, Masaaki Nagata
AAAI1
2014 Dependency-based Discourse Parser for Single-Document Summarization
abstract
The current state-of-the-art single-document summarization method gen-erates a summary by solving a Tree Knapsack Problem (TKP), which is the problem of finding the optimal rooted sub-tree of the dependency-based discourse tree (DEP-DT) of a document. We can obtain a gold DEP-DT by transforming a gold Rhetorical Structure Theory-based discourse tree (RST-DT). However, there is still a large difference between the ROUGE scores of a system with a gold DEP-DT and a system with a DEP-DT obtained from an automatically parsed RST-DT. To improve the ROUGE score, we propose a novel discourse parser that directly generates the DEP-DT. The evaluation results showed that the TKP with our parser outperformed that with the state-of-the-art RST-DT parser, and achieved almost equivalent ROUGE scores to the TKP with the gold DEP-DT. 1
Yasuhisa Yoshida, Jun Suzuki 0001, Tsutomu Hirao, Masaaki Nagata
EMNLP2
2014 Restructuring output layers of deep neural networks using minimum risk parameter clustering
Yotaro Kubo, Jun Suzuki 0001, Takaaki Hori, Atsushi Nakamura
INTERSPEECH2
2013 Text Summarization while Maximizing Multiple Objectives with Lagrangian Relaxation
Masaaki Nishino, Norihito Yasuda, Tsutomu Hirao, Jun Suzuki 0001, Masaaki Nagata
ECIR4
2013 Shift-Reduce Word Reordering for Machine Translation
abstract
This paper presents a novel word reordering model that employs a shift-reduce parser for inversion transduction grammars.Our model uses rich syntax parsing features for word reordering and runs in linear time.We apply it to postordering of phrase-based machine translation (PBMT) for Japanese-to-English patent tasks.Our experimental results show that our method achieves a significant improvement of +3.1 BLEU scores against 30.15BLEU scores of the baseline PBMT system.
Katsuhiko Hayashi 0001, Katsuhito Sudoh, Hajime Tsukada, Jun Suzuki 0001, Masaaki Nagata
EMNLP4
2011 Distributed Minimum Error Rate Training of SMT using Particle Swarm Optimization
Jun Suzuki 0001, Kevin Duh, Masaaki Nagata
IJCNLP1
2009 A Syntax-Free Approach to Japanese Sentence Compression
Tsutomu Hirao, Jun Suzuki 0001, Hideki Isozaki
ACL/IJCNLP2
2009 An Empirical Study of Semi-supervised Structured Conditional Models for Dependency Parsing
Jun Suzuki 0001, Hideki Isozaki, Xavier Carreras, Michael Collins 0001
EMNLP1
2008 Semi-Supervised Sequential Labeling and Segmentation Using Giga-Word Scale Unlabeled Data
Jun Suzuki 0001, Hideki Isozaki
ACL1
2008 Multi-label Text Categorization with Model Combination based on F1-score Maximization
Akinori Fujino, Hideki Isozaki, Jun Suzuki 0001
IJCNLP3
2007 Semi-Supervised Structured Output Learning Based on a Hybrid Generative and Discriminative Approach
Jun Suzuki 0001, Akinori Fujino, Hideki Isozaki
EMNLP-CoNLL1
2007 Online Large-Margin Training for Statistical Machine Translation
Taro Watanabe, Jun Suzuki 0001, Hajime Tsukada, Hideki Isozaki
EMNLP-CoNLL2
2006 Training Conditional Random Fields with Multivariate Evaluation Measures
abstract
This paper proposes a framework for training Conditional Random Fields (CRFs) to optimize multivariate evaluation measures, including non-linear measures such as F-score. Our proposed framework is derived from an error minimization approach that provides a simple solution for directly optimizing any evaluation measure. Specifically focusing on sequential segmentation tasks, i.e. text chunking and named entity recognition, we introduce a loss function that closely reflects the target evaluation measure for these tasks, namely, segmentation F-score. Our experiments show that our method performs better than standard CRF training.
Jun Suzuki 0001, Erik McDermott, Hideki Isozaki
ACL1
2005 Boosting-based Parse Reranking with Subtree Features
abstract
This paper introduces a new application of boosting for parse reranking.Several parsers have been proposed that utilize the all-subtrees representation (e.g., tree kernel and data oriented parsing).This paper argues that such an all-subtrees representation is extremely redundant and a comparable accuracy can be achieved using just a small set of subtrees.We show how the boosting algorithm can be applied to the all-subtrees representation and how it selects a small and relevant feature set efficiently.Two experiments on parse reranking show that our method achieves comparable or even better performance than kernel methods and also improves the testing efficiency.
Taku Kudo, Jun Suzuki 0001, Hideki Isozaki
ACL2
2005 Sequence and Tree Kernels with Statistical Feature Mining
abstract
This paper proposes a new approach to feature selection based on a sta- tistical feature mining technique for sequence and tree kernels. Since natural language data take discrete structures, convolution kernels, such as sequence and tree kernels, are advantageous for both the concept and accuracy of many natural language processing tasks. However, experi- ments have shown that the best results can only be achieved when lim- ited small sub-structures are dealt with by these kernels. This paper dis- cusses this issue of convolution kernels and then proposes a statistical feature selection that enable us to use larger sub-structures effectively. The proposed method, in order to execute efficiently, can be embedded into an original kernel calculation process by using sub-structure min- ing algorithms. Experiments on real NLP tasks confirm the problem in the conventional method and compare the performance of a conventional method to that of the proposed method.
Jun Suzuki 0001, Hideki Isozaki
NIPS1
2004 Convolution Kernels with Feature Selection for Natural Language Processing Tasks
abstract
Convolution kernels, such as sequence and tree kernels, are advantageous for both the concept and accuracy of many natural language processing (NLP) tasks. Experiments have, however, shown that the over-fitting problem often arises when these kernels are used in NLP tasks. This paper discusses this issue of convolution kernels, and then proposes a new approach based on statistical feature selection that avoids this issue. To enable the proposed method to be executed efficiently, it is embedded into an original kernel calculation process by using sub-structure mining algorithms. Experiments are undertaken on real NLP tasks to confirm the problem with a conventional method and to compare its performance with that of the proposed method.
Jun Suzuki 0001, Hideki Isozaki, Eisaku Maeda
ACL1
2004 Dependency-based Sentence Alignment for Multiple Document Summarization
Tsutomu Hirao, Jun Suzuki 0001, Hideki Isozaki, Eisaku Maeda
COLING2
2003 Hierarchical Directed Acyclic Graph Kernel: Methods for Structured Natural Language Data
abstract
This paper proposes the "Hierarchical Directed Acyclic Graph (HDAG) Kernel" for structured natural language data. The HDAG Kernel directly accepts several levels of both chunks and their relations, and then efficiently computes the weighed sum of the number of common attribute sequences of the HDAGs. We applied the proposed method to question classification and sentence alignment tasks to evaluate its performance as a similarity measure and a kernel function. The results of the experiments demonstrate that the HDAG Kernel is superior to other kernel functions and baseline methods.
Jun Suzuki 0001, Tsutomu Hirao, Yutaka Sasaki, Eisaku Maeda
ACL1
2003 Kernels for Structured Natural Language Data
abstract
This paper devises a novel kernel function for structured natural language data. In the field of Natural Language Processing, feature extraction consists of the following two steps: (1) syntactically and semantically analyzing raw data, i.e., character strings, then representing the results as discrete structures, such as parse trees and dependency graphs with part-of-speech tags; (2) creating (possibly high-dimensional) numerical feature vectors from the discrete structures. The new kernels, called Hier- archical Directed Acyclic Graph (HDAG) kernels, directly accept DAGs whose nodes can contain DAGs. HDAG data structures are needed to fully reflect the syntactic and semantic structures that natural language data inherently have. In this paper, we define the kernel function and show how it permits efficient calculation. Experiments demonstrate that the proposed kernels are superior to existing kernel functions, e.g., se- quence kernels, tree kernels, and bag-of-words kernels.
Jun Suzuki 0001, Yutaka Sasaki, Eisaku Maeda
NIPS1
2002 SVM Answer Selection for Open-Domain Question Answering
Jun Suzuki 0001, Yutaka Sasaki, Eisaku Maeda
COLING1