EDBT 2026 Demo / reviewers in the wild / expert
Kazuma Hashimoto
dblp:76/2653
· DBLP profile ↗
20ranked-venue papers
8as first author
7since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 7 first-author · 7 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
14 papers |
Question answering and dialogue systems · 42% Information extraction and text analysis · 13% Transfer learning and domain adaptation · 9% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 30 heaviest of 39, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Question answering and dialogue systems › reasoning-based question answering
multi-hop question answering |
1.0 | 2 | 2022 | Modeling Multi-hop Question Answering as Single Sequence Prediction · ACL (1) 2022 Learning to Retrieve Reasoning Paths over Wikipedia Graph for Question Answering · ICLR 2020 |
Natural language and speech › Question answering and dialogue systems
task-oriented dialogue |
1.0 | 2 | 2022 | [CASPI] Causal-aware Safe Policy Improvement for Task-oriented Dialogue · ACL (1) 2022 Discriminative Nearest Neighbor Few-Shot Intent Detection by Transferring Natural Language Inference · EMNLP (1) 2020 |
Natural language and speech › Question answering and dialogue systems › task-oriented dialogue
dialogue policy learning |
0.6 | 1 | 2022 | [CASPI] Causal-aware Safe Policy Improvement for Task-oriented Dialogue · ACL (1) 2022 |
Natural language and speech › Question answering and dialogue systems › answer generation
generative question answering |
0.6 | 1 | 2022 | Modeling Multi-hop Question Answering as Single Sequence Prediction · ACL (1) 2022 |
Natural language and speech › Question answering and dialogue systems
knowledge base question answering |
0.6 | 1 | 2022 | RNG-KBQA: Generation Augmented Iterative Ranking for Knowledge Base Question Answering · ACL (1) 2022 |
Machine learning › Reinforcement learning
offline reinforcement learning |
0.6 | 1 | 2022 | [CASPI] Causal-aware Safe Policy Improvement for Task-oriented Dialogue · ACL (1) 2022 |
Machine learning › Reinforcement learning › safe reinforcement learning
safe policy improvement |
0.6 | 1 | 2022 | [CASPI] Causal-aware Safe Policy Improvement for Task-oriented Dialogue · ACL (1) 2022 |
Natural language and speech › Information extraction and text analysis
sequence labeling |
0.6 | 1 | 2022 | Transforming Sequence Tagging Into A Seq2Seq Task · EMNLP 2022 |
Machine learning › Probabilistic and Bayesian machine learning
structured prediction |
0.6 | 1 | 2022 | Transforming Sequence Tagging Into A Seq2Seq Task · EMNLP 2022 |
Natural language and speech › Language models and text generation
controllable text generation |
0.5 | 1 | 2021 | CoCo: Controllable Counterfactuals for Evaluating Dialogue State Trackers · ICLR 2021 |
Natural language and speech › Question answering and dialogue systems › task-oriented dialogue
dialogue state tracking |
0.5 | 1 | 2021 | CoCo: Controllable Counterfactuals for Evaluating Dialogue State Trackers · ICLR 2021 |
Machine learning › Learning paradigms › semi-supervised learning
consistency regularization |
0.4 | 1 | 2020 | Simple Data Augmentation with the Mask Token Improves Domain Adaptation for Dialog Act Tagging · EMNLP (1) 2020 |
Machine learning › Transfer learning and domain adaptation
cross-task transfer |
0.4 | 1 | 2020 | Discriminative Nearest Neighbor Few-Shot Intent Detection by Transferring Natural Language Inference · EMNLP (1) 2020 |
Natural language and speech › Question answering and dialogue systems › dialogue understanding
dialogue act classification |
0.4 | 1 | 2020 | Simple Data Augmentation with the Mask Token Improves Domain Adaptation for Dialog Act Tagging · EMNLP (1) 2020 |
Machine learning › Transfer learning and domain adaptation
domain adaptation |
0.4 | 1 | 2020 | Simple Data Augmentation with the Mask Token Improves Domain Adaptation for Dialog Act Tagging · EMNLP (1) 2020 |
Natural language and speech › Question answering and dialogue systems › intent detection
few-shot intent detection |
0.4 | 1 | 2020 | Discriminative Nearest Neighbor Few-Shot Intent Detection by Transferring Natural Language Inference · EMNLP (1) 2020 |
Natural language and speech › Question answering and dialogue systems
intent detection |
0.4 | 1 | 2020 | Discriminative Nearest Neighbor Few-Shot Intent Detection by Transferring Natural Language Inference · EMNLP (1) 2020 |
Natural language and speech › Question answering and dialogue systems
open-domain question answering |
0.4 | 1 | 2020 | Learning to Retrieve Reasoning Paths over Wikipedia Graph for Question Answering · ICLR 2020 |
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation |
0.4 | 1 | 2020 | Simple Data Augmentation with the Mask Token Improves Domain Adaptation for Dialog Act Tagging · EMNLP (1) 2020 |
Machine learning › Deep learning architectures and training › encoder-decoder architecture
attention-based encoder-decoder |
0.3 | 1 | 2017 | Neural Machine Translation with Source-Side Latent Graph Parsing · EMNLP 2017 |
Machine learning › Learning paradigms
multi-task learning |
0.3 | 1 | 2017 | A Joint Many-Task Model: Growing a Neural Network for Multiple NLP Tasks · EMNLP 2017 |
Natural language and speech › Machine translation
neural machine translation |
0.3 | 1 | 2017 | Neural Machine Translation with Source-Side Latent Graph Parsing · EMNLP 2017 |
Natural language and speech › Machine translation › neural machine translation
syntax-based neural machine translation |
0.3 | 1 | 2017 | Neural Machine Translation with Source-Side Latent Graph Parsing · EMNLP 2017 |
Machine learning › Representation and self-supervised learning › text embedding
phrase embedding |
0.2 | 1 | 2016 | Adaptive Joint Learning of Compositional and Non-Compositional Phrase Embeddings · ACL (1) 2016 |
Machine learning › Representation and self-supervised learning › representation learning › compositional representation
compositional distributional semantics |
0.2 | 1 | 2014 | Jointly Learning Word Representations and Composition Functions Using Predicate-Argument Structures · EMNLP 2014 |
Natural language and speech › Information extraction and text analysis › semantic role labeling
predicate-argument structure |
0.2 | 1 | 2014 | Jointly Learning Word Representations and Composition Functions Using Predicate-Argument Structures · EMNLP 2014 |
Machine learning › Representation and self-supervised learning › representation learning › semantic representation learning
semantic composition |
0.2 | 1 | 2014 | Jointly Learning Word Representations and Composition Functions Using Predicate-Argument Structures · EMNLP 2014 |
Machine learning › Representation and self-supervised learning › word representation
word representation learning |
0.2 | 1 | 2014 | Jointly Learning Word Representations and Composition Functions Using Predicate-Argument Structures · EMNLP 2014 |
Information retrieval › document retrieval
passage retrieval |
0.2 | 1 | 2022 | Modeling Multi-hop Question Answering as Single Sequence Prediction · ACL (1) 2022 |
Machine learning › Deep learning architectures and training
recursive neural network |
0.2 | 1 | 2013 | Simple Customization of Recursive Neural Networks for Semantic Relation Classification · EMNLP 2013 |
Methods — techniques the papers use, named apart from their topics
sequence-to-sequence generation · 1.1fusion-in-decoder · 1.1cross-passage interaction encoding · 1.1safe policy improvement · 0.6pre-trained language model · 0.6multilingual transfer learning · 0.6generation model · 0.6contrastive ranking · 0.6causal reward learning · 0.6batch reinforcement learning · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Take One Step at a Time to Know Incremental Utility of Demonstration: An Analysis on Reranking for Few-Shot In-Context LearningabstractKazuma Hashimoto, Karthik Raman, Michael Bendersky. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Kazuma Hashimoto, Karthik Raman 0001, Michael Bendersky |
NAACL-HLT | 1 |
| 2022 | [CASPI] Causal-aware Safe Policy Improvement for Task-oriented DialogueabstractThe recent success of reinforcement learning (RL) in solving complex tasks is often attributed to its capacity to explore and exploit an environment.Sample efficiency is usually not an issue for tasks with cheap simulators to sample data online.On the other hand, Taskoriented Dialogues (ToD) are usually learnt from offline data collected using human demonstrations.Collecting diverse demonstrations and annotating them is expensive.Unfortunately, RL policy trained on off-policy data are prone to issues of bias and generalization, which are further exacerbated by stochasticity in human response and non-markovian nature of annotated belief state of a dialogue management system.To this end, we propose a batch-RL framework for ToD policy learning: Causal-aware Safe Policy Improvement (CASPI).CASPI includes a mechanism to learn fine-grained reward that captures intention behind human response and also offers guarantee on dialogue policy's performance against a baseline.We demonstrate the effectiveness of this framework on end-to-end dialogue task of the Multiwoz2.0dataset.The proposed method outperforms the current state of the art.Further more we demonstrate sample efficiency, where our method trained only on 20% of the data, are comparable to current state of the art method trained on 100% data on two out of there evaluation metrics. Govardana Sachithanandam Ramachandran, Kazuma Hashimoto, Caiming Xiong |
ACL (1) | 2 |
| 2022 | Modeling Multi-hop Question Answering as Single Sequence PredictionabstractFusion-in-decoder (FID) (Izacard and Grave, 2021) is a generative question answering (QA) model that leverages passage retrieval with a pre-trained transformer and pushed the state of the art on single-hop QA.However, the complexity of multi-hop QA hinders the effectiveness of the generative QA approach.In this work, we propose a simple generative approach (PATHFID) that extends the task beyond just answer generation by explicitly modeling the reasoning process to resolve the answer for multihop questions.By linearizing the hierarchical reasoning path of supporting passages, their key sentences, and finally the factoid answer, we cast the problem as a single sequence prediction task.To facilitate complex reasoning with multiple clues, we further extend the unified flat representation of multiple input documents by encoding cross-passage interactions.Our extensive experiments demonstrate that PATHFID leads to strong performance gains on two multihop QA datasets: HotpotQA and IIRC.Besides the performance gains, PATHFID is more interpretable, which in turn yields answers that are more faithfully grounded to the supporting passages and facts compared to the baseline FID model. Semih Yavuz, Kazuma Hashimoto, Yingbo Zhou 0002, Nitish Shirish Keskar, Caiming Xiong |
ACL (1) | 2 |
| 2022 | RNG-KBQA: Generation Augmented Iterative Ranking for Knowledge Base Question AnsweringabstractExisting KBQA approaches, despite achieving strong performance on i.i.d.test data, often struggle in generalizing to questions involving unseen KB schema items.Prior rankingbased approaches have shown some success in generalization, but suffer from the coverage issue.We present RnG-KBQA, a Rank-and-Generate approach for KBQA, which remedies the coverage issue with a generation model while preserving a strong generalization capability.Our approach first uses a contrastive ranker to rank a set of candidate logical forms obtained by searching over the knowledge graph.It then introduces a tailored generation model conditioned on the question and the top-ranked candidates to compose the final logical form.We achieve new state-ofthe-art results on GRAILQA and WEBQSP datasets.In particular, our method surpasses the prior state-of-the-art by a large margin on the GRAILQA leaderboard.In addition, RnG-KBQA outperforms all prior approaches on the popular WEBQSP benchmark, even including the ones that use the oracle entity linking.The experimental results demonstrate the effectiveness of the interplay between ranking and generation, which leads to the superior performance of our proposed approach across all settings with especially strong improvements in zero-shot generalization. 1 * Work done during internship at Salesforce Research. 1 Code available at https://github.com/salesforce/rng-kbqa. Xi Ye 0003, Semih Yavuz, Kazuma Hashimoto, Yingbo Zhou 0002, Caiming Xiong |
ACL (1) | 3 |
| 2022 | Transforming Sequence Tagging Into A Seq2Seq TaskabstractPretrained, large, generative language models (LMs) have had great success in a wide range of sequence tagging and structured prediction tasks.Casting a sequence tagging task as a Seq2Seq one requires deciding the formats of the input and output sequences.However, we lack a principled understanding of the tradeoffs associated with these formats (such as the effect on model accuracy, sequence length, multilingual generalization, hallucination).In this paper, we rigorously study different formats one could use for casting input text sentences and their output labels into the input and target (i.e., output) of a Seq2Seq model.Along the way, we introduce a new format, which we show to to be both simpler and more effective.Additionally the new format demonstrates significant gains in the multilingual settings -both zero-shot transfer learning and joint training.Lastly, we find that the new format is more robust and almost completely devoid of hallucination -an issue we find common in existing formats.With well over a 1000 experiments studying 14 different formats, over 7 diverse public benchmarksincluding 3 multilingual datasets spanning 7 languages -we believe our findings provide a strong empirical basis in understanding how we should tackle sequence tagging tasks.Dataset Lang # Train # Valid # Test Token/ex Tagged span/ex % tokens tagged # Tag classes Tag Entropy mATIS en 4478 500 893 11.28 3.32 36.50 79 3. Karthik Raman 0001, Iftekhar Naim, Jiecao Chen, Kazuma Hashimoto, Kiran Yalasangi, Krishna Srinivasan |
EMNLP | 4 |
| 2021 | CoCo: Controllable Counterfactuals for Evaluating Dialogue State Trackers
Semih Yavuz, Kazuma Hashimoto, Jia Li 0015, Nazneen Fatema Rajani, Xifeng Yan, Yingbo Zhou 0002, Caiming Xiong |
ICLR | 3 |
| 2021 | Focused Attention Improves Document-Grounded GenerationabstractShrimai Prabhumoye, Kazuma Hashimoto, Yingbo Zhou, Alan W Black, Ruslan Salakhutdinov. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Shrimai Prabhumoye, Kazuma Hashimoto, Yingbo Zhou 0002, Alan W. Black, Ruslan Salakhutdinov |
NAACL-HLT | 2 |
| 2020 | Simple Data Augmentation with the Mask Token Improves Domain Adaptation for Dialog Act TaggingabstractThe concept of Dialogue Act (DA) is universal across different task-oriented dialogue domains -the act of "request" carries the same speaker intention whether it is for restaurant reservation or flight booking.However, DA taggers trained on one domain do not generalize well to other domains, which leaves us with the expensive need for a large amount of annotated data in the target domain.In this work, we investigate how to better adapt DA taggers to desired target domains with only unlabeled data.We propose MASKAUGMENT, a controllable mechanism that augments text input by leveraging the pre-trained MASK token from BERT model.Inspired by consistency regularization, we use MASKAUGMENT to introduce an unsupervised teacher-student learning scheme to examine the domain adaptation of DA taggers.Our extensive experiments on the Simulated Dialogue (GSim) and Schema-Guided Dialogue (SGD) datasets show that MASKAUGMENT is useful in improving the cross-domain generalization for DA tagging. Semih Yavuz, Kazuma Hashimoto, Wenhao Liu 0003, Nitish Shirish Keskar, Richard Socher, Caiming Xiong |
EMNLP (1) | 2 |
| 2020 | Discriminative Nearest Neighbor Few-Shot Intent Detection by Transferring Natural Language InferenceabstractJianguo Zhang, Kazuma Hashimoto, Wenhao Liu, Chien-Sheng Wu, Yao Wan, Philip Yu, Richard Socher, Caiming Xiong. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Jianguo Zhang 0005, Kazuma Hashimoto, Wenhao Liu 0003, Chien-Sheng Wu, Yao Wan 0001, Philip S. Yu, Richard Socher, Caiming Xiong |
EMNLP (1) | 2 |
| 2020 | Learning to Retrieve Reasoning Paths over Wikipedia Graph for Question Answering
Akari Asai, Kazuma Hashimoto, Hannaneh Hajishirzi, Richard Socher, Caiming Xiong |
ICLR | 2 |
| 2020 | Parallelizing and optimizing neural Encoder-Decoder models without padding on multi-core architecture
Yuchen Qiao, Kazuma Hashimoto, Akiko Eriguchi, Haixia Wang 0001, Dongsheng Wang 0002, Yoshimasa Tsuruoka, Kenjiro Taura |
Future Gener. Comput. Syst. | 2 |
| 2019 | Incorporating Source-Side Phrase Structures into Neural Machine TranslationabstractNeural machine translation (NMT) has shown great success as a new alternative to the traditional Statistical Machine Translation model in multiple languages. Early NMT models are based on sequence-to-sequence learning that encodes a sequence of source words into a vector space and generates another sequence of target words from the vector. In those NMT models, sentences are simply treated as sequences of words without any internal structure. In this article, we focus on the role of the syntactic structure of source sentences and propose a novel end-to-end syntactic NMT model, which we call a tree-to-sequence NMT model, extending a sequence-to-sequence model with the source-side phrase structure. Our proposed model has an attention mechanism that enables the decoder to generate a translated word while softly aligning it with phrases as well as words of the source sentence. We have empirically compared the proposed model with sequence-to-sequence models in various settings on Chinese-to-Japanese and English-to-Japanese translation tasks. Our experimental results suggest that the use of syntactic structure can be beneficial when the training data set is small, but is not as effective as using a bi-directional encoder. As the size of training data set increases, the benefits of using a syntactic tree tends to diminish. Akiko Eriguchi, Kazuma Hashimoto, Yoshimasa Tsuruoka |
Comput. Linguistics | 2 |
| 2017 | Neural Machine Translation with Source-Side Latent Graph ParsingabstractThis paper presents a novel neural machine translation model which jointly learns translation and source-side latent graph representations of sentences.Unlike existing pipelined approaches using syntactic parsers, our end-to-end model learns a latent graph parser as part of the encoder of an attention-based neural machine translation model, and thus the parser is optimized according to the translation objective.In experiments, we first show that our model compares favorably with state-of-the-art sequential and pipelined syntax-based NMT models.We also show that the performance of our model can be further improved by pretraining it with a small amount of treebank annotations.Our final ensemble model significantly outperforms the previous best models on the standard Englishto-Japanese translation dataset. Kazuma Hashimoto, Yoshimasa Tsuruoka |
EMNLP | 1 |
| 2017 | A Joint Many-Task Model: Growing a Neural Network for Multiple NLP TasksabstractTransfer and multi-task learning have traditionally focused on either a single source-target pair or very few, similar tasks.Ideally, the linguistic levels of morphology, syntax and semantics would benefit each other by being trained in a single model.We introduce a joint many-task model together with a strategy for successively growing its depth to solve increasingly complex tasks.Higher layers include shortcut connections to lower-level task predictions to reflect linguistic hierarchies.We use a simple regularization term to allow for optimizing all model weights to improve one task's loss without exhibiting catastrophic interference of the other tasks.Our single end-to-end model obtains state-of-the-art or competitive results on five different tasks from tagging, parsing, relatedness, and entailment tasks. Kazuma Hashimoto, Caiming Xiong, Yoshimasa Tsuruoka, Richard Socher |
EMNLP | 1 |
| 2016 | Tree-to-Sequence Attentional Neural Machine TranslationabstractMost of the existing Neural Machine Translation (NMT) models focus on the conversion of sequential data and do not directly use syntactic information.We propose a novel end-to-end syntactic NMT model, extending a sequenceto-sequence model with the source-side phrase structure.Our model has an attention mechanism that enables the decoder to generate a translated word while softly aligning it with phrases as well as words of the source sentence.Experimental results on the WAT'15 Englishto-Japanese dataset demonstrate that our proposed model considerably outperforms sequence-to-sequence attentional NMT models and compares favorably with the state-of-the-art tree-to-string SMT system. Akiko Eriguchi, Kazuma Hashimoto, Yoshimasa Tsuruoka |
ACL (1) | 2 |
| 2016 | Adaptive Joint Learning of Compositional and Non-Compositional Phrase EmbeddingsabstractWe present a novel method for jointly learning compositional and noncompositional phrase embeddings by adaptively weighting both types of embeddings using a compositionality scoring function.The scoring function is used to quantify the level of compositionality of each phrase, and the parameters of the function are jointly optimized with the objective for learning phrase embeddings.In experiments, we apply the adaptive joint learning method to the task of learning embeddings of transitive verb phrases, and show that the compositionality scores have strong correlation with human ratings for verb-object compositionality, substantially outperforming the previous state of the art.Moreover, our embeddings improve upon the previous best model on a transitive verb disambiguation task.We also show that a simple ensemble technique further improves the results for both tasks. Kazuma Hashimoto, Yoshimasa Tsuruoka |
ACL (1) | 1 |
| 2016 | Topic detection using paragraph vectors to support active learning in systematic reviewsabstractSystematic reviews require expert reviewers to manually screen thousands of citations in order to identify all relevant articles to the review. Active learning text classification is a supervised machine learning approach that has been shown to significantly reduce the manual annotation workload by semi-automating the citation screening process of systematic reviews. In this paper, we present a new topic detection method that induces an informative representation of studies, to improve the performance of the underlying active learner. Our proposed topic detection method uses a neural network-based vector space model to capture semantic similarities between documents. We firstly represent documents within the vector space, and cluster the documents into a predefined number of clusters. The centroids of the clusters are treated as latent topics. We then represent each document as a mixture of latent topics. For evaluation purposes, we employ the active learning strategy using both our novel topic detection method and a baseline topic model (i.e., Latent Dirichlet Allocation). Results obtained demonstrate that our method is able to achieve a high sensitivity of eligible studies and a significantly reduced manual annotation cost when compared to the baseline method. This observation is consistent across two clinical and three public health reviews. The tool introduced in this work is available from https://nactem.ac.uk/pvtopic/. Kazuma Hashimoto, Georgios Kontonatsios, Makoto Miwa, Sophia Ananiadou |
J. Biomed. Informatics | 1 |
| 2015 | Task-Oriented Learning of Word Embeddings for Semantic Relation ClassificationabstractWe present a novel learning method for word embeddings designed for relation classification.Our word embeddings are trained by predicting words between noun pairs using lexical relation-specific features on a large unlabeled corpus.This allows us to explicitly incorporate relationspecific information into the word embeddings.The learned word embeddings are then used to construct feature vectors for a relation classification model.On a wellestablished semantic relation classification task, our method significantly outperforms a baseline based on a previously introduced word embedding method, and compares favorably to previous state-of-the-art models that use syntactic information or manually constructed external resources. Kazuma Hashimoto, Pontus Stenetorp, Makoto Miwa, Yoshimasa Tsuruoka |
CoNLL | 1 |
| 2014 | Jointly Learning Word Representations and Composition Functions Using Predicate-Argument StructuresabstractWe introduce a novel compositional lan-guage model that works on Predicate-Argument Structures (PASs). Our model jointly learns word representations and their composition functions using bag-of-words and dependency-based con-texts. Unlike previous word-sequence-based models, our PAS-based model com-poses arguments into predicates by using the category information from the PAS. This enables our model to capture long-range dependencies between words and to better handle constructs such as verb-object and subject-verb-object relations. We verify this experimentally using two phrase similarity datasets and achieve re-sults comparable to or higher than the pre-vious best results. Our system achieves these results without the need for pre-trained word vectors and using a much smaller training corpus; despite this, for the subject-verb-object dataset our model improves upon the state of the art by as much as 10 % in relative performance. 1 Kazuma Hashimoto, Pontus Stenetorp, Makoto Miwa, Yoshimasa Tsuruoka |
EMNLP | 1 |
| 2013 | Simple Customization of Recursive Neural Networks for Semantic Relation ClassificationabstractIn this paper, we present a recursive neural network (RNN) model that works on a syntactic tree.Our model differs from previous RNN models in that the model allows for an explicit weighting of important phrases for the target task.We also propose to average parameters in training.Our experimental results on semantic relation classification show that both phrase categories and task-specific weighting significantly improve the prediction accuracy of the model.We also show that averaging the model parameters is effective in stabilizing the learning and improves generalization capacity.The proposed model marks scores competitive with state-of-the-art RNN-based models. Kazuma Hashimoto, Makoto Miwa, Yoshimasa Tsuruoka, Takashi Chikayama |
EMNLP | 1 |