VLDB 2026 Research / reviewers in the wild / expert
Hiroshi Noji
dblp:136/9156
· DBLP profile ↗
19ranked-venue papers
4as first author
3since 2021 · last 2021
0000-0003-0061-6481ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Information extraction and text analysis · 50% Language models and text generation · 44% Knowledge representation and reasoning · 6% | |
| Theoretical computer science
2 papers |
Automata and formal languages · 82% Automated reasoning and model checking · 18% |
Topics — the 19 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis
syntactic parsing |
1.1 | 4 | 2019 | Automatic Generation of High Quality CCGbanks for Parser Domain Adaptation · ACL (1) 2019 A* CCG Parsing with a Supertag and Dependency Factored Model · ACL (1) 2017 Using Left-corner Parsing to Encode Universal Structural Constraints in Grammar Induction · EMNLP 2016 |
Natural language and speech › Information extraction and text analysis › syntactic parsing › grammar-based parsing
combinatory categorial grammar parsing |
0.7 | 2 | 2019 | Automatic Generation of High Quality CCGbanks for Parser Domain Adaptation · ACL (1) 2019 A* CCG Parsing with a Supertag and Dependency Factored Model · ACL (1) 2017 |
Natural language and speech › Language models and text generation
human sentence processing |
0.5 | 1 | 2021 | Modeling Human Sentence Processing with Left-Corner Recurrent Neural Network Grammars · EMNLP (1) 2021 |
Natural language and speech › Language models and text generation › natural language understanding › linguistic knowledge in language models
subject-verb agreement |
0.4 | 1 | 2020 | An Analysis of the Utility of Explicit Negative Examples to Improve the Syntactic Abilities of Neural Language Models · ACL 2020 |
Natural language and speech › Language models and text generation › linguistic generalization
syntactic generalization |
0.4 | 1 | 2020 | An Analysis of the Utility of Explicit Negative Examples to Improve the Syntactic Abilities of Neural Language Models · ACL 2020 |
Natural language and speech › Language models and text generation › text generation
content selection and planning |
0.4 | 1 | 2019 | Learning to Select, Track, and Generate for Data-to-Text · ACL (1) 2019 |
Natural language and speech › Language models and text generation › text generation
data-to-text generation |
0.4 | 1 | 2019 | Learning to Select, Track, and Generate for Data-to-Text · ACL (1) 2019 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge graph reasoning
knowledge base completion |
0.4 | 1 | 2019 | Combining Axiom Injection and Knowledge Base Completion for Efficient Natural Language Inference · AAAI 2019 |
Natural language and speech › Language models and text generation › natural language understanding › sentence pair modeling
natural language inference |
0.4 | 1 | 2019 | Combining Axiom Injection and Knowledge Base Completion for Efficient Natural Language Inference · AAAI 2019 |
Natural language and speech › Information extraction and text analysis › syntactic parsing
parser adaptation |
0.4 | 1 | 2019 | Automatic Generation of High Quality CCGbanks for Parser Domain Adaptation · ACL (1) 2019 |
Natural language and speech › Information extraction and text analysis
textual entailment |
0.4 | 1 | 2019 | Combining Axiom Injection and Knowledge Base Completion for Efficient Natural Language Inference · AAAI 2019 |
Natural language and speech › Information extraction and text analysis › syntactic parsing
a* parsing |
0.3 | 1 | 2017 | A* CCG Parsing with a Supertag and Dependency Factored Model · ACL (1) 2017 |
Natural language and speech › Language models and text generation
grammar induction |
0.2 | 1 | 2016 | Using Left-corner Parsing to Encode Universal Structural Constraints in Grammar Induction · EMNLP 2016 |
Automata and formal languages › parsing
left-corner parsing |
0.2 | 1 | 2016 | Using Left-corner Parsing to Encode Universal Structural Constraints in Grammar Induction · EMNLP 2016 |
Automata and formal languages
parsing |
0.2 | 1 | 2016 | Using Left-corner Parsing to Encode Universal Structural Constraints in Grammar Induction · EMNLP 2016 |
Natural language and speech › Information extraction and text analysis › syntactic parsing
constituency parsing |
0.2 | 1 | 2015 | Optimal Shift-Reduce Constituent Parsing with Structured Perceptron · ACL (1) 2015 |
Natural language and speech › Information extraction and text analysis › syntactic parsing › transition-based parsing
shift-reduce parsing |
0.2 | 1 | 2015 | Optimal Shift-Reduce Constituent Parsing with Structured Perceptron · ACL (1) 2015 |
Natural language and speech › Language models and text generation › language modeling
statistical language modeling |
0.2 | 1 | 2013 | Improvements to the Bayesian Topic N-Gram Models · EMNLP 2013 |
Automated reasoning and model checking › theorem proving
proof automation |
0.1 | 1 | 2019 | Combining Axiom Injection and Knowledge Base Completion for Efficient Natural Language Inference · AAAI 2019 |
Methods — techniques the papers use, named apart from their topics
knowledge base completion · 0.8axiom injection · 0.8abduction · 0.8recurrent neural network grammar · 0.5LSTM · 0.5margin loss · 0.4data augmentation · 0.4neural text generation · 0.4dependency tree conversion · 0.4automatic corpus generation · 0.4dependency model with valence · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Modeling Human Sentence Processing with Left-Corner Recurrent Neural Network GrammarsabstractIn computational linguistics, it has been shown that hierarchical structures make language models (LMs) more human-like.However, the previous literature has been agnostic about a parsing strategy of the hierarchical models.In this paper, we investigated whether hierarchical structures make LMs more human-like, and if so, which parsing strategy is most cognitively plausible.In order to address this question, we evaluated three LMs against human reading times in Japanese with head-final leftbranching structures: Long Short-Term Memory (LSTM) as a sequential model and Recurrent Neural Network Grammars (RNNGs) with top-down and left-corner parsing strategies as hierarchical models.Our computational modeling demonstrated that left-corner RNNGs outperformed top-down RNNGs and LSTM, suggesting that hierarchical and leftcorner architectures are more cognitively plausible than top-down or sequential architectures.In addition, the relationships between the cognitive plausibility and (i) perplexity, (ii) parsing, and (iii) beam size will also be discussed.1 Ryo Yoshida, Hiroshi Noji, Yohei Oseki |
EMNLP (1) | 2 |
| 2021 | Generating Racing Game Commentary from Vision, Language, and Structured DataabstractWe propose the task of automatically generating commentaries for races in a motor racing game, from vision, structured numerical, and textual data.Commentaries provide information to support spectators in understanding events in races.Commentary generation models need to interpret the race situation and generate the correct content at the right moment.We divide the task into two subtasks: utterance timing identification and utterance generation.Because existing datasets do not have such alignments of data in multiple modalities, this setting has not been explored in depth.In this study, we introduce a new large-scale dataset that contains aligned video data, structured numerical data, and transcribed commentaries that consist of 129,226 utterances in 1,389 races in a game.Our analysis reveals that the characteristics of commentaries change depending on time and viewpoints.Our experiments on the subtasks show that it is still challenging for a state-of-the-art vision encoder to capture useful information from videos to generate accurate commentaries.We make the dataset and baseline implementation publicly available for further research.1 Tatsuya Ishigaki, Goran Topic, Yumi Hamazono, Hiroshi Noji, Ichiro Kobayashi 0001, Yusuke Miyao, Hiroya Takamura |
INLG | 4 |
| 2021 | Controlling contents in data-to-document generation with human-designed topic labels
Kasumi Aoki, Akira Miyazawa, Tatsuya Ishigaki, Tatsuya Aoki, Hiroshi Noji, Keiichi Goshima, Hiroya Takamura, Yusuke Miyao, Ichiro Kobayashi 0001 |
Comput. Speech Lang. | 5 |
| 2020 | An Analysis of the Utility of Explicit Negative Examples to Improve the Syntactic Abilities of Neural Language ModelsabstractWe explore the utilities of explicit negative examples in training neural language models.Negative examples here are incorrect words in a sentence, such as barks in *The dogs barks.Neural language models are commonly trained only on positive examples, a set of sentences in the training data, but recent studies suggest that the models trained in this way are not capable of robustly handling complex syntactic constructions, such as long-distance agreement.In this paper, we first demonstrate that appropriately using negative examples about particular constructions (e.g., subject-verb agreement) will boost the model's robustness on them in English, with a negligible loss of perplexity.The key to our success is an additional margin loss between the log-likelihoods of a correct word and an incorrect word.We then provide a detailed analysis of the trained models.One of our findings is the difficulty of object-relative clauses for RNNs.We find that even with our direct learning signals the models still suffer from resolving agreement across an object-relative clause.Augmentation of training sentences involving the constructions somewhat helps, but the accuracy still does not reach the level of subjectrelative clauses.Although not directly cognitively appealing, our method can be a tool to analyze the true architectural limitation of neural models on challenging linguistic constructions. Hiroshi Noji, Hiroya Takamura |
ACL | 1 |
| 2020 | CharacterBERT: Reconciling ELMo and BERT for Word-Level Open-Vocabulary Representations From CharactersabstractDue to the compelling improvements brought by BERT, many recent representation models adopted the Transformer architecture as their main building block, consequently inheriting the wordpiece tokenization system despite it not being intrinsically linked to the notion of Transformers.While this system is thought to achieve a good balance between the flexibility of characters and the efficiency of full words, using predefined wordpiece vocabularies from the general domain is not always suitable, especially when building models for specialized domains (e.g., the medical domain).Moreover, adopting a wordpiece tokenization shifts the focus from the word level to the subword level, making the models conceptually more complex and arguably less convenient in practice.For these reasons, we propose CharacterBERT, a new variant of BERT that drops the wordpiece system altogether and uses a Character-CNN module instead to represent entire words by consulting their characters.We show that this new model improves the performance of BERT on a variety of medical domain tasks while at the same time producing robust, word-level, and open-vocabulary representations. Hicham El Boukkouri, Olivier Ferret, Thomas Lavergne, Hiroshi Noji, Pierre Zweigenbaum, Jun'ichi Tsujii |
COLING | 4 |
| 2020 | An empirical analysis of existing systems and datasets toward general simple question answeringabstractIn this paper, we evaluate the progress of our field toward solving simple factoid questions over a knowledge base, a practically important problem in natural language interface to database. As in other natural language understanding tasks, a common practice for this task is to train and evaluate a model on a single dataset, and recent studies suggest that SimpleQuestions, the most popular and largest dataset, is nearly solved under this setting. However, this common setting does not evaluate the robustness of the systems outside of the distribution of the used training data. We rigorously evaluate such robustness of existing systems using different datasets. Our analysis, including shifting of training and test datasets and training on a union of the datasets, suggests that our progress in solving SimpleQuestions dataset does not indicate the success of more general simple question answering. We discuss a possible future direction toward this goal. Namgi Han, Goran Topic, Hiroshi Noji, Hiroya Takamura, Yusuke Miyao |
COLING | 3 |
| 2020 | Learning with Contrastive Examples for Data-to-Text GenerationabstractYui Uehara, Tatsuya Ishigaki, Kasumi Aoki, Hiroshi Noji, Keiichi Goshima, Ichiro Kobayashi, Hiroya Takamura, Yusuke Miyao. Proceedings of the 28th International Conference on Computational Linguistics. 2020. Yui Uehara, Tatsuya Ishigaki, Kasumi Aoki, Hiroshi Noji, Keiichi Goshima, Ichiro Kobayashi 0001, Hiroya Takamura, Yusuke Miyao |
COLING | 4 |
| 2020 | Market Comment Generation from Data with Noisy AlignmentsabstractEnd-to-end models on data-to-text learn the mapping of data and text from the aligned pairs in the dataset.However, these alignments are not always obtained reliably, especially for the time-series data, for which real time comments are given to some situation and there might be a delay in the comment delivery time compared to the actual event time.To handle this issue of possible noisy alignments in the dataset, we propose a neural network model with multitimestep data and a copy mechanism, which allows the models to learn the correspondences between data and text from the dataset with noisier alignments.We focus on generating market comments in Japanese that are delivered each time an event occurs in the market.The core idea of our approach is to utilize multitimestep data, which is not only the latest market price data when the comment is delivered, but also the data obtained at several timesteps earlier.On top of this, we employ a copy mechanism that is suitable for referring to the content of data records in the market price data.We confirm the superiority of our proposal by two evaluation metrics and show the accuracy improvement of the sentence generation using the time series data by our proposed method. Yumi Hamazono, Yui Uehara, Hiroshi Noji, Yusuke Miyao, Hiroya Takamura, Ichiro Kobayashi 0001 |
INLG | 3 |
| 2019 | Combining Axiom Injection and Knowledge Base Completion for Efficient Natural Language InferenceabstractIn logic-based approaches to reasoning tasks such as Recognizing Textual Entailment (RTE), it is important for a system to have a large amount of knowledge data. However, there is a tradeoff between adding more knowledge data for improved RTE performance and maintaining an efficient RTE system, as such a big database is problematic in terms of the memory usage and computational complexity. In this work, we show the processing time of a state-of-the-art logic-based RTE system can be significantly reduced by replacing its search-based axiom injection (abduction) mechanism by that based on Knowledge Base Completion (KBC). We integrate this mechanism in a Coq plugin that provides a proof automation tactic for natural language inference. Additionally, we show empirically that adding new knowledge data contributes to better RTE performance while not harming the processing speed in this framework. Masashi Yoshikawa, Koji Mineshima, Hiroshi Noji, Daisuke Bekki |
AAAI | 3 |
| 2019 | Learning to Select, Track, and Generate for Data-to-TextabstractHayate Iso, Yui Uehara, Tatsuya Ishigaki, Hiroshi Noji, Eiji Aramaki, Ichiro Kobayashi, Yusuke Miyao, Naoaki Okazaki, Hiroya Takamura. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019. Hayate Iso, Yui Uehara, Tatsuya Ishigaki, Hiroshi Noji, Eiji Aramaki, Ichiro Kobayashi 0001, Yusuke Miyao, Naoaki Okazaki, Hiroya Takamura |
ACL (1) | 4 |
| 2019 | Automatic Generation of High Quality CCGbanks for Parser Domain AdaptationabstractWe propose a new domain adaptation method for Combinatory Categorial Grammar (CCG) parsing, based on the idea of automatic generation of CCG corpora exploiting cheaper resources of dependency trees. Our solution is conceptually simple, and not relying on a specific parser architecture, making it applicable to the current best-performing parsers. We conduct extensive parsing experiments with detailed discussion; on top of existing benchmark datasets on (1) biomedical texts and (2) question sentences, we create experimental datasets of (3) speech conversation and (4) math problems. When applied to the proposed method, an off-the-shelf CCG parser shows significant performance gains, improving from 90.7% to 96.6% on speech conversation, and from 88.5% to 96.8% on math problems. Masashi Yoshikawa, Hiroshi Noji, Koji Mineshima, Daisuke Bekki |
ACL (1) | 2 |
| 2019 | Controlling Contents in Data-to-Document Generation with Human-Designed Topic LabelsabstractKasumi Aoki, Akira Miyazawa, Tatsuya Ishigaki, Tatsuya Aoki, Hiroshi Noji, Keiichi Goshima, Ichiro Kobayashi, Hiroya Takamura, Yusuke Miyao. Proceedings of the 12th International Conference on Natural Language Generation. 2019. Kasumi Aoki, Akira Miyazawa, Tatsuya Ishigaki, Tatsuya Aoki, Hiroshi Noji, Keiichi Goshima, Ichiro Kobayashi 0001, Hiroya Takamura, Yusuke Miyao |
INLG | 5 |
| 2018 | Dynamic Feature Selection with Attention in Incremental ParsingabstractOne main challenge for incremental transition-based parsers, when future inputs are invisible, is to extract good features from a limited local context. In this work, we present a simple technique to maximally utilize the local features with an attention mechanism, which works as context- dependent dynamic feature selection. Our model learns, for example, which tokens should a parser focus on, to decide the next action. Our multilingual experiment shows its effectiveness across many languages. We also present an experiment with augmented test dataset and demon- strate it helps to understand the model’s behavior on locally ambiguous points. Ryosuke Kohita, Hiroshi Noji, Yuji Matsumoto 0001 |
COLING | 2 |
| 2018 | An Empirical Investigation of Error Types in Vietnamese ParsingabstractSyntactic parsing plays a crucial role in improving the quality of natural language processing tasks. Although there have been several research projects on syntactic parsing in Vietnamese, the parsing quality has been far inferior than those reported in major languages, such as English and Chinese. In this work, we evaluated representative constituency parsing models on a Vietnamese Treebank to look for the most suitable parsing method for Vietnamese. We then combined the advantages of automatic and manual analysis to investigate errors produced by the experimented parsers and find the reasons for them. Our analysis focused on three possible sources of parsing errors, namely limited training data, part-of-speech (POS) tagging errors, and ambiguous constructions. As a result, we found that the last two sources, which frequently appear in Vietnamese text, significantly attributed to the poor performance of Vietnamese parsing. Quy Nguyen, Yusuke Miyao, Hiroshi Noji, Nhung Nguyen |
COLING | 3 |
| 2017 | A* CCG Parsing with a Supertag and Dependency Factored ModelabstractWe propose a new A* CCG parsing model in which the probability of a tree is decomposed into factors of CCG categories and its syntactic dependencies both defined on bi-directional LSTMs.Our factored model allows the precomputation of all probabilities and runs very efficiently, while modeling sentence structures explicitly via dependencies.Our model achieves the stateof-the-art results on English and Japanese CCG parsing. 1 Masashi Yoshikawa, Hiroshi Noji, Yuji Matsumoto 0001 |
ACL (1) | 2 |
| 2016 | Using Left-corner Parsing to Encode Universal Structural Constraints in Grammar InductionabstractCenter-embedding is difficult to process and is known as a rare syntactic construction across languages.In this paper we describe a method to incorporate this assumption into the grammar induction tasks by restricting the search space of a model to trees with limited centerembedding.The key idea is the tabulation of left-corner parsing, which captures the degree of center-embedding of a parse via its stack depth.We apply the technique to learning of famous generative model, the dependency model with valence (Klein and Manning, 2004).Cross-linguistic experiments on Universal Dependencies show that often our method boosts the performance from the baseline, and competes with the current state-ofthe-art model in a number of languages. Hiroshi Noji, Yusuke Miyao, Mark Johnson 0001 |
EMNLP | 1 |
| 2015 | Optimal Shift-Reduce Constituent Parsing with Structured PerceptronabstractLe Quang Thang, Hiroshi Noji, Yusuke Miyao. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Le Quang Thang, Hiroshi Noji, Yusuke Miyao |
ACL (1) | 2 |
| 2014 | Left-corner Transitions on Dependency Parsing
Hiroshi Noji, Yusuke Miyao |
COLING | 1 |
| 2013 | Improvements to the Bayesian Topic N-Gram ModelsabstractOne of the language phenomena that n-gram language model fails to capture is the topic information of a given situation.We advance the previous study of the Bayesian topic language model by Wallach (2006) in two directions: one, investigating new priors to alleviate the sparseness problem caused by dividing all ngrams into exclusive topics, and two, developing a novel Gibbs sampler that enables moving multiple n-grams across different documents to another topic.Our blocked sampler can efficiently search for higher probability space even with higher order n-grams.In terms of modeling assumption, we found it is effective to assign a topic to only some parts of a document. Hiroshi Noji, Daichi Mochihashi, Yusuke Miyao |
EMNLP | 1 |