Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Jun-ya Norimatsu

dblp:136/9075 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Language models and text generation · 100%
Theoretical computer science
1 paper
Algorithms and data structures · 100%

Topics — the 1 heaviest of 2, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › language modeling
n-gram language model
0.212013
An Efficient Language Model Using Double-Array Structures · EMNLP 2013

Methods — techniques the papers use, named apart from their topics

word ID tuning · 0.3double-array representation · 0.3
YearPublicationVenuePosition
2016 A Fast and Compact Language Model Implementation Using Double-Array Structures
abstract
The language model is a widely used component in fields such as natural language processing, automatic speech recognition, and optical character recognition. In particular, statistical machine translation uses language models, and the translation speed and the amount of memory required are greatly affected by the performance of the language model implementation. We propose a fast and compact implementation of n -gram language models that increases query speed and reduces memory usage by using a double-array structure, which is known to be a fast and compact trie data structure. We propose two types of implementation: one for backward suffix trees and the other for reverse tries. The data structure is optimized for space efficiency by embedding model parameters into otherwise unused spaces in the double-array structure. We show that the reverse trie version of our method is among the smallest state-of-the-art implementations in terms of model size with almost the same speed as the implementation that performs fastest on perplexity calculation tasks. Similarly, we achieve faster decoding while keeping compact model sizes, and we confirm that our method can utilize the efficiency of the double-array structure to achieve a balance between speed and size on translation tasks.
Jun-ya Norimatsu, Makoto Yasuhara, Toru Tanaka, Mikio Yamamoto
ACM Trans. Asian Low Resour. Lang. Inf. Process.1
2013 An Efficient Language Model Using Double-Array Structures
abstract
N gram language models tend to increase in size with inflating the corpus size, and consume considerable resources.In this paper, we propose an efficient method for implementing ngram models based on doublearray structures.First, we propose a method for representing backwards suffix trees using double-array structures and demonstrate its efficiency.Next, we propose two optimization methods for improving the efficiency of data representation in the double-array structures.Embedding probabilities into unused spaces in double-array structures reduces the model size.Moreover, tuning the word IDs in the language model makes the model smaller and faster.We also show that our method can be used for building large language models using the division method.Lastly, we show that our method outperforms methods based on recent related works from the viewpoints of model size and query speed when both optimization methods are used.
Makoto Yasuhara, Toru Tanaka, Jun-ya Norimatsu, Mikio Yamamoto
EMNLP3