Daisuke Okanohara

dblp:60/4701 · DBLP profile ↗
← Back
16ranked-venue papers
9as first author
0since 2021 · last 2018
0000-0001-5969-587XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-authorDatabases, data management, data science and information retrieval · 5 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-authorTheory of computation · 3 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Theoretical computer science
1 paper
Approximation and online algorithms · 50% Algorithms and data structures · 50%
Artificial intelligence
2 papers
Language models and text generation · 54% Information extraction and text analysis · 46%
Databases, data mining, and information retrieval
1 paper
Data mining · 100%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Approximation and online algorithms
online learning
0.112009
Latent Variable Perceptron Algorithm for Structured Classification · IJCAI 2009
Algorithms and data structures › learning algorithms
perceptron
0.112009
Latent Variable Perceptron Algorithm for Structured Classification · IJCAI 2009
Natural language and speech › Language models and text generation › language modeling
discriminative language modeling
0.112007
A discriminative language model with pseudo-negative samples · ACL 2007
Natural language and speech › Information extraction and text analysis
named entity recognition
0.112006
Improving the Scalability of Semi-Markov Conditional Random Fields for Named Entity Recognition · ACL 2006
Data mining › predictive modeling › classification
structured classification
0.012009
Latent Variable Perceptron Algorithm for Structured Classification · IJCAI 2009

Methods — techniques the papers use, named apart from their topics

perceptron algorithm · 0.2latent variable models · 0.1latent variable model · 0.1pseudo-negative sampling · 0.1semi-markov conditional random field · 0.1feature forests · 0.1
YearPublicationVenuePosition
2018 Neural Multi-scale Image Compression
Ken Nakanishi, Shin-ichi Maeda, Takeru Miyato, Daisuke Okanohara
ACCV (6)4
2017 Compressed Bit vectors Based on Variable-to-Fixed Encodings
abstract
We consider practical implementations of compressed bitvectors, which support rank and select operations on a given bit-string, while storing the bit-string in compressed form. Our approach relies on variable-to-fixed encodings of the bit-string, an approach that has not yet been considered systematically for practical encodings of bitvectors. We show that this approach leads to fast practical implementations with low redundancy (i.e. the space used by the bitvector in addition to the compressed representation of the bit-string), and is a flexible and promising solution to the problem of supporting rank and select on moderately compressible bit-strings, such as those encountered in real-world applications.
Seungbum Jo, Stelios Joannou, Daisuke Okanohara, Rajeev Raman, S. Srinivasa Rao 0001
Comput. J.3
2014 Compressed Bit Vectors Based on Variable-to-Fixed Encodings
abstract
We consider practical implementations of compressed bit vectors, which support rank and select operations on a given bit-string, while storing thebit-string in compressed form. Our approach relies on variable-to-fixed (V2F) encodings of the bit-string, an approach that has not yet been considered systematically for practical encodings of bit-vectors. This approach leadsto fast practical implementations with low redundancy (i.e., the space used by the bit vector in addition to the compressed representation of the bit-string),and is a flexible and promising solution to the problem of supporting rank and select on moderately compressible bit-strings, such as those frequently found in real-world applications.
Seungbum Jo, Stelios Joannou, Daisuke Okanohara, Rajeev Raman, S. Srinivasa Rao 0001
DCC3
2011 LGM: Mining Frequent Subgraphs from Linear Graphs
Yasuo Tabei, Daisuke Okanohara, Shuichi Hirose, Koji Tsuda
PAKDD (2)2
2010 Conjunctive Filter: Breaking the Entropy Barrier
abstract
We consider a problem for storing a map that associates a key with a set of values. To store ( n) values from the m universe of size m, it requires log2 n bits of space, which can be approximated as (1.44 + n) log2 m/n bits when n ≪ m. If we allow ϵ fraction of errors in 1 outputs, we can store it with roughly n log2 ϵ bits, which matches the entropy bound. Bloom filter is a wellknown example for such data structures. Our objective is to break this entropy bound and construct more space-efficient data structures. In this paper, we propose a novel data structure called a conjunctive filter, which supports conjunctive queries on k distinct keys for fixed k. Although a conjunctive filter cannot return the set of values itself associated with a queried key, it can perform conjunctive queries with O(1 / √ m) fraction of errors. Also, the consumed space is n k log2 m bits and it is significantly smaller than the entropy bound n 2 log2 m when k ≥ 3. We will show that many problems can be solved by using a conjunctive filter such as full-text search and database join queries. Also, we conducted experiments using a real-world data set, and show that a conjunctive filter answers conjunctive queries almost correctly using about 1/2 ∼ 1/4 space as the entropy bound. 1
Daisuke Okanohara, Yuichi Yoshida
ALENEX1
2009 Latent Variable Perceptron Algorithm for Structured Classification
Xu Sun 0001, Takuya Matsuzaki, Daisuke Okanohara, Jun'ichi Tsujii
IJCAI3
2009 Text Categorization with All Substring Features
abstract
This paper presents a novel document classification method using all substrings as features. Although tokenized words are not enough for determining a class of a document, learning by using all substrings has a prohibitive computational cost because the number of all candidate substrings can be very large. We show that the idea of equivalent classes of substrings can help determine all effective substrings exhaustively in linear time. Moreover, by applying L1 regularization to our model, we obtain a compact result, which makes an inference extremely efficient in time and space, and robust even if we use substrings of all lengths. In experiments, we show that our method can extract effective substrings efficiently, and achieved more accurate results and the its inference was faster than the results using previous methods.
Daisuke Okanohara, Jun'ichi Tsujii
SDM1
2009 A Linear-Time Burrows-Wheeler Transform Using Induced Sorting
Daisuke Okanohara, Kunihiko Sadakane
SPIRE1
2008 Modeling Latent-Dynamic in Shallow Parsing: A Latent Conditional Model with Imrpoved Inference
Xu Sun 0001, Louis-Philippe Morency, Daisuke Okanohara, Yoshimasa Tsuruoka, Jun'ichi Tsujii
COLING3
2008 An Online Algorithm for Finding the Longest Previous Factors
Daisuke Okanohara, Kunihiko Sadakane
ESA1
2008 New challenges for text mining: mapping between text and manually curated pathways
abstract
BACKGROUND: Associating literature with pathways poses new challenges to the Text Mining (TM) community. There are three main challenges to this task: (1) the identification of the mapping position of a specific entity or reaction in a given pathway, (2) the recognition of the causal relationships among multiple reactions, and (3) the formulation and implementation of required inferences based on biological domain knowledge. RESULTS: To address these challenges, we constructed new resources to link the text with a model pathway; they are: the GENIA pathway corpus with event annotation and NF-kB pathway. Through their detailed analysis, we address the untapped resource, 'bio-inference,' as well as the differences between text and pathway representation. Here, we show the precise comparisons of their representations and the nine classes of 'bio-inference' schemes observed in the pathway corpus. CONCLUSIONS: We believe that the creation of such rich resources and their detailed analysis is the significant first step for accelerating the research of the automatic construction of pathway from text.
Kanae Oda, Jin-Dong Kim, Tomoko Ohta, Daisuke Okanohara, Takuya Matsuzaki, Yuka Tateisi, Jun'ichi Tsujii
BMC Bioinform.4
2007 A discriminative language model with pseudo-negative samples
Daisuke Okanohara, Jun'ichi Tsujii
ACL1
2007 Practical Entropy-Compressed Rank/Select Dictionary
abstract
Rank/Select dictionaries are data structures for an ordered set S ⊂ {0,1,…, n − 1} to compute runk(x, S) (the number of elements in S that are no greater than x), and select(i, S) (the i-th smallest element in S), which are the fundamental components of succinct data structures of strings, trees, graphs, etc‥ In these data structures, however, only asymptotic behavior has been considered and their performance for real data is not satisfactory. In this paper, we propose four novel Rank/Select dictionaries: esp, recrank, vcode and sdarray, each of which is small if the number of elements in S is small, and indeed close to nH0(S) (H0(S) < 1 is the zero-th order empirical entropy of S) in practice. Furthermore, their query times are superior to those of existing structures. Experimental results reveal the characteristics of our data structures and also show that these data structures are superior to existing implementations, both in terms of size and query time.
Daisuke Okanohara, Kunihiko Sadakane
ALENEX1
2006 Improving the Scalability of Semi-Markov Conditional Random Fields for Named Entity Recognition
abstract
This paper presents techniques to apply semi-CRFs to Named Entity Recognition tasks with a tractable computational cost. Our framework can handle an NER task that has long named entities and many labels which increase the computational cost. To reduce the computational cost, we propose two techniques: the first is the use of feature forests, which enables us to pack feature-equivalent states, and the second is the introduction of a filtering process which significantly reduces the number of candidate states. This framework allows us to use a rich set of features extracted from the chunk-based representation that can capture informative characteristics of entities. We also introduce a simple trick to transfer information about distant entities by embedding label information into non-entity labels. Experimental results show that our model achieves an F-score of 71.48% on the JNLPBA 2004 shared task without using any external resources or post-processing techniques.
Daisuke Okanohara, Yusuke Miyao, Yoshimasa Tsuruoka, Jun'ichi Tsujii
ACL1
2005 Partially Decodable Compression with Static PPM
abstract
Summary form only given. We propose a novel compression method, static PPM (SP), which supports partial decode (p-decode) with low memory and computation requirement. P-decode is a function that decodes data from an arbitrary position without decoding the whole, which is critical for exploiting large data in a compressed state. While conventional compression methods do not support p-decode, recent self-indexing data structures, such as compressed suffix arrays (CSA) (Sadakane, K., J. Algorithms, vol.48, no.2, p.294-313, 2003) and FM-index (Ferragina, P. and Manzini, G. Proc. ACM-SIAM SODA, p.269-78, 2001), support p-decode. However, they have to store the whole compressed data in memory, since the order of the data is not preserved after compression. This causes a serious problem when the compressed data is larger than memory size. In contrast, SP does not have to store compressed data in memory and, thus, is the first compression method that supports p-decode for very large compressed data. In order to show SP's high compression performance inherited from PPM and its p-decode performance, we compared SP, CSA, FM-index, gzip (with option -9) and bzip (with option -9) in terms of speed and compression ratio. The results show that SP achieves fast p-decode with high compression ratio.
Daisuke Okanohara
DCC1
2005 Assigning Polarity Scores to Reviews Using Machine Learning Techniques
Daisuke Okanohara, Jun'ichi Tsujii
IJCNLP1