VLDB 2026 Research / reviewers in the wild / expert
Daisuke Okanohara
dblp:60/4701
· DBLP profile ↗
16ranked-venue papers
9as first author
0since 2021 · last 2018
0000-0001-5969-587XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 3 first-authorDatabases, data management, data science and information retrieval · 5 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-authorTheory of computation · 3 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Theoretical computer science
1 paper |
Approximation and online algorithms · 50% Algorithms and data structures · 50% | |
| Artificial intelligence
2 papers |
Language models and text generation · 54% Information extraction and text analysis · 46% | |
| Databases, data mining, and information retrieval
1 paper |
Data mining · 100% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Approximation and online algorithms
online learning |
0.1 | 1 | 2009 | Latent Variable Perceptron Algorithm for Structured Classification · IJCAI 2009 |
Algorithms and data structures › learning algorithms
perceptron |
0.1 | 1 | 2009 | Latent Variable Perceptron Algorithm for Structured Classification · IJCAI 2009 |
Natural language and speech › Language models and text generation › language modeling
discriminative language modeling |
0.1 | 1 | 2007 | A discriminative language model with pseudo-negative samples · ACL 2007 |
Natural language and speech › Information extraction and text analysis
named entity recognition |
0.1 | 1 | 2006 | Improving the Scalability of Semi-Markov Conditional Random Fields for Named Entity Recognition · ACL 2006 |
Data mining › predictive modeling › classification
structured classification |
0.0 | 1 | 2009 | Latent Variable Perceptron Algorithm for Structured Classification · IJCAI 2009 |
Methods — techniques the papers use, named apart from their topics
perceptron algorithm · 0.2latent variable models · 0.1latent variable model · 0.1pseudo-negative sampling · 0.1semi-markov conditional random field · 0.1feature forests · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2018 | Neural Multi-scale Image Compression
Ken Nakanishi, Shin-ichi Maeda, Takeru Miyato, Daisuke Okanohara |
ACCV (6) | 4 |
| 2017 | Compressed Bit vectors Based on Variable-to-Fixed EncodingsabstractWe consider practical implementations of compressed bitvectors, which support rank and select operations on a given bit-string, while storing the bit-string in compressed form. Our approach relies on variable-to-fixed encodings of the bit-string, an approach that has not yet been considered systematically for practical encodings of bitvectors. We show that this approach leads to fast practical implementations with low redundancy (i.e. the space used by the bitvector in addition to the compressed representation of the bit-string), and is a flexible and promising solution to the problem of supporting rank and select on moderately compressible bit-strings, such as those encountered in real-world applications. Seungbum Jo, Stelios Joannou, Daisuke Okanohara, Rajeev Raman, S. Srinivasa Rao 0001 |
Comput. J. | 3 |
| 2014 | Compressed Bit Vectors Based on Variable-to-Fixed EncodingsabstractWe consider practical implementations of compressed bit vectors, which support rank and select operations on a given bit-string, while storing thebit-string in compressed form. Our approach relies on variable-to-fixed (V2F) encodings of the bit-string, an approach that has not yet been considered systematically for practical encodings of bit-vectors. This approach leadsto fast practical implementations with low redundancy (i.e., the space used by the bit vector in addition to the compressed representation of the bit-string),and is a flexible and promising solution to the problem of supporting rank and select on moderately compressible bit-strings, such as those frequently found in real-world applications. Seungbum Jo, Stelios Joannou, Daisuke Okanohara, Rajeev Raman, S. Srinivasa Rao 0001 |
DCC | 3 |
| 2011 | LGM: Mining Frequent Subgraphs from Linear Graphs
Yasuo Tabei, Daisuke Okanohara, Shuichi Hirose, Koji Tsuda |
PAKDD (2) | 2 |
| 2010 | Conjunctive Filter: Breaking the Entropy BarrierabstractWe consider a problem for storing a map that associates a key with a set of values. To store ( n) values from the m universe of size m, it requires log2 n bits of space, which can be approximated as (1.44 + n) log2 m/n bits when n ≪ m. If we allow ϵ fraction of errors in 1 outputs, we can store it with roughly n log2 ϵ bits, which matches the entropy bound. Bloom filter is a wellknown example for such data structures. Our objective is to break this entropy bound and construct more space-efficient data structures. In this paper, we propose a novel data structure called a conjunctive filter, which supports conjunctive queries on k distinct keys for fixed k. Although a conjunctive filter cannot return the set of values itself associated with a queried key, it can perform conjunctive queries with O(1 / √ m) fraction of errors. Also, the consumed space is n k log2 m bits and it is significantly smaller than the entropy bound n 2 log2 m when k ≥ 3. We will show that many problems can be solved by using a conjunctive filter such as full-text search and database join queries. Also, we conducted experiments using a real-world data set, and show that a conjunctive filter answers conjunctive queries almost correctly using about 1/2 ∼ 1/4 space as the entropy bound. 1 Daisuke Okanohara, Yuichi Yoshida |
ALENEX | 1 |
| 2009 | Latent Variable Perceptron Algorithm for Structured Classification
Xu Sun 0001, Takuya Matsuzaki, Daisuke Okanohara, Jun'ichi Tsujii |
IJCAI | 3 |
| 2009 | Text Categorization with All Substring FeaturesabstractThis paper presents a novel document classification method using all substrings as features. Although tokenized words are not enough for determining a class of a document, learning by using all substrings has a prohibitive computational cost because the number of all candidate substrings can be very large. We show that the idea of equivalent classes of substrings can help determine all effective substrings exhaustively in linear time. Moreover, by applying L1 regularization to our model, we obtain a compact result, which makes an inference extremely efficient in time and space, and robust even if we use substrings of all lengths. In experiments, we show that our method can extract effective substrings efficiently, and achieved more accurate results and the its inference was faster than the results using previous methods. Daisuke Okanohara, Jun'ichi Tsujii |
SDM | 1 |
| 2009 | A Linear-Time Burrows-Wheeler Transform Using Induced Sorting
Daisuke Okanohara, Kunihiko Sadakane |
SPIRE | 1 |
| 2008 | Modeling Latent-Dynamic in Shallow Parsing: A Latent Conditional Model with Imrpoved Inference
Xu Sun 0001, Louis-Philippe Morency, Daisuke Okanohara, Yoshimasa Tsuruoka, Jun'ichi Tsujii |
COLING | 3 |
| 2008 | An Online Algorithm for Finding the Longest Previous Factors
Daisuke Okanohara, Kunihiko Sadakane |
ESA | 1 |
| 2008 | New challenges for text mining: mapping between text and manually curated pathwaysabstractBACKGROUND: Associating literature with pathways poses new challenges to the Text Mining (TM) community. There are three main challenges to this task: (1) the identification of the mapping position of a specific entity or reaction in a given pathway, (2) the recognition of the causal relationships among multiple reactions, and (3) the formulation and implementation of required inferences based on biological domain knowledge. RESULTS: To address these challenges, we constructed new resources to link the text with a model pathway; they are: the GENIA pathway corpus with event annotation and NF-kB pathway. Through their detailed analysis, we address the untapped resource, 'bio-inference,' as well as the differences between text and pathway representation. Here, we show the precise comparisons of their representations and the nine classes of 'bio-inference' schemes observed in the pathway corpus. CONCLUSIONS: We believe that the creation of such rich resources and their detailed analysis is the significant first step for accelerating the research of the automatic construction of pathway from text. Kanae Oda, Jin-Dong Kim, Tomoko Ohta, Daisuke Okanohara, Takuya Matsuzaki, Yuka Tateisi, Jun'ichi Tsujii |
BMC Bioinform. | 4 |
| 2007 | A discriminative language model with pseudo-negative samples
Daisuke Okanohara, Jun'ichi Tsujii |
ACL | 1 |
| 2007 | Practical Entropy-Compressed Rank/Select DictionaryabstractRank/Select dictionaries are data structures for an ordered set S ⊂ {0,1,…, n − 1} to compute runk(x, S) (the number of elements in S that are no greater than x), and select(i, S) (the i-th smallest element in S), which are the fundamental components of succinct data structures of strings, trees, graphs, etc‥ In these data structures, however, only asymptotic behavior has been considered and their performance for real data is not satisfactory. In this paper, we propose four novel Rank/Select dictionaries: esp, recrank, vcode and sdarray, each of which is small if the number of elements in S is small, and indeed close to nH0(S) (H0(S) < 1 is the zero-th order empirical entropy of S) in practice. Furthermore, their query times are superior to those of existing structures. Experimental results reveal the characteristics of our data structures and also show that these data structures are superior to existing implementations, both in terms of size and query time. Daisuke Okanohara, Kunihiko Sadakane |
ALENEX | 1 |
| 2006 | Improving the Scalability of Semi-Markov Conditional Random Fields for Named Entity RecognitionabstractThis paper presents techniques to apply semi-CRFs to Named Entity Recognition tasks with a tractable computational cost. Our framework can handle an NER task that has long named entities and many labels which increase the computational cost. To reduce the computational cost, we propose two techniques: the first is the use of feature forests, which enables us to pack feature-equivalent states, and the second is the introduction of a filtering process which significantly reduces the number of candidate states. This framework allows us to use a rich set of features extracted from the chunk-based representation that can capture informative characteristics of entities. We also introduce a simple trick to transfer information about distant entities by embedding label information into non-entity labels. Experimental results show that our model achieves an F-score of 71.48% on the JNLPBA 2004 shared task without using any external resources or post-processing techniques. Daisuke Okanohara, Yusuke Miyao, Yoshimasa Tsuruoka, Jun'ichi Tsujii |
ACL | 1 |
| 2005 | Partially Decodable Compression with Static PPMabstractSummary form only given. We propose a novel compression method, static PPM (SP), which supports partial decode (p-decode) with low memory and computation requirement. P-decode is a function that decodes data from an arbitrary position without decoding the whole, which is critical for exploiting large data in a compressed state. While conventional compression methods do not support p-decode, recent self-indexing data structures, such as compressed suffix arrays (CSA) (Sadakane, K., J. Algorithms, vol.48, no.2, p.294-313, 2003) and FM-index (Ferragina, P. and Manzini, G. Proc. ACM-SIAM SODA, p.269-78, 2001), support p-decode. However, they have to store the whole compressed data in memory, since the order of the data is not preserved after compression. This causes a serious problem when the compressed data is larger than memory size. In contrast, SP does not have to store compressed data in memory and, thus, is the first compression method that supports p-decode for very large compressed data. In order to show SP's high compression performance inherited from PPM and its p-decode performance, we compared SP, CSA, FM-index, gzip (with option -9) and bzip (with option -9) in terms of speed and compression ratio. The results show that SP achieves fast p-decode with high compression ratio. Daisuke Okanohara |
DCC | 1 |
| 2005 | Assigning Polarity Scores to Reviews Using Machine Learning Techniques
Daisuke Okanohara, Jun'ichi Tsujii |
IJCNLP | 1 |