Takeshi Shinohara

dblp:89/21 · DBLP profile ↗
← Back
35ranked-venue papers
4as first author
4since 2021 · last 2025
0000-0002-7451-7374ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-authorDatabases, data management, data science and information retrieval · 8 · 1 since 2021Theory of computation · 8 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 7 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-authorHuman-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Human-computer interaction and pervasive computing
1 paper
Collaborative and social computing · 100%
Theoretical computer science
2 papers
Automata and formal languages · 55% Computational complexity · 31% Logic in computer science · 14%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 75% Data models and query languages · 25%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Parallel and multicore computing · 100%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Collaborative and social computing › computer-supported cooperative work
distributed collaboration
0.012000
Collaboration with Lean Media: how open-source software succeeds · CSCW 2000
Collaborative and social computing › peer production
open source software development
0.012000
Collaboration with Lean Media: how open-source software succeeds · CSCW 2000
Bioinformatics and computational biology
sequence analysis
0.011995
BONSAI Garden: Parallel Knowledge Discovery System for Amino Acid Sequences · ISMB 1995
Automata and formal languages › grammar formalisms
elementary formal systems
0.011994
Rich Classes Inferable from Positive Data: Length-Bounded Elementary Formal Systems · Inf. Comput. 1994
Computational complexity
inductive inference
0.011994
Rich Classes Inferable from Positive Data: Length-Bounded Elementary Formal Systems · Inf. Comput. 1994
Automata and formal languages
grammatical inference
0.011992
Polynomial Time Inference of a Subclass of Context-Free Transformations · COLT 1992
Information retrieval › document processing
document compression
0.011986
Efficient Storage and Retrieval of Very Large Document Databases · ICDE 1986
Data models and query languages › NoSQL database
document store
0.011986
Efficient Storage and Retrieval of Very Large Document Databases · ICDE 1986
Information retrieval
indexing
0.011986
Efficient Storage and Retrieval of Very Large Document Databases · ICDE 1986
Information retrieval › indexing
inverted file
0.011986
Efficient Storage and Retrieval of Very Large Document Databases · ICDE 1986
Logic in computer science
logic programming
0.011992
Polynomial Time Inference of a Subclass of Context-Free Transformations · COLT 1992
Logic in computer science › logic programming
prolog
0.011992
Polynomial Time Inference of a Subclass of Context-Free Transformations · COLT 1992

Methods — techniques the papers use, named apart from their topics

parallel computing · 0.0knowledge discovery · 0.0quantitative analysis · 0.0interviews · 0.0minimal multiple generalization · 0.0statistical word occurrence analysis · 0.0
YearPublicationVenuePosition
2025 Double Filtering Using Short and Long Quantized Projections
Naoya Higuchi, Yasunobu Imamura, Takeshi Shinohara, Kouichi Hirata, Tetsuji Kuboyama
SISAP3
2024 Fast Filtering for Similarity Search Using Conjunctive Enumeration of Sketches in Order of Hamming Distance
abstract
Sketches are compact bit-string representations of points, often employed for speeding up searches through the effects of dimensionality reduction and data compression. In this paper, we propose a novel sketch enumeration method and demonstrate its ability to realize fast filtering for approximate nearest neighbor search in metric spaces. Whereas the Hamming distance between the query’s sketch and sketches of points to be searched has been used for sketch prioritization traditionally, recent research has introduced asymmetric distances, enabling higher recall rates with fewer candidates. Additionally, sketch enumeration methods that speed up the filtering such that high-priority solution candidates are selected based on the priority of the sketch to the given query without the need for direct sketch comparisons have been proposed. Our primary goal in this paper is to further accelerate sketch enumeration through parallel processing. While Hamming distance-based enumeration can be parallelized relatively easily, achieving high recall rates requires a large number of candidates, and speeding up the filtering alone is insufficient for overall similarity search acceleration. Therefore, we introduce the conjunctive enumeration method, which concatenates two Hamming distance-based enumerations to approximate asymmetric distance-based enumeration. Then, we validate the effectiveness of the proposed method through experiments using large-scale public datasets. Our approach offers a significant acceleration effect, thereby enhancing the efficiency of similarity search operations.
Naoya Higuchi, Yasunobu Imamura, Vladimir Mic, Takeshi Shinohara, Kouichi Hirata, Tetsuji Kuboyama
ICPRAM4
2024 Fast Filtering by Conjunctive Enumeration of Sketches for Nearest Neighbor Search
Naoya Higuchi, Yasunobu Imamura, Vladimir Mic, Takeshi Shinohara, Kouichi Hirata, Tetsuji Kuboyama
ICPRAM4
2022 Nearest-neighbor Search from Large Datasets using Narrow Sketches
Naoya Higuchi, Yasunobu Imamura, Vladimir Mic, Takeshi Shinohara, Kouichi Hirata, Tetsuji Kuboyama
ICPRAM4
2020 Pivot Selection for Narrow Sketches by Optimization Algorithms
Naoya Higuchi, Yasunobu Imamura, Vladimir Mic, Takeshi Shinohara, Kouichi Hirata, Tetsuji Kuboyama
SISAP4
2019 Fast Nearest Neighbor Search with Narrow 16-bit Sketch
Naoya Higuchi, Yasunobu Imamura, Tetsuji Kuboyama, Kouichi Hirata, Takeshi Shinohara
ICPRAM5
2019 Annealing by Increasing Resampling in the Unified View of Simulated Annealing
abstract
Annealing by Increasing Resampling (AIR) is a stochastic hill-climbing optimization by resampling with increasing size for evaluating an objective function. In this paper, we introduce a unified view of the conventional Simulated Annealing (SA) and AIR. In this view, we generalize both SA and AIR to a stochastic hill-climbing for objective functions with stochastic fluctuations, i.e., logit and probit, respectively. Since the logit function is approximated by the probit function, we show that AIR is regarded as an approximation of SA. The experimental results on sparse pivot selection and annealing-based clustering also support that AIR is an approximation of SA. Moreover, when an objective function requires a large number of samples, AIR is much faster than SA without sacrificing the quality of the results.
Yasunobu Imamura, Naoya Higuchi, Takeshi Shinohara, Kouichi Hirata, Tetsuji Kuboyama
ICPRAM3
2018 Nearest Neighbor Search using Sketches as Quantized Images of Dimension Reduction
Naoya Higuchi, Yasunobu Imamura, Tetsuji Kuboyama, Kouichi Hirata, Takeshi Shinohara
ICPRAM5
2016 Fast Hilbert Sort Algorithm Without Using Hilbert Indices
Yasunobu Imamura, Takeshi Shinohara, Kouichi Hirata, Tetsuji Kuboyama
SISAP2
2015 High Dimensional Similarity Search with Bundled Query Processing on Hilbert R-Tree
Yohei Nasu, Naoki Kishikawa, Kei Tashima, Shin Kodama, Yasunobu Imamura, Takeshi Shinohara, Kouichi Hirata, Tetsuji Kuboyama
ICPRAM (1)6
2010 Accelerating Video Identification by Skipping Queries with a Compact Metric Cache
Takaaki Aoki, Daisuke Ninomiya, Arnoldo José Müller Molina, Takeshi Shinohara
ICCSA (4)4
2010 On the Configuration of the Similarity Search Data Structure D-Index for High Dimensional Objects
Arnoldo José Müller Molina, Takeshi Shinohara
ICCSA (3)2
2008 Foreword
John Case, Takeshi Shinohara, Thomas Zeugmann, Sandra Zilles
Theor. Comput. Sci.2
2008 Developments from enquiries into the learnability of the pattern languages from positive data
Yen Kaow Ng, Takeshi Shinohara
Theor. Comput. Sci.2
2006 Finding Consensus Patterns in Very Scarce Biosequence Samples from Their Minimal Multiple Generalizations
Yen Kaow Ng, Takeshi Shinohara
PAKDD2
2005 Inferring Unions of the Pattern Languages by the Most Fitting Covers
Yen Kaow Ng, Takeshi Shinohara
ALT2
2005 Measuring Over-Generalization in the Minimal Multiple Generalizations of Biosequences
Yen Kaow Ng, Hirotaka Ono 0001, Takeshi Shinohara
Discovery Science3
2002 Processing Text Files as Is: Pattern Matching over Compressed Texts, Multi-byte Character Texts, and Semi-structured Texts
Masayuki Takeda, Satoru Miyamoto, Takuya Kida, Ayumi Shinohara, Shuichi Fukamachi, Takeshi Shinohara, Setsuo Arikawa
SPIRE6
2001 An Efficient Derivation for Elementary Formal Systems Based on Partial Unification
Noriko Sugimoto, Hiroki Ishizaka, Takeshi Shinohara
Discovery Science3
2001 Speed-up of Aho-Corasick Pattern Matching Machines by Rearranging States
abstract
This article describes speed-up of string pattern matching by rearranging states in Aho-Corasick pattern matching machine, which is a kind of afinite automaton. We realized speed-up of string pattern matching using data compression. Although we obtain higher compression ratio using a finite state model, it doesn't lead speed-up of string pattern matching. Because the pattern matching machine becomes very large, when compression codes are complex. Random Access Memory (RAM) are scattered with states used frequently Such states are close to the initial state of pattern matching machine. We rearrange states so as to collecting states used frequently for CPU cache eficiency. We renumber states in breadth-first order. In experiments, the elapsed time is reduced to about 55% in case of a compressed English text.
T. Nishimura, Shuichi Fukamachi, Takeshi Shinohara
SPIRE3
2000 Speeding Up Pattern Matching by Text Compression
Yusuke Shibata, Takuya Kida, Shuichi Fukamachi, Masayuki Takeda, Ayumi Shinohara, Takeshi Shinohara, Setsuo Arikawa
CIAC6
2000 Collaboration with Lean Media: how open-source software succeeds
abstract
Open-source software, usually created by volunteer programmers dispersed worldwide, now competes with that developed by software firms. This achievement is particularly impressive as open-source programmers rarely meet. They rely heavily on electronic media, which preclude the benefits of face-to-face contact that programmers enjoy within firms. In this paper, we describe findings that address this paradox based on observation, interviews and quantitative analyses of two open-source projects. The findings suggest that spontaneous work coordinated afterward is effective, rational organizational culture helps achieve agreement among members and communications media moderately support spontaneous work. These findings can imply a new model of dispersed collaboration.
Yutaka Yamauchi, Makoto Yokozawa, Takeshi Shinohara, Toru Ishida 0001
CSCW3
2000 Speed-Up of Approximate String Matching Using Lossy Compression
Shuichi Fukamachi, Takeshi Shinohara
EJC2
2000 Inductive inference of unbounded unions of pattern languages from positive data
Takeshi Shinohara, Hiroki Arimura
Theor. Comput. Sci.1
1999 H-Map: A Dimension Reduction Mapping for Approximate Retrieval of Multi-dimensional Data
Takeshi Shinohara, Hiroki Ishizaka
Discovery Science1
1998 Approximate Retrieval of High-Dimensional Data by Spatial Indexing
Takeshi Shinohara, Jiyuan An, Hiroki Ishizaka
Discovery Science1
1997 Learning Unions of Tree Patterns Using Queries
Hiroki Arimura, Hiroki Ishizaka, Takeshi Shinohara
Theor. Comput. Sci.3
1995 Learning Unions of Tree Patterns Using Queries
Hiroki Arimura, Hiroki Ishizaka, Takeshi Shinohara
ALT3
1995 Editor's Introduction
Klaus P. Jantke, Takeshi Shinohara, Thomas Zeugmann
ALT2
1995 BONSAI Garden: Parallel Knowledge Discovery System for Amino Acid Sequences
Takayoshi Shoudai, Michael Lappe, Satoru Miyano, Ayumi Shinohara, Takeo Okazaki, Setsuo Arikawa, Tomoyuki Uchida, Shinichi Shimozono, Takeshi Shinohara, Satoru Kuhara
ISMB9
1994 Finding Minimal Generalizations for Unions of Pattern Languages and Its Application to Inductive Inference from Positive Data
Hiroki Arimura, Takeshi Shinohara, Setsuko Otsuki
STACS2
1994 Rich Classes Inferable from Positive Data: Length-Bounded Elementary Formal Systems
Takeshi Shinohara
Inf. Comput.1
1992 Polynomial Time Inference of a Subclass of Context-Free Transformations
abstract
This paper deals with a class of Prolog programs, called context-free term transformations (CTF). We present a polynomial time algorithm to identify a subclass of CFT, whose program consists of at most two clauses, from positive data; The algorithm uses 2-mmg (2-minimal multiple generalization) algorithm, which is natural extension of Plotkin's least generalization algorithm, to reconstruct the pair of heads of the unknown program. Using this algorithm, we show the consistent and conservative polynomial time identifiability of the class of tree languages defined by CFTFBuniq together with tree languages defined by pairs of two tree patterns, both of which are proper subclasses of CFT, in the limit from positive data.
Hiroki Arimura, Hiroki Ishizaka, Takeshi Shinohara
COLT3
1992 Learning Elementary Formal Systems
Setsuo Arikawa, Takeshi Shinohara, Akihiro Yamamoto
Theor. Comput. Sci.2
1986 Efficient Storage and Retrieval of Very Large Document Databases
abstract
The authors have developed an information retrieval system named AIR (Augmented Information Retrieval system), which might be one of the most efficient systems for very large document databases. AIR can store the document data compactly and retrieve them quickly. The techniques bringing AIR to the high efficiency, the data compression, the quick keyword index, and the automatic keyword selection, are discussed. These techniques, which are based on the statistical properties of word occurrence, are fairly simple, so that the information retrieval systems employing them can be implemented with ease. The data compression technique reduces English text by a factor of 4. The quick keyword index decreases the average number of disk accesses to retrieve a keyword to about 0.3. The automatic keyword selection technique roughly halves both the number of different keywords and the size of the inverted file with only 2% loss of retrieval power.
Fumihiro Matsuo, Shouichi Futamura, Takeshi Shinohara
ICDE3