EDBT 2026 Demo / reviewers in the wild / expert
Tetsuji Kuboyama
dblp:43/3007
· DBLP profile ↗
46ranked-venue papers
2as first author
7since 2021 · last 2025
0000-0003-1590-0231ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 16 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 1 since 2021Theory of computation · 2Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AI-Enhanced Two-Stage Clustering for COVID-19 Vaccine Discourse Analysis: Multi-Faceted Public Reaction Assessment
Takako Hashimoto, Tetsuji Kuboyama, Masashi Toyoda, Naoki Yoshinaga 0001, Masaru Kitsuregawa, Takeaki Uno |
IEEE Big Data | 2 |
| 2025 | Double Filtering Using Short and Long Quantized Projections
Naoya Higuchi, Yasunobu Imamura, Takeshi Shinohara, Kouichi Hirata, Tetsuji Kuboyama |
SISAP | 5 |
| 2024 | Fast Filtering for Similarity Search Using Conjunctive Enumeration of Sketches in Order of Hamming DistanceabstractSketches are compact bit-string representations of points, often employed for speeding up searches through the effects of dimensionality reduction and data compression. In this paper, we propose a novel sketch enumeration method and demonstrate its ability to realize fast filtering for approximate nearest neighbor search in metric spaces. Whereas the Hamming distance between the query’s sketch and sketches of points to be searched has been used for sketch prioritization traditionally, recent research has introduced asymmetric distances, enabling higher recall rates with fewer candidates. Additionally, sketch enumeration methods that speed up the filtering such that high-priority solution candidates are selected based on the priority of the sketch to the given query without the need for direct sketch comparisons have been proposed. Our primary goal in this paper is to further accelerate sketch enumeration through parallel processing. While Hamming distance-based enumeration can be parallelized relatively easily, achieving high recall rates requires a large number of candidates, and speeding up the filtering alone is insufficient for overall similarity search acceleration. Therefore, we introduce the conjunctive enumeration method, which concatenates two Hamming distance-based enumerations to approximate asymmetric distance-based enumeration. Then, we validate the effectiveness of the proposed method through experiments using large-scale public datasets. Our approach offers a significant acceleration effect, thereby enhancing the efficiency of similarity search operations. Naoya Higuchi, Yasunobu Imamura, Vladimir Mic, Takeshi Shinohara, Kouichi Hirata, Tetsuji Kuboyama |
ICPRAM | 6 |
| 2024 | Fast Filtering by Conjunctive Enumeration of Sketches for Nearest Neighbor Search
Naoya Higuchi, Yasunobu Imamura, Vladimir Mic, Takeshi Shinohara, Kouichi Hirata, Tetsuji Kuboyama |
ICPRAM | 6 |
| 2022 | Nearest-neighbor Search from Large Datasets using Narrow Sketches
Naoya Higuchi, Yasunobu Imamura, Vladimir Mic, Takeshi Shinohara, Kouichi Hirata, Tetsuji Kuboyama |
ICPRAM | 6 |
| 2021 | Random Number Generators in Training of Contextual Neural Networks
Maciej Huk, Kilho Shin 0001, Tetsuji Kuboyama, Takako Hashimoto |
ACIIDS | 3 |
| 2021 | Analyzing temporal patterns of topic diversity using graph clusteringabstractAbstract During a disaster, social media can be both a source of help and of danger: Social media has a potential to diffuse rumors, and officials involved in disaster mitigation must react quickly to the spread of rumor on social media. In this paper, we investigate how topic diversity (i.e., homogeneity of opinions in a topic) depends on the truthfulness of a topic (whether it is a rumor or a non-rumor) and how the topic diversity changes in time after a disaster. To do so, we develop a method for quantifying the topic diversity of the tweet data based on text content. The proposed method is based on clustering a tweet graph using Data polishing that automatically determines the number of subtopics. We perform a case study of tweets posted after the East Japan Great Earthquake on March 11, 2011. We find that rumor topics exhibit more homogeneity of opinions in a topic during diffusion than non-rumor topics. Furthermore, we evaluate the performance of our method and demonstrate its improvement on the runtime for data processing over existing methods. Takako Hashimoto, Dave Shepard 0001, Tetsuji Kuboyama, Kilho Shin 0001, Ryota Kobayashi, Takeaki Uno |
J. Supercomput. | 3 |
| 2020 | A Fast Algorithm for Unsupervised Feature Value Selection
Kilho Shin 0001, Kenta Okumoto, Dave Shepard 0001, Tetsuji Kuboyama, Takako Hashimoto, Hiroaki Ohshima |
ICAART (2) | 4 |
| 2020 | Twitter Topic Progress Visualization using Micro-clustering
Takako Hashimoto, Akira Kusaba, Dave Shepard 0001, Tetsuji Kuboyama, Kilho Shin 0001, Takeaki Uno |
ICPRAM | 4 |
| 2020 | Pivot Selection for Narrow Sketches by Optimization Algorithms
Naoya Higuchi, Yasunobu Imamura, Vladimir Mic, Takeshi Shinohara, Kouichi Hirata, Tetsuji Kuboyama |
SISAP | 6 |
| 2020 | Unsupervised Clustering based on Feature-value / Instance Transposition SelectionabstractThis paper presents FITS, or Feature-value / Instance Transposition Selection, a method for unsupervised clustering. FITS is a tractable, explicable clustering method, which leverages the unsupervised feature value selection algorithm known as UFVS in the literature. FITS combines repeated rounds of UFVS with alternating steps of matrix transposition to produce a set of homogenous clusters that describe data well. By repeatedly swapping the role of feature and instance and applying the same selection process to them, FITS leverages UFVS's speed and can perform clustering in our experiments in tens milliseconds for datasets of thousands of features and thousands of instances.We performed feature selection-based clustering on two real-world data sets. One is aimed at topic extraction from Twitter data, and the other is aimed at gaining awareness of energy conservation from time-series power consumption data. This study also proposes a novel method based on iterative feature extraction and transposition. The effectiveness of this method is shown in an application of Twitter data analysis. On the other hand, a more straightforward use of feature selection is adopted in the application of time series power consumption data analysis. Akira Kusaba, Takako Hashimoto, Kilho Shin 0001, Dave Shepard 0001, Tetsuji Kuboyama |
TENCON | 5 |
| 2020 | Guest Editorial: Special issue on Discovery Science
Takuya Kida, Tetsuji Kuboyama, Takeaki Uno, Akihiro Yamamoto |
Mach. Learn. | 2 |
| 2019 | Fast Nearest Neighbor Search with Narrow 16-bit Sketch
Naoya Higuchi, Yasunobu Imamura, Tetsuji Kuboyama, Kouichi Hirata, Takeshi Shinohara |
ICPRAM | 3 |
| 2019 | Annealing by Increasing Resampling in the Unified View of Simulated AnnealingabstractAnnealing by Increasing Resampling (AIR) is a stochastic hill-climbing optimization by resampling with increasing size for evaluating an objective function. In this paper, we introduce a unified view of the conventional Simulated Annealing (SA) and AIR. In this view, we generalize both SA and AIR to a stochastic hill-climbing for objective functions with stochastic fluctuations, i.e., logit and probit, respectively. Since the logit function is approximated by the probit function, we show that AIR is regarded as an approximation of SA. The experimental results on sparse pivot selection and annealing-based clustering also support that AIR is an approximation of SA. Moreover, when an objective function requires a large number of samples, AIR is much faster than SA without sacrificing the quality of the results. Yasunobu Imamura, Naoya Higuchi, Takeshi Shinohara, Kouichi Hirata, Tetsuji Kuboyama |
ICPRAM | 5 |
| 2019 | Bipartite Edge Correlation Clustering: Finding an Edge Biclique Partition from a Bipartite Graph with Minimum DisagreementabstractIn this paper, first we formulate the problem of a bipartite edge correlation clustering which finds an edge biclique partition with the minimum disagreement from a bipartite graph, by extending the bipartite correlation clustering which finds a biclique partition. Then, we design a simple randomized algorithm for bipartite edge correlation clustering, based on the randomized algorithm of bipartite correlation clustering. Finally, we give experimental results to evaluate the algorithms from both artificial data and real data. Mikio Mizukami, Kouichi Hirata, Tetsuji Kuboyama |
ICPRAM | 3 |
| 2018 | Nearest Neighbor Search using Sketches as Quantized Images of Dimension Reduction
Naoya Higuchi, Yasunobu Imamura, Tetsuji Kuboyama, Kouichi Hirata, Takeshi Shinohara |
ICPRAM | 3 |
| 2017 | A Context-Aware Fitness Function Based on Feature Selection for Evolutionary Learning of Characteristic Graph Patterns
Fumiya Tokuhara, Tetsuhiro Miyahara, Tetsuji Kuboyama, Yusuke Suzuki, Tomoyuki Uchida |
ACIIDS (1) | 3 |
| 2017 | Topic life cycle extraction from big Twitter data based on community detection in bipartite networksabstractThis paper is showing a time series topic life cycle extraction from millions of Tweets using our original community detection technique in bipartite networks. We suppose that the authors role that means who belong to what topics is important to extract quality topics from social media data. We already proposed the topic extraction method that considers the relationship between the authors and the words as bipartite networks and explores the authors role by forming clusters as topics. As the next step, this paper applies our method to the time series topic life cycle detection. We extract topics in different time slots and analyze the time series of topic transition using the coherence measure that expresses the semantic accuracy of topics. The paper demonstrates that our method can detect the topic life cycle such as the growth, the conflicts and so on over time from millions of Tweets. Takako Hashimoto, Hiroshi Okamoto, Tetsuji Kuboyama, Kilho Shin 0001 |
IEEE BigData | 3 |
| 2017 | Topic Extraction on Twitter Considering Author's Role Based on Bipartite Networks
Takako Hashimoto, Tetsuji Kuboyama, Hiroshi Okamoto, Kilho Shin 0001 |
DS | 2 |
| 2017 | Topic Extraction from Millions of Tweets Based on Community Detection in Bipartite NetworksabstractSocial media offers a wealth of insight into how significant topics such as the Great East Japan Earthquake, the Arab Spring, and the Boston Bombing affect individuals. The scale of available data, however, can be intimidating: during the Great East Japan Earthquake, over 8 million tweets were sent each day from Japan alone. Conventional word vector-based topic-detection techniques for social media that use Latent Semantic Analysis, Latent Dirichlet Allocation, or graph community detection often cannot extract appropriate topics from such a large volume of data with accuracy due to their space and time complexity. To alleviate this problem, we propose an effective topic extraction from millions of tweets based on community detection in bipartite networks. Our method is based on the bipartite community detection technique developed by Okamoto, one of the authors of this paper. The paper demonstrates our method effectiveness on social media analysis and identifies topics from millions of tweets after the Great East Japan Earthquake. To show our method's effectiveness, we compute the coherence measure that can evaluate the semantic accuracy and the running time, and compare the method with LDA that is the major topic model. Takako Hashimoto, Tetsuji Kuboyama, Hiroshi Okamoto, Kilho Shin 0001 |
EJC | 2 |
| 2016 | Fast Hilbert Sort Algorithm Without Using Hilbert Indices
Yasunobu Imamura, Takeshi Shinohara, Kouichi Hirata, Tetsuji Kuboyama |
SISAP | 4 |
| 2015 | Super-CWC and super-LCC: Super fast feature selection algorithmsabstractFeature selection is a useful tool for identifying which features, or attributes, of a dataset cause or explain phenomena, and improving the efficiency and accuracy of learning algorithms for discovering such phenomena. Consequently, feature selection has been studied intensively in machine learning research. However, advanced feature selection algorithms that can avoid redundant selection of features and can detect interacting features require heavy computation in general and hence are seldom used for big data analysis. To eliminate this limitation, we tried to improve the run-time performance of two of the most advanced feature selection algorithms known in the literature. We have developed two accurate and extremely fast algorithms, namely Super CWC and Super LCC. In experiments with multiple real datasets which are actually studied in big data research, we have demonstrated that our algorithms improve the performance of their original algorithms remarkably. For example, for two datasets, one with 15,568 instances and 15,741 features and another with 200,569 instances and 99,672 features, Super-CWC performed feature selection in 1.4 seconds and in 405 seconds, respectively. This is a remarkable improvement, because it is estimated that the original algorithms would need several hours to a few ten days to perform feature selection on the same datasets. Kilho Shin 0001, Tetsuji Kuboyama, Takako Hashimoto, Dave Shepard 0001 |
IEEE BigData | 2 |
| 2015 | Tree PCA for Extracting Dominant Substructures from Labeled Rooted Trees
Tomoya Yamazaki, Akihiro Yamamoto, Tetsuji Kuboyama |
Discovery Science | 3 |
| 2015 | Classifying Nucleotide Sequences and their Positions of Influenza A Viruses through Several Kernels
Issei Hamada, Takaharu Shimada, Daiki Nakata, Kouichi Hirata, Tetsuji Kuboyama |
ICPRAM (1) | 5 |
| 2015 | High Dimensional Similarity Search with Bundled Query Processing on Hilbert R-Tree
Yohei Nasu, Naoki Kishikawa, Kei Tashima, Shin Kodama, Yasunobu Imamura, Takeshi Shinohara, Kouichi Hirata, Tetsuji Kuboyama |
ICPRAM (1) | 8 |
| 2014 | Effects of External Information on Anonymity and Role of Transparency with Example of Social Network De-anonymisationabstractPersonal data to be used for sophisticated data-centric services is typically anonym zed to ensure the privacy of the underlying individuals. However, the degree of protection against de-anonymization is uncertain because de-anonymization has not been explicitly modelled for scientific analysis and because the information that attackers might use is not well defined given that they can use a wide variety of external information sources such as public ally accessible information on the web. We have developed a system that de-anonymizes anonymous posts on social networks with a considerably high precision rate. We analysed this system and its behaviour and, on the basis of our findings, we clarified a de-anonymization model and used it to clarify the effects of external information on anonymity. We also identified the limitation of anonymization under conditions of external information being available and clarified the role of transparency can play in controlling the use of personal data. Haruno Kataoka, Yohei Ogawa, Isao Echizen, Tetsuji Kuboyama, Hiroshi Yoshiura |
ARES | 4 |
| 2013 | Temporal Awareness of Needs after East Japan Great Earthquake using Latent Semantic AnalysisabstractThis paper proposes a time series topic extraction method to investigate the transitions of people's needs after the East Japan Great Earthquake using latent semantic analysis. Our target data is a blog about afflicted people's needs provided by a non-profit organization in Tohoku, Japan. The method crawls blog messages, extracts terms, and forms document-term matrix over time. Then, the method adopts the latent semantic analysis and extract hidden topics (people's needs) over time. In our previous work, we already proposed the graph-based topic extraction method using the modularity measure. Our previous method could visualize topic structure transition, but could not extract clear topics. In this paper, to show the effectiveness of our proposed method, we provide the experimental results, and compare them with our previous method's results. Takako Hashimoto, Basabi Chakraborty, Tetsuji Kuboyama, Yukari Shirota |
EJC | 3 |
| 2012 | Topic Detection about the East Japan Great Earthquake based on Emerging ModularityabstractOnce a disaster occurs, people discuss various topics in social media such as electronic bulletin boards, SNSs and video services, and their decision-making tends to be affected by discussions in social media. Under the circumstance, a mechanism to detect topics in social media has become important. This paper targets the East Japan Great Earthquake, and proposes a time series topic detection method based on modularity measure which shows the quality of a division of a network into modules or communities. Our proposed method clarifies emerging topics from social media messages by computing the modularity and analyzing them over time, and visualizes topic structures. An experimental result by actual social media data about the East Japan Great Earthquake is also shown. Takako Hashimoto, Tetsuji Kuboyama, Yukari Shirota |
EJC | 2 |
| 2011 | Improved MAX SNP-Hard Results for Finding an Edit Distance between Unordered Trees
Kouichi Hirata, Yoshiyuki Yamamoto, Tetsuji Kuboyama |
CPM | 3 |
| 2011 | Scalable Detection of Frequent Substrings by Grammar-Based Compression
Masaya Nakahara, Shirou Maruyama, Tetsuji Kuboyama, Hiroshi Sakamoto |
Discovery Science | 3 |
| 2011 | Graph-based Consumer Behavior Analysis from Buzz Marketing SitesabstractThis paper proposes a method to discover consumer behavior from buzz marketing sites. For example, in 2009, the super-flu virus spawned significant effects on various product marketing domains around the globe. Using text mining technology, we found a relationship between the flu pandemic and the reluctance of consumers to buy digital single-lens reflex camera. We could easily expect more air purifiers to be sold due to flu pandemic. However, the reluctance to buy digital single-lens reflex cameras because of the flu is not something we would have expected. This paper applies text mining techniques to analyze expected and unexpected consumer behavior caused by a current topic like the flu. The unforeseen relationship between a current topic and products is modeled and visualized using a directed graph that shows implicit knowledge. Consumer behavior is further analyzed based on the time series variation of directed graph structures. Takako Hashimoto, Tetsuji Kuboyama, Yukari Shirota |
EJC | 2 |
| 2011 | A Concept Model for Solving Bond Mathematics ProblemsabstractThis paper presents a concept model for solving bond mathematics problems; it is the manner in which our knowledge concerning bond mathematics is organized into a cognitive architecture. We build a concept model for bond mathematics to achieve good organization and to integrate the knowledge for bond mathematics. The ultimate goal is to enable many students to understand the solution process without difficulty. Our concept model comprises entity-relationship diagrams, and using our concept models, students can integrate financial theories and mathematical formulas. This paper illustrates concept models for bond mathematics by showing concrete examples of word mathematics problems. It also describes our principles in developing the concept model and the descriptive power of our model. Yukari Shirota, Takako Hashimoto, Tetsuji Kuboyama |
EJC | 3 |
| 2011 | Mapping kernels for trees
Kilho Shin 0001, Marco Cuturi, Tetsuji Kuboyama |
ICML | 3 |
| 2011 | Consumer Behavior Analysis from Buzz Marketing Sites over Time Series Concept Graphs
Tetsuji Kuboyama, Takako Hashimoto, Yukari Shirota |
KES (2) | 1 |
| 2011 | A Subpath Kernel for Rooted Unordered Trees
Daisuke Kimura, Tetsuji Kuboyama, Tetsuo Shibuya, Hisashi Kashima |
PAKDD (1) | 2 |
| 2010 | Approximating Tree Edit Distance through String Edit Distance for Binary Tree CodesabstractThis article proposes an approximation of the tree edit distance through the string edit distance for binary tree codes, instead of for Euler strings introduced by Akutsu (2006). Here, a binary tree code is a string obtained by traversing a binary tree representation with two kinds of dummy nodes of a tree in preorder. Then, we show that σ/2 ≤ τ ≤ (h + 1)σ + h, where τ is the tree edit distance between trees, and σ is the string edit distance between their binary tree codes and h is the minimum height of the trees. Taku Aratsu, Kouichi Hirata, Tetsuji Kuboyama |
Fundam. Informaticae | 3 |
| 2010 | A Generalization of Haussler's Convolution Kernel - Mapping Kernel and Its Application to Tree Kernels
Kilho Shin 0001, Tetsuji Kuboyama |
J. Comput. Sci. Technol. | 2 |
| 2009 | Extracting Research Communities by Improved Maximum Flow Algorithm
Toshihiko Horiike, Youhei Takahashi, Tetsuji Kuboyama, Hiroshi Sakamoto |
KES (2) | 3 |
| 2009 | Discovering Networks for Global Propagation of Influenza A (H3N2) Viruses by Clustering
Kazuya Sata, Kouichi Hirata, Kimihito Ito, Tetsuji Kuboyama |
KES (2) | 4 |
| 2009 | Approximating Tree Edit Distance through String Edit Distance for Binary Tree Codes
Taku Aratsu, Kouichi Hirata, Tetsuji Kuboyama |
SOFSEM | 3 |
| 2009 | Polynomial summaries of positive semidefinite kernels
Kilho Shin 0001, Tetsuji Kuboyama |
Theor. Comput. Sci. | 2 |
| 2008 | A generalization of Haussler's convolution kernel: mapping kernelabstractHaussler's convolution kernel provides a successful framework for engineering new positive semidefinite kernels, and has been applied to a wide range of data types and applications. In the framework, each data object represents a finite set of finer grained components. Then, Haussler's convolution kernel takes a pair of data objects as input, and returns the sum of the return values of the predetermined primitive positive semidefinite kernel calculated for all the possible pairs of the components of the input data objects. On the other hand, the mapping kernel that we introduce in this paper is a natural generalization of Haussler's convolution kernel, in that the input to the primitive kernel moves over a predetermined subset rather than the entire cross product. Although we have plural instances of the mapping kernel in the literature, their positive semidefiniteness was investigated in case-by-case manners, and worse yet, was sometimes incorrectly concluded. In fact, there exists a simple and easily checkable necessary and sufficient condition, which is generic in the sense that it enables us to investigate the positive semidefiniteness of an arbitrary instance of the mapping kernel. This is the first paper that presents and proves the validity of the condition. In addition, we introduce two important instances of the mapping kernel, which we refer to as the size-of-index-structure-distribution kernel and the editcost-distribution kernel. Both of them are naturally derived from well known (dis)similarity measurements in the literature (e.g. the maximum agreement tree, the edit distance), and are reasonably expected to improve the performance of the existing measures by evaluating their distributional features rather than their peak (maximum/minimum) features. Kilho Shin 0001, Tetsuji Kuboyama |
ICML | 2 |
| 2008 | An Efficient Unordered Tree Kernel and Its Application to Glycan Classification
Tetsuji Kuboyama, Kouichi Hirata, Kiyoko F. Aoki-Kinoshita |
PAKDD | 1 |
| 2007 | Polynomial Summaries of Positive Semidefinite Kernels
Kilho Shin 0001, Tetsuji Kuboyama |
ALT | 2 |
| 2005 | The q-Gram Distance for Ordered Unlabeled Trees
Nobuhito Ohkura, Kouichi Hirata, Tetsuji Kuboyama, Masateru Harao |
Discovery Science | 3 |
| 1999 | KD-FGS: A Knowledge Discovery System from Graph Data Using Formal Graph System
Tetsuhiro Miyahara, Tomoyuki Uchida, Tetsuji Kuboyama, Tasuya Yamamoto, Kenichi Takahashi, Hiroaki Ueda |
PAKDD | 3 |