Tetsuji Kuboyama

dblp:43/3007 · DBLP profile ↗
← Back
46ranked-venue papers
2as first author
7since 2021 · last 2025
0000-0003-1590-0231ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 16 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 1 since 2021Theory of computation · 2Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1
YearPublicationVenuePosition
2025 AI-Enhanced Two-Stage Clustering for COVID-19 Vaccine Discourse Analysis: Multi-Faceted Public Reaction Assessment
Takako Hashimoto, Tetsuji Kuboyama, Masashi Toyoda, Naoki Yoshinaga 0001, Masaru Kitsuregawa, Takeaki Uno
IEEE Big Data2
2025 Double Filtering Using Short and Long Quantized Projections
Naoya Higuchi, Yasunobu Imamura, Takeshi Shinohara, Kouichi Hirata, Tetsuji Kuboyama
SISAP5
2024 Fast Filtering for Similarity Search Using Conjunctive Enumeration of Sketches in Order of Hamming Distance
abstract
Sketches are compact bit-string representations of points, often employed for speeding up searches through the effects of dimensionality reduction and data compression. In this paper, we propose a novel sketch enumeration method and demonstrate its ability to realize fast filtering for approximate nearest neighbor search in metric spaces. Whereas the Hamming distance between the query’s sketch and sketches of points to be searched has been used for sketch prioritization traditionally, recent research has introduced asymmetric distances, enabling higher recall rates with fewer candidates. Additionally, sketch enumeration methods that speed up the filtering such that high-priority solution candidates are selected based on the priority of the sketch to the given query without the need for direct sketch comparisons have been proposed. Our primary goal in this paper is to further accelerate sketch enumeration through parallel processing. While Hamming distance-based enumeration can be parallelized relatively easily, achieving high recall rates requires a large number of candidates, and speeding up the filtering alone is insufficient for overall similarity search acceleration. Therefore, we introduce the conjunctive enumeration method, which concatenates two Hamming distance-based enumerations to approximate asymmetric distance-based enumeration. Then, we validate the effectiveness of the proposed method through experiments using large-scale public datasets. Our approach offers a significant acceleration effect, thereby enhancing the efficiency of similarity search operations.
Naoya Higuchi, Yasunobu Imamura, Vladimir Mic, Takeshi Shinohara, Kouichi Hirata, Tetsuji Kuboyama
ICPRAM6
2024 Fast Filtering by Conjunctive Enumeration of Sketches for Nearest Neighbor Search
Naoya Higuchi, Yasunobu Imamura, Vladimir Mic, Takeshi Shinohara, Kouichi Hirata, Tetsuji Kuboyama
ICPRAM6
2022 Nearest-neighbor Search from Large Datasets using Narrow Sketches
Naoya Higuchi, Yasunobu Imamura, Vladimir Mic, Takeshi Shinohara, Kouichi Hirata, Tetsuji Kuboyama
ICPRAM6
2021 Random Number Generators in Training of Contextual Neural Networks
Maciej Huk, Kilho Shin 0001, Tetsuji Kuboyama, Takako Hashimoto
ACIIDS3
2021 Analyzing temporal patterns of topic diversity using graph clustering
abstract
Abstract During a disaster, social media can be both a source of help and of danger: Social media has a potential to diffuse rumors, and officials involved in disaster mitigation must react quickly to the spread of rumor on social media. In this paper, we investigate how topic diversity (i.e., homogeneity of opinions in a topic) depends on the truthfulness of a topic (whether it is a rumor or a non-rumor) and how the topic diversity changes in time after a disaster. To do so, we develop a method for quantifying the topic diversity of the tweet data based on text content. The proposed method is based on clustering a tweet graph using Data polishing that automatically determines the number of subtopics. We perform a case study of tweets posted after the East Japan Great Earthquake on March 11, 2011. We find that rumor topics exhibit more homogeneity of opinions in a topic during diffusion than non-rumor topics. Furthermore, we evaluate the performance of our method and demonstrate its improvement on the runtime for data processing over existing methods.
Takako Hashimoto, Dave Shepard 0001, Tetsuji Kuboyama, Kilho Shin 0001, Ryota Kobayashi, Takeaki Uno
J. Supercomput.3
2020 A Fast Algorithm for Unsupervised Feature Value Selection
Kilho Shin 0001, Kenta Okumoto, Dave Shepard 0001, Tetsuji Kuboyama, Takako Hashimoto, Hiroaki Ohshima
ICAART (2)4
2020 Twitter Topic Progress Visualization using Micro-clustering
Takako Hashimoto, Akira Kusaba, Dave Shepard 0001, Tetsuji Kuboyama, Kilho Shin 0001, Takeaki Uno
ICPRAM4
2020 Pivot Selection for Narrow Sketches by Optimization Algorithms
Naoya Higuchi, Yasunobu Imamura, Vladimir Mic, Takeshi Shinohara, Kouichi Hirata, Tetsuji Kuboyama
SISAP6
2020 Unsupervised Clustering based on Feature-value / Instance Transposition Selection
abstract
This paper presents FITS, or Feature-value / Instance Transposition Selection, a method for unsupervised clustering. FITS is a tractable, explicable clustering method, which leverages the unsupervised feature value selection algorithm known as UFVS in the literature. FITS combines repeated rounds of UFVS with alternating steps of matrix transposition to produce a set of homogenous clusters that describe data well. By repeatedly swapping the role of feature and instance and applying the same selection process to them, FITS leverages UFVS's speed and can perform clustering in our experiments in tens milliseconds for datasets of thousands of features and thousands of instances.We performed feature selection-based clustering on two real-world data sets. One is aimed at topic extraction from Twitter data, and the other is aimed at gaining awareness of energy conservation from time-series power consumption data. This study also proposes a novel method based on iterative feature extraction and transposition. The effectiveness of this method is shown in an application of Twitter data analysis. On the other hand, a more straightforward use of feature selection is adopted in the application of time series power consumption data analysis.
Akira Kusaba, Takako Hashimoto, Kilho Shin 0001, Dave Shepard 0001, Tetsuji Kuboyama
TENCON5
2020 Guest Editorial: Special issue on Discovery Science
Takuya Kida, Tetsuji Kuboyama, Takeaki Uno, Akihiro Yamamoto
Mach. Learn.2
2019 Fast Nearest Neighbor Search with Narrow 16-bit Sketch
Naoya Higuchi, Yasunobu Imamura, Tetsuji Kuboyama, Kouichi Hirata, Takeshi Shinohara
ICPRAM3
2019 Annealing by Increasing Resampling in the Unified View of Simulated Annealing
abstract
Annealing by Increasing Resampling (AIR) is a stochastic hill-climbing optimization by resampling with increasing size for evaluating an objective function. In this paper, we introduce a unified view of the conventional Simulated Annealing (SA) and AIR. In this view, we generalize both SA and AIR to a stochastic hill-climbing for objective functions with stochastic fluctuations, i.e., logit and probit, respectively. Since the logit function is approximated by the probit function, we show that AIR is regarded as an approximation of SA. The experimental results on sparse pivot selection and annealing-based clustering also support that AIR is an approximation of SA. Moreover, when an objective function requires a large number of samples, AIR is much faster than SA without sacrificing the quality of the results.
Yasunobu Imamura, Naoya Higuchi, Takeshi Shinohara, Kouichi Hirata, Tetsuji Kuboyama
ICPRAM5
2019 Bipartite Edge Correlation Clustering: Finding an Edge Biclique Partition from a Bipartite Graph with Minimum Disagreement
abstract
In this paper, first we formulate the problem of a bipartite edge correlation clustering which finds an edge biclique partition with the minimum disagreement from a bipartite graph, by extending the bipartite correlation clustering which finds a biclique partition. Then, we design a simple randomized algorithm for bipartite edge correlation clustering, based on the randomized algorithm of bipartite correlation clustering. Finally, we give experimental results to evaluate the algorithms from both artificial data and real data.
Mikio Mizukami, Kouichi Hirata, Tetsuji Kuboyama
ICPRAM3
2018 Nearest Neighbor Search using Sketches as Quantized Images of Dimension Reduction
Naoya Higuchi, Yasunobu Imamura, Tetsuji Kuboyama, Kouichi Hirata, Takeshi Shinohara
ICPRAM3
2017 A Context-Aware Fitness Function Based on Feature Selection for Evolutionary Learning of Characteristic Graph Patterns
Fumiya Tokuhara, Tetsuhiro Miyahara, Tetsuji Kuboyama, Yusuke Suzuki, Tomoyuki Uchida
ACIIDS (1)3
2017 Topic life cycle extraction from big Twitter data based on community detection in bipartite networks
abstract
This paper is showing a time series topic life cycle extraction from millions of Tweets using our original community detection technique in bipartite networks. We suppose that the authors role that means who belong to what topics is important to extract quality topics from social media data. We already proposed the topic extraction method that considers the relationship between the authors and the words as bipartite networks and explores the authors role by forming clusters as topics. As the next step, this paper applies our method to the time series topic life cycle detection. We extract topics in different time slots and analyze the time series of topic transition using the coherence measure that expresses the semantic accuracy of topics. The paper demonstrates that our method can detect the topic life cycle such as the growth, the conflicts and so on over time from millions of Tweets.
Takako Hashimoto, Hiroshi Okamoto, Tetsuji Kuboyama, Kilho Shin 0001
IEEE BigData3
2017 Topic Extraction on Twitter Considering Author's Role Based on Bipartite Networks
Takako Hashimoto, Tetsuji Kuboyama, Hiroshi Okamoto, Kilho Shin 0001
DS2
2017 Topic Extraction from Millions of Tweets Based on Community Detection in Bipartite Networks
abstract
Social media offers a wealth of insight into how significant topics such as the Great East Japan Earthquake, the Arab Spring, and the Boston Bombing affect individuals. The scale of available data, however, can be intimidating: during the Great East Japan Earthquake, over 8 million tweets were sent each day from Japan alone. Conventional word vector-based topic-detection techniques for social media that use Latent Semantic Analysis, Latent Dirichlet Allocation, or graph community detection often cannot extract appropriate topics from such a large volume of data with accuracy due to their space and time complexity. To alleviate this problem, we propose an effective topic extraction from millions of tweets based on community detection in bipartite networks. Our method is based on the bipartite community detection technique developed by Okamoto, one of the authors of this paper. The paper demonstrates our method effectiveness on social media analysis and identifies topics from millions of tweets after the Great East Japan Earthquake. To show our method's effectiveness, we compute the coherence measure that can evaluate the semantic accuracy and the running time, and compare the method with LDA that is the major topic model.
Takako Hashimoto, Tetsuji Kuboyama, Hiroshi Okamoto, Kilho Shin 0001
EJC2
2016 Fast Hilbert Sort Algorithm Without Using Hilbert Indices
Yasunobu Imamura, Takeshi Shinohara, Kouichi Hirata, Tetsuji Kuboyama
SISAP4
2015 Super-CWC and super-LCC: Super fast feature selection algorithms
abstract
Feature selection is a useful tool for identifying which features, or attributes, of a dataset cause or explain phenomena, and improving the efficiency and accuracy of learning algorithms for discovering such phenomena. Consequently, feature selection has been studied intensively in machine learning research. However, advanced feature selection algorithms that can avoid redundant selection of features and can detect interacting features require heavy computation in general and hence are seldom used for big data analysis. To eliminate this limitation, we tried to improve the run-time performance of two of the most advanced feature selection algorithms known in the literature. We have developed two accurate and extremely fast algorithms, namely Super CWC and Super LCC. In experiments with multiple real datasets which are actually studied in big data research, we have demonstrated that our algorithms improve the performance of their original algorithms remarkably. For example, for two datasets, one with 15,568 instances and 15,741 features and another with 200,569 instances and 99,672 features, Super-CWC performed feature selection in 1.4 seconds and in 405 seconds, respectively. This is a remarkable improvement, because it is estimated that the original algorithms would need several hours to a few ten days to perform feature selection on the same datasets.
Kilho Shin 0001, Tetsuji Kuboyama, Takako Hashimoto, Dave Shepard 0001
IEEE BigData2
2015 Tree PCA for Extracting Dominant Substructures from Labeled Rooted Trees
Tomoya Yamazaki, Akihiro Yamamoto, Tetsuji Kuboyama
Discovery Science3
2015 Classifying Nucleotide Sequences and their Positions of Influenza A Viruses through Several Kernels
Issei Hamada, Takaharu Shimada, Daiki Nakata, Kouichi Hirata, Tetsuji Kuboyama
ICPRAM (1)5
2015 High Dimensional Similarity Search with Bundled Query Processing on Hilbert R-Tree
Yohei Nasu, Naoki Kishikawa, Kei Tashima, Shin Kodama, Yasunobu Imamura, Takeshi Shinohara, Kouichi Hirata, Tetsuji Kuboyama
ICPRAM (1)8
2014 Effects of External Information on Anonymity and Role of Transparency with Example of Social Network De-anonymisation
abstract
Personal data to be used for sophisticated data-centric services is typically anonym zed to ensure the privacy of the underlying individuals. However, the degree of protection against de-anonymization is uncertain because de-anonymization has not been explicitly modelled for scientific analysis and because the information that attackers might use is not well defined given that they can use a wide variety of external information sources such as public ally accessible information on the web. We have developed a system that de-anonymizes anonymous posts on social networks with a considerably high precision rate. We analysed this system and its behaviour and, on the basis of our findings, we clarified a de-anonymization model and used it to clarify the effects of external information on anonymity. We also identified the limitation of anonymization under conditions of external information being available and clarified the role of transparency can play in controlling the use of personal data.
Haruno Kataoka, Yohei Ogawa, Isao Echizen, Tetsuji Kuboyama, Hiroshi Yoshiura
ARES4
2013 Temporal Awareness of Needs after East Japan Great Earthquake using Latent Semantic Analysis
abstract
This paper proposes a time series topic extraction method to investigate the transitions of people's needs after the East Japan Great Earthquake using latent semantic analysis. Our target data is a blog about afflicted people's needs provided by a non-profit organization in Tohoku, Japan. The method crawls blog messages, extracts terms, and forms document-term matrix over time. Then, the method adopts the latent semantic analysis and extract hidden topics (people's needs) over time. In our previous work, we already proposed the graph-based topic extraction method using the modularity measure. Our previous method could visualize topic structure transition, but could not extract clear topics. In this paper, to show the effectiveness of our proposed method, we provide the experimental results, and compare them with our previous method's results.
Takako Hashimoto, Basabi Chakraborty, Tetsuji Kuboyama, Yukari Shirota
EJC3
2012 Topic Detection about the East Japan Great Earthquake based on Emerging Modularity
abstract
Once a disaster occurs, people discuss various topics in social media such as electronic bulletin boards, SNSs and video services, and their decision-making tends to be affected by discussions in social media. Under the circumstance, a mechanism to detect topics in social media has become important. This paper targets the East Japan Great Earthquake, and proposes a time series topic detection method based on modularity measure which shows the quality of a division of a network into modules or communities. Our proposed method clarifies emerging topics from social media messages by computing the modularity and analyzing them over time, and visualizes topic structures. An experimental result by actual social media data about the East Japan Great Earthquake is also shown.
Takako Hashimoto, Tetsuji Kuboyama, Yukari Shirota
EJC2
2011 Improved MAX SNP-Hard Results for Finding an Edit Distance between Unordered Trees
Kouichi Hirata, Yoshiyuki Yamamoto, Tetsuji Kuboyama
CPM3
2011 Scalable Detection of Frequent Substrings by Grammar-Based Compression
Masaya Nakahara, Shirou Maruyama, Tetsuji Kuboyama, Hiroshi Sakamoto
Discovery Science3
2011 Graph-based Consumer Behavior Analysis from Buzz Marketing Sites
abstract
This paper proposes a method to discover consumer behavior from buzz marketing sites. For example, in 2009, the super-flu virus spawned significant effects on various product marketing domains around the globe. Using text mining technology, we found a relationship between the flu pandemic and the reluctance of consumers to buy digital single-lens reflex camera. We could easily expect more air purifiers to be sold due to flu pandemic. However, the reluctance to buy digital single-lens reflex cameras because of the flu is not something we would have expected. This paper applies text mining techniques to analyze expected and unexpected consumer behavior caused by a current topic like the flu. The unforeseen relationship between a current topic and products is modeled and visualized using a directed graph that shows implicit knowledge. Consumer behavior is further analyzed based on the time series variation of directed graph structures.
Takako Hashimoto, Tetsuji Kuboyama, Yukari Shirota
EJC2
2011 A Concept Model for Solving Bond Mathematics Problems
abstract
This paper presents a concept model for solving bond mathematics problems; it is the manner in which our knowledge concerning bond mathematics is organized into a cognitive architecture. We build a concept model for bond mathematics to achieve good organization and to integrate the knowledge for bond mathematics. The ultimate goal is to enable many students to understand the solution process without difficulty. Our concept model comprises entity-relationship diagrams, and using our concept models, students can integrate financial theories and mathematical formulas. This paper illustrates concept models for bond mathematics by showing concrete examples of word mathematics problems. It also describes our principles in developing the concept model and the descriptive power of our model.
Yukari Shirota, Takako Hashimoto, Tetsuji Kuboyama
EJC3
2011 Mapping kernels for trees
Kilho Shin 0001, Marco Cuturi, Tetsuji Kuboyama
ICML3
2011 Consumer Behavior Analysis from Buzz Marketing Sites over Time Series Concept Graphs
Tetsuji Kuboyama, Takako Hashimoto, Yukari Shirota
KES (2)1
2011 A Subpath Kernel for Rooted Unordered Trees
Daisuke Kimura, Tetsuji Kuboyama, Tetsuo Shibuya, Hisashi Kashima
PAKDD (1)2
2010 Approximating Tree Edit Distance through String Edit Distance for Binary Tree Codes
abstract
This article proposes an approximation of the tree edit distance through the string edit distance for binary tree codes, instead of for Euler strings introduced by Akutsu (2006). Here, a binary tree code is a string obtained by traversing a binary tree representation with two kinds of dummy nodes of a tree in preorder. Then, we show that σ/2 ≤ τ ≤ (h + 1)σ + h, where τ is the tree edit distance between trees, and σ is the string edit distance between their binary tree codes and h is the minimum height of the trees.
Taku Aratsu, Kouichi Hirata, Tetsuji Kuboyama
Fundam. Informaticae3
2010 A Generalization of Haussler's Convolution Kernel - Mapping Kernel and Its Application to Tree Kernels
Kilho Shin 0001, Tetsuji Kuboyama
J. Comput. Sci. Technol.2
2009 Extracting Research Communities by Improved Maximum Flow Algorithm
Toshihiko Horiike, Youhei Takahashi, Tetsuji Kuboyama, Hiroshi Sakamoto
KES (2)3
2009 Discovering Networks for Global Propagation of Influenza A (H3N2) Viruses by Clustering
Kazuya Sata, Kouichi Hirata, Kimihito Ito, Tetsuji Kuboyama
KES (2)4
2009 Approximating Tree Edit Distance through String Edit Distance for Binary Tree Codes
Taku Aratsu, Kouichi Hirata, Tetsuji Kuboyama
SOFSEM3
2009 Polynomial summaries of positive semidefinite kernels
Kilho Shin 0001, Tetsuji Kuboyama
Theor. Comput. Sci.2
2008 A generalization of Haussler's convolution kernel: mapping kernel
abstract
Haussler's convolution kernel provides a successful framework for engineering new positive semidefinite kernels, and has been applied to a wide range of data types and applications. In the framework, each data object represents a finite set of finer grained components. Then, Haussler's convolution kernel takes a pair of data objects as input, and returns the sum of the return values of the predetermined primitive positive semidefinite kernel calculated for all the possible pairs of the components of the input data objects. On the other hand, the mapping kernel that we introduce in this paper is a natural generalization of Haussler's convolution kernel, in that the input to the primitive kernel moves over a predetermined subset rather than the entire cross product. Although we have plural instances of the mapping kernel in the literature, their positive semidefiniteness was investigated in case-by-case manners, and worse yet, was sometimes incorrectly concluded. In fact, there exists a simple and easily checkable necessary and sufficient condition, which is generic in the sense that it enables us to investigate the positive semidefiniteness of an arbitrary instance of the mapping kernel. This is the first paper that presents and proves the validity of the condition. In addition, we introduce two important instances of the mapping kernel, which we refer to as the size-of-index-structure-distribution kernel and the editcost-distribution kernel. Both of them are naturally derived from well known (dis)similarity measurements in the literature (e.g. the maximum agreement tree, the edit distance), and are reasonably expected to improve the performance of the existing measures by evaluating their distributional features rather than their peak (maximum/minimum) features.
Kilho Shin 0001, Tetsuji Kuboyama
ICML2
2008 An Efficient Unordered Tree Kernel and Its Application to Glycan Classification
Tetsuji Kuboyama, Kouichi Hirata, Kiyoko F. Aoki-Kinoshita
PAKDD1
2007 Polynomial Summaries of Positive Semidefinite Kernels
Kilho Shin 0001, Tetsuji Kuboyama
ALT2
2005 The q-Gram Distance for Ordered Unlabeled Trees
Nobuhito Ohkura, Kouichi Hirata, Tetsuji Kuboyama, Masateru Harao
Discovery Science3
1999 KD-FGS: A Knowledge Discovery System from Graph Data Using Formal Graph System
Tetsuhiro Miyahara, Tomoyuki Uchida, Tetsuji Kuboyama, Tasuya Yamamoto, Kenichi Takahashi, Hiroaki Ueda
PAKDD3