EDBT 2026 Demo / reviewers in the wild / expert
Takashi Washio
dblp:w/TakashiWashio
· DBLP profile ↗
46ranked-venue papers in the field
7as first author
4since 2021 · last 2023
0000-0001-6172-6401ORCID · corroborated
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 41 (6 first)Other / Interdisciplinary · 3 (1 first)Database Systems & Data Management · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Isolation Kernel Estimators
Kai Ming Ting, Takashi Washio, Jonathan R. Wells, Hang Zhang 0003, Ye Zhu 0002 |
Knowl. Inf. Syst. | 2 |
| 2023 | Isolation Distributional Kernel: A New Tool for Point and Group Anomaly DetectionsabstractWe introduce Isolation Distributional Kernel as a new way to measure the similarity between two distributions. Existing approaches based on kernel mean embedding, which convert a point kernel to a distributional kernel, have two key issues: the point kernel employed has a feature map with intractable dimensionality; and it is {\em data independent}. This paper shows that Isolation Distributional Kernel (IDK), which is based on a {\em data dependent} point kernel, addresses both key issues. We demonstrate IDK's efficacy and efficiency as a new tool for kernel based anomaly detection for both point and group anomalies. Without explicit learning, using IDK alone outperforms existing kernel based point anomaly detector OCSVM and other kernel mean embedding methods that rely on Gaussian kernel. For group anomaly detection,we introduce an IDK based detector called IDK$^2$. It reformulates the problem of group anomaly detection in input space into the problem of point anomaly detection in Hilbert space, without the need for learning. IDK$^2$ runs orders of magnitude faster than group anomaly detector OCSMM.We reveal for the first time that an effective kernel based anomaly detector based on kernel mean embedding must employ a characteristic kernel which is data dependent. Kai Ming Ting, Bi-Cun Xu, Takashi Washio, Zhi-Hua Zhou |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2021 | Isolation Kernel Density EstimationabstractThis paper shows that adaptive kernel density estimator (KDE) can be derived effectively from Isolation Kernel. Existing adaptive KDEs often employ a data independent kernel such as Gaussian kernel. Therefore, it requires an additional means to adapt its bandwidth locally in a given dataset. Because Isolation Kernel is a data dependent kernel which is derived directly from data, no additional adaptive operation is required. The resultant estimator called IKDE is the only KDE that is fast and adaptive. Existing KDEs are either fast but non-adaptive or adaptive but slow. In addition, using IKDE for anomaly detection, we identify two advantages of IKDE over LOF (Local Outlier Factor), contributing to significantly faster runtime. Kai Ming Ting, Takashi Washio, Jonathan R. Wells, Hang Zhang 0003 |
ICDM | 2 |
| 2021 | Isolation kernel: the X factor in efficient and effective large scale online kernel learning
Kai Ming Ting, Jonathan R. Wells, Takashi Washio |
Data Min. Knowl. Discov. | 3 |
| 2020 | Isolation Distributional Kernel: A New Tool for Kernel based Anomaly DetectionabstractWe introduce Isolation Distributional Kernel as a new way to measure the similarity between two distributions. Existing approaches based on kernel mean embedding, which converts a point kernel to a distributional kernel, have two key issues: the point kernel employed has a feature map with intractable dimensionality; and it is data independent. This paper shows that Isolation Distributional Kernel (IDK), which is based on a data dependent point kernel, addresses both key issues. We demonstrate IDK's efficacy and efficiency as a new tool for kernel based anomaly detection. Without explicit learning, using IDK alone outperforms existing kernel based anomaly detector OCSVM and other kernel mean embedding methods that rely on Gaussian kernel. We reveal for the first time that an effective kernel based anomaly detector based on kernel mean embedding must employ a characteristic kernel which is data dependent. Kai Ming Ting, Bi-Cun Xu, Takashi Washio, Zhi-Hua Zhou |
KDD | 3 |
| 2020 | A comparative study of data-dependent approaches without learning in measuring similarities of data objects
Sunil Aryal, Kai Ming Ting, Takashi Washio, Gholamreza Haffari |
Data Min. Knowl. Discov. | 3 |
| 2018 | Which Outlier Detector Should I use?abstractThis tutorial has four aims: (1) Providing the current comparative works on different outlier detectors, and analysing the strengths and weaknesses of these works and their recommendations. (2) Presenting non-obvious applications of outlier detectors. This provides examples of how outlier detectors are used in areas which are not normally considered to be the domains of outlier detection. (3) Inviting the research community to explore future research directions, in terms of both comparative study and outlier detection in general. (4) Giving an advice on the factors to consider when choosing an outlier detector, and strengths and weaknesses of some "top" recommended algorithms based on the current understanding in the literature. Kai Ming Ting, Sunil Aryal, Takashi Washio |
ICDM | 3 |
| 2018 | A Rare and Critical Condition Search Technique and its Application to Telescope Stray Light AnalysisabstractMany systems, including space satellites, cannot be upgraded or repaired easily during their missions. Simulation-based design techniques are often used to check conditions that can induce critical malfunctions in them, to ensure sufficient credibility and reliability during operation. However, critical conditions with a very low probability of occurring (e.g., 10−8 per trial) rarely appear within a tractable number of simulations. We propose herein a multicanonical Markov Chain Monte Carlo (MCMC) technique extended for the efficient search of rare but critical conditions, to significantly enhance simulation efficiency. Furthermore, we demonstrate an application of our proposed technique to an efficient search of “stray light” in a space telescope satellite. Keiichi Kisamori, Takashi Washio, Yoshio Kameda, Ryohei Fujimaki |
SDM | 2 |
| 2017 | Machine Learning Independent of Population Distributions for MeasurementabstractMany of recent advanced measurement techniques acquire highly complex information through elaborate measurement processes, and estimate the measurement objects from their corresponding patterns reflected in the outcome of these processes. The introduction of advanced statistical and machine learning methods to measurement techniques is now inevitable to solve these complicated inverse problems. However, the state of the art remains the straightforward application of problem settings and their solutions studied in statistics and machine learning which assume a steady population distribution providing the objective data, whereas every measurement is performed under a distinct distribution of disturbances and noises. As claimed and demonstrated in this paper, this fact largely degrades the accuracy and robustness of the measurements. To effectively overcome this issue, we present a framework of robust and accurate machine learning against deviations of population distributions between calibration data for the training and a new single observation in the measurement. This is achieved by properly reflecting generic measurement processes to their estimations. The significant advantages of the presented framework are demonstrated through a real-world application to olfactory sensing. Takashi Washio, Gaku Imamura, Genki Yoshikawa |
DSAA | 1 |
| 2017 | Data-dependent dissimilarity measure: an effective alternative to geometric distance measures
Sunil Aryal, Kai Ming Ting, Takashi Washio, Gholamreza Haffari |
Knowl. Inf. Syst. | 3 |
| 2016 | A Novel Continuous and Structural VAR Modeling Approach and Its Application to Reactor Noise AnalysisabstractA vector autoregressive model in discrete time domain (DVAR) is often used to analyze continuous time, multivariate, linear Markov systems through their observed time series data sampled at discrete timesteps. Based on previous studies, the DVAR model is supposed to be a noncanonical representation of the system, that is, it does not correspond to a unique system bijectively. However, in this article, we characterize the relations of the DVAR model with its corresponding Structural Vector AR (SVAR) and Continuous Time Vector AR (CTVAR) models through a finite difference method across continuous and discrete time domain. We further clarify that the DVAR model of a continuous time, multivariate, linear Markov system is canonical under a highly generic condition. Our analysis shows that we can uniquely reproduce its SVAR and CTVAR models from the DVAR model. Based on these results, we propose a novel Continuous and Structural Vector Autoregressive (CSVAR) modeling approach to derive the SVAR and the CTVAR models from their DVAR model empirically derived from the observed time series of continuous time linear Markov systems. We demonstrate its superior performance through some numerical experiments on both artificial and real-world data. Marina Demeshko, Takashi Washio, Yoshinobu Kawahara, Yuriy Pepyolyshev |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2014 | Mp-Dissimilarity: A Data Dependent Dissimilarity MeasureabstractNearest neighbour search is a core process in many data mining algorithms. Finding reliable closest matches of a query in a high dimensional space is still a challenging task. This is because the effectiveness of many dissimilarity measures, that are based on a geometric model, such as lp-norm, decreases as the number of dimensions increases. In this paper, we examine how the data distribution can be exploited to measure dissimilarity between two instances and propose a new data dependent dissimilarity measure called 'mp-dissimilarity'. Rather than relying on geometric distance, it measures the dissimilarity between two instances in each dimension as a probability mass in a region that encloses the two instances. It deems the two instances in a sparse region to be more similar than two instances in a dense region, though these two pairs of instances have the same geometric distance. Our empirical results show that the proposed dissimilarity measure indeed provides a reliable nearest neighbour search in high dimensional spaces, particularly in sparse data. Mp-dissimilarity produced better task specific performance than lp-norm and cosine distance in classification and information retrieval tasks. Sunil Aryal, Kai Ming Ting, Gholamreza Haffari, Takashi Washio |
ICDM | 4 |
| 2014 | Improving iForest with Relative Mass
Sunil Aryal, Kai Ming Ting, Jonathan R. Wells, Takashi Washio |
PAKDD (2) | 4 |
| 2013 | Efficiently rewriting large multimedia application execution traces with few event sequencesabstractThe analysis of multimedia application traces can reveal important information to enhance program execution comprehension. However typical size of traces can be in gigabytes, which hinders their effective exploitation by application developers. In this paper, we study the problem of finding a set of sequences of events that allows a reduced-size rewriting of the original trace. These sequences of events, that we call blocks, can simplify the exploration of large execution traces by allowing application developers to see an abstraction instead of low-level events. Christiane Kamdem Kengne, Léon Constantin Fopa, Alexandre Termier, Noha Ibrahim, Marie-Christine Rousset, Takashi Washio, Miguel Santana |
KDD | 6 |
| 2013 | DEMass: a new density estimator for big data
Kai Ming Ting, Takashi Washio, Jonathan R. Wells, Fei Tony Liu, Sunil Aryal |
Knowl. Inf. Syst. | 2 |
| 2011 | Density Estimation Based on MassabstractDensity estimation is the ubiquitous base modelling mechanism employed for many tasks such as clustering, classification, anomaly detection and information retrieval. Commonly used density estimation methods such as kernel density estimator and k-nearest neighbour density estimator have high time and space complexities which render them inapplicable in problems with large data size and even a moderate number of dimensions. This weakness sets the fundamental limit in existing algorithms for all these tasks. We propose the first density estimation method which stretches this fundamental limit to an extent that dealing with millions of data can now be done easily and quickly. We analyze the error of the new estimation (from the true density) using a bias-variance analysis. We then perform an empirical evaluation of the proposed method by replacing existing density estimators with the new one in two current density-based algorithms, namely, DBSCAN and LOF. The results show that the new density estimation method significantly improves the runtime of DBSCAN and LOF, while maintaining or improving their task-specific performances in clustering and anomaly detection, respectively. The new method empowers these algorithms, currently limited to small data size only, to process very large databases - setting a new benchmark for what density-based algorithms can achieve. Kai Ming Ting, Takashi Washio, Jonathan R. Wells, Fei Tony Liu |
ICDM | 2 |
| 2011 | Common Substructure Learning of Multiple Graphical Gaussian Models
Satoshi Hara 0001, Takashi Washio |
ECML/PKDD (2) | 2 |
| 2010 | GTRACE2: Improving Performance Using Labeled Union Graphs
Akihiro Inokuchi, Takashi Washio |
PAKDD (2) | 2 |
| 2010 | Mining Frequent Graph Sequence Patterns Induced by VerticesabstractThe mining of a complete set of frequent subgraphs from labeled graph data has been studied extensively. Furthermore, much attention has recently been paid to frequent pattern mining from graph sequences (dynamic graphs or evolving graphs). In this paper, we define a novel class of subgraph subsequence called an “induced subgraph subsequence” to enable efficient mining of a complete set of frequent patterns from graph sequences containing large graphs and long sequences. We also propose an efficient method to mine frequent patterns, called “FRISSs (Frequent Relevant, and Induced Subgraph Subsequences)”, from graph sequences. The fundamental performance of the method has been evaluated using artificial datasets, and its practicality has been confirmed through experiments using a real-world dataset. Akihiro Inokuchi, Takashi Washio |
SDM | 2 |
| 2010 | Best papers from the 12th Pacific-Asia conference on knowledge discovery and data mining (PAKDD2008)
Takashi Washio, Einoshin Suzuki, Kai Ming Ting |
Knowl. Inf. Syst. | 1 |
| 2008 | A Fast Method to Mine Frequent Subsequences from Graph Sequence DataabstractIn recent years, the mining of a complete set of frequent subgraphs from labeled graph data has been extensively studied.However, to our best knowledge, almost no methods have been proposed to find frequent subsequences of graphs from a set of graph sequences. In this paper, we define a novel class of graph subsequences by introducing axiomatic rules of graph transformation, their admissibility constraints and a union graph. Then we propose an efficient approach named "GTRACE'' to enumerate frequent transformation subsequences (FTSs) of graphs from a given set of graph sequences. Its fundamental performance has been evaluated by using artificial datasets, and its practicality has been confirmed through the experiments using real world datasets. Akihiro Inokuchi, Takashi Washio |
ICDM | 2 |
| 2008 | A Range Query Approach for High Dimensional Euclidean Space Based on EDM Estimation
Kentarou Kido, Hiroshi Kuwajima, Takashi Washio |
SDM | 3 |
| 2008 | DryadeParent, An Efficient and Robust Closed Attribute Tree Mining AlgorithmabstractIn this paper, we present a new tree mining algorithm, DryadeParent, based on the hooking principle first introduced in DRYADE. In the experiments, we demonstrate that the branching factor and depth of the frequent patterns to find are key factors of complexity for tree mining algorithms, even if often overlooked in previous work. We show that DryadeParent outperforms the current fastest algorithm, CMTreeMiner, by orders of magnitude on data sets where the frequent tree patterns have a high branching factor. Alexandre Termier, Marie-Christine Rousset, Michèle Sebag, Kouzou Ohara, Takashi Washio, Hiroshi Motoda |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2007 | Applications eligible for data mining
Takashi Washio |
Adv. Eng. Informatics | 1 |
| 2006 | Constructing Decision Trees for Graph-Structured Data by Chunkingless Graph-Based Induction
Phu Chien Nguyen, Kouzou Ohara, Akira Mogi, Hiroshi Motoda, Takashi Washio |
PAKDD | 5 |
| 2005 | Efficient Mining of High Branching Factor Attribute TreesabstractIn this paper, we present a new tree mining algorithm, DryadeParent, based on the hooking principle first introduced in Dryade (Termier et al, 2004). In the experiments, we demonstrate that the branching factor and depth of the frequent patterns to find are key factor of complexity for tree mining algorithms. We show that DryadeParent outperforms the current fastest algorithm, CMTreeMiner, by orders of magnitude on datasets where the frequent patterns have a high branching factor. Alexandre Termier, Marie-Christine Rousset, Michèle Sebag, Kouzou Ohara, Takashi Washio, Hiroshi Motoda |
ICDM | 5 |
| 2005 | Mining Quantitative Frequent Itemsets Using Adaptive Density-Based Subspace ClusteringabstractA novel approach to subspace clustering is proposed to exhaustively and efficiently mine quantitative frequent item-sets (QFIs) from massive transaction data. For the computational tractability, our approach introduces adaptive density-based and Apriori-like algorithm. Its outstanding performance is shown through numerical experiments. Takashi Washio, Yuki Mitsunaga, Hiroshi Motoda |
ICDM | 1 |
| 2005 | Cl-GBI: A Novel Approach for Extracting Typical Patterns from Graph-Structured Data
Phu Chien Nguyen, Kouzou Ohara, Hiroshi Motoda, Takashi Washio |
PAKDD | 4 |
| 2005 | Deriving Class Association Rules Based on Levelwise Subspace Clustering
Takashi Washio, Koutarou Nakanishi, Hiroshi Motoda |
PKDD | 1 |
| 2004 | Density-based spam detectorabstractThe volume of mass unsolicited electronic mail, often known as spam, has recently increased enormously and has become a serious threat to not only the Internet but also to society. This paper proposes a new spam detection method which uses document space density information. Although it requires extensive e-mail traffic to acquire the necessary information, an unsupervised learning engine with a short white list can achieve a 98% recall rate and 100% precision. A direct-mapped cache method contributes handling of over 13,000 e-mails per second. Experimental results, which were conducted using over 50 million actual e-mails of traffic, are also reported in this paper. Fuminori Adachi, Takashi Washio, Hiroshi Motoda, Teruaki Homma, Akihiro Nakashima, Hiromitsu Fujikawa, Katsuyuki Yamazaki |
KDD | 3 |
| 2004 | Using a Hash-Based Method for Apriori-Based Graph Mining
Phu Chien Nguyen, Takashi Washio, Kouzou Ohara, Hiroshi Motoda |
PKDD | 2 |
| 2003 | Classifier Construction by Graph-Based Induction for Graph-Structured Data
Warodom Geamsakul, Takashi Matsuda, Tetsuya Yoshida, Hiroshi Motoda, Takashi Washio |
PAKDD | 5 |
| 2002 | Adaptive Ripple Down Rules Method based on Minimum Description Length PrincipleabstractWhen class distribution changes, some pieces of knowledge previously acquired become worthless, and the existence of such knowledge may hinder acquisition of new knowledge. The paper proposes an adaptive ripple down rules (RDR) method based on the minimum description length principle aiming at knowledge acquisition in a dynamically changing environment. To cope with the change of class distribution, knowledge deletion is carried out as well as knowledge acquisition so that useless knowledge is properly discarded. To cope with the change of the source of knowledge, RDR knowledge based systems can be constructed adaptively by acquiring knowledge from both domain experts and data. By incorporating inductive learning methods, knowledge acquisition can be carried out even when only either data or experts are available by switching the source of knowledge from domain experts to data and vice versa at any time of knowledge acquisition. Since experts need not be available all the time, it contributes to reducing the cost of personnel expenses. Experiments were conducted by simulating the change of the source of knowledge and the change of class distribution using the datasets in UCI repository. The results are encouraging. Tetsuya Yoshida, Hiroshi Motoda, Takashi Washio |
ICDM | 3 |
| 2002 | Graph-based induction and its applications
Takashi Matsuda, Hiroshi Motoda, Takashi Washio |
Adv. Eng. Informatics | 3 |
| 2002 | Attribute Generation Based on Association Rules
Masahiro Terabe, Takashi Washio, Hiroshi Motoda, Osamu Katai, Tetsuo Sawaragi |
Knowl. Inf. Syst. | 2 |
| 2001 | Discovering Admissible Simultaneous Equation Models from Observed Data
Takashi Washio, Hiroshi Motoda, Yuji Niwa |
ECML | 1 |
| 2001 | S3Bagging: Fast Classifier Induction Method with Subsampling and Bagging
Masahiro Terabe, Takashi Washio, Hiroshi Motoda |
IDA | 2 |
| 2001 | Knowledge Acquisition from Both Human Expert and Data
Takuya Wada, Hiroshi Motoda, Takashi Washio |
PAKDD | 3 |
| 2001 | Automatic Web-Page Classification by Using Machine Learning Methods
Makoto Tsukada, Takashi Washio, Hiroshi Motoda |
Web Intelligence | 2 |
| 2001 | A Description Length-Based Decision Criterion for Default Knowledge in the Ripple Down Rules Method
Takuya Wada, Tadashi Horiuchi, Hiroshi Motoda, Takashi Washio |
Knowl. Inf. Syst. | 4 |
| 2000 | Extension of Graph-Based Induction for General Graph Structured Data
Takashi Matsuda, Tadashi Horiuchi, Hiroshi Motoda, Takashi Washio |
PAKDD | 4 |
| 2000 | An Apriori-Based Algorithm for Mining Frequent Substructures from Graph Data
Akihiro Inokuchi, Takashi Washio, Hiroshi Motoda |
PKDD | 2 |
| 1999 | Basket Analysis for Graph Structured Data
Akihiro Inokuchi, Takashi Washio, Hiroshi Motoda, Kouhei Kumasawa, Naohide Arai |
PAKDD | 2 |
| 1999 | A Data Pre-processing Method Using Association Rules of Attributes for Improving Decision Tree
Masahiro Terabe, Osamu Katai, Tetsuo Sawaragi, Takashi Washio, Hiroshi Motoda |
PAKDD | 4 |
| 1999 | Characterization of Default Knowledge in Ripple Down Rules Method
Takuya Wada, Tadashi Horiuchi, Hiroshi Motoda, Takashi Washio |
PAKDD | 4 |
| 1998 | Mining Association Rules for Estimation and Prediction
Takashi Washio, Hiroshi Motoda |
PAKDD | 1 |