Einoshin Suzuki

dblp:s/EinoshinSuzuki · DBLP profile ↗
← Back
54ranked-venue papers in the field
12as first author
5since 2021 · last 2024
0000-0001-7743-6177ORCID · verified

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 41 (11 first)Knowledge Engineering, Semantic Web & Information Systems · 7Information Retrieval & Web Search · 3Other / Interdisciplinary · 2 (1 first)Database Systems & Data Management · 1
YearPublicationVenuePosition
2024 SATJiP: Spatial and Augmented Temporal Jigsaw Puzzles for Video Anomaly Detection
Liheng Shen, Tetsu Matsukawa, Einoshin Suzuki
PAKDD (1)3
2023 Class-Specific Word Sense Aware Topic Modeling via Soft Orthogonalized Topics
abstract
We propose a word sense aware topic model for document classification based on soft orthogonalized topics. An essential problem for this task is to capture word senses related to classes, i.e., class-specific word senses. Traditional models mainly introduce semantic information of knowledge libraries for word sense discovery. However, this information may not align with the classification targets, because these targets are often subjective and task-related. We aim to model the class-specific word senses in topic space. The challenge is to optimize the class separability of the senses, i.e., obtaining sense vectors with (a) high intra-class and (b) low inter-class similarities. Most existing models predefine specific topics for each class to specify the class-specific sense vectors. We call them hard orthogonalization based methods. These methods can hardly achieve both (a) and (b) since they assume the conditional independence of topics to classes and inevitably lose topic information. To this problem, we propose soft orthogonalization for topics. Specifically, we reserve all the topics and introduce a group of class-specific weights for each word to handle the importance of topic dimensions to class separability. Besides, we detect and use highly class-specific words in each document to guide sense estimation. Our experiments on two standard datasets show that our proposal outperforms other state-of-the-art models in terms of accuracy of sense estimation, document classification, and topic modeling. In addition, our joint learning experiments with the pre-trained language model BERT showcased the best complementarity of our model in most cases compared to other topic models.
Wenbo Li 0011, Einoshin Suzuki
CIKM3
2022 GIAD-ST: Detecting anomalies in human monitoring based on generative inpainting via self-supervised multi-task learning
Ning Dong 0001, Einoshin Suzuki
J. Intell. Inf. Syst.2
2021 Adaptive and hybrid context-aware fine-grained word sense disambiguation in topic modeling based document representation
Wenbo Li 0011, Einoshin Suzuki
Inf. Process. Manag.2
2021 Topic modeling for sequential documents based on hybrid inter-document topic dependency
Wenbo Li 0011, Hiroto Saigo, Bin Tong, Einoshin Suzuki
J. Intell. Inf. Syst.4
2020 Hybrid Context-Aware Word Sense Disambiguation in Topic Modeling based Document Representation
abstract
We propose a hybrid context based topic model for word sense disambiguation in document representation. Document representation is an essential part of various document based tasks, and word sense disambiguation is to capture the distinctions of word senses in the representation. Traditional methods mainly rely on knowledge libraries for data enrichment; however, semantics division for a word may vary from different domain-specific datasets. We aim to discover more particular word semantic differences for each input dataset and handle the disambiguation problem without data enrichment. The challenge for this disambiguation is to (1) divide various senses for each polysemous word while (2) preserve the differences between synonyms. Most of the existing models are either based on separate context clusters or integrating an auxiliary module to specify word senses. They can hardly achieve both (1) and (2) since different senses of a word are assumed to be independent and their intrinsic relationships are ignored. To solve this problem, we estimate a word sense by both the context in which it occurs and the contexts of its other occurrences. Besides, we introduce the “Bag-of-Senses” (BoS) assumption: a document is a multiset of word senses, and the senses are generated instead of the words. Our experiments on three standard datasets show that our proposal outperforms other state-of-the-art methods in terms of accuracy of word sense estimation, topic modeling, and document classification.
Wenbo Li 0011, Einoshin Suzuki
ICDM2
2020 Context-Aware Latent Dirichlet Allocation for Topic Segmentation
Wenbo Li 0011, Tetsu Matsukawa, Hiroto Saigo, Einoshin Suzuki
PAKDD (1)4
2020 Detecting outliers with one-class selective transfer machine
Hirofumi Fujita, Tetsu Matsukawa, Einoshin Suzuki
Knowl. Inf. Syst.3
2019 Experimental validation for N-ary error correcting output codes for ensemble learning of deep neural networks
Kaikai Zhao, Tetsu Matsukawa, Einoshin Suzuki
J. Intell. Inf. Syst.3
2017 Skeleton clustering by multi-robot monitoring for fall risk discovery
Yutaka Deguchi, Daisuke Takayama, Shigeru Takano, Vasile-Marian Scuturici, Jean-Marc Petit, Einoshin Suzuki
J. Intell. Inf. Syst.6
2016 Minimizing response time in time series classification
Shin Ando, Einoshin Suzuki
Knowl. Inf. Syst.2
2015 Ensemble anomaly detection from multi-resolution trajectory features
Shin Ando, Theerasak Thanomphongphan, Yoichi Seki, Einoshin Suzuki
Data Min. Knowl. Discov.4
2015 Classifying actions based on histogram of oriented velocity vectors
Somar Boubou, Einoshin Suzuki
J. Intell. Inf. Syst.2
2014 Discriminative Learning on Exemplary Patterns of Sequential Numerical Data
abstract
One of the effective methodologies for time series classification is to identify informative subsequence patterns in time series and exploit them as discriminative features. Previous studies on this methodology have achieved promising results using a small number of individually selected patterns. However, there remain difficulties in finding a set of related patterns or patterns of a minor class, which can be critical in real-world applications. In this paper, we exploit the sparse learning technique for the support vector machine (SVM) to identify informative and exemplary patterns. We first present a representation of time series as a vector of distances to exemplary patterns. It allows a structural SVM to handle distance space data and function as the nearest neighbor classifier, the combination of which is known to be highly competitive in time series classification. We then extend the zero-norm approximation method for the structural SVM, which can eliminate non-essential patterns from the classification model. The resulting model makes predictions by a simple modified nearest neighbor rule, yet has a strong mathematical support for empirical risk minimization and feature selection. We conduct an empirical study on real-world behavior and sequential data to evaluate the effectiveness of the proposed method and graphically examine the exemplary patterns.
Shin Ando, Einoshin Suzuki
ICDM2
2014 Finding peculiar compositions of two frequent strings with background texts
Daisuke Ikeda, Einoshin Suzuki
Knowl. Inf. Syst.2
2014 Transfer dimensionality reduction by Gaussian process in parallel
Bin Tong, Junbin Gao, Thach Huy Nguyen, Hao Shao, Einoshin Suzuki
Knowl. Inf. Syst.5
2013 Time-sensitive Classification of Behavioral Data
abstract
In this paper, we address a classification task under a time-sensitive setting, in which the amount of observation required to make a prediction is viewed as a practical cost. Such a setting is intrinsic in many systems where the potential reward of the action against the predicted event depends on the response time, e.g., surveillance/warning and diagnostic applications. Meanwhile, predictions are usually less reliable when based on fewer observations, i.e., there exists a trade-off between such temporal cost and the accuracy. We address the task as a classification of subsequences in a time series. The goal is to predict the occurrences of events from subsequent observations and to learn when to commit to the prediction considering the trade-off. We propose an ensemble of classifiers which respectively makes predictions based on subsequences of different lengths. The prediction of the ensemble is given by the earliest confident prediction among the individual classifiers. We propose a cutting-plane algorithm for jointly training an ensemble of linear classifiers considering their temporal dependence. We compare the proposed algorithm against conventional approaches over a collection of behavioral trajectory data.
Shin Ando, Einoshin Suzuki
SDM2
2013 Transfer learning by centroid pivoted mapping in noisy environment
Thach Huy Nguyen, Bin Tong, Hao Shao, Einoshin Suzuki
J. Intell. Inf. Syst.4
2013 A feature-free and parameter-light multi-task clustering framework
Thach Huy Nguyen, Hao Shao, Bin Tong, Einoshin Suzuki
Knowl. Inf. Syst.4
2013 Extended MDL principle for feature-based inductive transfer learning
Hao Shao, Bin Tong, Einoshin Suzuki
Knowl. Inf. Syst.3
2012 Query by Committee in a Heterogeneous Environment
Hao Shao, Bin Tong, Einoshin Suzuki
ADMA3
2012 Intelligent Data Analysis by a Home-Use Human Monitoring Robot
Shinsuke Sugaya, Daisuke Takayama, Asuki Kouno, Einoshin Suzuki
IDA4
2012 Linear semi-supervised projection clustering by transferred centroid regularization
Bin Tong, Hao Shao, Bin-Hui Chou, Einoshin Suzuki
J. Intell. Inf. Syst.4
2011 Role Discovery for Graph Clustering
Bin-Hui Chou, Einoshin Suzuki
APWeb2
2011 Role-Behavior Analysis from Trajectory Data by Cross-Domain Learning
abstract
Behavior analysis using trajectory data presents a practical and interesting challenge for KDD. Conventional analyses address discriminative tasks of behaviors, e.g., classification and clustering typically using the subsequences extracted from the trajectory of an object as a numerical feature representation. In this paper, we explore further to identify the difference in the high-level semantics of behaviors such as roles and address the task in a cross-domain learning approach. The trajectory, from which the features are sampled, is intuitively viewed as a domain, and we assume that its intrinsic structure is characterized by the underlying role associated with the tracked object. We propose a novel hybrid method of spectral clustering and density approximation for comparing clustering structures of two independently sampled trajectory data and identifying patterns of behaviors unique to a role. We present empirical evaluations of the proposed method in two practical settings using real-world robotic trajectories.
Shin Ando, Einoshin Suzuki
ICDM2
2011 Compact Coding for Hyperplane Classifiers in Heterogeneous Environment
Hao Shao, Bin Tong, Einoshin Suzuki
ECML/PKDD (3)3
2011 ACE: Anomaly Clustering Ensemble for Multi-perspective Anomaly Detection in Robot Behaviors
abstract
This paper addresses an application of anomaly detection from subsequences of time series (STS) to autonomous robots' behaviors. An important aspect of mining sequential data is selecting the temporal parameters, such as the subsequence length and the degree of smoothing. For example in the task at hand, the patterns of the robot's velocity, which is one of its fundamental features, vary significantly subject to the interval for measuring the displacement. Selecting the time scale and resolution is difficult in unsupervised settings, and is often more critical than the choice of the method. In this paper, we propose an ensemble framework for aggregating anomaly detection from different perspectives, i.e., settings of user-defined, temporal parameters. In the proposed framework, each behavior is labeled whether it is an anomaly in multiple settings. The set of labels are used as meta-features of the respective behaviors. Cluster analysis in a meta-feature space partitions anomalous behaviors pertained to a specific range of parameters. The framework also includes a scalable implementation of the instance-based anomaly detection. We evaluate the proposed framework by ROC analysis, in comparison to conventional ensemble methods for anomaly detection.
Shin Ando, Einoshin Suzuki, Yoichi Seki, Theerasak Thanongphongphan, Daisuke Hoshino
SDM2
2011 Feature-based Inductive Transfer Learning through Minimum Encoding
abstract
This paper proposes an Extended Minimum Description Length Principle (EMDLP) for feature-based inductive transfer learning, in which both the source and the target data sets contain class labels and relevant features are transferred from the source domain to the target one. Despite numerous works on this topic, few of them have a solid theoretical framework and are parameter-free. Our EMDLP overcomes these flaws and allows us to evaluate the inferiority of the results of transfer learning with the add-sum of the code lengths of five components: the corresponding two hypotheses, the two data sets with the help of the hypotheses, and the set of the transferred features. We design a code book to build the connections between the source and the target tasks. Extensive experiments using both real and artificial data sets show that EMDLP is robust against noise and performs better on the classification accuracy than the state-of-the-art methods.
Hao Shao, Einoshin Suzuki
SDM2
2011 Gaussian Process for Dimensionality Reduction in Transfer Learning
abstract
Dimensionality reduction has been considered as one of the most significant tools for data analysis. In general, supervised information is helpful for dimensionality reduction. However, in typical real applications, supervised information in multiple source tasks may be available, while the data of the target task are unlabeled. An interesting problem of how to guide the dimensionality reduction for the unlabeled target data by exploiting useful knowledge, such as label information, from multiple source tasks arises in such a scenario. In this paper, we propose a new method for dimensionality reduction in the transfer learning setting. Unlike traditional paradigms where the useful knowledge from multiple source tasks is transferred through distance metric, our proposal firstly converts the dimensionality reduction problem into integral regression problems in parallel. Gaussian process is then employed to learn the underlying relationship between the original data and the reduced data. Such a relationship can be appropriately transferred to the target task by exploiting the prediction ability of the Gaussian process model and inventing different kinds of regularizers. Extensive experiments on both synthetic and real data sets show the effectiveness of our method.
Bin Tong, Junbin Gao, Thach Huy Nguyen, Einoshin Suzuki
SDM4
2010 Discovering Community-Oriented Roles of Nodes in a Social Network
Bin-Hui Chou, Einoshin Suzuki
DaWak2
2010 Subclass-Oriented Dimension Reduction with Constraint Transformation and Manifold Regularization
Bin Tong, Einoshin Suzuki
PAKDD (2)2
2010 Semi-supervised Projection Clustering with Transferred Centroid Regularization
Bin Tong, Hao Shao, Bin-Hui Chou, Einoshin Suzuki
ECML/PKDD (3)4
2010 Best papers from the 12th Pacific-Asia conference on knowledge discovery and data mining (PAKDD2008)
Takashi Washio, Einoshin Suzuki, Kai Ming Ting
Knowl. Inf. Syst.2
2009 Detection of unique temporal segments by information theoretic meta-clustering
abstract
The central challenge in temporal data analysis is to obtain knowledge about its underlying dynamics. In this paper, we address the observation of noisy, stochastic processes and attempt to detect temporal segments that are related to inconsistencies and irregularities in its dynamics. Many conventional anomaly detection approaches detect anomalies based on the distance between patterns, and often provide only limited intuition about the generative process of the anomalies. Meanwhile, model-based approaches have difficulty in identifying a small, clustered set of anomalies.
Shin Ando, Einoshin Suzuki
KDD2
2009 Negative Encoding Length as a Subjective Interestingness Measure for Groups of Rules
Einoshin Suzuki
PAKDD1
2009 Discovering Action Rules That Are Highly Achievable from Massive Data
Einoshin Suzuki
PAKDD1
2009 Mining Peculiar Compositions of Frequent Substrings from Sparse Text Data Using Background Texts
Daisuke Ikeda, Einoshin Suzuki
ECML/PKDD (1)2
2008 Unsupervised Cross-Domain Learning by Interaction Information Co-clustering
abstract
In real-world data mining applications, one often has access to multiple datasets that are relevant to the task at hand. However, learning from such datasets can be difficult as they are often drawn from different domains, i.e., not identically distributed or differ in class or feature sets. In this paper, we consider the problem of learning the class structures %, unique and shared, of related domains in an unsupervised manner. Its setting generalizes that of information filtering and novelty detection applications which addresses both known and unknown classes. We propose a co-clustering framework for estimating and adapting the class structures of two related domains, {enabling the analyses of shared and unique classes.} We define an objective function using interaction information to take account of the divergence between the corresponding clusters of respective domains. We present an iterative algorithm which alternates object and feature clustering and converges to a local minimum of the objective function. We present empirical results using text benchmarks, comparing the proposed algorithm and combinations of conventional approaches in problems of partitioning documents and detecting unknown topics.
Shin Ando, Einoshin Suzuki
ICDM2
2006 An Information Theoretic Approach to Detection of Minority Subsets in Database
abstract
Detection of rare and exceptional occurrences in large- scale databases have become an important practice in the field of knowledge discovery and information retrieval. Many databases include large amount of noise or irrelevant data, whose distribution often overlaps with the subsets of exceptional data containing useful knowledge. This paper addresses the problem of finding a small subset of "minority" data whose distribution overlaps with, but are exceptional to or inconsistent with that of the majority of the database. In such a case, conventional distance-based or density-based approaches in Outlier Detection are ineffective due to their dependence on the structure of the majority or the prerequisite of critical parameters. We formalize the task as an estimation of a model of the minority subset which provides a simple description of the subset and yet maintains divergence from that of the majority. This estimation is formalized as a minimization problem using an information theoretic framework of Rate Distortion theory. We further introduce conditions of the majority to derive an objective function which factorizes the property of the minority and dependence to the structure of the majority. The proposed method shows improvements from conventional approaches in artificial data and a promising result in document retrieval problem.
Shin Ando, Einoshin Suzuki
ICDM2
2005 Unified algorithm for undirected discovery of exception rules
abstract
This article presents an algorithm that seeks every possible exception rule that violates a commonsense rule and satisfies several assumptions of simplicity. Exception rules, which represent systematic deviation from commonsense rules, are often found interesting. Discovery of pairs that consist of a commonsense rule and an exception rule, resulting from undirected search for unexpected exception rules, was successful in various domains. In the past, however, an exception rule represented a change of conclusion caused by adding an extra condition to the premise of a commonsense rule. That approach formalized only one type of exception and failed to represent other types. To provide a systematic treatment of exceptions, we categorize exception rules into 11 categories, and we propose a unified algorithm for discovering all of them. Preliminary results on 15 real-world datasets provide an empirical proof of effectiveness of our algorithm in discovering interesting knowledge. The empirical results also match our theoretical analysis of exceptions, showing that the 11 types can be partitioned in three classes according to the frequency with which they occur in data. © 2005 Wiley Periodicals, Inc. Int J Int Syst 20: 673–691, 2005.
Einoshin Suzuki, Jan M. Zytkow
Int. J. Intell. Syst.1
2003 Detecting Interesting Exceptions from Medical Test Data with Visual Summarization
abstract
We propose a method which visualizes irregular multidimensional time-series data as a sequence of probabilistic prototypes for detecting exceptions from medical test data. Conventional visualization methods often require iterative analysis and considerable skill thus are not totally supported by a wide range of medical experts. Our PrototypeLines displays summarized information based on a probabilistic mixture model by using hue only thus is considered to exhibit novelty. The effectiveness of the summarization is pursued mainly through use of a novel information criterion. We report our endeavor with chronic hepatitis data, especially discoveries of interesting exceptions by a nonexpert and an untrained expert.
Einoshin Suzuki, Takeshi Watanabe, Hideto Yokoi, Katsuhiko Takabayashi
ICDM1
2003 Detecting Hostile Accesses through Incremental Subspace Clustering
abstract
We propose an incremental subspace clustering method for flexibly detecting hostile accesses to a Web site. Typical log data for Web accesses are huge, contain irrelevant information, and exhibit dynamic characteristics. We overcome these difficulties through data squashing, subspace clustering, and an incremental algorithm. We have improved, by modifying its data squashing functionality, our subspace clustering method SUBCCOM so that it can exploit previous results. Experimental evaluation confirms superiority of our I-SUBCCOM in terms of precision, recall, and computation time.
Masaki Narahashi, Einoshin Suzuki
Web Intelligence2
2002 Iterative Data Squashing for Boosting Based on a Distribution-Sensitive Distance
Yuta Choki, Einoshin Suzuki
PKDD2
2001 Bloomy Decision Tree for Multi-objective Classification
Einoshin Suzuki, Masafumi Gotoh, Yuta Choki
PKDD1
2000 Exception Rule Mining with a Relative Interestingness Measure
Farhad Hussain, Huan Liu 0001, Einoshin Suzuki, Hongjun Lu
PAKDD3
2000 Evaluating Hypothesis-Driven Exception-Rule Discovery with Medical Data Sets
Einoshin Suzuki, Shusaku Tsumoto
PAKDD1
2000 Unified Algorithm for Undirected Discovery of Execption Rules
Einoshin Suzuki, Jan M. Zytkow
PKDD1
1999 Prediction Rule Discovery Based on Dynamic Bias Selection
Einoshin Suzuki, Toru Ohno
PAKDD1
1999 Support Vector Machines for Knowledge Discovery
Shinsuke Sugaya, Einoshin Suzuki, Shusaku Tsumoto
PKDD2
1998 Simultaneous Reliability Evaluation of Generality and Accuracy for Rule Discovery in Databases
Einoshin Suzuki
KDD1
1998 Discovery of Surprising Exception Rules Based on Intensity of Implication
Einoshin Suzuki, Yves Kodratoff
PKDD1
1997 Autonomous Discovery of Reliable Exception Rules
Einoshin Suzuki
KDD1
1996 Exceptional Knowledge Discovery in Databases Based on Information Theory
Einoshin Suzuki, Masamichi Shimura
KDD1
1994 Knowledge-Based Handling of Design Expertise
abstract
Research issues in the domain of AI for design can be organized in three categories: decision making, representation and knowledge handling. In the area of knowledge handling, this paper addresses issues concerning the management of design experience to guide a priori the generation of candidate solutions. The approach is based on keeping the trace of a previous design experience as a hierarchical knowledge base. A level in the hierarchy can be viewed as a level of granularity of the description of the design process. A general framework for defining a partial order function between the granularity levels in the knowledge bases of design expertise is proposed. It is then possible to compute the sets of the elements belonging to smaller granularity levels, which are linked to any component of the hierarchy. Thus, it makes it possible to compute the level in the hierarchy that can be reused without modification for the design of a new product. Computation of the appropriate level is mainly based on matching the data corresponding to the new requirements with these sets. The approach has been tested by using a multiple expert systems structure based on using interactively two systems, an expert system development tool for design, KAUS, and an expert system development tool for diagnosing engineering processes, SUPER. The intrinsic properties of SUPER have also been used for improving the design procedure when qualitative and quantitative knowledge is involved.>
Pierre Morizet-Mahoudeaux, Einoshin Suzuki, Setsuo Ohsuga
ICDE2