Shusaku Tsumoto

dblp:66/3259 · DBLP profile ↗
← Back
68ranked-venue papers in the field
48as first author
10since 2021 · last 2025
0000-0001-6651-976XORCID · verified

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 30 (17 first)Big Data, Cloud & Distributed Data Systems · 22 (17 first)Knowledge Engineering, Semantic Web & Information Systems · 11 (9 first)Other / Interdisciplinary · 5 (5 first)
YearPublicationVenuePosition
2025 Medical Incident Reports Analysis
Tomohiro Kimura, Shusaku Tsumoto
IEEE Big Data2
2025 Topological Data Analysis in Rule Mining Space
Shusaku Tsumoto, Tomohiro Kimura, Shoji Hirano
IEEE Big Data1
2024 Toward the Implementation of the DPC Code Classification System in Hospital Information Systems
abstract
We have been conducting research to develop a system for selecting DPC codes from patient discharge summaries using text mining techniques. The performance demonstrated thus far indicates that the classification system is sufficiently practical for real-world use. However, with the limited dataset used for training, the possibility of overfitting cannot be ruled out, necessitating performance evaluation with new data.In this study, we propose a scheme in which the classification system is constructed using discharge summaries from one fiscal year and evaluated using data from a different fiscal year. Following this scheme, we conducted experiments to test its validity.The results showed that the evaluation metrics obtained through cross-validation were less than those obtained using new data. This finding suggests that constructing the classification system while bias estimation is crucial for its successful implementation in hospital information systems.
Shusaku Tsumoto, Tomohiro Kimura, Shoji Hirano
IEEE Big Data1
2024 Toward the Implementation of the DPC Code Classification System in Hospital Information Systems
abstract
We have been conducting research to develop a system for selecting DPC codes from patient discharge summaries using text mining techniques. The performance demonstrated thus far indicates that the classification system is sufficiently practical for real-world use. However, with the limited dataset used for training, the possibility of overfitting cannot be ruled out, necessitating performance evaluation with new data.In this study, we propose a scheme in which the classification system is constructed using discharge summaries from one fiscal year and evaluated using data from a different fiscal year. Following this scheme, we conducted experiments to test its validity.The results showed that the evaluation metrics obtained through cross-validation were less than those obtained using new data. This finding suggests that constructing the classification system while bias estimation is crucial for its successful implementation in hospital information systems.
Shusaku Tsumoto, Tomohiro Kimura, Shoji Hirano
IEEE Big Data1
2023 Using Hospital Information System Data to Estimate Nursing Care Needs
abstract
This paper proposes visualization and simulation methods which estimates clinical indices from the data stored in the hospital information system (HIS). The method is executed as follows. First, we construct DWH where needed variables are extracted from HIS. Second, the indices (nursing needs: Score A, which consists of the histories of monitoring and procedures and Score B, which evaluates of patient’s physical condition) are calculated from extracted data. Third, chronological changes of the indices are visualized. Finally, by using the model of stochastic differential equations, the nursing care needs and its chronological changes during the admission period are estimated. We evaluated the method with the data in HIS (from 2017 to 2022). The results shows that the method correctly simulated the chronological change of nursing care needs.
Hiroyuki Ikedo, Tomohiro Kimura, Shusaku Tsumoto, Shoji Hirano
IEEE Big Data3
2023 Analysis of Medical Incident Reports using Text Mining*
abstract
We conducted an analysis of incident reports from Shimane University School of Medicine Hospital using text mining techniques to efficiently analyze the cases based on the keywords they contained. The analysis targeted the text data of incident reports filed between fiscal years 2017 and 2022, with a particular focus on reports related to personal information. We employed morphological analysis on the collected text data, organizing it word by word to facilitate easier analysis. Subsequently, techniques such as word clouds and cluster analysis were utilized. Based on the results obtained, we conducted a thorough analysis and evaluation of the issues highlighted in the incidents. It was found that incidents related to USB memory often involved them being inadvertently placed in pockets and sent for laundering. Additionally, document-related incidents could be categorized into six distinct clusters.
Tomohiro Kimura, Shusaku Tsumoto, Shoji Hirano
IEEE Big Data2
2022 Similarity Analysis of Order Trajectory for Hospital Management
abstract
Last year we proposed the application of order trajectory curve analysis to hospital mangement, which will be useful to analyze the global behavior of hospital activities. This paper proposes multi-dimensional trajectories mining to analyze the temporal characteristics of hospital services. Order Trajectories mining method consists of the following two process. First the similarities between temporal trajectories of selected variables are calculated. Second, similarity-based analysis technique such as clustering and multidimensional scaling are applied. The method was evaluated on data on the number of orders extracted from hospital information system. The results showed that the method discovered several important characteristics of the divisions in the hospital..
Tomohiro Kimura, Shusaku Tsumoto, Shoji Hirano
IEEE Big Data2
2022 Temporal Data Mining in AI-based Patient Navigation Service
abstract
The pandemic of COVID-19 reminds us of the basic important principles for prevention of infection: avoid the "Three Cs": closed spaces, crowded places and close contact settings. Outpatient clinics in Japan are typical examples of three Cs, where some kinds of decision support system are required to solve the above situation. This paper proposes data mining based patient navigation support system to prevent the Three Cs. Behind the systems, temporal data mining units plays an important role in providing temporal information to the patients, such as waiting time and human densities in the waiting rooms. It analyzes the data stored in hospital information systems, including patient information, logs of clinical orders. The analysis results show that several aspects of patients’ waiting are visualized by temporal data mining.
Shusaku Tsumoto, Tomohiro Kimura, Shoji Hirano, Katsutoshi Yada
IEEE Big Data1
2021 Granular Computing based Comparison of Agglomerative Clustering
abstract
Empirical comparison of clustering methods is challenging because it is difficult to set up evaluation indices like supervised learning. In this paper, we propose weaker evaluation indices when at least one target group is given by domain experts. If so, we can introduce rough set based approximation for empirical comparison of dissimilarities and clustering strategies as follows. When a set of target (target set) is given, a level of clustering tree where one branch includes all the targets can be traced with the number of elements included. The pair (#clusters_of_a_level, #elements_of_a_cluster) can be viewed as indices-pair for a given clustering tree. Then, the algorithm for comparison of metric uses twofold comparison: first, it compares the number of partition, and then compares the ratio. Six distances and seven clustering methods were compared by using the above pair. The target items are based on nursing cares necessary for surgical operation of cataracts. Empirical results show that Euclidean distance and the Ward method is the best for obtaining a suitable clinical pathway for all clustering methods.
Shusaku Tsumoto
IEEE BigData1
2021 Empirical Rule Induction Methods Selection
abstract
For data mining in big data, selection of subsets from data and selection of rule induction are very critical issues. This paper proposes a method for selection of rule induction by using hierarchical sampling. First, the method generates training samples and test samples in a two-level hierarchical way.Second, it generates new training sub-samples and test sub-samples from training samples. Third, rule induction methods construct the models from training sub-samples and evaluate the model by new test samples. The second and third process are repeated for a given number of trials and evaluation indices are calculated as mean values of all the trials. Then, fifth, the method compared the results and select the best method. And finally, the selected method is used for construction of the model by using first training samples and the model is evaluated by first test samples. We evaluated this method on seven medical datasets. The results show that this method gives better performance than conventional methods.
Shusaku Tsumoto
IEEE BigData1
2020 Automated Dual Clustering for Clinical Pathway Mining*
abstract
One of the most important task of data mining in hospital is to discover structured knowledge about decision making, which is useful for management of clinical process. However, most of the data in hospital information are stored without classification labels or meaning of clinical actions. Thus, unsupervised learning techniques are required for analysis. This paper proposes a method which induces a clinical pathway by using sample and attribute clustering of the histories of nursing orders stored in hospital information system. The method consists of the following five steps: first, frequencies of nursing orders are extracted from hospital information system as a dataset in which row and column represents nursing orders and days. Second, orders are classified into several groups by using sample clustering. Then, attributes clustering is applied to the data for feature selection. Fourth, for each sample and attribute clustering, the number of clusters are obtained from the sequence of the height values and according to the results of attribute clustering, the original dataset is decomposed into subtables. Then, the second to fourth steps will be repeated in a recursive way until the grouping of attributes (days) are stable. Finally, a new pathway will be constructed from all the induced results. The proposed method was evaluated on datasets extracted from a hospital information system. The experiment results show that the method is useful for construction of a clinical pathway when the distribution of length of stay is uni-modular.
Shusaku Tsumoto, Tomohiro Kimura, Shoji Hirano
IEEE BigData1
2020 Order Trajectory Analysis in Hospital Information System*
abstract
Two of the most important roles of hospital information system (HIS) is to transfer clinical orders issued by doctors and nurses to other division and to store results of executed orders. Thus, the numbers of issued and executed orders will reflect the clinical activities in large hospitals. This paper proposes a visualization technique, called order trajectory analysis which visualize the temporal sequences of the number of orders. Then, clustering is applied to the order trajectory in order to show the similarities between clinical divisions.
Shusaku Tsumoto, Tomohiro Kimura, Shoji Hirano
IEEE BigData1
2019 Mining frequent temporal patterns from medical data based on fuzzy ranged relations
abstract
In this paper, we propose a method to mine frequent temporal patterns from medical time series based on fuzzy ranged interval relations. We firstly introduce the concept of ranged interval relations and then extend it to fuzzy relations in order to make it possible to work with fuzziness of duration like days or weeks and to generate pattens associated with abstracted durations. Through the experiments on a synthetic dataset we demonstrate that our approach enables a sequence to simultaneously belong to multiple relations and that it is possible to control the level of concordance for a case to support a pattern by changing the threshold of membership grade.
Shoji Hirano, Shusaku Tsumoto
IEEE BigData2
2019 Estimation of Disease Code from Electronic Patient Records
abstract
This paper proposes a method which classifies discharge summaries stored in hospital information system, which consists of the following four steps. First, a term matrix of the set of summaries is induced by morphological analysis (RMecab). Next, correspondence analysis is applied to the term matrix and numerical values of two dimensional coordinates are assigned to each keyword and each concept. By measuring the euclidean distance between categories and keywords, keywords are ordered. Then, keywords are selected as attributes according to the rank, and training examples for classifiers will be generated. Finally, learning methods are applied to the training examples. Experimental validation shows that random forest achieved the best performance and deep learning (multiple layer perceptron) is the second best.
Shusaku Tsumoto, Tomohiro Kimura, Haruko Iwata, Shoji Hirano
IEEE BigData1
2018 From Hospital Big Data to Clinical Process: A Granular Computing Approach
abstract
This paper proposes construction of clinical process plan from nursing order histories and discharge summaries stored in hospital information system. First, the system extracts subgrouping from clinical cases with the same Diagnostic Procedure Combination code (DPC) by mixture model clustering. Subgroups give different types of diseases with different temporal evolution. Then, classification models of each subgroup are constructed by the analysis of discharge summaries to capture the meaning of each subgroup. Finally, cases are classified by using the classification model and a clinical pathway is generated for each new subgroup. The proposed method was evaluated on the datasets extracted hospital information system, whose results show that plausible clinical pathways were obtained, compared with previously introduced methods.
Shusaku Tsumoto, Shoji Hirano, Tomohiro Kimura, Haruko Iwata
IEEE BigData1
2018 Empirical Comparison of Distances for Agglomerative Hierarchical Clustering
Shusaku Tsumoto, Tomohiro Kimura, Haruko Iwata, Shoji Hirano
IPMU (2)1
2017 Mining text for disease diagnosis in hospital information system
abstract
Electronic patient records (EPR) are rich in texts, where almost all the decision making processes of medical staff are written. Thus, mining in EPR is important for acquision of decision making process and diagnosis. In this paper, as a first step, we focus on text mining for discharge summaries, which include the compact explanation for the patient's admission. a record of her complaints, physical findings, laboratory results and radiographic studies while hospitalized; a list of changes in her medications at discharge; and recommendations for follow up care. Text mining process consists of the following four processes: first, morphological analysis is applied to a set of summaries and a term matrix is generated. Second, correspond analysis is applied to the classification labels and the term matrix and generates two dimensional coordinates. By measuring the distances between categories and the assigned points, ranking of key words will be generated. Then, keywords are selected as attributes according to the rank, and training examples for classifiers will be generated. Finally, learning methods are applied to the training examples. Experimental validation shows that random forest achieved the best performance and the second best was the deep learner with a small difference, but decision tree methods with many keywords performed only a little worse than neural network or deep learning methods.
Shusaku Tsumoto, Tomohiro Kimura, Haruko Iwata, Shoji Hirano
IEEE BigData1
2016 Construction of clinical pathway from histories of clinical actions in hospital information system
abstract
This paper proposes a method which induces a clinical pathway by using sample and attribute clustering of the histories of nursing orders stored in hospital information system. The method consists of the following five steps: first, frequencies of nursing orders are extracted from hospital information system. Second, orders are classified into several groups by using sample clustering. Then, attributes clustering is applied to the data for feature selection. Fourth, the method compares between generated functions for sample and attribute clustering which relate the number of clusters and calculated similarities. Fifth, if attribute clustering gives better performance with respect to the function, the dataset is decomposed into subtables by using the grouping of attribute clustering. Then, the first step will be repeated in a recursive way. After the grouping results are stable, a new pathway will be constructed from all the induced results. The method was applied to datasets of a disease extracted from a hospital information system. The results show that the proposed method is useful for construction of a clinical pathway.
Shusaku Tsumoto, Shoji Hirano, Haruko Iwata
IEEE BigData1
2016 Mining process for improvement of clinical process quality
abstract
This paper proposes an active mining process for improvement of quality of clinical process by using service logs in a hospital information system. First, datasets of temporal change of the number of orders are extracted from service logs stored in hospital information system. Then, since datasets of temporal change can be viewed as time-series of a statistic, clustering can be applied to the data. By using the groups obtained, datasets of command sequences are extracted from the logs and sequence mining process is applied. The results of sequence mining are interpreted with the results of clustering and hypothesis will be generated. The results show that the method improved the clinical process and waiting time in outpatient clinic.
Shusaku Tsumoto, Shoji Hirano, Haruko Iwata, Norio Yoshimoto, Tomohiro Kimura
IEEE BigData1
2016 A proposal of a privacy-preserving questionnaire by non-deterministic information and its analysis
abstract
We focus on a questionnaire consisting of three-choice question or multiple-choice question, and propose a privacy-preserving questionnaire by non-deterministic information. Each respondent usually answers one choice from the multiple choices, and each choice is stored as a tuple in a table data. The organizer of this questionnaire analyzes the table data set, and obtains rules and the tendency. If this table data set contains personal information, the organizer needs to employ the analytical procedures with the privacy-preserving functionality. In this paper, we propose a new framework that each respondent intentionally answers non-deterministic information instead of deterministic information. For example, he answers `either A, B, or C' instead of the actual choice A, and he intentionally dilutes his choice. This may be the similar concept on the k-anonymity. Non-deterministic information will be desirable for preserving each respondent's information. We follow the framework of Rough Non-deterministic Information Analysis (RNIA), and apply RNIA to the privacy-preserving questionnaire by non-deterministic information. In the current data mining algorithms, the tuples with non-deterministic information may be removed based on the data cleaning process. However, RNIA can handle such tuples as well as the tuples with deterministic information. By using RNIA, we can consider new types of privacy-preserving questionnaire.
Shusaku Tsumoto, Michinori Nakata, Hiroshi Sakai
IEEE BigData1
2015 Granular formalization of medical diagnostic process
abstract
This paper dicusses how to formalize medical diagnostic reasoing from the viewpoint of rule reasoning. Characteristics of rules shows that the rule model is closely related with rough set rule model. The important point is that medical diagnostic reasoning is characterized by focusing mechanism, composed of screening and differential diagnosis, which corresponds to upper approximation and lower approximation of a target concept in rough set theory. Furthremore, this paper focuses on detection of complications, which can be viewed as relations between rules of different diseases.
Shusaku Tsumoto, Shoji Hirano
IEEE BigData1
2015 Data decomposition and dual clustering for clinical care management
abstract
This paper proposes a method for construction of a clinical pathway based on attribute and sample clustering, called dual clustering. The method consists of the following five steps: first, histories of nursing orders are extracted from hospital information system. Second, orders are classified into several groups by using clustering on the pricipal components (sample clustering). Third, attributes clustering is applied to the data. Fourth, the method compares between generated functions for sample and attribute clustering which relate the number of clusters and calculated similarities. Fifth, if attribute clustering gives better performance with respect to the function, the dataset is decomposed into subtables by using the grouping of attribute clustering. Then, the first step will be repeated in a recursive way. After the grouping results are stable, a new pathway will be constructed from all the induced results. The method was applied to datasets of a disease extracted from a hospital information system. The results show that the proposed method is useful for construction of a clinical pathway.
Shusaku Tsumoto, Shoji Hirano, Haruko Iwata
IEEE BigData1
2013 Mining nursing care plan from data extracted from hospital information system
abstract
Schedule management of hospitalization is important to maintain or improve the quality of medical care. Application of a clinical pathway has been proposed as one of the important solutions for the management. This research proposed an data-oriented maintenance and construction of clinical pathways by using data on histories of nursing orders stored in hospital information system. The method was evaluated on data extracted from a hospital information system. The results show that the reuse of stored data will give a powerful tool for management of nursing schedule and lead to improvement of hospital services.
Shusaku Tsumoto, Shoji Hirano, Haruko Iwata
ASONAM1
2013 Granularity-based temporal data mining in hospital information system
abstract
This paper proposes granularity-based temporal data mining method which constructs clinical process conducted by nurses. The methods consist of three process. First, data on counting sum of executed orders are extracted from hospital informaton system with a given temporal granularity. Then, similarity-based methods, such as clustering and multidimensional scaling (MDS) are applied to the data and the labels for grouping are obtained. By using the labels, rule induction is applied, and classification power of each attribute is estimated. The attributes are sorted by an index of classification power, the original dataset is decomposed into subtables. Clustering, rule induction and table decomposition methods are applied to the subtables in a recursive way. The method was applied to datasets stored in hospital information system stored in 10 years. The results show that the reuse of stored data will give a powerful tool for construction of clinical process, which can be viewed as data-oriented management of nursing schedule.
Shusaku Tsumoto, Shoji Hirano, Haruko Iwata
IEEE BigData1
2013 Combinatorics of Information Granule in Contingency Table
abstract
This paper focuses on the degree of freedom and number of subdeterminants in a Pearson residual in a multiway contingency table. The results show that multidimensional residuals are represented as linear sum of determinants of 2 × 2 submatrices, which can be viewed as information granules measuring the degree of statistical dependence. Geometrical interpretation of Pearson residual is investigated. Furthermore, the number of subdeterminants in a residual is equal to the degree of freedom in χ2-test statistic. Since the way of calculation of the number of subdeterminants corresponds to the construction of a statistical model for a contingency table, it has been found that the combinatorics of the number subdeterminants is closely related with permutation of attributes in a given table, where symmetric group may play an important role.
Shusaku Tsumoto, Shoji Hirano
Int. J. Intell. Syst.1
2013 Clustering of non-metric proximity data based on bi-links with ϵ-indiscernibility
Shoji Hirano, Shusaku Tsumoto
J. Intell. Inf. Syst.2
2013 Special issue on challenges in knowledge discovery and data mining
Shusaku Tsumoto
J. Intell. Inf. Syst.1
2011 Special issue on data mining for decision making and risk management
Shusaku Tsumoto, Tzung-Pei Hong
J. Intell. Inf. Syst.1
2011 Detection of risk factors using trajectory mining
Shusaku Tsumoto, Shoji Hirano
J. Intell. Inf. Syst.1
2009 Contingency matrix theory: Statistical dependence in a contingency table
Shusaku Tsumoto
Inf. Sci.1
2006 Cluster Analysis of Time-Series Medical Data Based on the Trajectory Representation and Multiscale Comparison Techniques
abstract
This paper presents a cluster analysis method for multidimensional time-series data on clinical laboratory examinations. Our method represents the time series of test results as trajectories in multidimensional space, and compares their structural similarity by using the multiscale comparison technique. It enables us to find the part-to-part correspondences between two trajectories, taking into account the relationships between different tests. The resultant dissimilarity can be further used with clustering algorithms for finding the groups of similar cases. The method was applied to the cluster analysis of Albumin-Platelet data in the chronic hepatitis dataset. The results denonstrated that it could form interesting groups of cases that have high correspondence to the fibrotic stages.
Shoji Hirano, Shusaku Tsumoto
ICDM2
2006 Evaluating a Rule Evaluation Support Method Based on Objective Rule Evaluation Indices
Hidenao Abe, Shusaku Tsumoto, Miho Ohsaki, Takahira Yamaguchi
PAKDD2
2005 A Rule Evaluation Support Method with Learning Models Based on Objective Rule Evaluation Indexes
abstract
In this paper, we present a novel rule evaluation support method for post-processing of mined results with rule evaluation models based on objective indexes. Post-processing of mined results is one of the key issues to make a data mining process successfully. However, it is difficult for human experts to evaluate many thousands of rules from a large dataset with noises completely. To reduce the costs of rule evaluation procedures, we have developed the rule evaluation support method with rule evaluation models, which are obtained with objective rule evaluation indexes and evaluations of a human expert for each rule. Since the method is needed more accurate rule evaluation models, we have compared learning algorithms to construct rule evaluation models with the actual meningitis data mining result and actual rule sets from UCI datasets. Then we show the availability of our adaptive rule evaluation support method.
Hidenao Abe, Shusaku Tsumoto, Miho Ohsaki, Takahira Yamaguchi
ICDM2
2005 Clinical Decision Support Based on Mobile Telecommunication Systems
abstract
In this paper, we focus on the application of knowledge engineering techniques to medical mobile communication network, where the Web intelligence technologies are used for an efficient interface of medical expert system. Then, the system was put on the Internet to provide an intelligent decision support in telemedicine and is now being evaluated by region medical home doctors. The results show that such an Internet-based medical decision support enables home doctors to take a quick action to the applied domain.
Shusaku Tsumoto, Shoji Hirano, Hidenao Abe, Hideaki Nakakuni, Eisuke Hanada
Web Intelligence1
2005 Automated discovery of chronological patterns in long time-series medical datasets
abstract
Data mining in time-series medical databases has been receiving considerable attention because it provides a way of revealing useful information hidden in the database, for example, relationships between the temporal course of examination results and the onset time of diseases. This article presents a new method for finding similar patterns in temporal sequences. The method is a hybridization of phase-constraint multiscale matching and rough clustering. Multiscale matching enables us to cross-scale a comparison of the sequences, namely, it enables us to compare temporal patterns by partially changing observation scales. Rough clustering enables us to construct interpretable clusters of the sequences even if their similarities are given as relative similarities. We combine these methods and cluster the sequences according to the multiscale similarity of patterns. Experimental results on the chronic hepatitis dataset showed that clusters demonstrating interesting temporal patterns were successfully discovered. © 2005 Wiley Periodicals, Inc. Int J Int Syst 20: 737–757, 2005.
Shusaku Tsumoto, Shoji Hirano
Int. J. Intell. Syst.1
2004 Finding Interesting Pass Patterns from Soccer Game Records
Shoji Hirano, Shusaku Tsumoto
PKDD2
2004 Comparison of clustering methods for clinical databases
Shoji Hirano, Xiaoguang Sun, Shusaku Tsumoto
Inf. Sci.3
2004 Mining diagnostic rules from clinical databases using rough sets and medical diagnostic model
Shusaku Tsumoto
Inf. Sci.1
2003 Visualization of Rule's Similarity using Multidimensional Scaling
abstract
One of the most important problems with rule induction methods is that it is very difficult for domain experts to check millions of rules generated from large datasets. The discovery from these rules requires deep interpretation from domain knowledge. Although several solutions have been proposed in the studies on data mining and knowledge discovery, these studies are not focused on similarities between rules obtained. When one rule r/sub 1/ has reasonable features and the other rule r/sub 2/ with high similarity to r/sub 1/ includes unexpected factors, the relations between these rules will become a trigger to the discovery of knowledge. We propose a visualization approach to show the similar relations between rules based on multidimensional scaling, which assign a two-dimensional cartesian coordinate to each data point from the information about similarities between this data and others data. We evaluated this method on two medical data sets, whose experimental results show that knowledge useful for domain experts could be found.
Shusaku Tsumoto, Shoji Hirano
ICDM1
2003 Pattern Discovery based on Rule Induction and Taxonomy Generation
abstract
One of the most important problems with rule induction methods is that they cannot extract rules, which plausibly represent expert's decision processes. Here, the characteristics of expert's rules are closely examined and a new approach to extract plausible rules is introduced, which consists of the following three procedures. First, the characterization of decision attributes (given classes) is extracted from databases and the concept hierarchy for given classes is calculated. Second, based on the hierarchy, rules for each hierarchical level are induced from data. Then, for each given class, rules for all the hierarchical levels are integrated into one rule.
Shusaku Tsumoto, Shoji Hirano
ICDM1
2003 Dealing with Relative Similarity in Clustering: An Indiscernibility Based Approach
Shoji Hirano, Shusaku Tsumoto
PAKDD2
2003 An Indiscernibility-Based Clustering Method with Iterative Refinement of Equivalence Relations
Shoji Hirano, Shusaku Tsumoto
PKDD2
2003 Mining Rules of Multi-level Diagnostic Procedure from Databases
Shusaku Tsumoto
PKDD1
2003 Web Based Medical Decision Support System for Neurological Diseases
abstract
In early 1980s, many medical expert system were developed with knowledge bases which are acquired from medical experts, and their performance was almost as good as domain experts. However, they were not frequently used mainly due to the poor user interface and the lack in learning new knowledge. However, the solutions of these two problems have been introduced since 1990s. For the latter problem, machine learning methods have provided several solutions, and for the former problem, the rapid progress of Web technologies enables us to implement a good user interface. Furthermore, recent advances in computer resources strengthen these two solutions. We focus on the interface problem. The recent advances in Web technologies were used for an efficient interface of medical expert system. Then, the system was put on the Internet to provide an intelligent decision support in telemedicine and is now being evaluated by regional medical home doctors.
Shusaku Tsumoto
Web Intelligence1
2002 Mining Similar Temporal Patterns in Long Time-Series Data and Its Application to Medicine
abstract
Data mining in time-series medical databases has been receiving considerable attention since it provides a way of revealing useful information hidden in the database; for example relationships between temporal course of examination results and onset time of diseases. This paper presents a new method for finding similar patterns in temporal sequences. The method is a hybridization of phase-constraint multiscale matching and rough clustering. Multiscale matching enables us cross-scale comparison of the sequences, namely, it enable us to compare temporal patterns by partially changing observation scales. Rough clustering enable us to construct interpretable clusters of the sequences even if their similarities are given as relative similarities. We combine these methods and cluster the sequences according to multiscale similarity of patterns. Experimental results on the chronic hepatitis dataset showed that clusters demonstrating interesting temporal patterns were successfully discovered.
Shoji Hirano, Shusaku Tsumoto
ICDM2
2002 Multiscale Comparison of Temporal Patternsin Time-Series Medical Databases
Shoji Hirano, Shusaku Tsumoto
PKDD2
2002 Mining Hierarchical Decision Rules from Clinical Databases Using Rough Sets aaand Medical Diagnostic Model
Shusaku Tsumoto
PKDD1
2002 Analysis of amino-acid sequences by statistical technique
Shusaku Tsumoto, Shoji Hirano, Akira Yasuda, Kouhei Tsumoto
Inf. Sci.1
2001 Indiscernibility Degree of Objects for Evaluating Simplicity of Knowledge in the Clustering Procedure
abstract
The paper presents a novel, rough set-based clustering method that enables the evaluation of classification knowledge simplicity during the clustering procedure. The method iteratively refines equivalence relations so that they become a more simple set of relations that give adequate coarse classification to the objects. At each step of the iteration, the importance of the equivalence relation is evaluated on the basis of the newly introduced measure, indiscernibility degree. An indiscernibility degree is defined as a ratio of equivalence relations that classify the two objects into the same equivalence class. If an equivalence relation has the ability to discern two objects that have a high indiscernibility degree, a very fine classification is performed and then modified to regard them as indiscernible objects. The refinement is repeated, decreasing the threshold level of indiscernibility degree, and finally simple clusters can be obtained. Experimental results on the artificial data shows that iterative refinement of equivalence relation leads to successful generation of coarse clusters that can be represented by simple knowledge.
Shoji Hirano, Shusaku Tsumoto
ICDM2
2001 A Rough Set-Based Clustering Method with Modification of Equivalence Relations
Shoji Hirano, Tomohiro Okuzaki, Yutaka Hata, Shusaku Tsumoto, Kouhei Tsumoto
PAKDD4
2001 Discovery of Temporal Knowledge in Medical Time-Series Databases Using Moving Average, Multiscale Matching, and Rule Induction
Shusaku Tsumoto
PKDD1
2001 Mining Positive and Negative Knowledge in Clinical Databases Based on Rough Set Model
Shusaku Tsumoto
PKDD1
2000 Information Granules for Spatial Reasoning
Andrzej Skowron, Jaroslaw Stepaniuk, Shusaku Tsumoto
PAKDD3
2000 Evaluating Hypothesis-Driven Exception-Rule Discovery with Medical Data Sets
Einoshin Suzuki, Shusaku Tsumoto
PAKDD2
2000 Clinical Knowledge Discovery in Hospital Information Systems: Two Case Studies
Shusaku Tsumoto
PKDD1
2000 Knowledge discovery in clinical databases and evaluation of discovered knowledge in outpatient clinic
Shusaku Tsumoto
Inf. Sci.1
1999 Automated Discovery of Plausible Rules Based on Rough Sets and Rough Inclusion
Shusaku Tsumoto
PAKDD1
1999 Rule Discovery in Databases with Missing Values Based on Rough Set Model
Shusaku Tsumoto
PAKDD1
1999 Support Vector Machines for Knowledge Discovery
Shinsuke Sugaya, Einoshin Suzuki, Shusaku Tsumoto
PKDD3
1999 Rule Discovery in Large Time-Series Medical Databases
Shusaku Tsumoto
PKDD1
1999 Knowledge Discovery in Medical Multi-databases: A Rough Set Approach
Shusaku Tsumoto
PKDD1
1998 Discovery of Approximate Medical Knowledge Based on Rough Set Model
Shusaku Tsumoto
PKDD1
1998 Automated Extraction of Medical Expert System Rules from Clinical Databases on Rough Set Theory
Shusaku Tsumoto
Inf. Sci.1
1997 Extraction of Experts' Decision Process from Clinical Databases Using Rough Set Model
Shusaku Tsumoto
PKDD1
1996 Automated Discovery of Medical Expert System Rules from Clinical Databases Based on Rough Sets
Shusaku Tsumoto, Hiroshi Tanaka
KDD1
1995 Automated Selection of Rule Induction Methods Based on Recursive Iteration of Resampling Methods and Multiple Statistical Testing
Shusaku Tsumoto, Hiroshi Tanaka
KDD1
1995 Automated Discovery of Functional Components of Proteins from Amino-Acid Sequences Based on Rough Sets and Change of Representation
Shusaku Tsumoto, Hiroshi Tanaka
KDD1
1995 COBRA: Integration of Heterogeneous Knowledge-Bases in Medical Domain
abstract
Medical data consist of many kinds of data from different resources, such as natural language data, sound data from physical examinations, numerical data from laboratory examinations, time-series data from monitoring systems, and medical images (e.g. X-ray, Computer Tomography, and Magnetic Resonance Image). Therefore it has been pointed out that medical databases should be implemented as multidatabases. However, there have been few systems which integrate these data into multidatabases. In this paper, we report a system called COBRA (Computer-Operated Birth-defect Recognition Aid), which supports diagnosis and information retrieval of congenital malformation diseases and which also integrates natural language data, sound data, numerical data, and medical images into multidatabases on syndrome of congenital malformation. The results show that object-oriented scheme makes it easy to implement and integrate these knowledge-databases in COBRA, which suggests that these clinical databases should be implemented as object-oriented databases.
Shusaku Tsumoto, Hiroshi Tanaka, Hiromi Amano, Kimie Ohyama, Takayuki Kuroda
Int. J. Cooperative Inf. Syst.1