VLDB 2026 Research / reviewers in the wild / expert
Vincent S. Tseng
dblp:t/VincentSMTseng · also Shin-Mu Tseng, Vincent Shin-Mu Tseng
· DBLP profile ↗
191ranked-venue papers
34as first author
29since 2021 · last 2026
0000-0002-4853-1594ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 91 · 11 first-author · 13 since 2021Databases, data management, data science and information retrieval · 78 · 10 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 34 · 8 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 8 · 5 first-author · 1 since 2021Software engineering, systems software and programming languages · 7 · 3 first-authorSystems, architecture and hardware · 3 · 1 first-authorTheory of computation · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-source domain adaptive object detection under different privacy levels
Peggy Joy Lu, Wei-Yu Chen, Chia-Yung Jui, Vincent S. Tseng, Jen-Hui Chuang |
Multim. Syst. | 4 |
| 2025 | Prompt Learner for Industrial Anomaly Detection: Towards Generalizable Multimodal DefectabstractIndustrial Anomaly Detection (IAD) plays a key role in the field of smart manufacturing and quality control. However, due to the challenges of working with rare anomaly samples, high diversity, and complex data types, traditional methods often fail to yield stable results in practical applications. This study proposes a multimodal anomaly detection framework called Prompt Learner for Anomaly Detection, which integrates image and text data, and combines adaptive prompt learning with an efficient distributed training mechanism to improve the accuracy and flexibility of anomaly detection. This method makes three significant optimizations to the original AnomalyGPT architecture: First, the enhanced Llama MLP module is introduced to enhance feature screening capabilities and reduce overfitting. Second, the optimized self-attention mechanism is combined with RoPE position embedding and Past Key-Value Caching to improve long sequence processing performance and inference speed effectively. Thirdly, the adaptive Prompt Learner module can dynamically generate semantic prompts based on input image features to enhance the generalization ability of unknown anomalies.In addition, this study also designed three loss functions to balance the classification and segmentation tasks and used Poisson Image Editing technology for data augmentation purposes. The experiment utilized standard industrial datasets, including MVTec-AD, VisA, and AeBAD-S, for evaluation. The results showed that this method outperformed existing advanced techniques in both image and pixel-level AUC indicators, achieving a PRO score of 92.56% in the AeBAD-S dataset, which demonstrates excellent anomaly localization capabilities. Overall, this study demonstrated the practical potential of multimodal fusion, semantic adaptation, and efficient training strategies in industrial anomaly detection, laying a good foundation for practical applications and subsequent research. Yen-Han Chiang, Pablo Mollá, Yi-Lun Pan, Yi-Ju Tseng, Vincent S. Tseng |
ICMLA | 5 |
| 2025 | DL-KDD: Dual-Lightness Knowledge Distillation for Action Recognition in the DarkabstractHuman action recognition in dark videos is a challenging task for computer vision due to the low quality of the videos filmed in the dark. Recent studies focused on applying dark enhancement methods to improve the visibility of the video. However, such video processing results in the loss of critical information in the original (un-enhanced) video. Conversely, traditional two-stream methods are capable of learning information from both original and enhanced videos, but it can lead to a significant increase in the computational cost. To address these challenges, we propose a novel knowledge-distillation-based framework, named Dual-Lightness KnowleDge Distillation (DL-KDD), which simultaneously resolves the aforementioned issues by enabling a student model to obtain both original features and light-enhanced knowledge without additional complexity, thus improving the performance of the model and avoiding extra computational cost. Through comprehensive evaluations, the proposed DL-KDD, with only original video required as input during the inference phase, significantly outperforms state-of-the-art methods on the widely-used dark video datasets. The results highlight the excellence of our proposed knowledge-distillation-based framework for dark video human action recognition. Chi-Jui Chang, Oscar Tai-Yuan Chen, Vincent S. Tseng |
IJCAI | 3 |
| 2025 | Multi-Agent Transformer-based Automated Imbalanced Time Series Classification with Hyperparameter OptimizationabstractTime series classification is a prevalent research topic with wide applications, and Imbalanced Time Series Classification (ITSC) has emerged with high importance due to the nature of imbalanced class distribution in real-world applications. However, most existing ITSC methods rely heavily on manual adjustments that limit their scalability and performance. To address this, we introduce Multi-Agent Transformer-based Automated Imbalanced Time Series Classification (MAT-AITSC), a novel framework that includes an ITSC pipeline and a Multi-Agent Transformer-based Hyperparameter Optimization method named MAT-HPO method named MAT-HPO. Our pipeline uses cost-sensitive learning to address class imbalance while preserving the distribution of time series. MAT-HPO sequentially optimizes model architecture, class weights, and training hyperparameters to enhance classification performance on imbalanced time series datasets. Extensive evaluations on 52 diverse time series datasets demonstrate that MAT-AITSC significantly outperforms existing methods. To the best of our knowledge, this is the first work that offers an automated solution for imbalanced time series classification, reducing the need for manual intervention and paving the way for broader real-world applications. Nai-Hsin Cheng, Yujia Wu, Vincent S. Tseng |
IJCNN | 3 |
| 2025 | DiPPSI: Diffusion-Based Pulsative Physiological Signal Imputation
Su-Jung Wu, Jia-Ching Ying, Vincent S. Tseng |
PAKDD (1) | 3 |
| 2025 | HaGAR: Hardness-aware Generative Adversarial RecommenderabstractImplicit Collaborative filtering is a fundamental technique in recommendation systems, leveraging implicit user interactions to suggest items of interest. A significant challenge in this domain is the absence of explicit negative feedback, limiting the recommendation performance. Previous researchers have tried to tackle the challenge through the Generative Adversarial Network (GAN). The generator produces increasingly challenging samples for the discriminator, driving the optimization of the discrimination objective. Although GAN-style recommender systems can achieve decent performance by generating harder negative samples, the negatives selected by the generator may not always be ideal for training the discriminator. In this study, we focus on two types of undesirable negatives that persist in modern GAN-style recommenders: false negatives and uninformative negatives. In response to these issues, we propose a novel Hardness-aware Generative Adversarial Recommender (HaGAR). To the best of our knowledge, it is the first adversarial recommender that explicitly aims to alleviate the adverse impact of false and uninformative negatives. Our approach incorporates a relevance monitoring module and a hardness-aware weighting module to identify and address false and uninformative negatives during training with minimal additional computational cost. Our experimental results demonstrate that HaGAR significantly improves recommendation performance, achieving over a 21% increase in terms of NDCG@10 compared to the state-of-the-art GAN-style recommender. These findings highlight the efficacy of our improvement in providing more robust negative samples, leading to better-performing recommendation systems. Yuan-Heng Lee, Jia-Ching Ying, Vincent S. Tseng |
WSDM | 3 |
| 2025 | Forecasting leading economic indicators in the US from financial news using multi-task learning
Jia-Ching Ying, Chia-Chen Liu, Vincent S. Tseng, Wenbin Zhang 0002, Ji Zhang 0001 |
Soft Comput. | 3 |
| 2025 | Predicting Longitudinal Visual Field Progression With Class Imbalanced DataabstractGlaucoma is the leading cause of irreversible blindness worldwide. The clinical standard for glaucoma diagnosis and progression tracking remains visual field (VF) testing via standard automated perimetry. One outstanding challenge of many ophthalmic prediction tasks is the issue of class imbalance, where the majority class outnumbers the minority class(es). Although this issue has been reported in several prior studies on the prediction of VF progression or glaucoma, it has not been addressed in the context of longitudinal VF data. In this work, we proposed, VF-Transformer, a transformer-based framework for VF progression prediction based on longitudinal VF examination results. In particular, we addressed the class imbalance issue by incorporating our proposed inverted class-dependent temperature (ICDT) loss and weight normalization. The proposed framework was developed and evaluated on a public VF dataset and further validated on an external hospital dataset, using accuracy, sensitivity, specificity, and area under the receiver operating characteristic curve (AUC) as evaluation metrics. Extensive experiments and comparisons with existing state-of-the-art methods and class imbalance handling strategies confirmed the effectiveness of the proposed framework in predicting VF progression in the presence of class imbalance. Ling Chen 0004, Chun-Hung Chen, Da-Wen Lu, Vincent S. Tseng |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | HiTrace: Hierarchical Class Tracing Approach for Open-Set Recognition on Skin LesionsabstractIn the constantly evolving field of artificial intelligence (AI), identifying unknown or novel classes, known as Open-set Recognition (OSR), is critical for ensuring the reliability of AI models in various applications, especially in the vital field of medical diagnostics. This study introduces a novel approach to advance OSR through the proposed hierarchical class tracing approach, HiTrace, for skin lesion classification. HiTrace incorporates three key components that tackle the complex challenges in OSR: the Hierarchy-Aware Prototype (HAP) learning for an efficient training strategy, the Distribution Enhancement (DE) module for optimized post-processing feature adjustment, and the Potential Class Tracing Algorithm (PCTA) for a hierarchical classification decision-making process. HiTrace leverages a hierarchical taxonomy to simplify the identification of new skin conditions, reducing the need for extensive manual annotation and addressing the limitations of existing methods. This study also introduces two novel evaluation metrics, the Hierarchical Open-set Classification Score (HOC-Score) and Major-Type Accuracy for Open-set samples (MTACC-O), which provide robust criteria for assessing a model's performance in classifying closed-set, in-taxonomy, and out-of-taxonomy results. Notably, the experimental results demonstrate significant advancements in the PAD-UFES-20 and ISIC 2019 datasets, with relative improvements of 15.3% and 21.1% in HOC-Score and 12.3% and 5.8% in MTACC-O, respectively, without compromising on competitive closed-set performance. To the best of our knowledge, this is the first study to present a three-stage analysis (i.e., closed-set, in-taxonomy, and out-of-taxonomy classification results) of OSR applied to a practical medical field. This comprehensive approach represents an influential stride in enhancing patient care through the early detection and treatment of skin diseases, paving the way for future research and development in medical diagnostics and beyond. Benny Wei-Yun Hsu, Vincent S. Tseng |
IEEE J. Biomed. Health Informatics | 2 |
| 2024 | An MDL-Based Genetic Algorithm for Genome Sequence CompressionabstractThe exponential growth of genomic data has posed significant challenges for lossless compression of genome sequences. While recent reference-free genome compressors have shown promising results, they often fail to fully leverage the inherent sequential structure of genome sequences, require substantial computational resources and lack (or have limited) interpretability. This paper presents a novel genome compression method that employs the Minimum Description Length (MDL) principle, which is based on the idea that the best model for a given dataset is the one that provides the shortest description of that dataset. The proposed compressor, called GMG (Genetic algorithm for MDL-based Genome compression), integrates a genetic algorithm to identify optimal k-mers (patterns) in a model to best compress the genome data. Experimental results across various datasets demonstrate that GMG outperforms state-of-the-art genome compressors in terms of bits-per-base compression and computational efficiency. Furthermore, it is demonstrated that the optimal patterns identified by GMG for compression can also be utilized for genome classification, offering a multifunctional advantage over previous compressors. GMG is freely available at github.com/MuhammadzohaibNawaz/GMG Muhammad Zohaib Nawaz, M. Saqib Nawaz, Philippe Fournier-Viger, Vincent S. Tseng |
BIBM | 4 |
| 2024 | Self-Supervised Pulse-Aware Interpretable Disentangled ECG Representation LearningabstractElectrocardiography (ECG) is a widely used cardiac measurement for detecting cardiovascular conditions, while self-supervised learning leverages unlabeled data for model pre-training. However, current self-supervised frameworks for ECG signals generally lack a comprehensive understanding of intra-heartbeat and inter-heartbeat representation, which are crucial in the clinical interpretation of ECG data. In this work, a novel self-supervised pulse-aware interpretable disentangled ECG representation learning framework named SPIDER is proposed. The branch structure in SPIDER disentangles the general-purpose representation to encode specific information. The SPIDER framework notably enhances the performance in terms of area under the precision-recall curve (AUPRC), accompanied by a twofold improvement in training efficiency compared to alternative methods. Moreover, the design of the heartbeat branch provides interpretable heartbeat-level representations. The proposed SPIDER framework not only improves the performances on downstream tasks but also enhances training efficiency and interpretability, and these benefits are particularly valuable in real-world medical applications. Chun-Ti Chou, Vincent S. Tseng |
ICASSP | 2 |
| 2024 | Privacy-Preserving Attention-Weighted Multi-Source Domain Adaptation for EEG Motor ImageryabstractMotor Imagery (MI) is essential in the Brain-Computer Interface (BCI). Given the multi-source nature of EEG instability across subjects/sessions and the privacy concerns when dealing with data from other institutions, the importance of multi-source-free domain adaptation (MSFDA) becomes evident in conducting subject-independent applications. However, current MI approaches are lack of MSFDA for considering different sources’ importance, which may lead to negative transfer. In this work, we propose a novel two-stage MSFDA framework, namely Privacy-preserving Attention-Weighted domain adaptation (PAW), which is the first to emphasize different source importance in EEG MI by identifying crucial features across subjects and sessions. A domain discriminator is designed to find robust features within sessions in the training phase. In the adaptation phase, an attention weighted module determines the source importance, and a balanced confident set policy mitigates noise and imbalance samples. PAW delivers excellent performance via empirical evaluations on several public datasets. In particular, its adaptation phase is applicable across MSFDA scenarios, highlighting its broad application values. Yu-Mei Huang, Hui-Nien Hung, Vincent S. Tseng |
ICASSP | 3 |
| 2024 | Periodic Stacked Transformer-based Framework for Travel Time PredictionabstractTravel time analysis and prediction play critical roles in developing Intelligent Transportation Systems (ITS), which have attracted significant interests from the research community. Deep learning-based methodologies have proven to be powerful tools in utilizing big data for predicting travel times. However, while most studies have focused on short-term predictions, predicting travel times over longer periods is equally important for wide applications like traffic management and route planning. Long-term prediction, which often receives less attention due to its complexity, remains a gap in current researches. To address this challenge, we propose the Periodic Stacked Transformer (PS-Transformer), a novel Transformer-based framework designed to enhance both short and long-term traffic predictions. PS-Transformer consists of two primary modules: the Segment Encoding Integration (SEI) and the Periodic Stacked Encoder-Decoder (PSED). SEI module extracts periodic patterns from traffic data, while PSED effectively captures short-term and long-term dependencies from temporal attributes. Additionally, PSED tackles error accumulation, a common issue in extended prediction periods, through its non-autoregressive decoder design. Our PS-Transformer is validated through a series of experiments on a real-world dataset, demonstrating its capability in multi-step predictions that provide forecasts over an extended duration. Empirical evaluation results show that PS-Transformer outperforms state-of-the-art methods in both short and long-term travel time predictions across various metrics, including MAE, RMSE, and SMAPE. Hui-Ting Lin, Vincent S. Tseng |
IJCNN | 3 |
| 2024 | Learning Location Semantics and Dynamics for Traffic Origin-Destination Demand PredictionabstractTraffic Origin-Destination (OD) Demand Prediction is pivotal for real-time ride-hailing services and government traffic management, aiming to anticipate traffic patterns and volumes between specific locations. Traditional grid-based traffic OD demand prediction methods, however, often fail to accurately capture the intricate and context-driven demand patterns inherent in modern urban transportation systems. In this paper, we propose an innovative location semantics and dynamics learning framework for capturing location semantics and dynamics to improving origin-destination traffic demand prediction. Departing from conventional grid-based methods, our approach incorporates a semantic location generation module that dynamically organizes semantic locations based on Points of Interest (POI). Our semantic location generation module is designed to form semantic locations according to POI types, spatial proximity, and demand patterns. This model captures hierarchical relationships and varying importance levels across POI types and domains. To learn the context and function of each area, we design a POI context extractor which can analyze contextual information of POIs within a specified radius. The POI Type Encoder utilizes advanced word embedding techniques to encode the POIs, enabling the model to comprehend the semantic significance of diverse locations and their influence on traffic patterns. Extensive experiments conducted on real-world New York City Taxi datasets demonstrate the superior performance of the proposed framework over state-of-the-art methods, as indicated by experiment results across all evaluation measurement. These findings affirm that our proposed framework could improve effectiveness by learning location semantics and dynamics. Kuan-Hsuan Yung, Jia-Ching Ying, Hui-Ting Lin, Vincent S. Tseng |
IJCNN | 4 |
| 2024 | Federated Contrastive Domain Adaptation for Category-inconsistent Object DetectionabstractTo obtain diverse scenarios for collaboratively training a more generalized object detector, the multi-source domain adaptive object detection has been proposed. However, such scenario faces challenges related to data privacy, domain discrepancy and category inconsistency, thus we propose a framework called FedCoin: Federated Contrastive domain-adaptation for category-inconsistent object detection. On the client sides, a novel dynamic model contrastive strategy is proposed to reduce excessively domain-specific features from local models. On the server side, we design a two-stage teacher-student architecture to tackle the challenge of backbone aggregation and inconsistent categories integration. Our method outperforms SOTA methods across different domain adaptation tasks, with an average precision increase of 6% on various datasets, demonstrating its superiority over existing methods for category-inconsistent and privacy-preserving scenarios. The source code is available online: https://github.com/ccuvislab/FedCoin Wei-Yu Chen, Peggy Joy Lu, Vincent S. Tseng |
VCIP | 3 |
| 2024 | Snippet Policy Network V2: Knee-Guided Neuroevolution for Multi-Lead ECG Early ClassificationabstractEarly time series classification predicts the class label of a given time series before it is completely observed. In time-critical applications, such as arrhythmia monitoring in ICU, early treatment contributes to the patient's fast recovery, and early warning could even save lives. Hence, in these cases, it is worthy of trading, to some extent, classification accuracy in favor of earlier decisions when the time series data are collected over time. In this article, we propose a novel deep reinforcement learning-based framework, snippet policy network V2 (SPN-V2), for long and varied-length multi-lead electrocardiogram (ECG) early classification. The proposed SNP-V2 contains two main components: snippet representation learning (SRL) and early classification timing learning (ECTL). The SRL is proposed to encode inner-snippet spatial correlations and inter-snippet temporal correlations into the hidden representations of the subsegment (snippet) of the input ECG. ECTL aims to learn a decision agent to classify the time series early and accurately. To optimize the proposed framework, we design a novel knee-guided neuroevolution algorithm (KGNA) to solve cardiovascular diseases' early classification problem, automatically optimizing the proposed SPN-V2 regarding the tradeoff between accuracy and earliness. In addition, we conduct a series of experiments on two real-world ECG datasets. The experimental results show the superiority of the proposed algorithm over the state-of-the-art competing methods. Yu Huang 0018, Gary G. Yen, Vincent S. Tseng |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Continual learning with attentive recurrent neural networks for temporal data classification
Shao-Yu Yin, Yu Huang 0018, Tien-Yu Chang, Shih-Fang Chang, Vincent S. Tseng |
Neural Networks | 5 |
| 2023 | PoEMS: Policy Network-Based Early Warning Monitoring System for Sepsis in Intensive Care UnitsabstractSepsis is among the leading causes of morbidity and mortality in modern intensive care units (ICU). Due to accurate and early warning, the in-time antibiotic treatment of sepsis is critical for improving sepsis outcomes, contributing to saving lives, and reducing medical costs. However, the earlier prediction of sepsis onset is made, the fewer monitoring measurements can be processed, causing a lower prediction accuracy. In contrast, a more accurate prediction can be expected by analyzing more data but leading to the delayed warning associated with life-threatening events. In this study, we propose a novel deep reinforcement learning framework for solving early prediction of sepsis, called the Policy Network-based Early Warning Monitoring System (PoEMS). The proposed PoEMS provides accurate and early prediction results for sepsis onset based on analyzing varied-length electronic medical records (EMR). Furthermore, the system serves by monitoring the patient's health status consistently and provides an early warning only when a high risk of sepsis is detected. Additionally, a controlling parameter is designed for users to adjust the trade-off between earliness and accuracy, providing the adaptability of the model to meet various medical requirements in practical scenarios. Through a series of experiments on real-world medical data, the results demonstrate that our proposed PoEMS achieves a high AUROC result of more than 91% for early prediction, and predicts sepsis onset earlier and more accurately compared to other state-of-the-art competing methods. Hsin-Ginn Hwang, Vincent S. Tseng |
IEEE J. Biomed. Health Informatics | 3 |
| 2023 | Snippet Policy Network for Multi-Class Varied-Length ECG Early ClassificationabstractArrhythmia detection from ECG is an important research subject in the prevention and diagnosis of cardiovascular diseases. The prevailing studies formulate arrhythmia detection from ECG as a time series classification problem. Meanwhile, early detection of arrhythmia presents a real-world demand for early prevention and diagnosis. In this paper, we address a problem of cardiovascular diseases early classification, which is a varied-length and long-length time series early classification problem as well. For solving this problem, we propose a deep reinforcement learning-based framework, namely Snippet Policy Network (SPN), consisting of four modules, snippet generator, backbone network, controlling agent, and discriminator. Comparing to the existing approaches, the proposed framework features flexible input length, solves the dual-optimization solution of the earliness and accuracy goals. Experimental results demonstrate that SPN achieves an excellent performance of over 80% in terms of accuracy. Compared to the state-of-the-art methods, at least 7% improvement on different metrics, including the precision, recall, F1-score, and harmonic mean, is delivered by the proposed SPN. To the best of our knowledge, this is the first work focusing on solving the cardiovascular early classification problem based on varied-length ECG data. Based on these excellent features from SPN, it offers a good exemplification for addressing all kinds of varied-length time series early classification problems. Yu Huang 0018, Gary G. Yen, Vincent S. Tseng |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2022 | Contrastive Heartbeats: Contrastive Learning for Self-Supervised ECG Representation and PhenotypingabstractThe non-invasive and easily accessible characteristics of electrocardiogram (ECG) attract many studies targeting AI-enabled cardiovascular-related disease screening tools based on ECG. However, the high cost of manual labels makes high-performance deep learning models challenging to obtain. Hence, we propose a new self-supervised representation learning framework, contrastive heartbeats (CT-HB), which learns general and robust electrocardiogram representations for efficient training on various downstream tasks. We employ a novel heartbeat sampling method to define positive and negative pairs of heartbeats for contrastive learning by utilizing the periodic and meaningful patterns of electrocardiogram signals. Using the CT-HB framework, the self-supervised learning model learns personalized heartbeat representations representing the specific cardiology context of a patient. Evaluations on public benchmark datasets and a private large-scale real-world dataset with multiple tasks demonstrate that the learned semantic representations result in better performance on downstream tasks and retain high performance while supervised learning suffers performance degradation with fewer supervised labels in downstream tasks. Crystal T. Wei, Ming-En Hsieh, Chien-Liang Liu, Vincent S. Tseng |
ICASSP | 4 |
| 2022 | Periodic Attention-based Stacked Sequence to Sequence framework for long-term travel time prediction
Yu Huang 0018, Vincent S. Tseng |
Knowl. Based Syst. | 3 |
| 2022 | Mining frequent weighted utility itemsets in hierarchical quantitative databases
Ham Nguyen, Tuong Le, Philippe Fournier-Viger, Vincent S. Tseng, Bay Vo |
Knowl. Based Syst. | 5 |
| 2022 | A Novel Constraint-Based Knee- Guided Neuroevolutionary Algorithm for Context-Specific ECG Early ClassificationabstractCardiovascular diseases (CVDs) are considered the greatest threat to human life according to World Health Organization. Early classification of CVDs and the appropriate follow-up treatment are crucial for preventing sudden deaths. Electrocardiogram (ECG) is one of the most common non-invasive tools used to evaluate the state of the heart, which can be exploited to automatically diagnose as well. However, the importance of diagnosing CVDs is varying in different context-specific scenarios. For example, ST-segment elevation (STE) is an acute myocardial infarction indicator for patients associated with chest pain and cardiac biomarker. In in-hospital healthcare, STE should be diagnosed with a higher priority than the other phenotypes of ECG. Hence, the context-specific requirements should be considered in ECG early classification problems. We formalize the ECG early classification problem as the context-specific time series classification problem. We propose a novel Constraint-based Knee-guided Neuroevolutionary Algorithm (CKNA) based on the Snippet Policy Networks V2 to solve this problem. To validate the proposed method, we perform a series of experiments on two public ECG datasets under various context-specific simulated scenarios after consulting with physicians specializing in the area. Experimental results show that CKNA significantly improves the average recall of disease classification by 5.5% compared to the competing baseline under user-specified requirements. Moreover, experimental results prove that CKNA presents a feasible solution for the early classifying of cardiac arrhythmias under different user-specified scenarios. Yu Huang 0018, Gary G. Yen, Vincent S. Tseng |
IEEE J. Biomed. Health Informatics | 3 |
| 2021 | GAWD: graph anomaly detection in weighted directed graph databasesabstractGiven a set of node-labeled directed weighted graphs, how to find the most anomalous ones? How can we summarize the normal behavior in the database without losing information? We propose GAWD, for detecting anomalous graphs in directed weighted graph databases. The idea is to (1) iteratively identify the "best" substructure (i.e., subgraph or motif) that yields the largest compression when each of its occurrences is replaced by a super-node, and (2) score each graph by how much it compresses over iterations --- the more the compression, the lower the anomaly score. Different from existing work [1] on which we build, GAWD exhibits (i) a lossless graph encoding scheme, (ii) ability to handle numeric edge weights, (iii) interpretability by common patterns, and (iv) scalability with running time linear in input size. Experiments on four datasets injected with anomalies show that GAWD achieves significantly better results than state-of-the-art baselines. Meng-Chieh Lee, Hung T. Nguyen 0003, Dimitris Berberidis, Vincent S. Tseng, Leman Akoglu |
ASONAM | 4 |
| 2021 | Stable High Utility Itemset MiningabstractHigh Utility Itemset Mining (HUIM) aims at finding all sets of items that have high importance in a database, as measured by a utility function. Although HUIM has many applications, a key limitation is that the discovered patterns often have an unstable utility over time. For example, while a set of products may yield a high utility (profit) over a year, that utility may fluctuate from weeks to weeks. To discover patterns that have a stable utility and hence that are more suitable for decision-making, this paper redefines HUIM as the task of discovering Stable High Utility Itemsets (StableHUI). An efficient tree-based and pattern-growth algorithm named Stable-Growth is proposed to extract all the StableHUI. Several experiments on two real-world datasets and two synthetic datasets show that Stable-Growth is up to 60% faster than a baseline and that it can filter out numerous unstable HUI. Acquah Hackman, Yu Huang 0018, Philippe Fournier-Viger, Vincent S. Tseng |
iiWAS | 4 |
| 2021 | Spatio-attention embedded recurrent neural network for air quality prediction
Yu Huang 0018, Jia-Ching Ying, Vincent S. Tseng |
Knowl. Based Syst. | 3 |
| 2021 | Guest Editorial: Artificial Intelligence for Securing Industrial-Based Cyber-Physical SystemsabstractN/A Gautam Srivastava 0001, Jerry Chun-Wei Lin, Xuyun Zhang, Vincent S. Tseng |
IEEE Trans. Ind. Informatics | 4 |
| 2021 | Dynamic Graph Mining for Multi-weight Multi-destination Route Planning with Deadlines ConstraintsabstractRoute planning satisfied multiple requests is an emerging branch in the route planning field and has attracted significant attention from the research community in recent years. The prevailing studies focus only on seeking a route by minimizing a single kind of Travel Cost, such as trip time or distance, among others. In reality, most users would like to choose an appropriate route, neither fastest nor shortest route. Usually, a user may have multiple requirements, and an appropriate route would satisfy all requirements requested by the user. In fact, planning an appropriate route could be formulated as a problem of Multi-weight Multi-destination Route Planning with Deadlines Constraints (MWMDRP-DC). In this article, we propose a framework, namely, MWMD-Router, which addresses the MWMDRP-DC problem comprehensively. To consider the travel costs with time-variation, we propose not only four novel dynamic graph miner to extract travel costs that reveal users’ requirements but also two new algorithms, namely, Basic MWMD Route Planning and Advanced MWMD Route Planning , to plan a route that satisfies deadline requirements and optimizes another criterion like travel cost with time-variation efficiently. To the best of our knowledge, this is the first work on route planning that considers handling multiple deadlines for multi-destination planning as well as optimizing multiple travel costs with time-variation simultaneously. Experimental results demonstrate that our proposed algorithms deliver excellent performance with respect to efficiency and effectiveness. Yu Huang 0018, Jia-Ching Ying, Philip S. Yu, Vincent S. Tseng |
ACM Trans. Knowl. Discov. Data | 4 |
| 2021 | A Survey of Utility-Oriented Pattern MiningabstractThe main purpose of data mining and analytics is to find novel, potentially useful patterns that can be utilized in real-world applications to derive beneficial knowledge. For identifying and evaluating the usefulness of different kinds of patterns, many techniques and constraints have been proposed, such as support, confidence, sequence order, and utility parameters (e.g., weight, price, profit, quantity, satisfaction, etc.). In recent years, there has been an increasing demand for utility-oriented pattern mining (UPM, or called utility mining). UPM is a vital task, with numerous high-impact applications, including cross-marketing, e-commerce, finance, medical, and biomedical applications. This survey aims to provide a general, comprehensive, and structured overview of the state-of-the-art methods of UPM. First, we introduce an in-depth understanding of UPM, including concepts, examples, and comparisons with related concepts. A taxonomy of the most common and state-of-the-art approaches for mining different kinds of high-utility patterns is presented in detail, including Apriori-based, tree-based, projection-based, vertical-/horizontal-data-format-based, and other hybrid approaches. A comprehensive review of advanced topics of existing high-utility pattern mining techniques is offered, with a discussion of their pros and cons. Finally, we present several well-known open-source software packages for UPM. We conclude our survey with a discussion on open and practical challenges in this field. Wensheng Gan, Jerry Chun-Wei Lin, Philippe Fournier-Viger, Han-Chieh Chao, Vincent S. Tseng, Philip S. Yu |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2020 | AutoAudit: Mining Accounting and Time-Evolving GraphsabstractHow can we spot money laundering in large-scale graph-like accounting datasets? How to identify the most suspicious period in a time-evolving accounting graph? What kind of accounts and events should practitioners prioritize under time constraints? To tackle these crucial challenges in accounting and auditing tasks, we propose a flexible system called AutoAudit, which can be valuable for auditors and risk management professionals. To sum up, there are four major advantages of the proposed system: (a) "Smurfing" Detection, spots nearly 100% of injected money laundering transactions automatically in real-world datasets. (b) Attention Routing, attends to the most suspicious part of time-evolving graphs and provides an intuitive interpretation. (c) Insight Discovery, identifies similar month-pair patterns proved by "success stories" and patterns following Power Laws in log-logistic scales. (d) Scalability and Generality, ensures AutoAudit scales linearly and can be easily extended to other real-world graph datasets. Experiments on various real-world datasets illustrate the effectiveness of our method. To facilitate reproducibility and accessibility, we make the code, figure, and results public at https://github.com/mengchillee/AutoAudit. Meng-Chieh Lee, Yue Zhao 0016, Aluna Wang, Pierre Jinghong Liang, Leman Akoglu, Vincent S. Tseng, Christos Faloutsos |
IEEE BigData | 6 |
| 2020 | Efficient methods for mining weighted clickstream patterns
Huy Minh Huynh, Loan T. T. Nguyen, Bay Vo, Anh Nguyen 0006, Vincent S. Tseng |
Expert Syst. Appl. | 5 |
| 2019 | Mining Emerging High Utility Itemsets over Streaming Database
Acquah Hackman, Yu Huang 0018, Philip S. Yu, Vincent S. Tseng |
ADMA | 4 |
| 2019 | DeepIdentifier: A Deep Learning-Based Lightweight Approach for User Identity Recognition
Meng-Chieh Lee, Yu Huang 0018, Jia-Ching Ying, Chien Chen, Vincent S. Tseng |
ADMA | 5 |
| 2019 | Robust Sensor-based Human Activity Recognition with Snippet Consensus Neural NetworksabstractSensor-based human activity recognition is an important problem in pervasive computing, which has attracted lots of attention from the research community in the past few years. The existing relevant studies focused on using handcrafted features or machine learning-based methods to tackle this problem. However, these methods are usually limited to specific datasets, such that the generality is limited. Some methods are also limited to strict experimental environments, which do not take stability into consideration. In this paper, we propose a robust and novel deep learning-based framework, named Snippet Consensus Neural Networks (SCNet), which aims to conquer these challenges. Through a series of experiments, the proposed framework is verified to outperform seven state-of-the-art methods on five datasets in terms of not only accuracy but also generality and stability, averagely improving 10% on mean accuracy. Yu Huang 0018, Meng-Chieh Lee, Vincent S. Tseng, Ching-Jui Hsiao, Chi-Chiang Huang |
BSN | 3 |
| 2019 | Long-Term Traffic Time Prediction Using Deep Learning with Integration of Weather Effect
Chih-Hsin Chou, Yu Huang 0018, Chian-Yun Huang, Vincent S. Tseng |
PAKDD (2) | 4 |
| 2019 | Multivariate Time Series Early Classification with Interpretability Using Deep Learning and Attention Mechanism
En-Yu Hsu, Chien-Liang Liu, Vincent S. Tseng |
PAKDD (3) | 3 |
| 2019 | Parallel Mining of Top-k High Utility Itemsets in Spark In-Memory Computing Architecture
Chun-Han Lin, Cheng-Wei Wu, JianTao Huang, Vincent S. Tseng |
PAKDD (2) | 4 |
| 2019 | Mining high-utility itemsets in dynamic profit databases
Loan T. T. Nguyen, Trinh D. D. Nguyen, Bay Vo, Philippe Fournier-Viger, Vincent S. Tseng |
Knowl. Based Syst. | 6 |
| 2019 | Discovering negative comments by sentiment analysis on web forum
Wei-Yun Hsu, Hui-Huang Hsu, Vincent S. Tseng |
World Wide Web | 3 |
| 2018 | Music Recommendation Based on Information of User Profiles, Music Genres and User Ratings
Ja-Hwung Su, Chu-Yu Chin, Hsiao-Chuan Yang, Vincent S. Tseng, Sun-Yuan Hsieh |
ACIIDS (1) | 4 |
| 2018 | Mining Trending High Utility Itemsets from Temporal Transaction Databases
Acquah Hackman, Yu Huang 0018, Vincent S. Tseng |
DEXA (2) | 3 |
| 2018 | Multivariate Time Series Early Classification Using Multi-Domain Deep Neural NetworkabstractEarly classification on multivariate time series is an important research topic in data mining with wide applications to various domains like medical diagnosis, motion detection and financial prediction, etc. Shapelet is probably one of the most commonly used approaches to tackle early classification problem, but one drawback of shaplet is its inefficiency. More importantly, the extracted shapelets may not be applicable to every test case at any time point. This work focuses on early classification of multivariate time series and proposes a novel framework named Multi-Domain Deep Neural Network (MDDNN), in which convolutional neural network (CNN) and long-short term memory (LSTM) are incorporated to learn feature representation and relationship embedding in the long sequences with long time lags. The proposed model can make predictions at any time point of a multivariate time series with the help of a truncation process. We conducted experiments on four real datasets and compared with state-of-the-art algorithms. The experimental results indicate that the proposed method outperforms the alternatives significantly on both of earliness and accuracy. Detailed analysis about the proposed model is also provided in this work. To the best of our knowledge, this is the first work that incorporates deep neural network methods (CNN and LSTM) and multi-domain approach to boost the problem of early classification on multivariate time series. Huai-Shuo Huang, Chien-Liang Liu, Vincent S. Tseng |
DSAA | 3 |
| 2018 | Deep Discriminative Features Learning and Sampling for Imbalanced Data ProblemabstractThe imbalanced data problem occurs in many application domains and is considered to be a challenging problem in machine learning and data mining. Most resampling methods for synthetic data focus on minority class without considering the data distribution of major classes. In contrast to previous works, the proposed method considers both majority classes and minority classes to learn feature embeddings and utilizes appropriate loss functions to make feature embedding as discriminative as possible. The proposed method is a comprehensive framework and different deep learning feature extractors can be utilized for different domains. We conduct experiments utilizing seven numerical datasets and one image dataset based on multiclass classification tasks. The experimental results indicate that the proposed method provides accurate and stable results. Yi-Hsun Liu, Chien-Liang Liu, Vincent S. Tseng |
ICDM | 3 |
| 2018 | FrauDetector+: An Incremental Graph-Mining Approach for Efficient Fraudulent Phone Call DetectionabstractIn recent years, telecommunication fraud has become more rampant internationally with the development of modern technology and global communication. Because of rapid growth in the volume of call logs, the task of fraudulent phone call detection is confronted with big data issues in real-world implementations. Although our previous work, FrauDetector , addressed this problem and achieved some promising results, it can be further enhanced because it focuses only on fraud detection accuracy, whereas the efficiency and scalability are not top priorities. Other known approaches for fraudulent call number detection suffer from long training times or cannot accurately detect fraudulent phone calls in real time. However, the learning process of FrauDetector is too time-consuming to support real-world application. Although we have attempted to accelerate the the learning process of FrauDetector by parallelization, the parallelized learning process, namely PFrauDetector , still cannot afford the computing cost. In this article, we propose a highly efficient incremental graph-mining-based fraudulent phone call detection approach, namely FrauDetector + , which can automatically label fraudulent phone numbers with a “fraud” tag a crucial prerequisite for distinguishing fraudulent phone call numbers from nonfraudulent ones. FrauDetector + initially generates smaller, more manageable subnetworks from original graph and performs a parallelized weighted HITS algorithm for a significant speed increase in the graph learning module. It adopts a novel aggregation approach to generate a trust (or experience) value for each phone number (or user) based on their respective local values. After the initial procedure, we can incrementally update the trust (or experience) value for each phone number (or user) while a new fraud phone number is identified. An efficient fraud-centric hash structure is constructed to support fast real-time detection of fraudulent phone numbers in the detection module. We conduct a comprehensive experimental study based on real datasets collected through an antifraud mobile application called Whoscall . The results demonstrate a significantly improved efficiency of our approach compared with FrauDetector as well as superior performance against other major classifier-based methods. Jia-Ching Ying, Ji Zhang 0001, Che-Wei Huang, Kuan-Ta Chen, Vincent S. Tseng |
ACM Trans. Knowl. Discov. Data | 5 |
| 2017 | Long-Term User Location Prediction Using Deep Learning and Periodic Pattern Mining
Mun Hou Wong, Vincent S. Tseng, Jerry C. C. Tseng, Sun-Wei Liu, Cheng-Hung Tsai |
ADMA | 2 |
| 2017 | Efficient Multi-Destinations Route Planning with Deadlines and Cost ConstraintsabstractIn recent years, multi-destinations route planning has been the topic of much research, which is an emerging branch of the route planning problem. The existing works have been focusing on how to find routes that minimize a single kind of trip cost, such as trip time or distance, amongst others. In fact, users may have multiple requirements in real-life multi-destinations route planning applications, including for personal or business purposes (e.g., express delivery). We observed the fact that (i) there may exist a respective deadline in reaching each of the destinations, (ii) users may consider to reduce further kinds of trip costs, such as fuel, in addition to the deadline constraint. In this paper, we address a novel route planning problem named Multi-Destinations Route Planning with Deadlines and Cost Constraints and propose two approaches, namely BMDC (Basic Multi-Destinations Route Computation) and AMDC (Advanced Multi-Destinations Route Computation) to efficiently plan a route that satisfies deadline requirements and optimizes another criterion such as trip cost. To the best of our knowledge, this is the first work on route planning that considers multiple deadlines for multi-destinations as well as optimizing trip cost, simultaneously. Experimental results demonstrate that our proposed algorithms deliver excellent performance in terms of efficiency and effectiveness. Yu Huang 0018, Bo-Hau Lin, Vincent S. Tseng |
MDM | 3 |
| 2017 | Mining High-Utility Itemsets with Both Positive and Negative Unit Profits from Uncertain Databases
Wensheng Gan, Jerry Chun-Wei Lin, Philippe Fournier-Viger, Han-Chieh Chao, Vincent S. Tseng |
PAKDD (1) | 5 |
| 2017 | A Fast Fourier Transform-Coupled Machine Learning-Based Ensemble Model for Disease Risk Prediction Using a Real-Life Dataset
Raid Lafta, Ji Zhang 0001, Xiaohui Tao 0001, Yan Li 0002, Wessam Abbas, Yonglong Luo, Fulong Chen 0002, Vincent S. Tseng |
PAKDD (1) | 8 |
| 2017 | Effective social content-based collaborative filtering for music recommendationabstractRecently, music recommender systems have been proposed to help users obtain the interested music. Traditional recommender systems making attempts to discover users' musical preferences by ratings always suffer from problems of rating diversity, rating sparsity and lack of ratings. These problems re sult in unsatisfactory recommendation results. To deal with traditional problems, in this paper, we propose a novel music recommender system, namely Multi-modal Music Recommender system (MMR), which integrates social and collaborative information to predict users' preferences. In this work, the playcounts are transformed into collaborative information to cope with problem of lack of rating information, while item tags and artist tags are employed as social information to cope with problems of rating diversity and rating sparsity. Through optimizing the integrated social-and-collaborative information, the users' preferences can be inferred more accurately and efficiently. The experimental results reveal that, three problems can be alleviated significantly and our proposed method outperforms other state-of-the-art recommender systems in terms of RMSE (Root Mean Square Error) and NDCG (Normalized Discount Cumulative Gain). Ja-Hwung Su, Wei-Yi Chang, Vincent S. Tseng |
Intell. Data Anal. | 3 |
| 2017 | Efficiently mining high utility sequential patterns in static and streaming dataabstractHigh utility sequential pattern (HUSP) mining has emerged as a novel topic in data mining. Although some preliminary works have been conducted on this topic, they incur the problem of producing a large search space for high utility sequential patterns. In addition, they mainly focus on mining HUSPs in static databases and do not take streaming data into account, where unbounded data come continuously and often at a high speed. To efficiently deal with both problems, we propose a novel framework for mining high utility sequential patterns over static and streaming databases. In this regard, two efficient data structures named ItemUtilLists (Item Utility Lists) and HUSP-Tree (High Utility Sequential Pattern Tree) are proposed to maintain essential information for mining HUSPs in both offline and online fashions. In addition, a novel utility model called Sequence-Suffix Utility is proposed for effectively pruning the search space in HUSP mining. We propose an algorithm named HUSP-Miner (High Utility Sequential Pattern Miner) to find HUSPs in static databases efficiently. Then, a one-pass algorithm named HUSP-Stream (High Utility Sequential Pattern mining over Data Streams) is proposed to incrementally update ItemUtilLists and HUSP-Tree online and find HUSPs over data streams. To the best of our knowledge, HUSP-Stream is the first method to find HUSPs over data streams. Experimental results on both real and synthetic datasets show that HUSP-Miner outperforms the compared algorithms substantially in terms of execution time, memory usage and number of generated candidates. The experiments also demonstrate impressive performance of HUSP-Stream to update the data structures and discover HUSPs over data streams. Morteza Zihayat, Cheng-Wei Wu, Aijun An, Vincent S. Tseng, Chien Lin |
Intell. Data Anal. | 4 |
| 2017 | EFIM: a fast and memory efficient algorithm for high-utility itemset mining
Souleymane Zida, Philippe Fournier-Viger, Jerry Chun-Wei Lin, Cheng-Wei Wu, Vincent S. Tseng |
Knowl. Inf. Syst. | 5 |
| 2017 | Efficiently mining uncertain high-utility itemsets
Jerry Chun-Wei Lin, Wensheng Gan, Philippe Fournier-Viger, Tzung-Pei Hong, Vincent S. Tseng |
Soft Comput. | 5 |
| 2017 | Mining Sequential Risk Patterns From Large-Scale Clinical Databases for Early Assessment of Chronic Diseases: A Case Study on Chronic Obstructive Pulmonary DiseaseabstractChronic diseases have been among the major concerns in medical fields since they may cause a heavy burden on healthcare resources and disturb the quality of life. In this paper, we propose a novel framework for early assessment on chronic diseases by mining sequential risk patterns with time interval information from diagnostic clinical records using sequential rules mining, and classification modeling techniques. With a complete workflow, the proposed framework consists of four phases namely data preprocessing, risk pattern mining, classification modeling, and post analysis. For empiricasl evaluation, we demonstrate the effectiveness of our proposed framework with a case study on early assessment of COPD. Through experimental evaluation on a large-scale nationwide clinical database in Taiwan, our approach can not only derive rich sequential risk patterns but also extract novel patterns with valuable insights for further medical investigation such as discovering novel markers and better treatments. To the best of our knowledge, this is the first work addressing the issue of mining sequential risk patterns with time-intervals as well as classification models for early assessment of chronic diseases. Yi-Ting Cheng, Yu-Feng Lin, Kuo-Hwa Chiang, Vincent S. Tseng |
IEEE J. Biomed. Health Informatics | 4 |
| 2017 | An Efficient Framework for Multirequest Route Planning in Urban EnvironmentsabstractIn recent years, research on location-based services has received a lot of interest, in both industry and academia, due to a wide range of potential applications. Among them, one of the active topic areas is the constraint-based route planning on a point-of-interest (POI) network. Most of the previous studies on this topic primarily consider the geographic properties of the POIs in planning a route. However, we consider that the reason that a user visits a POI is that it provides some services that the user needs. In particular, in urban environments, a POI may provide various kinds of services. Hence, the user's requests should be considered. In this paper, we address a novel problem, which is called multirequest route planning, and propose a novel framework to efficiently plan a route for serving multiple user-specified requests. The framework consists of two major modules: planning module, in which four approaches with pruning and caching strategies are proposed for planning a preliminary route, and refinement module, in which two refinement mechanisms are proposed for further enhancing the quality of the route. To our best knowledge, this is the first work on route planning that considers multiple services provided by a POI and multiple requests specified by a user, simultaneously. Finally, we perform an extensive experimental evaluation based on three real-world POI data sets and deliver excellent performance. Eric Hsueh-Chan Lu, Huan-Sheng Chen, Vincent S. Tseng |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2016 | IRS-HD: An Intelligent Personalized Recommender System for Heart Disease Patients in a Tele-Health Environment
Raid Lafta, Ji Zhang 0001, Xiaohui Tao 0001, Yan Li 0002, Vincent S. Tseng |
ADMA | 5 |
| 2016 | A scalable complex event analytical system with incremental episode mining over data streamsabstractEpisode pattern mining is a very powerful technique to get high-valued information for people to solve real-life cross-disciplinary problems, such as for the analysis of manufacturing, stock markets, weather records and so on. As data grows, the mining process must be re-triggered again and again to obtain the most updated information. However, periodically re-mining the full dataset is not cost-effective, and thus a number of incremental mining approaches arise for the growing data. However, to our best knowledge, there exist few studies targeted on the problem of incremental episode mining. Moreover, streaming data of complex events is more and more popular because digital sensors always collect data around us in this big data age. Now the challenge is not only mining valuable episode patterns of incremental dataset, but also mining episode patterns over data streams of complex events. To address this research problem, we adopt the Lambda Architecture to design a scalable complex event analytical system that could be used to facilitate the incremental episode mining process over complex event sequences of data streams. Apache Spark and Apache Spark Streaming are applied as the development framework of the batch layer and the speed layer, respectively. To take both the efficiency and accuracy into consideration, we develop a series of modules and three algorithms, namely, batch episode mining, delta episode mining and pattern merging. Results from the experimental validation on a real dataset show that the proposed system carries high scalability and delivers excellent performance in terms of efficiency and accuracy. Jerry C. C. Tseng, Jia-Yuan Gu, Ping-Feng Wang, Ching-Yu Chen, Chu-Feng Li, Vincent S. Tseng |
CEC | 6 |
| 2016 | Mining Minimal High-Utility Itemsets
Philippe Fournier-Viger, Jerry Chun-Wei Lin, Cheng-Wei Wu, Vincent S. Tseng, Usef Faghihi |
DEXA (1) | 4 |
| 2016 | PFrauDetector: A Parallelized Graph Mining Approach for Efficient Fraudulent Phone Call DetectionabstractIn recent years, fraud is becoming more rampant internationally with the development of modern technology and global communication. Due to the rapid growth in the volume of call logs, the task of fraudulent phone call detection is confronted with Big Data issues in real-world implementations. While our previous work, FrauDetector, has addressed this problem and achieved some promising results, it can be further enhanced as it focuses on the fraud detection accuracy while the efficiency and scalability are not on the top priority. Meanwhile, other known approaches suffer from long training time and/or cannot accurately detect fraudulent phone calls in real time. In this paper, we propose a highly-efficient parallelized graph-mining-based fraudulent phone call detection framework, namely PFrauDetector, which is able to automatically label fraudulent phone numbers with a "fraud" tag, a crucial prerequisite for distinguishing fraudulent phone call numbers from the normal ones. PFrauDetector generates smaller, more manageable sub-networks from the original graph and performs a parallelized weighted HITS algorithm for significant speed acceleration in the graph learning module. It adopts a novel aggregation approach to generate the trust (or experience) value for each phone number (or user) based on their respective local values. We conduct a comprehensive experimental study based on a real dataset collected through an anti-fraud mobile application, Whoscall. The results demonstrate a significantly improved efficiency of our approach compared to FrauDetector and superior performance against other major classifier-based methods. Jia-Ching Ying, Ji Zhang 0001, Che-Wei Huang, Kuan-Ta Chen, Vincent S. Tseng |
ICPADS | 5 |
| 2016 | Efficient Mining of Uncertain Data for High-Utility Itemsets
Jerry Chun-Wei Lin, Wensheng Gan, Philippe Fournier-Viger, Tzung-Pei Hong, Vincent S. Tseng |
WAIM (1) | 5 |
| 2016 | Fast algorithms for mining high-utility itemsets with various discount strategies
Jerry Chun-Wei Lin, Wensheng Gan, Philippe Fournier-Viger, Tzung-Pei Hong, Vincent S. Tseng |
Adv. Eng. Informatics | 5 |
| 2016 | Weighted frequent itemset mining over uncertain databases
Jerry Chun-Wei Lin, Wensheng Gan, Philippe Fournier-Viger, Tzung-Pei Hong, Vincent S. Tseng |
Appl. Intell. | 5 |
| 2016 | Integrating tourist packages and tourist attractions for personalized trip planning based on travel constraints
Eric Hsueh-Chan Lu, Shih Hsin Fang, Vincent S. Tseng |
GeoInformatica | 3 |
| 2016 | Efficient algorithms for mining high-utility itemsets in uncertain databases
Jerry Chun-Wei Lin, Wensheng Gan, Philippe Fournier-Viger, Tzung-Pei Hong, Vincent S. Tseng |
Knowl. Based Syst. | 5 |
| 2016 | Efficient Algorithms for Mining Top-K High Utility ItemsetsabstractHigh utility itemsets (HUIs) mining is an emerging topic in data mining, which refers to discovering all itemsets having a utility meeting a user-specified minimum utility threshold min_util. However, setting min_util appropriately is a difficult problem for users. Generally speaking, finding an appropriate minimum utility threshold by trial and error is a tedious process for users. If min_util is set too low, too many HUIs will be generated, which may cause the mining process to be very inefficient. On the other hand, if min_util is set too high, it is likely that no HUIs will be found. In this paper, we address the above issues by proposing a new framework for top-k high utility itemset mining, where k is the desired number of HUIs to be mined. Two types of efficient algorithms named TKU (mining Top-K Utility itemsets) and TKO (mining Top-K utility itemsets in One phase) are proposed for mining such itemsets without the need to set min_util. We provide a structural comparison of the two algorithms with discussions on their advantages and limitations. Empirical evaluations on both real and synthetic datasets show that the performance of the proposed algorithms is close to that of the optimal case of state-of-the-art utility mining algorithms. Vincent S. Tseng, Cheng-Wei Wu, Philippe Fournier-Viger, Philip S. Yu |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2016 | An intelligent recommender system based on predictive analysis in telehealthcare environmentabstractThe use of intelligent technologies for providing useful recommendations to patients suffering chronic diseases may play a positive role in improving the general life quality of patients and help reduce the workload and cost involved in their daily healthcare. The objective of this study is to deve lop an intelligent recommender system based on predictive analysis for advising patients in the telehealth environment concerning whether they need to take the body test one day in advance by analyzing medical measurements of a patient for the past k days. The proposed algorithms supporting the recommender system have been validated using a time series telehealth data recorded from heart disease patients which were collected from May to January 2012, from our industry collaborator Tunstall. The experimental results show that the proposed system yields satisfactory recommendation accuracy and offer a promising way for saving the workload for patients to conduct body tests every day. This study highlights the possible usefulness of the computerized analysis of time series telehealth data in providing appropriate recommendations to patients suffering chronic diseases such as heart diseases patients. Raid Lafta, Ji Zhang 0001, Xiaohui Tao 0001, Yan Li 0002, Vincent S. Tseng, Yonglong Luo, Fulong Chen 0002 |
Web Intell. | 5 |
| 2015 | Mining high-utility itemsets with various discount strategiesabstractIn recent years, mining high-utility itemsets (HUIs) has become as a key topic in data mining. However, most of the developed algorithms assume the unrealistic situations that unit profits of items remain unchanged over time. But in real-life situations, the profit of an item or itemset varies as a function of cost prices, sales prices and sales strategies. In this paper, a novel framework for mining HUIs with two algorithms under various Discount strategies (HUID) are introduced. HUID-tp is based on various discount strategies and a novel downward closure property to mine the complete set of HUIs. HUID-Miner is an algorithm relying on a compact data structure (Positive-and-Negative Utility-list, PNU-list) and new pruning strategies to efficiently discover HUIs without candidate generation, while considerably reducing the size of the search space. Furthermore, a strategy named Estimated Utility Co-occurrence Strategy which stores the relationships between 2-itemsets is also adopted in the proposed improvement HUID-EMiner algorithm to speed up computation. An extensive experimental study carried on several real-life datasets shows the performance of the proposed algorithms. Jerry Chun-Wei Lin, Wensheng Gan, Philippe Fournier-Viger, Tzung-Pei Hong, Vincent S. Tseng |
DSAA | 5 |
| 2015 | FrauDetector: A Graph-Mining-based Framework for Fraudulent Phone Call DetectionabstractIn recent years, fraud is increasing rapidly with the development of modern technology and global communication. Although many literatures have addressed the fraud detection problem, these existing works focus only on formulating the fraud detection problem as a binary classification problem. Due to limitation of information provided by telecommunication records, such classifier-based approaches for fraudulent phone call detection normally do not work well. In this paper, we develop a graph-mining-based fraudulent phone call detection framework for a mobile application to automatically annotate fraudulent phone numbers with a "fraud" tag, which is a crucial prerequisite for distinguishing fraudulent phone calls from normal phone calls. Our detection approach performs a weighted HITS algorithm to learn the trust value of a remote phone number. Based on telecommunication records, we build two kinds of directed bipartite graph: i) CPG and ii) UPG to represent telecommunication behavior of users. To weight the edges of CPG and UPG, we extract features for each pair of user and remote phone number in two different yet complementary aspects: 1) duration relatedness (DR) between user and phone number; and 2) frequency relatedness (FR) between user and phone number. Upon weighted CPG and UPG, we determine a trust value for each remote phone number. Finally, we conduct a comprehensive experimental study based on a dataset collected through an anti-fraud mobile application, Whoscall. The results demonstrate the effectiveness of our weighted HITS-based approach and show the strength of taking both DR and FR into account in feature extraction. Vincent S. Tseng, Jia-Ching Ying, Che-Wei Huang, Yimin Kao, Kuan-Ta Chen |
KDD | 1 |
| 2015 | CPT+: Decreasing the Time/Space Complexity of the Compact Prediction Tree
Ted Gueniche, Philippe Fournier-Viger, Rajeev Raman, Vincent S. Tseng |
PAKDD (2) | 4 |
| 2015 | Reliable Early Classification on Multivariate Time Series with Numerical and Categorical Attributes
Yu-Feng Lin, Hsuan-Hsu Chen, Vincent S. Tseng, Jian Pei 0001 |
PAKDD (1) | 3 |
| 2015 | Mining High Utility Itemsets in Big Data
Ying Chun Lin, Cheng-Wei Wu, Vincent S. Tseng |
PAKDD (2) | 3 |
| 2015 | Efficient algorithms for mining up-to-date high-utility patterns
Jerry Chun-Wei Lin, Wensheng Gan, Tzung-Pei Hong, Vincent S. Tseng |
Adv. Eng. Informatics | 4 |
| 2015 | Discovering utility-based episode rules in complex event sequences
Yu-Feng Lin, Cheng-Wei Wu, Chien-Feng Huang, Vincent S. Tseng |
Expert Syst. Appl. | 4 |
| 2015 | Mining Partially-Ordered Sequential Rules Common to Multiple SequencesabstractSequential rule mining is an important data mining problem with multiple applications. An important limitation of algorithms for mining sequential rules common to multiple sequences is that rules are very specific and therefore many similar rules may represent the same situation. This can cause three major problems: (1) similar rules can be rated quite differently, (2) rules may not be found because they are individually considered uninteresting, and (3) rules that are too specific are less likely to be used for making predictions. To address these issues, we explore the idea of mining “partially-ordered sequential rules” (POSR), a more general form of sequential rules such that items in the antecedent and the consequent of each rule are unordered. To mine POSR, we propose the RuleGrowth algorithm, which is efficient and easily extendable. In particular, we present an extension (TRuleGrowth) that accepts a sliding-window constraint to find rules occurring within a maximum amount of time. A performance study with four real-life datasets show that RuleGrowth and TRuleGrowth have excellent performance and scalability compared to baseline algorithms and that the number of rules discovered can be several orders of magnitude smaller when the sliding-window constraint is applied. Furthermore, we also report results from a real application showing that POSR can provide a much higher prediction accuracy than regular sequential rules for sequence prediction. Philippe Fournier-Viger, Cheng-Wei Wu, Vincent S. Tseng, Longbing Cao, Roger Nkambou |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2015 | Efficient Algorithms for Mining the Concise and Lossless Representation of High Utility ItemsetsabstractMining high utility itemsets (HUIs) from databases is an important data mining task, which refers to the discovery of itemsets with high utilities (e.g. high profits). However, it may present too many HUIs to users, which also degrades the efficiency of the mining process. To achieve high efficiency for the mining task and provide a concise mining result to users, we propose a novel framework in this paper for mining closed+high utility itemsets(CHUIs), which serves as a compact and lossless representation of HUIs. We propose three efficient algorithms named AprioriCH (Apriori-based algorithm for mining High utility Closed+itemsets), AprioriHC-D (AprioriHC algorithm with Discarding unpromising and isolated items) and CHUD (Closed+High Utility Itemset Discovery) to find this representation. Further, a method called DAHU (Derive All High Utility Itemsets) is proposed to recover all HUIs from the set of CHUIs without accessing the original database. Results on real and synthetic datasets show that the proposed algorithms are very efficient and that our approaches achieve a massive reduction in the number of HUIs. In addition, when all HUIs can be recovered by DAHU, the combination of CHUD and DAHU outperforms the state-of-the-art algorithms for mining HUIs. Vincent S. Tseng, Cheng-Wei Wu, Philippe Fournier-Viger, Philip S. Yu |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2015 | Intelligent interaction, reasoning, and applications
Chao-Lin Liu, Mitsunori Matsushita, Yasufumi Takama, Min-Yuh Day, Vincent S. Tseng |
Web Intell. | 5 |
| 2014 | Novel Concise Representations of High Utility Itemsets Using Generator Patterns
Philippe Fournier-Viger, Cheng-Wei Wu, Vincent S. Tseng |
ADMA | 3 |
| 2014 | Location semantics prediction for living analytics by mining smartphone dataabstractAutomatic location semantics prediction for living analytics based on smartphone data has attracted extensive attention in just recent years. Basically, this task can be formulated as a multi-class classification problem, where different location/places are regarded as different labels. Previous studies were mostly based on common classification techniques directly, neglecting the critical challenging issue of class imbalance in such a problem (e.g., people go to offices much more often than they go to cinemas). It is also noteworthy that in contrast to common multi-class problems where the classes can be treated independently and interchangeably, the places for labeling usually have important correlations, which should be taken account in the classification/labeling process. Moreover, several activities may occur in the same place and thus the same place label might convey different semantics. In this paper, we address the above issues for location semantics prediction by proposing the FS-Mining (Frame-based Semantics Mining) approach. We treat the raw sensor data in the smartphone as a sequence of short and non-overlapping frames, based on which the user behavior at each place can be characterized and the place semantics can be modeled. To deal with the issues of label relation and class imbalance, a multi-level classification model with class-split and class-merge mechanisms was also developed. An ensemble strategy was also employed to further improve the performance. Experiments on the dataset of Nokia Mobile Data Challenge [1] demonstrate promising performances for the FS-Mining approach. Chi-Min Huang, Jia-Ching Ying, Vincent S. Tseng, Zhi-Hua Zhou |
DSAA | 3 |
| 2014 | ERMiner: Sequential Rule Mining Using Equivalence Classes
Philippe Fournier-Viger, Ted Gueniche, Souleymane Zida, Vincent S. Tseng |
IDA | 4 |
| 2014 | FHM: Faster High-Utility Itemset Mining Using Estimated Utility Co-occurrence Pruning
Philippe Fournier-Viger, Cheng-Wei Wu, Souleymane Zida, Vincent S. Tseng |
ISMIS | 4 |
| 2014 | WBPL: An Open-Source Library for Predicting Web Surfing Behaviors
Ted Gueniche, Philippe Fournier-Viger, Roger Nkambou, Vincent S. Tseng |
ISMIS | 4 |
| 2014 | Trip Recommendation with Multiple User Constraints by Integrating Point-of-Interests and Travel PackagesabstractWith the advances of mobile communication techniques in recent years, numerous kinds of Location-Based Services (LBSs) have been developed and one popular application of LBSs is trip recommendation. Although there exist already a number of studies on this topic in literatures, most of them focused on combining a set of point-of-interests (POIs, or say attractions) as a trip based on user-specific constraints. In another way, some few works discussed making recommendation in terms of travel packages, which have the benefits of lower cost and higher convenience. However, no prior work explores to integrate attractions and travel packages simultaneously for trip recommendation. In fact, such a hybrid-style recommender can provide higher benefits for users although there exist critical challenges here like the efficiency issue in such kind of real-time applications. In this paper, we propose a novel framework named Package-Attraction-based Trip Recommender (PATR) to efficiently recommend the personalized trips satisfying multiple constraints by effectively combining attractions and packages. In PATR, a Score Inference Model is proposed to infer the scores of attractions and packages by taking user-based preference and temporal-based properties into account. Then, the Hybrid Trip-Mine algorithm is proposed to efficiently discover the optimal trip which satisfies the multiple user-specific constraints with both of attractions and packages considered simultaneously. Furthermore, we propose two pruning strategies based on Hybrid Trip-Mine, named Score Estimation (SE) and Score Bound Tightening (SBT), to further improve the execution efficiency and memory utilization. To the best of our knowledge, this is the first work on travel recommendation that considers attractions and packages simultaneously. Through extensive experimental evaluations, our proposed approaches were shown to deliver excellent performance. Shih Hsin Fang, Eric Hsueh-Chan Lu, Vincent S. Tseng |
MDM (1) | 3 |
| 2014 | Network-based analysis identifies epigenetic biomarkers of esophageal squamous cell carcinoma progressionabstractMOTIVATION: A rapid progression of esophageal squamous cell carcinoma (ESCC) causes a high mortality rate because of the propensity for metastasis driven by genetic and epigenetic alterations. The identification of prognostic biomarkers would help prevent or control metastatic progression. Expression analyses have been used to find such markers, but do not always validate in separate cohorts. Epigenetic marks, such as DNA methylation, are a potential source of more reliable and stable biomarkers. Importantly, the integration of both expression and epigenetic alterations is more likely to identify relevant biomarkers. RESULTS: We present a new analysis framework, using ESCC progression-associated gene regulatory network (GRN escc), to identify differentially methylated CpG sites prognostic of ESCC progression. From the CpG loci differentially methylated in 50 tumor-normal pairs, we selected 44 CpG loci most highly associated with survival and located in the promoters of genes more likely to belong to GRN escc. Using an independent ESCC cohort, we confirmed that 8/10 of CpG loci in the promoter of GRN escc genes significantly correlated with patient survival. In contrast, 0/10 CpG loci in the promoter genes outside the GRN escc were correlated with patient survival. We further characterized the GRN escc network topology and observed that the genes with methylated CpG loci associated with survival deviated from the center of mass and were less likely to be hubs in the GRN escc. We postulate that our analysis framework improves the identification of bona fide prognostic biomarkers from DNA methylation studies, especially with partial genome coverage. Chun-Pei Cheng, I-Ying Kuo, Hakan Alakus, Kelly A. Frazer, Olivier Harismendy, Yi-Ching Wang, Vincent S. Tseng |
Bioinform. | 7 |
| 2014 | MiningABs: mining associated biomarkers across multi-connected gene expression datasetsabstractBACKGROUND: Human disease often arises as a consequence of alterations in a set of associated genes rather than alterations to a set of unassociated individual genes. Most previous microarray-based meta-analyses identified disease-associated genes or biomarkers independent of genetic interactions. Therefore, in this study, we present the first meta-analysis method capable of taking gene combination effects into account to efficiently identify associated biomarkers (ABs) across different microarray platforms. RESULTS: We propose a new meta-analysis approach called MiningABs to mine ABs across different array-based datasets. The similarity between paired probe sequences is quantified as a bridge to connect these datasets together. The ABs can be subsequently identified from an "improved" common logit model (c-LM) by combining several sibling-like LMs in a heuristic genetic algorithm selection process. Our approach is evaluated with two sets of gene expression datasets: i) 4 esophageal squamous cell carcinoma and ii) 3 hepatocellular carcinoma datasets. Based on an unbiased reciprocal test, we demonstrate that each gene in a group of ABs is required to maintain high cancer sample classification accuracy, and we observe that ABs are not limited to genes common to all platforms. Investigating the ABs using Gene Ontology (GO) enrichment, literature survey, and network analyses indicated that our ABs are not only strongly related to cancer development but also highly connected in a diverse network of biological interactions. CONCLUSIONS: The proposed meta-analysis method called MiningABs is able to efficiently identify ABs from different independently performed array-based datasets, and we show its validity in cancer biology via GO enrichment, literature survey and network analyses. We postulate that the ABs may facilitate novel target and drug discovery, leading to improved clinical treatment. Java source code, tutorial, example and related materials are available at "http://sourceforge.net/projects/miningabs/". Chun-Pei Cheng, Christopher M. DeBoever, Kelly A. Frazer, Vincent S. Tseng |
BMC Bioinform. | 5 |
| 2014 | On-shelf utility mining with negative item values
Guo-Cheng Lan, Tzung-Pei Hong, Jen-Peng Huang, Vincent S. Tseng |
Expert Syst. Appl. | 4 |
| 2014 | Applying the maximum utility measure in high utility sequential pattern mining
Guo-Cheng Lan, Tzung-Pei Hong, Vincent S. Tseng, Shyue-Liang Wang |
Expert Syst. Appl. | 3 |
| 2014 | Semantic trajectory-based high utility item recommendation system
Jia-Ching Ying, Huan-Sheng Chen, Kawuu Weicheng Lin, Eric Hsueh-Chan Lu, Vincent S. Tseng, Huan-Wen Tsai, Kuang Hung Cheng, Shun-Chieh Lin |
Expert Syst. Appl. | 5 |
| 2014 | SPMF: a Java open-source pattern mining library
Philippe Fournier-Viger, Antonio Gomariz, Ted Gueniche, Azadeh Soltani, Cheng-Wei Wu, Vincent S. Tseng |
J. Mach. Learn. Res. | 6 |
| 2014 | An efficient projection-based indexing approach for mining high utility itemsets
Guo-Cheng Lan, Tzung-Pei Hong, Vincent S. Tseng |
Knowl. Inf. Syst. | 3 |
| 2014 | Mining User Check-In Behavior with a Random Walk for Urban Point-of-Interest RecommendationsabstractIn recent years, research into the mining of user check-in behavior for point-of-interest (POI) recommendations has attracted a lot of attention. Existing studies on this topic mainly treat such recommendations in a traditional manner—that is, they treat POIs as items and check-ins as ratings. However, users usually visit a place for reasons other than to simply say that they have visited. In this article, we propose an approach referred to as Urban POI-Walk (UPOI-Walk), which takes into account a user's social-triggered intentions (SI), preference-triggered intentions (PreI), and popularity-triggered intentions (PopI), to estimate the probability of a user checking-in to a POI. The core idea of UPOI-Walk involves building a HITS-based random walk on the normalized check-in network, thus supporting the prediction of POI properties related to each user's preferences. To achieve this goal, we define several user--POI graphs to capture the key properties of the check-in behavior motivated by user intentions. In our UPOI-Walk approach, we propose a new kind of random walk model—Dynamic HITS-based Random Walk—which comprehensively considers the relevance between POIs and users from different aspects. On the basis of similitude, we make an online recommendation as to the POI the user intends to visit. To the best of our knowledge, this is the first work on urban POI recommendations that considers user check-in behavior motivated by SI, PreI, and PopI in location-based social network data. Through comprehensive experimental evaluations on two real datasets, the proposed UPOI-Walk is shown to deliver excellent performance. Jia-Ching Ying, Wen-Ning Kuo, Vincent S. Tseng, Eric Hsueh-Chan Lu |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2013 | Mining Maximal Sequential Patterns without Candidate Maintenance
Philippe Fournier-Viger, Cheng-Wei Wu, Vincent S. Tseng |
ADMA (1) | 3 |
| 2013 | Compact Prediction Tree: A Lossless Model for Accurate Sequence Prediction
Ted Gueniche, Philippe Fournier-Viger, Vincent S. Tseng |
ADMA (2) | 3 |
| 2013 | Mining high utility episodes in complex event sequencesabstractFrequent episode mining (FEM) is an interesting research topic in data mining with wide range of applications. However, the traditional framework of FEM treats all events as having the same importance/utility and assumes that a same type of event appears at most once at any time point. These simplifying assumptions do not reflect the characteristics of scenarios in real applications and thus the useful information of episodes in terms of utilities such as profits is lost. Furthermore, most studies on FEM focused on mining episodes in simple event sequences and few considered the scenario of complex event sequences, where different events can occur simultaneously. To address these issues, in this paper, we incorporate the concept of utility into episode mining and address a new problem of mining high utility episodes from complex event sequences, which has not been explored so far. In the proposed framework, the importance/utility of different events is considered and multiple events can appear simultaneously. Several novel features are incorporated into the proposed framework to resolve the challenges raised by this new problem, such as the absence of anti-monotone property and the huge set of candidate episodes. Moreover, an efficient algorithm named UP-Span (Utility ePisodes mining by Spanning prefixes) is proposed for mining high utility episodes with several strategies incorporated for pruning the search space to achieve high efficiency. Experimental results on real and synthetic datasets show that UP-Span has excellent performance and serves as an effective solution to the new problem of mining high utility episodes from complex event sequences. Cheng-Wei Wu, Yu-Feng Lin, Philip S. Yu, Vincent S. Tseng |
KDD | 4 |
| 2013 | Personalized Music Recommendation by Mining Social Media TagsabstractOver the past few years, the recommender system has been proposed as a critical role to help users choose the preferred product from a massive amount of data. For music recommendation, most recent recommender systems made attempts to associate music with the user's preferences primarily based on user ratings. However, this kind of recommendation mechanism encounters the problem called rating diversity that makes the prediction results unreliable. To cope with this problem, in this paper, we propose a novel music recommendation approach that utilizes social media tags instead of ratings to calculate the similarity between music pieces. Through the proposed tag-based similarity, the user preferences hidden in tags can be inferred effectively. The empirical evaluations on real social media datasets reveal that our proposed approach using social tags outperforms the existing ones using only ratings in terms of predicting the user's preferences to music. Ja-Hwung Su, Wei-Yi Chang, Vincent S. Tseng |
KES | 3 |
| 2013 | Efficient Approaches for Multi-requests Route Planning in Urban AreasabstractIn recent years, with the rapid developments of wireless technologies, researches on Location-Based Services (LBSs) have attracted extensive attentions and one active topic among them is constraint-based route planning on a Point-Of-Interest (POI) network. Although a number of studies on this topic have been proposed in literatures, most of them primarily consider the geographic properties of the POIs in planning a route. In fact, the motivation of a user to visit a POI is frequently due to that the POI can provide some services meeting the user's needs. Hence, user requests should be considered in route planning, especially in an urban area where a POI may provide various kinds of services. Besides, the efficiency of route planning is critical in such kind of real-time LBS applications. In this paper, we address a novel route planning problem named Multi-Requests Route Planning (MRRP) and propose four approaches, namely kNN-MS, kMD-MS, EMB and kRA-MS to efficiently plan a time-saving route based on the user-specific requests. Furthermore, we propose two refinement mechanisms, three pruning strategies and two caching techniques to further enhance the route quality and planning efficiency for MRRP, respectively. To the best of our knowledge, this is the first work on route planning that considers multiple services provided by a POI and multiple requests specified by a user, simultaneously. Through extensive experimental evaluations, our approaches were shown to deliver excellent performance. Eric Hsueh-Chan Lu, Huan-Sheng Chen, Vincent S. Tseng |
MDM (1) | 3 |
| 2013 | TripCloud: An Intelligent Cloud-Based Trip Recommendation System
Jia-Ching Ying, Eric Hsueh-Chan Lu, Bo-Nian Shi, Vincent S. Tseng |
SSTD | 4 |
| 2013 | Mining interesting user behavior patterns in mobile commerce environments
Bai-En Shie, Philip S. Yu, Vincent S. Tseng |
Appl. Intell. | 3 |
| 2013 | An efficient method for mining cross-timepoint gene regulation sequential patterns from time course gene expression datasetsabstractBACKGROUND: Observation of gene expression changes implying gene regulations using a repetitive experiment in time course has become more and more important. However, there is no effective method which can handle such kind of data. For instance, in a clinical/biological progression like inflammatory response or cancer formation, a great number of differentially expressed genes at different time points could be identified through a large-scale microarray approach. For each repetitive experiment with different samples, converting the microarray datasets into transactional databases with significant singleton genes at each time point would allow sequential patterns implying gene regulations to be identified. Although traditional sequential pattern mining methods have been successfully proposed and widely used in different interesting topics, like mining customer purchasing sequences from a transactional database, to our knowledge, the methods are not suitable for such biological dataset because every transaction in the converted database may contain too many items/genes. RESULTS: In this paper, we propose a new algorithm called CTGR-Span (Cross-Timepoint Gene Regulation Sequential pattern) to efficiently mine CTGR-SPs (Cross-Timepoint Gene Regulation Sequential Patterns) even on larger datasets where traditional algorithms are infeasible. The CTGR-Span includes several biologically designed parameters based on the characteristics of gene regulation. We perform an optimal parameter tuning process using a GO enrichment analysis to yield CTGR-SPs more meaningful biologically. The proposed method was evaluated with two publicly available human time course microarray datasets and it was shown that it outperformed the traditional methods in terms of execution efficiency. After evaluating with previous literature, the resulting patterns also strongly correlated with the experimental backgrounds of the datasets used in this study. CONCLUSIONS: We propose an efficient CTGR-Span to mine several biologically meaningful CTGR-SPs. We postulate that the biologist can benefit from our new algorithm since the patterns implying gene regulations could provide further insights into the mechanisms of novel gene regulations during a biological or clinical progression. The Java source code, program tutorial and other related materials used in this program are available at http://websystem.csie.ncku.edu.tw/CTGR-Span.rar. Chun-Pei Cheng, Yi-Lin Tsai, Vincent S. Tseng |
BMC Bioinform. | 4 |
| 2013 | Mining differential top-k co-expression patterns from time course comparative gene expression datasetsabstractBACKGROUND: Frequent pattern mining analysis applied on microarray dataset appears to be a promising strategy for identifying relationships between gene expression levels. Unfortunately, too many itemsets (co-expressed genes) are identified by this analysis method since it does not consider the importance of each gene within biological processes to a cellular response and does not take into account temporal properties under biological treatment-control matched conditions in a microarray dataset. RESULTS: We propose a method termed TIIM (Top-k Impactful Itemsets Miner), which only requires specifying a user-defined number k to explore the top k itemsets with the most significantly differentially co-expressed genes between 2 conditions in a time course. To give genes different weights, a table with impact degrees for each gene was constructed based on the number of neighboring genes that are differently expressed in the dataset within gene regulatory networks. Finally, the resulting top-k impactful itemsets were manually evaluated using previous literature and analyzed by a Gene Ontology enrichment method. CONCLUSIONS: In this study, the proposed method was evaluated in 2 publicly available time course microarray datasets with 2 different experimental conditions. Both datasets identified potential itemsets with co-expressed genes evaluated from the literature and showed higher accuracies compared to the 2 corresponding control methods: i) performing TIIM without considering the gene expression differentiation between 2 different experimental conditions and impact degrees, and ii) performing TIIM with a constant impact degree for each gene. Our proposed method found that several new gene regulations involved in these itemsets were useful for biologists and provided further insights into the mechanisms underpinning biological processes. The Java source code and other related materials used in this study are available at "http://websystem.csie.ncku.edu.tw/TIIM_Program.rar". Chun-Pei Cheng, Vincent S. Tseng |
BMC Bioinform. | 3 |
| 2013 | A hybrid scheme for energy-efficient object tracking in sensor networks
Ming-Hua Hsieh, Kawuu Weicheng Lin, Vincent S. Tseng |
Knowl. Inf. Syst. | 3 |
| 2013 | Efficient algorithms for discovering high utility user behavior patterns in mobile commerce environments
Bai-En Shie, Hui-Fang Hsiao, Vincent S. Tseng |
Knowl. Inf. Syst. | 3 |
| 2013 | Preference-oriented mining techniques for location-based store search
Jess Soo-Fong Tan, Eric Hsueh-Chan Lu, Vincent S. Tseng |
Knowl. Inf. Syst. | 3 |
| 2013 | Time series pattern discovery by a PIP-based evolutionary approach
Chun-Hao Chen, Vincent S. Tseng, Hsieh-Hui Yu, Tzung-Pei Hong |
Soft Comput. | 2 |
| 2013 | Mining geographic-temporal-semantic patterns in trajectories for location predictionabstractIn recent years, research on location predictions by mining trajectories of users has attracted a lot of attention. Existing studies on this topic mostly treat such predictions as just a type of location recommendation, that is, they predict the next location of a user using location recommenders. However, an user usually visits somewhere for reasons other than interestingness. In this article, we propose a novel mining-based location prediction approach called Geographic-Temporal-Semantic-based Location Prediction (GTS-LP), which takes into account a user's geographic-triggered intentions, temporal-triggered intentions, and semantic-triggered intentions, to estimate the probability of the user in visiting a location. The core idea underlying our proposal is the discovery of trajectory patterns of users, namely GTS patterns , to capture frequent movements triggered by the three kinds of intentions. To achieve this goal, we define a new trajectory pattern to capture the key properties of the behaviors that are motivated by the three kinds of intentions from trajectories of users. In our GTS-LP approach, we propose a series of novel matching strategies to calculate the similarity between the current movement of a user and discovered GTS patterns based on various moving intentions. On the basis of similitude, we make an online prediction as to the location the user intends to visit. To the best of our knowledge, this is the first work on location prediction based on trajectory pattern mining that explores the geographic, temporal, and semantic properties simultaneously. By means of a comprehensive evaluation using various real trajectory datasets, we show that our proposed GTS-LP approach delivers excellent performance and significantly outperforms existing state-of-the-art location prediction methods. Jia-Ching Ying, Wang-Chien Lee, Vincent S. Tseng |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2013 | Efficient Algorithms for Mining High Utility Itemsets from Transactional DatabasesabstractMining high utility itemsets from a transactional database refers to the discovery of itemsets with high utility like profits. Although a number of relevant algorithms have been proposed in recent years, they incur the problem of producing a large number of candidate itemsets for high utility itemsets. Such a large number of candidate itemsets degrades the mining performance in terms of execution time and space requirement. The situation may become worse when the database contains lots of long transactions or long high utility itemsets. In this paper, we propose two algorithms, namely utility pattern growth (UP-Growth) and UP-Growth+, for mining high utility itemsets with a set of effective strategies for pruning candidate itemsets. The information of high utility itemsets is maintained in a tree-based data structure named utility pattern tree (UP-Tree) such that candidate itemsets can be generated efficiently with only two scans of database. The performance of UP-Growth and UP-Growth+ is compared with the state-of-the-art algorithms on many types of both real and synthetic data sets. Experimental results show that the proposed algorithms, especially UP-Growth+, not only reduce the number of candidates effectively but also outperform other algorithms substantially in terms of runtime, especially when databases contain lots of long transactions. Vincent S. Tseng, Bai-En Shie, Cheng-Wei Wu, Philip S. Yu |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2012 | Using Partially-Ordered Sequential Rules to Generate More Accurate Sequence Prediction
Philippe Fournier-Viger, Ted Gueniche, Vincent S. Tseng |
ADMA | 3 |
| 2012 | CTGR-Span: Efficient mining of cross-timepoint gene regulation sequential patterns from microarray datasetsabstractSequential pattern mining techniques have been widely used in different topics of interest, such as mining customer purchasing sequences from a transactional database. Notably, observation of gene expressions to discover gene regulations during biological or clinical progression via microarray approaches has become the dominant trend. By converting microarray datasets into the format of transactional databases, sequential patterns implying gene regulations could be identified. However, there exists no effective method in current studies that can handle such kind of dataset as every transaction may contain too many items/genes and the resultant patterns are very susceptible to item order. We propose a new method called CTGR-Span (Cross-Timepoint Gene Regulation Sequential Patterns) to efficiently mine CTGR-SPs (cross-timepoint gene regulation sequential patterns). The proposed method was experimented with two publicly available human time course microarray datasets and it outperformed traditional methods over 2,000 times in terms of the execution efficiency. Furthermore, via a Gene Ontology enrichment analysis, the resultant patterns are more meaningful biologically compared to previous literature reports. Hence, it could provide biologists more insights into the mechanisms of novel gene regulations in certain disease progressions. Chun-Pei Cheng, Yi-Lin Tsai, Vincent S. Tseng |
BIBM | 3 |
| 2012 | Personalized trip recommendation with multiple constraints by mining user check-in behaviorsabstractIn recent years, researches on travel recommendation have attracted extensive attentions due to the wide applications. Among them, one of the active topics is constraint-based trip recommendation for meeting user's personal requirements. Although a number of studies on this topic have been proposed in literatures, most of them only regard the user-specific constraints as some filtering conditions for planning the trip. In fact, immersing the constraints into travel recommendation systems to provide a personalized trip is desired for users. Furthermore, time complexity of trip planning from a set of attractions is sensitive to the scalability of travel regions. Hence, how to reduce the computational cost by parallel cloud computing techniques is also a critical issue. In this paper, we propose a novel framework named Personalized Trip Recommendation (PTR) to efficiently recommend the personalized trips meeting multiple constraints of users by mining user's check-in behaviors. In PTR, a mining-based module is first proposed to estimate the scores of attractions by considering both of user-based preferences and temporal-based properties. Then, a trip planning algorithm named Parallel Trip-Mine+ is proposed to efficiently plan the trip that satisfies multiple user-specific constraints. To our best knowledge, this is the first work on travel recommendation that considers the issues of multiple constraints, social relationship, temporal property and parallel computing simultaneously. Through comprehensive experimental evaluations on a real check-in dataset obtained from Gowalla, PTR is shown to deliver excellent performance. Eric Hsueh-Chan Lu, Ching-Yu Chen, Vincent S. Tseng |
SIGSPATIAL/GIS | 3 |
| 2012 | Followee recommendation in asymmetrical location-based social networksabstractResearches on recommending followees in social networks have attracted a lot of attentions in recent years. Existing studies on this topic mostly treat this kind of recommendation as just a type of friend recommendation. However, apart from making friends, the reason of a user to follow someone in social networks is inherently to satisfy his/her information needs in asymmetrical manner. In this paper, we propose a novel mining-based recommendation approach named Geographic-Textual-Social Based Followee Recommendation (GTS-FR), which takes into account the user movements, online texting and social properties to discover the relationship between users' information needs and provided information for followee recommendation. The core idea of our proposal is to discover users' similarity in terms of all the three properties of information which are provided by the users in a Location-Based Social Network (LBSN). To achieve this goal, we define three kinds of features to capture the key properties of users' interestingness from their provided information. In GTS-FR approach, we propose a series of novel similarity measurements to calculate similarity of each pair of users based on various properties. Based on the similarity, we make on-line recommendation for the followee a user might be interested in following. To our best knowledge, this is the first work on followee recommendation in LBSNs by exploring the geographic, textual and social properties simultaneously. Through a comprehensive evaluation using a real LBSN dataset, we show that the proposed GTS-FR approach delivers excellent performance and outperforms existing stat-of-the-art friend recommendation methods significantly. Jia-Ching Ying, Eric Hsueh-Chan Lu, Vincent S. Tseng |
UbiComp | 3 |
| 2012 | A One-Phase Method for Mining High Utility Mobile Sequential Patterns in Mobile Commerce Environments
Bai-En Shie, Jihong Cheng, Kun-Ta Chuang, Vincent S. Tseng |
IEA/AIE | 4 |
| 2012 | Mining Top-K Non-redundant Association Rules
Philippe Fournier-Viger, Vincent S. Tseng |
ISMIS | 2 |
| 2012 | Mining top-K high utility itemsetsabstractMining high utility itemsets from databases is an emerging topic in data mining, which refers to the discovery of itemsets with utilities higher than a user-specified minimum utility threshold min_util. Although several studies have been carried out on this topic, setting an appropriate minimum utility threshold is a difficult problem for users. If min_util is set too low, too many high utility itemsets will be generated, which may cause the mining algorithms to become inefficient or even run out of memory. On the other hand, if min_util is set too high, no high utility itemset will be found. Setting appropriate minimum utility thresholds by trial and error is a tedious process for users. In this paper, we address this problem by proposing a new framework named top-k high utility itemset mining, where k is the desired number of high utility itemsets to be mined. An efficient algorithm named TKU (Top-K Utility itemsets mining) is proposed for mining such itemsets without setting min_util. Several features were designed in TKU to solve the new challenges raised in this problem, like the absence of anti-monotone property and the requirement of lossless results. Moreover, TKU incorporates several novel strategies for pruning the search space to achieve high efficiency. Results on real and synthetic datasets show that TKU has excellent performance and scalability. Cheng-Wei Wu, Bai-En Shie, Vincent S. Tseng, Philip S. Yu |
KDD | 3 |
| 2012 | Efficient algorithms for mining maximal high utility itemsets from data streams with different models
Bai-En Shie, Philip S. Yu, Vincent S. Tseng |
Expert Syst. Appl. | 3 |
| 2012 | Societally connected multimedia across culturesabstractThe advance of the Internet in the past decade has radically changed the way people communicate and collaborate with each other. Physical distance is no more a barrier in online social networks, but cultural differences (at the individual, community, as well as societal levels) still govern human-human interactions and must be considered and leveraged in the online world. The rapid deployment of high-speed Internet allows humans to interact using a rich set of multimedia data such as texts, pictures, and videos. This position paper proposes to define a new research area called ‘connected multimedia’, which is the study of a collection of research issues of the super-area social media that receive little attention in the literature. By connected multimedia, we mean the study of the social and technical interactions among users, multimedia data, and devices across cultures and explicitly exploiting the cultural differences. We justify why it is necessary to bring attention to this new research area and what benefits of this new research area may bring to the broader scientific research community and the humanity. Zhongfei Zhang, Zhengyou Zhang, Ramesh Jain 0001, Yueting Zhuang, Noshir S. Contractor, Alex Hauptmann 0001, Alejandro Jaimes, Wanqing Li 0001, Alexander C. Loui, Tao Mei 0001, Nicu Sebe, Yonghong Tian 0001, Vincent S. Tseng, Qing Wang 0015, Changsheng Xu, Shiwen Yu |
J. Zhejiang Univ. Sci. C | 13 |
| 2012 | A Framework for Personal Mobile Commerce Pattern Mining and PredictionabstractDue to a wide range of potential applications, research on mobile commerce has received a lot of interests from both of the industry and academia. Among them, one of the active topic areas is the mining and prediction of users' mobile commerce behaviors such as their movements and purchase transactions. In this paper, we propose a novel framework, called Mobile Commerce Explorer (MCE), for mining and prediction of mobile users' movements and purchase transactions under the context of mobile commerce. The MCE framework consists of three major components: 1) Similarity Inference Model (SIM) for measuring the similarities among stores and items, which are two basic mobile commerce entities considered in this paper; 2) Personal Mobile Commerce Pattern Mine (PMCP-Mine) algorithm for efficient discovery of mobile users' Personal Mobile Commerce Patterns (PMCPs); and 3) Mobile Commerce Behavior Predictor (MCBP) for prediction of possible mobile user behaviors. To our best knowledge, this is the first work that facilitates mining and prediction of mobile users' commerce behaviors in order to recommend stores and items previously unknown to a user. We perform an extensive experimental evaluation by simulation and show that our proposals produce excellent results. Eric Hsueh-Chan Lu, Wang-Chien Lee, Vincent S. Tseng |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2011 | Mining Top-K Sequential Rules
Philippe Fournier-Viger, Vincent S. Tseng |
ADMA (2) | 2 |
| 2011 | Mining High Utility Mobile Sequential Patterns in Mobile Commerce Environments
Bai-En Shie, Hui-Fang Hsiao, Vincent S. Tseng, Philip S. Yu |
DASFAA (1) | 3 |
| 2011 | A fast algorithm for mining frequent closed itemsets over stream sliding windowabstractMining frequent patterns refers to the discovery of the sets of items that frequently appear in a transaction database. Many approaches have been proposed for mining frequent itemsets from a large database, but a large number of frequent itemsets may be discovered. In order to present users fewer but more important patterns, researchers are interested in discovering frequent closed itemsets which is a well-known complete and condensed representation of frequent itemsets. In this paper, we propose an efficient algorithm for discovering frequent closed itemsets over a data stream. The previous approaches need to do a large number of searching operations and computations to maintain the closed itemsets when a transaction is added or deleted. Our approach only performs few intersection operations on the transaction and the closed itemsets related to the transaction without doing any searching operation on the previous closed itemsets. The experimental results show that our approach significantly outperforms the previous approaches. Show-Jane Yen, Cheng-Wei Wu, Yue-Shi Lee, Vincent S. Tseng, Chaur-Heh Hsieh |
FUZZ-IEEE | 4 |
| 2011 | Semantic trajectory mining for location predictionabstractResearch on predicting movements of mobile users has attracted a lot of attentions in recent years. Many of those prediction techniques are developed based only on geographic features of mobile users' trajectories. In this paper, we propose a novel approach for predicting the next location of a user's movement based on both the geographic and semantic features of users' trajectories. The core idea of our prediction model is based on a novel cluster-based prediction strategy which evaluates the next location of a mobile user based on the frequent behaviors of similar users in the same cluster determined by analyzing users' common behavior in semantic trajectories. Through a comprehensive evaluation by experiments, our proposal is shown to deliver excellent performance. Jia-Ching Ying, Wang-Chien Lee, Tz-Chiao Weng, Vincent S. Tseng |
GIS | 4 |
| 2011 | Efficient Mining of a Concise and Lossless Representation of High Utility ItemsetsabstractMining high utility item sets from transactional databases is an important data mining task, which refers to the discovery of item sets with high utilities (e.g. high profits). Although several studies have been carried out, current methods may present too many high utility item sets for users, which degrades the performance of the mining task in terms of execution and memory efficiency. To achieve high efficiency for the mining task and provide a concise mining result to users, we propose a novel framework in this paper for mining closed+ high utility item sets, which serves as a compact and loss less representation of high utility item sets. We present an efficient algorithm called CHUD (Closed+ High Utility item set Discovery) for mining closed+ high utility item sets. Further, a method called DAHU (Derive All High Utility item sets) is proposed to recover all high utility item sets from the set of closed+ high utility item sets without accessing the original database. Results of experiments on real and synthetic datasets show that CHUD and DAHU are very efficient with a massive reduction (up to 800 times in our experiments) in the number of high utility item sets. In addition, when all high utility item sets are recovered by DAHU, the approach combining CHUD and DAHU also outperforms the state-of-the-art algorithms in mining high utility item sets. Cheng-Wei Wu, Philippe Fournier-Viger, Philip S. Yu, Vincent S. Tseng |
ICDM | 4 |
| 2011 | Photosense: Make sense of your photos with enriched harmonic music via emotion associationabstractThis paper proposes a novel audiovisual presentation system, called PhotoSense, to enrich photo navigation experience by associating emotionally harmonic music with a given photo collection. Different from many conventional photo visualization systems which predominantly focus on the visual elements for presentation, we explore both visual and aural perspectives which can enhance the browsing experience from each other. This is achieved by building an emotion space shared by visual and aural domains, and a set of emotion classifiers which can associate each visual and aural element with this space. Furthermore, we design a sequence matching algorithm to associate a set of music with a photo collection by maximizing similarity in the emotion space. Photo-Sense represents one of the first mash-up applications which build a natural connection between the ever increasing personal photo collections on the Web and music-sharing sites. Experiments show that PhotoSense provides better browsing experience for photo collections. Ja-Hwung Su, Ming-Hua Hsieh, Tao Mei 0001, Vincent S. Tseng |
ICME | 4 |
| 2011 | Prediction of Essential Genes by Mining Gene Ontology Semantics
Po-I Chiu, Hsuan-Cheng Huang, Vincent S. Tseng |
ISBRA | 4 |
| 2011 | Effective Content-Based Music Retrieval with Pattern-Based Relevance Feedback
Ja-Hwung Su, Tzu-Shiang Hung, Chun-Jen Lee, Chung-Li Lu, Wei-Lun Chang, Vincent S. Tseng |
KES (2) | 6 |
| 2011 | Challenges for Mobile Data Management in the Era of Cloud and Social ComputingabstractThe mobile data management community is experiencing a rapid evolutionary change due to the worldwide diffusion of always-on mobile devices and to the increased popularity of location and context-aware mobile applications. Accordingly to recent studies, in two years from now one fourth of the total mobile data will come from audio and video streaming and nearly all the rest from other Internet services. A large part of the increase in mobile data will come from cloud computing applications that are massively used for storing personal data, for sharing data, as well as for utility software (such as maps) and productivity tools. Social networking will strongly influence the way mobile users choose, share and use content from mobile devices. On the other side mobile devices are changing the way social networks have been used till now introducing geo-tagging, location sharing, and many innovative location based services. Chatschik Bisdikian, Bernhard Mitschang, Dino Pedreschi, Vincent S. Tseng, Claudio Bettini |
Mobile Data Management (1) | 4 |
| 2011 | Trip-Mine: An Efficient Trip Planning Approach with Travel Time ConstraintsabstractWith the rapid development of wireless telecommunication technologies, a number of studies have been done on the Location-Based Services (LBSs) due to wide applications. Among them, one of the active topics is travel recommendation. Most of previous studies focused on recommendations of attractions or trips based on the user's location. However, such recommendation results may not satisfy the travel time constraints of users. Besides, the efficiency of trip planning is sensitive to the scalability of travel regions. In this paper, we propose a novel data mining-based approach, namely Trip-Mine, to efficiently find the optimal trip which satisfies the user's travel time constraint based on the user's location. Furthermore, we propose three optimization mechanisms based on Trip-Mine to further enhance the mining efficiency and memory storage requirement for optimal trip finding. To the best of our knowledge, this is the first work that takes efficient trip planning and travel time constraints into account simultaneously. Finally, we performed extensive experimental evaluations and show that our proposals deliver excellent results. Eric Hsueh-Chan Lu, Chih-Yuan Lin, Vincent S. Tseng |
Mobile Data Management (1) | 3 |
| 2011 | Discovering relational-based association rules with multiple minimum supports on microarray datasetsabstractMOTIVATION: Association rule analysis methods are important techniques applied to gene expression data for finding expression relationships between genes. However, previous methods implicitly assume that all genes have similar importance, or they ignore the individual importance of each gene. The relation intensity between any two items has never been taken into consideration. Therefore, we proposed a technique named REMMAR (RElational-based Multiple Minimum supports Association Rules) algorithm to tackle this problem. This method adjusts the minimum relation support (MRS) for each gene pair depending on the regulatory relation intensity to discover more important association rules with stronger biological meaning. RESULTS: In the actual case study of this research, REMMAR utilized the shortest distance between any two genes in the Saccharomyces cerevisiae gene regulatory network (GRN) as the relation intensity to discover the association rules from two S.cerevisiae gene expression datasets. Under experimental evaluation, REMMAR can generate more rules with stronger relation intensity, and filter out rules without biological meaning in the protein-protein interaction network (PPIN). Furthermore, the proposed method has a higher precision (100%) than the precision of reference Apriori method (87.5%) for the discovered rules use a literature survey. Therefore, the proposed REMMAR algorithm can discover stronger association rules in biological relationships dissimilated by traditional methods to assist biologists in complicated genetic exploration. Chun-Pei Cheng, Vincent S. Tseng |
Bioinform. | 3 |
| 2011 | Discovery of high utility itemsets from on-shelf time periods of products
Guo-Cheng Lan, Tzung-Pei Hong, Vincent S. Tseng |
Expert Syst. Appl. | 3 |
| 2011 | An Intelligent and Effective Mechanism for Mental Disorder Treatment by Using Biofeedback Analysis and Web TechnologiesabstractIn the medical science field, the treatment of mental disorders has been an important issue. Recently, the treatment of mental disorders through biofeedback therapy is emerging as an important topic. A number of studies on the integration of mental healthcare and the Internet have also been proposed, since the Internet plays a more and more crucial role in various healthcare applications. By performing biofeedback therapy via the Internet, the treatment time can be highly reduced and the medical costs can also be significantly cut down. In view of this, this research aims at developing an intelligent treatment system for the patients with mental disorders by integrating biofeedback therapy and web technology. The system provides not only a convenient mechanism for the patients to perform biofeedback therapy at home but also an effective communication channel for patients and medical professionals. Moreover, the functions which enable the therapists to manage their patients more conveniently are proposed. The results of this research are expected to bring a pivotal impact on the healthcare industry with increased and enhanced levels of technology and services. Bai-En Shie, Fong-Lin Jang, Vincent S. Tseng |
Int. J. Softw. Eng. Knowl. Eng. | 3 |
| 2011 | Hybrid data mining approaches for prevention of drug dispensing errors
Lien-Chin Chen, Chun-Hao Chen, Hsiao-Ming Chen, Vincent S. Tseng |
J. Intell. Inf. Syst. | 4 |
| 2011 | Mining fastest path from trajectories with multiple destinations in road networks
Eric Hsueh-Chan Lu, Wang-Chien Lee, Vincent S. Tseng |
Knowl. Inf. Syst. | 3 |
| 2011 | Genetic-fuzzy mining with multiple minimum supports based on fuzzy clustering
Chun-Hao Chen, Tzung-Pei Hong, Vincent S. Tseng |
Soft Comput. | 3 |
| 2011 | Mining Cluster-Based Temporal Mobile Sequential Patterns in Location-Based Service EnvironmentsabstractResearches on Location-Based Service (LBS) have been emerging in recent years due to a wide range of potential applications. One of the active topics is the mining and prediction of mobile movements and associated transactions. Most of existing studies focus on discovering mobile patterns from the whole logs. However, this kind of patterns may not be precise enough for predictions since the differentiated mobile behaviors among users and temporal periods are not considered. In this paper, we propose a novel algorithm, namely, Cluster-based Temporal Mobile Sequential Pattern Mine (CTMSP-Mine), to discover the Cluster-based Temporal Mobile Sequential Patterns (CTMSPs). Moreover, a prediction strategy is proposed to predict the subsequent mobile behaviors. In CTMSP-Mine, user clusters are constructed by a novel algorithm named Cluster-Object-based Smart Cluster Affinity Search Technique (CO-Smart-CAST) and similarities between users are evaluated by the proposed measure, Location-Based Service Alignment (LBS-Alignment). Meanwhile, a time segmentation approach is presented to find segmenting time intervals where similar mobile characteristics exist. To our best knowledge, this is the first work on mining and prediction of mobile behaviors with considerations of user relations and temporal property simultaneously. Through experimental evaluation under various simulated conditions, the proposed methods are shown to deliver excellent performance. Eric Hsueh-Chan Lu, Vincent S. Tseng, Philip S. Yu |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2011 | Efficient Relevance Feedback for Content-Based Image Retrieval by Mining User Navigation PatternsabstractNowadays, content-based image retrieval (CBIR) is the mainstay of image retrieval systems. To be more profitable, relevance feedback techniques were incorporated into CBIR such that more precise results can be obtained by taking user's feedbacks into account. However, existing relevance feedback-based CBIR methods usually request a number of iterative feedbacks to produce refined search results, especially in a large-scale image database. This is impractical and inefficient in real applications. In this paper, we propose a novel method, Navigation-Pattern-based Relevance Feedback (NPRF), to achieve the high efficiency and effectiveness of CBIR in coping with the large-scale image data. In terms of efficiency, the iterations of feedback are reduced substantially by using the navigation patterns discovered from the user query log. In terms of effectiveness, our proposed search algorithm NPRFSearch makes use of the discovered navigation patterns and three kinds of query refinement strategies, Query Point Movement (QPM), Query Reweighting (QR), and Query Expansion (QEX), to converge the search space toward the user's intention effectively. By using NPRF method, high quality of image retrieval on RF can be achieved in a small number of feedbacks. The experimental results reveal that NPRF outperforms other existing methods significantly in terms of precision, coverage, and number of feedbacks. Ja-Hwung Su, Wei-Jyun Huang, Philip S. Yu, Vincent S. Tseng |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2011 | Effective Semantic Annotation by Image-to-Concept Distribution ModelabstractImage annotation based on visual features has been a difficult problem due to the diverse associations that exist between visual features and human concepts. In this paper, we propose a novel approach called Annotation by Image-to-Concept Distribution Model (AICDM) for image annotation by discovering the associations between visual features and human concepts from image-to-concept distribution. Through the proposed image-to-concept distribution model, visual features and concepts can be bridged to achieve high-quality image annotation. In this paper, we propose to use “visual features”, “models”, and “visual genes” which represent analogous functions to the biological chromosome, DNA, and gene. Based on the proposed models using entropy, tf-idf, rules, and SVM, the goal of high-quality image annotation can be achieved effectively. Our empirical evaluation results reveal that the AICDM method can effectively alleviate the problem of visual-to-concept diversity and achieve better annotation results than many existing state-of-the-art approaches in terms of precision and recall. Ja-Hwung Su, Chien-Li Chou, Ching-Yung Lin, Vincent S. Tseng |
IEEE Trans. Multim. | 4 |
| 2010 | A Three-Scan Algorithm to Mine High On-Shelf Utility Itemsets
Guo-Cheng Lan, Tzung-Pei Hong, Vincent S. Tseng |
ACIIDS (2) | 3 |
| 2010 | A SPEA2-based genetic-fuzzy mining algorithmabstractIn this paper, we adopt a more sophisticated multi-objective approach, SPEA2, to find appropriate sets of membership functions for fuzzy data mining. Two objective functions are used to find the Pareto front. The first one is to minimize the suitability of membership functions and the second one is to maximize the total number of large 1-itemsets. An experimental comparison with the previous approach is also made to show the effectiveness of the proposed approach in finding the Pareto-front membership functions. Chun-Hao Chen, Tzung-Pei Hong, Vincent S. Tseng |
FUZZ-IEEE | 3 |
| 2010 | Automatic Chinese Text Classification Using N-Gram Model
Show-Jane Yen, Yue-Shi Lee, Yu-Chieh Wu, Jia-Ching Ying, Vincent S. Tseng |
ICCSA (3) | 5 |
| 2010 | Effective image semantic annotation by discovering visual-concept associations from image-concept distribution modelabstractUp to the present, the contemporary studies are not really successful in image annotation due to some critical problems like diverse regularities between visual features and human concepts. Such diverse regularities make it hard to annotate the image semantics correctly. In this paper, we propose a novel approach called AICDM (Annotation by Image-Concept Distribution Model) for image annotation by discovering the associations between visual features and human concepts from image-concept distribution. Through the proposed image-concept distribution model, the uncertain regularities between visual features and human concepts can be clarified for achieving high-quality image annotation. The empirical evaluation results also reveal that our proposed AICDM method can effectively alleviate the uncertain regularity problem and bring out better annotation results than other existing approaches in terms of precision and recall. Ja-Hwung Su, Chien-Li Chou, Ching-Yung Lin, Vincent S. Tseng |
ICME | 4 |
| 2010 | UP-Growth: an efficient algorithm for high utility itemset miningabstractMining high utility itemsets from a transactional database refers to the discovery of itemsets with high utility like profits. Although a number of relevant approaches have been proposed in recent years, they incur the problem of producing a large number of candidate itemsets for high utility itemsets. Such a large number of candidate itemsets degrades the mining performance in terms of execution time and space requirement. The situation may become worse when the database contains lots of long transactions or long high utility itemsets. In this paper, we propose an efficient algorithm, namely UP-Growth (Utility Pattern Growth), for mining high utility itemsets with a set of techniques for pruning candidate itemsets. The information of high utility itemsets is maintained in a special data structure named UP-Tree (Utility Pattern Tree) such that the candidate itemsets can be generated efficiently with only two scans of the database. The performance of UP-Growth was evaluated in comparison with the state-of-the-art algorithms on different types of datasets. The experimental results show that UP-Growth not only reduces the number of candidates effectively but also outperforms other algorithms substantially in terms of execution time, especially when the database contains lots of long transactions. Vincent S. Tseng, Cheng-Wei Wu, Bai-En Shie, Philip S. Yu |
KDD | 1 |
| 2010 | A novel two-level clustering method for time series data analysis
Cheng-Ping Lai, Pau-Choo Chung, Vincent S. Tseng |
Expert Syst. Appl. | 3 |
| 2010 | A novel prediction-based strategy for object tracking in sensor networks by mining seamless temporal movement patterns
Kawuu Weicheng Lin, Ming-Hua Hsieh, Vincent S. Tseng |
Expert Syst. Appl. | 3 |
| 2010 | Effective content-based video retrieval using pattern-indexing and matching techniques
Ja-Hwung Su, Hsin-Ho Yeh, Vincent S. Tseng |
Expert Syst. Appl. | 4 |
| 2010 | Personalized rough-set-based recommendation by integrating multiple contents and collaborative information
Ja-Hwung Su, Bo-Wen Wang, Chin-Yuan Hsiao, Vincent S. Tseng |
Inf. Sci. | 4 |
| 2009 | Speeding up genetic-fuzzy mining by fuzzy clusteringabstractIn the past, we proposed an algorithm for extracting appropriate multiple minimum support values, membership functions and fuzzy association rules from quantitative transactions. In this paper, an enhanced approach, called the fuzzy cluster-based genetic-fuzzy mining approach for items with multiple minimum supports (FCGFMMS), is proposed to speed up the evaluation process and keep nearly the same quality of solutions as the previous one. It divides the chromosomes in a population into several clusters by the fuzzy k-means clustering approach and evaluates each individual according to both their cluster and their own information. Experimental results also show the effectiveness and the efficiency of the proposed approach. Chun-Hao Chen, Tzung-Pei Hong, Vincent S. Tseng |
FUZZ-IEEE | 3 |
| 2009 | Mining Cluster-Based Mobile Sequential Patterns in Location-Based Service EnvironmentsabstractIn recent years, a number of studies have been done on Location-Based Service (LBS) due to their wide range of potential applications. In this paper, we propose a novel data mining algorithm named Cluster-based Mobile Sequential Pattern Mine (CMSP-Mine) for efficiently discovering the Cluster-based Mobile Sequential Patterns (CMSPs) of users in LBS environments. In CMSP-Mine, we first propose a transaction similarity measurement named Location-Based Service Alignment (LBS-Alignment) to evaluate the similarity between two mobile transaction sequences. Then, we propose a transaction clustering algorithm named Cluster-Object based Smart Cluster Affinity Search Technique (CO-Smart-CAST) to form a user cluster model of the mobile transactions based on LBS-Alignment. Furthermore, we proposed the novel prediction strategy that utilizes the discovered CMSPs to precisely predict the next movement of mobile users. To our best knowledge, this is the first work on mining the mobile sequential patterns associated with moving path and user clusters in LBS environments. Finally, through a series of experiments, our proposed methods were shown to deliver excellent performance in terms of efficiency, accuracy and applicability under various system conditions. Eric Hsueh-Chan Lu, Vincent S. Tseng |
Mobile Data Management | 2 |
| 2009 | An improved approach to find membership functions and multiple minimum supports in fuzzy data mining
Chun-Hao Chen, Tzung-Pei Hong, Vincent S. Tseng |
Expert Syst. Appl. | 3 |
| 2009 | Mining fuzzy frequent trends from time series
Chun-Hao Chen, Tzung-Pei Hong, Vincent S. Tseng |
Expert Syst. Appl. | 3 |
| 2009 | A novel method for personalized music recommendation
Cheng-Che Lu, Vincent S. Tseng |
Expert Syst. Appl. | 2 |
| 2009 | Effective temporal data classification by integrating sequential pattern mining and probabilistic induction
Vincent S. Tseng, Chao-Hui Lee |
Expert Syst. Appl. | 1 |
| 2009 | Trading decryption for speeding encryption in Rebalanced-RSA
Mu-En Wu, M. Jason Hinek, Cheng-Ta Yang, Vincent S. Tseng |
J. Syst. Softw. | 5 |
| 2009 | Energy-efficient real-time object tracking in multi-level sensor networks by mining and predicting movement patterns
Vincent S. Tseng, Eric Hsueh-Chan Lu |
J. Syst. Softw. | 1 |
| 2009 | Cluster-based genetic segmentation of time series with DWT
Vincent S. Tseng, Chun-Hao Chen, Pai-Chieh Huang, Tzung-Pei Hong |
Pattern Recognit. Lett. | 1 |
| 2009 | A genetic-fuzzy mining approach for items with multiple minimum supports
Chun-Hao Chen, Tzung-Pei Hong, Vincent S. Tseng, Chang-Shing Lee |
Soft Comput. | 3 |
| 2008 | An Integrated Data Mining System for Patient Monitoring with Applications on Asthma CareabstractIn this paper, we proposed an integrated data mining system for patient monitoring with applications on asthma care. In this system, two data mining methods named PBD and PBC are designed for predicting asthma attacks. The main methodology is to extract the significant information of asthma attacks and build classifiers by using users' daily bio-signal records and environmental data. Meanwhile, helpful medical information and suggestions supported by doctors are applied. In this way, the proposed system can predict the chances of asthma attacks and provide patients with the proper medical instructions or health messages. The experimental evaluation results proved that the proposed mechanism is effective and reliable in asthma attack prediction. Vincent S. Tseng, Chao-Hui Lee, Jessie Chia-Yu Chen |
CBMS | 1 |
| 2008 | A cluster-based genetic approach for segmentation of time series and pattern discoveryabstractIn the past, we proposed a time series segmentation approach by combining the clustering technique, the discrete wavelet transformation and the genetic algorithm to automatically find segments and patterns from a time series. In this paper, we propose an enhanced approach to solve the problems that may occur during the evolution process. Two factors, namely the density factor and the distortion factor, are used to solve them. The distortion factor is used to avoid the distortion of the segments and the density factor is used to avoid generation of meaningless patterns. The fitness value of a chromosome is then evaluated by the distances of segments and these two factors. Experimental results on a financial dataset also show the effectiveness of the proposed approach. Vincent S. Tseng, Chun-Hao Chen, Pai-Chieh Huang, Tzung-Pei Hong |
IEEE Congress on Evolutionary Computation | 1 |
| 2008 | Development of a Vital Sign Data Mining System for Chronic Patient MonitoringabstractIn recent years, the structure of global population keeps going towards highly-aged continuously. The development of chronic patient medical care system becomes important and meaningful since people paid a lot attention to medical prevention. The medical care system has to provide alerts in time before the severe chronic illness occurs, such as stroke, diabetics, heart disease. Thus, necessary procedures can be taken in short time to save one precious life. In this paper, we presented a data mining system for chronic patient monitoring with applications on caring of cardiovascular patients. By mining vital signs like ECG, the system can predict with a classification tree and inform doctors to take actions if any anomaly could happen. A series of experiments on PAF data showed that our system can stably predict the anomaly from patientspsila ECG data without coding of medical rules as done in other existing approaches. Vincent S. Tseng, Lee-Cheng Chen, Chao-Hui Lee, Jin-Shang Wu, Yu-Chia Hsu |
CISIS | 1 |
| 2008 | A divide-and-conquer genetic-fuzzy mining approach for items with multiple minimum supportsabstractSince items may have their own characteristics, different minimum support values and membership functions may be specified for different items. In this paper, an enhanced approach is proposed, which processes the items in a divide-and-conquer strategy. The approach is designed for finding minimum support values, membership functions, and fuzzy association rules. Possible solutions are evaluated by their requirement satisfaction divided by their suitability of derived membership functions. The proposed GA framework maintains multiple populations, each for one itempsilas minimum support value and membership functions. The final best minimum support values and membership functions in all the populations are then gathered together to be used for mining fuzzy association rules. Experimental results also show the effectiveness of the proposed approach. Chun-Hao Chen, Tzung-Pei Hong, Vincent S. Tseng |
FUZZ-IEEE | 3 |
| 2008 | Intelligent Concept-Oriented and Content-Based Image Retrieval by using data mining and query decomposition techniquesabstractTraditional image retrieval based on visual-based matching is not effective in multimedia applications. Consequently, the modeling of high-level human sense for image retrieval has been a challenging issue over the past few years. In fact, the concepts hidden in the images play key roles in semantic image retrieval. In this paper, we propose a novel method named Intelligent Concept-Oriented Search (ICOS) that can capture the high-level concepts in images by utilizing data mining and query decomposition techniques. The contributions of the proposed method lie in that we provide: 1) effective annotation for conceptual objects, 2) association mining for conceptual objects, 3) visual ranking for conceptual objects and 4) intelligent search method for enhancing high-level concept image retrieval. Through experimental evaluations, ICOS is shown to be very effective and efficient in capturing the implicit high-level concepts for image retrieval. Vincent S. Tseng, Ja-Hwung Su, Hao-Hua Ku, Bo-Wen Wang |
ICME | 1 |
| 2008 | A Cluster-Based Genetic-Fuzzy Mining Approach for Items with Multiple Minimum Supports
Chun-Hao Chen, Tzung-Pei Hong, Vincent S. Tseng |
PAKDD | 3 |
| 2008 | Constrained Clustering for Gene Expression Data Mining
Vincent S. Tseng, Lien-Chin Chen, Ching-Pin Kao |
PAKDD | 1 |
| 2008 | Semantic Video Annotation by Mining Association Patterns from Visual and Speech Features
Vincent S. Tseng, Ja-Hwung Su, Jhih-Hong Huang, Chih-Jen Chen |
PAKDD | 1 |
| 2008 | An efficient algorithm for mining temporal high utility itemsets from data streams
Chun-Jung Chu, Vincent S. Tseng, Tyne Liang |
J. Syst. Softw. | 2 |
| 2008 | Prediction of user navigation patterns by mining the temporal web usage evolution
Vincent S. Tseng, Kawuu Weicheng Lin, Jeng-Chuan Chang |
Soft Comput. | 1 |
| 2008 | Cluster-Based Evaluation in Fuzzy-Genetic Data MiningabstractData mining is commonly used in attempts to induce association rules from transaction data. Most previous studies focused on binary-valued transaction data. Transactions in real-world applications, however, usually consist of quantitative values. In the past, we proposed a fuzzy-genetic data-mining algorithm for extracting both association rules and membership functions from quantitative transactions. It used a combination of large 1-itemsets and membership-function suitability to evaluate the fitness values of chromosomes. The calculation for large 1-itemsets could take a lot of time, especially when the database to be scanned could not totally fed into main memory. In this paper, an enhanced approach, called the cluster-based fuzzy-genetic mining algorithm, is thus proposed to speed up the evaluation process and keep nearly the same quality of solutions as the previous one. It divides the chromosomes in a population into clusters by the - means clustering approach and evaluates each individual according to both cluster and their own information. Experimental results also show the effectiveness and efficiency of the proposed approach. Chun-Hao Chen, Vincent S. Tseng, Tzung-Pei Hong |
IEEE Trans. Fuzzy Syst. | 2 |
| 2008 | Integrated Mining of Visual Features, Speech Features, and Frequent Patterns for Semantic Video AnnotationabstractTo support effective multimedia information retrieval, video annotation has become an important topic in video content analysis. Existing video annotation methods put the focus on either the analysis of low-level features or simple semantic concepts, and they cannot reduce the gap between low-level features and high-level concepts. In this paper, we propose an innovative method for semantic video annotation through integrated mining of visual features, speech features, and frequent semantic patterns existing in the video. The proposed method mainly consists of two main phases: 1) Construction of four kinds of predictive annotation models, namely speech-association, visual-association, visual-sequential, and statistical models from annotated videos. 2) Fusion of these models for annotating un-annotated videos automatically. The main advantage of the proposed method lies in that all visual features, speech features, and semantic patterns are considered simultaneously. Moreover, the utilization of high-level rules can effectively complement the insufficiency of statistics-based methods in dealing with complex and broad keyword identification in video annotation. Through empirical evaluation on NIST TRECVID video datasets, the proposed approach is shown to enhance the performance of annotation substantially in terms of precision, recall, and F-measure. Vincent S. Tseng, Ja-Hwung Su, Jhih-Hong Huang, Chih-Jen Chen |
IEEE Trans. Multim. | 1 |
| 2007 | A modified approach to speed up genetic-fuzzy data mining with divide-and-conquer strategyabstractIn the past, we proposed a fuzzy data-mining algorithm for extracting both association rules and membership functions from quantitative transactions based on the divide-and-conquer strategy. In this paper, an enhanced approach, called the cluster-based genetic-fuzzy mining algorithm, is thus proposed to speed up the evaluation process and keep nearly the same quality of solutions as the previous one. It first divides the chromosomes in a population into k clusters by the A-means clustering approach and evaluates each individual according to its own information and the information of the cluster it belongs to. The final best sets of membership functions in all the populations are then gathered together for mining fuzzy association rules. Experimental results also show the effectiveness and efficiency of the proposed approach. Chun-Hao Chen, Tzung-Pei Hong, Vincent S. Tseng |
IEEE Congress on Evolutionary Computation | 3 |
| 2007 | Analysis and Prevention of Dispension Errors by Using Data Mining TechniquesabstractMedical treatment techniques have been improved continuously in the past years. However, the better approaches are still needed to solve medical treatment problems. One important topic in this field is the analysis and prevention of medication errors. In this paper, we focus on the problem of dispensing error that is one important problem of medication errors and we proposed a prevention model by using three approaches. The proposed dispensing error mining framework consists of two phases, namely the modeling and prediction phases. Firstly, Statistical approach (logistic regression) and data mining approaches (C4.5 and SVM) are used to analyze dispensing error problem and to build classification models. Three kinds of factors, namely drug-names factor, drug-properties factor and environmental factor, with totally thirteen attributes are used in the modeling phase. In prediction phase, new drugs thus can be analyzed for the probability of dispensing error by the model so as to prevent dispensing error. At last, experimental results on real dataset showed that the proposed approach is effective and the considered factors can actually increase the accuracy of the model Vincent S. Tseng, Chun-Hao Chen, Hsiao-Ming Chen, Hui-Jen Chang, Chin-Tai Yu |
CIBCB | 1 |
| 2007 | Gene Relation Discovery by Mining Similar Subsequences in Time-Series Microarray DataabstractTime-series microarray techniques are newly used to monitor large-scale gene expression profiles for studying biological systems. Previous studies have discovered novel regulatory relations among genes by analyzing time-series microarray data. In this study, we investigate the problem of mining similar subsequences in time-series microarray data so as to discover novel gene relations. A functional relationship among genes often presents itself by locally similar and potentially time-shifted patterns in their expression profiles. Although a number of studies have been done on time-series data analysis, they are insufficient in handling four important issues for time-series microarray data analysis, namely scaling, offset, shift, and noise. We proposed a novel method to address the four issues simultaneously, which consists of three phase, namely angular transformation, symbolic transformation and suffix-tree-based similar subsequences searching. Through experimental evaluation, it is shown that our method can effectively discover biological relations among genes by identifying the similar subsequences. Moreover, the execution efficiency of our method is much better than other approaches Vincent S. Tseng, Lien-Chin Chen, Jian-Jie Liu |
CIBCB | 1 |
| 2007 | A Genetic-Fuzzy Mining Approach for Items with Multiple Minimum SupportsabstractIn the past, we proposed a genetic-fuzzy data-mining algorithm for extracting both association rules and membership functions from quantitative transactions under a single minimum support. In real applications, different items may have different criteria to judge their importance. In this paper, we thus propose an algorithm which combines clustering, fuzzy and genetic concepts for extracting reasonable multiple minimum support values, membership functions and fuzzy association rules form quantitative transactions. It first uses the k-means clustering approach to gather similar items into groups. All items in the same cluster are considered to have similar characteristics and are assigned similar values for initializing a better population. Each chromosome is then evaluated by the criteria of requirement satisfaction and suitability of membership functions to estimate its fitness value. Experimental results also show the effectiveness and the efficiency of the proposed approach. Chun-Hao Chen, Tzung-Pei Hong, Vincent S. Tseng, Chang-Shing Lee |
FUZZ-IEEE | 3 |
| 2007 | Mining temporal mobile sequential patterns in location-based service environmentsabstractIn recent years, a number of studies have been done on Location-Based Service (LBS) due to the wide applications. One important research issue is the tracking and prediction of users’ mobile behavior. In this paper, we propose a novel data mining algorithm named TMSP-Mine for efficiently discovering the Temporal Mobile Sequential Patterns (TMSPs) of users in LBS environments. To our best knowledge, this is the first work on mining the mobile sequential patterns associated with moving paths and time intervals in LBS environments. Furthermore, we propose novel location prediction strategies that utilize the discovered TMSPs to effectively predict the next movement of mobile users. Finally, we conducted a series of experiments to evaluate the performance of the proposed method under different system conditions by varying various parameters. Vincent S. Tseng, Eric Hsueh-Chan Lu, Cheng-Hsien Huang |
ICPADS | 1 |
| 2007 | FCBIR: A Fuzzy Matching Technique for Content-Based Image Retrieval
Vincent S. Tseng, Ja-Hwung Su, Wei-Jyun Huang |
IFSA (2) | 1 |
| 2007 | Energy efficient strategies for object tracking in sensor networks: A data mining approach
Vincent S. Tseng, Kawuu Weicheng Lin |
J. Syst. Softw. | 1 |
| 2007 | A Novel Similarity-Based Fuzzy Clustering Algorithm by Integrating PCM and Mountain MethodabstractThe fuzzy c-means (FCM) and possibilistic c-means (PCM) algorithms have been utilized in a wide variety of fields and applications. Although many methods are derived from the FCM and PCM for clustering various types of spatial data, relational clustering has received much less attention. Most fuzzy clustering methods can only process the spatial data (e.g., in Euclidean space) instead of the nonspatial data (e.g., where the Pearson's correlation coefficient is used as similarity measure). In this paper, we propose a novel clustering method, similarity-based PCM (SPCM), which is fitted for clustering nonspatial data without requesting users to specify the cluster number. The main idea behind the SPCM is to extend the PCM for similarity-based clustering applications by integration with the mountain method. The SPCM has the merit that it can automatically generate clustering results without requesting users to specify the cluster number. Through performance evaluation on real and synthetic data sets, the SPCM method is shown to perform excellently for similarity-based clustering in clustering quality, even in a noisy environment with outliers. This complements the deficiency of other fuzzy clustering methods when applied to similarity-based clustering applications. Vincent S. Tseng, Ching-Pin Kao |
IEEE Trans. Fuzzy Syst. | 1 |
| 2006 | Feature Selection for Medical Data Mining: Comparisons of Expert Judgment and Automatic ApproachesabstractData mining refers to the process of automatic extracting previously unknown, valid, and actionable patterns or knowledge from large databases for crucial decision support. Among different data mining technique, classification analysis is widely adopted for healthcare applications for supporting medical diagnostic decisions, improving quality of patient care, etc. If a training dataset contains irrelevant features (i.e., attributes), classification analysis may produce less accurate and less understandable results. Two commonly employed feature selection approaches include use of automatic feature selection mechanisms (i.e., data-driven) or expert judgment (i.e., knowledgedriven). Due to differences in their underlying processes, the two prevailing feature selection approaches may have their unique biases that possibly lead to dissimilar classification effectiveness. In this study, we empirically evaluate the classification effectiveness resulted from the two feature selection approaches on a risk prediction of cardiovascular disease dataset. Our evaluation results suggest that the feature subsets selected domain experts improve the sensitivity of a classifier, while the feature subsets selected by an automatic feature selection mechanism improve the predictive power of a classifier on the majority class (i.e., the specificity in this study). Tsang-Hsiang Cheng, Chih-Ping Wei, Vincent S. Tseng |
CBMS | 3 |
| 2006 | Discovering Gene Clusters via Integrated Analysis on Time-Series and Group-Comparative Microarray DatasetsabstractIn this paper, we propose a novel gene clustering method named TGmix through integrated analysis on two types of datasets, namely the time-series and twogroup microarray datasets. The goal of the proposed method is to discover genes as biomarkers that have similar expression profiles in time-series conditions and are also significantly differentially expressed in two-group conditions. We applied the proposed method to microarray datasets for rat’s wound healing experiment, and the genes discovered in the same cluster conform to the analysis goal with related biological functions. Vincent S. Tseng, Lien-Chin Chen, Yao-Dung Hsieh |
CBMS | 1 |
| 2006 | A Less Domain-dependent Fuzzy Mining Algorithm for Frequent TrendsabstractTime series analysis has always been an important and interesting research field due to its frequent appearance in different applications. In the past, many mining approaches were proposed to find useful patterns from time-series data. Time-series data, however, are usually quantitative values and domain knowledge is needed to predefine crisp intervals of categories for a mining process to proceed. In this paper, we thus propose an algorithm based on Udechukwu et al.'s approach to mine fuzzy frequent trends from time series without referring to domain knowledge. The proposed approach first transforms data values into angles, and then uses a sliding window to generate continues subsequences from angular series. The a priori-like fuzzy mining algorithm is then used to generate frequent trends. Appropriate post-processing is also performed to remove redundant patterns. Finally, experiments are also made for different parameter settings. Chun-Hao Chen, Tzung-Pei Hong, Vincent S. Tseng |
FUZZ-IEEE | 3 |
| 2006 | A Cluster-Based Fuzzy-Genetic Mining Approach for Association Rules and Membership FunctionsabstractData mining is most commonly used in attempts to induce association rules from transaction data. Transactions in real-world applications, however, usually consist of quantitative values. Designing a sophisticated data-mining algorithm able to deal with various types of data presents a challenge to workers in this research field. In this paper, a cluster-based fuzzy-genetic mining algorithm is proposed for extracting both fuzzy association rules and membership functions from quantitative transactions. The proposed algorithm can dynamically adjust membership functions by genetic algorithms and uses them to fuzzify quantitative transactions. It can also speed up the evaluation process and keep good quality of solutions by clustering chromosomes. Experimental results show the effectiveness of the proposed approach. Chun-Hao Chen, Tzung-Pei Hong, Vincent S. Tseng |
FUZZ-IEEE | 3 |
| 2006 | Towards Parameter-less and Similarity-based Fuzzy Clustering based on PCM MethodabstractThe fuzzy clustering algorithms have been applied in a wide variety of fields. In this paper, we propose a novel fuzzy clustering method named Similarity-based PCM (SPCM), which is parameter-less and suitable for similarity-based clustering applications. The main idea behind SPCM is to integrate PCM clustering with the Mountain Method (MM) such that the good fuzzy clustering result can be generated automatically without requesting users to specify parameters like the cluster number. This complements the deficiency of other existing relational fuzzy clustering methods when applied to similarity-based clustering applications. For example, FANNY, RFCM, NERFCM, and FRC request the specification of the number of clusters and are severely sensitive to outliers. Although R-RFCM, R-NERFCM, and R-FRC are robust in noisy environments, they request the specification of the number of clusters and require good initialization. Through performance evaluation on both of real and synthetic data sets, the SPCM is shown to perform excellently in clustering quality with various kinds of similarity measures, even in a noisy environment with outliers. Therefore, the SPCM can serve as a promising method for parameter-less and similarity-based fuzzy clustering applications. Vincent S. Tseng, Ching-Pin Kao |
SMC | 1 |
| 2006 | Efficient mining and prediction of user behavior patterns in mobile web systems
Vincent S. Tseng, Kawuu Weicheng Lin |
Inf. Softw. Technol. | 1 |
| 2005 | Mining Sequential Mobile Access Patterns Efficiently in Mobile Web SystemsabstractThe rapid advance of wireless and Web technologies enable the mobile Web applications to provide plenty kinds of services for mobile users. Under a mobile Web system, analyzing mobile user's movement sequences and requested services is important for wide applications in wireless communication like data allocation, data replication, location-based and personalization services. The main challenge in this research issue is to effectively deal with the user's diverse behavior and the huge amount of data. However, to our best knowledge, no studies have been done on the problem of mining sequential mobile access patterns with both movement and service requests considered simultaneously. In this paper, we propose a novel data mining method, namely SMAP-Mine, that can discover patterns of sequential movement associated with requested services for mobile users in mobile Web systems. Through empirical evaluation on various simulation conditions, the proposed method is shown to deliver excellent performance in terms of accuracy, execution efficiency, and scalability. Vincent S. Tseng, Kawuu Weicheng Lin |
AINA | 1 |
| 2005 | An Energy-Efficient Approach for Real-Time Tracking of Moving Object in Multi-Level Sensor NetworksabstractIn this paper, we propose a new approach for efficient and real-time tracking of the moving objects in sensor networks by mining the movement log. In our approach, we first conduct the hierarchical clustering to form a hierarchical model for the sensor nodes. Secondly, the movement logs of the moving objects are analyzed by a data mining algorithm to obtain the movement rules, which are then used to predict the next position of a moving object. Besides, we use the multi-level structure to represent the hierarchical relations among sensors so as to achieve the goal of keeping tracking of moving objects in real-time. Through experimental evaluation on various simulation conditions, the proposed method is shown to deliver excellent performance in terms of both energy efficiency and timeliness. Vincent S. Tseng, Eric Hsueh-Chan Lu, Kawuu Weicheng Lin |
RTCSA | 1 |
| 2005 | CBS: A New Classification Method by Using Sequential PatternsabstractData classification is an important topic in data mining field due to the wide applications. A number of related methods have been proposed based on the well-known learning models like decision tree or neural network. However, these kinds of classification methods may not perform well in mining time sequence datasets like time-series gene expression data. In this paper, we propose a new data mining method, namely Classify-By-Sequence (CBS), for classifying large time-series datasets. The main methodology of CBS method is to integrate the sequential pattern mining with the probabilistic induction such that the inherent sequential patterns can be extracted efficiently and the classification task be done more accurately. Meanwhile, CBS method has the merit of simplicity in implementation. Through experimental evaluation, the CBS method is shown to outperform other methods greatly in the classification accuracy. Vincent S. Tseng, Chao-Hui Lee |
SDM | 1 |
| 2005 | On the security of some proxy blind signature schemes
Bin-Tsan Hsieh, Vincent S. Tseng |
J. Syst. Softw. | 3 |
| 2005 | Efficiently Mining Gene Expression Data via a Novel Parameterless Clustering MethodabstractClustering analysis has been an important research topic in the machine learning field due to the wide applications. In recent years, it has even become a valuable and useful tool for in-silico analysis of microarray or gene expression data. Although a number of clustering methods have been proposed, they are confronted with difficulties in meeting the requirements of automation, high quality, and high efficiency at the same time. In this paper, we propose a novel, parameterless and efficient clustering algorithm, namely, Correlation Search Technique (CST), which fits for analysis of gene expression data. The unique feature of CST is it incorporates the validation techniques into the clustering process so that high quality clustering results can be produced on the fly. Through experimental evaluation, CST is shown to outperform other clustering methods greatly in terms of clustering quality, efficiency, and automation on both of synthetic and real data sets. Vincent S. Tseng, Ching-Pin Kao |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2004 | A Multi-Information based Gene Scoring Model with Applications on Analysis of Hepatocellular CarcinomaabstractHepatitis B virus (HBV) infection is a major health problem worldwide, with more than 1 million people died each year from liver cirrhosis and hepatocellular carcinoma (HCC). The recent development of DNA microarray technology allows us to simultaneously analyze thousands of genes that are differentially expressed in different clinical status. However, the large amount of data has far exceeded our human ability for comprehension without powerful analysis tools. We proposed a methodology to score genes based on the microarray expressions, with the aim to extract interesting genes related to targeted diseases. The methodology consists of preprocessing steps and multi-information based gene scoring methods that can rank the genes according to the degree of relevance to the analysis target. The proposed methodology is applied for analysis of liver cirrhosis and hepatocellular carcinoma (HCC). The experimental results show that our approach has high predictive power through the assessment of QRT-PRC results. Hsieh-Hui Yu, Vincent S. Tseng, Jiin-Haur Chuang |
BIBE | 2 |
| 2004 | An Efficient Approach for Partial-Sum Queries in Data Cubes Using Hamming-Based Codes
Chien-I Lee, Yu-Chiang Li, Vincent S. Tseng |
DASFAA | 3 |
| 2004 | A Novel Parameter-Less Clustering Method for Mining Gene Expression Data
Vincent S. Tseng, Ching-Pin Kao |
PAKDD | 1 |
| 2004 | Mining multilevel and location-aware service patterns in mobile web environmentsabstractIn this correspondence, we address the issue of efficiently mining multilevel and location-aware associated service patterns in a mobile web environment. In terms of multilevel concept, we consider the complex problem that locations and services are of hierarchical structures. We propose a new data mining method named two-dimensional multilevel (2-DML) association rules mining, which can efficiently discover the associated service request patterns by taking into account the multilevel properties of locations and services. The discovered patterns can be effectively utilized in real applications like location-based and personalized services. To the best of our knowledge, this is the first work addressing this research issue. Some variations of the 2-DML method with different properties in terms of execution efficiency and memory efficiency were also developed. Through empirical evaluation, the proposed methods are shown to deliver good performance in terms of efficiency and scalability under various system conditions. Vincent S. Tseng, Ching-Fu Tsui |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2002 | Efficiently Mining Gene Expression Data via Integrated Clustering and Validation Techniques
Vincent S. Tseng, Ching-Pin Kao |
PAKDD | 1 |
| 2001 | Scheduling value-based transactions in distributed real-time database systemsabstractMany noticeable studies have focussed on scheduling flat transactions in a distributed real-time database system (RTDBS). However, a nested transaction model has been widely adopted in many real-life applications such as Internet stock trading systems and telecommunications. This work concerns efficiently scheduling real-time nested transactions in a distributed RTDBS. A new real-time scheduler called flexible high reward for nested transactions (FHRN) is proposed. FHRN consists of (1) FHRNp1 policy to schedule real-time nested transactions and (2) 2PL_HPN to resolve the concurrent data-accessing problem among interleaved nested transactions. Simulation results show that FHRN outperforms these existent real-time schedulers such as random priority (RP), earliest deadline (ED), highest value (HV), hierarchical earliest deadline (HED), and highest reward and urgency (HRU) when an application requires a nested transaction model. Hong-Ren Chen, Yeh-Hao Chin, Vincent S. Tseng |
IPDPS | 3 |
| 2000 | Efficient mining of categorized association rules in large databasesabstractA number of studies have been made on discovering association rules in a large database due to the wide applications. The common goal of the studies focused on finding the associated occurrence patterns between all items in a database. In practice, mining the association rules with the granularity as fine as a single item could result in a huge number of rules that are too large to utilize efficiently. In practical applications, the users may be more interested in the associations between the categories the items belong to. In this paper, we propose a new method for mining categorized association rules efficiently by using compressed feature vectors. With the proposed method, at most one scan of the database is needed to produce the categorized association rules in each user query, even under different mining parameters. Furthermore, the calculation time during the mining process is also reduced greatly by using only simple logic operations on feature vectors. Hence, the overall performance in mining categorized association rules could be improved substantially. Vincent S. Tseng |
SMC | 1 |
| 1998 | Efficient Mining of Association Rules with Item Constraints
Vincent S. Tseng |
Discovery Science | 1 |