EDBT 2026 Demo / reviewers in the wild / expert
Shen-Shyang Ho
dblp:60/6527
· DBLP profile ↗
54ranked-venue papers
18as first author
14since 2021 · last 2026
0000-0002-0353-7159ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 15 first-author · 5 since 2021Databases, data management, data science and information retrieval · 16 · 6 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-authorSystems, architecture and hardware · 3 · 3 since 2021Computer networks · 2 · 1 since 2021Security and privacy · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Detecting and explaining structural changes in an evolving graph using a martingaleabstractDynamic systems such as sensor networks, social networks, computer networks, and power grids can be modeled as evolving graphs, requiring effective methods to monitor structural changes in real-time. In this paper, we present a change detection framework for detecting global structural changes in evolving graphs. This framework combines multiple martingale-based change detectors using different graph features. We establish a mathematical relationship between an additive martingale (derived from the multiple graph features) and the Shapley Additive Explanations (SHAP) method, demonstrating that martingale values at detected change-points directly quantify each feature’s contribution to the detected change. Thus, graph features with significantly high martingale values provide meaningful explanations for detected structural changes. We demonstrate our approach using three synthetic graph types (random topology, scale-free, and small-world), showing how different graph features vary in detection performance across network types. This highlights the importance of using multiple graph features, as relying on any single feature might not detect certain changes. Additionally, we apply our approach to the MIT Reality dataset to identify changes in social interactions among dormitory students at some special time periods. Our results reveal which graph features best characterize the real-world network changes during these special time periods. Shen-Shyang Ho, Tarun Teja Kairamkonda, Izhar Ali |
Pattern Recognit. | 1 |
| 2025 | Early Detection and Attribution of Structural Changes in Dynamic NetworksabstractDynamic networks evolve through structural changes that significantly impact their functionality. Early detection of these changes remains a fundamental challenge due to the inherent tradeoffs between detection delay and false alarm control. Current methods accumulate evidence of these changes only after they occur, creating unavoidable detection delays that limit practical utility in time-sensitive applications. We address this limitation through horizon martingales that leverage statistical forecasting to accumulate evidence from predicted future network states, enabling detection before changes fully manifest. We establish two fundamental results: (1) horizon martingales preserve the martingale property under proper forecasting calibration, and (2) they maintain rigorous false alarm control with probability bounds of$1/\lambda$for detection threshold$\lambda$. Comprehensive evaluation across diverse synthetic network types and real-world social interaction data demonstrates consistent improvements in detection speed with 13–25% reductions in detection delay while maintaining strict statistical guarantees. Izhar Ali, Shen-Shyang Ho |
ICDM | 2 |
| 2024 | Poster: A Hybrid-Cloud Autoencoder Ensemble Method for BotNets Detection on Edge DevicesabstractWe propose a lightweight unsupervised hybrid-cloud ensemble anomaly detection system. We utilize transfer learning to create a model that uses multiple IoT device sources to create a generalized model that requires minimal training to learn new network traffic. These devices feed their output to the cloud enabling more computation while keeping the network traffic secure on the device itself maintaining data privacy. We test this system by creating a simulation testbed to conduct attacks on the IoT Devices to evaluate how well the detection system works. We also compare multiple transfer learned sources to a single source to show how the learning of a target device is impacted. Steven E. Arroyo, Shen-Shyang Ho |
ICFEC | 2 |
| 2024 | Poster: SplitTracer: A Cooperative Inference Evaluation Toolkit for Computation Offloading on the EdgeabstractSplitTracer is an experimental test-bed we are developing to evaluate the efficacy of computation offloading for cooperative inference between edge and fog devices without indepth architectural changes to the models under test for any user-defined application scenario. We describe our test-bed design, functionalities, and support for architectures with splits that require input of multiple layers such as MobileNet and YOLO. Nicholas Bovee, Stephen Piccolo, Suraj Bitla, Gopi Krishna Patapanchala, Shen-Shyang Ho |
ICFEC | 5 |
| 2024 | Poster: Computation Offloading for Precision Agriculture using Cooperative InferenceabstractPrecision Agriculture leverages technology and data-driven machine learning methods for optimized utilization of farm inputs. A challenge for such technology is the need for fast computational resources and limited power supply that translate to expensive computing equipment for the farmers. We propose the use of a cooperative inference framework to supporting edge devices with limited power and computational resources on agricultural tasks. Our preliminary results demonstrate the feasibility of applying computation offloading using cooperative inference for Precision Agriculture. Nicholas Bovee, Paolo Rommel Sanchez, Shen-Shyang Ho, Suraj Bitla, Gopi Krishna Patapanchala, Stephen Piccolo |
ICFEC | 3 |
| 2023 | Experimental Test-Bed for Computation Offloading for Cooperative Inference on Edge DevicesabstractIn this paper, we describe the experimental test-bed we are developing to evaluate the efficacy of computation offloading for cooperative inference without in-depth architectural changes to the models under consideration. We describe our test-bed design, functionalities, and work in progress features. We demonstrate a simple use case of the test-bed, consisting of the identification of a split layer in an AlexNet for a particular hardware scenario and the validation of those results. Nicholas Bovee, Stephen Piccolo, Shen-Shyang Ho, Ning Wang 0018 |
SEC | 3 |
| 2023 | Distributed Tracking and Verifying: A Real-Time and High-Accuracy Visual Tracking Edge Computing Framework for Internet of ThingsabstractWe observe that accurate and fast tracking in Internet of Things (IoT) devices is still a challenging problem. Several deep learning models have emerged which provide higher accuracy scores in object detection and tracking, however, due to their computationally expensive nature they are not useful in enabling real-time tracking at IoT devices. Correlation filters have emerged to show better speed in real-time tracking and provide good tracking results in cases of occlusion, rotation, illumination and other distractions. To get better speed as well as accuracy we use combination of correlation filter and deep learning methods. We propose a distributed tracking and verifying (DTAV) framework. Specifically, we run two object tracking algorithms, one on the client and another on the server. The algorithm run on the client is referred to as the Tracker, which is based on correlation filter and runs easily in real-time. The server hosts the verifier algorithm which performs high accuracy verification. Thus, while the client performs fast object tracking, the server's tracking algorithm verifies the output and corrects the server whenever required to maintain the accuracy of the model. We present our edge computing-based framework and discuss the motivation, system setup and series of experiments performed for the framework and present our experimental results. DTAV achieved 7.78% improvement on accuracy and 15% improvement in FPS. Purva Makarand Mhasakar, Kevin Bhadresh Doshi, Ning Wang 0018, Shen-Shyang Ho, Haibin Ling |
SEC | 4 |
| 2023 | Personalized Learning with Limited Data on Edge Devices Using Federated Learning and Meta-LearningabstractThe efficient and effective handling of few-shot learning tasks on mobile devices is challenging due to the small training set issue and the physical limitations in power and computational resources on these devices. We propose a framework that combines federated learning and meta-learning to handle independent few-shot learning tasks on multiple devices. In particular, we utilize the Prototypical Networks to perform meta-learning on all devices to learn multiple independent few-shot learning models and to aggregate the device models using federated learning which can be reused by the devices subsequently. We perform extensive experiments to (1) compare three different federated learning approaches, namely Federated Averaging (FedAvg), Federated Proximal (FedProx), and Federated Personalization (FedPer) on the proposed framework, and (2) investigate the effect of data heterogeneity issue on multiple devices on their few-shot learning performance. Our empirical results show that our proposed framework is feasible and is able to improve the devices' individual prediction performance and significant performance improvement using the aggregated model using any of the federated learning approaches when the few-shot learning tasks are from the same source and data heterogeneity continues to be a challenging issue to overcome. Kousalya Soumya Lahari Voleti, Shen-Shyang Ho |
SEC | 2 |
| 2022 | Hierarchical Bitmap Indexing for Range Queries on Multidimensional Arrays
Lubos Krcál, Shen-Shyang Ho, Jan Holub 0001 |
DASFAA (1) | 2 |
| 2022 | Autoencoder Ensemble Method for Botnets Detection on IOT DevicesabstractLike anything else on the internet, IoT devices are very susceptible to cyber-attacks that could take out the device or install spyware. In this paper, we propose an anomaly detection solution driven by an autoencoder ensemble to detect botnets on IOT devices. In particular, the ensemble size is determined by hierarchical clustering of the features in the packet header. Moreover, one does not require an additional neural network to combine the decisions. The proposed approach is a more efficient solution for IOT problem setting and hence, overcomes the issue of lacking computational resources and memory on IOT devices, as well as run-time performance problems. Empirical results on two datasets, one from the 2016 Mirai botnet attacks on IoT devices and the other from Gafgyt malware attacks on various IOT devices, show the competitiveness and feasibility of our proposed solution. Steven E. Arroyo, Shen-Shyang Ho |
ICMLA | 2 |
| 2022 | Representation recovery via L1-norm minimization with corrupted data
Woon Huei Chai, Shen-Shyang Ho, Hiok Chai Quek |
Inf. Sci. | 2 |
| 2022 | A Novel Quasi-Newton Method for Composite Convex Minimization
Woon Huei Chai, Shen-Shyang Ho, Hiok Chai Quek |
Pattern Recognit. | 2 |
| 2021 | Energy Demand Prediction with Optimized Clustering-Based Federated LearningabstractThe rapid growth in pervasive Internet-of-Things (IoT) and Deep Learning (DL) is creating a huge demand for applying DL on IoT systems. However, it is non-trivial to train highly accurate DL models in such scenarios due to the following two challenges: (1) individual IoT devices may not have sufficient training data, and (2) simply combining all sensory data across all devices may cause performance degradation due to data imbalance and varying temporal patterns across different devices. The objective of this paper is to achieve high-accurate prediction models for each device in an IoT system. We propose a federated learning approach for IoT systems driven by trend-based clustering for energy demand prediction for Electric Vehicle (EV) charging station network. We first apply a time-series clustering method to identify stations with similar temporal demand patterns. Using time-series data from stations in a cluster, a single Long-Short Term Memory (LSTM) network is trained using FedAvg algorithm for energy demand prediction for all the stations in the cluster. Experimental results on a real-world energy usage dataset from an EV charging station network show that our proposed approach is very competitive against baseline federated learning approaches. In particular, the energy demand prediction error decreases by 80%. Dylan Perry, Ning Wang 0018, Shen-Shyang Ho |
GLOBECOM | 3 |
| 2021 | Ensemble Learning using Error Correcting Output Codes: New Classification Error BoundsabstractNew bounds on classification error rates for the error-correcting output code (ECOC) approach in machine learning are presented. These bounds have exponential decay complexity with respect to codeword length and theoretically validate the effectiveness of the ECOC approach. Bounds are derived for two different models: the first assumes that all base classifiers are independent and the second assumes that all base classifiers are mutually correlated up to first-order. Moreover, we perform ECOC classification on six datasets and compare their error rates with our bounds to experimentally validate our work and show the effect of correlation on classification accuracy. Hieu D. Nguyen, Mohammed Sarosh Khan, Nicholas Kaegi, Shen-Shyang Ho, Jonathan Moore, Logan Borys, Lucas Lavalva |
ICTAI | 4 |
| 2020 | An Error-Correcting Output Code Framework for Lifelong Learning without a TeacherabstractAn intelligent system should learn new concepts continuously and autonomously. The system should recognize that a concept is new and learn the concept without any guidance. In this paper, a novel error correcting output code (ECOC)-based framework, motivated by the complementary learning systems (CLS) theory, is proposed to perform (i) rapid detection of a new concept and (ii) learning and storing of the new concept via reinforced encoding. Experimental results on six datasets show the feasibility and competitive performance of the proposed ECOC-based framework for life-long learning using four different base classifiers against two baseline approaches. Moreover, we demonstrate the performance of our proposed ECOC-based framework on a continual learning scenario without any label feedback using a realistic but stringent cumulative performance measure, which combines detection error and classification error. Shen-Shyang Ho, Mathew Marchiano, Scott Zockoll, Hieu D. Nguyen |
ICTAI | 1 |
| 2020 | Assessing Accident Risk using Ordinal Regression and Multinomial Logistic Regression Data GenerationabstractRobust and accurate modeling of motor vehicle accident and injury severities have significant impact on transportation safety and economy. The capability to assess accident risk based on external driving conditions (e.g., weather, road condition, etc.) and driver behavior and characteristics can reduce accident occurrences by alerting drivers to alleviated risk. In this paper, we propose a novel accident risk assessment framework driven by ordinal regression. One challenge of the risk assessment problem is that non-accident data are not collected by any agency in their study of transportation safety. Hence, we also propose a realistic negative data generation scheme based on feature weighs derived from multinomial logistic regression to overcome this challenge. Experimental results on two different real-world datasets from the US National Highway Traffic Safety Administration and UK Transport for Greater Manchester are used to demonstrate the feasibility and robustness of our proposed ordinal regression framework. Performance on four ordinal regression algorithms, namely: logistic all-threshold, logistic immediate-threshold, ordinal ridge, and least absolute deviations are compared. In addition, for US dataset, we investigate the effect of random oversampling and undersampling on the proposed risk assessment framework. We empirically show that bagging with random oversampling using logistic all-threshold ordinal regression method has the best prediction performance among ordinal regression models. Gulsum Alicioglu, Shen-Shyang Ho |
IJCNN | 3 |
| 2019 | A Martingale-Based Approach for Flight Behavior Anomaly DetectionabstractThe timely detection of anomalous flight behavior is critical to ensure a prompt and appropriate response to mitigate any dangers to flight safety or hindrance of logistics operations. Most previous approaches focused on anomaly detection, leading them to only be able to raise an alert after an occurrence of an anomaly. A more effective approach is to predict a potential anomaly based on current observations, thus cutting down on detection time and allowing for a more expedient response. We propose a novel martingale-based approach to predict anomalous flight behavior in the near future as data points are observed one by one in real-time. The proposed anomaly prediction method consists of two components: (i) utilization of regression to model the historical full flight behavior and (ii) monitoring of the real-time flight behavior using a martingale (stochastic) process. The latter component consists of two prediction steps: (i) first to predict future values of multiple target variables (e.g., latitude, longitude, and altitude) using regression models, and (ii) then to decide whether the predicted values exhibit anomalies. In particular, our proposed method uses martingale tests on multiple Gaussian process regression (GPR) predictive models of target variables. The main advantages of the proposed method are: (i) the use of multiple martingale tests allows one to have a tighter false positive bound for anomaly detection/prediction, and (ii) the prediction steps reduce the delay time for anomaly detection. Experimental results on real-world data show that the performance (mean delay time, recall, and precision) of our proposed approach is competitive against other compared methods. Shen-Shyang Ho, Matthew Schofield, Jason Snouffer, Jean Kirschner |
MDM | 1 |
| 2019 | Improving Bayesian network local structure learning via data-driven symmetry correction methods
Shen-Shyang Ho |
Int. J. Approx. Reason. | 2 |
| 2019 | N-ary decomposition for multi-class classification
Joey Tianyi Zhou, Ivor W. Tsang, Shen-Shyang Ho, Klaus-Robert Müller |
Mach. Learn. | 3 |
| 2018 | Visual Analytics for Real-Time Flight Behavior Threat AssessmentabstractWe propose integrating data visualization and machine learning techniques to support Air Traffic Controller (ATC) systems for detecting and identifying friendly/unfriendly aircraft. Our platform composes data-driven decisions to optimize strategic, and operative elements of an ATC system and mitigates its drawbacks by analyzing real-time data from a radar system. A threat assessment approach that incorporates flight behavior assessment, based on visualizing flight data together with flight anomaly prediction and flight origin/destination prediction can be used in cases where the system fails. Furthermore, our proposed integrated tool lets users quickly identify flights and security concerns by analyzing and visualizing data. Bo Beth Sun, Eric Zielonka, Aleksandr Fritz, Matthew Schofield, Brennan Ringel, Brendan Armstrong, Shen-Shyang Ho, Anthony F. Breitzman, Jason Snouffer, Jean Kirschner, Kimberly Davis |
IEEE BigData | 7 |
| 2017 | Structural knowledge transfer for learning Sum-Product Networks
Shen-Shyang Ho |
Knowl. Based Syst. | 2 |
| 2016 | Transfer Learning for Cross-Language Text Categorization through Active Correspondences ConstructionabstractMost existing heterogeneous transfer learning (HTL) methods for cross-language text classification rely on sufficient cross-domain instance correspondences to learn a mapping across heterogeneous feature spaces, and assume that such correspondences are given in advance. However, in practice, correspondences between domains are usually unknown. In this case, extensively manual efforts are required to establish accurate correspondences across multilingual documents based on their content and meta-information. In this paper, we present a general framework to integrate active learning to construct correspondences between heterogeneous domains for HTL, namely HTL through active correspondences construction (HTLA). Based on this framework, we develop a new HTL method. On top of the new HTL method, we further propose a strategy to actively construct correspondences between domains. Extensive experiments are conducted on various multilingual text classification tasks to verify the effectiveness of HTLA. Joey Tianyi Zhou, Sinno Jialin Pan, Ivor W. Tsang, Shen-Shyang Ho |
AAAI | 4 |
| 2016 | Exploiting sparsity for image-based object surface anomaly detectionabstractThe anomaly detection task plays an important role in quality control in many industrial or manufacturing processes. However, in many such processes, anomaly detection is done visually by human experts who have in-depth knowledge and vast experience on a product in order to perform well in the detection task. In this paper, we present an approach that (i) identifies anomalies in an image based on the sparse residuals (or errors) obtained during image reconstruction using sparse representation and (ii) learns the threshold to classify an image pixel based on its residual value. The intuitions for our proposed sparse approximation driven approach are, namely: (i) anomalies are infrequent and (ii) anomalies are unwanted portions of an image reconstruction. Empirical results on a real-world image dataset for an industrial surface defect detection task are used to demonstrate the feasibility of our proposed approach. Woon Huei Chai, Shen-Shyang Ho, Chi Keong Goh |
ICASSP | 2 |
| 2016 | Is overfeat useful for image-based surface defect classification tasks?abstractOne of the challenges for real-world image-based surface defect classification task is the lack of labeled training samples to extract useful features to confidently classify defects. In this paper, we present results on our investigation on whether features derived from OverFeat, a variant of Convolution Neural Network, can be used directly for image-based surface defect classification task. We show that the classification performance of two real-world defect images datasets can be significantly different. For the harder classification task, OverFeat features are useful for some types of surface defects, but performs poorly when the defects demonstrate characteristics beyond texture patterns. We propose a simple heuristic approach called Approximate Surface Roughness (ASR) that provides auxiliary information on the relationship between spatial regions in the defect image to be used together with the OverFeat features. Empirical results show improvement in classification performance for those defect types that do not classify well using only OverFeat features. Pei-Hung Chen, Shen-Shyang Ho |
ICIP | 2 |
| 2016 | ParkGauge: Gauging the Occupancy of Parking Garages with Crowdsensed Parking CharacteristicsabstractFinding available parking spaces in dense urban areas is a globally recognized issue in urban mobility. Whereas prior studies have focused on outdoor/street parking due to a common belief that parking garages are capable of delivering real-time occupancy information, we specifically target at (indoor) parking garages as this belief is far from true. This problem is very challenging as all the infrastructure supports (e.g., GPS and Wi-Fi) assumed by existing proposals are not available to parking garages, so counting how many vehicles are using a parking garage by crowd sensing can be extremely difficult. To this end, we present Park Gauge, a method to gauge the occupancy of parking garages, along with a reference system prototype for performance evaluation, it infers parking occupancy from crowd sensed parking characteristics instead of counting the parked vehicles. Park Gauge adopts low-power sensors (e.g., accelerometer and barometer) in the driver's smartphone to determine the driving states (e.g., turning and braking). A sequence of such states further allows the inference of driving contexts (e.g., driving, queuing and parked) that in turn yield temporal parking characteristics of a parking garage, including time-to-park and time-in-cruising/queuing. Mining such mobile data opportunistically collected from a crowd of drivers arriving at various garages yields a good measure of their occupancies and hence useful recommendations can be generated (in real-time) to inform drivers coming toward these venues. Through extensive experiments, we demonstrate that our method fully explores these parking characteristics to efficiently infer occupancies of parking garages with high accuracy. Jim Cherian, Jun Luo 0001, Shen-Shyang Ho, Richard Wisbrun |
MDM | 4 |
| 2016 | Robust energy-based least squares twin support vector machines
Muhammad Tanveer 0001, Mohammad Asif Khan 0001, Shen-Shyang Ho |
Appl. Intell. | 3 |
| 2016 | An efficient regularized K-nearest neighbor based weighted twin support vector regression
Muhammad Tanveer 0001, K. Shubham, Mujahed Aldhaifallah, Shen-Shyang Ho |
Knowl. Based Syst. | 4 |
| 2016 | Manifold Learning for Multivariate Variable-Length Sequences With an Application to Similarity SearchabstractMultivariate variable-length sequence data are becoming ubiquitous with the technological advancement in mobile devices and sensor networks. Such data are difficult to compare, visualize, and analyze due to the nonmetric nature of data sequence similarity measures. In this paper, we propose a general manifold learning framework for arbitrary-length multivariate data sequences driven by similarity/distance (parameter) learning in both the original data sequence space and the learned manifold. Our proposed algorithm transforms the data sequences in a nonmetric data sequence space into feature vectors in a manifold that preserves the data sequence space structure. In particular, the feature vectors in the manifold representing similar data sequences remain close to one another and far from the feature points corresponding to dissimilar data sequences. To achieve this objective, we assume a semisupervised setting where we have knowledge about whether some of data sequences are similar or dissimilar, called the instance-level constraints. Using this information, one learns the similarity measure for the data sequence space and the distance measures for the manifold. Moreover, we describe an approach to handle the similarity search problem given user-defined instance level constraints in the learned manifold using a consensus voting scheme. Experimental results on both synthetic data and real tropical cyclone sequence data are presented to demonstrate the feasibility of our manifold learning framework and the robustness of performing similarity search in the learned manifold. Shen-Shyang Ho, Peng Dai 0002, Frank Rudzicz |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2015 | Stochastic location optimization in a dynamic environmentabstractExisting location optimization solutions only consider the positioning of resouces/seeds at the best location in a target area by minimizing a certain metric over distance. In reality, what really matters is time. In this paper, the location optimization problem is formulated as the expected response time minimization problem rather than a distance minimization problem. Moreover, we propose an algorithm that takes into consideration various stochastic factors which affect the location optimization problem, such as non-uniform probability distribution of the demands, road congestion level, and vehicles' maximum speed. Our proposed algorithm shows promising performance when the disparity among vehicle capabilities (e.g., maximum speed) are large and the environment constraints (e.g., traffic jam) are taken into consideration. Yubo Dong, Shen-Shyang Ho |
SIGSPATIAL/GIS | 3 |
| 2015 | Poster: ParkGauge: Gauging the Congestion Level of Parking Garages with Crowdsensed Parking CharacteristicsabstractFinding available parking spaces in dense urban areas is a globally recognized issue in urban mobility. Whereas prior studies have focused on outdoor/street parking, we target at (indoor) parking garages where the infrastructure supports (e.g., GPS and Wi-Fi) assumed by existing proposals are unavailable and counting vehicles by crowdsensing is difficult. To this end, we present ParkGauge as a system gauging the congestion level of parking garages; it infers (coarse-grained) parking occupancy from crowdsensed parking characteristics instead of counting the parked vehicles. ParkGauge adopts mostly low-power sensors in the driver's smartphone to determine driving states, contexts and temporal parking characteristics of a garage, including time-to-park and time-in-cruising/queuing. Mining such data collected from a crowd of drivers at various garages yields a good measure of their congestion levels and provide recommendations (in real-time) to drivers coming to these venues. Jim Cherian, Jun Luo 0001, Shen-Shyang Ho, Richard Wisbrun |
SenSys | 4 |
| 2015 | Sequential behavior prediction based on hybrid similarity and cross-user activity transfer
Peng Dai 0002, Shen-Shyang Ho, Frank Rudzicz |
Knowl. Based Syst. | 2 |
| 2015 | ML-TREE: A Tree-Structure-Based Approach to Multilabel LearningabstractMultilabel learning aims to predict labels of unseen instances by learning from training samples that are associated with a set of known labels. In this paper, we propose to use a hierarchical tree model for multilabel learning, and to develop the ML-Tree algorithm for finding the tree structure. ML-Tree considers a tree as a hierarchy of data and constructs the tree using the induction of one-against-all SVM classifiers at each node to recursively partition the data into child nodes. For each node, we define a predictive label vector to represent the predictive label transmission in the tree model for multilabel prediction and automatic discovery of the label relationships. If two labels co-occur frequently as predictive labels at leaf nodes, these labels are supposed to be relevant. The amount of predictive label co-occurrence provides an estimation of the label relationships. We examine the ML-Tree method on 11 real data sets of different domains and compare it with six well-established multilabel learning algorithms. The performances of these approaches are evaluated by 16 commonly used measures. We also conduct Friedman and Nemenyi tests to assess the statistical significance of the differences in performance. Experimental results demonstrate the effectiveness of our method. Qingyao Wu, Yunming Ye, Haijun Zhang 0002, Tommy W. S. Chow, Shen-Shyang Ho |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2014 | Traffic incident validation and correlation using text alerts and imagesabstractOne of the major challenges during the process of extracting information from multiple spatio-temporal data sources of diverse data types is the matching and fusion of extracted knowledge (e.g. interesting nearby events detected from text, estimated density or flow from a set of geo-coded images). In this demonstration, we present PETRINA ("PErsonalized TRaffic INformation Analytics"), a system that provides traffic-related incident monitoring, mapping, and analytics services. In particular, we showcase two main functionalities: (1) text traffic alert validation based on traffic condition information derived from traffic camera images and (2) traffic incident correlation based on spatio-temporal proximity of different incident types (e.g., accidents and heavy traffic). Despite the fact that the images are sparse (available every three minutes), the regularity makes it possible to validate whether a text traffic alert is outdated or not, and to more accurately estimate the time elapsed and total incident time. Multiple traffic incidents can be grouped together as a single event based on the traffic incident correlation to reduce information redundancy. Such enhanced real-time traffic information enables PETRINA to offer services such as dynamic routing with traffic incident advices, spatiotemporal traffic incident visual analytics, and congestion analysis. Wye Huong Yan, Justin Ong, Shen-Shyang Ho, Jim Cherian |
SIGSPATIAL/GIS | 3 |
| 2014 | Robust prediction in nearly periodic time series using motifsabstractIn this paper, we consider the prediction task for a process with nearly periodic property, i.e., patterns occur with some regularities but no exact periodicity. We propose an inference approach based on probabilistic Markov framework utilizing motif-driven transition probabilities for sequential prediction. In particular, a Markov-based weighting framework utilizing fully the information from recent historical data and sequential pattern regularities is developed for nearly periodic time series prediction. Preliminary experimental results show that our prediction approach is competitive against the moving average and multi-layer perceptron neural network approaches on synthetic data. Moreover, our proposed method is shown to be empirically robust on time-series with missing data and noise. We also demonstrate the usefulness of our proposed approach on a real-world vehicle parking lot availability prediction task. Woon Huei Chai, Shen-Shyang Ho |
IJCNN | 3 |
| 2014 | A Smartphone User Activity Prediction Framework Utilizing Partial Repetitive and Landmark BehaviorsabstractIn this paper, we propose a general smartphone user activity prediction framework utilizing the general concept of partial repetitive behavior (instead of the stronger periodicity condition) for similarity scoring and the landmark behaviors (representative behaviors to identify groups of similar behavior vectors). Prediction of the next-day(s) behavior is based on a weighted sum of the most similar behavior vectors related to the landmark behavior of the next-day(s) behavior. These behavior vectors are selected based on the likely partial repetition of the next-day behavior and similarity in the eigen behavior feature space. Our proposed prediction algorithm allows one to categorically quantify the frequency of a target behavior, such as no behavior, normal behavior, and high frequency behavior, or other more refined categorization based on user preference. Extensive experiments are carried out using the Nokia Mobile Data Challenge (MDC) dataset to demonstrate the feasibility of our proposed approach and its generality using arbitrary call activity, voice call activity, short message activity, media consumption, and apps usage data types. Peng Dai 0002, Shen-Shyang Ho |
MDM (1) | 2 |
| 2014 | A Generative Model with Network Regularization for Semi-Supervised Collective ClassificationabstractIn recent years much effort has been devoted to Collective Classification (CC) techniques for predicting labels of linked instances. Given a large number of labeled data, conventional CC algorithms make use of local labeled neighbours to increase accuracy. However, in many real-world applications, labeled data are limited and very expensive to obtain. In this situation, most of the data have no connection to labeled data, and supervision knowledge cannot be obtained from the local connections. Recently, Semi-Supervised Collective Classification (SSCC) has been examined to leverage unlabeled data for enhancing the classification performance of CC. In this paper we propose a probabilistic generative model with network regularization (GMNR) for SSCC. Our main idea is to compute label probability distributions for unlabeled instances by maximizing both the log-likelihood in the generative model and the label smoothness on the network topology of data. The proposed generative model is based on the Probabilistic Latent Semantic Analysis (PLSA) method using attribute features of all instances. A network regularizer is employed to smooth the label probability distributions on the network topology of data. Finally, we develop an effective EM algorithm to compute the label probability distributions for label prediction. Experimental results on three real sparsely-labeled network datasets show that the proposed model GMNR outperforms state-of-the-art CC algorithms and other SSCC algorithms. Ruichao Shi, Qingyao Wu, Yunming Ye, Shen-Shyang Ho |
SDM | 4 |
| 2014 | Collective prediction of protein functions from protein-protein interaction networksabstractBACKGROUND: Automated assignment of functions to unknown proteins is one of the most important task in computational biology. The development of experimental methods for genome scale analysis of molecular interaction networks offers new ways to infer protein function from protein-protein interaction (PPI) network data. Existing techniques for collective classification (CC) usually increase accuracy for network data, wherein instances are interlinked with each other, using a large amount of labeled data for training. However, the labeled data are time-consuming and expensive to obtain. On the other hand, one can easily obtain large amount of unlabeled data. Thus, more sophisticated methods are needed to exploit the unlabeled data to increase prediction accuracy for protein function prediction. RESULTS: In this paper, we propose an effective Markov chain based CC algorithm (ICAM) to tackle the label deficiency problem in CC for interrelated proteins from PPI networks. Our idea is to model the problem using two distinct Markov chain classifiers to make separate predictions with regard to attribute features from protein data and relational features from relational information. The ICAM learning algorithm combines the results of the two classifiers to compute the ranks of labels to indicate the importance of a set of labels to an instance, and uses an ICA framework to iteratively refine the learning models for improving performance of protein function prediction from PPI networks in the paucity of labeled data. CONCLUSION: Experimental results on the real-world Yeast protein-protein interaction datasets show that our proposed ICAM method is better than the other ICA-type methods given limited labeled training data. This approach can serve as a valuable tool for the study of protein function prediction from PPI networks. Qingyao Wu, Yunming Ye, Michael Kwok-Po Ng, Shen-Shyang Ho, Ruichao Shi |
BMC Bioinform. | 4 |
| 2014 | ForesTexter: An efficient random forest algorithm for imbalanced text categorization
Qingyao Wu, Yunming Ye, Haijun Zhang 0002, Michael Kwok-Po Ng, Shen-Shyang Ho |
Knowl. Based Syst. | 5 |
| 2013 | Cluster tree based multi-label classification for protein function predictionabstractAutomatically assigning functions for unknown proteins is a key task in computational biology. Proteins in nature have multiple classes according to the functions they perform. Many efforts have been made to cast the protein function prediction into a multi-label learning problem. This paper proposes a novel Cluster Tree based Multi-label Learning algorithm (CTML) for protein function prediction. The main idea is to compute a set of predictive labels associated at each node for multi-label prediction by using the k-means clustering techniques and the predictive functions via the learning data at the nodes. With the propagation of the predictive labels from the root node to the leaf node, the correlations between labels can be preserved. Experimental results on benchmark data (genbase and yeast datasets) show that the proposed CTML algorithm is effective in predicting protein functions. Moreover, the classification performance of the CTML algorithm is competitive against the other baseline multi-label learning algorithms. Qingyao Wu, Yunming Ye, Xiaofeng Zhang 0002, Shen-Shyang Ho |
BIBM | 4 |
| 2012 | Mining multivariate spatiotemporal patterns from heterogeneous mobility dataabstractMobility data mining in the form of trajectory data mining has been extensively investigated in recent years. Predictive modeling and pattern discovery approaches have been proposed to predict movements and locations, and to extract useful trajectory and location patterns. Nowadays, mobility data consist of not only trajectory data. Mobility data from smart phones include measurements such as call duration/time, call type, digital media consumption, calendar information, apps usage, social interactions, and mobile browsing. These heterogeneous multivariate data allow one to discover interesting and more complex behavioral patterns and rules in terms of space and time. Shen-Shyang Ho |
SIGSPATIAL/GIS | 1 |
| 2012 | An effective vortex detection approach for velocity vector field
Shen-Shyang Ho |
ICPR | 1 |
| 2012 | Preserving privacy for moving objects data miningabstractThe prevalence of mobile devices with geopositioning capability has resulted in the rapid growth in the amount of moving object trajectories. These data have been collected and analyzed for both commercial (e.g., recommendation system) and security (e.g. surveillance and monitoring system) purposes. One needs to ensure the privacy of these raw trajectory data and the derived knowledge by not disclosing or releasing them to adversary. In this paper, we propose a practical implementation of a (ε; δ)-differentially private mechanism for moving objects data mining; in particular, we apply it to the frequent location pattern mining algorithm. Experimental results on the real-world GeoLife dataset are used to compare the performance of the (ε; δ)-differential privacy mechanism with the standard ε-differential privacy mechanism. Shen-Shyang Ho |
ISI | 1 |
| 2010 | Tropical cyclone event sequence similarity search via dimensionality reduction and metric learningabstractThe Earth Observing System Data and Information System (EOSDIS) is a comprehensive data and information system which archives, manages, and distributes Earth science data from the EOS spacecrafts. One non-existent capability in the EOSDIS is the retrieval of satellite sensor data based on weather events (such as tropical cyclones) similarity query output. Shen-Shyang Ho, Wenqing Tang, W. Timothy Liu |
KDD | 1 |
| 2010 | A Framework for Moving Sensor Data Query and Retrieval of Dynamic Atmospheric Events
Shen-Shyang Ho, Wenqing Tang, W. Timothy Liu, Markus Schneider 0001 |
SSDBM | 1 |
| 2010 | A Martingale Framework for Detecting Changes in Data Streams by Testing ExchangeabilityabstractIn a data streaming setting, data points are observed sequentially. The data generating model may change as the data are streaming. In this paper, we propose detecting this change in data streams by testing the exchangeability property of the observed data. Our martingale approach is an efficient, nonparametric, one-pass algorithm that is effective on the classification, cluster, and regression data generating models. Experimental results show the feasibility and effectiveness of the martingale methodology in detecting changes in the data generating model for time-varying data streams. Moreover, we also show that: 1) An adaptive support vector machine (SVM) utilizing the martingale methodology compares favorably against an adaptive SVM utilizing a sliding window, and 2) a multiple martingale video-shot change detector compares favorably against standard shot-change detection algorithms. Shen-Shyang Ho, Harry Wechsler |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2009 | Cyclone tracking using multiple satellite image sourcesabstractWe present an automated cyclone tracking system that uses images from multiple satellite sources. The system tracks cyclones using infrared images from a Geostationary Operational Environmental Satellite (GOES), precipitation images derived from five satellite sources, and ocean surface wind field satellite images. The system consists of three main components: (i) data preprocessing steps for each data source, (ii) cyclone eye detection algorithms for each data source, and (iii) a filter-based tracker that integrates the eye detection results from each data source. Experimental results show that our prototype system is operationally feasible and has better performance than our prior cyclone tracking system. Anand V. Panangadan, Shen-Shyang Ho, Ashit Talukder |
GIS | 2 |
| 2008 | Automated cyclone discovery and tracking using knowledge sharing in multiple heterogeneous satellite dataabstractCurrent techniques for cyclone detection and tracking employ NCEP (National Centers for Environmental Prediction) models from in-situ measurements. This solution does not provide true global coverage, unlike remote satellite observations. However it is impractical to use a single Earth orbiting satellite to detect and track events such as cyclones in a continuous manner due to limited spatial and temporal coverage. One solution to alleviate such persistent problems is to utilize heterogeneous sensor data from multiple orbiting satellites. However, this solution requires overcoming other new challenges such as varying spatial and temporal resolution between satellite sensor data, the need to establish correspondence between features from different satellite sensors, and the lack of definitive indicators for cyclone events in some sensor data. Shen-Shyang Ho, Ashit Talukder |
KDD | 1 |
| 2008 | Query by TransductionabstractThere has been recently a growing interest in the use of transductive inference for learning. We expand here the scope of transductive inference to active learning in a stream-based setting. Towards that end this paper proposes Query-by-Transduction (QBT) as a novel active learning algorithm. QBT queries the label of an example based on the p-values obtained using transduction. We show that QBT is closely related to Query-by-Committee (QBC) using relations between transduction, Bayesian statistical testing, Kullback-Leibler divergence, and Shannon information. The feasibility and utility of QBT is shown on both binary and multi-class classification tasks using SVM as the choice classifier. Our experimental results show that QBT compares favorably, in terms of mean generalization, against random sampling, committee-based active learning, margin-based active learning, and QBC in the stream-based setting. Shen-Shyang Ho, Harry Wechsler |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2007 | Confident Identification of Relevant Objects Based on Nonlinear Rescaling Method and Transductive InferenceabstractWe present a novel machine learning algorithm to identify relevant objects from a large amount of data. This approach is driven by linear discrimination based on nonlinear rescaling (NR) method and transductive inference. The NR algorithm for linear discrimination (NRLD) computes both the primal and the dual approximation at each step. The dual variables associated with the given labeled data-set provide important information about the objects in the data-set and play the key role in ordering these objects. A confidence score based on a transductive inference procedure using NRLD is used to rank and identify the relevant objects from a pool of unlabeled data. Experimental results on an unbalanced protein data-set for the drug target prioritization and identification problem are used to illustrate the feasibility of the proposed identification algorithm. Shen-Shyang Ho, Roman A. Polyak |
ICDM | 1 |
| 2007 | Detecting Changes in Unlabeled Data Streams Using Martingale
Shen-Shyang Ho, Harry Wechsler |
IJCAI | 1 |
| 2005 | A martingale framework for concept change detection in time-varying data streamsabstractIn a data streaming setting, data points are observed one by one. The concepts to be learned from the data points may change infinitely often as the data is streaming. In this paper, we extend the idea of testing exchangeability online (Vovk et al., 2003) to a martingale framework to detect concept changes in time-varying data streams. Two martingale tests are developed to detect concept changes using: (i) martingale values, a direct consequence of the Doob’s Maximal Inequality, and (ii) the martingale difference, justified using the Hoeffding-Azuma Inequality. Under some assumptions, the second test theoretically has a lower probability than the first test of rejecting the null hypothesis, “no concept change in the data stream”, when it is in fact correct. Experiments show that both martingale tests are effective in detecting concept changes in time-varying data streams simulated using two synthetic data sets and three benchmark data sets. 1. Shen-Shyang Ho |
ICML | 1 |
| 2005 | Adaptive Support Vector Machine for Time-Varying Data Streams Using Martingale
Shen-Shyang Ho, Harry Wechsler |
IJCAI | 1 |
| 2005 | On the Detection of Concept Changes in Time-Varying Data Stream by Testing Exchangeability
Shen-Shyang Ho, Harry Wechsler |
UAI | 1 |
| 2003 | Transductive confidence machine for active learningabstractThis paper describes a novel active learning strategy using universal p-value measures of confidence based on algorithmic randomness, and transconductive inference. The early stopping criterion for active learning is based on the bias-variance tradeoff for classification. This corresponds to that learning instance when the boundary bias becomes positive, and requires one to switch from active to random selection of learning examples. The sign for the boundary and the increase in the classification error are two manifestations of the same phenomena, i.e., over-training. The experimental results presented show the feasibility and usefulness of our novel approach using a non-separable two-class classification problem. Our hybrid learning strategy achieves competitive performance against standard nearest neighbor methods using much fewer training examples. Shen-Shyang Ho, Harry Wechsler |
IJCNN | 1 |