EDBT 2026 Demo / reviewers in the wild / expert
Wei Ding 0003
dblp:59/622-3
· DBLP profile ↗
58ranked-venue papers in the field
8as first author
13since 2021 · last 2026
0000-0002-3383-551XORCID · conflict
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 40 (6 first)Database Systems & Data Management · 9Big Data, Cloud & Distributed Data Systems · 5Other / Interdisciplinary · 2 (1 first)Information Retrieval & Web Search · 1 (1 first)Knowledge Engineering, Semantic Web & Information Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Transductive Model-Agnostic Contrastive Learning Framework for Few-Shot Learning
Tianyu Kang, Chengjie Zheng, Ping Chen 0001, Wei Ding 0003 |
PAKDD (1) | 4 |
| 2025 | FACT: Gated Fusion-Augmented Causal Mask Transformer for Pseudotime Analysis
Chengjie Zheng, Iris Shen, John Quackenbush, Viola Fanfani, Wei Ding 0003, Ping Chen 0001 |
IEEE Big Data | 6 |
| 2025 | KDD Health Day 2025: Harnessing AI Opportunities in Biomedicine and HealthcareabstractThe ACM KDD 2025 Health Day theme, ''Harnessing AI Opportunities in Biomedicine and Healthcare'' highlights the transformative potential of AI-driven applications in healthcare, translational biomedical research, and basic biological research. This extended abstract discusses recent advancements, challenges, and future directions, focusing on integrating AI-ready data sets, interdisciplinary collaborations, and ethical AI practices. It aims to catalyze discussions on the potential of AI ecosystems in revolutionizing healthcare and related fields. Peipei Ping, Wei Ding 0003, Carl Yang 0001 |
KDD (2) | 2 |
| 2025 | Improving Generalization in Deep Neural Networks by Mitigating Memorization
Yong Zhuang, Tianyu Kang, Wei Ding 0003, Ping Chen 0001 |
PAKDD (2) | 4 |
| 2025 | Horizon Forcing: Improving the Recurrent Forecasting of Chaotic SystemsabstractChaotic dynamics are ubiquitous in many real-world systems, ranging from biological and industrial processes to climate dynamics and the spread of viruses. These systems are characterized by high sensitivity to initial conditions, making it challenging to predict their future behavior confidently. In this study, we propose a novel deep-learning framework that addresses this challenge by directly exploiting the long-term compounding of local prediction errors during model training, aiming to extend the time horizon for reliable predictions of chaotic systems. Our approach observes the future trajectories of initial errors at a time horizon, modeling the evolution of the loss to that point through the use of two major components: (1) a recurrent architecture (Error Trajectory Tracing) designed to trace the trajectories of predictive errors through phase space, and (2) a training regime, Horizon Forcing, that pushes the model’s focus out to a predetermined time horizon. We validate our method on three classic chaotic systems and six real-world time series prediction tasks with chaotic characteristics. The results show that our approach outperforms the state-of-the-art methods. Yong Zhuang, Matthew Almeida, Wei Ding 0003, Ping Chen 0001 |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2024 | SteLLA: A Structured Grading System Using LLMs with RAGabstractLarge Language Models (LLMs) have shown strong general capabilities in many applications. However, how to make them reliable tools for some specific tasks such as automated short answer grading (ASAG) remains a challenge. We present SteLLA (Structured Grading System Using LLMs with RAG) in which a) Retrieval Augmented Generation (RAG) approach is used to empower LLMs specifically on the ASAG task by extracting structured information from the highly relevant and reliable external knowledge based on the instructor-provided reference answer and rubric, b) an LLM performs a structured and question-answering-based evaluation of student answers to provide analytical grades and feedback. A real-world dataset which contains students’ answers in an exam was collected from a college-level Biology course. Experiments show that our proposed system can achieve substantial agreement with the human grader while providing break-down grades and feedback on all the knowledge points examined in the problem. A qualitative and error analysis of the feedback generated by GPT4 shows that GPT4 is good at capturing facts while may prone to inferring too much implication from the given text in the grading task which provides insights into the usage of LLMs in the ASAG system. Hefei Qiu, Ashley Ding, Reinaldo Costa, Ali Hachem, Wei Ding 0003, Ping Chen 0001 |
IEEE Big Data | 6 |
| 2024 | Animal-JEPA: Advancing Animal Behavior Studies Through Joint Embedding Predictive Architecture in Video AnalysisabstractAnalyzing animal behavior from video data is crucial for understanding brain function, assessing pharmacological interventions, and examining genetic modifications. Traditional methods often struggle to accurately analyze group behaviors in complex environments. To address these challenges, we introduce the Animal Joint Embedded Prediction Architecture (Animal-JEPA), a novel self-supervised learning model designed for studying animal behavior from video data. Animal-JEPA leverages a dynamic scaling mechanism and an elliptical masking strategy to enhance feature extraction and behavioral analysis without the need for labeled data. Our approach significantly outperforms existing models, including Separate 3D ConvNet (S3D) [3], Video Vision Transformer (ViViT) [4], and the original V-JEPA [6], particularly in multi-category and multi-objective classification tasks on our newly developed Mice-Behavior3 (MB3) dataset. The results highlight Animal-JEPA’s potential to improve the accuracy and adaptability of behavioral analysis in animal research, providing a powerful tool for neuroscientists and researchers. Chengjie Zheng, Tewodros Mulugeta Dagnew, Liuyue Yang, Wei Ding 0003, Shiqian Shen, Changning Wang, Ping Chen 0001 |
IEEE Big Data | 4 |
| 2024 | Overview of ACM SIGKDD 2024 AI4Science4AI Special DayabstractThis paper provides an overview of the ACM SIGKDD 2024 AI4Science4AI special day. It includes information about the organizers, invited speakers, keynote speakers, the event agenda, and insights from related workshops. The AI4Science4AI special day aims to bring together experts in artificial intelligence (AI) and science to discuss the latest developments, challenges, and future directions. Wei Ding 0003, Gustau Camps-Valls |
KDD | 1 |
| 2024 | A Multi-view Feature Construction and Multi-Encoder-Decoder Transformer Architecture for Time Series Classification
Wei Ding 0003, Inal Mashukov, Scott E. Crouter, Ping Chen 0001 |
PAKDD (6) | 2 |
| 2023 | CASTLE: A Cascaded Spatio-Temporal Approach for Long-lead Streamflow ForecastingabstractEffective early warning systems for extreme flood events in large river basins necessitate reliable long-lead streamflow forecasts. However, the inherent uncertainty within each phase of the weather system-rainfall prediction, runoff generation, and streamflow prediction-amplifies with each stage, rendering accurate long-lead streamflow estimations challenging. In response to this, our study introduces a novel deep-learning-based model, the Cascaded Spatio-Temporal Learning Deep Network (CASTLE). CASTLE synergistically integrates observed upstream precipitation, recent streamflow data, and short-term precipitation forecasts derived from a selection of quantitative climate models to produce an accurate streamflow estimate. Specifically, we employ deep residual architectures on both observed and forecasted precipitation data to model the cascading spatio-temporal processes, which begin with upstream rainfall, move to rainfall-runoff, and finally conclude with downstream discharge. Our aim is to identify hidden space-time patterns that can be used to forecast future downstream flow over extended periods. We assess CASTLE’s efficacy by forecasting the downstream discharge of the Ganges River over a long lead time. Results show that our approach outperforms the current state-of-the-art streamflow forecasting models. Yong Zhuang, David L. Small, Patrick D. Flynn, Wahid Palash, Ping Chen 0001, Wei Ding 0003 |
IEEE Big Data | 7 |
| 2022 | Widening the Time Horizon: Predicting the Long-Term Behavior of Chaotic SystemsabstractThe understanding of chaotic systems is challenging not only for theoretical research but also for many important applications. Chaotic behavior is found in many nonlinear dynamical systems, such as those found in climate dynamics, weather, the stock market, and the space-time dynamics of virus spread. A reliable solution for these systems must handle their complex space-time dynamics and sensitive dependence on initial conditions. We develop a deep learning framework to push the time horizon at which reliable predictions can be made further into the future by better evaluating the consequences of local errors when modeling nonlinear systems. Our approach observes the future trajectories of initial errors at a time horizon to model the evolution of the loss to that point with two major components: 1) a recurrent architecture, Error Trajectory Tracing, that is designed to trace the trajectories of predictive errors through phase space, and 2) a training regime, Horizon Forcing, that pushes the model’s focus out to a predetermined time horizon. We validate our method on classic chaotic systems and real-world time series prediction tasks with chaotic characteristics, and show that our approach outperforms the current state-of-the-art methods. Yong Zhuang, Matthew Almeida, Wei Ding 0003, Patrick D. Flynn, Ping Chen 0001 |
ICDM | 3 |
| 2022 | Causal Feature Selection with Missing DataabstractCausal feature selection aims at learning the Markov blanket (MB) of a class variable for feature selection. The MB of a class variable implies the local causal structure among the class variable and its MB and all other features are probabilistically independent of the class variable conditioning on its MB, this enables causal feature selection to identify potential causal features for feature selection for building robust and physically meaningful prediction models. Missing data, ubiquitous in many real-world applications, remain an open research problem in causal feature selection due to its technical complexity. In this article, we discuss a novel multiple imputation MB (MimMB) framework for causal feature selection with missing data. MimMB integrates Data Imputation with MB Learning in a unified framework to enable the two key components to engage with each other. MB Learning enables Data Imputation in a potentially causal feature space for achieving accurate data imputation, while accurate Data Imputation helps MB Learning identify a reliable MB of the class variable in turn. Then, we further design an enhanced kNN estimator for imputing missing values and instantiate the MimMB. In our comprehensively experimental evaluation, our new approach can effectively learn the MB of a given variable in a Bayesian network and outperforms other rival algorithms using synthetic and real-world datasets. Kui Yu, Yajing Yang, Wei Ding 0003 |
ACM Trans. Knowl. Discov. Data | 3 |
| 2021 | Mitigating Class-Boundary Label Uncertainty to Reduce Both Model Bias and VarianceabstractThe study of model bias and variance with respect to decision boundaries is critically important in supervised learning and artificial intelligence. There is generally a tradeoff between the two, as fine-tuning of the decision boundary of a classification model to accommodate more boundary training samples (i.e., higher model complexity) may improve training accuracy (i.e., lower bias) but hurt generalization against unseen data (i.e., higher variance). By focusing on just classification boundary fine-tuning and model complexity, it is difficult to reduce both bias and variance. To overcome this dilemma, we take a different perspective and investigate a new approach to handle inaccuracy and uncertainty in the training data labels, which are inevitable in many applications where labels are conceptual entities and labeling is performed by human annotators. The process of classification can be undermined by uncertainty in the labels of the training data; extending a boundary to accommodate an inaccurately labeled point will increase both bias and variance. Our novel method can reduce both bias and variance by estimating the pointwise label uncertainty of the training set and accordingly adjusting the training sample weights such that those samples with high uncertainty are weighted down and those with low uncertainty are weighted up. In this way, uncertain samples have a smaller contribution to the objective function of the model’s learning algorithm and exert less pull on the decision boundary. In a real-world physical activity recognition case study, the data present many labeling challenges, and we show that this new approach improves model performance and reduces model variance. Matthew Almeida, Yong Zhuang, Wei Ding 0003, Scott E. Crouter, Ping Chen 0001 |
ACM Trans. Knowl. Discov. Data | 3 |
| 2020 | Catalysis Clustering with GAN by Incorporating Domain KnowledgeabstractClustering is an important unsupervised learning method with serious challenges when data is sparse and high-dimensional. Generated clusters are often evaluated with general measures, which may not be meaningful or useful for practical applications and domains. Using a distance metric, a clustering algorithm searches through the data space, groups close items into one cluster, and assigns far away samples to different clusters. In many real-world applications, the number of dimensions is high and data space becomes very sparse. Selection of a suitable distance metric is very difficult and becomes even harder when categorical data is involved. Moreover, existing distance metrics are mostly generic, and clusters created based on them will not necessarily make sense to domain-specific applications. One option to address these challenges is to integrate domain-defined rules and guidelines into the clustering process. In this work we propose a GAN-based approach called Catalysis Clustering to incorporate domain knowledge into the clustering process. With GANs we generate catalysts, which are special synthetic points drawn from the original data distribution and verified to improve clustering quality when measured by a domain-specific metric. We then perform clustering analysis using both catalysts and real data. Final clusters are produced after catalyst points are removed. Experiments on two challenging real-world datasets clearly show that our approach is effective and can generate clusters that are meaningful and useful for real-world applications. Olga Andreeva, Wei Li 0121, Wei Ding 0003, Marieke L. Kuijjer, John Quackenbush, Ping Chen 0001 |
KDD | 3 |
| 2020 | A Novel Deep Learning Model by Stacking Conditional Restricted Boltzmann Machine and Deep Neural NetworkabstractA real-world system often exhibits complex dynamics arising from interaction among its subunits. In machine learning and data mining, these interactions are usually formulated as dependency and correlation among system variables. Similar to Convolution Neural Network dealing with spatially correlated features and Recurrent Neural Network with temporally correlated features, in this paper we present a novel deep learning model to tackle functionally interactive features by stacking a Conditional Restricted Boltzmann Machine and a Deep Neural Network (CRBM-DNN). Variables with their dependency relationships are organized into a bipartite graph, which is further converted into a Restricted Boltzmann Machine conditioned by domain knowledge. We integrate this CRBM and a DNN into one deep learning model constrained by one overall cost function. CRBM-DNN can solve both supervised and unsupervised learning problems. Compared to a regular neural network of the same size, CRBM-DNN has fewer parameters so they require fewer training samples. We perform extensive comparative studies with a large number of supervised learning and unsupervised learning methods using several challenging real-world datasets, and achieve significant superior performance. Tianyu Kang, Ping Chen 0001, John Quackenbush, Wei Ding 0003 |
KDD | 4 |
| 2020 | Introducing time series snippets: a new primitive for summarizing long time series
Shima Imani, Frank Madrid, Wei Ding 0003, Scott E. Crouter, Eamonn J. Keogh |
Data Min. Knowl. Discov. | 3 |
| 2019 | Domain agnostic online semantic segmentation for multi-dimensional time seriesabstractUnsupervised semantic segmentation in the time series domain is a much studied problem due to its potential to detect unexpected regularities and regimes in poorly understood data. However, the current techniques have several shortcomings, which have limited the adoption of time series semantic segmentation beyond academic settings for four primary reasons. First, most methods require setting/learning many parameters and thus may have problems generalizing to novel situations. Second, most methods implicitly assume that all the data is segmentable and have difficulty when that assumption is unwarranted. Thirdly, many algorithms are only defined for the single dimensional case, despite the ubiquity of multi-dimensional data. Finally, most research efforts have been confined to the batch case, but online segmentation is clearly more useful and actionable. To address these issues, we present a multi-dimensional algorithm, which is domain agnostic, has only one, easily-determined parameter, and can handle data streaming at a high rate. In this context, we test the algorithm on the largest and most diverse collection of time series datasets ever considered for this task and demonstrate the algorithm's superiority over current solutions. Shaghayegh Gharghabi, Chin-Chia Michael Yeh, Yifei Ding, Wei Ding 0003, Paul Hibbing, Samuel LaMunion, Andrew Kaplan, Scott E. Crouter, Eamonn J. Keogh |
Data Min. Knowl. Discov. | 4 |
| 2019 | Correction to: Domain agnostic online semantic segmentation for multi-dimensional time seriesabstractThe article Domain agnostic online semantic segmentation for multi-dimensional time series, written by Shaghayegh Gharghabi, Chin-Chia Michael Yeh, Yifei Ding, Wei Ding, Paul Hibbing, Samuel LaMunion, Andrew Kaplan, Scott E. Crouter, Eamonn Keogh was originally published electronically on the publisher’s internet portal (currently SpringerLink) on 25 September 2018 without open access. Shaghayegh Gharghabi, Chin-Chia Michael Yeh, Yifei Ding, Wei Ding 0003, Paul Hibbing, Samuel LaMunion, Andrew Kaplan, Scott E. Crouter, Eamonn J. Keogh |
Data Min. Knowl. Discov. | 4 |
| 2019 | BAMB: A Balanced Markov Blanket Discovery Approach to Feature SelectionabstractThe discovery of Markov blanket (MB) for feature selection has attracted much attention in recent years, since the MB of the class attribute is the optimal feature subset for feature selection. However, almost all existing MB discovery algorithms focus on either improving computational efficiency or boosting learning accuracy, instead of both. In this article, we propose a novel MB discovery algorithm for balancing efficiency and accuracy, called BAlanced Markov Blanket (BAMB) discovery. To achieve this goal, given a class attribute of interest, BAMB finds candidate PC (parents and children) and spouses and removes false positives from the candidate MB set in one go. Specifically, once a feature is successfully added to the current PC set, BAMB finds the spouses with regard to this feature, then uses the updated PC and the spouse set to remove false positives from the current MB set. This makes the PC and spouses of the target as small as possible and thus achieves a trade-off between computational efficiency and learning accuracy. In the experiments, we first compare BAMB with 8 state-of-the-art MB discovery algorithms on 7 benchmark Bayesian networks, then we use 10 real-world datasets and compare BAMB with 12 feature selection algorithms, including 8 state-of-the-art MB discovery algorithms and 4 other well-established feature selection methods. On prediction accuracy, BAMB outperforms 12 feature selection algorithms compared. On computational efficiency, BAMB is close to the IAMB algorithm while it is much faster than the remaining seven MB discovery algorithms. Zhaolong Ling, Kui Yu, Hao Wang 0008, Lin Liu 0003, Wei Ding 0003, Xindong Wu 0001 |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2019 | Heterogeneous-Length Text Topic Modeling for Reader-Aware Multi-Document SummarizationabstractMore and more user comments like Tweets are available, which often contain user concerns. In order to meet the demands of users, a good summary generating from multiple documents should consider reader interests as reflected in reader comments. In this article, we focus on how to generate a summary from multi-document documents by considering reader comments, named as reader-aware multi-document summarization (RA-MDS). We present an innovative topic-based method for RA-MDA, which exploits latent topics to obtain the most salient and lessen redundancy summary from multiple documents. Since finding latent topics for RA-MDS is a crucial step, we also present a Heterogeneous-length Text Topic Modeling (HTTM) to extract topics from the corpus that includes both news reports and user comments, denoted as heterogeneous-length texts. In this case, the latent topics extract by HTTM cover not only important aspects of the event, but also aspects that attract reader interests. Comparisons on summary benchmark datasets also confirm that the proposed RA-MDS method is effective in improving the quality of extracted summaries. In addition, experimental results demonstrate that the proposed topic modeling method outperforms existing topic modeling algorithms. Jipeng Qiang, Ping Chen 0001, Wei Ding 0003, Tong Wang 0007, Fei Xie 0002, Xindong Wu 0001 |
ACM Trans. Knowl. Discov. Data | 3 |
| 2018 | Clustering on Sparse Data in Non-overlapping Feature Space with Applications to Cancer SubtypingabstractThis paper presents a new algorithm, Reinforced and Informed Network-based Clustering(RINC), for finding unknown groups of similar data objects in sparse and largely non-overlapping feature space where a network structure among features can be observed. Sparse and non-overlapping unlabeled data become increasingly common and available especially in text mining and biomedical data mining. RINC inserts a domain informed model into a modelless neural network. In particular, our approach integrates physically meaningful feature dependencies into the neural network architecture and soft computational constraint. Our learning algorithm efficiently clusters sparse data through integrated smoothing and sparse auto-encoder learning. The informed design requires fewer samples for training and at least part of the model becomes explainable. The architecture of the reinforced network layers smooths sparse data over the network dependency in the feature space. Most importantly, through back-propagation, the weights of the reinforced smoothing layers are simultaneously constrained by the remaining sparse auto-encoder layers that set the target values to be equal to the raw inputs. Empirical results demonstrate that RINC achieves improved accuracy and renders physically meaningful clustering results. Tianyu Kang, Kourosh Zarringhalam, Marieke L. Kuijjer, Ping Chen 0001, John Quackenbush, Wei Ding 0003 |
ICDM | 6 |
| 2016 | Scalable and Accurate Online Feature Selection for Big DataabstractFeature selection is important in many big data applications. Two critical challenges closely associate with big data. First, in many big data applications, the dimensionality is extremely high, in millions, and keeps growing. Second, big data applications call for highly scalable feature selection algorithms in an online manner such that each feature can be processed in a sequential scan. We present SAOLA, a Scalable and Accurate OnLine Approach for feature selection in this paper. With a theoretical analysis on bounds of the pairwise correlations between features, SAOLA employs novel pairwise comparison techniques and maintains a parsimonious model over time in an online manner. Furthermore, to deal with upcoming features that arrive by groups, we extend the SAOLA algorithm, and then propose a new group-SAOLA algorithm for online group feature selection. The group-SAOLA algorithm can online maintain a set of feature groups that is sparse at the levels of both groups and individual features simultaneously. An empirical study using a series of benchmark real datasets shows that our two algorithms, SAOLA and group-SAOLA, are scalable on datasets of extremely high dimensionality and have superior performance over the state-of-the-art feature selection methods. Kui Yu, Xindong Wu 0001, Wei Ding 0003, Jian Pei 0001 |
ACM Trans. Knowl. Discov. Data | 3 |
| 2016 | Hierarchical Spatio-Temporal Pattern Discovery and Predictive ModelingabstractWe propose a new approach, CCRBoost, to identify the hierarchical structure of spatio-temporal patterns at different resolution levels and subsequently construct a predictive model based on the identified structure. To accomplish this, we first obtain indicators within different spatio-temporal spaces from the raw data. A distributed spatio-temporal pattern (DSTP) is extracted from a distribution, which consists of the locations with similar indicators from the same time period, generated by multi-clustering. Next, we use a greedy searching and pruning algorithm to combine the DSTPs in order to form an ensemble spatio-temporal pattern (ESTP). An ESTP can represent the spatio-temporal pattern of various regularities or a non-stationary pattern. To consider all the possible scenarios of a real-world ST pattern, we then build a model with layers of weighted ESTPs. By evaluating all the indicators of one location, this model can predict whether a target event will occur at this location. In the case study of predicting crime events, our results indicate that the predictive model can achieve 80 percent accuracy in predicting residential burglary, which is better than other methods. Chung-Hsien Yu, Wei Ding 0003, Melissa Morabito, Ping Chen 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2016 | Online Learning from Trapezoidal Data StreamsabstractIn this paper, we study a new problem of continuous learning from doubly-streaming data where both data volume and feature space increase over time. We refer to the doubly-streaming data as trapezoidal data streams and the corresponding learning problem as online learning from trapezoidal data streams. The problem is challenging because both data volume and data dimension increase over time, and existing online learning[1],[2], online feature selection[3], and streaming feature selection algorithms[4],[5]are inapplicable. We propose a new Online Learning with Streaming Features algorithm (OL$_{SF}$for short) and its two variants, which combine online learning[1],[2]and streaming feature selection[4],[5]to enable learning from trapezoidal data streams with infinite training instances and features. When a new training instance carrying new features arrives, a classifier updates the existing features by following the passive-aggressive update rule[2]and updates the new features by following the structural risk minimization principle. Feature sparsity is then introduced by using the projected truncation technique. We derive performance bounds of the OL$_{SF}$algorithm and its variants. We also conduct experiments on real-world data sets to show the performance of the proposed algorithms. Qin Zhang 0011, Peng Zhang 0001, Guodong Long, Wei Ding 0003, Chengqi Zhang, Xindong Wu 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2015 | Spatio-temporal asynchronous co-occurrence pattern for big climate data towards long-lead flood predictionabstractRecent research efforts aim at utilizing Big Climate Data to predict floods 5 to 15 days in advance. Improvements in the prediction of heavy precipitation, a major factor related with flood occurrences, have lagged behind due to the high-dimensionality and non-linearity in the weather, hydriology and dydraulic systems. In this paper, we introduce Spatio-Temporal Asynchronous Co-Occurrence Pattern to associate heavy precipitation with dense precipitable water and explore long-lead flood prediction from the machine learning perspective. Our model predicts one location's flooding risk by connecting the heavy precipitation with its preceding precipitable water through an association mining method. We discover asynchronous co-occurrence location and discuss a spatio-temporal ensemble learning method for predictive modeling. Our framework requires less computational cost and smaller train data compared to other existing approaches. In addition, the framework is designed to be scalable and allows distributed computing. Our real-world case study in the state of Iowa has achieved 87% accuracy on predicting the heavy precipitations which trigger severe floods at least 9 days in advance. Chung-Hsien Yu, Wei Ding 0003, Joseph Paul Cohen, David L. Small |
IEEE BigData | 3 |
| 2015 | A Hierarchical Pattern Learning Framework for Forecasting Extreme Weather EventsabstractExtreme weather events, like extreme rainfalls, are severe weather hazards and also the triggers for other natural disasters like floods and tornadoes. Accurate forecasting of such events relies on the understanding of the spatiotemporal evolution processes in climate system. Learning from climate science data has been a challenging task, because the variations among spatial, temporal and multivariate spaces have created a huge amount of features and complex regularities within the data. In this study we developed a framework for learning patterns from the spatiotemporal system and forecasting extreme weather events. In this framework, we learned patterns in a hierarchical manner: in each level, new features were learned from data and used as the input for the next level. Firstly, we summarized the temporal evolution process of individual variables by learning the location-based patterns. Secondly, we developed an optimization algorithm for summarizing the spatial regularities, SCOT, by growing spatial clusters from the location-based patterns. Finally, we developed an instance-based algorithm, SPC, to forecast the extreme events through classification. We applied this framework to forecasting extreme rainfall events in the eastern Central Andes area. Our experiments show that this method was able to find climatic process patterns similar to those found in domain studies, and our forecasting results outperformed the state-of-art model. Dawei Wang 0008, Wei Ding 0003 |
ICDM | 2 |
| 2015 | Towards Mining Trapezoidal Data StreamsabstractWe study a new problem of learning from doubly-streaming data where both data volume and feature space increase over time. We refer to the problem as mining trapezoidal data streams. The problem is challenging because both data volume and feature space are increasing, to which existing online learning, online feature selection and streaming feature selection algorithms are inapplicable. We propose a new Sparse Trapezoidal Streaming Data mining algorithm (STSD) and its two variants which combine online learning and online feature selection to enable learning trapezoidal data streams with infinite training instances and features. Specifically, when new training instances carrying new features arrive, the classifier updates the existing features by following the passive-aggressive update rule used in online learning and updates the new features with the structural risk minimization principle. Feature sparsity is also introduced using the projected truncation techniques. Extensive experiments on the demonstrated UCI data sets show the performance of the proposed algorithms. Qin Zhang 0011, Peng Zhang 0001, Guodong Long, Wei Ding 0003, Chengqi Zhang, Xindong Wu 0001 |
ICDM | 4 |
| 2015 | Tornado Forecasting with Multiple Markov BoundariesabstractReliable tornado forecasting with a long-lead time can greatly support emergency response and is of vital importance for the economy and society. The large number of meteorological variables in spatiotemporal domains and the complex relationships among variables remain the top difficulties for a long-lead tornado forecasting. Kui Yu, Dawei Wang 0008, Wei Ding 0003, Jian Pei 0001, David L. Small, Xindong Wu 0001 |
KDD | 3 |
| 2015 | Classification with Streaming Features: An Emerging-Pattern Mining ApproachabstractMany datasets from real-world applications have very high-dimensional or increasing feature space. It is a new research problem to learn and maintain a classifier to deal with very high dimensionality or streaming features. In this article, we adapt the well-known emerging-pattern--based classification models and propose a semi-streaming approach. For streaming features, it is computationally expensive or even prohibitive to mine long-emerging patterns, and it is nontrivial to integrate emerging-pattern mining with feature selection. We present an online feature selection step, which is capable of selecting and maintaining a pool of effective features from a feature stream. Then, in our offline step, separated from the online step, we periodically compute and update emerging patterns from the pool of selected features from the online step. We evaluate the effectiveness and efficiency of the proposed method using a series of benchmark datasets and a real-world case study on Mars crater detection. Our proposed method yields classification performance comparable to the state-of-art static classification methods. Most important, the proposed method is significantly faster and can efficiently handle datasets with streaming features. Kui Yu, Wei Ding 0003, Dan A. Simovici, Hao Wang 0008, Jian Pei 0001, Xindong Wu 0001 |
ACM Trans. Knowl. Discov. Data | 2 |
| 2014 | Towards Scalable and Accurate Online Feature Selection for Big DataabstractFeature selection is important in many big data applications. There are at least two critical challenges. Firstly, in many applications, the dimensionality is extremely high, in millions, and keeps growing. Secondly, feature selection has to be highly scalable, preferably in an online manner such that each feature can be processed in a sequential scan. In this paper, we develop SAOLA, a Scalable and Accurate On Line Approach for feature selection. With a theoretical analysis on a low bound on the pair wise correlations between features in the currently selected feature subset, SAOLA employs novel online pair wise comparison techniques to address the two challenges and maintain a parsimonious model over time in an online manner. An empirical study using a series of benchmark real data sets shows that SAOLA is scalable on data sets of extremely high dimensionality, and has superior performance over the state-of-the-art feature selection methods. Kui Yu, Xindong Wu 0001, Wei Ding 0003, Jian Pei 0001 |
ICDM | 3 |
| 2014 | Crime Forecasting Using Spatio-temporal Pattern with Ensemble Learning
Chung-Hsien Yu, Wei Ding 0003, Ping Chen 0001, Melissa Morabito |
PAKDD (2) | 2 |
| 2014 | Bipart: Learning Block Structurefor Activity DetectionabstractPhysical activity consists complex behavior, typically structured in bouts which can consist of one continuous movement (e.g. exercise) or many sporadic movements (e.g. household chores). Each bout can be represented as a block of feature vectors corresponding to the same activity type. This paper introduces a general distance metric technique to use this block representation to first predict activity type, and then uses the predicted activity to estimate energy expenditure within a novel framework. This distance metric, dubbed Bipart, learns block-level information from both training and test sets, combining both to form a projection space which materializes block-level constraints. Thus, Bipart provides a space which can improve the bout classification performance of all classifiers. We also propose an energy expenditure estimation framework which leverages activity classification in order to improve estimates. Comprehensive experiments on waist-mounted accelerometer data, comparing Bipart against many similar methods as well as other classifiers, demonstrate the superior activity recognition of Bipart, especially in low-information experimental settings. Yang Mu, Henry Z. Lo, Wei Ding 0003, Kevin Michael Amaral, Scott E. Crouter |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2014 | Data Mining with Big DataabstractBig Data concern large-volume, complex, growing data sets with multiple, autonomous sources. With the fast development of networking, data storage, and the data collection capacity, Big Data are now rapidly expanding in all science and engineering domains, including physical, biological and biomedical sciences. This paper presents a HACE theorem that characterizes the features of the Big Data revolution, and proposes a Big Data processing model, from the data mining perspective. This data-driven model involves demand-driven aggregation of information sources, mining and analysis, user interest modeling, and security and privacy considerations. We analyze the challenging issues in the data-driven model and also in the Big Data revolution. Xindong Wu 0001, Xingquan Zhu 0001, Gong-Qing Wu, Wei Ding 0003 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2013 | Group Feature Selection with Streaming FeaturesabstractGroup feature selection makes use of structural information among features to discover a meaningful subset of features. Existing group feature selection algorithms only deal with pre-given candidate feature sets and they are incapable of handling streaming features. On the other hand, feature selection algorithms targeted for streaming features can only perform at the individual feature level without considering intrinsic group structures of the features. In this paper, we perform group feature selection with streaming features. We propose to perform feature selection at the group and individual feature levels simultaneously in a manner of a feature stream rather than a pre-given candidate feature set. In our approach, the group structures are fully utilized to reduce the cost of evaluating streaming features. We have extensively evaluated the proposed method. Experimental results have demonstrated that our proposed algorithms statistically outperform state-of-the-art methods of feature selection in terms of classification accuracy. Hai-Guang Li, Xindong Wu 0001, Zhao Li 0007, Wei Ding 0003 |
ICDM | 4 |
| 2013 | Markov Blanket Feature Selection with Non-faithful Data DistributionsabstractIn faithful Bayesian networks, the Markov blanket of the class attribute is a unique and minimal feature subset for optimal feature selection. However, little attention has been paid to Markov blanket feature selection in a non-faithful environment which widely exists in the real world. To tackle this issue, in this paper, we deal with non-faithful data distributions and propose the concept of representative sets instead of Markov blankets. With a standard sparse group lasso for selection of features from the representative sets, we design an effective algorithm, SRS, for Markov blanket feature Selection via Representative Sets with non-faithful data distributions. Empirical studies demonstrate that SRS outperforms the state-of-the-art Markov blanket feature selectors and other well-established feature selection methods. Kui Yu, Xindong Wu 0001, Zan Zhang 0002, Yang Mu, Hao Wang 0008, Wei Ding 0003 |
ICDM | 6 |
| 2013 | Constrained stochastic gradient descent for large-scale least squares problemabstractThe least squares problem is one of the most important regression problems in statistics, machine learning and data mining. In this paper, we present the Constrained Stochastic Gradient Descent (CSGD) algorithm to solve the large-scale least squares problem. CSGD improves the Stochastic Gradient Descent (SGD) by imposing a provable constraint that the linear regression line passes through the mean point of all the data points. It results in the best regret bound $O(\log{T})$, and fastest convergence speed among all first order approaches. Empirical studies justify the effectiveness of CSGD by comparing it with SGD and other state-of-the-art approaches. An example is also given to show how to use CSGD to optimize SGD based least squares problems to achieve a better performance. Yang Mu, Wei Ding 0003, Tianyi Zhou 0001, Dacheng Tao |
KDD | 2 |
| 2013 | Towards long-lead forecasting of extreme flood events: a data mining framework for precipitation cluster precursors identificationabstractThe development of disastrous flood forecasting techniques able to provide warnings at a long lead-time (5-15 days) is of great importance to society. Extreme Flood is usually a consequence of a sequence of precipitation events occurring over from several days to several weeks. Though precise short-term forecasting the magnitude and extent of individual precipitation event is still beyond our reach, long-term forecasting of precipitation clusters can be attempted by identifying persistent atmospheric regimes that are conducive for the precipitation clusters. However, such forecasting will suffer from overwhelming number of relevant features and high imbalance of sample sets. In this paper, we propose an integrated data mining framework for identifying the precursors to precipitation event clusters and use this information to predict extended periods of extreme precipitation and subsequent floods. We synthesize a representative feature set that describes the atmosphere motion, and apply a streaming feature selection algorithm to online identify the precipitation precursors from the enormous feature space. A hierarchical re-sampling approach is embedded in the framework to deal with the imbalance problem. An extensive empirical study is conducted on historical precipitation and associated flood data collected in the State of Iowa. Utilizing our framework a few physically meaningful precipitation cluster precursor sets are identified from millions of features. More than 90% of extreme precipitation events are captured by the proposed prediction model using precipitation cluster precursors with a lead time of more than 5 days. Dawei Wang 0008, Wei Ding 0003, Kui Yu, Xindong Wu 0001, Ping Chen 0001, David L. Small |
KDD | 2 |
| 2013 | Feature Selection by Joint Graph Sparse CodingabstractThis paper takes manifold learning and regression simultaneously into account to perform unsupervised spectral feature selection. We first extract the bases of the data, and then represent the data sparsely using the extracted bases by proposing a novel joint graph sparse coding model, JGSC for short. We design a new algorithm TOSC to compute the resulting objective function of JGSC, and then theoretically prove that the proposed objective function converges to its global optimum via the proposed TOSC algorithm. We repeat the extraction and the TOSC calculation until the value of the objective function of JGSC satisfies pre-defined conditions. Eventually the derived new representation of the data may only have a few non-zero rows, and we delete the zero rows (a.k.a. zero-valued features) to conduct feature selection on the new representation of the data. Our empirical studies demonstrate that the proposed method outperforms several state-of-the-art algorithms on real datasets in term of the kNN classification performance. Wei Ding 0003, Xindong Wu 0001, Shichao Zhang 0001, Xiaofeng Zhu 0001 |
SDM | 1 |
| 2013 | Bridging Causal Relevance and Pattern Discriminability: Mining Emerging Patterns from High-Dimensional DataabstractIt is a nontrivial task to build an accurate emerging pattern (EP) classifier from high-dimensional data because we inevitably face two challenges 1) how to efficiently extract a minimal set of strongly predictive EPs from an explosive number of candidate patterns, and 2) how to handle the highly sensitive choice of the minimal support threshold. To address these two challenges, we bridge causal relevance and EP discriminability (the predictive ability of emerging patterns) to facilitate EP mining and propose a new framework of mining EPs from high-dimensional data. In this framework, we study the relationships between causal relevance in a causal Bayesian network and EP discriminability in EP mining, and then reduce the pattern space of EP mining to direct causes and direct effects, or the Markov blanket (MB) of the class attribute in a causal Bayesian network. The proposed framework is instantiated by two EPs-based classifiers, CE-EP and MB-EP, where CE stands for direct Causes and direct Effects, and MB for Markov Blanket. Extensive experiments on a broad range of data sets validate the effectiveness of the CE-EP and MB-EP classifiers against other well-established methods, in terms of predictive accuracy, pattern numbers, running time, and sensitivity analysis. Kui Yu, Wei Ding 0003, Hao Wang 0008, Xindong Wu 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2012 | Self-Taught Active Learning from CrowdsabstractThe emergence of social tagging and crowdsourcing systems provides a unique platform where multiple weak labelers can form a crowd to fulfill a labeling task. Yet crowd labelers are often noisy, inaccurate, and have limited labeling knowledge, and worst of all, they act independently without seeking complementary knowledge from each other to improve labeling performance. In this paper, we propose a Self-Taught Active Learning (STAL) paradigm, where imperfect labelers are able to learn complementary knowledge from one another to expand their knowledge sets and benefit the underlying active learner. We employ a probabilistic model to characterize the knowledge of each labeler through which a weak labeler can learn complementary knowledge from a stronger peer. As a result, the self-taught active learning process eventually helps achieve high classification accuracy with minimized labeling costs and labeling errors. Xingquan Zhu 0001, Bin Li 0015, Wei Ding 0003, Xindong Wu 0001 |
ICDM | 4 |
| 2012 | Coupled behavior analysis for capturing coupling relationships in group-based market manipulationsabstractIn stock markets, an emerging challenge for surveillance is that a group of hidden manipulators collaborate with each other to manipulate the price movement of securities. Recently, the coupled hidden Markov model (CHMM)-based coupled behavior analysis (CBA) has been proposed to consider the coupling relationships in the above group-based behaviors for manipulation detection. From the modeling perspective, however, this requires overall aggregation of the behavioral data to cater for the CHMM modeling, which does not differentiate the coupling relationships presented in different forms within the aggregated behaviors and degrade the capability for further anomaly detection. Thus, this paper suggests a general CBA framework for detecting group-based market manipulation by capturing more comprehensive couplings and proposes two variant implementations, which are hybrid coupling (HC)-based and hierarchical grouping (HG)-based respectively. The proposed framework consists of three stages. The first stage, qualitative analysis, generates possible qualitative coupling relationships between behaviors with or without domain knowledge. In the second stage, quantitative representation of coupled behaviors is learned via proper methods. For the third stage, anomaly detection algorithms are proposed to cater for different application scenarios. Experimental results on data from a major Asian stock market show that the proposed framework outperforms the CHMM-based analysis in terms of detecting abnormal collaborative market manipulations. Additionally, the two different implementations are compared with their effectiveness for different application scenarios. Yin Song, Longbing Cao, Xindong Wu 0001, Wu Ye, Wei Ding 0003 |
KDD | 6 |
| 2012 | Mining emerging patterns by streaming feature selectionabstractBuilding an accurate emerging pattern classifier with a high-dimensional dataset is a challenging issue. The problem becomes even more difficult if the whole feature space is unavailable before learning starts. This paper presents a new technique on mining emerging patterns using streaming feature selection. We model high feature dimensions with streaming features, that is, features arrive and are processed one at a time. As features flow in one by one, we online evaluate each coming feature to determine whether it is useful for mining predictive emerging patterns (EPs) by exploiting the relationship between feature relevance and EP discriminability (the predictive ability of an EP). We employ this relationship to guide an online EP mining process. This new approach can mine EPs from a high-dimensional dataset, even when its entire feature set is unavailable before learning. The experiments on a broad range of datasets validate the effectiveness of the proposed approach against other well-established methods, in terms of predictive accuracy, pattern numbers and running time. Kui Yu, Wei Ding 0003, Dan A. Simovici, Xindong Wu 0001 |
KDD | 2 |
| 2011 | Bernoulli trials based feature selection for crater detectionabstractCounting craters is a fundamental task of planetary science because it provides the only tool for measuring relative ages of planetary surfaces. However, advances in surveying craters present in data gathered by planetary probes have not kept up with advances in data collection. One challenge of auto-detecting craters in images is to identify an image's features that discriminate it between craters and other surface objects. The problem of optimal feature selection is known to be NP-hard and the search is computationally intractable. In this paper we propose a wrapper based randomized feature selection method to efficiently select relevant features for crater detection. We design and implement a dynamic programming algorithm to search for a relevant feature subset by removing irrelevant features and minimizing a cost objective function simultaneously. In order to only remove irrelevant features we use Bernoulli Trials to calculate the probability of such a case using the cost function. Our proposed algorithms are empirically evaluated on a large high-resolution Martian image exhibiting a heavily cratered Martian terrain characterized by heterogeneous surface morphology. The experimental results demonstrate that the proposed approach achieves a higher accuracy than other existing randomized approaches to a large extent with less runtime. Wei Ding 0003, Joseph Paul Cohen, Dan A. Simovici, Tomasz F. Stepinski |
GIS | 2 |
| 2011 | Causal Associative ClassificationabstractAssociative classifiers have received considerable attention due to their easy to understand models and promising performance. However, with a high dimensional dataset, associative classifiers inevitably face two challenges: (1) how to extract a minimal set of strong predictive rules from an explosive number of generated association rules, and (2) how to deal with the highly sensitive choice of the minimal support threshold. In order to address these two challenges, we introduce causality into associative classification, and propose a new framework of causal associative classification. In this framework, we use causal Bayesian networks to bridge irrelevant and redundant features with irrelevant and redundant rules in associative classification. Without loss of prediction power, the feature space involved with the antecedent of a classification rule is reduced to the space of the direct causes, direct effects, and direct causes of the direct effects, a.k.a. the Markov blanket, of the consequent of the rule in causal Bayesian networks. The proposed framework is instantiated via baseline classifiers using emerging patterns. Experimental results show that our framework significantly reduces the model complexity while outperforming the other state-of-the-art algorithms. Kui Yu, Xindong Wu 0001, Wei Ding 0003, Hao Wang 0008, Hongliang Yao |
ICDM | 3 |
| 2011 | Empirical Discriminative Tensor Analysis for Crime Forecasting
Yang Mu, Wei Ding 0003, Melissa Morabito, Dacheng Tao |
KSEM | 2 |
| 2011 | A framework for regional association rule mining and scoping in spatial datasets
Wei Ding 0003, Christoph F. Eick, Xiaojing Yuan, Jing Wang 0007, Jean-Philippe Nicot |
GeoInformatica | 1 |
| 2011 | Controlling patterns of geospatial phenomena
Tomasz F. Stepinski, Wei Ding 0003, Christoph F. Eick |
GeoInformatica | 2 |
| 2011 | One-class learning and concept summarization for data streams
Xingquan Zhu 0001, Wei Ding 0003, Philip S. Yu, Chengqi Zhang |
Knowl. Inf. Syst. | 2 |
| 2011 | Subkilometer crater discovery with boosting and transfer learningabstractCounting craters in remotely sensed images is the only tool that provides relative dating of remote planetary surfaces. Surveying craters requires counting a large amount of small subkilometer craters, which calls for highly efficient automatic crater detection. In this article, we present an integrated framework on autodetection of subkilometer craters with boosting and transfer learning. The framework contains three key components. First, we utilize mathematical morphology to efficiently identify crater candidates , the regions of an image that can potentially contain craters. Only those regions occupying relatively small portions of the original image are the subjects of further processing. Second, we extract and select image texture features, in combination with supervised boosting ensemble learning algorithms, to accurately classify crater candidates into craters and noncraters. Third, we integrate transfer learning into boosting, to enhance detection performance in the regions where surface morphology differs from what is characterized by the training set. Our framework is evaluated on a large test image of 37,500 × 56,250 m 2 on Mars, which exhibits a heavily cratered Martian terrain characterized by nonuniform surface morphology. Empirical studies demonstrate that the proposed crater detection framework can achieve an F1 score above 0.85, a significant improvement over the other crater detection algorithms. Wei Ding 0003, Tomasz F. Stepinski, Yang Mu, Lourenço P. C. Bandeira, Ricardo Vilalta, Youxi Wu, Tianyu Cao 0001, Xindong Wu 0001 |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2010 | Automatic detection of craters in planetary images: an embedded framework using feature selection and boostingabstractIdentifying impact craters on planetary surfaces is one fundamental task in planetary science. In this paper, we present an embedded framework on auto-detection of craters, using feature selection and boosting strategies. The paradigm aims at building a universal and practical crater detector. This methodology addresses three issues that such a tool must possess: (i) it utilizes mathematical morphology to efficiently identify the regions of an image that can potentially contain craters; only those regions, defined as crater candidates, are the subjects of further processing; (ii) it selects Haar-like image texture features in combination with boosting ensemble supervised learning algorithms to accurately classify candidates into craters and non-craters; (iii) it uses transfer learning, at a minimum additional cost, to enable maintaining an accurate auto-detection of craters on new images, having morphology different from what has been captured by the original training set. All three aforementioned components of the detection methodology are discussed, and the entire framework is evaluated on a large test image of 37,500 x 56,250$ m2 on Mars, showing heavily cratered Martian terrain characterized by nonuniform surface morphology. Our study demonstrates that this methodology provides a robust and practical tool for planetary science, in terms of both detection accuracy and efficiency. Wei Ding 0003, Tomasz F. Stepinski, Lourenço P. C. Bandeira, Ricardo Vilalta, Youxi Wu, Tianyu Cao 0001 |
CIKM | 1 |
| 2010 | Exploring labeled spatial datasets using association analysisabstractWe use an association analysis-based strategy for exploration of multi-attribute spatial datasets possessing naturally arising classification. In this demonstration, we present a prototype system, ESTATE (Exploring Spatial daTa Association patTErns), inverting such classification by interpreting different classes found in the dataset in terms of sets of discriminative patterns of its attributes. The system consists of several core components including discriminative data mining, similarity between transactional patterns, and visualization. An algorithm for calculating similarity measure between patterns is the major original contribution that facilitates summarization of discovered information and makes the entire framework practical for real life applications. We demonstrate two applications of ESTATE in the domains of ecology and sociology. The ecology application is to discover the associations of between environmental factors and the spatial distribution of biodiversity across the contiguous United States, and the sociology application aims to discover different spatio-social motifs of support for Barack Obama in the 2008 presidential election. Tomasz F. Stepinski, Josue Salazar, Wei Ding 0003 |
GIS | 3 |
| 2010 | Causal Discovery from Streaming FeaturesabstractIn this paper, we study a new research problem of causal discovery from streaming features. A unique characteristic of streaming features is that not all features can be available before learning begins. Feature generation and selection often have to be interleaved. Managing streaming features has been extensively studied in classification, but little attention has been paid to the problem of causal discovery from streaming features. To this end, we propose a novel algorithm to solve this challenging problem, denoted as CDFSF (Causal Discovery From Streaming Features) which consists of two phases: growing and shrinking. In the growing phase, CDFSF finds candidate parents or children for each feature seen so far, while in the shrinking phase the algorithm dynamically removes false positives from the current sets of candidate parents and children. In order to improve the efficiency of CDFSF, we present S-CDFSF, a faster version of CDFSF, using two symmetry theorems. Experimental results validate our algorithms in comparison with other state-of-art algorithms of causal discovery. Kui Yu, Xindong Wu 0001, Hao Wang 0008, Wei Ding 0003 |
ICDM | 4 |
| 2009 | Discovery of Geospatial Discriminating Patterns from Remote Sensing DatasetsabstractLarge amounts of remotely sensed data calls for data mining techniques to fully utilize their rich information content. In this paper, we study new means of discovery and summarization of knowledge contained in the spatial patterns of remote sensing datasets. Several geospatial feature variables are fused together, and the vector of their values at each spatial cell is considered as a transaction to be used in association analysis. The concept of emerging patterns is applied to ascertain the variables that exert dominant influence on the distribution of a selected class variable. A new value-iteration method is introduced to optimally split the spatial domain of the selected variable into two classes. This division is used to calculate the set of patterns that are emerging with respect to the two classes; these patterns are the controlling factors—they are responsible for the spatial distribution of the class variable. A method for a concise summarization of controlling factors is introduced using a similarity measure that is custom-made for the type of patterns stemmed from remote sensing measurements. Using such a similarity measure, controlling factors are clustered providing brief description of different manners, in which the class variable is constrained by the explanatory variables. We evaluate our method in a real-world application pertaining to the density of vegetation within the continental United States. Examination of patterns related to the high vegetation cover provides a summary of data dependencies that helps to develop a better empirical model of the vegetation growth. Wei Ding 0003, Tomasz F. Stepinski, Josue Salazar |
SDM | 1 |
| 2008 | Finding regional co-location patterns for sets of continuous variables in spatial datasetsabstractThis paper proposes a novel framework for mining regional co-location patterns with respect to sets of continuous variables in spatial datasets. The goal is to identify regions in which multiple continuous variables with values from the wings of their statistical distribution are co-located. A co-location mining framework is introduced that operates in the continuous domain and which views regional co-location mining as a clustering problem in which an externally given fitness function has to be maximized. Interestingness of co-location patterns is assessed using products of z-scores of the relevant continuous variables. The proposed framework is evaluated by a domain expert in a case study that analyzes Arsenic contamination in Texas water wells centering on regional co-location patterns. Our approach is able to identify known and unknown regional co-location patterns, and different sets of algorithm parameters lead to the characterization of Arsenic distribution at different scales. Moreover, inconsistent colocation sets are found for regions in South Texas and West Texas that can be clearly attributed to geological differences in the two regions, emphasizing the need for regional co-location mining techniques. Moreover, a novel, prototype-based region discovery algorithm named CLEVER is introduced that uses randomized hill climbing, and searches a variable number of clusters and larger neighborhood sizes. Christoph F. Eick, Rachana Parmar, Wei Ding 0003, Tomasz F. Stepinski, Jean-Philippe Nicot |
GIS | 3 |
| 2008 | Discovering controlling factors of geospatial variablesabstractEfficient means of determining factors controlling spatial distribution of an environmental class variable are of significant interest in Earth science. In this paper, we present a method for automated discovery of controlling factors by mining for emerging patterns in a database constructed from the fusion of several explanatory datasets. We introduce a new definition of pattern support to account for spatial character of the data and systematically evaluate the effectiveness of our technique using a real-world application pertaining to density of vegetation cover. Experimental results show that our method can successfully identify controlling factors for the presence of high vegetation cover. Tomasz F. Stepinski, Wei Ding 0003, Christoph F. Eick |
GIS | 2 |
| 2008 | Towards Region Discovery in Spatial Datasets
Wei Ding 0003, Rachsuda Jiamthapthaksin, Rachana Parmar, Tomasz F. Stepinski, Christoph F. Eick |
PAKDD | 1 |
| 2006 | A Framework for Regional Association Rule Mining in Spatial DatasetsabstractThe immense explosion of geographically referenced data calls for efficient discovery of spatial knowledge. One of the special challenges for spatial data mining is that information is usually not uniformly distributed in spatial datasets. Consequently, the discovery of regional knowledge is of fundamental importance for spatial data mining. This paper centers on discovering regional association rules in spatial datasets. In particular, we introduce a novel framework to mine regional association rules relying on a given class structure. A reward-based regional discovery methodology is introduced, and a divisive, grid-based supervised clustering algorithm is presented that identifies interesting subregions in spatial datasets. Then, an integrated approach is discussed to systematically mine regional rules. The proposed framework is evaluated in a real-world case study that identifies spatial risk patterns from arsenic in the Texas water supply. Wei Ding 0003, Christoph F. Eick, Jing Wang 0007, Xiaojing Yuan |
ICDM | 1 |
| 2003 | Icon-based Visualization of Large High-Dimensional DatasetsabstractHigh dimensional data visualization is critical to data analysts since it gives a direct view of original data. We present a method to visualize large amount of high dimensional data. We divide dimensions of data into several groups. Then, we use one icon to represent each group, and associate visual properties of each icon with dimensions in each group. A high dimensional data record will be represented by multiple different types of icons located in the same position. Furthermore, we use summary icons to display local details of viewer's interests and the whole data set at meantime. We show its effectiveness and efficiency through a case study on a real large data set. Ping Chen 0001, Chenyi Hu, Wei Ding 0003, Heloise Lynn, Yves Simon |
ICDM | 3 |