EDBT 2026 Demo / reviewers in the wild / expert
Rafal A. Angryk
dblp:25/2162
· DBLP profile ↗
56ranked-venue papers in the field
2as first author
6since 2021 · last 2024
0000-0001-9598-8207ORCID · verified
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 32 (1 first)Database Systems & Data Management · 11Data Mining & Knowledge Discovery · 8Other / Interdisciplinary · 4 (1 first)Information Retrieval & Web Search · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Embedding Ordinality to Binary Loss Function for Improving Solar Flare ForecastingabstractSeveral natural phenomena, such as floods, earth-quakes, volcanic eruptions, or extreme space weather events often come with severity indexes. While these indexes, whether linear or logarithmic are vital, data-driven predictive models for these events rather use a fixed threshold. In this paper, we explore encoding this ordinality to enhance the performance of data-driven models, with specific application in solar flare forecasting. The prediction of solar flares is commonly approached as a binary forecasting problem, categorizing events as either Flare (FL) or No-Flare (NF) based on a chosen threshold (e.g., >C-class, > M-class, or >X-class). However, this binary formulation overlooks the inherent ordinality between the sub-classes within each binary class (FL and NF). In this paper, we propose a novel loss function aimed at optimizing the binary flare prediction problem by embedding the intrinsic ordinal flare characteristics into the binary cross-entropy (BCE) loss function. This modification is intended to provide the model with better guidance based on the ordinal characteristics of the data and improve the overall performance of the models. For our experiments, we employ a ResNet34-based model with transfer learning to predict 2:M-class flares by utilizing the shape-based features of magnetograms of active region (AR) patches spanning from -90° to +90°of solar longitude as our input data. We use a composite skill score (CSS) as our evaluation metric, which is calculated as the geometric mean of the True Skill Score (TSS) and the Heidke Skill Score (HSS) to rank and compare our models' performance. The primary contributions of this work are as follows: (i) We introduce a novel approach to encode ordinality into a binary loss function showing an application to solar flare prediction, (ii) We enhance solar flare forecasting by enabling flare predictions for each AR across the entire solar disk, without any longitudinal restrictions, and evaluate and compare performance. (iii) Our candidate model, optimized with the proposed loss function, shows an improvement of (~17%, (~14%, and (~13% for AR patches within ±30°, ±60°, and ±90° of solar longitude, respectively in terms of CSS, when compared with standard BCE. Additionally, we demonstrate the ability to issue flare forecasts for ARs in near-limb regions (regions between ±60° to ±90°) with a CSS=0.34 (TSS=0.50 and HSS=0.23), expanding the scope of AR-based models for solar flare prediction. This advances the reliability of solar flare forecasts, leading to more effective prediction capabilities. Chetraj Pandey, Anli Ji, Jinsu Hong, Rafal A. Angryk, Berkay Aydin |
DSAA | 4 |
| 2024 | Advancing Solar Flare Prediction Using Deep Learning with Active Region Patches
Chetraj Pandey, Temitope Adeyeha, Jinsu Hong, Rafal A. Angryk, Berkay Aydin |
ECML/PKDD (10) | 4 |
| 2023 | Exploring Deep Learning for Full-disk Solar Flare Prediction with Empirical Insights from Guided Grad-CAM ExplanationsabstractThis study progresses solar flare prediction research by presenting a full-disk deep-learning model to forecast $\geq\mathrm{M}$-class solar flares and evaluating its efficacy on both central (within ±70°) and near-limb beyond ±70°) events, showcasing qualitative assessment of post hoc explanations for the model’s predictions, and providing empirical findings fro human-centered quantitative assessments of these explanations. Our model is trained using hourly full-disk line-of-sight magnetogram images to predict $\geq{\mathrm{M}}$-class solar flares within the subsequent 24-hour prediction window. Additionally, we apply the Guided Gradient-weighted Class Activation Mapping (Guided Grad-CAM) attribution method to interpret our model’s predictions and evaluate the explanations. Our analysis unveils that full-disk solar flare predictions correspond with active region characteristics. The following points represent the most important findings of our study: ❨1❩ Our deep learning models achieved an average true skill statistic (TSS) of $\sim 0.51$ and a Heidke skill score (HSS) of $\sim.38$, exhibiting skill to predict solar flares where for central locations the average recall is $\sim 0.75$ (recall values for X- and M-class are 0.95 and 0.73 respectively) and for the near-limb flares the average recall is $\sim 0.52$ (recall values for X- and M- class are 0.74 and 0.50 respectively); ❨2❩ qualitative examination of the model’s explanations reveals that it discerns and leverages features linked to active regions in both central and near-limb locations within full-disk magnetograms to produce respective predictions. In essence, our models grasp the shape and texture-based properties of flaring active regions, even in proximity to limb areas—a novel and essential capability with considerable significance for operational forecasting systems. Chetraj Pandey, Anli Ji, Trisha Nandakumar, Rafal A. Angryk, Berkay Aydin |
DSAA | 4 |
| 2022 | TS-MIoU: A Time Series Similarity Metric Without Mapping
Azim Ahmadzadeh, Krishna Rukmini Puthucode, Ruizhe Ma, Rafal A. Angryk |
ECML/PKDD (6) | 5 |
| 2021 | Solar Flare Forecasting with Deep Neural Networks using Compressed Full-disk HMI MagnetogramsabstractPrediction of solar flares is a challenging problem in space weather forecasting that has piqued the interest of many researchers in recent years due to improved data availability and the advancements in the field of machine learning and deep learning. In this paper, we present a solution to full-disk flare prediction using compressed magnetogram images, which was performed by training a set of Convolutional Neural Networks to perform operations-ready flare forecasts. W e s elected two prediction modes, which are both binary for predicting the occurrence of ≥M1.0 and ≥C1.0 class flares within the next 24 hours. For this, we use a simple yet powerful pre-trained AlexNet model and we collect compressed images derived from solar magnetograms provided by the Helioseismic and Magnetic Imager (HMI) instrument onboard Solar Dynamics Observatory (SDO). We followed two time-segmented cross-validation strategies: chronological and non-chronological, to effectively understand the predictive skill of our models. We also trained our models using data-augmentation and oversampling to address the existing class-imbalance issue and used true skill statistic (TSS) and Heidke skill score (HSS) as metrics to compare and evaluate. The major results of this study are (1) we successfully implemented an efficient and effective full-disk flare predictor ready for operational forecasting using 8-bit compressed images of solar magnetograms without further preprocessing; (2) Our candidate model achieves an average TSS of 0.47±0.06 for ≥M1.0 mode and 0.63±0.05 for ≥C1.0 mode, and HSS of 0.35±0.05 for ≥M1.0 and 0.62±0.05 for ≥C1.0 mode. Our experimental evaluation also suggests that training a flare prediction model is heavily influenced by the sampling strategies involved due to the imbalanced nature of the datasets and predicting ≥M1.0 class flares is a more challenging task compared to ≥C1.0 ones. Chetraj Pandey, Rafal A. Angryk, Berkay Aydin |
IEEE BigData | 2 |
| 2021 | Spatiotemporal event sequence discovery without thresholds
Berkay Aydin, Soukaina Filali Boubrahimi, Ahmet Küçük, Bita Nezamdoust, Rafal A. Angryk |
GeoInformatica | 5 |
| 2020 | On the Mining of the Minimal Set of Time Series Data ShapeletsabstractShapelets, also known as motifs, are time series sequences that have the property of discriminating between time series classes. Lately, shapelets studies have gained a lot of momentum due to their interpretable nature. As opposed to traditional time series classifiers, shapelet-based learners provide a visual representation of the pattern that triggers the classification decision. One of the most challenging issues of shapelet-based classifiers is the generation of a large number of shapelet outputs. To the best of our knowledge, this is the first effort that addresses the high numerosity problem of mined shapelets issue by mining the minimal set of discriminative shapelets for time series data. We propose a new shapelet mining learner, 1DCNN, that has the property of learning shapelets of different lengths using a black-box neural network model. 1DCNN optimizes the entire classification schema by learning the shapes of the representative patterns. Our proposed model uses network pruning to sparsify the network and keep only the most discriminative shapelets without compromising the classification accuracy. We validated our model using 59 real-world time series datasets from the UCR repository. Our experimental results show the effectiveness and efficiency of our approach in comparison with other competing baselines models. For fairness purposes, we did not compare 1DCNN with ensemble based approaches that encapsulates many learners. Our results show that the performance of our model is superior to all other baselines pertaining to the shapelet-based classifier category, with up to 95% less Floating Points Operations per Second (FLOPs) required by the network. Soukaina Filali Boubrahimi, Shah Muhammad Hamdi, Ruizhe Ma, Rafal A. Angryk |
IEEE BigData | 4 |
| 2020 | A Framework for Detecting Polarity Inversion Lines from Longitudinal MagnetogramsabstractMagnetic polarity inversion line (PIL) in solar active regions have been recognized as essential features for the occurrence of solar flares and the prediction of the flaring phenomenon. In this work, we provide a software framework that detects PILs from the line-of-sight (LoS) or the radial component of the magnetic field vector in active region magnetogram patches. The PIL detection procedure is based on an edge detection technique along with magnetic field strength and PIL size filter. First, we identify positive and negative polarity regions with a magnetic field strength threshold. Then, we utilize the Canny edge detector and morphological operations to both positive and negative regions to identify coarse PILs. Finally, we generate PILs by applying magnetic field strength and PIL size filter to the coarse PILs as mentioned above. Moreover, we provide feature extraction functions to obtain the properties of PILs (i.e., PIL size, the area of polarity inversion, the masked unsigned flux of enclosing PIL, convexity, eigenvalues, fractal dimension, and Hu moments of PIL shape), and produce three PIL-related binary masks (i.e., PIL, the region of polarity inversion, and the convex hull of PIL) for each Longitudinal magnetogram patch. We also provide a qualitative evaluation of our module. The detection results for a well known active region (NOAA AR 11158) are shown as proof of concept. Xumin Cai, Berkay Aydin, Anli Ji, Manolis K. Georgoulis, Rafal A. Angryk |
IEEE BigData | 5 |
| 2020 | Local Outlier Detection for Multi-type Spatio-temporal TrajectoriesabstractOutlier detection has become one of the core tasks in spatio-temporal data mining. It plays an essential role in data quality improvement for the machine learning models and recognizing the anomalous patterns, which may remarkably deviate from expected patterns among the trajectory datasets. In this work, we propose a clustering-based technique to detect local outliers in trajectory datasets by utilizing spatial and temporal attributes of moving objects. This local outlier detection involves three phases. In the first phase, we apply a temporal partition procedure to divide the raw trajectory into multiple trajectory segments and extract trajectory features from spatial and temporal attributes for each trajectory segment. Then, we generate template features of trajectory segments by applying a clustering schema in the second phase. Finally, we use the abnormal score - a novel dissimilarity measure, which quantifies the disparity among the query and template trajectory segments in terms of trajectory features and hence determines the local outliers based on the distribution of abnormal score. To demonstrate the effectiveness of our method, we conduct three case studies on the real-life spatio-temporal trajectory datasets from the solar astroinformatics domain (i.e., solar active regions, coronal mass ejections, polarity inversion lines (PIL)). Our experimental results show that our local outlier detection approach can effectively discover the erroneous reports from the reporting module and abnormal phenomenon in various spatio-temporal trajectory datasets. Xumin Cai, Berkay Aydin, Saurabh Maydeo, Anli Ji, Rafal A. Angryk |
IEEE BigData | 5 |
| 2020 | On The Effectiveness of Imaging of Time Series for Flare Forecasting ProblemabstractForecasting the occurrence of solar flares is a typical 21st century rare-event classification task. Over the past two decades, many studies have implemented various techniques and approaches for classification of strong and weak solar flares. The release of the recent flare forecasting benchmark dataset, named SWAN-SF, has opened the door for taking advantage of multivariate time series (MVTS) of pre-flare magnetic fields’ activity in order to potentially achieve higher performance and increase the robustness of the new forecasting models. In this study, we take a new approach and explore the effectiveness of imaging algorithms on the time series. We convert MVTS data into multi-channel image data using Gramian Angular Fields (GAF) and Markov Transition Fields (MTF) to explore the proven strength of deep neural networks in the Image Processing domain, on flares’ MVTS data. Inspired by GAF, we propose another imaging matrix that our experiments show that it significantly improves the performance of the CNN, by 120% in terms of TSS and 440% in terms of HSS2. Finally, we juxtapose two approaches to tackle the flare forecasting problem: one, to utilize a time-series specific Support Vector Classifier for classification of flares, and the other, to train a Convolutional Neural Network (CNN) on the derived images using GAF, MTF, and our modified GAF. Anli Ji, Pavan Ajit Babajiyavar, Azim Ahmadzadeh, Rafal A. Angryk |
IEEE BigData | 5 |
| 2020 | Deep Neural Network-based Active Region Magnetogram Patch Super ResolutionabstractImage super-resolution is a branch of image processing that is concerned with enhancing the spatial resolution and quality of images by learning the intrinsic details and relations between the lower resolution input and the higher resolution output images. It is widely accepted as an ill-posed problem, which has seen tremendous advancements with deep learning-based models. In this work, we present two super resolution models, Sub-Pixel Convolutional Neural Network (CNN) and Enhanced Deep Residual Networks (ResNet), which can be used for improving the spatial resolution of solar magnetograms. While the ill-posed nature of problem is still a challenge, there are several application areas, including space weather prediction, which can greatly benefit from the improved spatial resolution of solar magnetograms. Along with classical raster inputs we try to improve the model objective by giving HMI Active Region Patches. We show that through our experimental evaluation our models perform better than baselines and CNN-based super resolution model provides viable results for magnetogram super resolution. Mohammed Shoebuddin Habeeb, Berkay Aydin, Azim Ahmadzadeh, Manolis K. Georgoulis, Rafal A. Angryk |
IEEE BigData | 5 |
| 2020 | Solar Line-of-Sight Magnetograms Super-Resolution Using Deep Neural NetworksabstractImage super-resolution is a branch of image processing that is concerned with enhancing the spatial resolution and quality of images by learning the intrinsic details and relations between the lower resolution input and the higher resolution output images. It is widely accepted as an ill-posed problem, which has seen tremendous advancements with deep learning based models. In this work, we present two magnetogram super resolution models, Sub-Pixel Convolutional Neural Network (CNN) and Enhanced Deep Residual Networks (ResNet), which can be used for improving the spatial resolution of solar magnetograms. While the ill-posed nature of problem is still a challenge, there are several application areas, including space weather prediction, which can greatly benefit from the improved spatial resolution of solar magnetograms. We show that through our experimental evaluation our models perform better than baselines and Sub-Pixel CNN super resolution model provides viable results for magnetogram super resolution. Mohammed Shoebuddin Habeeb, Berkay Aydin, Azim Ahmadzadeh, Manolis K. Georgoulis, Rafal A. Angryk |
IEEE BigData | 5 |
| 2020 | First Steps Toward Synthetic Sample Generation for Machine Learning Based Flare ForecastingabstractThe imbalanced class problem is intrinsic to solar flare forecasting, as are other issues we find in data-driven forecasting problems that are often hidden within an imbalanced dataset. One method of dealing with imbalanced data is to balance the data by using synthetic oversampling to create synthetic examples of the minority class. Though synthetic oversampling techniques have been applied to problems in medicine, finance, security, and other areas, we have not seen these approaches used in solar flare forecasting. We investigate two methods of synthetic oversampling, Rapidly Converging Gibbs Sampler (RACOG) and Synthetic Minority Oversampling Technique (SMOTE). We devise three naive synthetic oversampling techniques for compar-ison. We rely on data provided by the Space Weather ANalytics for Solar Flares (SWAN- SF) benchmark dataset. Our results indicate that synthetic oversampling can be effective for machine learning based solar flare forecasting. Maxwell Hostetter, Rafal A. Angryk |
IEEE BigData | 2 |
| 2020 | All-Clear Flare Prediction Using Interval-based Time Series ClassifiersabstractAn all-clear flare prediction is a type of solar flare forecasting that puts more emphasis on predicting non-flaring instances (often relatively small flares and flare quiet regions) with high precision while still maintaining valuable predictive results. While many flare prediction studies do not address this problem directly, all-clear predictions can be useful in operational context. However, in all-clear predictions, finding the right balance between avoiding false negatives (misses) and reducing the false positives (false alarms) is often challenging. Our study focuses on training and testing a set of interval-based time series named Time Series Forest (TSF). These classifiers will be used towards building an all-clear flare prediction system by utilizing multivariate time series data. Throughout this paper, we demonstrate our data collection, predictive model building and evaluation processes, and compare our time series classification models with baselines using our benchmark datasets. Our results show that time series classifiers provide better forecasting results in terms of skill scores, precision and recall metrics, and they can be further improved for more precise all-clear forecasts by tuning model hyperparameters. Anli Ji, Berkay Aydin, Manolis K. Georgoulis, Rafal A. Angryk |
IEEE BigData | 4 |
| 2019 | Challenges with Extreme Class-Imbalance and Temporal Coherence: A Study on Solar Flare DataabstractIn analyses of rare-events, regardless of the domain of application, class-imbalance issue is intrinsic. Although the challenges are known to data experts, their explicit impact on the analytic and the decisions made based on the findings are often overlooked. This is in particular prevalent in interdisciplinary research where the theoretical aspects are sometimes overshadowed by the challenges of the application. To show-case these undesirable impacts, we conduct a series of experiments on a recently created benchmark data, named Space Weather ANalytics for Solar Flares (SWAN-SF). This is a multivariate time series dataset of magnetic parameters of active regions. As a remedy for the imbalance issue, we study the impact of data manipulation (undersampling and oversampling) and model manipulation (using class weights). Furthermore, we bring to focus the auto-correlation of time series that is inherited from the use of sliding window for monitoring flares' history. Temporal coherence, as we call this phenomenon, invalidates the randomness assumption, thus impacting all sampling practices including different cross-validation techniques. We illustrate how failing to notice this concept could give an artificial boost in the forecast performance and result in misleading findings. Throughout this study we utilized Support Vector Machine as a classifier, and True Skill Statistics as a verification metric for comparison of experiments. We conclude our work by specifying the correct practice in each case, and we hope that this study could benefit researchers in other domains where time series of rare events are of interest. Azim Ahmadzadeh, Maxwell Hostetter, Berkay Aydin, Manolis K. Georgoulis, Dustin Kempton, Sushant S. Mahajan, Rafal A. Angryk |
IEEE BigData | 7 |
| 2019 | Toward Filament Segmentation Using Deep Neural NetworksabstractWe use a well-known deep neural network framework, called Mask R-CNN, for identification of solar filaments in full-disk H-$\alpha$ images from Big Bear Solar Observatory (BBSO). The image data, collected from BBSO's archive, are integrated with the spatiotemporal metadata of filaments retrieved from the Heliophysics Events Knowledgebase (HEK) system. This integrated data is then treated as the ground-truth in the training process of the model. The available spatial metadata are the output of a currently running filament-detection module developed and maintained by the Feature Finding Team; an international consortium selected by NASA. Despite the known challenges in the identification and characterization of filaments by the existing module, which in turn are inherited into any other module that intends to learn from such outputs, Mask R-CNN shows promising results. Trained and validated on two years worth of BBSO data, this model is then tested on the three following years. Our case-by-case and overall analyses show that Mask R-CNN can clearly compete with the existing module and in some cases even perform better. Several cases of false positives and false negatives, that are correctly segmented by this model are also shown. The overall advantages of using the proposed model are two-fold: First, deep neural networks' performance generally improves as more annotated data, or better annotations are provided. Second, such a model can be scaled up to detect other solar events, as well as a single multi-purpose module. The results presented in this study introduce a proof of concept in benefits of employing deep neural networks for detection of solar events, and in particular, filaments. Azim Ahmadzadeh, Sushant S. Mahajan, Dustin Kempton, Rafal A. Angryk, Shihao Ji 0001 |
IEEE BigData | 4 |
| 2019 | An Application of Spatio-temporal Co-occurrence Analyses for Integrating Solar Active Region Data from Multiple Reporting ModulesabstractSpatio-temporal co-occurrence analysis captures the spatial and temporal relations between events that occur at the same time and location. In this paper, we utilize spatio-temporal co-occurrence relations to integrate solar active region (AR) data detected and reported by three feature recognition methods, namely, human labeling by forecasters at National Oceanic and Atmospheric Administration (NOAA), Spaceweather HMI Active Region Patch (SHARP) detection pipeline, and Spatial Possibilistic Clustering Algorithm (SPoCA). We determine the associations between individual reports by identifying the spatio-temporal co-occurrences among the reports from these modules. We compare our findings with the data from the Joint Science Operations Center (JSOC), analyzing the discrepancies in different circumstances. We found 105 SHARP series not properly associated with the NOAA-labeled ARs. In the end, we provide detailed movement analyses for the AR trajectories, create an updated SHARP-to-NOAA AR associations, that is crucial for space weather predictions utilizing magnetic field information, and make the ternary associations between SHARP, NOAA ARs, and SPoCA ARs available to the public. Xumin Cai, Berkay Aydin, Manolis K. Georgoulis, Rafal A. Angryk |
IEEE BigData | 4 |
| 2019 | Understanding the Impact of Statistical Time Series Features for Flare Prediction AnalysisabstractMachine learning-based space weather analytics has attracted much attention due to the potential damages that can be caused by the extreme space weather events. Using a recently released data benchmark, named SWAN-SF, designed for solar flare forecasting based on the pre-flare time series of solar magnetic field parameters, we conduct a case study on the impacts of statistical features derived from the multivariate time series. We investigate the relationship between the number of needed statistical features extracted from the multi-variate time series and the performance of flare forecast models. To that end, we employ random forest and mean decrease impurity to determine a feature selection methodology along with an evaluation procedure. The proposed evaluation method delivers a balance between the two frequently used metrics in this domain, namely True Skill Statistic and Heidke Skill Score. Our approach allows to introduce a generic feature selection and evaluation procedure that is independent from the minor and often obscured decisions that must be made for having a binary forecast model, while presenting interpretable and actionable tools that can help non-data experts make more informed and realistic decisions. Maxwell Hostetter, Azim Ahmadzadeh, Berkay Aydin, Manolis K. Georgoulis, Dustin Kempton, Rafal A. Angryk |
IEEE BigData | 6 |
| 2019 | Solar Pre-Flare Classification with Time Series ProfilingabstractSpace weather encapsulates the impact of variable solar activity on the vicinity of Earth and elsewhere in the solar system. A major agent of space weather, with significant effort already devoted to its prediction, is solar flares. Most existing analysis in this direction focus on the instantaneous (point-in-time) magnitude of various pre-flare parameters in flare host locations, solar active regions. Nonetheless, a recent trend places data-intensive studies, focusing on the pre-flare time series of these parameters, to the forefront. We take on this task in this study, focusing on the shape of pre-flare active region parameter time series by introducing a data-driven class profiling and clustering of these time series. We rely on data provided by the Space Weather ANalytics for Solar Flares (SWAN-SF) benchmark dataset. Our results indicate some potentially interesting temporal patterns that are unrelated to parameter magnitudes and may be used, both in tandem and independently from magnitudes, for future flare forecasting efforts. Our analysis also provides flexibility to define custom flare classes relying on pre-flare time series behavior and relate them to the existing, conventional NOAA / GOES flare classes. Ruizhe Ma, Azim Ahmadzadeh, Soukaina Filali Boubrahimi, Manolis K. Georgoulis, Rafal A. Angryk |
IEEE BigData | 5 |
| 2019 | Tensor Decomposition-based Node EmbeddingabstractIn recent years, node embedding algorithms, which learn low dimensional vector representations for nodes in a graph, have been one of the key research interests of the graph mining community. The existing algorithms either rely on computationally expensive eigendecomposition of the large matrices, or require tuning of the word embedding-based hyperparameters as a result of representing the graph as a node sequence similar to the sentences in a document. Moreover, the latent features produced by these algorithms are hard to interpret. In this paper, we present Tensor Decomposition-based Node Embedding (TDNE), a novel model for learning node representations for arbitrary types of graphs: undirected, directed, and/or weighted. Our model preserves the local and global structural properties of a graph by constructing a third-order tensor using the k-step transition probability matrices and decomposing the tensor through CANDECOMP/PARAFAC (CP) decomposition in order to produce an interpretable, low dimensional vector space for the nodes. Our experimental evaluation using two well-known social network datasets proves TDNE to be interpretable with respect to the understandability of the feature space, and precise with respect to the network reconstruction. Shah Muhammad Hamdi, Soukaina Filali Boubrahimi, Rafal A. Angryk |
CIKM | 3 |
| 2019 | Interpretable Feature Learning of Graphs using Tensor DecompositionabstractIn recent years, node embedding algorithms, which learn low dimensional vector representations for nodes in a graph, have been one of the key research interests of the graph mining community. The existing algorithms either rely on computationally expensive eigendecomposition of the large matrices, or require tuning of the word embedding-based hyperparameters as a result of representing the graph as a node sequence similar to the sentences in a document. Moreover, the latent features produced by these algorithms are hard to interpret. In this paper, we present two novel tensor decomposition-based node embedding algorithms, that can learn node features from arbitrary types of graphs: undirected, directed, and/or weighted, without relying on eigendecomposition or word embedding-based hyperparameters. Both algorithms preserve the local and global structural properties of the graph by using k-step transition probability matrices to construct third-order multidimensional arrays or tensors and perform CANDECOMP/PARAFAC (CP) decomposition in order to produce an interpretable and low dimensional vector space for the nodes. Our experiments encompass different types of graphs (undirected/directed, unweighted/weighted, sparse/dense) of different domains such as social networking and neuroscience. Our experimental evaluation proves our models to be interpretable with respect to the understandability of the feature space, precise with respect to the network reconstruction and link prediction, and accurate with respect to node classification and graph classification. Shah Muhammad Hamdi, Rafal A. Angryk |
ICDM | 2 |
| 2018 | Heuristics Significance of Neuro-Ensemble-based Time Series ClassificationabstractEnsemble learning is a popular paradigm for improving the predictive performance of individual classifiers. In this work, we approach the problem of ensemble learning from an optimization perspective applied on time series data. We propose Neuro-Ensemble, a classifier fusion model based on a shallow Multi-Layer Perceptron (MLP) meta-learner. The neural network learns the expertise of each classifier in the ensemble and optimizes the entire classification schema based on class-level expertise weights. We defined and compared three different classifiers ordering heuristics: Random, BestFirst, and BestLast, that we coupled with our new ensemble technique. We validated our Neuro-Ensemble on 43 real-world time series datasets from the UCR repository. Our experimental results shows the competitiveness of our approach with respect to Evaluation and Selection and that the use of heuristics with Neuro-Ensemble model is insignificant. Soukaina Filali Boubrahimi, Rafal A. Angryk |
IEEE BigData | 2 |
| 2018 | Segmentation of Time Series in Improving Dynamic Time WarpingabstractSince its introduction to the computer science community, the Dynamic Time Warping (DTW) algorithm has demonstrated good performance with time series data. While this elastic measure is known for its effectiveness with time series sequence comparisons, the possibility of pathological warping paths weakens the algorithms potential considerably. Techniques centering on pruning off impossible mappings or lowering data dimensions such as windowing, slope weighting, step pattern, and approximation have been proposed over the years to reduce the possibility of pathological warping paths with Dynamic Time Warping. However, because the current DTW improvement techniques are mostly global methods, they are either limited in effect or limit the warping path excessively. We believe segmenting time series at significant feature points will alleviate some of the pathological warpings, and at the same time allowing us to obtain more intuitive warpings. Our heuristic approaches the problem from the human perspective of sequence comparison: by identifying global similarity before local similarities. We use easily identifiable peaks as the significant feature. The final distance is the DTW distance sum of all segments of time series. In this paper, we explore the impact of different peak identification parameters on Dynamic Time Warping and demonstrate how segmentation can help to avoid pathological warpings. Ruizhe Ma, Azim Ahmadzadeh, Soukaina Filali Boubrahimi, Rafal A. Angryk |
IEEE BigData | 4 |
| 2018 | Time Series Distance Density Cluster with Statistical Preprocessing
Ruizhe Ma, Soukaina Filali Boubrahimi, Rafal A. Angryk |
DaWaK | 3 |
| 2018 | Neuro-Ensemble for Time Series Data ClassificationabstractCombining a set of classification algorithms is a powerful technique in improving the accuracy of individual classifiers. There are two main paradigms in combining classifiers: classifier selection, where each classifier is considered as an expert in some local area of the feature space, and classifier fusion, where all classifiers are trained over the entire feature space and they are considered as competitive and complementary to each other. In this paper, we propose a new ensemble technique, NeuroEnsemble, that follows the classifier fusion paradigm applied on time series data. The Neuro-Ensemble exploits the idea that different classifiers participating in the ensemble have varying degrees of expertise on learning different class labels and it optimizes the ensemble using a shallow Multi-Layer Perceptron (MLP) based meta-learner to capture the expertise of individual classifiers. Every neuron in the MLP represents a classifier that contributes with a vote and performs activation and state computations. This work is the first attempt to train a neural network for learning the expertise of each classifier in an ensemble and optimize the entire classification schema based on class-level expertise weights. We validated our Neuro-Ensemble on 43 real-world time series datasets from the UCR repository. Our experimental results show the effectiveness and efficiency of our approach in comparison with individual baseline learners and ensemble techniques. Soukaina Filali Boubrahimi, Ruizhe Ma, Rafal A. Angryk |
DSAA | 3 |
| 2017 | Improving the functionality of tamura directionality on solar imagesabstractDirectionality as a textural parameter is a great tool in many image-based tasks including the automated analysis of medical, geographical, or astronomical images. Tamura directionality is a popular parameters that measures the degree of texture directionality of images. In this study, we review the original idea and also address some ambiguities present in the original definition that play an important role in the effectiveness of this parameter. Then, we propose different ideas to attack each of the addressed ambiguities and find an optimal setting that maximizes the effectiveness of this parameter. Lastly, we evaluate the effect of our modifications on a large collection of solar images captured by the Solar Dynamic Observatory mission. We show that such modifications significantly improve the effectiveness of this parameter. Although, different settings might be obtained for images with different sort of texture, the methodology for finding an optimal setting remains the same regardless. The goal of our study is to make sure that parameters that are being frequently used in several areas, especially in the interdisciplinary research between solar physicists and computer scientists, are calculated accurately and effectively. Azim Ahmadzadeh, Dustin Kempton, Michael A. Schuh, Rafal A. Angryk |
IEEE BigData | 4 |
| 2017 | Parallel computation of magnetic field parameters from HMI active region patchesabstractMagnetic parameters are crucial to analyze and forecast solar events such as solar flares and coronal mass ejections. These parameters can be computed from vector magnetogram data collected from the solar surface. We propose a tool to compute these magnetic parameters from solar image data files in parallel using multi-threading and GPUs. The architecture of the proposed solar magnetic parameter computational tool is discussed in detail. We use the images from Spaceweather HMI Active Region Patches (SHARP) data available from JSOC to perform our experiments. We perform exhaustive analysis on the parameters used in the architecture. Run times of magnetic parameter generation are compared across serial execution codes (implemented in Python and C++) and parallel code (implemented in python and C++ using OpenACC). We conclude by showing that parallel computation of magnetic parameters using GPUs is much faster when compared to serial execution. Sunitha Basodi, Berkay Aydin, Rafal A. Angryk |
IEEE BigData | 3 |
| 2017 | On the prediction of >100 MeV solar energetic particle events using GOES satellite dataabstractSolar energetic particles are a result of intense solar events such as solar flares and Coronal Mass Ejections (CMEs). These latter events all together can cause major disruptions to spacecraft that are in Earth's orbit and outside of the magnetosphere. In this work we are interested in establishing the necessary conditions for a major geo-effective solar particle storm immediately after a major flare, namely the existence of a direct magnetic connection. To our knowledge, this is the first work that explores not only the correlations of GOES X-ray and proton channels, but also the correlations that happen across all the proton channels. We found that proton channels autocorrelations and cross-correlations may also be precursors to the occurrence of an SEP event. In this paper, we tackle the problem of predicting >100 MeV SEP events from a multivariate time series perspective using easily interpretable decision tree models. Soukaina Filali Boubrahimi, Berkay Aydin, Petrus C. Martens, Rafal A. Angryk |
IEEE BigData | 4 |
| 2017 | A time series classification-based approach for solar flare predictionabstractSolar flare prediction is an important task because of their potential impacts on both space and terrestrial infrastructure. This prediction task can be modeled as a binary classification between flaring and non-flaring Active Regions. Previous works on flare prediction focused on representing flaring and non-flaring Active Region examples in vector space, where the feature space was found from the Active Region magnetic field parameters. We extract time series samples of these Active Region parameters and present a flare prediction method based on the k-NN classification of the univariate time series. We find that, for our classification task, using a statistical summarization on the time series of a single Active Region parameter, called total unsigned current helicity, outperforms the use of all Active Region parameters at a single instant of time. Additionally, we present a data model of the flaring/non-flaring Active Regions using multivariate time series. Shah Muhammad Hamdi, Dustin Kempton, Ruizhe Ma, Soukaina Filali Boubrahimi, Rafal A. Angryk |
IEEE BigData | 5 |
| 2017 | Multi-wavelength solar event detection using faster R-CNNabstractThe automated detection of solar phenomena from high resolution images became important for solar physics researchers after the launch of the Solar Dynamics Observatory. The solar event detections help researchers find and track relevant regions and eventually facilitate the discovery of trends and patterns between different types of events. We address the problem of automated detection of solar events from multi-wavelength solar images using deep learning-based Faster R-CNN method. Earlier work on solar event detection primarily use the observed models to locate the events on the solar images in an unsupervised fashion and each detection algorithm targets specific solar event type. Here, we will present a data-driven methodology to facilitate solar physics research. While this work presents a proof of concept that supervised deep learning-based event detection methodology for solar images is possible, our results show that data-driven detection using deep learning can successfully detect the multiple types of solar events and it can be used for validating the results from existing modules or training modules to detect new event types. Ahmet Küçük, Berkay Aydin, Rafal A. Angryk |
IEEE BigData | 3 |
| 2017 | Solar flare prediction using multivariate time series decision treesabstractSpace Weather is of rising importance in scientific discipline that describes the way in which the Sun and space impact a myriad of activities down on Earth as well as the safety of the space crew members on board of the space stations. Consequently, it is imperative to better quantify the risk of future space weather events. Most of the flare prediction models in literature use physical parameters of the potentially flaring active regions during a limited interval to gain insights on whether a flare will happen or not. This limits our perception of how an event evolves for an extended duration across multiple parameters. In this paper we followed a data-driven approach to address the problem of flare prediction from a multivariate time series analysis perspective and attempt to cluster potential flaring active regions by applying Distance Density clustering on individual parameters and further organize the clustering results into a multivariate time series decision tree. We compared different data extraction priors and spans, and ranked the importance for different parameters through univariate clustering. To the best of our knowledge, this is the first attempt to predict solar flares using a tree structure. Ruizhe Ma, Soukaina Filali Boubrahimi, Shah Muhammad Hamdi, Rafal A. Angryk |
IEEE BigData | 4 |
| 2017 | An Integrated Solar Database (ISD) with Extended Spatiotemporal Querying Capabilities
Ahmet Küçük, Berkay Aydin, Soukaina Filali Boubrahimi, Dustin Kempton, Rafal A. Angryk |
SSTD | 5 |
| 2016 | The SMART approach to comprehensive quality assessment of site-based spatial-temporal dataabstractThere is a need for comprehensive solutions to address the challenges of spatio-temporal data quality assessment. Emphasis is often placed on the quality assessment of individual observations from sensors but not on the sensors themselves nor upon site metadata such as location and timestamps. The focus of this paper is on the development and evaluation of a representative and comprehensive, interpolation-based methodology for assessment of spatio-temporal data quality. We call our method the SMART method, short for Simple Mappings for the Approximation and Regression of Time series. When applied to a real-world, meteorological data set, we show that our method outperforms standard interpolators and we identify numerous problematic sites that otherwise would not have been flagged as bad. We further identify sites for which metadata is incorrect. We believe that there are many problems with real data sets like these and, in the absence of an approach like ours, these problems have largely gone unidentified. Our results bring into question the validity of provider-based quality control indicators. In addition to providing a comprehensive solution, our approach is novel for the simple but effective way that it accounts for spatial and temporal variation. Rafal A. Angryk, Douglas E. Galarus |
IEEE BigData | 1 |
| 2016 | Indexing spatiotemporal relations in solar event datasetsabstractWith the advancements in spatiotemporal co-occurrence pattern and event sequence mining algorithms, spatiotemporal knowledge discovery from solar event datasets has been prominent in solar data mining. This work presents an efficient and extensible data access mechanism specifically designed for spatiotemporal relationships among the solar event instances. Previous indexing strategies primarily focus on indexing the trajectories of solar event instances. We propose a graph-based indexing structure for spatiotemporal relationships appearing among the event instances. In our graph structure, the vertices correspond to trajectory-based solar event instances, and the edges represent a particular spatiotemporal relationship. Furthermore, it forms the groundwork for semantic representations of solar event instances. Berkay Aydin, Ahmet Küçük, Rafal A. Angryk |
IEEE BigData | 3 |
| 2016 | Spatio-temporal interpolation methods for solar events metadataabstractThis paper introduces three interpolation methods that enrich complex evolving region trajectories that are captured every day from numerous ground-based and space-based solar observatories. The interpolation module takes a trajectory as its input and generates an enriched trajectory with interpolated time-geometry pairs. we created three different interpolation techniques that are: MBR-Interpolation (Minimum Bounding Rectangle Interpolation), CP-Interpolation (Complex Polygon Interpolation), and FP-Interpolation (Filament Polygon Interpolation). The methods combine K-means clustering algorithm, shape signature representation, and linear interpolation to generate the missing polygons. This is the first research of this kind that attempts to address the problem of solar big data interpolation. Finally, we outline future improvements and opportunities for solar data interpolation. Soukaina Filali Boubrahimi, Berkay Aydin, Dustin Kempton, Rafal A. Angryk |
IEEE BigData | 4 |
| 2016 | Describing solar images with sparse coding for similarity searchabstractIn this work, we present a method of producing image descriptors that is based on max-pooling of sparse codes. We use this method on images from the Solar Dynamics Observatory (SDO). The SDO produces over 70,000 images of the Sun each day, and with so many images being archived, an efficient method for finding similar images in this ever growing dataset is critical. Our method for producing descriptors is advantageous because the results are of a reasonable size for indexing, and are more selective than other methods used in the past. We use sparse coding on learned dictionaries to produce linear decompositions of the input signals. These decompositions are unlike decompositions based on principal component analysis, as we do not impose the constraint that basis vectors be orthogonal. By removing the orthogonal constraint, we are able to more easily adapt the representation to the data, and we show that our initial retrieval results alleviate a problem found to be an issue for this dataset. Specifically, the problem of the immediate temporal neighbor being the most similar in virtually every case. Dustin Kempton, Michael A. Schuh, Rafal A. Angryk |
IEEE BigData | 3 |
| 2016 | A data-driven analysis of interplanetary coronal mass ejecta and magnetic flux ropesabstractScientists have observed the occurrence of two distinctive subsets in Interplanetary Coronal Mass Ejections (ICMEs): magnetic clouds (MCs) and non-magnetic clouds (non-MCs). While we are aware of some of the distinctive features of MCs and non-MCs, we cannot draw a precise line between them. Features such as large magnetic field, low plasma-beta, low proton temperature, etc. suggest when an ICME event is also an MC event, however, this categorization is far from an automated process. In addition to being time-consuming, the results differ depending on the precision of definition. In this paper, we approach the MC and non-MC class distinction from a data analysis perspective and show a data-driven taxonomy of ICME events. We use a time series dataset from the Ulysses spacecraft combined with a list of labeled MC and non-MC events. The time series data are hierarchically clustered with Euclidean distance and Dynamic Time Warping algorithm, and we compare our MC and non-MC clusters with the results from classifications generated by domain experts. Ruizhe Ma, Rafal A. Angryk, Pete Riley |
IEEE BigData | 2 |
| 2016 | On the theory and practice of high-dimensional data indexing with iDistanceabstractAn important and challenging task for modern large-scale and high-dimensional data is k-nearest neighbor (kNN) retrieval. Using iDistance as the current state-of-the-art high-dimensional indexing algorithm, this work discusses the theoretical bounds associated with a distance-based indexing technique and proposes several simple optimizations to improve the efficiency and stability of query retrieval. We then present practical analysis of these bounds and optimizations through experiments on synthetic and real world datasets over a wide variety of data characteristics. Results indicate overall greatly improved retrieval performance and especially promising results in high dimensions. Michael A. Schuh, Rafal A. Angryk |
IEEE BigData | 2 |
| 2016 | SOLEV: a video generation framework for solar events from mixed data sources (demo paper)abstractOne of the main strengths of Geographical Information Systems (GIS) is the analysis of spatial and attributive data. Spatiotemporal interpolation techniques allow the expansion of the collected data to the sites where no samples are available. In the context of GIS, the data, be it interpolated or collected, are visual in nature and hard to understand in raw forms. Visualization of complex evolving region trajectories is often times used as an aid to better understand the data and its underlying patterns. In this work, we created SOLEV, a solar event video generation framework that integrates multiple data sources of solar images. This is the first framework of this kind that not only visualizes spatial solar event boundaries, but also the tracked and interpolated spatiotemporal trajectories they form over time. Soukaina Filali Boubrahimi, Berkay Aydin, Dustin Kempton, Rafal A. Angryk |
SIGSPATIAL/GIS | 4 |
| 2016 | A SMART approach to quality assessment of site-based spatio-temporal dataabstractA significant challenge we face in assessing spatio-temporal data quality is a lack of ground-truth data. Error is by definition the deviation of observation from ground truth. In the absence of ground truth, we depend on our own or provider quality assessment to evaluate our methods. The focus of this paper is the development of a representative, weather-like spatio- temporal dataset and the use of this dataset to develop and evaluate a robust, interpolation-based method for assessment of data quality. We call our method the SMART method, short for Simple Mappings for the Approximation and Regression of Time series. We present this method as a representative approach to demonstrate and overcome the challenges of spatio- temporal data quality assessment. Our results bring into question the validity of provider-based quality control indicators. Douglas E. Galarus, Rafal A. Angryk |
SIGSPATIAL/GIS | 2 |
| 2016 | Mining spatiotemporal co-occurrence patterns in non-relational databases
Berkay Aydin, Vijay Akkineni, Rafal A. Angryk |
GeoInformatica | 3 |
| 2015 | Time-efficient significance measure for discovering spatiotemporal co-occurrences from data with unbalanced characteristicsabstractMining spatiotemporal co-occurrence patterns requires assessing the strength of co-occurrences among the instances of different feature types. Currently, a spatiotemporal version of the Jaccard measure is used for measuring the strength of spatiotemporal co-occurrences. We present an extended spatiotemporal version of the Jaccard measure (J*) that is more relevant and efficient for the task of STCOP mining. We also demonstrate the space and time efficiency of the J* with experimental evaluation. Berkay Aydin, Vijay Akkineni, Rafal A. Angryk |
SIGSPATIAL/GIS | 3 |
| 2014 | Storing Long-Lived Concurrent Schema and Data Versions in Relational Databases
Bob Wall, Rafal A. Angryk |
ADBIS (2) | 2 |
| 2014 | Spatiotemporal indexing techniques for efficiently mining spatiotemporal co-occurrence patternsabstractIn this paper, we investigate using specifically-designated spatiotemporal indexing techniques for mining cooccurrence patterns from spatiotemporal datasets with evolving polygon-based representations. Previously, suggested techniques for spatiotemporal pattern mining algorithms did not take spatiotemporal indexing techniques into account. We present a new framework for mining spatiotemporal co-occurrence patterns that can use various indexing techniques for efficiently accessing data. Two well-studied spatiotemporal indexing structures, Scalable and Efficient Trajectory Index (SETI) and Chebyshev Polynomial Indexing are currently implemented and available in our framework. Berkay Aydin, Dustin Kempton, Vijay Akkineni, Shaktidhar Reddy Gopavaram, Karthik Ganesan Pillai, Rafal A. Angryk |
IEEE BigData | 6 |
| 2014 | Scalable solar image Retrieval with LuceneabstractIn this work we present an alternative approach for large-scale retrieval of solar images using the highly-scalable retrieval engine Lucene. While Lucene is widely popular among text- based search engines, significant adjustments need to be made to take advantage of its fast indexing mechanism and highly-scalable architecture to enable search on image repositories. In this work we describe a novel way of representing image feature vectors in order to enable Lucene to perform search and retrieval of similar images. We compare our proposed method with other popular alternatives and provide commentary of the performance as well as the benefits and caveats of the proposed method. Juan M. Banda, Rafal A. Angryk |
IEEE BigData | 2 |
| 2014 | Iterative refinement of multiple targets tracking of solar eventsabstractIn this paper, we combine two approaches to multiple-target tracking: the first is a hierarchical approach to iteratively growing track fragments across gaps in detections, and the second is a network flow based optimization method for data association. We introduce a new parallel algorithm for initial track fragment formation as the base of the hierarchical approach. The network flow based optimization method is then utilized for the remaining levels of the hierarchy. This process is applied to solar data retrieved from the Heliophysics Event Knowledgebase (HEK). We compare our results to labeled data from the same, and show improvements over a non-hierarchical sequential approach. Dustin Kempton, Karthik Ganesan Pillai, Rafal A. Angryk |
IEEE BigData | 3 |
| 2014 | Massive labeled solar image data benchmarks for automated feature recognitionabstractThis paper introduces standard benchmarks for automated feature recognition using solar image data from the Solar Dynamics Observatory (SDO) mission. We combine general purpose image parameters extracted in-line from this massive data stream of images with reported solar event metadata records from automated detection modules to create a variety of event-labeled image datasets. These new large-scale datasets can be used for computer vision and machine learning benchmarks as-is, or as the starting point for further data mining research and investigations, the results of which can also aide understanding and knowledge discovery in the solar science community. Here we present an overview of the dataset creation process, including data collection, analysis, and labeling, which currently spans over two years of data and continues to grow with the ongoing mission. We then highlight two case studies to evaluate several data labeling methodologies and provide real world examples of our dataset benchmarks. Preliminary results show promising capability for the recognition of solar flare events and the classification of active and quiet regions of the Sun. Michael A. Schuh, Rafal A. Angryk |
IEEE BigData | 2 |
| 2014 | Quality control from the perspective of a near-real-time, spatial-temporal data aggregator and (re)distributorabstractQuality control for near-real-time spatial-temporal data is often presented from the perspective of the original owner and provider of the data, and focuses on general techniques for outlier detection or uses domain-specific knowledge and rules to assess quality. The impact of quality control on the data aggregator and redistributor is neglected. The focus of this paper is to define and demonstrate quality control measures for real-time, spatial-temporal data from the perspective of the aggregator to provide tools for assessment and optimization of system operation and data redistribution. We define simple measures that account for temporal completeness and spatial coverage. The measures and methods developed are tested on real-world data and applications. Douglas E. Galarus, Rafal A. Angryk |
SIGSPATIAL/GIS | 2 |
| 2013 | Big Data New Frontiers: Mining, Search and Management of Massive Repositories of Solar Image Data and Solar Events
Juan M. Banda, Michael A. Schuh, Rafal A. Angryk, Karthik Ganesan Pillai, Patrick McInerney |
ADBIS (2) | 3 |
| 2013 | When Too Similar Is Bad: A Practical Example of the Solar Dynamics Observatory Content-Based Image-Retrieval System
Juan M. Banda, Michael A. Schuh, Tim Wylie, Patrick McInerney, Rafal A. Angryk |
ADBIS (2) | 5 |
| 2013 | Spatiotemporal Co-occurrence Rules
Karthik Ganesan Pillai, Rafal A. Angryk, Juan M. Banda, Tim Wylie, Michael A. Schuh |
ADBIS (2) | 2 |
| 2013 | Improving the Performance of High-Dimensional kNN Retrieval through Localized Dataspace Segmentation and Hybrid Indexing
Michael A. Schuh, Tim Wylie, Rafal A. Angryk |
ADBIS | 3 |
| 2013 | A filter-and-refine approach to mine spatiotemporal co-occurrencesabstractSpatiotemporal co-occurrence patterns (STCOPs) represent the subsets of event types that occur together in both space and time. However, the discovery of STCOPs in data sets with extended spatial representations that evolve over time is computationally expensive because of the necessity to calculate interest measures to assess the co-occurrence strength, and the number of candidates for STCOPs growing exponentially with the number of spatiotemporal event types. In this paper, we introduce a novel and effective filter-and-refine algorithm to efficiently find prevalent STCOPs in massive spatiotemporal data repositories with polygon shapes that move and evolve over time. We provide theoretical analysis of our approach, and follow this investigation with a practical evaluation of our algorithm effectiveness on three real-life data sets and one artificial data set. Karthik Ganesan Pillai, Rafal A. Angryk, Berkay Aydin |
SIGSPATIAL/GIS | 2 |
| 2013 | Abstracting for Dimensionality Reduction in Text ClassificationabstractThere is a growing interest in efficient models of text mining and an emergent need for new data structures that address word relationships. Detailed knowledge about the taxonomic environment of keywords that are used in text documents can provide valuable insight into the nature of the subject matter contained therein. Such insight may be used to enhance the data structures used in the text data mining task as relationships become usefully apparent. A popular scalable technique used to infer these relationships, while reducing dimensionality, has been Latent Semantic Analysis. We present a new approach, which uses an ontology of lexical abstractions to create abstraction profiles of documents and uses these profiles to perform text organization based on a process that we call frequent abstraction analysis. We introduce TATOO, the Text Abstraction TOOlkit, which is a full implementation of this new approach. We present our data model via an example of how taxonomically derived abstractions can be used to supplement semantic data structures for the text classification task. Richard McAllister 0001, Rafal A. Angryk |
Int. J. Intell. Syst. | 2 |
| 2007 | Distributed Document Clustering Using Word-clustersabstractDocument clustering has become an increasingly important task in analyzing huge numbers of documents distributed among various sites. The challenging aspect is to analyze this enormous number of extremely high dimensional distributed documents and to organize them in such a way that results in better search and knowledge extraction without introducing much extra cost and complexity. This paper presents a distributed document clustering approach called distributed information bottleneck (DIB). DIB adopts a two stage agglomerative information bottleneck (aIB) algorithm to generate local clusters. At the first stage, the high-dimensional document vector is significantly reduced by finding word-clusters. These word-clusters are then used to obtain document-clusters in the second stage. DIB then extracts compact but informative local models from these document-clusters and transfers them to a central site. At the global site, the local models, that are likely to describe the same document set, are first combined. The resultant local models are then clustered by using the aIB algorithm to produce a hierarchical organization of all distributed documents. Our experimental results demonstrate the robustness, efficiency and effectiveness of DIB approach to cluster distributed documents. Debzani Deb, Rafal A. Angryk |
CIDM | 2 |
| 2007 | Attribute-oriented fuzzy generalization in proximity- and similarity-based relational database systemsabstractIn this article we investigate an attribute-oriented induction approach for acquisition of abstract knowledge from data stored in a fuzzy database environment. We utilize a proximity-based fuzzy database schema as the medium carrying the original information, where lack of precise information about an entity can be reflected via multiple attribute values, and the classical equivalence relation is replaced with the broader fuzzy proximity relation. We analyze in detail the process of attribute-oriented induction by concept hierarchies, utilizing the original properties of fuzzy databases to support this established data mining technique. In our approach we take full advantage of the implicit knowledge about the similarity of original attribute values, included by default in the investigated fuzzy database schemas. © 2007 Wiley Periodicals, Inc. Int J Int Syst 22: 763–779, 2007. Rafal A. Angryk, Fred Petry |
Int. J. Intell. Syst. | 1 |