Stan Matwin

dblp:m/StanMatwin · DBLP profile ↗
← Back
48ranked-venue papers in the field
3as first author
9since 2021 · last 2023
0000-0001-6629-8434ORCID · conflict

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 16Big Data, Cloud & Distributed Data Systems · 10Other / Interdisciplinary · 9 (2 first)Database Systems & Data Management · 6Knowledge Engineering, Semantic Web & Information Systems · 6 (1 first)Information Retrieval & Web Search · 1
YearPublicationVenuePosition
2023 Charting the Course of Ship Track Prediction: A Novel Approach for Maritime Traffic Analysis and Enhanced Situational Awareness
abstract
Accurate ship track prediction plays a pivotal role in maritime operations, enabling proactive decision-making, enhancing safety, and optimizing vessel routing. We propose ship trajectory prediction - a threefold technique for facilitating accurate predictions, which involves clustering historical AIS trajectories into maritime de facto routes, classifying new trajectory to one of these routes, and conducting predictions along the identified route. To overcome the challenges of capturing the latent structure in high-dimensional and heterogeneous space imposed by AIS data, we introduce a new similarity technique that automatically determines the number of clusters. Furthermore, we introduce a method to automatically annotate the feature space, enhancing the efficiency of data analysis tasks like clustering. This not only improves performance but also ensures transparency, allowing for effective performance evaluation. Our approach demonstrates an accuracy of over 88% and an accuracy of 78% in predicting routes for tanker vessels.
Lubna Eljabu, Mohammad Etemad, Stan Matwin
IEEE Big Data3
2023 Assessing compression algorithms to improve the efficiency of clustering analysis on AIS vessel trajectories
abstract
In the maritime environment, the Automatic Identification System (AIS) is used to monitor vessel activity concerning security and safety ocean-wide. AIS data has been used to detect anomalous behaviors related to suspicious activities and hazardous events. Typically, clustering analysis is used to investigate anomalous events within the AIS data stream. However, the main challenge in this approach is to determine and execute the dissimilarity measure between trajectories since they differ in size and time. In addition, these calculations are computationally expensive and not scalable. To tackle this issue, compression algorithms can be applied to perform clustering analysis since they are typically used to reduce storage and processing time. Therefore, the proposed analysis will assess how compression algorithms affect clustering results with respect to detecting anomalous vessel trajectories. The analysis results show that a suitable compression algorithm can reduce the overall processing time with little impact on the clustering results while supporting the scalability of this type of analysis.
Martha Dais Ferreira, Jessica N. A. Campbell, Evan Purney, Amílcar Soares Júnior 0001, Stan Matwin
Int. J. Geogr. Inf. Sci.5
2022 Spatial Clustering Method of Historical AIS Data for Maritime Traffic Routes Extraction
abstract
The automated extraction of maritime routes that accurately resembles the real traffic of vessels is crucial for intelligent traffic management systems to understand vessel behaviour in sea areas, identify events, and support decision-making. Available solutions for maritime traffic route extraction utilize traditional clustering algorithms, which have high computational costs. Data reduction methods are proposed for use with these clustering methods to improve clustering performance, which involves a loss in movement pattern quality. Such solutions often result in low quality representations of traffic routes, which poorly estimate sailing distances for long journeys and time of arrivals and poorly identify non-conformities. In this paper, we propose a spatial clustering method (SPTCLUST-II) to extract spatial representations of sailing routes from historical Automatic Identification System (AIS) data. Our method can cluster huge volumes of trajectory data in a minimal amount of time without using any of the traditional clustering algorithms and with no reduction or modification of the spatio-temporal predicates of the original trajectories. A real-world AIS dataset captured in the area of the Gulf of Mexico is used for the evaluation of the proposed method. The results demonstrate that the proposed method extracts tankers maritime traffic routes with an accuracy of 97% and a f1-measure of 98.5% and cargoes maritime traffic routes with an accuracy of 98.1% and a f1-measure of 99%. This method can be utilized by surveillance authorities for stable and sustainable vessel traffic management.
Lubna Eljabu, Mohammad Etemad, Stan Matwin
IEEE Big Data3
2022 From multiple aspect trajectories to predictive analysis: a case study on fishing vessels in the Northern Adriatic sea
abstract
Abstract In this paper we model spatio-temporal data describing the fishing activities in the Northern Adriatic Sea over four years. We build, implement and analyze a database based on the fusion of two complementary data sources: trajectories from fishing vessels (obtained from terrestrial Automatic Identification System, or AIS, data feed) and fish catch reports (i.e., the quantity and type of fish caught) of the main fishing market of the area. We present all the phases of the database creation, starting from the raw data and proceeding through data exploration, data cleaning, trajectory reconstruction and semantic enrichment. We implement the database by using MobilityDB, an open source geospatial trajectory data management and analysis platform. Subsequently, we perform various analyses on the resulting spatio-temporal database, with the goal of mapping the fishing activities on some key species, highlighting all the interesting information and inferring new knowledge that will be useful for fishery management. Furthermore, we investigate the use of machine learning methods for predicting the Catch Per Unit Effort (CPUE), an indicator of the fishing resources exploitation in order to drive specific policy design. A variety of prediction methods, taking as input the data in the database and environmental factors such as sea temperature, waves height and Clorophill-a, are put at work in order to assess their prediction ability in this field. To the best of our knowledge, our work represents the first attempt to integrate fishing ships trajectories derived from AIS data, environmental data and catch data for spatio-temporal prediction of CPUE – a challenging task.
Bruno Brandoli Machado, Alessandra Raffaetà, Marta Simeoni, Pedram Adibi, Fateha Khanam Bappee, Fabio Pranovi, Giulia Rovinelli, Elisabetta Russo, Claudio Silvestri, Amílcar Soares Júnior 0001, Stan Matwin
GeoInformatica11
2022 Understanding evolution of maritime networks from automatic identification system data
Emanuele Carlini 0001, Vinicius Monteiro de Lira, Amílcar Soares Júnior 0001, Mohammad Etemad, Bruno Brandoli Machado, Stan Matwin
GeoInformatica6
2021 MTLV: a library for building deep multi-task learning architectures
abstract
Multi-Task Learning (MTL) for text classification takes advantage of the data to train a single shared model with multiple task-specific layers on multiple related classification tasks to improve its generalization performance. We choose pre-trained language models (BERT-family) as the shared part of this architecture. Although they have achieved noticeable performance in different downstream NLP tasks, their performance in an MTL setting for the biomedical domain is not thoroughly investigated. In this work, we investigate the performance of BERT-family models in different MTL settings with Open-I (radiology reports) and OHSUMED (PubMed abstracts) datasets. We introduce the MTLV (Multi-Task Learning Visualizer) library for building Multi-task learning-related architectures which use existing infrastructure (e.g., Hugging Face Transformers and MLflow Tracking). Following previous work in computer vision, we clustered tasks and trained a separate model on each cluster (Grouped Multi-Task Learning (GMTL)). Contextual representation of the class labels (Tasks) and their descriptions was used by the library as features to cluster the tasks. We observed that grouping tasks for training with few models (GMTL) outperforms the MTL also GMTL is computationally more efficient than the STL setting (a separate model is trained for each task).
Fatemeh Rahimi, Evangelos E. Milios, Stan Matwin
DocEng3
2021 SWS: an unsupervised trajectory segmentation algorithm based on change detection with interpolation kernels
Mohammad Etemad, Amílcar Soares Júnior 0001, Elham Etemad, Jordan Rose, Luís Torgo, Stan Matwin
GeoInformatica6
2021 Building navigation networks from multi-vessel trajectory data
Iraklis Varlamis, Ioannis Kontopoulos, Konstantinos Tserpes, Mohammad Etemad, Amílcar Soares Júnior 0001, Stan Matwin
GeoInformatica6
2021 Multiple-aspect analysis of semantic trajectories(MASTER)
abstract
A plethora of applications and devices reporting their locations generate massive amounts of spatiotemporal data along with other useful information. These data can form trajectories with sequences...
Chiara Renso, Vania Bogorny, Konstantinos Tserpes, Stan Matwin, José A. F. de Macêdo
Int. J. Geogr. Inf. Sci.4
2019 VISTA: A visual analytics platform for semantic annotation of trajectories
abstract
Most of the trajectory datasets only record the spatio-temporal position of the moving object, thus lacking semantics and this is due to the fact that this information mainly depends on the domain expert labeling, a time-consuming and complex process. This paper is a contribution in facilitating and supporting the manual annotation of trajectory data thanks to a visual-analytics-based platform named VISTA. VISTA is designed to assist the user in the trajectory annotation process in a multi-role user environment. A session manager creates a tagging session selecting the trajectory data and the semantic contextual information. The VISTA platform also supports the creation of several features that will assist the tagging users in identifying the trajectory segments that will be annotated. A distinctive feature of VISTA is the visual analytics functionalities that support the users in exploring and processing the trajectory data, the associated features and the semantic information for a proper comprehension of how to properly label trajectories.
Amílcar Soares Júnior 0001, Jordan Rose, Mohammad Etemad, Chiara Renso, Stan Matwin
EDBT5
2019 Automatic Fusion of Satellite Imagery and AIS data for Vessel Detection
Aristides Milios, Konstantina Bereta, Konstantinos Chatzikokolakis 0002, Dimitrios Zissis, Stan Matwin
FUSION5
2019 Black Box Explanation by Learning Image Exemplars in the Latent Feature Space
Riccardo Guidotti, Anna Monreale, Stan Matwin, Dino Pedreschi
ECML/PKDD (1)3
2019 Marine Mammal Species Classification Using Convolutional Neural Networks and a Novel Acoustic Representation
Bruce Martin, Katie Kowarski, Briand J. Gaudet, Stan Matwin
ECML/PKDD (3)5
2019 Computational modelling and data-driven techniques for systems analysis
Stan Matwin, Luca Tesei, Roberto Trasarti
J. Intell. Inf. Syst.1
2018 A Semi-Supervised Approach for the Semantic Segmentation of Trajectories
abstract
A first fundamental step in the process of analyzing movement data is trajectory segmentation, i.e., splitting trajectories into homogeneous segments based on some criteria. Although trajectory segmentation has been the object of several approaches in the last decade, a proposal based on a semi-supervised approach remains inexistent. A semi-supervised approach means that a user labels manually a small set of trajectories with meaningful segments and, from this set, the method infers in an unsupervised way the segments of the remaining trajectories. The main advantage of this method compared to pure supervised ones is that it reduces the human effort to label the number of trajectories. In this work, we propose the use of the Minimum Description Length (MDL) principle to measure homogeneity inside segments. We also introduce the Reactive Greedy Randomized Adaptive Search Procedure for semantic Semi-supervised Trajectory Segmentation (RGRASP-SemTS) algorithm that segments trajectories by combining a limited user labeling phase with a low number of input parameters and no predefined segmenting criteria. The approach and the algorithm are presented in detail throughout the paper, and the experiments are carried out on two real-world datasets. The evaluation tests prove how our approach outperforms state-of-the-art competitors when compared to ground truth.
Amílcar Soares Júnior 0001, Valéria Cesário Times, Chiara Renso, Stan Matwin, Lucídio A. F. Cabral
MDM4
2018 Incremental anomaly detection using two-layer cluster-based structure
Elnaz Bigdeli, Mahdi Mohammadi, Bijan Raahemi, Stan Matwin
Inf. Sci.4
2016 Predicting annual average daily highway traffic from large data and very few measurements
abstract
This paper is an early report from research undertaken to meet the needs of the General Directorate for National Roads and Motorways of Poland. They have defined the task of estimating the annual average Daily Traffic on a class of highways in the country, based on a very small number of daily traffic measurements undertaken throughout the year (typically one or two such measurements). We report the data available to us, and the data preprocessing step, including the generation of additional attributes and generation of synthetic data. We use a deep neural network model of the annual count number, and we report encouraging early result. In the conclusion, we discuss the next steps of this research.
Tomasz Tajmajer, Malwina Splawinska, Piotr Wasilewski, Stan Matwin
IEEE BigData4
2016 Fast Unsupervised Online Drift Detection Using Incremental Kolmogorov-Smirnov Test
abstract
Data stream research has grown rapidly over the last decade. Two major features distinguish data stream from batch learning: stream data are generated on the fly, possibly in a fast and variable rate; and the underlying data distribution can be non-stationary, leading to a phenomenon known as concept drift. Therefore, most of the research on data stream classification focuses on proposing efficient models that can adapt to concept drifts and maintain a stable performance over time. However, specifically for the classification task, the majority of such methods rely on the instantaneous availability of true labels for all already classified instances. This is a strong assumption that is rarely fulfilled in practical applications. Hence there is a clear need for efficient methods that can detect concept drifts in an unsupervised way. One possibility is the well-known Kolmogorov-Smirnov test, a statistical hypothesis test that checks whether two samples differ. This work has two main contributions. The first one is the Incremental Kolmogorov-Smirnov algorithm that allows performing the Kolmogorov-Smirnov hypothesis test instantly using two samples that change over time, where the change is an insertion and/or removal of an observation. Our algorithm employs a randomized tree and is able to perform the insertion and removal operations in O(log N) with high probability and calculate the Kolmogorov-Smirnov test in O(1), where N is the number of sample observations. This is a significant speed-up compared to the O(N log N) cost of the non-incremental implementation. The second contribution is the use of the Incremental Kolmogorov-Smirnov test to detect concept drifts without true labels. Classification algorithms adapted to use the test rely on a limited portion of those labels just to update the classification model after a concept drift is detected.
Denis Moreira dos Reis, Peter A. Flach, Stan Matwin, Gustavo Batista
KDD3
2016 Ensembles of label noise filters: a ranking approach
Luís Paulo F. Garcia, Ana Carolina Lorena, Stan Matwin, André C. P. L. F. de Carvalho
Data Min. Knowl. Discov.3
2015 Ship movement anomaly detection using specialized distance measures
Erico N. de Souza, Casey Hilliard, Stan Matwin
FUSION4
2015 Anomaly detection in maritime data based on geometrical analysis of trajectories
Behrouz Haji Soleimani, Erico N. de Souza, Casey Hilliard, Stan Matwin
FUSION4
2015 GRASP-UTS: an algorithm for unsupervised trajectory segmentation
abstract
An important problem in the knowledge discovery of trajectories is segmentation in subparts (subtrajectories). Existing algorithms for trajectory segmentation generally use explicit criteria to create segments. In this article, we propose segmenting trajectories using a novel, unsupervised approach, in which no explicit criteria are predetermined. To achieve this, we apply the Minimum Description Length (MDL) principle, which can measure homogeneity in the trajectory data by computing the similarities between landmarks (i.e. representative points of the trajectory) and the points in their neighborhood. Based on the homogeneity measurements, we propose an algorithm named Greedy Randomized Adaptive Search Procedure for Unsupervised Trajectory Segmentation (GRASP-UTS), which is a meta-heuristic that builds segments by modifying the number and positions of landmarks. We perform experiments with GRASP-UTS in two real-world datasets, using segment purity and coverage metrics to evaluate its efficiency. Experimental results demonstrate that GRASP-UTS correctly segmented sample trajectories without predetermined criteria, by computing similarities between landmarks and other trajectory points.
Amílcar Soares Júnior 0001, Bruno Moreno, Valéria Cesário Times, Stan Matwin, Lucídio A. F. Cabral
Int. J. Geogr. Inf. Sci.4
2014 Privacy-aware filter-based feature selection
abstract
A large amount of digital information collected and stored in databases creates new opportunities for knowledge discovery and data mining. The datasets, however, may contain personally identifiable information that needs to be protected. With high dimensionality of many large datasets, dimensionality reduction such as feature selection becomes indispensible. In this work, we aim at incorporating privacy into the very process of feature selection and as such, propose a privacy-aware filter-based feature selection method (PF-IFR). Our method enables data custodians to define a trade-off measure for controlling the amount of privacy and efficacy using filter-based feature selection techniques.
Yasser Jafer, Stan Matwin, Marina Sokolova
IEEE BigData2
2014 Knowledge-based clustering of ship trajectories using density-based approach
abstract
Maritime traffic monitoring is an important aspect of safety and security, particularly in close to port operations. While there is a large amount of data with variable quality, decision makers need reliable information about possible situations or threats. To address this requirement, we propose extraction of normal ship trajectory patterns that builds clusters using, besides ship tracing data, the publicly available International Maritime Organization (IMO) rules. The main result of clustering is a set of generated lanes that can be mapped to those defined in the IMO directives. Since the model also takes non-spatial attributes (speed and direction) into account, the results allow decision makers to detect abnormal patterns - vessels that do not obey the normal lanes or sail with higher or lower speeds.
Erico N. de Souza, Stan Matwin, Marcin Sydow
IEEE BigData3
2014 Vessel route anomaly detection with Hadoop MapReduce
abstract
We present a two-level approach to detect abnormal activities for vessels' routes. The data is obtained from the Automatic Identification System (AIS) which is required to be installed on vessels over specific gross tonnage. In the first level, we develope a Clustering algorithm: Density-based Spatial Clustering of Applications with Noise considering Speed and Direction (DBSCAN_SD). This algorithm is applied to pre-cluster the data points. Using domain knowledge in maritime, experts adjust the results produced by DBSCAN_SD with extra features. In this way, we get the optimal labeling result about whether a data point is normal or abnormal. In the second level, we use the labeled data generated in the first level to train the Parallel Meta-Learning (PML) algorithm on Hadoop. The results show that both accuracy and time complexity results are improved when we increase the number of nodes in a cluster.
Xiaoguang Wang 0001, Xuan Liu 0007, Erico N. de Souza, Stan Matwin
IEEE BigData5
2014 A distributed instance-weighted SVM algorithm on large-scale imbalanced datasets
abstract
When huge amounts of data are processed to extract knowledge, the situation becomes a challenge because the data mining techniques are not adapted to the space and time requirements. This challenge is more significant when the data is class imbalanced. Like many other machine learning algorithms, the success of the support vector machine (SVM) is limited when it is applied to the problem of learning from imbalanced datasets, especially on big datasets. In this paper, we are trying to apply an instance-weighted variant of the SVM, with a parallel Meta-learning algorithm using MapReduce, to deal with the big data class imbalance problem. We develop a symmetric weight boosting method to optimize the instance-weighted SVM. Experimental results on benchmark datasets and real application big datasets show that the proposed algorithm not only is effective on big data class imbalanced problem, but also reduces the training computational complexity significantly when the number of computing nodes increases.
Xiaoguang Wang 0001, Xuan Liu 0007, Stan Matwin
IEEE BigData3
2014 Applying instance-weighted support vector machines to class imbalanced datasets
abstract
Learning with class imbalance is always a challenging task in many real world applications such as the Internet, surveillance, security, and finance. Like many other successful machine learning algorithms, the success of the support vector machine (SVM) is limited when it is applied to the problem of learning from imbalanced datasets. SVM with different error costs has been widely used to deal with the class imbalanced problem. In this paper, we are trying to apply an instance-weighted variant of the SVM with both 1-norm and 2-norm format to deal with the class imbalance problem. We develop an asymmetric boosting method on the weights of the tradeoff parameters to optimize the instance-weighted SVM. The experimental results on the benchmark datasets show that the proposed algorithm is effective on the class imbalanced problem.
Xiaoguang Wang 0001, Xuan Liu 0007, Stan Matwin, Nathalie Japkowicz
IEEE BigData3
2014 A multi-view two-level classification method for generalized multi-instance problems
abstract
Multi-instance (MI) learning is different than standard propositional classification, as it uses a set of bags containing many instances as input. While the instances in each bag are not labeled, the bags themselves are, as positive or negative. In this paper, we present a novel multi-view, two-level classification framework to address the generalized multi-instance problems. We first apply supervised and unsupervised learning methods to transform a MI dataset into a multi-view, single meta-instance dataset. Then we develop a multi-view learning approach that can integrate the information acquired by individual view learners on the meta-instance dataset from the previous step, and construct a final model. Our empirical studies show that the proposed method performs well compared to other popular MI learning methods.
Xiaoguang Wang 0001, Xuan Liu 0007, Stan Matwin, Nathalie Japkowicz
IEEE BigData3
2014 Processing OLAP Queries over an Encrypted Data Warehouse Stored in the Cloud
Claudivan Cruz Lopes, Valéria Cesário Times, Stan Matwin, Ricardo Rodrigues Ciferri, Cristina Dutra de Aguiar Ciferri
DaWaK3
2014 Dream sentiment analysis using second order soft co-occurrences (SOSCO) and time course representations
Amir Hossein Razavi, Stan Matwin, Joseph De Koninck, Ray Reza Amini
J. Intell. Inf. Syst.2
2013 Meta-learning for large scale machine learning with MapReduce
abstract
We have entered the big data age. Knowledge extraction from massive data is becoming more and more rewarding and urgent. MapReduce has provided a feasible framework for programming machine learning algorithms in Map and Reduce functions. The relatively simple programming interface has helped to solve machine learning algorithms' scalability problems. However, this framework suffers from an obvious weakness: it does not support iterations. This makes those algorithms requiring iterations difficult to fully explore the efficiency of MapReduce. In this paper, we propose to apply Meta-learning programmed with MapReduce to avoid parallelizing machine learning algorithms while also improving their scalability to big datasets. The experiments conducted on Hadoop fully distributed mode on Amazon EC2 demonstrate that our algorithm PML reduces the training computational complexity significantly when the number of computing nodes increases while gaining smaller error rates than those on one single node. The comparison of PML with the contemporary parallelized AdaBoost algorithm: AdaBoost.PL shows that PML has lower error rates.
Xuan Liu 0007, Xiaoguang Wang 0001, Stan Matwin, Nathalie Japkowicz
IEEE BigData3
2013 Inner Ensembles: Using Ensemble Methods Inside the Learning Algorithm
Houman Abbasian, Chris Drummond, Nathalie Japkowicz, Stan Matwin
ECML/PKDD (3)4
2011 Smooth Receiver Operating Characteristics (smROC) Curves
William Klement, Peter A. Flach, Nathalie Japkowicz, Stan Matwin
ECML/PKDD (2)4
2008 Generation of Globally Relevant Continuous Features for Classification
Sylvain Létourneau, Stan Matwin, Fazel Famili
PAKDD2
2008 Proper Model Selection with Significance Test
Charles Ling 0001, Harry Zhang, Stan Matwin
ECML/PKDD (1)4
2006 Evaluating Misclassifications in Imbalanced Data
William Elazmeh, Nathalie Japkowicz, Stan Matwin
ECML3
2005 STochFS: A Framework for Combining Feature Selection Outcomes Through a Stochastic Process
Jerffeson Teixeira de Souza, Nathalie Japkowicz, Stan Matwin
PKDD3
2004 A formal approach to using data distributions for building causal polytree structures
M. Ouerd, B. John Oommen, Stan Matwin
Inf. Sci.3
2004 Filtering Multi-Instance Problems to Reduce Dimensionality in Relational Learning
Érick Alphonse, Stan Matwin
J. Intell. Inf. Syst.2
2002 Privacy-Oriented Data Mining by Proof Checking
Amy P. Felty, Stan Matwin
PKDD2
1998 A Normalization Method for Contextual Data: Experience from a Large-Scale Application
Sylvain Létourneau, Stan Matwin, Fazel Famili
ECML2
1997 Learning When Negative Examples Abound
Miroslav Kubat, Robert C. Holte, Stan Matwin
ECML3
1994 Inverting Implication with Small Training Sets
David W. Aha, Stephane Lapointe, Charles Ling 0001, Stan Matwin
ECML4
1993 Learning Domain Theories using Abstract Beckground Knowledge
Peter Clark, Stan Matwin
ECML2
1993 A Case-Based Approach to Software Reuse
Gilles Fouqué, Stan Matwin
J. Intell. Inf. Syst.2
1985 A Logic-Based Knowledge Source System for Natural Language Document
Douglas R. Skuce, Stan Matwin, Branka Tauzovich, Franz Oppacher, Stan Szpakowicz
Data Knowl. Eng.2
1977 On the Completeness of a Set of Transformations Optimizing Linear Programs
Stan Matwin
Inf. Process. Lett.1
1977 An Experimental Investigation of Geschke's Method of Global Program Optimization
Stan Matwin
Inf. Process. Lett.1