VLDB 2026 Research / reviewers in the wild / expert
Petre Lameski
dblp:22/10118
· DBLP profile ↗
15ranked-venue papers
0as first author
4since 2021 · last 2024
0000-0002-5336-1796ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 3 since 2021Software engineering, systems software and programming languages · 9 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Benchmarking OpenAI's APIs and Large Language Models for Repeatable, Efficient Question Answering Across Multiple DocumentsabstractThe rapid growth of document volumes and complexity in various domains necessitates advanced automated methods to enhance the efficiency and accuracy of information extraction and analysis.This paper aims to evaluate the efficiency and repeatability of OpenAI's APIs and other Large Language Models (LLMs) in automating question-answering tasks across multiple documents, specifically focusing on analyzing Data Privacy Policy (DPP) documents of selected EdTech providers.We test how well these models perform on large-scale text processing tasks using the OpenAI's LLM models (GPT 3.5 Turbo, GPT 4, GPT 4o) and APIs in several frameworks: direct API calls (i.e., one-shot learning), LangChain, and Retrieval Augmented Generation (RAG) systems.We also evaluate a local deployment of quantized versions (with FAISS) of LLM models (Llama-2-13B-chat-GPTQ).Through systematic evaluation against predefined use cases and a range of metrics, including response format, execution time, and cost, our study aims to provide insights into the optimal practices for document analysis.Our findings demonstrate that using OpenAI's LLMs via API calls is a workable workaround for accelerating document analysis when using a local GPU-powered infrastructure is not a viable solution, particularly for long texts.On the other hand, the local deployment is quite valuable for maintaining the data within the private infrastructure.Our findings show that the quantized models retain substantial relevance even with fewer parameters than ChatGPT and do not impose processing restrictions on the number of tokens.This study offers insights on maximizing the use of LLMs for better efficiency and data governance in addition to confirming their usefulness in improving document analysis procedures. Elena Filipovska, Ana Mladenovska, Merxhan Bajrami, Jovana Dobreva, Vellislava Hillman, Petre Lameski, Eftim Zdravevski |
FedCSIS | 6 |
| 2024 | Approaches and Opportunities of Using Machine Learning Methods in Telecommunications and Industry 4.0
Ivan Cvitic, Aleksandar Jevremovic, Petre Lameski |
Mob. Networks Appl. | 3 |
| 2023 | Towards Industry 4.0: Machine malfunction prediction based on IIoT streaming dataabstractThe manufacturing industry relies on continuous optimization to meet quality and safety standards, which is part of the Industry 4.0 concept.Predicting when a specific part of a product will fail to meet these standards is of utmost importance and requires vast amounts of data, which often are collected from variety of sensors, often reffered to as Industrial Internet of Things (IIoT).Using a published dataset from Bosch, that describes the process at every step of production, we aim to train a machine learning model that can accurately predict faults in the manufacturing process.The dataset provides two years of production data across four production lines and 52 stations.Considering that the data generated from each production part includes more than four thousand features, we investigate various feature selection and data preprocessing methods.The obtained results exhibit Area Under the Receiver Operating Characteristic Curve (AUC ROC) of up to 0.997, which is remarkable and promising even for real-life production use. Dragana Nikolova, Petre Lameski, Ivan Miguel Pires, Eftim Zdravevski |
FedCSIS | 2 |
| 2021 | Are central-zone restaurants better for consumers? -An analytical approachabstractAnalysis of restaurant location has significant value both for owners, and for consumers. In this paper we examine the statistical relationships of restaurant properties and qualities in order to establish their relevance for the success of a restaurant. For this purpose, we use information-based methods and a dataset obtained from the TripAdvisor platform containing data about restaurant facilities situated in London. The dataset is rich with numerous qualitative features describing the restaurants. Through the use of statistical and geographical methods we explore the usefulness of geographic visualisation of restaurant qualities, and examine the statistical relationships among restaurant descriptors. Our results demonstrate the effectiveness of information methods for researching economic and non-economic phenomena. Filip Markoski, Lasko Basnarkov, Biljana Risteska Stojkoska, Petre Lameski, Eftim Zdravevski |
IEEE BigData | 4 |
| 2020 | Short-term air pollution forecasting based on environmental factors and deep learning modelsabstractThe effects of air pollution on people, the environment, and the global economy are profound -and often under-recognized.Air pollution is becoming a global problem.Urban areas have dense populations and a high concentration of emission sources: vehicles, buildings, industrial activity, waste, and wastewater.Tackling air pollution is an immediate problem in developing countries, such as North Macedonia, especially in larger urban areas.This paper exploits Recurrent Neural Network (RNN) models with Long Short-Term Memory units to predict the level of PM10 particles in the near future (+3 hours), measured with sensors deployed in different locations in the city of Skopje.Historical air quality measurements data were used to train the models.In order to capture the relation of air pollution and seasonal changes in meteorological conditions, we introduced temperature and humidity data to improve the performance.The accuracy of the models is compared to PM10 concentration forecast using an Autoregressive Integrated Moving Average (ARIMA) model.The obtained results show that specific deep learning models consistently outperform the ARIMA model, particularly when combining meteorological and air pollution historical data.The benefit of the proposed models for reliable predictions of only 0.01 MSE could facilitate preemptive actions to reduce air pollution, such as temporarily shutting main polluters, or issuing warnings so the citizens can go to a safer environment and minimize exposure. Mirche Arsov, Eftim Zdravevski, Petre Lameski, Roberto Corizzo, Nikola Koteli, Kosta Mitreski, Vladimir Trajkovik |
FedCSIS | 3 |
| 2020 | Explorations into Deep Learning Text Architectures for Dense Image CaptioningabstractImage captioning is the process of generating a textual description that best fits the image scene.It is one of the most important tasks in computer vision and natural language processing and has the potential to improve many applications in robotics, assistive technologies, storytelling, medical imaging and more.This paper aims to analyse different encoder-decoder architectures for dense image caption generation while focusing on the text generation component.Already trained models for image feature generation are utilized with transfer learning.These features are used for describing the regions using three different models for text generation.We propose three deep learning architectures for generating one-sentence captions of Regions of Interest (RoIs).The proposed architectures reflect several ways of integrating features from images and text.The proposed models were evaluated and compared with several metrics for natural language generation.The experimental results demonstrate that injecting image features into a decoder RNN while generating a caption word by word is the best performing architecture among the architectures explored in this paper. Martina Toshevska, Frosina Stojanovska, Eftim Zdravevski, Petre Lameski, Sonja Gievska |
FedCSIS | 4 |
| 2019 | Cluster-size optimization within a cloud-based ETL framework for Big DataabstractThe ability to analyze the available data is a valuable asset for any successful business, especially when the analysis yields meaningful knowledge. One of the key processes for acquiring such ability is the Extract-Transform-Load (ETL) process. For Big Data, ETL requires a significant effort and it is a very challenging task to be performed in a cost-effective way. There are quite a few examples in the literature that describe an architecture for cost-effective ETL but none of the available examples are complete enough and they are usually evaluated in narrow problem domains. The ones that are more general, require specific implementation details. In this paper we propose a cloud-based ETL framework where we use a general cluster-size optimization algorithm, while providing implementation details, and is able to perform the required job within a predefined, and thus known, time. We evaluated the algorithm by executing three scenarios regarding data aggregation during ETL: (i) ETL with no aggregation; (ii) aggregation based on predefined columns or time intervals; and (iii) aggregation within single user sessions spanning over arbitrary time intervals. The execution of the three ETL scenarios in a production setting showed that the cluster size could be optimized so it can process the required data volume within a predefined and thus, expected, latency. The scalability was evaluated on Amazon AWS Hadoop clusters by processing user logs collected with Kinesis streams with datasets ranging from 30 GB to 2.6 TB. Eftim Zdravevski, Petre Lameski, Ace Dimitrievski, Marek Grzegorowski, Cas Apanowicz |
IEEE BigData | 2 |
| 2017 | Firearms training simulator based on low cost motion tracking sensor
Dimitar Bogatinov, Petre Lameski, Vladimir Trajkovik, Katerina Mitkovska Trendova |
Multim. Tools Appl. | 2 |
| 2016 | Automatic Feature Engineering for Prediction of Dangerous Seismic Activities in Coal MinesabstractIn this paper we present our submission to the AAIA'16 Data Mining Challenge, where the objective was to predict dangerous seismic events based on hourly aggregated readings from different sensor and recent mining expert assessment of the conditions in the mine.During the course of the competition we have exploited a framework for automatic feature extraction from time series data that did not require any manual tuning.Furthermore, we have analyzed the impact of overlapping of input data on model robustness.We argue that training an ensemble of classifiers with distinct (i.e.nonoverlapping) chronological data rather than one classifier with all available data can produce more reliable and robust prediction models.By doing that, we were able to avoid overfitting and obtain the same score performance on the evaluation and test datasets, despite the significant data drift in the datasets. Eftim Zdravevski, Petre Lameski, Andrea Kulakov |
FedCSIS | 2 |
| 2016 | Row Key Designs of NoSQL Database Tables and Their Impact on Write PerformanceabstractIn several NoSQL database systems, among which is HBase, only one index is available for the tables, which is also the row key and the clustered index. Using other indexes does not come out of the box. As a result, the row key design is the most important thing when designing tables, because an inappropriate design can lead to detrimental consequences on performances and costs. Particular row key designs are suitable for different problems, and in this paper we analyze the performance, characteristics and applicability of each of them. In particular we investigate the effect of using various techniques for modeling row keys: sequences, salting, padding, hashing, and modulo operations. We propose four different designs based on these techniques and we analyze their performance on different HBase clusters when loading HDFS files with various sizes. The experiments show that particular designs consistently outperform others on differently sized clusters in both execution time and even load distribution across nodes. Eftim Zdravevski, Petre Lameski, Andrea Kulakov |
PDP | 2 |
| 2015 | Parallel computation of information gain using Hadoop and MapReduceabstractNowadays, companies collect data at an increasingly high rate to the extent that traditional implementation of algorithms cannot cope with it in reasonable time.On the other hand, analysis of the available data is a key to the business success.In a Big Data setting tasks like feature selection, finding discretization thresholds of continuous data, building decision threes, etc are especially difficult.In this paper we discuss how a parallel implementation of the algorithm for computing the information gain can address these issues.Our approach is based on writing Pig Latin scripts that are compiled into MapReduce jobs which then can be executed on Hadoop clusters.In order to implement the algorithm first we define a framework for developing arbitrary algorithms and then we apply it for the task at hand.With intent to analyze the impact of the parallelization, we have processed the FedCSIS AAIA'14 dataset with the proposed implementation of the information gain.During the experiments we evaluate the speedup of the parallelization compared to a one-node cluster.We also analyze how to optimally determine the number of map and reduce tasks for a given cluster.To demonstrate the portability of the implementation we present results using an on-premises and Amazon AWS clusters.Finally, we illustrate the scalability of the implementation by evaluating it on a replicated version of the same dataset which is 80 times larger than the original. Eftim Zdravevski, Petre Lameski, Andrea Kulakov, Sonja Filiposka, Dimitar Trajanov, Boro Jakimovski |
FedCSIS | 2 |
| 2015 | Transformation of nominal features into numeric in supervised multi-class problems based on the weight of evidence parameterabstractMachine learning has received increased interest by both the scientific community and the industry. Most of the machine learning algorithms rely on certain distance metrics that can only be applied to numeric data. This becomes a problem in complex datasets that contain heterogeneous data consisted of numeric and nominal (i.e. categorical) features. Thus the need of transformation from nominal to numeric data. Weight of evidence (WoE) is one of the parameters that can be used for transformation of the nominal features to numeric. In this paper we describe a method that uses WoE to transform the features. Although the applicability of this method is researched to some extent, in this paper we extend its applicability for multi-class problems, which is a novelty. We compared it with the method that generates dummy features. We test both methods on binary and multi-class classification problems with different machine learning algorithms. Our experiments show that the WoE based transformation generates smaller number of features compared to the technique based on generation of dummy features while also improving the classification accuracy, reducing memory complexity and shortening the execution time. Be that as it may, we also point out some of its weaknesses and make some recommendations when to use the method based on dummy features generation instead. Eftim Zdravevski, Petre Lameski, Andrea Kulakov, Slobodan Kalajdziski |
FedCSIS | 2 |
| 2015 | Robust histogram-based feature engineering of time series dataabstractCollecting data at regular time nowadays is ubiquitous.The most widely used type of data that is being collected and analyzed is financial data and sensor readings.Various businesses have realized that financial time series analysis is a powerful analytical tool that can lead to competitive advantages.Likewise, sensor networks generate time series and if they are properly analyzed can give a better understanding of the processes that are being monitored.In this paper we propose a novel generic histogram-based method for feature engineering of time series data.The preprocessing phase consists of several steps: deseansonalyzing the time series data, modeling the speed of change with first derivatives, and finally calculating histograms.By doing all of those steps the goal is three-fold: achieve invariance to different factors, good modeling of the data and preform significant feature reduction.This method was applied to the AAIA Data Mining Competition 2015, which was concerned with recognition of activities carried out by firefighters by analyzing body sensor network readings.By doing that we were able to score the third place with predictive accuracy of about 83%, which was about 1% worse than the winning solution. Eftim Zdravevski, Petre Lameski, Riste Mingov, Andrea Kulakov, Dejan Gjorgjevikj |
FedCSIS | 2 |
| 2014 | Feature selection and allocation to diverse subsets for multi-label learning problems with large datasetsabstractFeature selection is important phase in machine learning and in the case of multi-label classification, it can be considerably challenging.In like manner, finding the best subset of good features is involved and difficult when the dataset has significantly large number of features (more than a thousand).In this paper we address the problem of feature selection for multilabel classification with large number of features.The proposed method is a hybrid of two phases -preliminary feature selection based on the information value and additional correlation-based selection.We show how with the first phase we can do preliminary selection of features from tens of thousands to couple of hundred, and then with the second phase we can make fine-grained feature selection with more sophisticated but computationally intensive methods.Finally, we analyze the ways of allocating the selected features to diverse subsets, which are suitable for training of ensembles of classifiers. Eftim Zdravevski, Petre Lameski, Andrea Kulakov, Dejan Gjorgjevikj |
FedCSIS | 2 |
| 2011 | Weight of evidence as a tool for attribute transformation in the preprocessing stage of supervised learning algorithmsabstractTransformation of features is a common task in the data preprocessing stage while solving data mining and classification problems. Many classification algorithms have preference of continual attributes over nominal attributes, and sometimes the distance between different data points cannot be estimated if the values of the attributes are not continual and normalized. The Weight of Evidence has some very desirable properties that make it very useful tool for the transformation of attributes, but unfortunately there are some preconditions that need to be met in order to calculate it. In this paper we propose a modified calculation of the Weight of Evidence that overcomes these preconditions, and additionally makes it usable for test examples that were not present in the training set. The proposed transformation can be used for all supervised learning problems. At the end, we present the results from the proposed transformation and discuss the benefits of the transformed nominal and continual attributes from the PAKDD 2009 dataset. The results show that the proposed transformation contributes towards a better performance in all tested classification algorithms than the method that generates dummy (i.e. binary) variables for each value of the nominal attributes. Eftim Zdravevski, Petre Lameski, Andrea Kulakov |
IJCNN | 2 |