Marek Gagolewski

dblp:53/7259 · DBLP profile ↗
← Back
39ranked-venue papers
13as first author
10since 2021 · last 2024
0000-0003-0637-6028ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 29 · 9 first-author · 7 since 2021Databases, data management, data science and information retrieval · 18 · 6 first-author · 3 since 2021Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2024 Random generation of linearly constrained fuzzy measures and domain coverage performance evaluation
abstract
The random generation of fuzzy measures under complex linear constraints holds significance in various fields, including optimization solutions, machine learning, decision making, and property investigation. However, most existing random generation methods primarily focus on addressing the monotonicity and normalization conditions inherent in the construction of fuzzy measures, rather than the linear constraints that are crucial for representing special families of fuzzy measures and additional preference information. In this paper, we present two categories of methods to address the generation of linearly constrained fuzzy measures using linear programming models. These methods enable a comprehensive exploration and coverage of the entire feasible convex domain. The first category involves randomly selecting a subset and assigning measure values within the allowable range under given linear constraints. The second category utilizes convex combinations of constrained extreme fuzzy measures and vertex fuzzy measures. Then we employ some indices of fuzzy measures, objective functions, and distances to domain boundaries to evaluate the coverage performance of these methods across the entire feasible domain. We further provide enhancement techniques to improve the coverage ratios. Finally, we discuss and demonstrate potential applications of these generation methods in practical scenarios.
Jianzhang Wu 0001, Gleb Beliakov, Simon James, Marek Gagolewski
Inf. Sci.4
2023 A benchmark-type generalization of the Sugeno integral with applications in bibliometrics
Michal Boczek, Marek Gagolewski, Marek Kaluszka, Andrzej Okolewski
Fuzzy Sets Syst.2
2023 Hierarchical clustering with OWA-based linkages, the Lance-Williams formula, and dendrogram inversions
abstract
Agglomerative hierarchical clustering based on Ordered Weighted Averaging (OWA) operators not only generalises the single, complete, and average linkages, but also includes intercluster distances based on a few nearest or farthest neighbours, trimmed and winsorised means of pairwise point similarities, amongst many others. We explore the relationships between the famous Lance–Williams update formula and the extended OWA-based linkages with weights generated via infinite coefficient sequences. Furthermore, we provide some conditions for the weight generators to guarantee the resulting dendrograms to be free from unaesthetic inversions.
Marek Gagolewski, Anna Cena, Simon James, Gleb Beliakov
Fuzzy Sets Syst.1
2022 Hierarchical data fusion processes involving the Möbius representation of capacities
Gleb Beliakov, Marek Gagolewski, Simon James
Fuzzy Sets Syst.2
2022 Reduction of variables and constraints in fitting antibuoyant fuzzy measures to data using linear programming
Gleb Beliakov, Marek Gagolewski, Simon James
Fuzzy Sets Syst.2
2022 A critique of the bounded fuzzy possibilistic method
Marek Gagolewski
Fuzzy Sets Syst.1
2022 Time to vote: Temporal clustering of user activity on Stack Overflow
abstract
Abstract Question‐and‐answer (Q&A) sites improve access to information and ease transfer of knowledge. In recent years, they have grown in popularity and importance, enabling research on behavioral patterns of their users. We study the dynamics related to the casting of 7 M votes across a sample of 700 k posts on Stack Overflow, a large community of professional software developers. We employ log‐Gaussian mixture modeling and Markov chains to formulate a simple yet elegant description of the considered phenomena. We indicate that the interevent times can naturally be clustered into 3 typical time scales: those which occur within hours, weeks, and months and show how the events become rarer and rarer as time passes. It turns out that the posts' popularity in a short period after publication is a weak predictor of its overall success, contrary to what was observed, for example, in case of YouTube clips. Nonetheless, the sleeping beauties sometimes awake and can receive bursts of votes following each other relatively quickly.
Agnieszka Geras, Grzegorz Siudem, Marek Gagolewski
J. Assoc. Inf. Sci. Technol.3
2021 Are cluster validity measures (in) valid?
Marek Gagolewski, Maciej Bartoszuk, Anna Cena
Inf. Sci.1
2021 T-norms or t-conorms? How to aggregate similarity degrees for plagiarism detection
Maciej Bartoszuk, Marek Gagolewski
Knowl. Based Syst.2
2021 A Characterization of the Property of Orthomonotonicity
abstract
In this letter, a characterization of the property of orthomonotnoicity for aggregation functions for multidimensional data is provided.
Raúl Pérez-Fernández, Bernard De Baets, Marek Gagolewski, Juan Jesús Salamanca
IEEE Trans. Fuzzy Syst.3
2020 Constrained ordered weighted averaging aggregation with multiple comonotone constraints
Lucian C. Coroianu, Robert Fullér, Marek Gagolewski, Simon James
Fuzzy Sets Syst.3
2020 Robust fitting for the Sugeno integral with respect to general fuzzy measures
Gleb Beliakov, Marek Gagolewski, Simon James
Inf. Sci.2
2020 Genie+OWA: Robustifying hierarchical clustering with OWA-based linkages
Anna Cena, Marek Gagolewski
Inf. Sci.2
2020 Should we introduce a dislike button for academic articles?
abstract
There is a mutual resemblance between the behavior of users of the Stack Exchange and the dynamics of the citations accumulation process in the scientific community, which enabled us to tackle the outwardly intractable problem of assessing the impact of introducing “negative” citations. Although the most frequent reason to cite an article is to highlight the connection between the 2 publications, researchers sometimes mention an earlier work to cast a negative light. While computing citation‐based scores, for instance, the h‐index, information about the reason why an article was mentioned is neglected. Therefore, it can be questioned whether these indices describe scientific achievements accurately. In this article we shed insight into the problem of “negative” citations, analyzing data from Stack Exchange and, to draw more universal conclusions, we derive an approximation of citations scores. Here we show that the quantified influence of introducing negative citations is of lesser importance and that they could be used as an indicator of where the attention of the scientific community is allocated.
Agnieszka Geras, Grzegorz Siudem, Marek Gagolewski
J. Assoc. Inf. Sci. Technol.3
2020 An Inherent Difficulty in the Aggregation of Multidimensional Data
abstract
In the field of information fusion, the problem of data aggregation has been formalized as an order-preserving process that builds upon the property of monotonicity. However, fields such as computational statistics, data analysis, and geometry usually emphasize the role of equivariances to various geometrical transformations in aggregation processes. Admittedly, if we consider a unidimensional data fusion task, both requirements are often compatible with each other. Nevertheless, in this paper, we show that, in the multidimensional setting, the only idempotent functions that are monotone and orthogonal equivariant are the over-simplistic weighted centroids. Even more, this result still holds after replacing monotonicity and orthogonal equivariance by the weaker property of orthomonotonicity. This implies that the aforementioned approaches to the aggregation of multidimensional data are irreconcilable, and that, if a weighted centroid is to be avoided, we must choose between monotonicity and a desirable behavior with regard to orthogonal transformations.
Marek Gagolewski, Raúl Pérez-Fernández, Bernard De Baets
IEEE Trans. Fuzzy Syst.1
2019 Aggregation on ordinal scales with the Sugeno integral for biomedical applications
Gleb Beliakov, Marek Gagolewski, Simon James
Inf. Sci.2
2019 Piecewise linear approximation of fuzzy numbers: algorithms, arithmetic operations and stability of characteristics
abstract
The problem of the piecewise linear approximation of fuzzy numbers giving outputs nearest to the inputs with respect to the Euclidean metric is discussed. The results given in Coroianu et al. (Fuzzy Sets Syst 233:26–51, 2013) for the 1-knot fuzzy numbers are generalized for arbitrary n-knot ( $$n\ge 2$$ ) piecewise linear fuzzy numbers. Some results on the existence and properties of the approximation operator are proved. Then, the stability of some fuzzy number characteristics under approximation as the number of knots tends to infinity is considered. Finally, a simulation study concerning the computer implementations of arithmetic operations on fuzzy numbers is provided. Suggested concepts are illustrated by examples and algorithms ready for the practical use. This way, we throw a bridge between theory and applications as the latter ones are so desired in real-world problems.
Lucian C. Coroianu, Marek Gagolewski, Przemyslaw Grzegorzewski
Soft Comput.2
2019 Supervised Learning to Aggregate Data With the Sugeno Integral
abstract
The problem of learning symmetric capacities (or fuzzy measures) from data is investigated toward applications in data analysis and prediction as well as decision making. Theoretical results regarding the solution minimizing the mean absolute error are exploited to develop an exact branch-refine-and-bound-type algorithm for fitting Sugeno integrals (weighted lattice polynomial functions, max-min operators) with respect to symmetric capacities. The proposed method turns out to be particularly suitable for acting on ordinal data. In addition to providing a model that can be used for the general data regression task, the results can be used, among others, to calibrate generalized h-indices to bibliometric data.
Marek Gagolewski, Simon James, Gleb Beliakov
IEEE Trans. Fuzzy Syst.1
2018 Least Median of Squares (LMS) and Least Trimmed Squares (LTS) Fitting for the Weighted Arithmetic Mean
Gleb Beliakov, Marek Gagolewski, Simon James
IPMU (2)2
2017 Binary aggregation functions in software plagiarism detection
abstract
Supervised learning is of key interest in data science. Even though there exist many approaches to solving, among others, classification as well as ordinal and standard regression tasks, most of them output models that do not possess useful formal properties, like nondecreasingness in each independent variable, idempotence, symmetry, etc. This makes them difficult to interpret and analyze. For instance, it might be impossible to determine the importances of individual features or to assess the effects of increasing the values of predictors on the behavior of a chosen response variable. Such properties are especially important in software plagiarism detection, where we are faced with the combination of degrees to which how much a code chunk A is similar to (or contained in) B as well as how much B is similar to A. Therefore, in this paper we consider a new method for fitting B-spline tensor product-based aggregation functions to empirical data. An empirical study indicates a highly competitive performance of the resulting models. Additionally, they possess an intuitive interpretation which is highly desirable for end-users.
Maciej Bartoszuk, Marek Gagolewski
FUZZ-IEEE2
2017 OWA-based linkage and the genie correction for hierarchical clustering
abstract
In this paper we thoroughly investigate various OWA-based linkages in hierarchical clustering on numerous benchmark data sets. The inspected setting generalizes the well-known single, complete, and average linkage schemes, among others. The incorporation of weights into the cluster merge procedure creates an opportunity to make use of experts' knowledge about a particular data domain so as to generate partitions of a given data set that better reflect the true underlying cluster structure. Moreover, we introduce a correction for the inequality of cluster size distribution - similar to the one proposed in our recently introduced Genie algorithm - which results in a significant performance boost in terms of clustering quality.
Anna Cena, Marek Gagolewski
FUZZ-IEEE2
2017 Penalty-based aggregation of multidimensional data
Marek Gagolewski
Fuzzy Sets Syst.1
2016 Fitting Aggregation Functions to Data: Part I - Linearization and Regularization
Maciej Bartoszuk, Gleb Beliakov, Marek Gagolewski, Simon James
IPMU (2)3
2016 Fitting Aggregation Functions to Data: Part II - Idempotization
Maciej Bartoszuk, Gleb Beliakov, Marek Gagolewski, Simon James
IPMU (2)3
2016 Fuzzy K-Minpen Clustering and K-nearest-minpen Classification Procedures Incorporating Generic Distance-Based Penalty Minimizers
Anna Cena, Marek Gagolewski
IPMU (2)2
2016 Hierarchical Clustering via Penalty-Based Aggregation and the Genie Approach
Marek Gagolewski, Anna Cena, Maciej Bartoszuk
MDAI1
2016 Penalty-Based and Other Representations of Economic Inequality
abstract
Economic inequality measures are employed as a key component in various socio-demographic indices to capture the disparity between the wealthy and poor. Since their inception, they have also been used as a basis for modelling spread and disparity in other contexts. While recent research has identified that a number of classical inequality and welfare functions can be considered in the framework of OWA operators, here we propose a framework of penalty-based aggregation functions and their associated penalties as measures of inequality.
Gleb Beliakov, Marek Gagolewski, Simon James
Int. J. Uncertain. Fuzziness Knowl. Based Syst.2
2016 Genie: A new, fast, and outlier-resistant hierarchical clustering algorithm
Marek Gagolewski, Maciej Bartoszuk, Anna Cena
Inf. Sci.1
2016 H-Index and Other Sugeno Integrals: Some Defects and Their Compensation
abstract
The famous Hirsch index has been introduced just ca. ten years ago. Despite that, it is already widely used in many decision-making tasks, like in evaluation of individual scientists, research grant allocation, or even production planning. It is known that the h-index is related to the discrete Sugeno integral and the Ky Fan metric introduced in the 1940s. The aim of this paper is to propose a few modifications of this index as well as other fuzzy integrals-also on bounded chains-that lead to better discrimination of some types of data that are to be aggregated. All of the suggested compensation methods try to retain the simplicity of the original measure.
Radko Mesiar, Marek Gagolewski
IEEE Trans. Fuzzy Syst.2
2015 The winning solution to the AAIA'15 data mining competition: Tagging Firefighter Activities at a Fire Scene
abstract
© 2015, IEEE. Multi-sensor based classification of professionals' activities plays a key role in ensuring the success of an his/her goals. In this paper we present the winning solution to the AAIA'15 Tagging Firefighter Activities at a Fire Scene data mining competition. The approach is based on a Random Forest classifier trained on an input data set with almost 5000 features describing the underlying time series of sensory data.
Jan Lasek, Marek Gagolewski
FedCSIS2
2015 OM3: Ordered maxitive, minitive, and modular aggregation operators - Axiomatic and probabilistic properties in an arity-monotonic setting
Anna Cena, Marek Gagolewski
Fuzzy Sets Syst.2
2014 A Fuzzy R Code Similarity Detection Algorithm
Maciej Bartoszuk, Marek Gagolewski
IPMU (3)2
2014 Piecewise Linear Approximation of Fuzzy Numbers Preserving the Support and Core
Lucian C. Coroianu, Marek Gagolewski, Przemyslaw Grzegorzewski, M. Adabitabar Firozja, Tahereh Houlari
IPMU (2)2
2014 Monotone measures and universal integrals in a uniform framework for the scientific impact assessment problem
Marek Gagolewski, Radko Mesiar
Inf. Sci.1
2013 Nearest piecewise linear approximation of fuzzy numbers
Lucian C. Coroianu, Marek Gagolewski, Przemyslaw Grzegorzewski
Fuzzy Sets Syst.2
2013 On the relationship between symmetric maxitive, minitive, and modular aggregation operators
Marek Gagolewski
Inf. Sci.1
2012 On the Relation between Effort-Dominating and Symmetric Minitive Aggregation Operators
Marek Gagolewski
IPMU (3)1
2011 Possibilistic analysis of arity-monotonic aggregation operators and its relation to bibliometric impact assessment of individuals
Marek Gagolewski, Przemyslaw Grzegorzewski
Int. J. Approx. Reason.1
2010 Arity-Monotonic Extended Aggregation Operators
Marek Gagolewski, Przemyslaw Grzegorzewski
IPMU (1)1