Filip Habarta

dblp:276/0409 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
6since 2021 · last 2024
0000-0003-4477-2224ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 5 since 2021Software engineering, systems software and programming languages · 6 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 5 since 2021Theory of computation · 2 · 1 since 2021
YearPublicationVenuePosition
2024 Non-parametric comparison of survival functions with censored data: A computational analysis of greedy and Monte Carlo approaches
abstract
Comparison of two survival functions, which describe the probability of not experiencing an event of interest by a given time point in two different groups, is a typical task in survival analysis.There are several well-established methods for comparing survival functions, such as the log-rank test and its variants.However, these methods often come with rigid statistical assumptions.In this work, we introduce a non-parametric alternative for comparing survival functions that is nearly free of assumptions.Unlike the log-rank test, which requires the estimation of hazard functions derived from (or facilitating the derivation of) survival functions and assumes a minimum number of observations to ensure asymptotic properties, our method models all possible scenarios based on observed data.These scenarios include those in which the compared survival functions differ in the same way or even more significantly, thus allowing us to calculate the p-value directly.Individuals in these groups may experience an event of interest at specific time points or may be censored, i.e., they might experience the event outside the observed time points.Focusing on all scenarios where survival probabilities differ at least as much as observed usually requires computationally intensive calculations.Censoring is treated as a form of noise, increasing the range of scenarios that need to be calculated and evaluated.Therefore, to estimate the p-value, we compare a computationally exhaustive approach that computes all possible scenarios in which groups' survival functions differ as observed or more, with a Monte Carlo simulation of these scenarios, alongside a traditional approach based on the log-rank test.Our proposed method reduces the first type error rate, enhancing its utility in studies where robustness against false positives is critical.We also analyze the asymptotic time complexity of both proposed approaches.
Lubomír Stepánek, Filip Habarta, Ivana Malá, Lubos Marek
FedCSIS2
2023 Let's estimate all parameters as probabilities: Precise estimation using Chebyshev's inequality, Bernoulli distribution, and Monte Carlo simulations
abstract
Regarding the parameter estimation task, besides the time effectiveness of the simulation, parameter estimates are required to be precise enough.Usually, the estimates are Monte Carlo-simulated using a prior estimated variability within a small sample.However, the problem with pre-estimated variability is that it can be estimated imprecisely or, even worse, underestimated, resulting in estimation bias.In this work, we address the abovementioned issue and suggest estimating all parameters as probabilities.Since the probability is not only finite but has its theoretical maximum as 1, using outcomes of Bernoulli and binomial distribution's upper-bounded variance and Chebyshev's inequality, the estimator's variability is theoretically upperbounded within the Monte Carlo simulation and estimation process.It cannot be underestimated or estimated inaccurately; thus, its precision is ensured till a given decimal digit, with very high probability.If there is a known process that treats the parameter of interest in terms of probability, we can estimate how many iterations of the Monte Carlo simulation are needed to ensure parameter estimate on a given level of precision.Also, we analyze the asymptotic time complexity of the proposed estimation strategy and illustrate the approach on a short case study of π constant estimation.
Lubomír Stepánek, Filip Habarta, Ivana Malá, Lubos Marek
FedCSIS2
2023 A lower bound for proportion of visibility polygon's surface to entire polygon's surface: Estimated by Art Gallery Problem and proven that cannot be greatly improved
abstract
Assuming a bounded polygon and a point inside the polygon or on its boundary, the visibility polygon, also called the visibility region, is a polygon reachable, i.e., visible by straight lines from the point without hitting the polygon's edges or vertices.If the polygon is bounded, then the visibility polygon is bounded, and the proportion of the visibility polygon's surface area to the given polygon's surface area could be enumerated.Many papers investigate applications of the visibility polygons in robotics and computer graphics or focus on computationally effective finding the visibility region for a given polygon.However, surprisingly, there seems to be no work estimating the proportion of a visibility polygon's surface to an entire polygon's surface or its bounds.Thus, in this paper, we search for a lower bound of the surface proportion of a visibility polygon to a given one.Assuming n-sided simple polygon, i.e., a polygon without holes and edge intersections, we apply the well-known art gallery problem and derive there is always a point inside the polygon or on its boundary that guarantees the proportion of the visibility polygon's surface to the entire polygon's surface is at least
Lubomír Stepánek, Filip Habarta, Ivana Malá, Lubos Marek
FedCSIS2
2022 A short note on post-hoc testing using random forests algorithm: Principles, asymptotic time complexity analysis, and beyond
abstract
When testing whether a continuous variable differs between categories of a factor variable or their combinations, taking into account other continuous covariates, one may use an analysis of covariance.Several post-hoc methods, such as Tukey's honestly significant difference test, Scheffé's, Dunn's, or Nemenyi's test are well-established when the analysis of covariance rejects the hypothesis there is no difference between any categories.However, these methods are statistically rigid and usually require meeting statistical assumptions.In this work, we address the issue using a random forest-based algorithm, practically assumption-free, classifying individual observations into the factor's categories using the dependent continuous variable and covariates on input.The higher the proportion of trees classifying the observations into two different categories is, the more likely a statistical difference between the categories is.To adjust the method's first-type error rate, we change random forest trees' complexity by pruning to modify the proportions of highly complex trees.Besides simulations that demonstrate a relationship between the tree pruning level, tree complexity, and first-type error rate, we analyze the asymptotic time complexity of the proposed random forest-based method compared to established techniques.
Lubomír Stepánek, Filip Habarta, Ivana Malá, Lubos Marek
FedCSIS2
2022 Application of an Inverse Dirichlet's Principle to Discrete Recreational Problems: Bound Estimation's Optimization Using Combinatorial Probability and Comparison of Numerical Bound Estimation Using Various Algorithms, Including Recursive Inclusion-Exclusion Principle
Lubomír Stepánek, Filip Habarta, Ivana Malá, Lubos Marek, Stefka Fidanova
WCO2
2021 A random forest-based approach for survival curves comparing: principles, computational aspects and asymptotic time complexity analysis
abstract
The log-rank test and Cox's proportional hazard model can be used to compare survival curves but are limited by strict statistical assumptions.In this study, we introduce a novel, assumption-free method based on a random forest algorithm able to compare two or more survival curves.A proportion of the random forest's trees with sufficient complexity is close to the test's p-value estimate.The pruning of trees in the model modifies trees' complexity and, thus, both the method's robustness and statistical power.The discussed results are confirmed using a simulation study, varying the survival curves and the tree pruning level.
Lubomír Stepánek, Filip Habarta, Ivana Malá, Lubos Marek
FedCSIS2
2020 Analysis of asymptotic time complexity of an assumption-free alternative to the log-rank test
abstract
Comparison of two time-event survival curves representing two groups of individuals' evolution in time is relatively usual in applied biostatistics.Although the log-rank test is the suggested tool how to face the above-mentioned problem, there is a rich statistical toolbox used to overcome some of the properties of the log-rank test.However, all of these methods are limited by relatively rigorous statistical assumptions.In this study, we introduce a new robust method for comparing two time-event survival curves.We briefly discuss selected issues of the robustness of the log-rank test and analyse a bit more some of the properties and mostly asymptotic time complexity of the proposed method.The new method models individual time-event survival curves in a discrete combinatorial way as orthogonal monotonic paths, which enables direct estimation of the p-value as it was originally defined.We also gently investigate how the surface of an area, bounded by two survival curves plotted onto a plane chart, is related to the test's p-value.Finally, using simulated time-event data, we check the robustness of the introduced method in comparison with the log-rank test.Based on the theoretical analysis and simulations, the introduced method seems to be a promising and valid alternative to the log-rank test, particularly in case on how to compare two time-event curves regardless of any statistical assumptions.
Lubomír Stepánek, Filip Habarta, Ivana Malá, Lubos Marek
FedCSIS2
2020 Reducing the First-Type Error Rate of the Log-Rank Test: Asymptotic Time Complexity Analysis of An Optimized Test's Alternative
Lubomír Stepánek, Filip Habarta, Ivana Malá, Lubos Marek
WCO@FedCSIS2