VLDB 2026 Research / reviewers in the wild / expert
Janusz Wojtusiak
dblp:41/5352
· DBLP profile ↗
28ranked-venue papers
9as first author
6since 2021 · last 2026
0000-0003-2238-0588ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 17 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 10 · 4 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Case-Based Framework for Explaining AI-Driven Bruise Detection Model Behavior Using Structured Data
Dharmi Desai, Mehrdad Ghyabi, David Lattanzi, Katherine Scafide, Janusz Wojtusiak |
AIME (1) | 5 |
| 2025 | Influence of Stratified Variable Encoding on Quality of Mortality Prediction in Systolic Heart FailureabstractThis study investigates how stratified variable encoding strategies affect the performance of machine learning (ML) models in predicting in-hospital mortality among ICU patients with systolic heart failure (HF). Effective data preprocessing, particularly encoding of clinical variables, is essential for transforming raw data into formats suitable for analysis. Despite its significance, optimal encoding methods remain poorly defined, with implications for model accuracy and robustness. Data were sourced from the MIMIC-IV dataset, focusing on 1,037 adult ICU patients within the first 24 hours of admission. Forty-three patient characteristics, including 39 laboratory test results and demographic data, were encoded using four approaches: binary (present/not present), mean/maximum/minimum, first/last, and a combined method integrating all of the encodings. Gradient boosting, random forest, and logistic regression models were trained and evaluated using the area under the curve (AUC), accuracy, precision, and recall metrics. Feature importance was measured using the random forest method with stratified analyses to evaluate optimal encoding techniques. Model performance varied by the encoding method and algorithm. Random forest achieved the highest AUC with binary (0.833) and first/last (0.920), while gradient boosting achieved the highest AUC with average/minimum/maximum (0.911) and combination (0.931). In stratified evaluation, the highest average AUC further improved performance (AUC = 0.939) in comparison to the highest minimum (AUC = 0.930) and the highest maximum-minimum (AUC = 0.907). These findings emphasize the importance of encoding selection in ML model development. Tailored encoding strategies can enhance model performance, providing critical insights for clinical decision-making and improved patient outcomes. Bhumi Patel, Lemba Priscille Ngana, Atefehsadat Haghighathoseini, Janusz Wojtusiak |
ICMLA | 4 |
| 2024 | Big Data Decision-Making and Racial Disparities: A Case Study Among COVID-19 Inpatient VisitsabstractThe COVID-19 pandemic has had a disproportionate impact on certain racial and ethnic groups, resulting in significant health outcome disparities. The National COVID Cohort Collaborative (N3C) provides a valuable resource for exploring these disparities through big data analytics. This study belongs to a broader work that examines decisions made during data processing and their impact on the analyses performed. Central to our analysis is the introduction of the Continuous Inpatient Encounter (CIE) concept—a novel method we propose for aggregating inpatient visits. By utilizing big data analytics, we aim to identify potential disparities in CIE rates among different racial groups. The results of this study are critical for enhancing the equity of data-driven decision-making in healthcare and for addressing the racial disparities observed in COVID-19 outcomes. Atefehsadat Haghighathoseini, Janusz Wojtusiak, Nirup M. Menon, Hua Min, Cara Frankenfeld, Timothy Leslie |
IEEE Big Data | 2 |
| 2021 | Dashboard for Machine Learning Models in Health Care
Wejdan Bagais, Janusz Wojtusiak |
AMIA | 2 |
| 2021 | Active Learning Based User-Defined Chest X-Ray Diagnosis System Leveraging 5G Infrastructure for COVID-19 Variant Detection
Ying Wang 0062, Janusz Wojtusiak |
AMIA | 2 |
| 2021 | Do Clusters and Sequence of Symptoms Predict COVID-19 Test Results?
Janusz Wojtusiak, Wejdan Bagais, Hedyeh Mobahi, Elina Guralnik, Amira Roess, Farrokh Alemi |
AMIA | 1 |
| 2020 | Using Landmark Information to Enhance Location Prediction for Missing People with Alzheimer's Disease
Reyhaneh Mogharab Nia, Janusz Wojtusiak |
AMIA | 2 |
| 2019 | Generic vs. Specialized Models for Assessing Disabilities among People with Dementia
Negin Asadzadehzanjani, Janusz Wojtusiak, Allison E. Williams, Cari Levy |
AMIA | 2 |
| 2019 | Synthetic Data for Teaching Data Integration in Informatics Graduate Program
Hedyeh Mobahi, Hua Min, Janusz Wojtusiak |
AMIA | 3 |
| 2019 | Using GPS to Locate the Missing and Track Wandering in People with Alzheimer's Disease
Janusz Wojtusiak, Reyhaneh Mogharab Nia, U. Nalla B. Durai, Beverly Middle, Robert Koester, Catherine Tompkins |
AMIA | 1 |
| 2018 | Weighted Itemsets Error (WIE) Approach for Evaluating Generated Synthetic Patient DataabstractPatient data are regarded as highly sensitive and protected information by federal, state and local policies that make it available to only those who have been given access to Protected Health Information (PHI). In many applications, the access to PHI and real patient data can be substituted with generated realistic synthetic data used instead of real patient data. While methods exist that can generate synthetic data, it is unclear how to evaluate synthetic data quality. The objective of this paper is to present investigation of a new method for statistically testing the quality of synthetic patient data. Weighted Itemsets Error (WIE) measure compares frequent itemsets in the synthetic data with expected itemsets in real data, thus allowing for evaluating cooccurrence of data items. The derived measure is tested in the context of synthetic data comprising of medical diagnoses. The results demonstrate the effects of parameters that control WIE measure, and indicate that WIE is a simple yet powerful approach for evaluating synthetic datasets. Mojtaba Zare, Janusz Wojtusiak |
ICMLA | 2 |
| 2016 | Towards Automated Selection of Patient-Specific Education Materials in Ambulatory Care Settings
Fatemah M. Aloudah, Janusz Wojtusiak |
AMIA | 2 |
| 2016 | Applying Machine Learning Methods to Predict Activities of Daily Living for Cancer Patients
Hua Min, Talha Oz, Sava Vukomanovic, Hedyeh Mobahi, Katherine Irvin, Ilirjeta Krasniqi, Janusz Wojtusiak |
AMIA | 7 |
| 2016 | Visualizing the Effects of Cancers on Relationships Between Comorbidities and Activities of Daily Living
Hua Min, Talha Oz, Sava Vukomanovic, Hedyeh Mobahi, Katherine Irvin, Ilirjeta Krasniqi, Janusz Wojtusiak |
AMIA | 7 |
| 2014 | Creating Clinically Homogeneous Groups of Prostate Cancer Patients
Janusz Wojtusiak, Che Ngufor, Lorens Helmchen, Jack Hadley |
AMIA | 1 |
| 2013 | Machine Learning-based Detection of Health Data Elements
Andrej Kolacevski, Janusz Wojtusiak |
AMIA | 2 |
| 2013 | Mining Progress Notes for Prediction of Activities of Daily Living
Talha Oz, Che Ngufor, Janusz Wojtusiak |
AMIA | 3 |
| 2012 | Comparison of Classification Learning Methods for Medical Claims Payments
Katherine Irvin, Janusz Wojtusiak |
AMIA | 2 |
| 2012 | Recent Advances in AQ21 Rule Learning System for Healthcare Data
Janusz Wojtusiak |
AMIA | 1 |
| 2012 | Semantic Data Types in Machine Learning from Healthcare DataabstractHealthcare is particularly rich in semantic information and background knowledge describing data. This paper discusses a number of semantic data types that can be found in healthcare data, presents how the semantics can be extracted from existing sources including the Unified Medical Language System (UMLS), discusses how the semantics can be used in both supervised and unsupervised learning, and presents an example rule learning system that implements several of these types. Results from three example applications in the healthcare domain are used to further exemplify semantic data types. Janusz Wojtusiak |
ICMLA (1) | 1 |
| 2012 | Reasoning with unknown, not-applicable and irrelevant meta-values in concept learning and pattern discovery
Ryszard S. Michalski, Janusz Wojtusiak |
J. Intell. Inf. Syst. | 2 |
| 2011 | Agent-based Pickup and Delivery Planning: The Learnable Evolution Model ApproachabstractThe Dynamic Vehicle Routing Problem (DVRP) is an optimization problem in which agents deliver orders that are not known in advance to the routing. Partial solutions need to be adapted to continuously accommodate new orders within dynamically changing conditions. This research focuses on using a combination of multiagent-based autonomous control with non-Darwinian evolutionary optimization. In order to compile transport plans and render optimized decisions agents managing transport vehicles employ a guided evolutionary computation method, called the learnable evolution model (LEM). Implementation and experimental evaluation of the method is performed within the Plasma multiagent simulation platform. Janusz Wojtusiak, Tobias Warden, Otthein Herzog |
CISIS | 1 |
| 2010 | Combining Rule Induction and Reinforcement Learning: An Agent-based Vehicle RoutingabstractReinforcement learning suffers from inefficiency when the number of potential solutions to be searched is large. This paper describes a method of improving reinforcement learning by applying rule induction in multi-agent systems. Knowledge captured by learned rules is used to reduce search space in reinforcement learning, allowing it to shorten learning time. The method is particularly suitable for agents operating in dynamically changing environments, in which fast response to changes is required. The method has been tested in transportation logistics domain in which agents represent vehicles being routed in a simple road network. Experimental results indicate that in this domain the method performs better than traditional Q-learning, as indicated by statistical comparison. Bartlomiej Sniezynski, Wojciech Wójcik, Jan D. Gehrke, Janusz Wojtusiak |
ICMLA | 4 |
| 2008 | Computational intelligence virtual community: Framework and implementation issuesabstractThis paper discusses the framework for virtual collaborative environment for researchers, practitioners, users and learners in the areas of computational intelligence and machine learning (CIML) that is currently developed by our group. It also outlines main features of the community portal under construction that will support communication and sharing of computational resources. In particular, selected aspects of structure of the portal such as common formats of data, models, software, publications and software documentation are discussed. Jacek M. Zurada, Janusz Wojtusiak, Fahmida Chowdhury, James E. Gentle, Cedric J. Jeannot, Maciej A. Mazurowski |
IJCNN | 2 |
| 2007 | The Natural Induction System AQ21 and its Application to Data Describing Patients with Metabolic Syndrome: Initial ResultsabstractThis paper briefly describes the AQ21 learning system that implements a simple form of natural induction, an approach to learning that generates hypotheses in forms resembling natural language descriptions, and by that easy to understand and interpret. The system was applied to the analysis of aggregated data obtained from non-invasive tests performed on different groups of patients with metabolic syndrome. The discovered patterns were very simple and were evaluated by an expert as potentially medically significant. Janusz Wojtusiak, Ryszard S. Michalski, Thipkesone Simanivanh, Ancha V. Baranova |
ICMLA | 1 |
| 2006 | The LEM3 implementation of learnable evolution model and its testing on complex function optimization problemsabstractLearnable Evolution Model (LEM) is a form of non-Darwinian evolutionary computation that employs machine learning to guide evolutionary processes. Its main novelty are new type of operators for creating new individuals, specifically, hypothesis generation, which learns rules indicating subareas in the search space that likely contain the optimum, and hypothesis instantiation, which populates these subspaces with new individuals. This paper briefly describes the newest and most advanced implementation of learnable evolution, LEM3, its novel features, and results from its comparison with a conventional, Darwinian-type evolutionary computation program (EA), a cultural evolution algorithm (CA), and the estimation of distribution algorithm (EDA) on selected function optimization problems (with the number of variables varying up to 1000). In every experiment, LEM3 outperformed the compared programs in terms of the evolution length (the number of fitness evaluations needed to achieved a desired solution), sometimes more than by one order of magnitude. Janusz Wojtusiak, Ryszard S. Michalski |
GECCO | 1 |
| 2006 | Intelligent Optimization via Learnable Evolution ModelabstractA new method for optimizing complex functions and systems is described that employs learnable evolution model (LEM), a form of non-Darwinian evolutionary computation guided by machine learning. LEM's main novelties are operators for creating new individuals that include hypothesis generation, which learns rules indicating subareas in the search space likely containing the optimum, and hypothesis instantiation, which populates these subareas with new candidate solutions. LEM3, the newest and most advanced implementation of learnable evolution, is briefly described and experimentally compared with other evolutionary computation programs on selected function optimization problems. We also describe two specialized LEM-based systems for heat exchanger optimization Ryszard S. Michalski, Janusz Wojtusiak, Kenneth A. Kaufman |
ICTAI | 2 |
| 2006 | The AQ21 Natural Induction Program for Pattern Discovery: Initial Version and its Novel FeaturesabstractThe AQ21 program aims to perform natural induction, a process of generating inductive hypotheses in human-oriented forms that are easy to interpret and understand. This is achieved by employing a highly expressive representation language, attributional calculus, whose statements resemble natural language descriptions. This paper focuses on the pattern discovery mode of AQ21, which produces attributional rules that capture strong regularities in the data, but may not be fully consistent or complete with regard to the training data. AQ21 integrates several novel features, such as optimizing patterns according to multiple criteria, learning attributional rules with exceptions, generating optimized sets of alternative hypotheses, and handling data with unknown, irrelevant and/or non-applicable meta-values Janusz Wojtusiak, Ryszard S. Michalski, Kenneth A. Kaufman, Jaroslaw Pietrzykowski |
ICTAI | 1 |