EDBT 2026 Demo / reviewers in the wild / expert
Patryk Orzechowski
dblp:130/6513
· DBLP profile ↗
9ranked-venue papers
6as first author
2since 2021 · last 2025
0000-0003-3578-9809ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 4 first-author · 1 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 1 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
4 papers |
Bioinformatics and computational biology · 54% Medical and health informatics · 46% | |
| Artificial intelligence
1 paper |
Question answering and dialogue systems · 77% Language models and text generation · 23% | |
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
GPUs and heterogeneous computing · 68% Parallel and multicore computing · 32% |
Topics — the 5 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › gene expression analysis
biclustering |
1.0 | 3 | 2019 | EBIC: an open source software for high-dimensional and big data analyses · Bioinform. 2019 EBIC: an evolutionary-based parallel biclustering algorithm for pattern discovery · Bioinform. 2018 runibic: a Bioconductor package for parallel row-based biclustering of gene expression data · Bioinform. 2018 |
Medical and health informatics › mental health informatics
mental health conversational agents |
0.9 | 1 | 2025 | MentalChat16K: A Benchmark Dataset for Conversational Mental Health Assistance · KDD (2) 2025 |
Natural language and speech › Language models and text generation
large language model evaluation |
0.3 | 1 | 2025 | MentalChat16K: A Benchmark Dataset for Conversational Mental Health Assistance · KDD (2) 2025 |
GPUs and heterogeneous computing
multi-GPU computing |
0.2 | 2 | 2019 | EBIC: an open source software for high-dimensional and big data analyses · Bioinform. 2019 EBIC: an evolutionary-based parallel biclustering algorithm for pattern discovery · Bioinform. 2018 |
Parallel and multicore computing › parallel computing
parallel implementation |
0.1 | 1 | 2018 | runibic: a Bioconductor package for parallel row-based biclustering of gene expression data · Bioinform. 2018 |
Methods — techniques the papers use, named apart from their topics
dataset construction · 1.7evolutionary computation · 1.4GPU acceleration · 1.4parallelization · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MentalChat16K: A Benchmark Dataset for Conversational Mental Health AssistanceabstractWe introduce MentalChat16K, an English benchmark dataset combining a synthetic mental health counseling dataset and a dataset of anonymized transcripts from interventions between Behavioral Health Coaches and Caregivers of patients in palliative or hospice care. Covering a diverse range of conditions like depression, anxiety, and grief, this curated dataset is designed to facilitate the development and evaluation of large language models for conversational mental health assistance. By providing a high-quality resource tailored to this critical domain, MentalChat16K aims to advance research on empathetic, personalized AI solutions to improve access to mental health support services. The dataset prioritizes patient privacy, ethical considerations, and responsible data usage. MentalChat16K presents a valuable opportunity for the research community to innovate AI technologies that can positively impact mental well-being. The dataset is available at https://huggingface.co/datasets/ShenLab/MentalChat16K and the code and documentation are hosted on GitHub at https://github.com/PennShenLab/MentalChat16K. Tianyi Wei, Bojian Hou, Patryk Orzechowski, Shu Yang 0009, Ruochen Jin, Rachael Paulbeck, Joost B. Wagenaar, George Demiris, Li Shen 0001 |
KDD (2) | 4 |
| 2025 | QOT: Quantized Optimal Transport for sample-level distance matrix in single-cell omicsabstractSingle-cell technologies have enabled the high-dimensional characterization of cell populations at an unprecedented scale. The innate complexity and increasing volume of data pose significant computational and analytical challenges, especially in comparative studies delineating cellular architectures across various biological conditions (i.e. generation of sample-level distance matrices). Optimal Transport is a mathematical tool that captures the intrinsic structure of data geometrically and has been applied to many bioinformatics tasks. In this paper, we propose QOT (Quantized Optimal Transport), a new method enabling efficient computation of sample-level distance matrix from large-scale single-cell omics data through a quantization step. We apply our algorithm to real-world single-cell genomics and pathomics datasets, aiming to extrapolate cell-level insights to inform sample-level categorizations. Our empirical study shows that QOT outperforms existing two OT-based algorithms in accuracy and robustness when obtaining a distance matrix from high throughput single-cell measures at the sample level. Moreover, the sample level distance matrix could be used in the downstream analysis (i.e. uncover the trajectory of disease progression), highlighting its usage in biomedical informatics and data science. Zexuan Wang, Qipeng Zhan, Shu Yang 0009, Shizhuo Mu, Sumita Garai, Patryk Orzechowski, Joost B. Wagenaar, Li Shen 0001 |
Briefings Bioinform. | 7 |
| 2020 | Benchmarking Manifold Learning Methods on a Large Collection of Datasets
Patryk Orzechowski, Franciszek Magiera, Jason H. Moore |
EuroGP | 1 |
| 2019 | EBIC: an open source software for high-dimensional and big data analysesabstractMOTIVATION: In this paper, we present an open source package with the latest release of Evolutionary-based BIClustering (EBIC), a next-generation biclustering algorithm for mining genetic data. The major contribution of this paper is adding a full support for multiple graphics processing units (GPUs) support, which makes it possible to run efficiently large genomic data mining analyses. Multiple enhancements to the first release of the algorithm include integration with R and Bioconductor, and an option to exclude missing values from the analysis. RESULTS: Evolutionary-based BIClustering was applied to datasets of different sizes, including a large DNA methylation dataset with 436 444 rows. For the largest dataset we observed over 6.6-fold speedup in computation time on a cluster of eight GPUs compared to running the method on a single GPU. This proves high scalability of the method. AVAILABILITY AND IMPLEMENTATION: The latest version of EBIC could be downloaded from http://github.com/EpistasisLab/ebic. Installation and usage instructions are also available online. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Patryk Orzechowski, Jason H. Moore |
Bioinform. | 1 |
| 2018 | Where are we now?: a large benchmark study of recent symbolic regression methodsabstractIn this paper we provide a broad benchmarking of recent genetic programming approaches to symbolic regression in the context of state of the art machine learning approaches. We use a set of nearly 100 regression benchmark problems culled from open source repositories across the web. We conduct a rigorous benchmarking of four recent symbolic regression approaches as well as nine machine learning approaches from scikit-learn. The results suggest that symbolic regression performs strongly compared to state-of-the-art gradient boosting algorithms, although in terms of running times is among the slowest of the available methodologies. We discuss the results in detail and point to future research directions that may allow symbolic regression to gain wider adoption in the machine learning community. Patryk Orzechowski, William G. La Cava, Jason H. Moore |
GECCO | 1 |
| 2018 | runibic: a Bioconductor package for parallel row-based biclustering of gene expression dataabstractMotivation: Biclustering is an unsupervised technique of simultaneous clustering of rows and columns of input matrix. With multiple biclustering algorithms proposed, UniBic remains one of the most accurate methods developed so far. Results: In this paper we introduce a Bioconductor package called runibic with parallel implementation of UniBic. For the convenience the algorithm was reimplemented, parallelized and wrapped within an R package called runibic. The package includes: (i) a couple of times faster parallel version of the original sequential algorithm, (ii) much more efficient memory management, (iii) modularity which allows to build new methods on top of the provided one and (iv) integration with the modern Bioconductor packages such as SummarizedExperiment, ExpressionSet and biclust. Availability and implementation: The package is implemented in R and is available from Bioconductor (starting from version 3.6) at the following URL http://bioconductor.org/packages/runibic with installation instructions and tutorial. Supplementary information: Supplementary data are available at Bioinformatics online. Patryk Orzechowski, Artur Panszczyk, Xiuzhen Huang, Jason H. Moore |
Bioinform. | 1 |
| 2018 | EBIC: an evolutionary-based parallel biclustering algorithm for pattern discoveryabstractMotivation: Biclustering algorithms are commonly used for gene expression data analysis. However, accurate identification of meaningful structures is very challenging and state-of-the-art methods are incapable of discovering with high accuracy different patterns of high biological relevance. Results: In this paper, a novel biclustering algorithm based on evolutionary computation, a sub-field of artificial intelligence, is introduced. The method called EBIC aims to detect order-preserving patterns in complex data. EBIC is capable of discovering multiple complex patterns with unprecedented accuracy in real gene expression datasets. It is also one of the very few biclustering methods designed for parallel environments with multiple graphics processing units. We demonstrate that EBIC greatly outperforms state-of-the-art biclustering methods, in terms of recovery and relevance, on both synthetic and genetic datasets. EBIC also yields results over 12 times faster than the most accurate reference algorithms. Availability and implementation: EBIC source code is available on GitHub at https://github.com/EpistasisLab/ebic. Supplementary information: Supplementary data are available at Bioinformatics online. Patryk Orzechowski, Moshe Sipper, Xiuzhen Huang, Jason H. Moore |
Bioinform. | 1 |
| 2016 | Hybrid Biclustering Algorithms for Data Mining
Patryk Orzechowski, Krzysztof Boryczko |
EvoApplications (1) | 1 |
| 2016 | QtBiVis: a software toolbox for visual analysis of biclustering experimentabstractIn this article we introduce QtBiVis -a novel software intended for the comparative analysis of biclustering results.This modular tool has been efficiently implemented in C++ with Qt framework GUI.It may be successfully used for coverage analysis of the results of biclustering as well filtering or sorting biclusters by Gene Ontology (GO) identifiers or bicluster enrichment values.It may also be useful for parameter studies of biclustering algorithms.In future releases we plan to add different modules for visualizing and comparing different GO terms and biclusters. Artur Panszczyk, Patryk Orzechowski |
FedCSIS | 2 |