VLDB 2026 Research / reviewers in the wild / expert
Jan Hula
dblp:200/9929
· DBLP profile ↗
16ranked-venue papers
5as first author
14since 2021 · last 2026
0000-0001-7639-864XORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 5 first-author · 12 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Neural approaches to SAT solving: Design choices and interpretabilityabstractIn this contribution, we provide a comprehensive evaluation of graph neural networks applied to Boolean satisfiability problems, accompanied by an intuitive explanation of the mechanisms enabling the model to generalize to different instances. We introduce several training improvements, particularly a novel closest assignment supervision method that dynamically adapts to the model’s current state, significantly enhancing performance on problems with larger solution spaces. Our experiments demonstrate the suitability of variable-clause graph representations with recurrent neural network updates, which achieve good accuracy on SAT assignment prediction while reducing computational demands. We extend the base graph neural network into a diffusion model that facilitates incremental sampling and can be effectively combined with classical techniques like unit propagation. Through analysis of embedding space patterns and optimization trajectories, we show how these networks implicitly perform a process very similar to continuous relaxations of MaxSAT, offering an interpretable view of their reasoning process. This understanding guides our design choices and explains the ability of recurrent architectures to scale effectively at inference time beyond their training distribution, which we demonstrate with test-time scaling experiments. David Mojzísek, Jan Hula, Ziyu Zhou 0013, Mikolás Janota |
Int. J. Approx. Reason. | 2 |
| 2025 | BenCzechMark : A Czech-Centric Multitask and Multimetric Benchmark for Large Language Models with Duel Scoring MechanismabstractAbstract We present BenCzechMark (BCM), the first comprehensive Czech language benchmark designed for large language models, offering diverse tasks, multiple task formats, and multiple evaluation metrics. Its duel scoring system is grounded in statistical significance theory and uses aggregation across tasks inspired by social preference theory. Our benchmark encompasses 50 challenging tasks, with corresponding test datasets, primarily in native Czech, with 14 newly collected ones. These tasks span 8 categories and cover diverse domains, including historical Czech news, essays from pupils or language learners, and spoken word. Furthermore, we collect and clean BUT-Large Czech Collection, the largest publicly available clean Czech language corpus, and use it for (i) contamination analysis and (ii) continuous pretraining of the first Czech-centric 7B language model with Czech-specific tokenization. We use our model as a baseline for comparison with publicly available multilingual models. Lastly, we release and maintain a leaderboard with existing 50 model submissions, where new model submissions can be made at https://huggingface.co/spaces/CZLC/BenCzechMark. Martin Fajcik, Martin Docekal, Jan Dolezal, Karel Ondrej, Karel Benes, Jan Kapsa, Pavel Smrz, Alexander Polok, Michal Hradis, Zuzana Neverilová, Ales Horák, Radoslav Sabol, Michal Stefánik, Adam Jirkovsky, David Adamczyk, Petr Hyner, Jan Hula, Hynek Kydlícek |
Trans. Assoc. Comput. Linguistics | 17 |
| 2024 | Understanding GNNs for Boolean Satisfiability through Approximation AlgorithmsabstractThis paper delves into the interpretability of Graph Neural Networks in the context of Boolean Satisfiability. The goal is to demystify the internal workings of these models and provide insightful perspectives into their decision-making processes. This is done by uncovering connections to two approximation algorithms studied in the domain of Boolean Satisfiability: Belief Propagation and Semidefinite Programming Relaxations. Revealing these connections has empowered us to introduce a suite of impactful enhancements. The first significant enhancement is a curriculum training procedure, which incrementally increases the problem complexity in the training set, together with increasing the number of message passing iterations of the Graph Neural Network. We show that the curriculum, together with several other optimizations, reduces the training time by more than an order of magnitude compared to the baseline without the curriculum. Furthermore, we apply decimation and sampling of initial embeddings, which significantly increase the percentage of solved problems. Jan Hula, David Mojzísek, Mikolás Janota |
CIKM | 1 |
| 2024 | Towards an Understanding of Stepwise Inference in Transformers: A Synthetic Graph Navigation ModelabstractStepwise inference protocols, such as scratchpads and chain-of-thought, help language models solve complex problems by decomposing them into a sequence of simpler subproblems. To unravel the underlying mechanisms of stepwise inference we propose to study autoregressive Transformer models on a synthetic task that embodies the multi-step nature of problems where stepwise inference is generally most useful. Specifically, we define a graph navigation problem wherein a model is tasked with traversing a path from a start to a goal node on the graph. We find we can empirically reproduce and analyze several phenomena observed at scale: (i) the stepwise inference reasoning gap, the cause of which we find in the structure of the training data; (ii) a diversity-accuracy trade-off in model generations as sampling temperature varies; (iii) a simplicity bias in the model’s output; and (iv) compositional generalization and a primacy bias with in-context exemplars. Overall, our work introduces a grounded, synthetic framework for studying stepwise inference and offers mechanistic hypotheses that can lay the foundation for a deeper understanding of this phenomenon. Mikail Khona, Maya Okawa, Jan Hula, Rahul Ramesh, Kento Nishi, Robert P. Dick, Ekdeep Singh Lubana, Hidenori Tanaka |
ICML | 3 |
| 2024 | Efficient Use of Large Language Models for Analysis of Text Corpora
David Adamczyk, Jan Hula |
ICPRAM | 2 |
| 2024 | Efficient Solver Scheduling and Selection for Satisfiability Modulo Theories (SMT) Problems
David Mojzísek, Jan Hula |
ICPRAM | 2 |
| 2024 | Stealing Brains: From English to Czech Language Model
Petr Hyner, Petr Marek, David Adamczyk, Jan Hula, Jan Sedivý |
IJCCI | 4 |
| 2023 | Fast Heuristic for Ricochet Robots
Jan Hula, David Adamczyk, Mikolás Janota |
ICAART (1) | 1 |
| 2023 | Molecule Builder: Environment for Testing Reinforcement Learning Agents
Petr Hyner, Jan Hula, Mikolás Janota |
IJCCI | 2 |
| 2022 | 3D Shapes Classification Using Intermediate Parts Representation
Jan Hula, David Mojzísek, David Adamczyk |
IPMU (2) | 1 |
| 2022 | Targeted Configuration of an SMT Solver
Jan Hula, Jan Jakubuv, Mikolás Janota, Lukás Kubej |
CICM | 1 |
| 2022 | Poly-YOLO: higher speed, more precise detection and instance segmentation for YOLOv3
Petr Hurtík, Vojtech Molek, Jan Hula, Marek Vajgl, Pavel Vlasánek, Tomas Nejezchleba |
Neural Comput. Appl. | 3 |
| 2022 | Binary cross-entropy with dynamical clipping
Petr Hurtík, Stefania Tomasiello, Jan Hula, David Hynar |
Neural Comput. Appl. | 3 |
| 2021 | Graph Neural Networks for Scheduling of SMT SolversabstractThis paper develops an approach to the scheduling of solvers in the domain of Satisfiability Modulo Theories (SMT) using a Graph Neural Network (GNN). In contrast to related methods, GNNs do not require manual feature design as they enable discovering relevant features in the raw data. We train them to predict the effectivity of individual solvers on a given problem. Rather than choosing only one solver with the best prediction, we schedule the solvers by ordering them according to the predicted runtime and dividing the overall runtime into all solvers uniformly. We compare our approach to several baselines. In the selected benchmarks, we show a substantial improvement over these baselines in terms of the number of solved problems and overall solving time. Jan Hula, David Mojzísek, Mikolás Janota |
ICTAI | 1 |
| 2020 | Data Preprocessing Technique for Neural Networks Based on Image Represented by a Fuzzy FunctionabstractAlthough data preprocessing is a universal technique that can be widely used in neural networks (NNs), most research in this area is focused on designing new NN architectures. This paper, we propose a preprocessing technique that enriches the original image data using local intensity information; this technique is motivated by human perception. To encode this information into an image, we introduce a new image structure named image represented by a fuzzy function. When using this structure, a crisp intensity value of each pixel is replaced by a fuzzy set given by a membership function constructed with the usage of extremal values from the particular neighborhood of that pixel. We describe this structure and its properties and propose a way in which it can be used as an input into existing NNs without any modifications. Based on our benchmark consisting of three well-known datasets and five NN architectures, we show that the proposed preprocessing can, in most cases, decrease classification error compared with a baseline and two other preprocessing methods. To support our claim, we have also selected several publicly available projects and tested the impact of the preprocessing with a positive result. Petr Hurtík, Vojtech Molek, Jan Hula |
IEEE Trans. Fuzzy Syst. | 3 |
| 2019 | Can You Tell Me How to Get Past Sesame Street? Sentence-Level Pretraining Beyond Language ModelingabstractAlex Wang, Jan Hula, Patrick Xia, Raghavendra Pappagari, R. Thomas McCoy, Roma Patel, Najoung Kim, Ian Tenney, Yinghui Huang, Katherin Yu, Shuning Jin, Berlin Chen, Benjamin Van Durme, Edouard Grave, Ellie Pavlick, Samuel R. Bowman. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019. Jan Hula, Patrick Xia 0002, Raghavendra Pappagari, Tom McCoy 0001, Roma Patel, Najoung Kim, Ian Tenney, Katherin Yu, Shuning Jin, Berlin Chen, Benjamin Van Durme, Edouard Grave, Ellie Pavlick, Samuel R. Bowman |
ACL (1) | 2 |