Denis Janiak

dblp:306/8791 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2025
0000-0003-1859-9093ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Language models and text generation · 74% Trustworthy machine learning · 26%

Topics — the 3 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
hallucination detection
1.722025
The Illusion of Progress: Re-evaluating Hallucination Detection in LLMs · EMNLP 2025
Hallucination Detection in LLMs Using Spectral Features of Attention Maps · EMNLP 2025
Machine learning › Trustworthy machine learning › AI safety
human-aligned evaluation
0.912025
The Illusion of Progress: Re-evaluating Hallucination Detection in LLMs · EMNLP 2025
Natural language and speech › Language models and text generation › low-resource language processing
low-resource language models
0.212022
This is the way: designing and compiling LEPISZCZE, a comprehensive NLP benchmark for Polish · NeurIPS 2022

Methods — techniques the papers use, named apart from their topics

laplacian eigenvalues · 0.9graph spectral analysis · 0.9attention map analysis · 0.9ROUGE · 0.9LLM-as-judge · 0.9model tracking · 0.6data versioning · 0.6
YearPublicationVenuePosition
2025 Hallucination Detection in LLMs Using Spectral Features of Attention Maps
abstract
Large Language Models (LLMs) have demonstrated remarkable performance across various tasks but remain prone to hallucinations.Detecting hallucinations is essential for safetycritical applications, and recent methods leverage attention map properties to this end, though their effectiveness remains limited.In this work, we investigate the spectral features of attention maps by interpreting them as adjacency matrices of graph structures.We propose the LapEigvals method, which utilizes the topk eigenvalues of the Laplacian matrix derived from the attention maps as an input to hallucination detection probes.Empirical evaluations demonstrate that our approach achieves stateof-the-art hallucination detection performance among attention-based methods.Extensive ablation studies further highlight the robustness and generalization of LapEigvals, paving the way for future advancements in the hallucination detection domain.
Jakub Binkowski, Denis Janiak, Albert Sawczyn, Bogdan Gabrys, Tomasz Kajdanowicz
EMNLP2
2025 The Illusion of Progress: Re-evaluating Hallucination Detection in LLMs
abstract
Large language models (LLMs) have revolutionized natural language processing, yet their tendency to hallucinate poses serious challenges for reliable deployment.Despite numerous hallucination detection methods, their evaluations often rely on ROUGE, a metric based on lexical overlap that misaligns with human judgments.Through comprehensive human studies, we demonstrate that while ROUGE exhibits high recall, its extremely low precision leads to misleading performance estimates.In fact, several established detection methods show performance drops of up to 45.9% when assessed using human-aligned metrics like LLM-as-Judge.Moreover, our analysis reveals that simple heuristics based on response length can rival complex detection techniques, exposing a fundamental flaw in current evaluation practices.We argue that adopting semantically aware and robust evaluation frameworks is essential to accurately gauge the true performance of hallucination detection methods, ultimately ensuring the trustworthiness of LLM outputs.
Denis Janiak, Jakub Binkowski, Albert Sawczyn, Bogdan Gabrys, Ravid Shwartz-Ziv, Tomasz Kajdanowicz
EMNLP1
2022 This is the way: designing and compiling LEPISZCZE, a comprehensive NLP benchmark for Polish
abstract
The availability of compute and data to train larger and larger language models increases the demand for robust methods of benchmarking the true progress of LM training. Recent years witnessed significant progress in standardized benchmarking for English. Benchmarks such as GLUE, SuperGLUE, or KILT have become a de facto standard tools to compare large language models. Following the trend to replicate GLUE for other languages, the KLEJ benchmark\ (klej is the word for glue in Polish) has been released for Polish. In this paper, we evaluate the progress in benchmarking for low-resourced languages. We note that only a handful of languages have such comprehensive benchmarks. We also note the gap in the number of tasks being evaluated by benchmarks for resource-rich English/Chinese and the rest of the world.In this paper, we introduce LEPISZCZE (lepiszcze is the Polish word for glew, the Middle English predecessor of glue), a new, comprehensive benchmark for Polish NLP with a large variety of tasks and high-quality operationalization of the benchmark.We design LEPISZCZE with flexibility in mind. Including new models, datasets, and tasks is as simple as possible while still offering data versioning and model tracking. In the first run of the benchmark, we test 13 experiments (task and dataset pairs) based on the five most recent LMs for Polish. We use five datasets from the Polish benchmark and add eight novel datasets. As the paper's main contribution, apart from LEPISZCZE, we provide insights and experiences learned while creating the benchmark for Polish as the blueprint to design similar benchmarks for other low-resourced languages.
Lukasz Augustyniak, Kamil Tagowski, Albert Sawczyn, Denis Janiak, Roman Bartusiak, Adrian Szymczak, Arkadiusz Janz, Piotr Szymanski, Marcin Watroba, Mikolaj Morzy, Tomasz Kajdanowicz, Maciej Piasecki
NeurIPS4
2021 Fact-checking: relevance assessment of references in the Polish political domain
abstract
The prevalence of fake news could be observed in circumstances of emotion-causing events, like elections or pandemics. In fear of the potential impact, many fact-checking organisations were established. However, fact-checking requires a large amount of human labor, and hence there is a strong demand for complete automation of this process. Nevertheless, this milestone has not been achieved yet, even for English. The problem grows for the less popular languages that suffer from a scarcity of available resources. To address this problem for the Polish language domain, we propose a solution for automating one of the fact-checking stages - relevance assessment, which is crucial when searching for evidence. Leveraging recent advancements in natural language processing, we have acquired relevant data and developed classifiers of evidence relevance with respect to claims in Polish. Our approach can assess the evidence relevance with a performance at a level of a 0.778 F1-score.
Albert Sawczyn, Jakub Binkowski, Denis Janiak, Lukasz Augustyniak, Tomasz Kajdanowicz
KES3