Petr Hájek 0002

dblp:75/3354-2 · DBLP profile ↗
← Back
38ranked-venue papers
24as first author
13since 2021 · last 2026
0000-0001-5579-1215ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 36 · 22 first-author · 12 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Financial statement fraud detection using topic-driven financial sentiment analysis
abstract
Financial statement fraud undermines market integrity and incurs substantial costs for investors, regulators, and companies. Text-based detection methods have emerged as useful complements to traditional financial indicators, but many fail to incorporate domain-specific topics or sentiment cues, often missing subtle changes in deceptive communication. To overcome this problem, this study proposes a topic-driven financial sentiment analysis (TDFSA) model that detects corporate fraud by analyzing linguistic patterns in the Management Discussion & Analysis (MD&A) sections of annual reports. Our approach captures contextual sentiment within financially relevant topics using FinBERT embeddings. To evaluate these signals in fraud detection, we integrate the TDFSA outputs into a broader cost-sensitive evaluation framework. This framework combines text-based indicators with financial ratios to balance the need to avoid false alarms with the high cost of undetected fraud. Using data from U.S. firms flagged in SEC Accounting and Auditing Enforcement Releases from 2014 to 2024 and matched non-fraud peers, we examine trends in financial ratios, textual complexity, and sentiment dynamics in the three years preceding fraud events. The results show that models leveraging TDFSA achieve higher detection accuracy and lower cost than dictionary-based sentiment, generic topic models, and deep learning baselines. • Topic-driven financial sentiment analysis (TDFSA) improves financial statement fraud detection. • FinBERT embeddings capture both topic-level and sentiment context in MD&A disclosures. • Cost-sensitive learning prioritizes preventing undetected fraud over false alarms, using a 6.46:1 ratio. • The proposed model provides accurate and fair decision support for auditors and investors.
Petr Hájek 0002, Josef Novotny, Michal Munk
Decis. Support Syst.1
2026 Responsible cross-lingual hate speech moderation with context-adaptive transformers
Andrew Asante, Petr Hájek 0002
Eng. Appl. Artif. Intell.2
2026 A multi-objective framework for predicting public opinion trends on infectious diseases using NSGA-II and interval predictions
abstract
Predicting public opinion trends during major infectious disease outbreaks is critical for guiding effective public health responses. However, predicting public opinion remains challenging because it is influenced by socio-economic, psychological, and media factors. This paper presents a novel framework for predicting public opinion trends related to significant infectious diseases, with a focus on COVID-19 as a case study. The proposed framework identifies the key factors influencing public opinion development and enables both point and interval predictions. The framework uses information ecology theory and applies the NSGA-II algorithm to select the features that best drive public opinion trends. By incorporating this framework, accurate point forecasts are produced alongside prediction intervals, effectively quantifying the uncertainty inherent in public opinion dynamics. This approach minimizes the quality-driven loss function to generate precise prediction intervals, providing decision-makers with critical insights into public opinion fluctuations during epidemics. The results offer valuable, real-time public sentiment warnings, supporting timely and effective interventions in epidemic prevention and control efforts.
Futian Weng, Petr Hájek 0002, Mohammad Zoynul Abedin
Expert Syst. Appl.3
2026 Two-Stage feature selection for early warning of default risk
Zhe Li 0047, Mohammad Zoynul Abedin, Petr Hájek 0002, Brian Lucey
Knowl. Based Syst.4
2025 Combining Large-Scale and Domain-Specific Datasets for Hate Speech Severity Modeling: A Regression-Based Approach
Andrew Asante, Petr Hájek 0002
IJCCI (3)2
2025 Multimodal Financial Sentiment for Stock Return Prediction
abstract
This paper proposes a novel multimodal deep learning framework for stock return prediction that integrates heterogeneous data sources: technical indicators, market investor sentiment indices, and textual sentiment extracted from earnings conference call transcripts. The proposed model employs a hybrid architecture combining transformer encoder for the technical modality and neural networks for market and textual modalities. A modality-level attention mechanism is used in a late fusion setup to dynamically weight the contributions of each modality. We evaluate our model on a large-scale dataset comprising 24,821 samples from 497 S&P 500 companies over the period 2010–2022. The results show that our model outperforms traditional models (LSTM, BiL-STM, CNN-LSTM) and alternative fusion strategies, achieving a directional accuracy of 59.94% on the test set. Attention weight analysis confirms that all three modalities contribute meaningfully to prediction performance. These results demonstrate the overall effectiveness of the proposed framework in accurately predicting abnormal stock returns in a multimodal setting.
Petr Hájek 0002, Josef Novotny, Michal Munk, Dasa Munková
KES1
2025 Enhancing IoT Intrusion Detection Performance Using Autoencoder-Based Feature Optimization and K-Means Clustering
abstract
Deep learning plays a critical role in designing intrusion detection systems (IDS) to protect Internet of Things (IoT) environments against cyberattacks. However, the performance of DL-based IDS models heavily depends on the quality and balance of the training data. Real-world intrusion detection datasets often suffer from severe class imbalance, causing models to become biased toward majority attack types and underperform in detecting rare threats. While various techniques have been proposed to address class imbalance, the high dimensionality and complexity of IoT datasets remain significant challenges in building effective classifiers. Due to the dynamic nature of cyberattacks, no single method can fully address the diverse security threats in IoT networks. As a result, hybrid approaches have gained traction in enhancing cybersecurity solutions. This paper presents a novel hybrid method that combines an autoencoder and K-means clustering to generate a synthetic dataset that balances minority classes in the training set. The autoencoder reduces feature dimensionality, while K-means clustering supports oversampling of underrepresented classes. DL models are then employed for multi-class attack classification. The proposed approach is evaluated on the recent CICIoT2023 dataset. Experimental results demonstrate substantial improvements in recall, precision, and F1-score, especially in detecting minority class attacks. These findings indicate that the proposed method improves detection accuracy for rare intrusions, reduces false alarms, and supports administrators in deploying more effective IoT security measures.
Zeru Kifle Kebede, Petr Hájek 0002
KES2
2024 Enhancing cardiovascular risk assessment with advanced data balancing and domain knowledge-driven explainability
abstract
In medical risk prediction, such as predicting heart disease, machine learning (ML) classifiers must achieve high accuracy, precision, and recall to minimize the chances of incorrect diagnoses or treatment recommendations. However, real-world datasets often have imbalanced data, which can affect classifier performance. Traditional data balancing methods can lead to overfitting and underfitting, making it difficult to identify potential health risks accurately. Early prediction of heart attacks is of paramount importance, and researchers have developed ML-based systems to address this problem. However, much of the existing ML research is based on a single dataset, often ignoring performance evaluation across multiple datasets. As the demand for interpretable ML models grows, model interpretability becomes central to revealing insights and feature effects within predictive models. To address these challenges, we present a novel data balancing technique that uses a divide-and-conquer strategy with the K-Means clustering algorithm to segment the dataset. The performance of our approach is highlighted through comparisons with established techniques, which demonstrate the superiority of our proposed method. To address the challenge of inter-dataset discrepancies, we use two different datasets. Our holistic pipeline, strengthened by the innovative balancing technique, effectively addresses performance discrepancies, culminating in a significant improvement from 81% to 90%. Furthermore, through advanced statistical analysis, it has been determined that the 95% confidence interval for the AUC metric of our method ranges from 0.8187 to 0.8411. This observation serves to underscore the consistency and reliability of our approach, demonstrating its ability to achieve high performance across a range of scenarios. Incorporating Explainable AI (XAI), we examine the feature rankings and their contributions within the best performing Random Forest model. While the domain expert feedback is consistent with the explanatory power of XAI, some differences remain. Nevertheless, a remarkable convergence in feature ranking and weighting is observed, bridging the insights from XAI tools and domain expert perspectives.
Fan Yang 0033, Yanan Qiao, Petr Hájek 0002, Mohammad Zoynul Abedin
Expert Syst. Appl.3
2024 Corporate financial distress prediction using the risk-related information content of annual reports
Petr Hájek 0002, Michal Munk
Inf. Process. Manag.1
2023 Speech emotion recognition and text sentiment analysis for financial distress prediction
abstract
Abstract In recent years, there has been an increasing interest in text sentiment analysis and speech emotion recognition in finance due to their potential to capture the intentions and opinions of corporate stakeholders, such as managers and investors. A considerable performance improvement in forecasting company financial performance was achieved by taking textual sentiment into account. However, far too little attention has been paid to managerial emotional states and their potential contribution to financial distress prediction. This study seeks to address this problem by proposing a deep learning architecture that uniquely combines managerial emotional states extracted using speech emotion recognition with FinBERT-based sentiment analysis of earnings conference call transcripts. Thus, the obtained information is fused with traditional financial indicators to achieve a more accurate prediction of financial distress. The proposed model is validated using 1278 earnings conference calls of the 40 largest US companies. The findings of this study provide evidence on the essential role of managerial emotions in predicting financial distress, even when compared with sentiment indicators obtained from text. The experimental results also demonstrate the high accuracy of the proposed model compared with state-of-the-art prediction models.
Petr Hájek 0002, Michal Munk
Neural Comput. Appl.1
2022 Fuzzy Rule-Based Prediction of Gold Prices using News Affect
abstract
Because of gold’s value, systems for predicting its price have attracted extensive interest in the scientific and industrial communities. Diverse artificial intelligence methods outperform traditional statistical methods in predicting short- and long-term gold price. However, previous research has neglected the transparency of these systems, nor have these systems incorporated the potentially important effect of media sentiment on investment decisions. Therefore, we here propose a fuzzy rule-based prediction system with a component that processes various aspects of news stories. This system is trained on historical data to provide investors with one- and five-days-ahead gold price predictions while achieving a highly interpretable trading strategy in terms of rule complexity. We demonstrate that the proposed system is effective in terms of both prediction accuracy and interpretability compared with state-of-the-art models, such as extreme learning machines and neural networks with deep learning. Our findings suggest that the component of news affect is particularly important for one-day-ahead predictions. We also show that the proposed system performs well in terms of average annual return while providing an interpretable set of linguistic trading rules. This has important implications for investors.
Petr Hájek 0002, Josef Novotny
Expert Syst. Appl.1
2022 Neural intuitionistic fuzzy system with justified granularity
Petr Hájek 0002, Wojciech Froelich, Vladimír Olej, Josef Novotny
Neural Comput. Appl.1
2021 Neural Networks with Emotion Associations, Topic Modeling and Supervised Term Weighting for Sentiment Analysis
abstract
-commerce platform review websites. Deep neural networks outperform traditional lexicon-based and machine learning methods by effectively exploiting contextual word embeddings to generate dense document representation. However, this representation model is not fully adequate to capture topical semantics and the sentiment polarity of words. To overcome these problems, a novel sentiment analysis model is proposed that utilizes richer document representations of word-emotion associations and topic models, which is the main computational novelty of this study. The sentiment analysis model integrates word embeddings with lexicon-based sentiment and emotion indicators, including negations and emoticons, and to further improve its performance, a topic modeling component is utilized together with a bag-of-words model based on a supervised term weighting scheme. The effectiveness of the proposed model is evaluated using large datasets of Amazon product reviews and hotel reviews. Experimental results prove that the proposed document representation is valid for the sentiment analysis of product and hotel reviews, irrespective of their class imbalance. The results also show that the proposed model improves on existing machine learning methods.
Petr Hájek 0002, Aliaksandr Barushka, Michal Munk
Int. J. Neural Syst.1
2020 Combining Rough Set-based Relevance and Redundancy for the Ranking and Selection of Nominal Features
abstract
In this paper, we propose a new method for features ranking and selection. Our approach is based on ranking nominal features in terms of their relevance to the assigned class and mutual redundancy with the other features. To calculate the relevance and redundancy, we propose to use a rough-set based approach. After performing the ranking, features filtering is carried out in a supervised way enabling the user to decide on the number of the retained features. The experiments revealed that thanks to our method, it is possible to filter out numerous features describing data while still maintaining satisfactory classification accuracy achieved by the classifier trained using the reduced dataset. The comparative experiments performed with the use of publicly available datasets proved the high efficiency and competitiveness of our approach.
Wojciech Froelich, Petr Hájek 0002
KES2
2020 Intuitionistic fuzzy grey cognitive maps for forecasting interval-valued time series
Petr Hájek 0002, Wojciech Froelich, Ondrej Prochazka
Neurocomputing1
2020 Spam detection on social networks using cost-sensitive feature selection and ensemble-based regularized deep neural networks
Aliaksandr Barushka, Petr Hájek 0002
Neural Comput. Appl.2
2020 Fake consumer review detection using deep neural networks integrating word embeddings and emotion mining
Petr Hájek 0002, Aliaksandr Barushka, Michal Munk
Neural Comput. Appl.1
2019 IVIFCM-TOPSIS for Bank Credit Risk Assessment
Wojciech Froelich, Petr Hájek 0002
KES-IDT (1)2
2019 Modelling Loss Given Default in Peer-to-Peer Lending Using Random Forests
Monika Papousková, Petr Hájek 0002
KES-IDT (1)2
2019 Two-stage consumer credit risk modelling using heterogeneous ensemble learning
Monika Papousková, Petr Hájek 0002
Decis. Support Syst.2
2019 Integrating TOPSIS with interval-valued intuitionistic fuzzy cognitive maps for effective group decision making
Petr Hájek 0002, Wojciech Froelich
Inf. Sci.1
2018 Interval-Valued Intuitionistic Fuzzy Inference System for Supporting Corporate Financial Decisions
abstract
Representing the inherent uncertainty in the corporate financial environment is critical for effective decision-making in this domain. This is attributed to the increasing complexity of such an environment. One way in which to address this issue is to represent financial attributes in terms of interval-valued intuitionistic fuzzy sets. In this paper, a novel interval-valued intuitionistic fuzzy inference system (IVIFIS) of the Takagi-Sugeno-Kang type is proposed. To calculate the output of the IVIFIS system, a defuzzification method is developed based on the weighted average of the consequents of if-then rules. To adapt the consequent parameters of the IVIFIS, a gradient algorithm is used. Then, by using two regression problems from the corporate financial domain, the dominance of the system over other state-of-the-art extensions of fuzzy inference systems is experimentally shown.
Petr Hájek 0002, Vladimír Olej
FUZZ-IEEE1
2018 Spam filtering using integrated distribution-based balancing approach and regularized deep neural networks
Aliaksandr Barushka, Petr Hájek 0002
Appl. Intell.2
2018 Combining bag-of-words and sentiment features of annual reports to predict abnormal stock returns
Petr Hájek 0002
Neural Comput. Appl.1
2017 Interval-Valued Intuitionistic Fuzzy Cognitive Maps for Supplier Selection
Petr Hájek 0002, Ondrej Prochazka
KES-IDT (1)1
2017 Mining corporate annual reports for intelligent detection of financial statement fraud - A comparative study of machine learning methods
Petr Hájek 0002, Roberto Henriques
Knowl. Based Syst.1
2016 Predicting Abnormal Bank Stock Returns Using Textual Analysis of Annual Reports - a Neural Network Approach
Petr Hájek 0002, Jana Bohácová
EANN1
2016 Interval-valued fuzzy cognitive maps for supporting business decisions
abstract
Fuzzy cognitive maps (FCMs) are used to model uncertainty in complex causal relationships among concepts. This is an important issue for supporting business decisions. However, determining the precise values of these relationships is difficult in the domain of business decision-making. To overcome this problem, we introduce interval-valued FCMs representing a generalization of FCMs. Interval-valued FCMs provide desirable additional freedom in the design of concepts and their causal relationships. This feature enables decision makers to model highly complex business decision-making problems. We demonstrate this for two case studies: (1) a supplier-selection task, and (2) business-performance modelling.
Petr Hájek 0002, Ondrej Prochazka
FUZZ-IEEE1
2015 Intuitionistic Fuzzy Neural Network: The Case of Credit Scoring Using Text Information
Petr Hájek 0002, Vladimír Olej
EANN1
2013 Prediction of Air Quality Indices by Neural Networks and Fuzzy Inference Systems - The Case of Pardubice Microregion
Petr Hájek 0002, Vladimír Olej
EANN (1)1
2013 Evaluating Sentiment in Annual Reports for Financial Distress Prediction Using Neural Networks and Support Vector Machines
Petr Hájek 0002, Vladimír Olej
EANN (2)1
2013 Feature selection in corporate credit rating prediction
Petr Hájek 0002, Krzysztof Michalak
Knowl. Based Syst.1
2011 Municipal credit rating modelling by neural networks
Petr Hájek 0002
Decis. Support Syst.1
2011 Credit rating modelling by kernel-based approaches with supervised and semi-supervised learning
Petr Hájek 0002, Vladimír Olej
Neural Comput. Appl.1
2010 IF-Inference Systems Design for Prediction of Ozone Time Series: The Case of Pardubice Micro-region
Vladimír Olej, Petr Hájek 0002
ICANN (1)2
2009 Municipal Creditworthiness Modelling by Kernel-Based Approaches with Supervised and Semi-supervised Learning
Petr Hájek 0002, Vladimír Olej
EANN1
2009 Municipal Creditworthiness Modelling by Radial Basis Function Neural Networks and Sensitive Analysis of Their Input Parameters
Vladimír Olej, Petr Hájek 0002
ICANN (2)2
2008 Municipal Creditworthiness Modelling by Kohonen's Self-Organizing Feature Maps and Fuzzy Logic Neural Networks
Petr Hájek 0002, Vladimír Olej
ICANN (1)1