Hossein Keshavarz

dblp:123/5684 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
3since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2024 JITGNN: A deep graph neural network framework for Just-In-Time bug prediction
abstract
Just-In-Time (JIT) bug prediction is the problem of predicting software failure immediately after a change is submitted to the code base. JIT bug prediction is often preferred to other types of bug prediction (subsystem, module, file, class, or function-level) because changes are associated with one developer and the predictions can be applied when the design decisions are fresh in the developer’s mind. Many approaches have been proposed to predict correctly whether a software change is bug-inducing. These approaches mainly rely on change metrics such as the size, the number of modified files, and the developer’s experience. Although there has been extensive work on employing deep learning models for other forms of bug prediction, there are few deep models for JIT bug prediction. Furthermore, none of the existing JIT models that use the changed source code consider the graph structure of source codes. In this paper, we propose a JIT model that incorporates both the content and metadata of changes leveraging the graph structure of programs. We designed and built JITGNN, a deep graph neural network (GNN) framework for JIT bug prediction. JITGNN uses the abstract syntax trees (ASTs) of changed programs. We evaluate the performance of JITGNN on two datasets and compare it to a baseline and state-of-the-art JIT model. We hypothesize that by including the graph structure of source codes in the JIT bug prediction process, we can improve the performance of the JIT models. Our study, however, shows that JITGNN achieves the same AUC as the state-of-the-art model (JITLine), and they both have the same discriminatory power.
Hossein Keshavarz, Gema Rodríguez-Pérez
J. Syst. Softw.1
2022 Named Entity Recognition in Long Documents: An End-to-end Case Study in the Legal Domain
abstract
Named entity recognition (NER) is a fundamental task for several important applications such as knowledge base construction and semantic search. So far, the focus has been on building machine learning models, which identify generic named entities (e.g., person, date). Such models can be used off-the-shelf without requiring ground truth labels for training. However, such models cannot generalize to specialized domains that have domain-specific named entities (e.g., the legal domain). In these cases, it is inevitable to generate ground truth data and experiment with a variety of models in order to achieve good performance. Motivated by a real use case from the financial sector, we discuss the approach and lessons learned when solving the NER problem in the legal domain. This task is particularly challenging because it requires extensive human expertise to produce high quality ground-truth labels. For solving the legal domain NER problem, we first crawl a large dataset of legal documents and then introduce a semi-automated process to generate high-quality labels for a set of eleven predefined named entities. We validate that the proposed approach achieves high quality labels that outperform popular out-of-the-box NER methods. On top of that, our method once followed, can generate ground truth labels for the pre-defined named entities for an unbounded number of documents. Next, we experiment with a set of models and training procedures and report their performance on the NER task. Our experimental evaluation confirms that most of the models can generalize very well, achieving F1-score between 86% and 98.9%. The dataset, the labels produced by human annotators and our semi-supervised approach, as well as our code are made available to the research community.
Hossein Keshavarz, Zografoula Vagena, Pigi Kouki, Ilias Fountalis, Mehdi Mabrouki, Aziz Belaweid, Nikolaos Vasiloglou
IEEE Big Data1
2022 ApacheJIT: A Large Dataset for Just-In-Time Defect Prediction
abstract
In this paper, we present ApacheJIT, a large dataset for Just-In-Time (JIT) defect prediction. ApacheJIT consists of clean and bug-inducing software changes in 14 popular Apache projects. ApacheJIT has a total of 106,674 commits (28,239 bug-inducing and 78,435 clean commits). Having a large number of commits makes ApacheJIT a suitable dataset for machine learning JIT models, especially deep learning models that require large training sets to effectively generalize the patterns present in the historical data to future data.
Hossein Keshavarz, Meiyappan Nagappan
MSR1
2020 Sequential change-point detection in high-dimensional Gaussian graphical models
abstract
High dimensional piecewise stationary graphical models represent a versatile class for modelling time varying networks arising in diverse application areas, including biology, economics, and social sciences. There has been recent work in offline detection and estimation of regime changes in the topology of sparse graphical models. However, the online setting remains largely unexplored, despite its high relevance to applications in sensor networks and other engineering monitoring systems, as well as financial markets. To that end, this work introduces a novel scalable online algorithm for detecting an unknown number of abrupt changes in the inverse covariance matrix of sparse Gaussian graphical models with small delay. The proposed algorithm is based upon monitoring the conditional log-likelihood of all nodes in the network and can be extended to a large class of continuous and discrete graphical models. We also investigate asymptotic properties of our procedure under certain mild regularity conditions on the graph size, sparsity level, number of samples, and pre- and post-changes in the topology of the network. Numerical works on both synthetic and real data illustrate the good performance of the proposed methodology both in terms of computational and statistical efficiency across numerous experimental settings.
Hossein Keshavarz, George Michailidis, Yves F. Atchadé
J. Mach. Learn. Res.1
2012 Impact of Cognition and Cooperation on MAC Layer Performance Metrics, Part I: Maximum Stable Throughput
abstract
In this paper, we consider a broadband secondary transmitter-receiver pair which interferes with N narrowband primary users and study the effect of cognition and cooperation on the maximum stable throughput. In our study we focus on four transmission protocols as well as two channel types, i.e., flat fading and frequency selective fading. In the cooperative protocols, the broadband transmitter relays the packets of the primary users which have not correctly decoded at the primary receiver. The analysis includes random packet arrivals at the transmitters which may impact on the maximum stable throughput. Moreover, sensing errors at the secondary user are considered. In this paper, we derive the exact stability region for non-cooperative protocols and inner bounds for cooperative ones. The results reveal that depending on the channel states, the cooperative protocols may provide significant performance gains over non-cooperative protocols. Numerical results indicate that in the frequency selective channel, independent and simultaneous parallel transmissions in the available subbands are preferred, while in the flat fading channel, a non-parallel transmission in the total available bandwidth is recommended. In part II of our paper, we investigate the impact of cognition and cooperation on the delay performance of the protocols introduced in this paper.
Ali Asghar Sharifi, Farid Ashtiani, Hossein Keshavarz, Masoumeh Nasiri-Kenari
IEEE Trans. Wirel. Commun.3