VLDB 2026 Research / reviewers in the wild / expert
Sofonias Yitagesu
dblp:219/2227
· DBLP profile ↗
10ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0002-9247-7521ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 6 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Domain-constrained synthesis of inconsistent key aspects in textual vulnerability descriptions
Linyi Han, Shidong Pan, Zhenchang Xing, Sofonias Yitagesu, Xiaowang Zhang, Zhiyong Feng 0002, Jiamou Sun |
Autom. Softw. Eng. | 4 |
| 2026 | Systematic Literature Review on Software Security Vulnerability Information ExtractionabstractBackground . Software vulnerabilities are increasing in complexity and scale, posing great security risks to many software systems. Extracting information about software vulnerabilities is a critical area of research that aims to identify and create a structured representation of vulnerability-related information. These structured data help software systems better understand vulnerabilities and provide security professionals with timely information to mitigate the impact of rapidly growing vulnerabilities while guiding future research to develop more secure systems. However, this process relies on the effectiveness of information extraction to transform manual vulnerability analysis from security experts to digital solutions. Despite its importance, the unique nature of vulnerability information and the fast pace at which machine learning-based extraction methods and techniques have evolved make it challenging to assess the current successes, failures, challenges, and opportunities within this research area. This study presents a systematic literature review aimed at clarifying this complex landscape. Methods . In this study, we conduct a systematic literature review (SLR) to explore existing research focusing on extracting information about software security vulnerabilities. We search for 829 primary studies on security vulnerability information extraction from seven widely used online digital libraries, focusing on top peer-reviewed journals and conferences published between 2001 and 2024. After applying our inclusion and exclusion criteria and the snowballing technique, we narrowed our selection to 87 studies for in-depth analysis and addressed four main research questions. We collect qualitative and quantitative data from each study, identifying 34 components such as research problems, methods, contributions, evaluation metrics, results, types of extracted vulnerability information, challenges, and limitations. We use meta-analysis, statistical machine learning, and text-mining techniques to identify themes, patterns, and trends across the primary studies and visualize findings. Result : The study provides an overview of the security vulnerability data landscape, identifies key resources, and guides efforts to improve vulnerability information extraction and analysis. The study finds a diverse landscape of learning algorithms used in security vulnerability information extraction, with Bidirectional Encoder Representations from Transformers (BERT), Long Short-term Memory (LSTM), and Support Vector Machine (SVM) being the most dominant. The study identifies key challenges, including feature engineering complexity, lack of a gold-standard corpus, preprocessing errors, generating accurate training data, addressing imbalanced data, multimodality fusion, and graph sparsity in security knowledge graphs. Insights for Future Research Directions . The study underscores the need for advanced extraction approaches, robust datasets, automated annotation methods, and advanced machine learning algorithms to improve the extraction of security vulnerability information. This study also suggests using large language models (LLMs) and transformer models to facilitate the automatic extraction of security-related words, terms, concepts, and phrases and introduce new filtering parameters for user requirements. We provide all our implementations; it can be found at https://bitbucket.org/slr-svie/vulnerability-information-extraction/src/master/ . Sofonias Yitagesu, Zhenchang Xing, Xiaowang Zhang, Zhiyong Feng 0002, Tingting Bi, Linyi Han, Xiaohong Li 0001 |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2025 | Leveraging Machine-Translated Data for Sentiment Analysis in Low-Resource Languages: A Case Study on Bengali
Nur-A-Alam Abir, Xiaowang Zhang, Rafiul Haq, Sofonias Yitagesu |
ICANN (3) | 4 |
| 2025 | Towards Roman Urdu Named Entity Recognition: Standardized Dataset and Transformer-Based ApproachesabstractThe recent progress in data and modeling techniques has significantly improved Named Entity Recognition (NER) for well-structured texts and high-resource languages. However, Roman Urdu, written in Latin script, poses unique challenges due to its informal structure, code-switching tendencies, and lack of standardized resources. The existing NER dataset for Roman Urdu suffers from inconsistent annotations and undefined entity boundaries, while traditional deep-learning models struggle with informal and multilingual text. To address these challenges, this study makes several key contributions. First, we re-annotate the Roman Urdu dataset to ensure tagging consistency, adopting the BIOES tagging scheme for entity boundary representation. Second, we fine-tuned multilingual pre-trained language models (PLMs), including mBERT, MuRIL, and XLM-R. Finally, we introduce XLM-R-IDCNN-BiLSTM-H, a novel architecture that combines transformer-based embeddings with advanced neural components, including Iterated Dilated Convolutional Neural Networks (IDCNN) for local dependencies, Bidirectional Long Short-Term Memory (BiLSTM) networks for bidirectional context, and highway networks for smooth information flow. We employ a conditional random field (CRF) to obtain the optimal tag sequence for classification. According to experimental findings, the proposed approach delivers state-of-the-art performance, which yielded an F1 score of 84.26% on the re-annotated Roman Urdu dataset. Rafiul Haq, Xiaowang Zhang, Sofonias Yitagesu, Wahab Khan, Zhiyong Feng 0002 |
IJCNN | 3 |
| 2025 | DeepSarc: A Transformer-Based Deep Learning Approach for Sarcasm Detection in Social Media TextabstractABSTRACT Sarcasm detection is a critical and challenging task in sentiment analysis, particularly for low‐resource languages like Urdu, where limited annotated data, linguistic complexity, and subtle contextual cues hinder accurate classification. Traditional machine learning methods often fail to capture the nuanced and often contradictory nature of sarcastic expression. To address these challenges, this paper presents a comprehensive and computationally efficient framework for Urdu sarcasm detection. We first mitigate severe class imbalance through strategic down‐sampling and back‐translation‐based data augmentation. We then conduct extensive benchmarking of traditional deep learning architectures against fine‐tuned pre‐trained language models, including multilingual, monolingual, and Twitter‐specific variants. Building on these insights, we propose DeepSarc, a novel hybrid model that integrates the contextual embeddings from XLM‐T, the multi‐scale feature extraction capabilities of dilated convolutional neural networks (DCNNs), and the sequential dependency modeling of bidirectional long short‐term memory (BiLSTM) networks. While slightly more computationally intensive than simpler alternatives, DeepSarc achieves a state‐of‐the‐art F1‐score, significantly outperforming existing approaches. Our results establish a new benchmark for sarcasm detection in low‐resource languages and provide a scalable, high‐performance framework adaptable to diverse linguistic contexts. Rafiul Haq, Xiaowang Zhang, Sofonias Yitagesu, Wahab Khan, Zhiyong Feng 0002 |
Concurr. Comput. Pract. Exp. | 3 |
| 2025 | Do Chase Your Tail! Missing Key Aspects Augmentation in Textual Vulnerability Descriptions of Long-Tail Software Through Feature InferenceabstractAugmenting missing key aspects in Textual Vulnerability Descriptions (TVDs) is crucial for effective vulnerability analysis. For instance, in TVDs, key aspects includeAttack Vector,Vulnerability Type, among others. These key aspects help security engineers understand and address the vulnerability in a timely manner. For software with a large user base (non-long-tail software), augmenting these missing key aspects has significantly advanced vulnerability analysis and software security research. However, software instances with a limited user base (long-tail software) often get overlooked due to inconsistency software names, TVD limited avaliability, and domain-specific jargon, which complicates vulnerability analysis and software repairs. In this paper, we introduce a novel software feature inference framework designed to augment the missing key aspects of TVDs for long-tail software. Firstly, we tackle the issue of non-standard software names found in community-maintained vulnerability databases by cross-referencing government databases with Common Vulnerabilities and Exposures (CVEs). Next, we employ Large Language Models (LLMs) to generate the missing key aspects. However, the limited availability of historical TVDs restricts the variety of examples. To overcome this limitation, we utilize the Common Weakness Enumeration (CWE) to classify all TVDs and select cluster centers as representative examples. To ensure accuracy, we present Natural Language Inference (NLI) models specifically designed for long-tail software. These models identify and eliminate incorrect responses. Additionally, we use a wiki repository to provide explanations for proprietary terms. Our evaluations demonstrate that our approach significantly improves the accuracy of augmenting missing key aspects of TVDs for log-tail software from 0.27 to 0.56 (+107%). Interestingly, the accuracy of non-long-tail software also increases from 64% to 71%. As a result, our approach can be useful in various downstream tasks that require complete TVD information. Linyi Han, Shidong Pan, Zhenchang Xing, Jiamou Sun, Sofonias Yitagesu, Xiaowang Zhang, Zhiyong Feng 0002 |
IEEE Trans. Software Eng. | 5 |
| 2023 | Extraction of Phrase-based Concepts in Vulnerability Descriptions through Unsupervised LabelingabstractSoftware vulnerabilities, once disclosed, can be documented in vulnerability databases, which have great potential to advance vulnerability analysis and security research. People describe the key characteristics of software vulnerabilities in natural language mixed with domain-specific names and concepts. This textual nature poses a significant challenge for the automatic analysis of vulnerability knowledge embedded in text. Automatic extraction of key vulnerability aspects is highly desirable but demands significant effort to manually label data for model training. In this article, we propose unsupervised methods to label and extract important vulnerability concepts in textual vulnerability descriptions (TVDs). We focus on six types of phrase-based vulnerability concepts (vulnerability type, vulnerable component, root cause, attacker type, impact, and attack vector) as they are much more difficult to label and extract than name- or number-based entities (i.e., vendor, product, and version). Our approach is based on a key observation that the same-type of phrases, no matter how they differ in sentence structures and phrase expressions, usually share syntactically similar paths in the sentence parsing trees. Specifically, we present a source-target neural architecture that learns the Part-of-Speech (POS) tagging to identify a token’s functional role within TVDs, where the source neural model is trained to capture common features found in the TVD corpus, and the target model is trained to identify linguistically malformed words specific to the security domain. Our evaluation confirms that the proposed tagger outperforms (4.45%–5.98%) the taggers designed on natural language notions and identifies a broad set of TVDs and natural language contents. Then, based on the key observations, we propose two path representations (absolute paths and relative paths) and use an auto-encoder to encode such syntactic similarities. To address the discrete nature of our paths, we enhance the traditional Variational Auto-encoder (VAE) with Gumble-Max trick for categorical data distribution and thus create a Categorical VAE (CaVAE). In the latent space of absolute and relative paths, we further apply unsupervised clustering techniques to generate clusters of the same-type of concepts. Our evaluation confirms the effectiveness of our CaVAE, which achieves a small (85.85) log-likelihood for encoding path representations and the accuracy (83%–89%) of vulnerability concepts in the resulting clusters. The resulting clusters accurately label six types of vulnerability concepts from a TVD corpus in an unsupervised way. Furthermore, these labeled vulnerability concepts can be mapped back to the corresponding phrases in the original TVDs, which produce labels of six types of vulnerability concepts. The resulting labeled TVDs can be used to train concept extraction models for other TVD corpora. In this work, we present two concept extraction methods (concept classification and sequence labeling model) to demonstrate the utility of the unsupervisedly labeled concepts. Our study shows that models trained with our unsupervisedly labeled vulnerability concepts outperform (3.9%–5.14%) those trained with the two manually labeled TVD datasets from previous work due to the consistent boundary and typing by our unsupervised labeling method. Sofonias Yitagesu, Zhenchang Xing, Xiaowang Zhang, Zhiyong Feng 0002, Xiaohong Li 0001, Linyi Han |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2023 | Phrase-level attention network for few-shot inverse relation classification in knowledge graph
Shaojuan Wu, Chunliu Dou, Dazhuang Wang, Jitong Li, Xiaowang Zhang, Zhiyong Feng 0002, Kewen Wang 0001, Sofonias Yitagesu |
World Wide Web (WWW) | 8 |
| 2021 | Unsupervised Labeling and Extraction of Phrase-based Concepts in Vulnerability DescriptionsabstractPeople usually describe the key characteristics of software vulnerabilities in natural language mixed with domain-specific names and concepts. This textual nature poses a significant challenge for the automatic analysis of vulnerabilities. Automatic extraction of key vulnerability aspects is highly desirable but demands significant effort to manually label data for model training. In this paper, we propose an unsupervised approach to label and extract important vulnerability concepts in textural vulnerability descriptions (TVDs). We focus on three types of phrase-based vulnerability concepts (root cause, attack vector, and impact) as they are much more difficult to label and extract than name- or number-based entities (i.e., vendor, product, and version). Our approach is based on a key observation that the same-type of phrases, no matter how they differ in sentence structures and phrase expressions, usually share syntactically similar paths in the sentence parsing trees. Therefore, we propose two path representations (absolute paths and relative paths) and use an auto-encoder to encode such syntactic similarities. To address the discrete nature of our paths, we enhance traditional Variational Auto-encoder (VAE) with Gumble-Max trick for categorical data distribution, and thus creates a Categorical VAE (CaVAE). In the latent space of absolute and relative paths, we further use FIt-TSNE and clustering techniques to generate clusters of the same-type of concepts. Our evaluation confirms the effectiveness of our CaVAE for encoding path representations and the accuracy of vulnerability concepts in the resulting clusters. In a concept classification task, our unsupervisedly labeled vulnerability concepts outperform the two manually labeled datasets from previous work. Sofonias Yitagesu, Zhenchang Xing, Xiaowang Zhang, Zhiyong Feng 0002, Xiaohong Li 0001, Linyi Han |
ASE | 1 |
| 2021 | Automatic Part-of-Speech Tagging for Security Vulnerability DescriptionsabstractIn this paper, we study the problem of part-of-speech (POS) tagging for security vulnerability descriptions (SVD). In contrast to newswire articles, SVD often contains a high-level natural language description of the text composed of mixed language studded with codes, domain-specific jargon, vague language, and abbreviations. Moreover, training data dedicated to security vulnerability research is not widely available. Existing neural network-based POS tagging has often relied on manually annotated training data or applying natural language processing (NLP) techniques, suffering from two significant drawbacks. The former is extremely time-consuming and requires labor-intensive feature engineering and expertise. The latter is inadequate to identify linguistically-informed words specific to the SVD domain. In this paper, we propose an automatic approach to assign POS tags to tokens in SVD. Our approach uses the character-level representation to automatically extract orthographic features and unsupervised word embeddings to capture meaningful syntactic and semantic regularities from SVD. The character level representations are then concatenated with the word embedding as a combined feature, which is then learned and used to predict the POS tagging. To deal with the issue of the poor availability of annotated security vulnerability data, we implement a finetuning approach. Our approach provides public access to a POS annotated corpus of ~8M tokens, which serves as a training dataset in this domain. Our evaluation results show a significant improvement in accuracy (17.72%-28.22%) of POS tagging in SVD over the current approaches. Sofonias Yitagesu, Xiaowang Zhang, Zhiyong Feng 0002, Xiaohong Li 0001, Zhenchang Xing |
MSR | 1 |