Michael Bewong

dblp:198/8680 · DBLP profile ↗
← Back
15ranked-venue papers in the field
2as first author
13since 2021 · last 2026
0000-0002-5848-7451ORCID · verified

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 6Database Systems & Data Management · 3 (1 first)Data Mining & Knowledge Discovery · 3 (1 first)Knowledge Engineering, Semantic Web & Information Systems · 2Big Data, Cloud & Distributed Data Systems · 1
YearPublicationVenuePosition
2026 RED: Rule Guided Prompt Engineering for Graph Data Imputation
Xinyao Huang, Jiang Hua, Michael Bewong, Selasi Kwashie, Zaiwen Feng
PAKDD (4)3
2026 A Unified and Time-Efficient Multi-Agent Framework for Data Discovery
Yunhao Xiao, Michael Bewong, Selasi Kwashie, Zaiwen Feng
WWW3
2025 Social Engineering Attacks: A Systemisation of Knowledge on People Against Humans
Scott Thomson, Michael Bewong, Arash Mahboubi, Tanveer A. Zia
IEEE Big Data2
2025 RAE: A Rule-Driven Approach for Attribute Embedding in Property Graph Recommendation
Sibo Zhao, Michael Bewong, Selasi Kwashie, Zaiwen Feng
ECML/PKDD (6)2
2025 FastER: On-demand Entity Resolution in Property Graphs
Shujing Wang 0013, Sibo Zhao, Shiqi Miao, Selasi Kwashie, Michael Bewong, Vincent Mwintieru Nofong, Zaiwen Feng
ISWC (1)5
2025 When GDD meets GNN: A knowledge-driven neural connection for effective entity resolution in property graphs
abstract
This paper studies the entity resolution (ER) problem in property graphs. ER is the task of identifying and linking different records that refer to the same real-world entity. It is commonly used in data integration, data cleansing, and other applications where it is important to have accurate and consistent data. In general, two predominant approaches exist in the literature: rule-based and learning-based methods. On the one hand, rule-based techniques are often desired due to their explainability and ability to encode domain knowledge. Learning-based methods, on the other hand, are preferred due to their effectiveness in spite of their black-box nature. In this work, we devise a hybrid ER solution, GraphER , that leverages the strengths of both systems for property graphs. In particular, we adopt graph differential dependency (GDD) for encoding the so-called record-matching rules , and employ them to guide a graph neural network (GNN) based representation learning for the task. We conduct extensive empirical evaluation of our proposal on benchmark ER datasets including 17 graph datasets and 7 relational datasets in comparison with 10 state-of-the-art (SOTA) techniques. The results show that our approach provides a significantly better solution to addressing ER in graph data, both quantitatively and qualitatively, while attaining highly competitive results on the benchmark relational datasets w.r.t. the SOTA solutions.
Michael Bewong, Selasi Kwashie, Vincent Mwintieru Nofong, John Wondoh, Zaiwen Feng
Inf. Syst.2
2024 MAPX: An Explainable Model-Agnostic Framework for Detecting False Information on Social Media Networks
Sarah Condran, Michael Bewong, Selasi Kwashie, Md Zahidul Islam 0001, Irfan Altas, Joshua Condran
WISE (2)2
2024 A Graph-Based Approach for Software Functionality Classification on the Web
Yinhao Jiang, Michael Bewong, Arash Mahboubi, Sajal Halder, Md. Rafiqul Islam 0001, Md Zahidul Islam 0001, Ryan H. L. Ip, Praveen Gauravaram, Minhui Xue 0001
WISE (5)2
2024 FUD-LDP: Fully User Driven Local Differential Privacy
Gnana Thedchanamoorthy, Michael Bewong, Meisam Mohammady, Tanveer A. Zia, Md Zahidul Islam 0001
WISE (5)2
2024 Malicious Package Detection using Metadata Information
abstract
Protecting software supply chains from malicious packages is paramount in the evolving landscape of software development. Attacks on the software supply chain involve attackers injecting harmful software into commonly used packages or libraries in a software repository. For instance, JavaScript uses Node Package Manager (NPM), and Python uses Python Package Index (PyPi) as their respective package repositories. In the past, NPM has had vulnerabilities such as the event-stream incident, where a malicious package was introduced into a popular NPM package, potentially impacting a wide range of projects. As the integration of third-party packages becomes increasingly ubiquitous in modern software development, accelerating the creation and deployment of applications, the need for a robust detection mechanism has become critical. On the other hand, due to the sheer volume of new packages being released daily, the task of identifying malicious packages presents a significant challenge. To address this issue, in this paper, we introduce a metadata-based malicious package detection model, MeMPtec. This model extracts a set of features from package metadata information. These extracted features are classified as either easy-to-manipulate (ETM) or difficult-to-manipulate (DTM) features based on monotonicity and restricted control properties. By utilising these metadata features, not only do we improve the effectiveness of detecting malicious packages, but also we demonstrate its resistance to adversarial attacks in comparison with existing state-of-the-art. Our experiments indicate a significant reduction in both false positives (up to 97.56%) and false negatives (up to 91.86%).
Sajal Halder, Michael Bewong, Arash Mahboubi, Yinhao Jiang, Md. Rafiqul Islam 0001, Md Zahidul Islam 0001, Ryan H. L. Ip, M. Ejaz Ahmed, Gowri Sankar Ramachandran, Muhammad Ali Babar 0001
WWW2
2024 An efficient approach for discovering Graph Entity Dependencies (GEDs)
abstract
Graph entity dependencies (GEDs) are novel graph constraints, unifying keys and functional dependencies, for property graphs. They have been found useful in many real-world data quality and data management tasks, including fact checking on social media networks and entity resolution. In this paper, we study the discovery problem of GEDs—finding a minimal cover of valid GEDs in a given graph data. We formalise the problem, and propose an effective and efficient approach to overcome major bottlenecks in GED discovery. In particular, we leverage existing graph partitioning algorithms to enable fast GED-scope discovery, and employ effective pruning strategies over the prohibitively large space of candidate dependencies. Furthermore, we define an interestingness measure for GEDs based on the minimum description length principle, to score and rank the mined cover set of GEDs. Finally, we demonstrate the scalability and effectiveness of our GED discovery approach through extensive experiments on real-world benchmark graph data sets; and present the usefulness of the discovered rules in different downstream data quality management applications.
Dehua Liu, Selasi Kwashie, Guangtong Zhou, Michael Bewong, Keqing He 0002, Zaiwen Feng
Inf. Syst.5
2023 FastAGEDs: Fast Approximate Graph Entity Dependency Discovery
Guangtong Zhou, Selasi Kwashie, Michael Bewong, Vincent Mwintieru Nofong, Debo Cheng, Keqing He 0002, Shanmei Liu, Zaiwen Feng
WISE4
2021 BDF: A new decision forest algorithm
Md. Nasim Adnan, Ryan H. L. Ip, Michael Bewong, Md Zahidul Islam 0001
Inf. Sci.3
2019 Privacy preserving serial publication of transactional data
Michael Bewong, Jixue Liu, Lin Liu 0003, Jiuyong Li
Inf. Syst.1
2017 Utility Aware Clustering for Publishing Transactional Data
Michael Bewong, Jixue Liu, Lin Liu 0003, Jiuyong Li
PAKDD (2)1