Muhammad Syafiq Mohd Pozi

dblp:132/1763 · also Muhammad Syafiq Bin Mohd Pozi · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
4since 2021 · last 2025
0000-0001-9379-7351ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 A data-augmented model routing framework for efficient LLM deployment in edge-cloud environments
abstract
Abstract Large language model (LLM)-based program generation tasks are hindered by high computational demands. These challenges, along with high deployment costs, often pose a barrier to practical applications. To address these, we propose a novel data-augmented multi-LLM model routing approach that classifies prompts based on whether they should be processed on a weak LLM engine or a strong LLM. Experimental results show up to 16 times better efficiency compared to the existing cascaded approaches, while preserving the inference accuracy. Thus, the proposed method optimally allocates prompts across multiple LLMs, reducing computational costs while maintaining inference accuracy.
Muhammad Syafiq Mohd Pozi, Yukinori Sato
J. Supercomput.1
2024 An Investigation of SMOTE Based Methods for Imbalanced Datasets with Data Complexity Analysis (Extended Abstract)
abstract
This extended abstract highlights challenges with imbalanced datasets in real-world applications, where issues like noise, class overlap, and small subsets of data impact classification accuracy. While the Synthetic Minority Oversampling Technique (SMOTE) addresses imbalanced datasets by increasing minority class examples, it struggles with handling these data complexities and might worsen the situation. As a result, several SMOTE variants have emerged, aiming to improve its effectiveness by integrating it with other methods or altering its approach. This paper offers a comparative analysis of these variants, examining how each tackles specific data complexities. Through experiments on 24 imbalanced datasets, changes in complexity measures resulting from these SMOTE variants, in terms of F1-Score and data complexity metrics are observed and demonstrated.
Nur Athirah Azhar, Muhammad Syafiq Mohd Pozi, Aniza Mohamed Din, Adam Jatowt
ICDE2
2024 Using Causal Inference to Solve Uncertainty Issues in Dataset Shift
abstract
Dataset shift will lead to uncertainty issues, and then the models will accurately decline. Using causality instead of correlation to find the invariant characteristic and solve the uncertainty issues between different dataset distributions (eg. Domain Adaptation). Summarizing datasets can be used in current domain training, building a benchmarking framework of causal learning that combines the causal inference and traditional model to detect, address, and determine the characteristic of the dataset shift.
Song Shuang, Muhammad Syafiq Mohd Pozi
WSDM2
2023 An Investigation of SMOTE Based Methods for Imbalanced Datasets With Data Complexity Analysis
abstract
Many binary class datasets in real-life applications are affected by class imbalance problem. Data complexities like noise examples, class overlap and small disjuncts problems are observed to play a key role in producing poor classification performance. These complexities tend to exist in tandem with class imbalance problem. Synthetic Minority Oversampling Technique (SMOTE) is a well-known method to re-balance the number of examples in imbalanced datasets. However, this technique cannot effectively tackle data complexities and it also has the capability of magnifying the degree of complexities. Also, the performance of the SMOTE is still not satisfactory. Therefore, various SMOTE variants have been proposed to overcome the downsides of SMOTE either by combining SMOTE with other algorithms or modifying the existing SMOTE algorithm. This paper aims to comparatively review the algorithms applied in SMOTE variants and investigate which data complexities are being addressed in what variants. Series of experiments are conducted on 24 binary class imbalanced datasets to observe the changes in the data complexity measures after SMOTE variants were applied in these datasets. The evaluation metrics like G-Mean and F1-Score are also analyzed to investigate the difference in classification performance between SMOTE variants.
Nur Athirah Azhar, Muhammad Syafiq Mohd Pozi, Aniza Mohamed Din, Adam Jatowt
IEEE Trans. Knowl. Data Eng.2
2017 Visualization of Spatio-Temporal Events in Geo-Tagged Social Media
Yuanyuan Wang 0003, Muhammad Syafiq Mohd Pozi, Gouki Yasui, Yukiko Kawai, Kazutoshi Sumiya, Toyokazu Akiyama
W2GIS2
2016 Improving Anomalous Rare Attack Detection Rate for Intrusion Detection System Using Support Vector Machine and Genetic Programming
Muhammad Syafiq Mohd Pozi, Md Nasir Sulaiman, Norwati Mustapha, Thinagaran Perumal
Neural Process. Lett.1