Muhammad Asif Naeem

dblp:82/8568 · DBLP profile ↗
← Back
31ranked-venue papers
12as first author
9since 2021 · last 2025
0000-0001-6785-7875ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 18 · 11 first-author · 4 since 2021Artificial intelligence and machine learning · 17 · 6 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 since 2021
YearPublicationVenuePosition
2025 CoverGAN: cover photo generation from text story using layout guided GAN
Adeel Cheema, Muhammad Asif Naeem
Soft Comput.2
2024 Dynamic Quantification With Constrained Error Under Unknown General Dataset Shift
abstract
Quantification research has sought to accurately estimate class distributions under dataset shift. While existing methods perform well under assumed conditions of shift, it is not always clear whether such assumptions will hold in a given application. This work extends the analysis and experimental evaluation of our Gain-Some-Lose-Some (GSLS) model for quantification under general dataset shift and incorporates it into a method for dynamically selecting the most appropriate quantification method. Selection by a Kolmogorov-Smirnov test for any shift followed by a newly proposed “Adjusted Kolmogorov-Smirnov” test for non-prior shift is found to best balance quantification and runtime performance. We also present a framework for constraining quantification prediction intervals to user-specified limits by requesting a smaller set of instance class labels from the user than required with confidence-based rejection.
Benjamin Denham, Edmund M.-K. Lai, Roopak Sinha, Muhammad Asif Naeem
IEEE Trans. Knowl. Data Eng.4
2023 Predicting Peak Demand Days for Asthma-Related Emergency Hospitalisations: A Machine Learning Approach
Rashi Bhalla, Farhaan Mirza, Muhammad Asif Naeem, Amy Hai Yan Chan
PKAW3
2022 Novel method for optimizing performance in resource constrained distributed data streams
Rashi Bhalla, Russel Pears, Muhammad Asif Naeem, Farhaan Mirza
Appl. Intell.3
2022 Witan: Unsupervised Labelling Function Generation for Assisted Data Programming
abstract
Effective supervised training of modern machine learning models often requires large labelled training datasets, which could be prohibitively costly to acquire for many practical applications. Research addressing this problem has sought ways to leverage weak supervision sources, such as the user-defined heuristic labelling functions used in the data programming paradigm, which are cheaper and easier to acquire. Automatic generation of these functions can make data programming even more efficient and effective. However, existing approaches rely on initial supervision in the form of small labelled datasets or interactive user feedback. In this paper, we propose Witan, an algorithm for generating labelling functions without any initial supervision. This flexibility affords many interaction modes, including unsupervised dataset exploration before the user even defines a set of classes. Experiments in binary and multi-class classification demonstrate the efficiency and classification accuracy of Witan compared to alternative labelling approaches.
Benjamin Denham, Edmund M.-K. Lai, Roopak Sinha, Muhammad Asif Naeem
Proc. VLDB Endow.4
2022 TinyLFU-based semi-stream cache join for near-real-time data warehousing
Muhammad Asif Naeem, Wasiullah Waqar, Farhaan Mirza, Ali Tahir
Soft Comput.1
2021 Gain-Some-Lose-Some: Reliable Quantification Under General Dataset Shift
abstract
When applying supervised learning to estimate class distributions of unlabelled samples (so-called quantification), dataset shift is an expected yet challenging problem. Existing quantification methods make strong assumptions on the nature of dataset shift that often will not hold in practice. We propose a novel Gain-Some-Lose-Some (GSLS) model that accounts for more general conditions of dataset shift. We present a method for fitting the GSLS model without any labelled instances from the target sample, and experimentally demonstrate that GSLS can produce reliable quantification prediction intervals under broader conditions of shift than existing quantification methods.
Benjamin Denham, Edmund M.-K. Lai, Roopak Sinha, Muhammad Asif Naeem
ICDM4
2021 Topical affinity in short text microblogs
Herman Masindano Wandabwa, Muhammad Asif Naeem, Farhaan Mirza, Russel Pears
Inf. Syst.2
2021 Multi-interest semantic changes over time in short-text microblogs
Herman Masindano Wandabwa, Muhammad Asif Naeem, Farhaan Mirza, Russel Pears
Knowl. Based Syst.2
2020 Null-Labelling: A Generic Approach for Learning in the Presence of Class Noise
abstract
Class noise in datasets presents a significant challenge to accurate classification, requiring classifiers that can refuse to classify noisy instances. We demonstrate the inability of the popular confidence-thresholding rejection method to learn from relationships between input features and not-at-random class noise. To take advantage of these relationships, we propose a novel null-labelling scheme based on iterative re-training with relabelled datasets that enables a classifier to learn to reject instances that are likely to be misclassified. We demonstrate the ability of null-labelling to achieve a significantly better tradeoff between classification error and coverage than confidence-thresholding. Models generated by the null-labelling scheme have the added advantage of interpretability, in that they are able to identify features correlated with class noise. We also unify prior theories for combining and evaluating sets of rejecting classifiers.
Benjamin Denham, Russel Pears, Muhammad Asif Naeem
ICDM3
2020 Enhancing random projection with independent and cumulative additive noise for privacy-preserving data stream mining
Benjamin Denham, Russel Pears, Muhammad Asif Naeem
Expert Syst. Appl.3
2020 Optimizing Semi-Stream CACHEJOIN for Near-Real- Time Data Warehousing
abstract
Streaming data join is a critical process in the field of near-real-time data warehousing. For this purpose, an adaptive semi-stream join algorithm called CACHEJOIN (Cache Join) focusing non-uniform stream data is provided in the literature. However, this algorithm cannot exploit the memory and CPU resources optimally and consequently it leaves its service rate suboptimal due to sequential execution of both of its phases, called stream-probing (SP) phase and disk-probing (DP) phase. By integrating the advantages of CACHEJOIN, this article presents two modifications for it. The first is called P-CACHEJOIN (Parallel Cache Join) that enables the parallel processing of two phases in CACHEJOIN. This increases number of joined stream records and therefore improves throughput considerably. The second is called OP-CACHEJOIN (Optimized Parallel Cache Join) that implements a parallel loading of stored data into memory while the DP phase is executing. This research presents the performance analysis of both of the approaches defined within the paper existing CACHEJOIN empirically using synthetic skewed dataset.
Muhammad Asif Naeem, Erum Mehmood, Muhammad Ghulam Abbas Malik, Noreen Jamil
J. Database Manag.1
2020 HDSM: A distributed data mining approach to classifying vertically distributed data streams
Benjamin Denham, Russel Pears, Muhammad Asif Naeem
Knowl. Based Syst.3
2019 A memory-optimal many-to-many semi-stream join
Muhammad Asif Naeem, Gerald Weber, Christof Lutteroth
Distributed Parallel Databases1
2018 The incremental Fourier classifier: Leveraging the discrete Fourier transform for classifying high speed data streams
Chamari I. Kithulgoda, Russel Pears, Muhammad Asif Naeem
Expert Syst. Appl.3
2017 Uncovering useful patterns in shopping cart data
abstract
Understanding the shopping and purchasing behaviours of customers is an essential task for business and retail organizations. While customers look for useful information from retailers as they shop, businesses seek to collect increasing amounts of data in order to deliver added value to their customers. This requires an intensive analysis of sales data. Extracting shopping patterns across the many levels of information is a non-trivial task as datasets on sales transactions can contain many levels of information such as item category, brand name, colour, and price. This paper examines the use of multi-level association rules to uncover purchasing patterns at multiple levels of detail. It shows how different kinds of purchasing patterns can emerge at different association levels of analysis. This type of analysis is indeed helpful in assisting retailers to make wise decisions for their customers.
Ali Haider Hussein Ghazala, Muhammad Asif Naeem, Farhaan Mirza, Noreen Jamil
CSCWD2
2017 A review on IoT healthcare monitoring applications and a vision for transforming sensor data into real-time clinical feedback
abstract
Ageing populations and the increase in chronic diseases all over the world demand efficient healthcare solutions for maintaining well-being of people. One strategy that has drawn significant research attention is a focus on remote health monitoring systems based on Internet of Things (IoT) technology. This concept can help decrease pressure on hospital systems and healthcare providers, reduce healthcare costs, and improve homecare especially for patients with chronic diseases and the elderly. This paper explores the use of IoT-based applications in medical field and proposes an IoT Tiered Architecture (IoTTA) towards an approach for transforming sensor data into real-time clinical feedback. This approach considers a range of aspects including sensing, sending, processing, storing, and mining and learning. Using this approach will help to develop useful and effective solutions for pursuing systems development in IoT healthcare applications. The result of the review found that the growth of IoT applications for healthcare is in areas of self-care, data mining, and machine learning.
Hoa Hong Nguyen, Farhaan Mirza, Muhammad Asif Naeem
CSCWD3
2017 Document level semantic comprehension of noisy text streams via convolutional neural networks
abstract
Content comprehension in text is one of the challenges in natural language processing. Understanding text at a low level has become increasingly relevant due to the surge in the amount of content on the web space, where most of it is stream data. In our case, data streams are considered to be an ordered sequence of short and noisy textual messages that are read once or fewer number of times, for example tweets. Our approach entails processing and interpreting streaming texts at document level in mini-batches via deep convolutional networks for opinion, semantic or relationship analysis. Training our model is iterative and incremental, where documents are learnt by understanding the sentence structure and content in vector form based on a known offline model. The model however, incrementally adapts to the changing textual patterns. Our conceptual framework design is distributed in nature such that a pipeline of inputs, deep processing framework and output will be coordinated by the Apache Storm framework. This design is applicable in real time sentiment analysis, opinion mining, and stance detection among others.
Herman Masindano Wandabwa, Muhammad Asif Naeem, Farhaan Mirza
CSCWD2
2017 Detecting Falls Using a Wearable Accelerometer Motion Sensor
abstract
This research aims to early detect falls based on the rapid acceleration changes using the threshold based approach, using a single accelerometer. We propose the Acceleration Change-based Falls Detection Algorithm (ACFDA). The ACFDA observes and detects the rapid change of acceleration in vertical axis and the average value of signal magnitude vector of acceleration to differentiate falls from other activities of daily life (ADL). Initial results demonstrates that our algorithm achieved 100% of sensitivity, 95.65% of specificity and 96.35% of accuracy when tested with a total of 44 intentional falls and 230 ADLs in 32 datasets. Future work will focus on developing other strategies to reduce false alarms for improving both specificity and accuracy of the algorithm while still maintaining 100% of sensitivity.
Hoa Hong Nguyen, Farhaan Mirza, Muhammad Asif Naeem, Mirza Mansoor Baig
MobiQuitous3
2017 Skewed distributions in semi-stream joins: How much can caching help?
Muhammad Asif Naeem, Gillian Dobbie, Christof Lutteroth, Gerald Weber
Inf. Syst.1
2015 S3J: A Parallel Semi-Stream Similarity Join
abstract
Semi-stream join algorithms join a continuous stream with a large disk-based relation. While there are efficient semi-stream equijoins for exact matches in the joined data, there are currently no semi-stream similarity joins for approximate matches. The existing similarity join algorithms work either offline (on datasets that are fully known) or on several streams (using a join window), and are less suitable for applications where continuous, immediate and complete similarity join results are required. To address this gap we propose S3J, the first semi-stream similarity join algorithm. To utilize disk and CPU optimally, S3J combines a disk-intensive queue-based semi-stream join approach with a CPU-intensive similarity matching algorithm. The similarity matching algorithm is based on tries to minimize the memory footprint. Moreover, it supports parallel execution to utilize modern multicore CPUs. We provide a cost model for S3J and evaluate its performance empirically.
Muhammad Asif Naeem, Christof Lutteroth, Gerald Weber
DOLAP2
2014 Optimizing Queue-Based Semi-Stream Joins with Indexed Master Data
Muhammad Asif Naeem, Gerald Weber, Christof Lutteroth, Gillian Dobbie
DaWaK1
2014 Efficient processing of streaming updates with archived master data in near-real-time data warehousing
Muhammad Asif Naeem, Gillian Dobbie, Gerald Weber
Knowl. Inf. Syst.1
2013 Tuned X-HYBRIDJOIN for Near-Real-Time Data Warehousing
Muhammad Asif Naeem
APWeb1
2013 A generic front-stage for semi-stream processing
abstract
Recently, a number of semi-stream join algorithms have been published. The typical system setup for these consists of one fast stream input that has to be joined with a disk-based relation R. These semi-stream join approaches typically perform the join with a limited main memory partition assigned to them, which is generally not large enough to hold the whole relation R. We propose a caching approach that can be used as a front-stage for different semi-stream join algorithms, resulting in significant performance gains for common applications. We analyze our approach in the context of a seminal semi-stream join, MESHJOIN (Mesh Join), and provide a cost model for the resulting semi-stream join algorithm, which we call CMESHJOIN (Cached Mesh Join). The algorithm takes advantage of skewed distributions; this article presents results for Zipfian distributions of the type that appears in many applications.
Muhammad Asif Naeem, Gerald Weber, Gillian Dobbie, Christof Lutteroth
CIKM1
2013 SSCJ: A Semi-Stream Cache Join Using a Front-Stage Cache Module
Muhammad Asif Naeem, Gerald Weber, Gillian Dobbie, Christof Lutteroth
DaWaK1
2012 A Lightweight Stream-Based Join with Limited Resource Consumption
Muhammad Asif Naeem, Gillian Dobbie, Gerald Weber
DaWaK1
2012 Interacting with Data Warehouse by Using a Natural Language Interface
Muhammad Asif Naeem, Imran Sarwar Bajwa
NLDB1
2011 On Specifying Requirements Using a Semantically Controlled Representation
Imran Sarwar Bajwa, Muhammad Asif Naeem
NLDB2
2010 A swarm intelligence based clustering approach for outlier detection
abstract
Outlier detection is an important field in data mining and knowledge discovery, which aims to identify abnormal observations in a large dataset. Common application areas of outlier detection are intrusion detection in computer networks, credit cards fraud detection, detecting abnormal changes in stock prices, and identifying abnormal health conditions. We propose the use of a novel swarm intelligence based clustering technique called Hierarchical Particle Swarm Optimization Based Clustering (HPSO-clustering) for outlier detection. The proposed technique is able to perform Hierarchical Agglomerative Clustering (HAC) as well as outlier detection. In the proposed approach a swarm of particles evolves through different stages to identify outliers and normal clusters. The experimentation of the proposed approach is performed on benchmark datasets which show that the efficiency of the approach is better than some other popular outlier detection techniques.
Shafiq Alam, Gillian Dobbie, Patricia J. Riddle, Muhammad Asif Naeem
IEEE Congress on Evolutionary Computation4
2010 R-MESHJOIN for near-real-time data warehousing
abstract
To fulfill the increasing demand of business for the latest information, current data integration approaches are moving towards real-time updates. One important element in real-time data integration is the join of a continuous incoming data stream with a disk-based relation. In this paper we investigate a stream-based join algorithm, called mesh join (MESHJOIN), and propose an improved version called reduced MESHJOIN (R-MESHJOIN). Both algorithms tune the memory, allocating parts of the memory to key components. In MESHJOIN there is a dependency between the size of partitions in an internal queue for the stream data and the number of iterations required to bring the disk-based relation into memory. This dependency hampers the optimal distribution of memory among the join components. In particular the size of the disk-buffer varies with the size of the disk-based relation which is unnecessary. On the other hand the R-MESHJOIN algorithm removes this dependency. This enables an optimal distribution of available memory among the join components. In R-MESHJOIN a change in the size of the disk-based relation does not affect the size of the disk-buffer. An experimental study is conducted in order to validate the arguments.
Muhammad Asif Naeem, Gillian Dobbie, Gerald Weber, Shafiq Alam
DOLAP1