VLDB 2026 Research / reviewers in the wild / expert
Umair ul Hassan
dblp:77/10503
· DBLP profile ↗
11ranked-venue papers
5as first author
3since 2021 · last 2025
0000-0002-3647-9020ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-authorSystems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Stacking-Based Ensemble Deep Learning Approach to Forecast Offshore Wind Farm Power Output
Long Hoang, Umair ul Hassan |
IEEE Big Data | 2 |
| 2022 | Efficient Data Analytics on Augmented Similarity TripletsabstractData analysis requires a pairwise proximity measure over objects. Recent work has extended this to situations where the distance information between objects is given as comparison results of distances between three objects (triplets). Humans find comparison tasks much easier than the exact distance computation, and such data can be easily obtained in big quantities via crowdsourcing. In this work, we propose triplets augmentation, an efficient method to extend the triplets data by inferring the hidden implicit information from the existing data. Triplets augmentation improves the quality of kernel-based and kernel-free data analytics. We also propose a novel set of algorithms for common data analysis tasks based on triplets. These methods work directly with triplets and avoid kernel evaluations, thus are scalable to big data. We demonstrate that our methods outperform the current best-known techniques and are robust to noisy data. Sarwan Ali, Muhammad Ahmad 0005, Umair ul Hassan, Muhammad Asad Khan, Shafiq Alam |
IEEE Big Data | 3 |
| 2022 | Impact Of Missing Data Imputation On The Fairness And Accuracy Of Graph Node ClassifiersabstractAnalysis of the fairness of machine learning (ML) algorithms has attracted many researchers’ interest. Several studies have shown that ML methods produce a bias toward different groups, which limits the applicability of ML models in many applications, such as crime rate prediction. The data used for ML may have missing values, which, if not appropriately handled, are known to further harmfully affect fairness. To address this issue, many imputation methods have been proposed to deal with missing data. However, research on the effect of missing data imputation on fairness is still rather limited. In this paper, we analyze the impact of imputation on fairness in the context of graph data (node attributes) using different embedding and neural network methods. Extensive experiments on six datasets demonstrate several issues of fairness in graph node classification when dealing with missing data and various imputation techniques. We find that the choice of the imputation method affects both fairness and accuracy. Our results provide valuable insights into fairness ML over graph data and how to handle missingness in graphs efficiently. Haris Mansoor, Sarwan Ali, Shafiq Alam, Muhammad Asad Khan, Umair ul Hassan |
IEEE Big Data | 5 |
| 2020 | Creating Expert Knowledge by Relying on Language Learners: a Generic Approach for Mass-Producing Language Resources by Combining Implicit Crowdsourcing and Language LearningabstractWe introduce in this paper a generic approach to combine implicit crowdsourcing and language learning in order to mass-produce language resources (LRs) for any language for which a crowd of language learners can be involved. We present the approach by explaining its core paradigm that consists in pairing specific types of LRs with specific exercises, by detailing both its strengths and challenges, and by discussing how much these challenges have been addressed at present. Accordingly, we also report on on-going proof-of-concept efforts aiming at developing the first prototypical implementation of the approach in order to correct and extend an LR called ConceptNet based on the input crowdsourced from language learners. We then present an international network called the European Network for Combining Language Learning with Crowdsourcing Techniques (enetCollect) that provides the context to accelerate the implementation of this generic approach. Finally, we exemplify how it can be used in several language learning scenarios to produce a multitude of NLP resources and how it can therefore alleviate the long-standing NLP issue of the lack of LRs. Lionel Nicolas, Verena Lyding, Claudia Borg, Corina Forascu, Karën Fort, Katerina Zdravkova, Iztok Kosem, Jaka Cibej, Spela Arhar Holdt, Alice Millour, Christos T. Rodosthenous, Federico Sangati, Umair ul Hassan, Anisia Katinskaia, Anabela Barreiro, Lavinia Aparaschivei, Yaakov HaCohen-Kerner |
LREC | 14 |
| 2020 | Using Crowdsourced Exercises for Vocabulary Training to Expand ConceptNetabstractIn this work, we report on a crowdsourcing experiment conducted using the V-TREL vocabulary trainer which is accessed via a Telegram chatbot interface to gather knowledge on word relations suitable for expanding ConceptNet. V-TREL is built on top of a generic architecture implementing the implicit crowdsourding paradigm in order to offer vocabulary training exercises generated from the commonsense knowledge-base ConceptNet and – in the background – to collect and evaluate the learners’ answers to extend ConceptNet with new words. In the experiment about 90 university students learning English at C1 level, based on Common European Framework of Reference for Languages (CEFR), trained their vocabulary with V-TREL over a period of 16 calendar days. The experiment allowed to gather more than 12,000 answers from learners on different question types. In this paper we present in detail the experimental setup and the outcome of the experiment, which indicates the potential of our approach for both crowdsourcing data as well as fostering vocabulary skills. Christos T. Rodosthenous, Verena Lyding, Federico Sangati, Umair ul Hassan, Lionel Nicolas, Jolita Horbacauskiene, Anisia Katinskaia, Lavinia Aparaschivei |
LREC | 5 |
| 2019 | A Real-time Linked Dataspace for the Internet of Things: Enabling "Pay-As-You-Go" Data Management in Smart EnvironmentsabstractAs smart environments move from a research vision to concrete manifestations in real-world enabled by the Internet of Things , they are encountering a number of very practical challenges in data management in terms of the flexibility needed to bring together contextual and real-time data, the interface between new digital infrastructures and existing information systems , and how to easily share data between stakeholders in the environment. Therefore, data management approaches for smart environments need to support flexibility, dynamicity, incremental change, while keeping costs to a minimum. A Dataspace is an emerging approach to data management that has proved fruitful for personal information and scientific data management. However, their use within smart environments and for real-time data remains largely unexplored. This paper introduces a Real-time Linked Dataspace (RLD) as an enabling platform for data management within smart environments. This paper identifies common data management requirements for smart energy and water environments, details the RLD architecture and the key support services and their tiered support levels, and a principled approach to “Pay-As-You-Go” data management. The paper presents a dataspace query service for real-time data streams and entities to enable unified entity-centric queries across live and historical stream data. The RLD was validated in 5 real-world pilot smart environments following the OODA (Observe, Orient, Decide, and Act) Loop to build real-time analytics, decisions support, and smart apps for energy and water management. The pilots demonstrate that the RLD enables incremental pay-as-you-go data management with support services that simplify the development of applications and analytics for smart environments. Finally, the paper discusses experiences, lessons learnt, and future directions. Edward Curry, Wassim Derguech, Souleiman Hasan, Christos Kouroupetroglou, Umair ul Hassan |
Future Gener. Comput. Syst. | 5 |
| 2016 | ACRyLIQ: Leveraging DBpedia for Adaptive Crowdsourcing in Linked Data Quality Assessment
Umair ul Hassan, Amrapali Zaveri, Edgard Marx, Edward Curry, Jens Lehmann 0001 |
EKAW | 1 |
| 2016 | Efficient task assignment for spatial crowdsourcing: A combinatorial fractional optimization approach with semi-bandit learningabstractSpatial crowdsourcing has emerged as a new paradigm for solving problems in the physical world with the help of human workers. A major challenge in spatial crowdsourcing is to assign reliable workers to nearby tasks. The goal of such task assignment process is to maximize the task completion in the face of uncertainty. This process is further complicated when tasks arrivals are dynamic and worker reliability is unknown. Recent research proposals have tried to address the challenge of dynamic task assignment. Yet the majority of the proposals do not consider the dynamism of tasks and workers. They also make the unrealistic assumptions of known deterministic or probabilistic workers’ reliabilities. In this paper, we propose a novel approach for dynamic task assignment in spatial crowdsourcing. The proposed approach combines bi-objective optimization with combinatorial multi-armed bandits. We formulate an online optimization problem to maximize task reliability and minimize travel costs in spatial crowdsourcing. We propose the distance-reliability ratio (DRR) algorithm based on a combinatorial fractional programming approach. The DRR algorithm reduces travel costs by 80% while maximizing reliability when compared to existing algorithms. We extend the DRR algorithm for the scenario when worker reliabilities are unknown. We propose a novel algorithm (DRR-UCB) that uses an interval estimation heuristic to approximate worker reliabilities. Experimental results demonstrate that the DRR-UCB achieves high reliability in the face of uncertainty. The proposed approach is particularly suited for real-life dynamic spatial crowdsourcing scenarios. This approach is generalizable to the similar problems in other areas in expert systems . First, it encompasses online assignment problems when the objective function is a ratio of two linear functions. Second, it considers situations when intelligent and repeated assignment decisions are needed under uncertainty. Umair ul Hassan, Edward Curry |
Expert Syst. Appl. | 1 |
| 2015 | Flag-verify-fix: adaptive spatial crowdsourcing leveraging location-based social networksabstractThis paper introduces the flag-verify-fix pattern that employs spatial crowdsourcing for city maintenance. The patterns motivates the need for appropriate assignment of dynamically arriving spatial tasks to a pool for workers on the ground. The assignment is aimed at maximizing the coverage of tasks spread over spatial locations; however, the coverage depends of willingness of workers to perform tasks assigned to them. We introduce the maximum coverage assignment problem that formulates two design issues of dynamic assignment. The quantity issue determines the number of worker required for a task and selection issue determines the set of workers. We propose an adaptive algorithm that uses location diversity based on a location-based social network to address the quantity issue and employs Thompson sampling for selecting the workers by learning their willingness. We evaluate the performance of the proposed algorithm in terms of coverage and number of assignments using real world datasets. The results show that our proposed algorithm achieves 30%--50% more coverage than the baseline algorithms, while requiring less workers per task. Umair ul Hassan, Edward Curry |
SIGSPATIAL/GIS | 1 |
| 2013 | A collaborative approach for metadata management for Internet of Things: Linking micro tasks with physical objectsabstractThere has been considerable efforts in modelling the semantics of Internet of Things and their specific context. Acquiring and managing metadata related to the physical devices and their surrounding environment becomes challenging due to the dynamic nature of environment. This paper focuses on manag Umair ul Hassan, Murilo Bassora, Ali Hosseinzadeh Vahid, Seán O'Riain, Edward Curry |
CollaborateCom | 1 |
| 2013 | A capability requirements approach for predicting worker performance in crowdsourcingabstractAssigning heterogeneous tasks to workers is an important challenge of crowdsourcing platforms. Current approaches to task assignment have primarily focused on contentbased approaches, qualifications, or work history. We propose an alternative and complementary approach that focuses on what capabil Umair ul Hassan, Edward Curry |
CollaborateCom | 1 |