EDBT 2026 Demo / reviewers in the wild / expert
Shazia Sadiq
dblp:s/SWSadiq · also Shazia W. Sadiq, Shazia Wasim Sadiq
· DBLP profile ↗
71ranked-venue papers in the field
6as first author
24since 2021 · last 2026
0000-0001-6739-4145ORCID · verified
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 27 (3 first)Information Retrieval & Web Search · 24 (1 first)Business Process & Enterprise Data · 9 (1 first)Data Mining & Knowledge Discovery · 8Knowledge Engineering, Semantic Web & Information Systems · 2Other / Interdisciplinary · 1 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Decomposition-Driven Multi-Table Retrieval and Reasoning for Numerical Question AnsweringabstractIn this paper, we study the problem of numerical multi-table question answering (MTQA) over large-scale table collections (e.g., online data repositories). This task is essential in many analytical applications. Existing MTQA solutions, such as text-to-SQL or open-domain MTQA methods, are designed for databases and struggle when applied to large-scale table collections. The key limitations include: (1) Limited support for complex table relationships; (2) Ineffective retrieval of relevant tables at scale; (3) Inaccurate answer generation. To overcome these limitations, we propose DMRAL, a Decomposition-driven Multi-table Retrieval and Answering framework for MTQA over large-scale table collections, which consists of: (1) constructing a table relationship graph to capture complex relationships among tables; (2) Table-Aligned Question Decomposer and Coverage-Aware Retriever, which jointly enable the effective identification of relevant tables from large-scale corpora by enhancing the question decomposition quality and maximizing the question coverage of retrieved tables; and (3) Sub-question Guided Reasoner, which produces correct answers by progressively generating and refining the reasoning program based on sub-questions. Experiments on two MTQA datasets demonstrate that DMRAL significantly outperforms existing state-of-the-art MTQA methods, with an average improvement of 24% in table retrieval and 55% in answer accuracy. Feng Luo 0005, Hui Luo 0001, Zhifeng Bao, Xiaoli Wang 0002, J. Shane Culpepper, Shazia Sadiq |
ICDE | 7 |
| 2026 | When Graph Contrastive Learning Backfires: Spectral Vulnerability and Defense in RecommendationabstractGraph Contrastive Learning (GCL) has demonstrated substantial promise in enhancing the robustness and generalization of recommender systems, particularly by enabling models to leverage large-scale unlabeled data for improved representation learning. However, in this article, we reveal an unexpected vulnerability: the integration of GCL inadvertently increases the susceptibility of a recommender to targeted promotion attacks. Through both theoretical investigation and empirical validation, we identify the root cause as the spectral smoothing effect induced by contrastive optimization, which disperses item embeddings across the representation space and unintentionally enhances the exposure of target items. Building on this insight, we introduce a bi-level optimization attack method, named graph Contrastive Learning Recommendation Attack (CLeaR), which deliberately amplifies spectral smoothness and enables a systematic investigation of the susceptibility of GCL-based recommendation models to targeted promotion attacks. Our findings highlight the urgent need for robust countermeasures; in response, we further propose a Spectral-Irregularity Mitigation framework, named SIM, which accurately detects and suppresses targeted items without compromising model performance. Extensive experiments on multiple benchmark datasets demonstrate that, compared to existing targeted promotion attacks, GCL-based recommendation models exhibit greater susceptibility when evaluated with CLeaR, while SIM effectively mitigates these vulnerabilities. Zongwei Wang 0002, Min Gao 0001, Junliang Yu, Shazia Sadiq, Hongzhi Yin, Ling Liu 0001 |
ACM Trans. Inf. Syst. | 4 |
| 2026 | Missing Value Imputation in Tabular Data Lakes Unleashed: A Hybrid ApproachabstractAbstract Missing values in tabular data lakes can severely impact data analysis and diminish the performance in downstream applications. We highlight that a robust imputation strategy should properly take three aspects of variety into consideration: source of imputed value, the types of tables involved, and the data types of the missing value. Existing imputation methods rely on estimation-based approaches (using a model trained on data from the same table to estimate missing values) or search-based approaches (retrieving values from other tables). Unfortunately, none of these approaches effectively incorporate all three aspects of variety. To address this gap, we propose , a novel framework that uses a C ombination of E stimation-based and S earch-based methods for missing value I mputation in D ata lakes. contains three core modules: (1) the , which efficiently discovers candidate values from tables by exploiting the contextual information; (2) the , which introduces an influence function and a sampling-based exploration strategy to yield accurate estimated values; (3) the , which determines the most suitable method based on table-level and column-level statistics. Extensive experiments conducted on three data lakes demonstrate that effectively and efficiently addresses the missing value problem. Feng Luo 0005, Hui Luo 0001, Zhifeng Bao, J. Shane Culpepper, Shazia Sadiq, Xiaoli Wang 0002 |
VLDB J. | 6 |
| 2025 | How Do Experts Make Sense of Integrated Process Models?
Tianwa Chen, Barbara Weber, Graeme G. Shanks, Gianluca Demartini, Marta Indulska, Shazia Sadiq |
CAiSE (2) | 6 |
| 2025 | LLM-Based Semantic Augmentation for Harmful Content DetectionabstractRecent advances in large language models (LLMs) have demonstrated strong performance on simple text classification tasks, frequently under zero-shot settings. However, their efficacy declines when tackling complex social media challenges such as propaganda detection, hateful meme classification, and toxicity identification. Much of the existing work has focused on using LLMs to generate synthetic training data, overlooking the potential of LLM-based text preprocessing and semantic augmentation. In this paper, we introduce an approach that prompts LLMs to clean noisy text and provide context-rich explanations, thereby enhancing training sets without substantial increases in data volume. We systematically evaluate on the SemEval 2024 multi-label Persuasive Meme dataset and further validate on the Google Jigsaw toxic comments and Facebook hateful memes datasets to assess generalizability. Our results reveal that zero-shot LLM classification underperforms on these high-context tasks compared to supervised models. In contrast, integrating LLM-based semantic augmentation yields performance on par with approaches that rely on human-annotated data, at a fraction of the cost. These findings underscore the importance of strategically incorporating LLMs into machine learning (ML) pipeline for social media classification tasks, offering broad implications for combating harmful content online. Disclaimer: This paper contains examples of explicit language that may be disturbing to some readers. Elyas Meguellati, Assaad Oussama Zeghina, Shazia Sadiq, Gianluca Demartini |
ICWSM | 3 |
| 2025 | Progressive Generalization Risk Reduction for Data-Efficient Causal Effect EstimationabstractCausal effect estimation (CEE) provides a crucial tool for predicting the unobserved counterfactual outcome for an entity. As CEE relaxes the requirement for "perfect'' counterfactual samples (e.g., patients with identical attributes and only differ in treatments received) that are impractical to obtain and can instead operate on observational data, it is usually used in high-stake domains like medical treatment effect prediction. Nevertheless, in those high-stake domains, gathering a decently sized, fully labelled observational dataset remains challenging due to hurdles associated with costs, ethics, expertise and time needed, etc., of which medical treatment surveys are a typical example. Consequently, if the training dataset is small in scale, low generalization risks can hardly be achieved on any CEE algorithms. Hechuan Wen, Tong Chen 0005, Guanhua Ye, Li Kheng Chai, Shazia Sadiq, Hongzhi Yin |
KDD (1) | 5 |
| 2025 | ID-Free Not Risk-Free: LLM-Powered Agents Unveil Risks in ID-Free Recommender SystemsabstractRecent advances in ID-free recommender systems have attracted significant attention for effectively addressing the cold start problem. However, their vulnerability to malicious attacks remains largely unexplored. In this paper, we unveil a critical yet overlooked risk: LLM-powered agents can be strategically deployed to attack ID-free recommenders, stealthily promoting low-quality items in black-box settings. This attack exploits a novel rewriting-based deception strategy, where malicious agents synthesize deceptive textual descriptions by simulating the characteristics of popular items. To achieve this, the attack mechanism integrates two primary components: (1) a popularity extraction component that captures essential characteristics of popular items and (2) a multi-agent collaboration mechanism that enables iterative refinement of promotional textual descriptions through independent thinking and team discussion. To counter this risk, we further introduce a detection method to identify suspicious text generated by our discovered attack. By unveiling this risk, our work aims to underscore the urgent need to enhance the security of ID-free recommender systems. Zongwei Wang 0002, Min Gao 0001, Junliang Yu, Xinyi Gao 0001, Nguyen Quoc Viet Hung, Shazia Sadiq, Hongzhi Yin |
SIGIR | 6 |
| 2025 | Continual Text-to-Video Retrieval with Frame Fusion and Task-Aware RoutingabstractText-to-Video Retrieval (TVR) aims to retrieve relevant videos based on textual queries.However, as video content evolves continuously, adapting TVR systems to new data remains a critical yet underexplored challenge.In this paper, we introduce the first benchmark for Continual Text-to-Video Retrieval (CTVR) to address the limitations of existing approaches.Current Pre-Trained Model (PTM)based TVR methods struggle with maintaining model plasticity when adapting to new tasks, while existing Continual Learning (CL) methods suffer from catastrophic forgetting, leading to semantic misalignment between historical queries and stored video features.To address these two challenges, we propose FrameFu-sionMoE, a novel CTVR framework that comprises two key components: (1) the Frame Fusion Adapter (FFA), which captures temporal video dynamics while preserving model plasticity, and (2) the Task-Aware Mixture-of-Experts (TAME), which ensures consistent semantic alignment between queries across tasks and the stored video features.Thus, FrameFusionMoE enables effective adaptation to new video content while preserving historical textvideo relevance to mitigate catastrophic forgetting.We comprehensively evaluate FrameFusionMoE on two benchmark datasets under various task settings.Results demonstrate that FrameFusionMoE outperforms existing CL and TVR methods, achieving superior retrieval performance with minimal degradation on earlier tasks when handling continuous video streams.Our code is available at: https://github.com/JasonCodeMaker/CTVR. Zecheng Zhao, Zhi Chen 0010, Zi Huang, Shazia Sadiq, Tong Chen 0005 |
SIGIR | 4 |
| 2025 | Towards Secure and Robust Recommender Systems: A Data-Centric PerspectiveabstractAs recommender systems (RS) continue to evolve, the field has seen a pivotal shift from model-centric to data-centric paradigms, where the quality, integrity, and security of data are increasingly becoming the key drivers of system performance and personalization. This transformation has unlocked new avenues for more precise recommendations, yet it also introduces significant challenges. As reliance on data intensifies, RS face mounting threats that can compromise both their effectiveness and user trust. These challenges include (1) Malicious Data Manipulation, where adversaries corrupt or tamper with datasets, distorting recommendation outcomes and undermining system reliability; (2) Data Privacy Leakage, where adversarial actors exploit system outputs to infer sensitive user information, leading to serious privacy concerns; and (3) Erroneous Data Noise, where inaccuracies, inconsistencies, and redundant data obscure the true user preferences, degrading recommendation quality and user satisfaction. By focusing on these critical data-centric challenges, this tutorial aims to equip participants with the knowledge to build RS that are secure, privacy-preserving, and resilient to data-driven threats, ensuring reliable and trustworthy performance in real-world environments. In addition, attendees will gain hands-on experience with our newly released toolkit for RS-based attacks and defenses, providing them with practical, actionable insights into safeguarding RS against emerging vulnerabilities. Zongwei Wang 0002, Junliang Yu, Tong Chen 0005, Hongzhi Yin, Shazia Sadiq, Min Gao 0001 |
WSDM | 5 |
| 2024 | EMIT - Event-Based Masked Auto Encoding for Irregular Time SeriesabstractIrregular time series, where data points are recorded at uneven intervals, are prevalent in healthcare settings, such as emergency wards where vital signs and laboratory results are captured at varying times. This variability, which reflects critical fluctuations in patient health, is essential for informed clinical decision-making. Existing self-supervised learning research on irregular time series often relies on generic pretext tasks like forecasting, which may not fully utilise the signal provided by irregular time series. There is a significant need for specialised pretext tasks designed for the characteristics of irregular time series to enhance model performance and robustness, especially in scenarios with limited data availability. This paper proposes a novel pretraining framework, EMIT, an event-based masking for irregular time series. EMIT focuses on masking-based reconstruction in the latent space, selecting masking points based on the rate of change in the data. This method preserves the natural variability and timing of measurements while enhancing the model's ability to process irregular intervals without losing essential information. Extensive experiments on the MIMIC-III and PhysioNet Challenge datasets demonstrate the superior performance of our event-based masking strategy. The code has been released at https://github.com/hrishi-ds/EMIT. Hrishikesh Patel, Ruihong Qiu, Adam Irwin, Shazia Sadiq, Sen Wang 0001 |
ICDM | 4 |
| 2024 | On the Role of Large Language Models in Crowdsourcing Misinformation AssessmentabstractThe proliferation of online misinformation significantly undermines the credibility of web content. Recently, crowd workers have been successfully employed to assess misinformation to address the limited scalability of professional fact-checkers. An alternative approach to crowdsourcing is the use of large language models (LLMs). These models are however also not perfect. In this paper, we investigate the scenario of crowd workers working in collaboration with LLMs to assess misinformation. We perform a study where we ask crowd workers to judge the truthfulness of statements under different conditions: with and without LLMs labels and explanations. Our results show that crowd workers tend to overestimate truthfulness when exposed to LLM-generated information. Crowd workers are misled by wrong LLM labels, but, on the other hand, their self-reported confidence is lower when they make mistakes due to relying on the LLM. We also observe diverse behaviors among crowd workers when the LLM is presented, indicating that leveraging LLMs can be considered a distinct working strategy. Jiechen Xu, Lei Han 0003, Shazia Sadiq, Gianluca Demartini |
ICWSM | 3 |
| 2024 | Unveiling Vulnerabilities of Contrastive Recommender Systems to Poisoning AttacksabstractContrastive learning (CL) has recently gained prominence in the domain of recommender systems due to its great ability to enhance recommendation accuracy and improve model robustness. Despite its advantages, this paper identifies a vulnerability of CL-based recommender systems that they are more susceptible to poisoning attacks aiming to promote individual items. Our analysis indicates that this vulnerability is attributed to the uniform spread of representations caused by the InfoNCE loss. Furthermore, theoretical and empirical evidence shows that optimizing this loss favors smooth spectral values of representations. This finding suggests that attackers could facilitate this optimization process of CL by encouraging a more uniform distribution of spectral values, thereby enhancing the degree of representation dispersion. With these insights, we attempt to reveal a potential poisoning attack against CL-based recommender systems, which encompasses a dual-objective framework: one that induces a smoother spectral value distribution to amplify the InfoNCE loss's inherent dispersion effect, named dispersion promotion; and the other that directly elevates the visibility of target items, named rank promotion. We validate the threats of our attack model through extensive experimentation on four datasets. By shedding light on these vulnerabilities, our goal is to advance the development of more robust CL-based recommender systems. The code is available at https://github.com/CoderWZW/ARLib. Zongwei Wang 0002, Junliang Yu, Min Gao 0001, Hongzhi Yin, Bin Cui 0001, Shazia Sadiq |
KDD | 6 |
| 2024 | Variational Counterfactual Prediction Under Runtime Domain CorruptionabstractTo date, various neural methods have been proposed for causal effect estimation based on observational data, where a default assumption is the same distribution and availability of variables at both training and inference (i.e., runtime) stages. However, distribution shift (i.e., domain shift) could happen during runtime, and bigger challenges arise from the impaired accessibility of variables. This is commonly caused by increasing privacy and ethical concerns, which can make arbitrary variables unavailable in the entire runtime data and imputation impractical. We term the co-occurrence of domain shift and inaccessible variablesruntime domain corruption, which seriously impairs the generalizability of a trained counterfactual predictor. To counter runtime domain corruption, we subsume counterfactual prediction under the notion of domain adaptation. Specifically, we upper-bound the error w.r.t. the target domain (i.e., runtime covariates) by the sum of source domain error and inter-domain distribution distance. In addition, we build an adversarially unified variational causal effect model, named VEGAN, with a novel two-stage adversarial domain adaptation scheme to reduce the latent distribution disparity between treated and control groups first, and between training and runtime variables afterwards. We demonstrate that VEGAN outperforms other state-of-the-art baselines on individual-level treatment effect estimation in the presence of runtime domain corruption on benchmark datasets. Hechuan Wen, Tong Chen 0005, Li Kheng Chai, Shazia Sadiq, Junbin Gao, Hongzhi Yin |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2024 | On the Impact of Showing Evidence from Peers in Crowdsourced Truthfulness AssessmentsabstractMisinformation has been rapidly spreading online. The common approach to dealing with it is deploying expert fact-checkers who follow forensic processes to identify the veracity of statements. Unfortunately, such an approach does not scale well. To deal with this, crowdsourcing has been looked at as an opportunity to complement the work done by trained journalists. In this article, we look at the effect of presenting the crowd with evidence from others while judging the veracity of statements. We implement variants of the judgment task design to understand whether and how the presented evidence may or may not affect the way crowd workers judge truthfulness and their performance. Our results show that, in certain cases, the presented evidence and the way in which it is presented may mislead crowd workers who would otherwise be more accurate if judging independently from others. Those who make appropriate use of the provided evidence, however, can benefit from it and generate better judgments. Jiechen Xu, Lei Han 0003, Shazia Sadiq, Gianluca Demartini |
ACM Trans. Inf. Syst. | 3 |
| 2023 | To Predict or to Reject: Causal Effect Estimation with Uncertainty on Networked DataabstractDue to the imbalanced nature of networked observational data, the causal effect predictions for some individuals can severely violate the positivity/overlap assumption, rendering unreliable estimations. Nevertheless, this potential risk of individual-level treatment effect estimation on networked data has been largely under-explored. To create a more trustworthy causal effect estimator, we propose the uncertainty-aware graph deep kernel learning (GraphDKL) framework with Lipschitz constraint to model the prediction uncertainty with Gaussian process and identify unreliable estimations. To the best of our knowledge, GraphDKL is the first framework to tackle the violation of positivity assumption when performing causal effect estimation with graphs. With extensive experiments, we demonstrate the superiority of our proposed method in uncertainty-aware causal effect estimation on networked data. The code of GraphDKL is available at https://github.com/uqhwen2/GraphDKL. Hechuan Wen, Tong Chen 0005, Li Kheng Chai, Shazia Sadiq, Kai Zheng 0001, Hongzhi Yin |
ICDM | 4 |
| 2023 | Human-in-the-loop Regular Expression Extraction for Single Column Format InconsistencyabstractFormat inconsistency is one of the most frequently appearing data quality issues encountered during data cleaning. Existing automated approaches commonly lack applicability and generalisability, while approaches with human inputs typically require specialized skills such as writing regular expressions. This paper proposes a novel hybrid human-machine system, namely “Data-Scanner-4C”, which leverages crowdsourcing to address syntactic format inconsistencies in a single column effectively. We first ask crowd workers to create examples from single-column data through “data selection” and “result validation” tasks. Then, we propose and use a novel rule-based learning algorithm to infer the regular expressions that propagate formats from created examples to the entire column. Our system integrates crowdsourcing and algorithmic format extraction techniques in a single workflow. Having human experts write regular expressions is no longer required, thereby reducing both the time as well as the opportunity for error. We conducted experiments through both synthetic and real-world datasets, and our results show how the proposed approach is applicable and effective across data types and formats. Shaochen Yu, Lei Han 0003, Marta Indulska, Shazia Sadiq, Gianluca Demartini |
WWW | 4 |
| 2023 | On the role of human and machine metadata in relevance judgment tasks
Jiechen Xu, Lei Han 0003, Shazia Sadiq, Gianluca Demartini |
Inf. Process. Manag. | 3 |
| 2023 | DataOps-4G: On Supporting Generalists in Data Quality DiscoveryabstractData preparation has become a necessary but labor and resource intensive step to perform data analytics. To date, such activities still require considerable manual effort from experts. In this paper, we focus on a specific data preparation activity, namely data quality discovery. We explore different settings in which data workers undertake data quality discovery tasks and the implications of those settings for the efficiency and effectiveness of data workers. To this end, we propose DataOps-4G, a data curation platform for generalists, that allows users to interact with data without the need to write code. We wrap up pre-defined code snippets that implement useful functionalities to explore data quality and bundle the code into so-called DataOps. Then, we conduct a lab-based user study to evaluate our DataOps-4G platform from two perspectives: (i) effectiveness, the accuracy of the outcomes achieved by participants; and (ii) efficiency, their effort and strategies in task completion. Our experimental results uncover how effectiveness and efficiency can be affected by their task completion patterns and strategies. This opens up the possibility of popularizing data curation processes by employing non-experts (e.g., from crowdsourcing platforms) and consequently allowing experts to focus on more complex activities (e.g., building machine learning models). Shaochen Yu, Tianwa Chen, Lei Han 0003, Gianluca Demartini, Shazia Sadiq |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2023 | A Data-Driven Analysis of Behaviors in Data Curation ProcessesabstractUnderstanding how data workers interact with data, and various pieces of information related to data preparation, is key to designing systems that can better support them in exploring datasets. To date, however, there is a paucity of research studying the strategies adopted by data workers as they carry out data preparation activities. In this work, we investigate a specific data preparation activity, namely data quality discovery , and aim to (i) understand the behaviors of data workers in discovering data quality issues, (ii) explore what factors (e.g., prior experience) can affect their behaviors, as well as (iii) understand how these behavioral observations relate to their performance. To this end, we collect a multi-modal dataset through a data-driven experiment that relies on the use of eye-tracking technology with a purpose-designed platform built on top of iPython Notebook. The experiment results reveal that: (i) ‘copy–paste–modify’ is a typical strategy for writing code to complete tasks; (ii) proficiency in writing code has a significant impact on the quality of task performance, while perceived difficulty and efficacy can influence task completion patterns; and (iii) searching in external resources is a prevalent action that can be leveraged to achieve better performance. Furthermore, our experiment indicates that providing sample code within the system can help data workers get started with their task, and surfacing underlying data is an effective way to support exploration. By investigating data worker behaviors prior to each search action, we also find that the most common reasons that trigger external search actions are the need to seek assistance in writing or debugging code and to search for relevant code to reuse. Based on our experiment results, we showcase a systematic approach to select from the top best code snippets created by data workers and assemble them to achieve better performance than the best individual performer in the dataset. By doing so, our findings not only provide insights into patterns of interactions with various system components and information resources when performing data curation tasks, but also build effective and efficient data curation processes through data workers’ collective intelligence. Lei Han 0003, Tianwa Chen, Gianluca Demartini, Marta Indulska, Shazia Sadiq |
ACM Trans. Inf. Syst. | 5 |
| 2022 | Workshop on Human-in-the-loop Data CurationabstractAlthough data quality is a long-standing and enduring problem, it has recently received a resurgence of attention due to the fast proliferation of data analytics, machine learning, and decision-support applications built upon the wide-scale availability and accessibility of (big) data. The success of such applications heavily relies on not only the quantity, but also the quality of data. Data curation, which may include annotation, cleaning, transformation, integration, etc., is a critical step to provide adequate assurances on the quality of analytics and machine learning results. Such data preparation activities are recognised as time and resource intensive for data scientists as data often comes with a number of challenges that need to be tackled before it can be used in practice. Data re-purposing and the resulting distance between design and use intentions of the data, is a fundamental issue behind many of these challenges. These challenges include a variety of data issues such as noise and outliers, incompleteness, representativeness or biases, heterogeneity of format or semantics, etc. Mishandling these challenges can lead to negative and sometimes damaging effects, especially in critical domains like healthcare, transport, and finance. An observable distinct feature of data quality in these contexts is the increasingly important role played by humans, being often the source of data generation and the active players in data curation. This workshop will provide an opportunity to explore the interdisciplinary overlap between manual, automated, and hybrid human-machine methods of data curation. Gianluca Demartini, Jie Yang 0028, Shazia Sadiq |
CIKM | 3 |
| 2022 | A Behavioural Analysis of Metadata Use in Evaluating the Quality of Repurposed Data
Lei Han 0003, Gianluca Demartini, Marta Indulska, Shazia Sadiq |
ER | 5 |
| 2022 | Exploring Data Literacy Levels in the Crowd - the Case of COVID-19
Shaoyang Fan, Lei Han 0003, Gianluca Demartini, Shazia Sadiq |
ICWSM | 4 |
| 2022 | Business process and rule integration approaches - An empirical analysis of model understanding
Wei Wang 0186, Tianwa Chen, Marta Indulska, Shazia Sadiq, Barbara Weber |
Inf. Syst. | 4 |
| 2022 | Information Resilience: the nexus of responsible and agile approaches to information useabstractAbstract The appetite for effective use of information assets has been steadily rising in both public and private sector organisations. However, whether the information is used for social good or commercial gain, there is a growing recognition of the complex socio-technical challenges associated with balancing the diverse demands of regulatory compliance and data privacy, social expectations and ethical use, business process agility and value creation, and scarcity of data science talent. In this vision paper, we present a series of case studies that highlight these interconnected challenges, across a range of application areas. We use the insights from the case studies to introduce Information Resilience, as a scaffold within which the competing requirements of responsible and agile approaches to information use can be positioned. The aim of this paper is to develop and present a manifesto for Information Resilience that can serve as a reference for future research and development in relevant areas of responsible data management. Shazia Sadiq, Amir Aryani, Gianluca Demartini, Wen Hua, Marta Indulska, Andrew Burton-Jones, Hassan Khosravi, Diana Benavides-Prado, Timos K. Sellis, Ida Asadi Someh, Rhema Vaithianathan, Sen Wang 0001, Xiaofang Zhou 0001 |
VLDB J. | 1 |
| 2020 | Modelling User Behavior Dynamics with EmbeddingsabstractUnderstanding user interaction behaviors remains a challenging problem. Quantifying behavior dynamics over time as users complete tasks has only been done in specific domains. In this paper, we present a user behavior model built using behavior embeddings to compare behaviors and their change over time. To this end, we first define the formal model and train the model using both action (e.g., copy/paste) embeddings and user interaction feature (e.g., length of the copied text) embeddings. Having obtained vector representations of user behaviors, we then define three measurements to model behavior dynamics over time, namely: behavior position, displacement, and velocity. To evaluate the proposed methodology, we use three real world datasets: (i) tens of users completing complex data curation tasks in a lab setting, (ii) hundreds of crowd workers completing structured tasks in a crowdsourcing setting, and (iii) thousands of editors completing unstructured editing tasks on Wikidata. Through these datasets, we show that the proposed methodology can: (i) surface behavioral differences among users; (ii) recognize relative behavioral changes; and (iii) discover directional deviations of user behaviors. Our approach can be used (i) to capture behavioral semantics from data in a consistent way, (ii) to quantify behavioral diversity for a task and among different users, and (iii) to explore the temporal behavior evolution with respect to various task properties (e.g., structure and difficulty). Lei Han 0003, Alessandro Checco, Djellel Eddine Difallah, Gianluca Demartini, Shazia Sadiq |
CIKM | 5 |
| 2020 | Sensemaking in Dual Artefact Tasks - The Case of Business Process Models and Business Rules
Tianwa Chen, Shazia Sadiq, Marta Indulska |
ER | 2 |
| 2020 | Identifying Cohorts: Recommending Drill-Downs Based on Differences in Behaviour for Process Mining
Sander J. J. Leemans, Shiva Shabaninejad, Kanika Goel 0002, Hassan Khosravi, Shazia Sadiq, Moe Thandar Wynn |
ER | 5 |
| 2020 | On Understanding Data Worker Interaction BehaviorsabstractUnderstanding how data workers interact with data and various pieces of information (e.g., code snippet examples) is key to design systems that can better support them in exploring a given dataset. To date, however, there is a paucity of research studying information seeking patterns and the strategies adopted by data workers as they carry out data curation activities. In this work, we aim at understanding the behaviors of data workers in discovering data quality issues, and how these behavioral observations relate to their performance. Specifically, we investigate how data workers use information resources and tools to support their task completion. To this end, we collect a multi-modal dataset through a data-driven experiment that relies on the use of eye-tracking technology with a purpose-designed platform built on top of iPython Notebook. The collected data reveals that: (i) searching in external resources is a prevalent action that can be leveraged to achieve better performance; (ii) 'copy-paste-modify' is a typical strategy for writing code to complete tasks; (iii) providing sample code within the system could help data workers to get started with their task; and (iv) surfacing underlying data is an effective way to support exploration. By investigating the behaviors prior to each search action, we also find that the most common reasons that trigger external search actions are the need to seek assistance in writing or debugging code and to search for relevant code to reuse. Our findings provide insights into patterns of interactions with various system components and information resources to perform data curation tasks. This bears implications on the design of domain-specific IR systems for data workers like code-base search. Lei Han 0003, Tianwa Chen, Gianluca Demartini, Marta Indulska, Shazia Sadiq |
SIGIR | 5 |
| 2020 | Factors influencing effective use of big data: A research framework
Feliks Sejahtera, Wei Wang 0186, Marta Indulska, Shazia Sadiq |
Inf. Manag. | 4 |
| 2018 | Special Issue of DASFAA 2018abstractWe are pleased to present a special issue of Data Science and Engineering (DSE), which contains a collection of five extended papers from the DASFAA 2018 conference.Besides these five papers, this DSE issue also has one invited paper.DASFAA 2018 is the 23rd International Conference on Database Systems for Advanced Applications.DASFAA is an annual international database conference, which provides a forum for technical presentations and discussions among database researchers, developers, and users from academia, business, and industry.This year the dominant topics for the selected papers included learning models, graph and network data processing, and social network analysis, followed by text and data mining, recommendation, data quality and crowd sourcing, and trajectory and stream data.Selected papers also included topics relating to network embedding, sequence and temporal data processing, RDF and knowledge graphs, security and privacy, medical data mining, query processing and optimization, search and information retrieval, multimedia data processing, and distributed computing.The 2018 edition of DASFAA was held in Gold Coast, Australia, and attracted a total of 360 research paper submissions, spanning over numerous active and emerging topic areas.The conference program committee selected 83 full research papers and 21 short papers, six industry papers, and eight demo papers to be presented at the conference and published in the conference proceedings [1,2].The conference program also included keynote presentations by Dr. Shazia Sadiq, Jianxin Li 0001 |
Data Sci. Eng. | 1 |
| 2018 | Guidelines for Business Rule Modeling DecisionsabstractBusiness process models are used heavily in practice as a basis for process improvement, systems development, and understanding business operations. While prior research has identified a clear need for integrating business rules into graphical business process models, there is little guidance on the circumstances under which business rules should be integrated into business process models. Unnecessary integration may hamper business rule reuse, increase business process model complexity, and lead to difficulties with business rule modification, to name a few. Accordingly, it is important to understand when such integration is appropriate. The aim of this article is to address this need for guidance on when business rules should be integrated in process models, and when they should remain separate. To this end, we explain 12 factors posited to influence such modeling decisions, conduct an empirical study to identify their importance, and develop empirically based modeling guidelines that inform business rule modeling decisions. Wei Wang 0186, Marta Indulska, Shazia Sadiq |
J. Comput. Inf. Syst. | 3 |
| 2017 | Jointly Modeling Heterogeneous Temporal Properties in Location Recommendation
Saeid Hosseini, Hongzhi Yin, Meihui Zhang 0001, Xiaofang Zhou 0001, Shazia Sadiq |
DASFAA (1) | 5 |
| 2017 | ST-SAGE: A Spatial-Temporal Sparse Additive Generative Model for Spatial Item RecommendationabstractWith the rapid development of location-based social networks (LBSNs), spatial item recommendation has become an important mobile application, especially when users travel away from home. However, this type of recommendation is very challenging compared to traditional recommender systems. A user may visit only a limited number of spatial items, leading to a very sparse user-item matrix. This matrix becomes even sparser when the user travels to a distant place, as most of the items visited by a user are usually located within a short distance from the user’s home. Moreover, user interests and behavior patterns may vary dramatically across different time and geographical regions. In light of this, we propose ST-SAGE, a spatial-temporal sparse additive generative model for spatial item recommendation in this article. ST-SAGE considers both personal interests of the users and the preferences of the crowd in the target region at the given time by exploiting both the co-occurrence patterns and content of spatial items. To further alleviate the data-sparsity issue, ST-SAGE exploits the geographical correlation by smoothing the crowd’s preferences over a well-designed spatial index structure called the spatial pyramid . To speed up the training process of ST-SAGE, we implement a parallel version of the model inference algorithm on the GraphLab framework. We conduct extensive experiments; the experimental results clearly demonstrate that ST-SAGE outperforms the state-of-the-art recommender systems in terms of recommendation effectiveness, model training efficiency, and online recommendation efficiency. Weiqing Wang 0001, Hongzhi Yin, Ling Chen 0006, Yizhou Sun, Shazia Sadiq, Xiaofang Zhou 0001 |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2016 | To Integrate or Not to Integrate - The Business Rules Question
Wei Wang 0186, Marta Indulska, Shazia Sadiq |
CAiSE | 3 |
| 2016 | Big data quality - whose problem is it?abstractThe increased reliance on data driven enterprise has seen an unprecedented investment in big data initiatives. Organizations averaged US$8M in investments in big data-related initiatives and programs in 2014, with 70% of large enterprises and 56% of small and medium enterprises (SMEs) having already deployed, or planning to deploy, big-data projects [1]. As companies intensify their efforts to get value from big data, the growth in the amount of data being managed continues at an exponential rate, leaving organizations with a massive footprint of unexplored, unfamiliar datasets. On February 8th, 2015, a group of global thought leaders from the database research community outlined the grand challenges in getting value from big data [2]. The key message was the need to develop the capacity to `understand how the quality of data affects the quality of the insight we derive from it'. Shazia Sadiq, Paolo Papotti |
ICDE | 1 |
| 2016 | SPORE: A sequential personalized spatial item recommender systemabstractWith the rapid development of location-based social networks (LBSNs), spatial item recommendation has become an important way of helping users discover interesting locations to increase their engagement with location-based services. Although human movement exhibits sequential patterns in LBSNs, most current studies on spatial item recommendations do not consider the sequential influence of locations. Leveraging sequential patterns in spatial item recommendation is, however, very challenging, considering 1) users' check-in data in LBSNs has a low sampling rate in both space and time, which renders existing prediction techniques on GPS trajectories ineffective; 2) the prediction space is extremely large, with millions of distinct locations as the next prediction target, which impedes the application of classical Markov chain models; and 3) there is no existing framework that unifies users' personal interests and the sequential influence in a principled manner. In light of the above challenges, we propose a sequential personalized spatial item recommendation framework (SPORE) which introduces a novel latent variable topic-region to model and fuse sequential influence with personal interests in the latent and exponential space. The advantages of modeling the sequential effect at the topic-region level include a significantly reduced prediction space, an effective alleviation of data sparsity and a direct expression of the semantic meaning of users' spatial activities. Furthermore, we design an asymmetric Locality Sensitive Hashing (ALSH) technique to speed up the online top-k recommendation process by extending the traditional LSH. We evaluate the performance of SPORE on two real datasets and one large-scale synthetic dataset. The results demonstrate a significant improvement in SPORE's ability to recommend spatial items, in terms of both effectiveness and efficiency, compared with the state-of-the-art methods. Weiqing Wang 0001, Hongzhi Yin, Shazia Sadiq, Ling Chen 0006, Xiaofang Zhou 0001 |
ICDE | 3 |
| 2016 | Discovering interpretable geo-social communities for user behavior predictionabstractSocial community detection is a growing field of interest in the area of social network applications, and many approaches have been developed, including graph partitioning, latent space model, block model and spectral clustering. Most existing work purely focuses on network structure information which is, however, often sparse, noisy and lack of interpretability. To improve the accuracy and interpretability of community discovery, we propose to infer users' social communities by incorporating their spatiotemporal data and semantic information. Technically, we propose a unified probabilistic generative model, User-Community-Geo-Topic (UCGT), to simulate the generative process of communities as a result of network proximities, spatiotemporal co-occurrences and semantic similarity. With a well-designed multi-component model structure and a parallel inference implementation to leverage the power of multicores and clusters, our UCGT model is expressive while remaining efficient and scalable to growing large-scale geo-social networking data. We deploy UCGT to two application scenarios of user behavior predictions: check-in prediction and social interaction prediction. Extensive experiments on two large-scale geo-social networking datasets show that UCGT achieves better performance than existing state-of-the-art comparison methods. Hongzhi Yin, Zhiting Hu, Xiaofang Zhou 0001, Hao Wang 0005, Kai Zheng 0001, Nguyen Quoc Viet Hung, Shazia Sadiq |
ICDE | 7 |
| 2016 | Preface to BPM 2014
Mathias Weske, Shazia Sadiq, Pnina Soffer, Hagen Völzer |
Inf. Syst. | 2 |
| 2016 | A Spatial-Temporal Topic Model for the Semantic Annotation of POIs in LBSNsabstractSemantic tags of points of interest (POIs) are a crucial prerequisite for location search, recommendation services, and data cleaning. However, most POIs in location-based social networks (LBSNs) are either tag-missing or tag-incomplete. This article aims to develop semantic annotation techniques to automatically infer tags for POIs. We first analyze two LBSN datasets and observe that there are two types of tags, category-related ones and sentimental ones, which have unique characteristics. Category-related tags are hierarchical, whereas sentimental ones are category-aware. All existing related work has adopted classification methods to predict high-level category-related tags in the hierarchy, but they cannot apply to infer either low-level category tags or sentimental ones. In light of this, we propose a latent-class probabilistic generative model, namely the spatial-temporal topic model (STM), to infer personal interests, the temporal and spatial patterns of topics/semantics embedded in users’ check-in activities, the interdependence between category-topic and sentiment-topic, and the correlation between sentimental tags and rating scores from users’ check-in and rating behaviors. Then, this learned knowledge is utilized to automatically annotate all POIs with both category-related and sentimental tags in a unified way. We conduct extensive experiments to evaluate the performance of the proposed STM on a real large-scale dataset. The experimental results show the superiority of our proposed STM, and we also observe that the real challenge of inferring category-related tags for POIs lies in the low-level ones of the hierarchy and that the challenge of predicting sentimental tags are those with neutral ratings. Tieke He, Hongzhi Yin, Zhenyu Chen 0001, Xiaofang Zhou 0001, Shazia Sadiq, Bin Luo 0003 |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2016 | Joint Modeling of User Check-in Behaviors for Real-time Point-of-Interest RecommendationabstractPoint-of-Interest (POI) recommendation has become an important means to help people discover attractive and interesting places, especially when users travel out of town. However, the extreme sparsity of a user-POI matrix creates a severe challenge. To cope with this challenge, we propose a unified probabilistic generative model, the Topic-Region Model (TRM) , to simultaneously discover the semantic, temporal, and spatial patterns of users’ check-in activities, and to model their joint effect on users’ decision making for selection of POIs to visit. To demonstrate the applicability and flexibility of TRM, we investigate how it supports two recommendation scenarios in a unified way, that is, hometown recommendation and out-of-town recommendation. TRM effectively overcomes data sparsity by the complementarity and mutual enhancement of the diverse information associated with users’ check-in activities (e.g., check-in content, time, and location) in the processes of discovering heterogeneous patterns and producing recommendations. To support real-time POI recommendations, we further extend the TRM model to an online learning model, TRM-Online, to track changing user interests and speed up the model training. In addition, based on the learned model, we propose a clustering-based branch and bound algorithm (CBB) to prune the POI search space and facilitate fast retrieval of the top- k recommendations. We conduct extensive experiments to evaluate the performance of our proposals on two real-world datasets, including recommendation effectiveness, overcoming the cold-start problem, recommendation efficiency, and model-training efficiency. The experimental results demonstrate the superiority of our TRM models, especially TRM-Online, compared with state-of-the-art competitive methods, by making more effective and efficient mobile recommendations. In addition, we study the importance of each type of pattern in the two recommendation scenarios, respectively, and find that exploiting temporal patterns is most important for the hometown recommendation scenario, while the semantic patterns play a dominant role in improving the recommendation effectiveness for out-of-town users. Hongzhi Yin, Bin Cui 0001, Xiaofang Zhou 0001, Weiqing Wang 0001, Zi Huang, Shazia Sadiq |
ACM Trans. Inf. Syst. | 6 |
| 2015 | Joint Modeling of User Check-in Behaviors for Point-of-Interest RecommendationabstractPoint-of-Interest (POI) recommendation has become an important means to help people discover attractive and interesting locations, especially when users travel out of town. However, extreme sparsity of user-POI matrix creates a severe challenge. To cope with this challenge, a growing line of research has exploited the temporal effect, geographical-social influence, content effect and word-of-mouth effect. However, current research lacks an integrated analysis of the joint effect of the above factors to deal with the issue of data-sparsity, especially in the out-of-town recommendation scenario which has been ignored by most existing work. Hongzhi Yin, Xiaofang Zhou 0001, Yingxia Shao, Hao Wang 0005, Shazia Sadiq |
CIKM | 5 |
| 2015 | Making Sense of Spatial TrajectoriesabstractSpatial trajectory data is widely available today. Over a sustained period of time, trajectory data has been collected from numerous GPS devices, smartphones, sensors and social media applications. Daily increases of real-time trajectory data have also been phenomenal in recent years. More and more new applications have emerged to derive business values from both trajectory data warehouses and real-time trajectory data. Due to their very large volumes, their nature of streaming, their highly variable levels of data quality, as well as many possible links with other types of data, making sense of spatial trajectory data becomes one of the crucial areas for big data analytics. In this paper we will present a review of the extensive work in spatiotemporal data management and trajectory mining, and discuss new challenges and new opportunities in the context of new applications, focusing on recent advances in trajectory data management and trajectory mining from their foundations to high performance processing with modern computing infrastructure. Xiaofang Zhou 0001, Kai Zheng 0001, Hoyoung Jeung, Jiajie Xu 0001, Shazia Sadiq |
CIKM | 5 |
| 2015 | Making sense of trajectory data: A partition-and-summarization approachabstractDue to the prevalence of GPS-enabled devices and wireless communication technology, spatial trajectories that describe the movement history of moving objects are being generated and accumulated at an unprecedented pace. However, a raw trajectory in the form of sequence of timestamped locations does not make much sense for humans without semantic representation. In this work we aim to facilitate human's understanding of a raw trajectory by automatically generating a short text to describe it. By formulating this task as the problem of adaptive trajectory segmentation and feature selection, we propose a partition-and-summarization framework. In the partition phase, we first define a set of features for each trajectory segment and then derive an optimal partition with the aim to make the segments within each partition as homogeneous as possible in terms of their features. In the summarization phase, for each partition we select the most interesting features by comparing against the common behaviours of historical trajectories on the same route and generate short text description for these features. For empirical study, we apply our solution to a real trajectory dataset and have found that the generated text can effectively reflect the important parts in a trajectory. Han Su 0001, Kai Zheng 0001, Kai Zeng 0002, Jiamin Huang, Shazia Sadiq, Nicholas Jing Yuan, Xiaofang Zhou 0001 |
ICDE | 5 |
| 2015 | Approximate keyword search in semantic trajectory databaseabstractDriven by the advances in location positioning techniques and the popularity of location sharing services, semantic enriched trajectory data have become unprecedentedly available. While finding relevant Point-of-Interest (POIs) based on users' locations and query keywords has been extensively studied in the past years, it is largely untouched to explore the keyword queries in the context of semantic trajectory database. In this paper, we study the problem of approximate keyword search in massive semantic trajectories. Given a set of query keywords, an approximate keyword query of semantic trajectory (AKQST) returns k trajectories that contain the most relevant keywords to the query and yield the least travel effort in the meantime. The main difference between AKQST and conventional spatial keyword queries is that there is no query location in AKQST, which means the search area cannot be localized. To capture the travel effort in the context of query keywords, a novel utility function, called spatio-textual utility function, is first defined. Then we develop a hybrid index structure called GiKi to organize the trajectories hierarchically, which enables pruning the search space by spatial and textual similarity simultaneously. Finally an efficient search algorithm and fast evaluation of the minimum value of spatio-textual utility function are proposed. The results of our empirical studies based on real check-in datasets demonstrate that our proposed index and algorithms can achieve good scalability. Bolong Zheng, Nicholas Jing Yuan, Kai Zheng 0001, Xing Xie 0001, Shazia Sadiq, Xiaofang Zhou 0001 |
ICDE | 5 |
| 2015 | Geo-SAGE: A Geographical Sparse Additive Generative Model for Spatial Item RecommendationabstractWith the rapid development of location-based social networks (LBSNs), spatial item recommendation has become an important means to help people discover attractive and interesting venues and events, especially when users travel out of town. However, this recommendation is very challenging compared to the traditional recommender systems. A user can visit only a limited number of spatial items, leading to a very sparse user-item matrix. Most of the items visited by a user are located within a short distance from where he/she lives, which makes it hard to recommend items when the user travels to a far away place. Moreover, user interests and behavior patterns may vary dramatically across different geographical regions. In light of this, we propose Geo-SAGE, a geographical sparse additive generative model for spatial item recommendation in this paper. Geo-SAGE considers both user personal interests and the preference of the crowd in the target region, by exploiting both the co-occurrence pattern of spatial items and the content of spatial items. To further alleviate the data sparsity issue, Geo-SAGE exploits the geographical correlation by smoothing the crowd's preferences over a well-designed spatial index structure called spatial pyramid. We conduct extensive experiments and the experimental results clearly demonstrate our Geo-SAGE model outperforms the state-of-the-art. Weiqing Wang 0001, Hongzhi Yin, Ling Chen 0006, Yizhou Sun, Shazia Sadiq, Xiaofang Zhou 0001 |
KDD | 5 |
| 2015 | SharkDB: An In-Memory Storage System for Massive Trajectory DataabstractAn increasing amount of motion history data, which is called trajectory, is being collected from different sources such as GPS-enabled mobile devices, surveillance cameras and social networks. However it is hard to store and manage trajectory data in traditional database systems, since its variable lengths and asynchronous sampling rates do not fit disk-based and tuple-oriented structures, which are the fundamental structures of traditional database systems. We implement a novel trajectory storage system that is motivated by the success of column store and recent development of in-memory based databases. In this storage design, we try to explore the potential opportunities, which can boost the performance of query processing for trajectory data. To achieve this, we partition the trajectories into frames as column-oriented storage in order to store the sample points of a moving object, which are aligned by the time interval, within the main memory. Furthermore, the frames can be highly compressed and well structured to increase the memory utilization ratio and reduce the CPU-cache missing. It is also easier for parallelizing data processing on the multi-core server since the frames are mutually independent. Haozhou Wang, Kai Zheng 0001, Xiaofang Zhou 0001, Shazia Sadiq |
SIGMOD Conference | 4 |
| 2014 | EISA: An Efficient Information Theoretical Approach to Value Segmentation in Large Databases
Weiqing Wang 0001, Shazia Sadiq, Xiaofang Zhou 0001 |
APWeb | 2 |
| 2014 | SharkDB: An In-Memory Column-Oriented Trajectory StorageabstractThe last decade has witnessed the prevalence of sensor and GPS technologies that produce a high volume of trajectory data representing the motion history of moving objects. However some characteristics of trajectories such as variable lengths and asynchronous sampling rates make it difficult to fit into traditional database systems that are disk-based and tuple-oriented. Motivated by the success of column store and recent development of in-memory databases, we try to explore the potential opportunities of boosting the performance of trajectory data processing by designing a novel trajectory storage within main memory. In contrast to most existing trajectory indexing methods that keep consecutive samples of the same trajectory in the same disk page, we partition the database into frames in which the positions of all moving objects at the same time instant are stored together and aligned in main memory. We found this column-wise storage to be surprisingly well suited for in-memory computing since most frames can be stored in highly compressed form, which is pivotal for increasing the memory throughput and reducing CPU-cache miss. The independence between frames also makes them natural working units when parallelizing data processing on a multi-core environment. Lastly we run a variety of common trajectory queries on both real and synthetic datasets in order to demonstrate advantages and study the limitations of our proposed storage. Haozhou Wang, Kai Zheng 0001, Jiajie Xu 0001, Bolong Zheng, Xiaofang Zhou 0001, Shazia Sadiq |
CIKM | 6 |
| 2014 | Location Oriented Phrase Detection in Microblogs
Saeid Hosseini, Sayan Unankard, Xiaofang Zhou 0001, Shazia Sadiq |
DASFAA (1) | 4 |
| 2014 | ORange: Objective-Aware Range Query RefinementabstractIn this demo paper we present Orange, a system prototype for objective-aware range query refinement. Orange essentially refines a range query to meet a pre-specified cardinality constraint while taking into account the (dis)similarity between the initial query and its corresponding refined version. To achieve this goal, Orange employes the novel scheme SAQR for efficient similarity-aware query refinement. The main idea underlying SAQR is to utilize the pre-defined constraints on cardinality and similarity in order to bound the search space and quickly find a refined query, which meets the user's expectations. We showcase Orange in a web-based application which aims to guide planners in allocating service zones for police patrol units using real and historical dataset of crime incidents. Abdullah M. Albarrak, Tatiana Noboa, Hina A. Khan, Mohamed A. Sharaf, Xiaofang Zhou 0001, Shazia Sadiq |
MDM (1) | 6 |
| 2014 | Efficient Retrieval of Top-K Most Similar Users from Travel Smart Card DataabstractUnderstanding the dynamics of human daily mobility patterns is essential for the management and planning of urban facilities and services. Travel smart cards, which record users' public transporting histories, capture rich information of users' mobility pattern. This provides the opportunity to discover valuable knowledge from these transaction records. In recent years, research on measuring user similarity for behavior analysis has attracted a lot of attention in applications such as recommendation systems, crowd behavior analysis applications, and numerous data mining tasks. In this paper, our goal is to estimate the similarity between users' travel patterns according to their travel smart card data. The core of our proposal is a novel user similarity measurement, namely, Travel Spatial-Temporal Similarity (TST), which measures the spatial range and temporal similarity between users. Moreover, we also propose a hybrid index structure, which integrates inverted files and cluster-based partitioning, to allow for efficient retrieval of the top-K most similar users. Through experimental evaluation, our proposed approach is shown to deliver scalable performance. Bolong Zheng, Kai Zheng 0001, Mohamed A. Sharaf, Xiaofang Zhou 0001, Shazia Sadiq |
MDM (1) | 5 |
| 2014 | A framework for data quality aware query systems
Naiem Khodabandehloo Yeganeh, Shazia Sadiq, Mohamed A. Sharaf |
Inf. Syst. | 2 |
| 2013 | Exploiting Structural Similarity for Automatic Information Extraction from Lists
Dat T. Huynh, Jiajie Xu 0001, Shazia Sadiq, Xiaofang Zhou 0001 |
WISE (2) | 3 |
| 2012 | A Compliance Management Ontology: Developing Shared Understanding through Models
Norris Syed Abdullah, Shazia Sadiq, Marta Indulska |
CAiSE | 2 |
| 2012 | Efficient provenance storage for relational queriesabstractProvenance information is vital in many application areas as it helps explain data lineage and derivation. However, storing fine-grained provenance information can be expensive. In this paper, we present a framework for storing provenance information relating to data derived via database queries. In particular, we first propose a provenance tree data structure which matches the query structure and thereby presents a possibility to avoid redundant storage of information regarding the derivation process. Then we investigate two approaches for reducing storage costs. The first approach utilizes two ingenious rules to achieve reduction on provenance trees. The second one is a dynamic programming solution, which provides a way of optimizing the selection of query tree nodes where provenance information should be stored. The optimization algorithm runs in polynomial time in the query size and is linear in the size of the provenance information, thus enabling provenance tracking and optimization without incurring large overheads. Experiments show that our approaches guarantee significantly lower storage costs than existing approaches. Zhifeng Bao, Henning Köhler, Liwei Wang 0011, Xiaofang Zhou 0001, Shazia Sadiq |
CIKM | 5 |
| 2012 | WebPut: Efficient Web-Based Data Imputation
Zhixu Li, Mohamed A. Sharaf, Laurianne Sitbon, Shazia Sadiq, Marta Indulska, Xiaofang Zhou 0001 |
WISE | 4 |
| 2012 | On Group Nearest Group Query ProcessingabstractGiven a data point set D, a query point set Q, and an integer k, the Group Nearest Group (GNG) query finds a subset ω (|ω| ≤ k)of points from Dsuch that the total distance from all points in Q to the nearest point in ω is not greater than any other subset ω' (|ω'| ≤ k) of points in D. GNG query is a partition-based clustering problem which can be found in many real applications and is NP-hard. In this paper, Exhaustive Hierarchical Combination (EHC) algorithm and Subset Hierarchial Refinement (SHR) algorithm are developed for GNG query processing. While EHC is capable to provide the optimal solution for k = 2, SHR is an efficient approximate approach that combines database techniques with local search heuristic. The processing focus of our approaches is on minimizing the access and evaluation of subsets of cardinality k in D since the number of such subsets is exponentially greater than |D|. To do that, the hierarchical blocks of data points at high level are used to find an intermediate solution and then refined by following the guided search direction at low level so as to prune irrelevant subsets. The comprehensive experiments on both real and synthetic data sets demonstrate the superiority of SHR in terms of efficiency and quality. Shazia Sadiq, Xiaofang Zhou 0001, Gabriel Pui Cheong Fung, Yansheng Lu |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2010 | Emerging Challenges in Information Systems Research for Regulatory Compliance Management
Norris Syed Abdullah, Shazia Sadiq, Marta Indulska |
CAiSE | 2 |
| 2010 | Active Duplicate Detection
Liwei Wang 0011, Xiaofang Zhou 0001, Shazia Sadiq, Gabriel Pui Cheong Fung |
DASFAA (1) | 4 |
| 2010 | Sampling dirty data for matching attributesabstractWe investigate the problem of creating and analyzing samples of relational databases to find relationships between string-valued attributes. Our focus is on identifying attribute pairs whose value sets overlap, a pre-condition for typical joins over such attributes. However, real-world data sets are often 'dirty', especially when integrating data from different sources. To deal with this issue, we propose new similarity measures between sets of strings, which not only consider set based similarity, but also similarity between strings instances. To make the measures effective, we develop efficient algorithms for distributed sample creation and similarity computation. Test results show that for dirty data our measures are more accurate for measuring value overlap than existing sample-based methods, but we also observe that there is a clear tradeoff between accuracy and speed. This motivates a two-stage filtering approach, with both measures operating on the same samples. Henning Köhler, Xiaofang Zhou 0001, Shazia Sadiq, Yanfeng Shu, Kerry L. Taylor |
SIGMOD Conference | 3 |
| 2009 | Processing Group Nearest Group QueryabstractGiven a data point set D, a query point set Q and an integer k, the group nearest group (GNG) query finds a subset of points from D, omega (|omega| les k), such that the total distance from all points in Q to the nearest point in omega is no greater than any other subset of points in D, omega(|omega| les k). GNG query can be found in many real applications. In this paper, exhaustive hierarchical combination algorithm (EHC) and subset hierarchical refinement algorithm (SHR) are developed for GNG query processing. The superiority of SHR in terms of efficiency and quality compared to existing algorithms developed originally for data clustering is demonstrated. Shazia Sadiq, Yansheng Lu, Gabriel Pui Cheong Fung, Heng Tao Shen |
ICDE | 3 |
| 2009 | On managing business processes variants
Ruopeng Lu, Shazia Sadiq, Guido Governatori |
Data Knowl. Eng. | 2 |
| 2009 | Instance optimal query processing in spatial networks
Xiaofang Zhou 0001, Heng Tao Shen, Shazia Sadiq, Xue Li 0001 |
VLDB J. | 4 |
| 2008 | Research and Practice in Data Quality
Shazia Sadiq, Xiaofang Zhou 0001 |
APWeb | 1 |
| 2008 | Data Quality in Web Information Systems
Xiaofang Zhou 0001, Shazia Sadiq |
WISE | 2 |
| 2007 | On the Discovery of Preferred Work Practice Through Business Process Variants
Ruopeng Lu, Shazia Sadiq |
ER | 2 |
| 2007 | On the Optimal Robot Routing Problem in Wireless Sensor NetworksabstractGiven a set of sparsely distributed sensors in the Euclidean plane, a mobile robot is required to visit all sensors to download the data and finally return to its base. The effective range of each sensor is specified by a disk, and the robot must at least reach the boundary to start communication. The primary goal of optimization in this scenario is to minimize the traveling distance by the robot. This problem can be regarded as a special case of the traveling salesman problem with neighborhoods (TSPN), which is known to be NP-hard. In this paper, we present a novel TSPN algorithm for this class of TSPN, which can yield significantly improved results compared to the latest approximation algorithm. Bo Yuan 0003, Maria E. Orlowska, Shazia Sadiq |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2006 | On Sensor Network Segmentation for Urban Water Distribution Monitoring
Sudarsanan Nesamony, Madhan Karky Vairamuthu, Maria E. Orlowska, Shazia Sadiq |
APWeb | 4 |
| 2005 | Collaborative business process technologies
Maria E. Orlowska, Shazia Sadiq |
Data Knowl. Eng. | 2 |
| 2005 | Specification and validation of process constraints for flexible workflows
Shazia Sadiq, Maria E. Orlowska, Wasim Sadiq |
Inf. Syst. | 1 |
| 2001 | Pockets of Flexibility in Workflow Specification
Shazia Sadiq, Wasim Sadiq, Maria E. Orlowska |
ER | 1 |