Hye-Young Paik

dblp:20/11 · also Helen Paik, Hye-young Paik · DBLP profile ↗
← Back
75ranked-venue papers
5as first author
32since 2021 · last 2026
0000-0003-4425-7388ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 31 · 1 first-author · 11 since 2021Databases, data management, data science and information retrieval · 28 · 3 first-author · 7 since 2021Security and privacy · 10 · 8 since 2021Artificial intelligence and machine learning · 9 · 7 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Systems, architecture and hardware · 3 · 3 since 2021Computer networks · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2026 ParaMETA: Towards Learning Disentangled Paralinguistic Speaking Styles Representations from Speech
abstract
Learning representative embeddings for different types of speaking styles, such as emotion, age, and gender, is critical for both recognition tasks (e.g., cognitive computing and human-computer interaction) and generative tasks (e.g., style-controllable speech generation). In this work, we introduce ParaMETA, a unified and flexible framework for learning and controlling speaking styles directly from speech. Unlike existing methods that rely on single-task models or cross-modal alignment, ParaMETA learns disentangled, task-specific embeddings by projecting speech into dedicated subspaces for each style type. This design reduces inter-task interference, mitigates negative transfer, and allows a single model to handle multiple paralinguistic tasks such as emotion, gender, age, and nationality classification. Beyond recognition, ParaMETA enables fine-grained style control in Text-To-Speech (TTS) generative models. It supports both speech- and text-based prompting and allows users to modify one speaking style while preserving others. Extensive experiments demonstrate that ParaMETA outperforms strong baselines in classification accuracy and generates more natural and expressive speech, while maintaining a lightweight and efficient model suitable for real-world applications.
Haowei Lou, Hye-Young Paik, Wen Hu 0001, Lina Yao 0001
AAAI2
2026 AMMs in Tokenized Real-World Asset Markets: A Market Functionality and Sustainability Assessment
Zhonghao Liu, H. M. N. Dilum Bandara, Hye-Young Paik
ICBC3
2026 Zero trust-driven access control delegation using blockchain
abstract
As digital ecosystems become more complex with decentralized technologies like the Internet of Things (IoT) and blockchain, traditional access control models fail to meet the security needs of dynamic, high-risk environments. The need for dynamic, fine-grained access control mechanisms has become critical, particularly in environments where trust must be continuously evaluated, and access decisions must adapt to real-time conditions. Traditional models often rely on static identity management and centralized trust assumptions, which are inadequate for modern, decentralized, and highly dynamic environments such as IoT ecosystems. Consequently, existing solutions lack fine-grained identity management, flexible delegation, and continuous trust evaluation, highlighting the need for a more robust, adaptive, and decentralized access control architecture. To address these gaps, this paper presents a novel access control architecture that integrates self-sovereign identity (SSI) and decentralized identifier (DID)-based access control with zero trust principles, enhanced by a flexible capability-based access control (CapBAC) approach. Leveraging SSI and DID allows entities to manage their identities without relying on a central authority, aligning with zero-trust principles. The integration of CapBAC ensures flexible, context-aware, and attribute-based access control, where access rights are dynamically granted based on the requester's capabilities. This enables fine-grained delegation of access rights, allowing trusted entities to delegate specific privileges to others without compromising overall security. Continuous trust evaluation is employed to assess the authenticity of access requests, mitigating the risks posed by compromised devices or users. The proposed architecture also incorporates blockchain technology to ensure transparent, immutable, and secure management of access logs, providing traceability and accountability for all access events. We demonstrate the feasibility and effectiveness of this solution through performance evaluations and comparisons with existing access control schemes, showing its superior security, scalability, and adaptability in real-world scenarios. Our work demonstrates a comprehensive, decentralized, and scalable solution for secure access control delegation using zero trust-driven principles.
Rahma Mukta, Shantanu Pal, Kowshik Chowdhury, Michael Hitchens, Hye-Young Paik, Salil S. Kanhere
Blockchain Res. Appl.5
2025 ParaStyleTTS: Toward Efficient and Robust Paralinguistic Style Control for Expressive Text-to-Speech Generation
abstract
Controlling speaking style in text-to-speech (TTS) systems has become a growing focus in both academia and industry. While many existing approaches rely on reference audio to guide style generation, such methods are often impractical due to privacy concerns and limited accessibility. More recently, large language models (LLMs) have been used to control speaking style through natural language prompts; however, their high computational cost, lack of interpretability, and sensitivity to prompt phrasing limit their applicability in real-time and resource-constrained environments. In this work, we propose ParaStyleTTS, a lightweight and interpretable TTS framework that enables expressive style control from text prompts alone. ParaStyleTTS features a novel two-level style adaptation architecture that separates prosodic and paralinguistic speech style modeling. It allows fine-grained and robust control over factors such as emotion, gender, and age. Unlike LLM-based methods, ParaStyleTTS maintains consistent style realization across varied prompt formulations and is well-suited for real-world applications, including on-device and low-resource deployment. Experimental results show that ParaStyleTTS generates high-quality speech with performance comparable to state-of-the-art LLM-based systems while being 30x faster, using 8x fewer parameters, and requiring 2.5x less CUDA memory. Moreover, ParaStyleTTS exhibits superior robustness and controllability over paralinguistic speaking styles, providing a practical and efficient solution for style-controllable text-to-speech generation. Demo can be found at https://parastyletts.github.io/ParaStyleTTS_Demo/. Code can be found at https://github.com/haoweilou/ParaStyleTTS.
Haowei Lou, Hye-Young Paik, Wen Hu 0001, Lina Yao 0001
CIKM2
2025 Achoio: A Skill-Aware Evaluation Management System for Text-To-Speech Research
abstract
Human subjective evaluation plays a crucial role in evaluating speech-related generative tasks such as text-to-speech (TTS) generation. However, current practices are often constrained by limited scalability, fragmented workflows, and inconsistent rating reliability. Researchers frequently rely on manual methods or general-purpose crowdsourcing systems, where recruiting appropriately skilled listeners is challenging, and result analysis is labor-intensive. In this work, we introduce Achoio, a dedicated end-to-end online system designed to streamline and scale human evaluation for the TTS research community. Achoio allows researchers to create and manage evaluation projects, upload synthesized speech samples, and automatically match them with qualified listeners based on linguistic proficiency and domain knowledge. The system provides built-in tools for project status tracking, result aggregation and visualization. In this demonstration, we will walk through the core features of Achoio, including intuitive project setup, skill-based listener matching algorithm, and automated analytics. By addressing the limitations of existing workflows, Achoio offers a scalable, domain-aware, and analysis-ready solution for conducting high-quality subjective TTS evaluations. Our system is live and can be found at https://www.achoio.com. Demo is available on YouTube at https://youtu.be/Ugjj3_YooSM.
Haowei Lou, Hye-Young Paik, Basem Suleiman, Wen Hu 0001, Lina Yao 0001
CIKM2
2025 MetaLogic: Robustness Evaluation of Text-to-Image Models via Logically Equivalent Prompts
Yangyang Shu, Hye-Young Paik, Yulei Sui
ICFEM3
2025 FedSIG: Privacy-Preserving Federated Recommendation via Synthetic Interaction Generation
abstract
Recommendation Systems (RS) play an important role in our everyday life in this data-driven digital era by providing users with the convenience of navigating the plethora of available choices. An RS collects user behavioural data to provide them with valuable suggestions. The growing privacy concerns regarding private data collection have led to the use of Federated Learning (FL) to implement RS. However, many research works have exposed the privacy leakages in FL gradient sharing. The embedding gradients shared by FL users during the RS model training can be used to infer the items that users have interacted with. Existing defences, such as random noise injection or pseudo-interaction sampling to obfuscate the privacysensitive information reflected by the shared gradients. However, these techniques provide limited protection and often result in substantial degradation of recommendation performance, leading to an unfavourable privacy–utility trade-off. In this paper, we propose FedSIG (Federated Synthetic Interaction Generation), a defence mechanism that mitigates useritem interaction inference in federated recommendation systems by generating synthetic interaction data using generative models. The generated items are selectively used to replace or augment real user interactions, thereby obfuscating sensitive data while preserving user preference signals. To further enhance utility, we design an item selection module based on an attention mechanism to identify less contributive interactions for replacement. Extensive experiments conducted on five real-world datasets and two state-of-the-art recommendation models demonstrate that FedSIG achieves a significantly improved privacy–utility balance compared to existing approaches, effectively reducing inference success rates while maintaining competitive recommendation accuracy.
Thirasara Ariyarathna, Salil S. Kanhere, Meisam Mohammady, Hye-Young Paik
RAID4
2025 LatentSpeech: Latent Diffusion for Text-To-Speech Generation
abstract
Text-To-Speech (TTS) generation plays a crucial role in human-robot interaction by allowing robots to communicate naturally with humans. Researchers have developed various TTS models to enhance speech generation. More recently, diffusion models have emerged as a powerful generative framework, achieving state-of-the-art performance in tasks such as image and video generation. However, their application in TTS has been limited by its slow inference speeds due to their iterative denoising process. Previous work has applied diffusion models to Mel-Spectrograms with an additional vocoder to convert them into waveforms. To address these limitations, we propose LatentSpeech, a novel diffusion-based TTS framework that operates directly in a latent space. This space is significantly more compact and information-rich than raw Mel-Spectrograms. Furthermore, we introduce an alternative latent space of Pseudo-Quadrature Mirror Filters (PQMF), which decomposes speech into multiple subbands. By leveraging PQMF’s near-perfect waveform reconstruction capability, LatentSpeech eliminates the need for a separate vocoder and reduces both model size and inference time. Our PQMF-based LatentSpeech model reduces inference time by 45% and model size by 77% compared to Mel-Spectrogram diffusion models. On benchmark datasets, it achieves 25% lower WER and 58% higher MOS using the same training data. These results highlight LatentSpeech as an efficient, high-quality TTS solution for real-time and human-robot interaction. Code and models are available here.
Haowei Lou, Hye-Young Paik, Pari Delir Haghighi, Sheng Li 0010, Wen Hu 0001, Lina Yao 0001
RO-MAN2
2025 DeepSneak: User GPS Trajectory Reconstruction from Federated Route Recommendation Models
abstract
Decentralized machine learning, such as Federated Learning (FL), is widely adopted in many application domains. Especially in domains like recommendation systems, sharing gradients instead of private data has recently caught the research community’s attention. Personalized travel route recommendation utilizes users’ location data to recommend optimal travel routes. Location data is extremely privacy sensitive, presenting increased risks of exposing behavioral patterns and demographic attributes. FL for route recommendation can mitigate the sharing of location data. However, this article shows that an adversary can recover the user trajectories used to train the federated recommendation models with high proximity accuracy. To this effect, we propose a novel attack called DeepSneak, which uses shared gradients obtained from global model training in FL to reconstruct private user trajectories. We formulate the attack as a regression problem and train a generative model by minimizing the distance between gradients. We validate the success of DeepSneak on two real-world trajectory datasets. The results show that we can recover the location trajectories of users with reasonable spatial and semantic accuracy.
Thirasara Ariyarathna, Meisam Mohommady, Hye-Young Paik, Salil S. Kanhere
ACM Trans. Intell. Syst. Technol.3
2024 VLIA: Navigating Shadows with Proximity for Highly Accurate Visited Location Inference Attack against Federated Recommendation Models
abstract
Personalized location recommendation allows users to enjoy a seamless travel experience by suggesting the optimal travel locations/routes based on user preferences. Most service providers collect users' location data centrally to develop accurate route recommendation applications. Federated learning (FL) can be used as an inherent privacy-preserving mechanism in these applications to prevent users from sharing private data. However, recent research shows that FL is still vulnerable to privacy leakages. Therefore, many FL-based recommendation systems use Local Differential Privacy (LDP) to defend against such attacks. In this paper, we propose the Visited Location Inference Attack (VLIA), a novel attack for federated location recommendation systems through the lens of Membership Inference Attack (MIA). Specifically, we focus on inferring user behaviour data (visited locations) even when the federated recommendation system is protected with LDP. We design and implement VLIA leveraging both embedding and proximity information of locations, making the inference more accurate. Our extensive experiments with two state-of-the-art personalized route recommendation (PRR) systems implemented in the FL setting and two real-world trajectory datasets showcase the effectiveness of the VLIA attack. Our results show that LDP cannot defend VLIA unless the recommendation performance is significantly compromised.
Thirasara Ariyarathna, Meisam Mohammady, Hye-Young Paik, Salil S. Kanhere
AsiaCCS3
2024 CredAct: Privacy-Preserving Activity Verification for Benefits Schemes in Self-Sovereign Identity
abstract
We propose CredAct, a user activity verification designed with data minimisation to protect privacy. Many Benefits Schemes, such as discount offers, loyalty programs, and incentive systems, require verification of user activity (e.g., buying healthy food, step counts) in their business processes. These service providers can collect a large amount of users’ personal information, and often users do not have fine-grained control over the scope of data disclosure. In CredAct, we propose a Self-Sovereign Identity based framework implemented on blockchain that enables users participating in a benefits scheme to minimise data sharing during the submission and verification of data. We use a smart contract-based function along with a Zero-Knowledge Proof cryptographic commitment scheme, that forces the entities involved in the business process to collect or disclose only the required (minimum) data to fulfill the intended purpose. The evaluation shows that the system is feasible with minimal operational overheads compared to traditional cryptographic techniques. We also perform a qualitative privacy and security analysis considering relevant threats to CredAct.
Rahma Mukta, Hye-Young Paik, Qinghua Lu 0001, Salil S. Kanhere
ICBC2
2024 Prompt Perturbation in Retrieval-Augmented Generation based Large Language Models
abstract
The robustness of large language models (LLMs) becomes increasingly important as their use rapidly grows in a wide range of domains.Retrieval-Augmented Generation (RAG) is considered as a means to improve the trustworthiness of text generation from LLMs.However, how the outputs from RAG-based LLMs are affected by slightly different inputs is not well studied.In this work, we find that the insertion of even a short prefix to the prompt leads to the generation of outputs far away from factually correct answers.We systematically evaluate the effect of such prefixes on RAG by introducing a novel optimization technique called Gradient Guided Prompt Perturbation (GGPP).GGPP achieves a high success rate in steering outputs of RAG-based LLMs to targeted wrong answers.It can also cope with instructions in the prompts requesting to ignore irrelevant context.We also exploit LLMs' neuron activation difference between prompts with and without GGPP perturbations to give a method that improves the robustness of RAG-based LLMs through a highly effective detector trained on neuron activation triggered by GGPP generated prompts.Our evaluation on open-sourced LLMs demonstrates the effectiveness of our methods.
Zhibo Hu, Chen Wang 0008, Yanfeng Shu, Hye-Young Paik, Liming Zhu 0001
KDD4
2024 StyleSpeech: Parameter-efficient Fine Tuning for Pre-trained Controllable Text-to-Speech
Haowei Lou, Hye-Young Paik, Wen Hu 0001, Lina Yao 0001
MMAsia2
2024 Establishing a Data-Efficient Witness Protocol for Connected Autonomous Vehicles
Siriboon Chaisawat, Hye-Young Paik, Salil S. Kanhere
MobiQuitous2
2024 SoK: Trusting Self-Sovereign Identity
abstract
Digital identity is evolving from centralized systems to a decentralized approach known as Self-Sovereign Identity (SSI). SSI empowers individuals to control their digital identities, eliminating reliance on third-party data custodians and reducing the risk of data breaches. However, the concept of trust in SSI remains complex and fragmented. This paper systematically analyzes trust in SSI in light of its components and threats posed by various actors in the system. As a result, we derive three distinct trust models that capture the threats and mitigations identified across SSI literature and implementations. Our work provides a foundational framework for future SSI research and development, including a comprehensive catalogue of SSI components and design requirements for trust, shortcomings in existing SSI systems and areas for further exploration.
Evan Krul, Hye-Young Paik, Sushmita Ruj, Salil S. Kanhere
Proc. Priv. Enhancing Technol.2
2023 SID: Service Identification and Discovery Framework for Decentralised Supply Chains
abstract
The traditional decentralised supply chains are facing challenges to meet the increasing demands for openness, transparency, trust and efficiency. As a result, blockchain-based decentralised supply chains are emerging. However, how to identify and discover various services in a decentralised supply chain remains an unsolved issue. This paper presents a new framework for modelling and discovering services in decentralised supply chain systems. In the framework, W3C Decentralised Identifier (DID) is adopted to describe service attributes to meet different business and technical requirements for a supply chain service. While blockchain is used for publishing and storing DID, a graphic-database-based service repository is proposed to organize these DID strings in a natural way as a graph for better service discovery. An Ethereum-based prototype is implemented as a proof of concept to demonstrate its feasibility and usefulness.
Hye-Young Paik, Qin Wang 0008, Shining Chen
SSE2
2023 Efficient Credential Revocation Using Cryptographic Accumulators
abstract
Credentials must be revoked on time to avoid legal issues, information misuse, and potential economic damage. A cryptographic accumulator is the most efficient data structure to record whether a credential is valid. However, credential holders must update their verification proof, or witness whenever a credential is revoked in the system. The witness update is expensive and not feasible for applications with limited computational power, such as mobile phones. We propose a credential revocation management framework that removes the need for witness updates, while still utilising the benefits of the cryptographic accumulator. We use three accumulators, the notion of epoch and replication of accumulator transition from one state to another for verification. The evaluation shows our method is efficient and suitable for credential holders with limited computing resources.
Daria Schumm, Rahma Mukta, Hye-Young Paik
ICBC3
2023 A Selection Model of Privacy Patterns
abstract
Privacy has become an increasingly essential quality to consider in a software system. Privacy practices should be adopted from the first stage of software design to safeguard personal data from unintentional information leakage. Privacy patterns have been investigated by academic and industry practitioners to address different privacy issues. The presence of these patterns is both helpful and challenging for the designer. On one hand, the privacy patterns are valuable as reusable design solutions to solve common privacy problems. On the other hand, the multitude of privacy patterns makes the choice of patterns as solutions difficult for the designer. In this paper, we propose a selection model that can assist architects in deciding suitable patterns for a software system. The selection is based on the regulatory entities and architectural characteristics implicit in the patterns. We evaluate the proposed selection model through case studies and interviews with practitioners. Our evaluation accesses the applicability and usefulness of the selection model in guiding the pattern selection for architectural design and understanding the rationale of different design decisions.
Su Yen Chia, Xiwei Xu 0001, Ming Ding 0001, David B. Smith 0001, Hye-Young Paik, Liming Zhu 0001
ICSA5
2023 A Pattern-Oriented Reference Architecture for Governance-Driven Blockchain Systems
abstract
Blockchain technology has been integrated into diverse software applications by enabling a decentralised architecture design. However, the defects of on-chain algorithmic mechanisms, and tedious disputes and debates in off-chain communities may affect the operation of blockchain systems. Accordingly, blockchain governance has received great interest for supporting the design, use, and maintenance of blockchain systems, hence improving the overall trustworthiness. Although much effort has been put into this research topic, there is a distinct lack of consideration for blockchain governance from the perspective of software architecture design. In this study, we propose a pattern-oriented reference architecture for governance-driven blockchain systems, which can provide guidance for future blockchain architecture design. We design the reference architecture based on an extensive review of architectural patterns for blockchain governance in academic literature and industry implementation. The reference architecture consists of four layers. We demonstrate the components in each layer, annotating with the identified patterns. A qualitative analysis of mapping two concrete blockchain architectures, Polkadot and Quorum, on the reference architecture is conducted, to evaluate the correctness and utility of proposed reference architecture.
Yue Liu 0010, Qinghua Lu 0001, Guangsheng Yu, Hye-Young Paik, Liming Zhu 0001
ICSA4
2023 A Blockchain-Based Interoperable Architecture for IoT with Selective Disclosure of Information
abstract
With the improvement of Internet of Things (IoT) technologies, services, and applications, there is a proliferation of access to smart devices in everyday life. However, granting access and controlling access rights for each resource is challenging in highly dynamic and large-scale IoT deployments. In particular, multiple access information may need to be provided to an entity when granting access rights to several resources. The situation becomes more complex when an entity is required to share its identity attribute to receive the access information. These raise the question of what identity information an entity needs to provide to obtain the required access to a particular resource and, subsequently, what access information needs to be provided when accessing that resource. That said, there is a need for a flexible approach where an entity can share a distinct identity and access attributes for accessing a resource without revealing additional information. Such flexibility in sharing information is significant given the privacy risk of an entity’s identity. This paper presents an architecture that delivers access rights to an entity with selective disclosure of information. Our approach ensures the minimum exchange of information (identity and access attribute) to enhance an entity’s privacy when granting access rights to an entity. We use blockchain to provide data authenticity (i.e., tamper-proof), transparency and automatic execution of access rights based on shared attributes using smart contracts. We implement a proof of concept of the proposed system using Hyperledger fabric as a permissioned blockchain network. Our results demonstrate the feasibility of the proposed system showing efficiency in granting access rights.
Rahma Mukta, Shantanu Pal, Hye-Young Paik, Salil S. Kanhere, Michael Hitchens
PRDC4
2023 Toward Trustworthy AI: Blockchain-Based Architecture Design for Accountability and Fairness of Federated Learning Systems
abstract
Federated learning is an emerging privacy-preserving AI technique where clients (i.e., organizations or devices) train models locally and formulate a global model based on the local model updates without transferring local data externally. However, federated learning systems struggle to achieve trustworthiness and embody responsible AI principles. In particular, federated learning systems face accountability and fairness challenges due to multistakeholder involvement and heterogeneity in client data distribution. To enhance the accountability and fairness of federated learning systems, we present a blockchain-based trustworthy federated learning architecture. We first design a smart contract-based data-model provenance registry to enable accountability. Additionally, we propose a weighted fair data sampler algorithm to enhance fairness in training data. We evaluate the proposed approach using a COVID-19 X-ray detection use case. The evaluation results show that the approach is feasible to enable accountability and improve fairness. The proposed algorithm can achieve better performance than the default federated learning setting in terms of the model’s generalization and accuracy.
Sin Kit Lo, Yue Liu 0010, Qinghua Lu 0001, Chen Wang 0008, Xiwei Xu 0001, Hye-Young Paik, Liming Zhu 0001
IEEE Internet Things J.6
2023 A systematic literature review on blockchain governance
Yue Liu 0010, Qinghua Lu 0001, Liming Zhu 0001, Hye-Young Paik, Mark Staples
J. Syst. Softw.4
2023 Uncertainty Estimation With Neural Processes for Meta-Continual Learning
abstract
The ability to evaluate uncertainties in evolving data streams has become equally, if not more, crucial than building a static predictor. For instance, during the pandemic, a model should consider possible uncertainties such as governmental policies, meteorological features, and vaccination schedules. Neural process families (NPFs) have recently shone a light on predicting such uncertainties by bridging Gaussian processes (GPs) and neural networks (NNs). Their abilities to output average predictions and the acceptable variances, i.e., uncertainties, made them suitable for predictions with insufficient data, such as meta-learning or few-shot learning. However, existing models have not addressed continual learning which imposes a stricter constraint on the data access. Regarding this, we introduce a member meta-continual learning with neural process (MCLNP) for uncertainty estimation. We enable two levels of uncertainty estimations: the local uncertainty on certain points and the global uncertainty p(z) that represents the function evolution in dynamic environments. To facilitate continual learning, we hypothesize that the previous knowledge can be applied to the current task, hence adopt a coreset as a memory buffer to alleviate catastrophic forgetting. The relationships between the degree of global uncertainties with the intratask diversity and model complexity are discussed. We have estimated prediction uncertainties with multiple evolving types including abrupt/gradual/recurrent shifts. The applications encompass meta-continual learning in the 1-D, 2-D datasets, and a novel spatial-temporal COVID dataset. The results show that our method outperforms the baselines on the likelihood and can rebound quickly even for heavily evolved data streams.
Xuesong Wang 0002, Lina Yao 0001, Xianzhi Wang 0001, Hye-Young Paik, Sen Wang 0001
IEEE Trans. Neural Networks Learn. Syst.4
2022 Profiler: Distributed Model to Detect Phishing
abstract
Many Machine Learning (ML) based phishing detection algorithms are not adept to recognise "concept drift"; attackers introduce small changes in the statistical characteristics of their phishing attempts to successfully bypass detection. This leads to the classification problem of frequent false positives and false negatives, and a reliance on manual reporting of phishing by users. Profiler is a distributed phishing risk assessment tool that combines three email profiling dimensions: (1) threat level, (2) cognitive manipulation, and (3) email content type to detect email phishing. Unlike pure ML-based approaches, Profiler does not require large data sets to be effective and evaluations on real-world data sets show that it can be useful in conjunction with ML algorithms to mitigate the impact of concept drift.
Mariya Shmalko, Alsharif Abuadbba, Raj Gaire 0001, Tingmin Wu, Hye-Young Paik, Surya Nepal
ICDCS5
2022 A survey of data minimisation techniques in blockchain-based healthcare
Rahma Mukta, Hye-Young Paik, Qinghua Lu 0001, Salil S. Kanhere
Comput. Networks2
2022 Defining blockchain governance principles: A comprehensive framework
Yue Liu 0010, Qinghua Lu 0001, Guangsheng Yu, Hye-Young Paik, Liming Zhu 0001
Inf. Syst.4
2022 Chain or DAG? Underlying data structures, architectures, topologies and consensus in distributed ledger technology: A review, taxonomy and research issues
Huanyu Wu, Chentao Yue, Hye-Young Paik, Salil S. Kanhere
J. Syst. Archit.4
2022 Architectural patterns for the design of federated learning systems
Sin Kit Lo, Qinghua Lu 0001, Liming Zhu 0001, Hye-Young Paik, Xiwei Xu 0001, Chen Wang 0008
J. Syst. Softw.4
2021 FLRA: A Reference Architecture for Federated Learning Systems
Sin Kit Lo, Qinghua Lu 0001, Hye-Young Paik, Liming Zhu 0001
ECSA3
2021 Global Convolutional Neural Processes
abstract
The ability to deal with uncertainty in machine learning models has become equally, if not more, crucial to their predictive ability itself. For instance, during the pandemic, governmental policies and personal decisions are constantly made around uncertainties. Targeting this, Neural Process Families (NPFs) have recently shone a light on prediction with uncertainties by bridging Gaussian processes and neural networks. Latent neural process, a member of NPF, is believed to be capable of modelling the uncertainty on certain points (local uncertainty) as well as the general function priors (global uncertainties). Nonetheless, some critical questions remain unresolved, such as a formal definition of global uncertainties, the causality behind global uncertainties, and the manipulation of global uncertainties for generative models. Regarding this, we build a member GloBal Convolutional Neural Process(GBCoNP) that achieves the SOTA log-likelihood in latent NPFs. It designs a global uncertainty representation p(z), which is an aggregation on a discretized input space. The causal effect between the degree of global uncertainty and the intra-task diversity is discussed. The learnt prior is analyzed on a variety of scenarios, including 1D, 2D, and a newly proposed spatial-temporal COVID dataset. Our manipulation of the global uncertainty not only achieves generating the desired samples to tackle few-shot learning, but also enables the probability evaluation on the functional priors.
Xuesong Wang 0002, Lina Yao 0001, Xianzhi Wang 0001, Hye-Young Paik, Sen Wang 0001
ICDM4
2021 MEChain: A Multi-layer Blockchain Structure with Hierarchical Consensus for Secure EHR System
abstract
Although Electronic Health Record (EHR) systems are widely used in health care organisations, they still face security and privacy problems. Blockchain is considered a promising technology that could help overcome many of these problems. However, efforts to incorporate blockchain into EHR systems so far have shown heavy overhead, unsuitable software architectural structures and inefficient consensus protocols. We present ME Chain, a multi-layer blockchain structure, which aims to solve the adoption, storage and consensus problems when implementing blockchain in EHR systems. The multi-layer structure is better suited for the operational hierarchy in health organisations and is supported by the following novel features: (i) a consensus implementation optimised for the multi-layer operations, which solves the byzantine failure in both layers, hence guarantees consistency, it also reduces request confirmation latency and overhead and provides every layer with the ability to correct the faults, (ii) a data synchronisation method for some nodes that does not receive the correct information or simply get offline to catch up with the system, data verification and retrieve mechanisms to protect data integrity. Experiments are conducted to evaluate ME Chain in terms of security, performance and storage cost. The results show that ME Chain can establish a secure EHR with high performance and acceptable storage cost.
Huanyu Wu, Lunjie Li, Hye-Young Paik, Salil S. Kanhere
TrustCom3
2021 DR-BFT: A consensus algorithm for blockchain-based multi-layer data integrity framework in dynamic edge computing system
Yuqi Fan 0001, Huanyu Wu, Hye-Young Paik
Future Gener. Comput. Syst.3
2020 Blockchain-based Verifiable Credential Sharing with Selective Disclosure
abstract
Sharing credentials could raise privacy concerns. For digital credentials to be widely accepted, there is a need for an end-to-end system that provides (i) secure verification of the participant identities and credentials to increase trust, and (ii) a data minimisation mechanism to reduce the risk of oversharing the credential data. This paper proposes CredChain, a blockchain-based Self-Sovereign Identity (SSI) platform architecture that allows secure creation, sharing and verification of credentials. Beyond the verification of identities and credentials, a flexible selective disclosure solution is proposed using redactable signatures. The credentials are managed through a decentralised application/wallet which allows users to store their credential data privately under their full control and re-use as necessary. Our evaluation results show that CredChain architecture is feasible, secure and exhibits the level of performance that is within the expected benchmarks of the well-known blockchain platform, Parity Ethereum.
Rahma Mukta, James Martens, Hye-Young Paik, Qinghua Lu 0001, Salil S. Kanhere
TrustCom3
2019 Optimising Architectures for Performance, Cost, and Security
Rajitha Yasaweerasinghelage, Mark Staples, Hye-Young Paik, Ingo Weber
ECSA3
2019 A Case Based Deep Neural Network Interpretability Framework and Its User Study
Rimmal Nadeem, Huijun Wu 0001, Hye-Young Paik, Chen Wang 0008
WISE3
2019 TEXUS: A unified framework for extracting and understanding tables in PDF documents
Roya Rastan, Hye-Young Paik, John Shepherd 0001
Inf. Process. Manag.2
2019 Business process improvement with the AB-BPM methodology
Suhrid Satyal, Ingo Weber, Hye-Young Paik, Claudio Di Ciccio, Jan Mendling
Inf. Syst.3
2018 AB Testing for Process Versions with Contextual Multi-armed Bandit Algorithms
Suhrid Satyal, Ingo Weber, Hye-Young Paik, Claudio Di Ciccio, Jan Mendling
CAiSE3
2018 Predicting the Performance of Privacy-Preserving Data Analytics Using Architecture Modelling and Simulation
abstract
Privacy-preserving data analytics is an emerging technology which allows multiple parties to perform joint data analytics without disclosing source data to each other or a trusted third-party. A variety of platforms and protocols have been proposed in this domain. However, these systems are not yet widely used, and little is known about them from a software architecture and performance perspective. Here we investigate the feasibility of using architectural performance modelling and simulation tools for predicting the performance of privacy-preserving data analytics systems. We report on a lab-based experimental study of a privacy-preserving credit scoring application that uses an implementation of a partial homomorphic encryption scheme. The main experiments are on the impact of analytic problem size (number of data items and number of features), and cryptographic key length for the overall system performance. Our modelling approach performed with a relative error consistently under 5\% when predicting the median learning time for the scoring application. We find that the use of this approach is feasible in this technology domain, and discuss how it can support architectural decision making on trade-offs between properties such as performance, cost, and security. We expect this to enable the evaluation and optimisation of privacy-preserving data analytics technologies.
Rajitha Yasaweerasinghelage, Mark Staples, Ingo Weber, Hye-Young Paik
ICSA4
2017 AB-BPM: Performance-Driven Instance Routing for Business Process Improvement
Suhrid Satyal, Ingo Weber, Hye-Young Paik, Claudio Di Ciccio, Jan Mendling
BPM3
2017 Multi-Level Privacy-Preserving Access Control as a Service for Personal Healthcare Monitoring
abstract
The Internet of Things (IoT) changes many sectors of our lives. In the healthcare domain, IoT presents as mobile medical applications over various sensors that update healthcare professionals on patients' health information. However, IoT-based healthcare systems also face major challenges in protecting patients' privacy via an effective access control system. This paper presents an ambient home solution framework for privacy-preserving monitoring of patients' health status. We focus on two major points: 1) how to use the data collected from ambient and biometric sensors, to perform the high-level task of activity recognition, and 2) how to secure the collected healthcare data via effective access control. We achieve multi-level access control by using Public Key Infrastructure (PKI) for authentication and Attribute-Based Access Control (ABAC) for authorisation. Our access control system regulates access to healthcare data by classification over healthcare professionals and data. Our system provides guidelines to define data classes and healthcare professional groups and specifies security policies to control access to the data classes. The system is flexible and can incorporate more policy rules, professionals, and data classes.
Usama Salama, Lina Yao 0001, Xianzhi Wang 0001, Hye-Young Paik, Amin Beheshti
ICWS4
2016 Automated Table Understanding Using Stub Patterns
Roya Rastan, Hye-Young Paik, John Shepherd 0001, Armin Haller
DASFAA (1)2
2016 Aggregated Search over Personal Process Description Graph
Jing Ouyang Hsu, Hye-Young Paik, Liming Zhan, Anne H. H. Ngu
DEXA (2)2
2016 A PDF Wrapper for Table Processing
abstract
We propose a PDF document wrapper system that is specifically targeted at table processing applications. We (i) review the PDF specifications and identify particular challenges from the table processing point of view, (ii) specify a table-oriented document model containing the required atomic elements for table extraction and understanding applications. Our evaluation showed that the wrapper was able to detect important features such as page columns, bullets and numbering in all measures, recording over 90% accuracy, leading to better table locating and segmenting.
Roya Rastan, Hye-Young Paik, John Shepherd 0001
DocEng2
2016 Towards a Common Understanding of Business Process Instance Data
abstract
In an organisation several Business Process Management System (BPMS) products can co-exist and work alongside each other. Each one of these BPM tools has its own definition of process instances, creating a heterogeneous environment. This reduces interoperability between business process management systems and increases the effort involved in analysing the data. In this paper, we propose a common model for business process instances, named Business Process Instance Model (BPIM), which provides a holistic view of business process instances generated from multiple systems. BPIM consists of visual notations and their metadata schema. It captures three dimensions of process instances: process execution paths, instance data provenance and meta-data. BPIM aims to provide an abstract layer between the process instance repository and BPM engines, leading to common understanding of business process instances.
Nima N. Moghadam, Hye-Young Paik
MODELSWARD2
2016 Building a Process Description Repository with Knowledge Acquisition
Diyin Zhou, Hye-Young Paik, Seung Hwan Ryu, John Shepherd 0001, Paul Compton
PKAW2
2015 TEXUS: A Task-based Approach for Table Extraction and Understanding
abstract
In this paper, we propose a precise, comprehensive model of table processing which aims to remedy some of the problems in the discussion of table processing in the literature. The model targets application-independent, end-to-end table processing, and thus encompasses a large subset of the work in the area. The model can be used to aid the design of table processing systems (We provide an example of such a system), can be considered as a reference framework for evaluating the performance of table processing systems, and can assist in clarifying terminological differences in the table processing literature.
Roya Rastan, Hye-Young Paik, John Shepherd 0001
DocEng2
2015 7th International Workshop on Principles of Engineering Service-Oriented and Cloud Systems (PESOS 2015)
abstract
PESOS has established itself as a forum that brings together software engineering researchers and practitioners working in the areas of service-oriented systems to discuss research challenges, new developments and applications, as well as methods, techniques, experiences, and tools to support engineering, evolution and adaptation of service-oriented systems. The technical advances and growing adoption of Cloud computing is creating new challenges for the PESOS the software services community to explore the approaches to better engineer software systems that are designed, developed, operated and governed in the context of the Cloud. We again attracted high-quality submissions on a diverse set of relevant topics such as better approaches to engineering service-based collaborative systems, Infrastructure as a Service (IaaS), Platform as a Service (PaaS), and Software as a Service (SaaS) models of cloud computing and associated software quality attributes. PESOS 2015 will continue to be the key forum for collecting case studies and artifacts for educators and researchers in this area.
Muhammad Ali Babar 0001, Hye-Young Paik, Malolan Chetlur, Amir Molzam Sharifloo
ICSE (2)2
2015 Similarity Search over Personal Process Description Graph
Jing Ouyang Hsu, Hye-Young Paik, Liming Zhan
WISE (1)2
2013 5th international workshop on principles of engineering service-oriented systems (PESOS 2013)
abstract
PESOS 2013 is a forum that brings together software engineering researchers from academia and industry, as well as practitioners working in the areas of service-oriented systems to discuss research challenges, recent developments, novel application scenarios, as well as methods, techniques, experiences, and tools to support engineering, evolution and adaptation of service-oriented systems. The special theme of the 5th edition of PESOS is “Service Engineering for the Cloud” The goal is to explore approaches to better engineer service-oriented systems, to either take advantage of the qualities offered by cloud infrastructures or to account for lack of full control over important quality attributes. PESOS 2013 also continues to be the key forum for collecting case studies and artifacts for educators and researchers in this area.
Domenico Bianculli, Patricia Lago, Grace A. Lewis, Hye-Young Paik
ICSE4
2013 Form-Based Web Service Composition for Domain Experts
abstract
In many cases, it is not cost effective to automate business processes which affect a small number of people and/or change frequently. We present a novel approach for enabling domain experts to model and deploy such processes from their respective domain as Web service compositions. The approach builds on user-editable service, naming and representing Web services as forms. On this basis, the approach provides a visual composition language with a targeted restriction of control-flow expressivity, process simulation, automated process verification mechanisms, and code generation for executing orchestrations. A Web-based service composition prototype implements this approach, including a WS-BPEL code generator. A small lab user study with 14 participants showed promising results for the usability of the system, even for nontechnical domain experts.
Ingo Weber, Hye-Young Paik, Boualem Benatallah
ACM Trans. Web2
2012 Automating Form-Based Processes through Annotation
Sung Wook Kim, Hye-Young Paik, Ingo Weber
ICSOC2
2011 Similarity Function Recommender Service Using Incremental User Knowledge Acquisition
Seung Hwan Ryu, Boualem Benatallah, Hye-Young Paik, Yang Sok Kim, Paul Compton
ICSOC3
2011 Forms-based Service Composition
Ingo Weber, Hye-Young Paik, Boualem Benatallah
ICSOC2
2011 MashSheet: Mashups in Your Spreadsheet
Dat Dac Hoang, Hye-Young Paik
WISE2
2010 Spreadsheet as a Generic Purpose Mashup Development Environment
Dat Dac Hoang, Hye-Young Paik, Anne H. H. Ngu
ICSOC2
2010 Managing Long-Tail Processes Using FormSys
Ingo Weber, Hye-Young Paik, Boualem Benatallah, Corren Vorwerk, Zifei Gong, Liangliang Zheng, Sung Wook Kim
ICSOC2
2010 FormSys: form-processing web services
abstract
In this paper we present FormSys, a Web-based system that service-enables form documents. It offers two main services: filling in forms based on Web services' incoming SOAP messages, and invoking Web services based on filled-in forms. This can be applied to benefit individuals to reduce the number of often repetitive form fields they have to complete manually in many scenarios. It can also help organisations to remove the need for manual data entry by automatically triggering business process implementations based on incoming case data from filled-in forms. While the concept applies to forms of any type of document, our implementation uses Adobe AcroForms due to its universal applicability, availability of a usable API, and end-user appeal. In the demo, we will show the two core functions, namely soap2pdf and pdf2soap, along with use case applications of the services developed from real world scenarios. Essentially, this work demonstrates how PDFs can be used as a channel for interacting with Web services.
Ingo Weber, Hye-Young Paik, Boualem Benatallah, Zifei Gong, Liangliang Zheng, Corren Vorwerk
WWW2
2010 Semantic-Based Mashup of Composite Applications
abstract
The need for integration of all types of client and server applications that were not initially designed to interoperate is gaining popularity. One of the reasons for this popularity is the capability to quickly reconfigure a composite application for a task at hand, both by changing the set of components and the way they are interconnected. Service-Oriented Architecture (SOA) has recently become a popular platform in the IT industry for building such composite applications with the integrated components being provided as Web services. A key limitation of solely Web-service-based integration is that it requires extra programming efforts when integrating non-Web service components, which is not cost-effective. Moreover, with the emergence of new standards, such as Open Service Gateway Initiative (OSGi), the components used in composite applications have grown to include more than just Web services. Our work enables progressive composition of non-Web-service-based components such as portlets, Web applications, native widgets, legacy systems, and Java Beans. Further, we proposed a novel application of semantic annotation together with the standard semantic Web matching algorithm for finding sets of functionally equivalent components out of a large set of available non-Web-service-based components. Once such a set is identified, the user can drag and drop the most suitable component into an Eclipse-based composition canvas. After a set of components has been selected in such a way, they can be connected by data-flow arcs, thus forming an integrated, composite application without any low-level programming and integration efforts. We implemented and conducted extensive experimental study on the above progressive composition framework on IBM's Lotus Expeditor, an extension of an SOA platform called the Eclipse Rich Client Platform (RCP) that complies with the OSGi standard.
Anne H. H. Ngu, Michael Pierre Carlson, Quan Z. Sheng, Hye-Young Paik
IEEE Trans. Serv. Comput.4
2009 Risk Identification and Mitigation Processes for Using Scrum in Global Software Development: A Conceptual Framework
abstract
There is growing interest in applying agile practices in Global Software Development (GSD) projects. But project stakeholder distribution in GSD creates a number of challenges that make it difficult to use some agile practices. Moreover, little is known about what the key challenges or risks are, and how GSD project mangers deal with these risks while using agile practices. We conduct a Systematic Literature Review (SLR) following existing guidelines to identify primary papers that discuss the use of Scrum practices in GSD projects. We identify key challenges, due to global project distribution, that restrict the use of Scrum and explore the strategies used by project managers to deal with these challenges. Our findings are consolidated into a conceptual framework and we discuss various elements of this framework. This research is relevant to project managers who are seeking ways to use Scrum in their globally distributed projects.
Emam Hossain, Muhammad Ali Babar 0001, Hye-Young Paik, June M. Verner
APSEC3
2009 Using Scrum in Global Software Development: A Systematic Literature Review
abstract
There is a growing interest in applying agile practices in global software development (GSD) projects. The literature on using Scrum, one of the most popular agile approaches, in distributed development projects has steadily been growing. However, there has not been any effort to systematically select, review, and synthesize the literature on this topic. We have conducted a systematic literature review of the primary studies that report using Scrum practices in GSD projects. Our search strategy identified 366 papers, of which 20 were identified as primary papers relevant to our research. We extracted data from these papers to identify various challenges of using Scrum in GSD. Current strategies to deal with the identified challenges have also been extracted. This paper presents the reviewpsilas findings that are expected to help researchers and practitioners to understand the challenges involved in using Scrum for GSD projects and the strategies available to deal with them.
Emam Hossain, Muhammad Ali Babar 0001, Hye-Young Paik
ICGSE3
2007 Conceptual Modeling of Privacy-Aware Web Service Protocols
Rachid Hamadi, Hye-Young Paik, Boualem Benatallah
CAiSE2
2007 WS-Advisor: A Task Memory for Service Composition Frameworks
abstract
With the proliferation of Web services, it is becoming increasingly important to support the users in selecting the most appropriate compositions of services for a task. We propose a new service discovery and selection framework that utilises the concept of task memories and a social network of task memories. A task memory captures the service composition history and their meta-data such as associated context and user rating. A network of task memories is formed to realise an effective task memory sharing platform among the users.
Rosanna Bova, Hye-Young Paik, Salima Hassas, Salima Benbernou, Boualem Benatallah
ICCCN2
2007 Task Memories and Task Forums: A Foundation for Sharing Service-Based Personal Processes
Rosanna Bova, Hye-Young Paik, Boualem Benatallah, Liangzhao Zeng, Salima Benbernou
ICSOC2
2007 On Embedding Task Memory in Services Composition Frameworks
Rosanna Bova, Hye-Young Paik, Salima Hassas, Salima Benbernou, Boualem Benatallah
ICWE2
2007 Privacy Inspection and Monitoring Framework for Automated Business Processes
Yin Hua Li, Hye-Young Paik
WISE2
2006 Formal consistency verification between BPEL process and privacy policy
abstract
Despite the increased privacy concerns in the Internet, not much attention has been paid into enforcing privacy policies of organisations who collect and consume personal data using automatic means (e.g., Web services). In this paper, we propose a graph-transformation based framework to check whether an internal business process (implemented using a standard Web service composition language such as BPEL) adheres to the organisation's privacy policies. The graph-based specification formalism combines the advantages of an intuitive visual framework with rigorous semantical foundation that allows consistency checking between a business process and privacy policy. The privacy consistency verification framework is defined by a set of rules to build the system state and sets of constraints (positive and negative) to specify the wanted and unwanted substates.
Yin Hua Li, Hye-Young Paik, Boualem Benatallah
PST2
2006 Towards semantic-driven, flexible and scalable framework for peering and querying e-catalog communities
Boualem Benatallah, Mohand-Said Hacid, Hye-Young Paik, Christophe Rey, Farouk Toumani
Inf. Syst.3
2006 Building and querying e-catalog networks using P2P and data summarisation techniques
Hye-Young Paik, Noureddine Mouaddib, Boualem Benatallah, Farouk Toumani, Mahbub Hassan
J. Intell. Inf. Syst.1
2005 Toward self-organizing service communities
abstract
This paper discusses a framework in which catalog service communities are built, linked for interaction, and constantly monitored and adapted over time. A catalog service community (represented as a peer node in a peer-to-peer network) in our system can be viewed as domain specific data integration mediators representing the domain knowledge and the registry information. The query routing among communities is performed to identify a set of data sources that are relevant to answering a given query. The system monitors the interactions between the communities to discover patterns that may lead to restructuring of the network (e.g., irrelevant peers removed, new relationships created, etc.).
Hye-Young Paik, Boualem Benatallah, Farouk Toumani
IEEE Trans. Syst. Man Cybern. Part A1
2004 WS-CatalogNet: Building Peer-to-Peer e-Catalog
Hye-Young Paik, Boualem Benatallah, Farouk Toumani
FQAS1
2004 Peering and Querying e-Catalog Communities
abstract
More and more suppliers are offering access to their product or information portals (also called e-catalogs) via the Web. The key issue is how to efficiently integrate and query large, intricate, heterogeneous information sources such as e-catalogs. Traditional data integration approach, where the development of an integrated schema requires the understanding of both structure and semantics of all schemas of sources to be integrated, is hardly applicable because of the dynamic nature and size of the Web. We present WS-CatalogNet: a Web services based data sharing middleware infrastructure whose aims is to enhance the potential of e-catalogs by focusing on scalability and flexible aspects of their sharing and access.
Boualem Benatallah, Mohand-Said Hacid, Hye-Young Paik, Christophe Rey, Farouk Toumani
ICDE3
2004 WS-CatalogNet: An Infrastructure for Creating, Peering, and Querying e-Catalog Communities
Karim Baïna, Boualem Benatallah, Hye-Young Paik, Farouk Toumani, Christophe Rey, Agnieszka Rutkowska, Bryan Harianto
VLDB3
2002 Usage-Centric Adaptation of Dynamic E-Catalogs
Hye-Young Paik, Boualem Benatallah, Rachid Hamadi
CAiSE1
2002 Dynamic Restructuring of E-Catalog Communities Based on User Interaction Patterns
Hye-Young Paik, Boualem Benatallah, Rachid Hamadi
World Wide Web1