Xiaoying Bai

dblp:67/2196 · DBLP profile ↗
← Back
48ranked-venue papers
11as first author
13since 2021 · last 2026
0000-0003-3989-4075ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 23 · 9 first-author · 1 since 2021Artificial intelligence and machine learning · 12 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 5 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Databases, data management, data science and information retrieval · 5 · 4 since 2021Systems, architecture and hardware · 1Security and privacy · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Global-Local Confidence Fusion for Hallucination Detection in Mathematical Reasoning Task
abstract
Large Reasoning Models (LRMs) achieve promising results on complex reasoning tasks but remain susceptible to hallucinations. Existing hallucination detection methods based on Large Language Models (LLMs) often focus solely on final answers, overlooking inconsistencies between the answer and reasoning process. This limitation reduces their ability to detect hallucinations during inference. Moreover, training-free approaches lack mechanisms for confidence estimation, resulting in an unquantified detection output. In contrast, training-based methods can provide fine-grained assessments but often neglect the self-correction capability of LRMs, where earlier errors may be corrected in subsequent steps, leading to inaccurate hallucination detection. To address these challenges, we propose ConfFuse, a unified framework that fuses global and local confidence scores for hallucination detection. A Global Hallucination Detection Model (GHDM) is trained using Direct Preference Optimization (DPO) to assess hallucinations at the level of entire reasoning chains, yielding global confidence estimates. Simultaneously, a Process Reward Model (PRM) estimates step-wise confidence scores to capture local logical flaws. A weighted fusion strategy combines the global confidence score with the minimum local score to jointly reflect overall reasoning consistency and local soundness. Experimental evaluations demonstrate that ConfFuse surpasses Qwen3-1.7B and Qwen3-8B by up to 11.86% and 5.46% in F1 score on in-distribution datasets, and achieves average improvements of 4.65% and 2.80% on out-of-distribution datasets. These results verify the effectiveness and generalizability of the proposed framework.
Linkang Yang, Bingxu Han, Zhunchen Luo, Guotong Geng, Xiaoying Bai
AAAI8
2026 Diffusion-based Kriging Model with Graph-enhanced Attention
abstract
In web-based systems, elements are commonly organized within a graph structure, with each node collecting essential spatio-temporal data. Examples include websites on the World Wide Web, traffic monitors in transportation networks, or sensors in the Internet of Things (IoT). However, sensors are typically deployed sparsely and unevenly, leaving the remaining nodes unobserved. The spatio-temporal kriging task, which infers values at unobserved nodes from observed ones, has thus attracted significant research interest. Due to limitations such as reliance on static graph structures and iterative Graph Convolution Network (GCN) frameworks, accurate kriging remains challenging. To address these issues, we propose a Diffusion-based Kriging Model with Graph-enhanced Attention (DKM-GA). Our approach first introduces a graph-enhanced attention mechanism that dynamically learns more accurate graph structures by combining predefined graph knowledge with global node value similarities. It is then integrated into a diffusion-based framework, which is tailored for the reliance of attention on known values. Therefore, the framework progressively refines the target values using correlated nodes, and the graph-enhanced attention selects more relevant neighbors based on the refined values. Furthermore, a node-based rescaling strategy is introduced to align the inference phase graphs to the training ones. Experiments on eight real-world datasets demonstrate that DKM-GA achieves superior performance, reducing estimation errors by up to 12.66%. Moreover, our analysis identifies three practical scenarios where the model delivers greater performance gains, even achieving 19.51% improvements on datasets that show minor gains under standard settings. These results highlight the effectiveness and potential of our model, while the scenarios provide settings for more comprehensive evaluations in terms of performance and robustness.
Guoli Yang, Zhanxing Zhu, Guangyin Jin, Mengzhu Wang, Xiaoying Bai
WWW6
2026 DeepMEL: A multi-agent collaboration framework for multimodal entity linking
abstract
Multimodal Entity Linking (MEL) aims to associate textual and visual mentions with entities in a multimodal knowledge graph. Despite its importance, current methods face challenges such as incomplete contextual information, coarse cross-modal fusion, and the difficulty of jointly large language models (LLMs) and large visual models (LVMs). To address these issues, we propose DeepMEL, a novel framework based on multi-agent collaborative reasoning, which achieves efficient alignment and disambiguation of textual and visual modalities through a role-specialized division strategy. DeepMEL integrates four specialized agents, namely Modal-Fuser, Candidate-Adapter, Entity-Clozer and Role-Orchestrator, to complete end-to-end cross-modal linking through specialized roles and dynamic coordination. DeepMEL adopts a dual-modal alignment path, and combines the fine-grained text semantics generated by the LLM with the structured image representation extracted by the LVM, significantly narrowing the modal gap. We design an adaptive iteration strategy, combines tool-based retrieval and semantic reasoning capabilities to dynamically optimize the candidate set and balance recall and precision. DeepMEL also unifies MEL tasks into a structured cloze prompt to reduce parsing complexity and enhance semantic comprehension. Extensive experiments on five public benchmark datasets demonstrate that DeepMEL achieves state-of-the-art performance, improving ACC by 1 %-57 %. Ablation studies verify the effectiveness of all modules.
Fang Wang 0011, Tianwei Yan 0001, Zonghao Yang, Minghao Hu 0001, Zhunchen Luo, Xiaoying Bai
Inf. Process. Manag.7
2026 WizardEvent: Empowering Event Reasoning by Hybrid Event-Aware Data Synthesizing
abstract
Event reasoning is to reason with events and certain inter-event relations. These cutting-edge techniques possess crucial and fundamental capabilities that underlie various applications. Large language models (LLMs) have made advances in event reasoning owing to their wealth of training. However, the LLMs commonly used today still do not consistently demonstrate proficiency in managing event reasoning as humans. This discrepancy arises from not explicitly modeling events and their relations and insufficient knowledge of event relations. In addition, the different reasoning paradigms of the LLMs are trained in an imbalanced way. In this paper, we propose WIZARDEVENT, to synthesize data from the unlabeled corpus with the proposed hybrid event-aware instruction tuning. Specifically, we first represent the events and their relation in a novel structure and then extract the knowledge from the raw text. Second, we introduce hybrid event reasoning paradigms with four reasoning formats. Lastly, we wrap our constructed WIZARDEVENT with the paradigms to create the instruction tuning dataset. We fine-tune the model with this enriched dataset, significantly improving the event reasoning. The performance of WIZARDEVENT is rigorously evaluated through extensive experiments. The results demonstrate that WIZARDEVENT substantially outperforms baselines, indicating the effectiveness of our approach.
Zhengwei Tao, Xiancai Chen, Zhi Jin 0001, Xiaoying Bai, Haiyan Zhao 0001, Wenpeng Hu, Chongyang Tao, Shuai Ma 0001
IEEE Trans. Knowl. Data Eng.4
2025 M^3EL: A Multi-task Multi-topic Dataset for Multi-modal Entity Linking
abstract
Multi-modal Entity Linking (MEL) is a fundamental component for various downstream tasks. However, existing MEL datasets suffer from small scale, scarcity of topic types and limited coverage of tasks, making them incapable of effectively enhancing the entity linking capabilities of multi-modal models. To address these obstacles, we propose a dataset construction pipeline and publish M^3EL, a large-scale dataset for MEL. M^3EL includes 79,625 instances, covering 9 diverse multi-modal tasks, and 5 different topics. In addition, to further improve the model's adaptability to multi-modal tasks, We propose a modality-augmented training strategy. Utilizing M^3EL as a corpus, train the CLIP_ND model based on CLIP (ViT-B-32), and conduct a comparative analysis with an existing multi-modal baselines. Experimental results show that the existing models perform far below expectations (ACC of 49.4%-75.8%), After analysis, it was obtained that small dataset sizes, insufficient modality task coverage, and limited topic diversity resulted in poor generalization of multi-modal models. Our dataset effectively addresses these issues, and the CLIP_ND model fine-tuned with M^3EL shows a significant improvement in accuracy, with an average improvement of 9.3% to 25% across various tasks. Our dataset publicly available to facilitate future research.
Fang Wang 0011, Shenglin Yin, Xiaoying Bai, Minghao Hu 0001, Tianwei Yan 0001
AAAI3
2025 Poplar: Efficient Scaling of Distributed DNN Training on Heterogeneous GPU Clusters
abstract
Scaling Deep Neural Networks (DNNs) requires significant computational resources in terms of GPU quantity and compute capacity. In practice, there usually exists a large number of heterogeneous GPU devices due to the rapid release cycle of GPU products. It is highly needed to efficiently and economically harness the power of heterogeneous GPUs, so that it can meet the requirements of DNN research and development. The paper introduces Poplar, a distributed training system that extends Zero Redundancy Optimizer (ZeRO) with heterogeneous-aware capabilities. We explore a broader spectrum of GPU heterogeneity, including compute capability, memory capacity, quantity and a combination of them. In order to achieve high computational efficiency across all heterogeneous conditions, Poplar conducts fine-grained measurements of GPUs in each ZeRO stage. We propose a novel batch allocation method and a search algorithm to optimize the utilization of heterogeneous GPUs clusters. Furthermore, Poplar implements fully automated parallelism, eliminating the need for deploying heterogeneous hardware and finding suitable batch size. Extensive experiments on three heterogeneous clusters, comprising six different types of GPUs, demonstrate that Poplar achieves a training throughput improvement of 1.02-3.92x over current state-of-the-art heterogeneous training systems.
Xiaoying Bai
AAAI4
2025 An LLM Agent-Based Complex Semantic Table Annotation Approach
Shujing Wang 0013, Keqing He 0002, Yanfei Lv, Zaiwen Feng, Xiaoying Bai
ADMA (2)8
2025 BlockAlign: Fair Performance Testing for Blockchains based on Configuration Alignment
abstract
As blockchain technology grows more prevalent, its performance limitations have become a critical barrier to large-scale adoption. The rise of optimized heterogeneous blockchain systems has significantly increased the demand for fair performance testing frameworks. However, existing work, whether simulator-based or system-based, often relies on default settings, overlooking the impact of detailed configuration parameters, which can affect the fairness of performance evaluations. To fill this gap, we propose BlockAlign, a configuration alignment tool based on a rule tree, designed to enhance the fairness of performance testing across heterogeneous blockchain systems. Firstly, based on architectural analysis, we design a rule tree to filter, classify, and semantically align configurations. Secondly, we introduce a metric called fluctuation rate to measure performance differences before and after configuration alignment. Finally, we conduct experiments on Geth, Besu, and Conflux, demonstrating that BlockAlign significantly improves the fairness and credibility of performance comparisons.
Chenglin Xie, Peilun Li, Guoli Yang, Xiaoying Bai
APSEC6
2025 Uncovering Argumentative Flow: A Question-Focus Discourse Structuring Framework
abstract
Understanding the underlying argumentative flow in analytic argumentative writing is essential for discourse comprehension, especially in complex argumentative discourse such as think-tank commentary.However, existing structure modeling approaches often rely on surface-level topic segmentation, failing to capture the author's rhetorical intent and reasoning process.To address this limitation, we propose a Question-Focus discourse structuring framework that explicitly models the underlying argumentative flow by anchoring each argumentative unit to a guiding question (reflecting the author's intent) and a set of attentional foci (highlighting analytical pathways).To assess its effectiveness, we introduce an argument reconstruction task in which the modeled discourse structure guides both evidence retrieval and argument generation.We construct a high-quality dataset comprising 600 authoritative Chinese think-tank articles for experimental analysis.To quantitatively evaluate performance, we propose two novel metrics: (1) Claim Coverage, measuring the proportion of original claims preserved or similarly expressed in reconstructions, and (2) Evidence Coverage, assessing the completeness of retrieved supporting evidence.Experimental results show that our framework uncovers the author's argumentative logic more effectively and offers better structural guidance for reconstruction, yielding up to a 10% gain in claim coverage and outperforming strong baselines across both curated and LLM-based metrics.
Yini Wang, Xian Zhou 0003, Shengan Zheng, Linpeng Huang, Zhunchen Luo, Xiaoying Bai
EMNLP7
2025 DiffMEL: A large-scale difficulty-graded dataset for Multimodal Entity Linking
abstract
Multimodal Large Language Models (MLLMs) have shown tremendous potential in Multimodal Entity Linking (MEL). However, they are still far from achieving the expected effectiveness in practical applications. This could be due to limitations in the MEL dataset used for training. Existing MEL datasets primarily focus on simple tasks and only consider the direct matching of mentions with labeled entities within a multimodal context, ignoring mentions of unmatched entities. Factors such as the presence of the ground-truth entity within the candidate set and its position directly impact the performance of MLLMs on MEL tasks. To tackle these obstacles, we constructed DiffMEL, the first large-scale difficulty-graded dataset for MEL of MLLMs. DiffMEL contains 79,625 instances and 318.5K instance-related high-resolution images, covering 3 various difficulty graded linking tasks and 5 different entity themes. We utilize DiffMEL to train several open-source MLLMs. Experiment results demonstrate DiffMEL empowers MLLMs with stronger capabilities in MEL by a large-margin (5%-56.1%). our dataset is now available at https://github.com/ww-ffff/DiffMEL.
Fang Wang 0011, Xiaoying Bai, Tianwei Yan 0001, Minghao Hu 0001
ICASSP2
2024 From Toxic to Trustworthy: Using Self-Distillation and Semi-supervised Methods to Refine Neural Networks
abstract
Despite the tremendous success of deep neural networks (DNNs) across various fields, their susceptibility to potential backdoor attacks seriously threatens their application security, particularly in safety-critical or security-sensitive ones. Given this growing threat, there is a pressing need for research into purging backdoors from DNNs. However, prior efforts on erasing backdoor triggers not only failed to withstand increasingly powerful attacks but also resulted in reduced model performance. In this paper, we propose From Toxic to Trustworthy (FTT), an innovative approach to eliminate backdoor triggers while simultaneously enhancing model accuracy. Following the stringent and practical assumption of limited availability of clean data, we introduce a self-attention distillation (SAD) method to remove the backdoor by aligning the shallow and deep parts of the network. Furthermore, we first devise a semi-supervised learning (SSL) method that leverages ubiquitous and available poisoned data to further purify backdoors and improve accuracy. Extensive experiments on various attacks and models have shown that our FTT can reduce the attack success rate from 97% to 1% and improve the accuracy of 4% on average, demonstrating its effectiveness in mitigating backdoor attacks and improving model performance. Compared to state-of-the-art (SOTA) methods, our FTT can reduce the attack success rate by 2 times and improve the accuracy by 5%, shedding light on backdoor cleansing.
Baolin Zheng, Jianbao Hu, Chengyang Li 0001, Xiaoying Bai
AAAI5
2024 LLM Assists Hypothesis Generation and Testing for Deliberative Questions
Fuchun Wang, Xian Zhou 0003, Wenpeng Hu, Zhunchen Luo, Xiaoying Bai
NLPCC (2)6
2022 Mixed graph convolution and residual transformation network for skeleton-based action recognition
Xiaoying Bai, Ming Fang 0006, Lanting Li, Chih-Cheng Hung
Appl. Intell.2
2020 Integrating Gaussian mixture model and dilated residual network for action recognition in videos
Ming Fang 0006, Xiaoying Bai, Fengqin Yang, Chih-Cheng Hung
Multim. Syst.2
2020 Advances in test automation for software with special focus on artificial intelligence and machine learning
J. Jenny Li 0001, Andreas Ulrich, Xiaoying Bai, Antonia Bertolino
Softw. Qual. J.3
2019 Decentralized Authorization and Authentication Based on Consortium Blockchain
Xiaoying Bai
BlockSys2
2018 The DevOps Lab Platform for Managing Diversified Projects in Educating Agile Software Engineering
abstract
This Research Work-in-Progress paper presents the design of a Software Engineering (SE) course to support project-based practical training. Group projects, especially projects from industry partners, are deemed to be necessary for students to gain hands-on experiences. With projects from the real world, students learn not only practical engineering solutions, but also the context, constraints, and social aspects of SE. For a course having over 100 students with different interests and experiences, it is desired to provide diversified choices of projects to stimulate enthusiasm for learning. However, management and evaluation of diversified projects are challenging. Following the Agile principles, we need to continuously track progress and activities of each group, to provide quick feedback of deliveries, and to periodically evaluate students’ performance. Therefore, we built a DevOps platform based on GitLab version control and continuous integration framework. Commits to GitLab code repositories automatically trigger build, testing, and analysis functions (which provide both qualitative and quantitative feedback to the students). This system has been in operations since 2014 for an undergraduate SE course, with over 500 students participating in over 130 project teams in total. The preliminary research showed promising results in improving SE education.
Xiaoying Bai, Dan Pei, Mingjie Li 0005
FIE1
2017 A Model-Based Framework for Cloud API Testing
abstract
Following the Service-Oriented Architecture, a large number of diversified Cloud services are exposed as Web APIs (Application Program Interface), which serve as the contracts between the service providers and service consumers. Due to their massive and broad applications, any flaw in the cloud APIs may lead to serious consequences. API testing is thus necessary to ensure the availability, reliability, and stability of cloud services. The research proposes a model-based approach to automating API testing. The semi-structured API specifications, like XML/HTML specifications, are gathered from the Web sites using web crawlers, and translated into YAML-encoded standard representations. A scenario editor is designed to specify the dependencies among API operations. Test generators are built to derive test scripts from the specifications and scenarios, including test data, test cases for individual operations as well as operations sequences. Various algorithms can be used for test generation, such as combinatorial data generation, heuristic graph search, and optimization algorithms. The produced test scripts, together with a load model, can be deployed on Cloud and scheduled for execution. A prototype system, called ATCloud, was constructed to illustrate the process of API understanding, test scenario modeling using directed diagraph annotated with transfer probabilities between operations, cloud-based test resources management, distributed workload simulation, and performance monitoring.
Xiaoying Bai
COMPSAC (2)2
2016 Cloud Performance Modeling with Benchmark Evaluation of Elastic Scaling Strategies
abstract
In this paper, we present generic cloud performance models for evaluating Iaas, PaaS, SaaS, and mashup or hybrid clouds. We test clouds with real-life benchmark programs and propose some new performance metrics. Our benchmark experiments are conducted mainly on IaaS cloud platforms over scale-out and scale-up workloads. Cloud benchmarking results are analyzed with the efficiency, elasticity, QoS, productivity, and scalability of cloud performance. Five cloud benchmarks were tested on Amazon IaaS EC2 cloud: namely YCSB, CloudSuite, HiBench, BenchClouds, and TPC-W. To satisfy production services, the choice of scale-up or scale-out solutions should be made primarily by the workload patterns and resources utilization rates required. Scaling-out machine instances have much lower overhead than those experienced in scale-up experiments. However, scaling up is found more cost-effective in sustaining heavier workload. The cloud productivity is greatly attributed to system elasticity, efficiency, QoS and scalability. We find that auto-scaling is easy to implement but tends to over provision the resources. Lower resource utilization rate may result from auto-scaling, compared with using scale-out or scale-up strategies. We also demonstrate that the proposed cloud performance models are applicable to evaluate PaaS, SaaS and hybrid clouds as well.
Kai Hwang 0001, Xiaoying Bai, Wen-Guang Chen, Yongwei Wu 0001
IEEE Trans. Parallel Distributed Syst.2
2014 Scale-Out vs. Scale-Up Techniques for Cloud Performance and Productivity
abstract
An elastic cloud provisions machine instances upon user demand. Auto-scaling, scale-out, scale-up, or any mixture techniques are used to reconfigure the user cluster as workload changes. We evaluate three scaling strategies to upgrade the performance, efficiency and productivity of elastic clouds like EC2, Rack space, etc. We developed new performance models and run the Hi Bench benchmark to test Hadoop performance on various EC2 configurations. The strengths and shortcomings of three scaling strategies are revealed in our Hi Bench experiments: (1). Scale-out overhead is shown lower than that experienced in scale-up or mixed scaling clouds. Scale-out to a larger cluster of small nodes demonstrated high scalability. (2). Scaling up and mixed scaling have high performance in using smaller clusters with a few powerful machine instances. (3). With a mixed scaling mode, the cloud productivity is shown upgradable with higher flexibility in applications with performance/cost tradeoffs.
Kai Hwang 0001, Xiaoying Bai
CloudCom3
2014 Data Driven Testing of Open Source Software
Inbal Yahav, Ron S. Kenett, Xiaoying Bai
ISoLA (2)3
2014 Software-as-a-service (SaaS): perspectives and challenges
Wei-Tek Tsai, Xiaoying Bai, Yu Huang 0008
Sci. China Inf. Sci.2
2014 Special issue on "Trustworthy Software Systems for the Digital Society"
Xiaoying Bai, Atilla Elçi, Mohammad Zulkernine
J. Syst. Softw.1
2013 Adaptive Fault Detection for Testing Tenant Applications in Multi-tenancy SaaS Systems
abstract
SaaS (Software-as-a-Service) often uses multi-tenancy architecture (MTA) where tenant developers compose their applications online using the components stored in the SaaS database. Tenant applications need to be tested, and combinatorial testing can be used. While numerous combinatorial testing techniques are available, most of them produce static sequences of test configurations and their goal is often to provide sufficient coverage such as 2-way interaction coverage. But the goal of SaaS testing is to identify those compositions that are faulty for tenant applications. This paper proposes an adaptive test configuration generation algorithm AR (Adaptive Reasoning) that can rapidly identify those faulty combinations so that those faulty combinations cannot be selected by tenant developers for composition. The AR algorithm has been evaluated by both simulation and real experimentation using a MTA SaaS sample running on GAE (Google App Engine). Both the simulation and experiment showed show that the AR algorithm can identify those faulty combinations rapidly. Whenever a new component is submitted to the SaaS database, the AR algorithm can be applied so that any faulty interactions with new components can be identified to continue to support future tenant applications.
Wei-Tek Tsai, Qingyang Li 0001, Charles J. Colbourn, Xiaoying Bai
IC2E4
2012 A cloud-based TaaS infrastructure with tools for SaaS validation, performance and scalability evaluation
abstract
With the fast advancements in cloud computing and software- as-a-service (SaaS), testing and evaluation of cloud-based software and SaaS applications became an important task for engineers. Since most existing tools are not developed to support cloud-based software testing and SaaS evaluation, there is a strong demand for a new cloud-based testing infrastructure and evaluation environment for SaaS applications. This paper proposes a testing-as-service (TaaS) infrastructure and reports a cloud-based TaaS environment with tools (known as CTaaS) developed to meet the needs in SaaS testing, performance and scalability evaluation. The paper presents TaaS concepts and CTaaS, including their infrastructure, design and implementation. In addition, the paper demonstrates the application results of our previously proposed graphic models and metrics for SaaS performance and scalability evaluation. Moreover, the paper reports one case study for a selected SaaS (OrangeHRM) using the developed TaaS environment.
Jerry Zeyu Gao, K. Manjula, P. Roopa, E. Sumalatha, Xiaoying Bai, Wei-Tek Tsai, Tadahiro Uehara
CloudCom5
2012 Risk Assessment and Adaptive Group Testing of Semantic Web Services
abstract
Testing is necessary to ensure the quality of web services that are loosely coupled, dynamic bound and integrated through standard protocols. Exhaustive testing of web services is usually impossible due to unavailable source code, diversified user requirements and large number of possible service combinations delivered by the open platform. This paper proposes a risk-based approach for selecting and prioritizing test cases for testing service-based systems. We specially address the problem in the context of semantic web services. Semantic web services introduce semantics to service integration and interoperation using ontology models and specifications. Semantic errors are considered more difficult to detect than syntactic errors. Due to the complexity of conceptual uniformity, it is hard to ensure the completeness, consistency and unified quality of ontology model. A failure of the semantic service-based software may result from many factors such as misused data, unsuccessful service binding, and unexpected usage scenarios. This work analyzes the two factors of risk estimation: failure probability and importance, from three aspects: ontology data, service and composite service. With this approach, test cases are associated to semantic features, and are scheduled based on the risks of their target features. Risk assessment is used to control the process of Web Services progressive group testing, including test case ranking, test case selection and service ruling out. This paper discusses the control architecture and adaptive measurement mechanism for adaptive group testing. As a statistical testing technique, the proposed approach aims to detect, as early as possible, the problems with highest impact on the users.
Xiaoying Bai, Ron S. Kenett
Int. J. Softw. Eng. Knowl. Eng.1
2011 Semantic-Based Test Oracles
abstract
Test oracle is one of the most difficult parts for test automation. For software with a large number of test cases, it is always both expensive and error prone to develop and maintain test oracles. The research is motivated by industry needs of automated testing on software with standard interfaces in an open system architecture. In counter to test oracle challenges, it proposes an innovative method to represent and calculate test oracles based on the semantic model of standard interface service specification of the software under test (SUT). Semantic model provides well-defined domain knowledge of service data, functionalities and constraints. Rules are created to model the expected SUT behavior in terms of antecedents and consequents. For each service, it captures both direct input-output relations and service interactions, that is, how the execution of a service may be affected by (pre-condition) or impact (post-condition) the SUT system state. As rule languages are neutral to programming languages, oracles specified in this way are independent of SUT implementations and can be reused across different systems conforming to the same interface standards. With the support of semantic techniques and tools like ontology modeler and rule engine, the proposed approach can enhance test oracle automation based on sophisticated defined domain model. Experiments and analysis show promising improvements in test productivity and quality.
Xiaoying Bai, Kejia Hou, Linping Hu
COMPSAC1
2011 An Approach for Service Composition and Testing for Cloud Computing
abstract
As cloud services proliferate, it becomes difficult to facilitate service composition and testing in clouds. In traditional service-oriented computing, service composition and testing are carried out independently. This paper proposes a new approach to manage services on the cloud so that it can facilitate service composition and testing. The paper uses service implementation selection to facilitate service composition similar to Google's Guice and Spring tools, and apply the group testing technique to identify the oracle, and use the established oracle to perform continuous testing for new services or compositions. The paper extends the existing concept of template based service composition and focus on testing the same workflow of service composition. In addition, all these testing processes can be executed in parallel, and the paper illustrates how to apply service-level MapReduce technique to accelerate the testing process.
Wei-Tek Tsai, Peide Zhong, Janaka Balasooriya, Yinong Chen 0004, Xiaoying Bai, Jay Elston
ISADS5
2011 Service Replication Strategies with MapReduce in Clouds
abstract
The current implementations of cloud environment do not have suitable mechanism through which services can be managed to make use of cloud resources. The services in these environments can passively serve users' request only. If a service receives more requests than it can handle in a certain time period, it is subject to malfunctioning. This paper proposes a new approach to service replications that allows a cloud to adjust its service instance deployments in response to existing and projected service request loads. This approach defines an optimal service replication strategy based on MapReduce, a processing model used extensively in GAE (Google App Engine) to solve huge data processing tasks. This service replication strategy is implemented by Service Level MapReduce (SLMR). To better support for SLMR, service replication technology is introduced, which include dynamic service replication and pre-deployed service replication. Furthermore, a passive SLMR approach that depends on the cloud management service (CMS) and a positive SLMR approach that does not need the support from CMS will be introduced.
Wei-Tek Tsai, Peide Zhong, Jay Elston, Xiaoying Bai, Yinong Chen 0004
ISADS4
2011 Testing Configurable Component-Based Software - Configuration Test Modeling and Complexity Analysis
Jerry Zeyu Gao, Jing Guan, Alex Ma, Chuanqi Tao, Xiaoying Bai, David Chenho Kung
SEKE5
2011 Guest Editors' Introduction
Jerry Zeyu Gao, Henry Muccini, Xiaoying Bai
Int. J. Softw. Eng. Knowl. Eng.3
2010 Towards a scalable and robust multi-tenancy SaaS
abstract
Software-as-as-Service (SaaS) is a new approach for developing software, and it is characterized by its multi-tenancy architecture and its ability to provide flexible customization to individual tenant. However, the multi-tenancy architecture and customization requirements have brought up new issues in software, such as database design, database partition, scalability, recovery, and continuous testing. This paper proposes a hybrid test database design to support SaaS customization with two-layer database partitioning. The database is further extended with a new built-in redundancy with ontology so that the SaaS can recover from ontology, data or meta-data failures. Furthermore, constraints in metadata can be used either as test cases or policies to support SaaS continuous testing and policy enforcement.
Wei-Tek Tsai, Qihong Shao, Yu Huang 0008, Xiaoying Bai
Internetware4
2010 Ontology-Based Dependency-Guided Service Composition for User-Centric SOA
Wei-Tek Tsai, Peide Zhong, Jay Elston, Yinong Chen 0004, Xiaoying Bai
SEKE5
2009 Risk-Based Adaptive Group Testing of Semantic Web Services
abstract
Comprehensive testing is necessary to ensure the quality of complex Web services that are loosely coupled, dynamic bound and integrated through standard protocols. Testing of such web services can be however very expensive due to the diversified user requirements and the large numbers of service combinations delivered by the open platform. Group testing was introduced in our previous research as a selective testing technique to reduce test cost and improve test efficiencies. It applies test cases efficiently so that the largest percent of problematic web service is detected as early as possible. The paper proposes a risk-based approach to group test selection. With this approach, test cases are categorized and scheduled with respect to the risks of their target service features. The approach is based on the assumption that for a service-based system, the tolerance to a featurepsilas failure is an inverse ratio to its risk. The risky features should be tested earlier and with more tests. We specially address the problem in the context of semantic Web Services and report a first attempt for an ontology-based quantitative risk assessment. The paper also discusses risk-based group testing process and strategies for ranking and ruling-out services of the test groups, at each risk level. Runtime monitoring mechanism is incorporated to detect the dynamic changes in service configuration and composition so that the risks can be continuously adjusted online.
Xiaoying Bai, Ron S. Kenett
COMPSAC (2)1
2009 First international workshop on service-oriented architecture testing (SOAT 2009)
abstract
Service-Oriented Architecture (SOA) is a way of designing, developing, deploying, and managing enterprise systems where business needs and technical solutions are closely aligned. SOA offers a number of potential benefits, such as cost-efficiency and agility. However, adopting SOA is not without considerable challenges. For example, the most common way to implement a SOA-based system is with Web services, but the standards that define Web services are evolving rapidly and many of the Web services tools are still somewhat immature. There is also the question of how to leverage existing legacy assets within a SOA context. Perhaps most importantly, there are serious challenges related to the testing of SOA-based systems that must be addressed before the SOA paradigm will enjoy broadbased success.
Scott R. Tilley, Xiaoying Bai, Grace A. Lewis
ICSM2
2009 Internetware computing: issues and perspective
abstract
The Internetware is a new initiative to develop software on the web for web applications. The open and dynamic nature of Internet applications suggest new ways of thinking will be needed for this initiative. This paper discusses several important issues in Internetware and put forward to some relevant research directions. The relevant issues include lifecycle models, ontology and context systems, modeling and simulation, social networking, and adaptive control.
Wei-Tek Tsai, Xiaoying Bai
Internetware3
2008 Collaborative Web Services Monitoring with Active Service Broker
abstract
This paper proposes a collaborative runtime monitoring framework to enhance the dependability of the software developed in traditional Web services architecture. The enabling mechanism is an active service broker (ASB) architecture which allows the service broker not only to serve as a passive service repository, but also to involve itself in service interactions and thus to play an active role in service execution and monitoring. The ASB communicates remotely with the distributed monitoring agents deployed at the service providerspsila sites. Sensors are instrumented in services at different levels, including the composition level, interface level, and component level. The model-based approach is discussed for automatic sensor generation and runtime enforcement based on service's process model and verification model. A prototype is implemented for the research of the user and science data center of the astronomic satellite system at Tsinghua University.
Xiaoying Bai, Shufang Lee, Wei-Tek Tsai, Yinong Chen 0004
COMPSAC1
2008 Ontology-Based Test Modeling and Partition Testing of Web Services
abstract
Testing is useful to establish trust between service providers and clients. To test the service-oriented applications, automated and specification-based test generation and test collaboration are necessary. The paper proposes an ontology-based approach for Web services (WS) testing. A test ontology model (TOM) is defined to specify the test concepts, relationships, and semantics from two aspects: test design (such as test data, test behavior, and test cases) and test execution (such as test plan, schedule and configuration). The TOM specification using OWL (Web ontology language) can serve as test contracts among test components. Based on the WS semantic specification in OWL-S, the paper discusses the techniques to generate the sub-domains for input partition testing. Data pools are established for each parameter of the specified service. Data partitions are derived by class property and relationship analysis. Completeness and consistency (C&C) checking can be performed on the data partitions and data values, both within the TOM and against the OWL-S, by ontology class computation and reasoning. A prototype tool is implemented to support OWL-S analysis, test ontology generation and C&C checking.
Xiaoying Bai, Shufang Lee, Wei-Tek Tsai, Yinong Chen 0004
ICWS1
2008 SyncTest: a Tool to Synchronize Source Code, Model and Testing
Xiaoying Bai
SEKE1
2008 Translating OWL Specified Domain Knowledge to Aspect Oriented Model
Juan-Zi Li, Xinyu You, Xiaoying Bai
SEKE3
2007 Adaptive Web Services Testing
abstract
Web services (WS) and service-oriented architecture (SOA) present a set of unique testing challenges. As services are distributed, it is necessary to test them using a distributed architecture. Furthermore, as these services may keep on changing, testing needs to be adaptive. This paper proposes an adaptive testing framework which can continuously learn and improve the built-in test strategies. The framework allows different test cases to be selected based on the recent test results. The framework also has a windowing mechanism to evaluate and select test cases.
Xiaoying Bai, Yinong Chen 0004, Zhongkui Shao
COMPSAC (2)1
2007 Dynamic Reconfigurable Testing of Service-Oriented Architecture
abstract
SOA (Service-Oriented Architecture) presents unique requirements and challenges for testing. Dynamic reconfiguration in SOA software means that testing need to be adaptive to the changes of the service-oriented applications at runtime. This paper presents a ConfigTest approach to enable the online change of test organization, test scheduling, test deployment, test case binding, and service binding. ConfigTest is based on our previous research on the MAST (Multi-Agents-based Service Testing) framework. It extends MAST with a new test broker architecture, configuration management and event-based subscription/notification mechanism. The test broker decouples test case definition from its implementation and usage. It also decouples the testing system from the services under test. With the configuration management, ConfigTest allows the test agents to bind dynamically to each other and build up their collaborations at runtime. The event mechanism enables that a change in one test artifact can be notified to all the others which subscribe their interests to the change event. This paper presents and analyzes the collaboration diagrams of various testing reconfiguration scenarios and illustrates the ConfigTest approach with an example of service-based book ordering system.
Xiaoying Bai, Dezheng Xu, Guilan Dai
COMPSAC (1)1
2007 Contract-Based Testing for Web Services
abstract
This paper examines the use of Design by Contract for web service descriptions, and explores the issues and solutions of automatic test case generation and test oracle generation in the context of WS testing based on contracts. In our approach, the traditional concept of contracts (pre-condition, post-condition, and invariant) is extended to contain richer information, such as process control, to support automatic test generation. Contracts are used to specify the relation between a component and its clients as a formal agreement, expressing each party's rights and obligations. Contracts can be expressed in the OWL-S process model. By checking whether the web service respects its contracts, we can ascertain its validity. Therefore, contracts provide the basis for the automation of the testing process.
Guilan Dai, Xiaoying Bai, Fengjun Dai
COMPSAC (1)2
2007 Ontology-Based Test Case Generation for Testing Web Services
abstract
Web services (WS) enables agile application development by orchestrating the existing service components. However, the dynamically constructed service-based system has to be tested dynamically and automatically at runtime without human intervention. To address the challenges of automatic WS test case generation, this paper proposes a model driven ontology-based approach with the purpose of improving test formalism and test intelligence. The semantic WS specification OWL-S is used to describe the application logic of composite service process. A Petri-Net model is created to provide a formal representation of the OWL-S (Web Ontology Language for Web service) process model. The Petri-net ontology is defined to incorporate the operation and IOPE (inputs, outputs, preconditions, and effects) semantics for test generation. Test cases are generated from two aspects. Test steps are generated by traversing various execution paths of the Petri-net graph. Test data are generated by reasoning over the IOPE ontology
Xiaoying Bai, Juan-Zi Li, Ruobo Huang
ISADS2
2001 Scenario-Based Functional Regression Testing
abstract
Regression testing has been a popular quality-assurance technique. Most regression testing techniques are based on code or software design. This paper proposes a scenario-based functional regression testing, which is based on end-to-end (E2E) integration test scenarios. The test scenarios are first represented in a template model that embodies both test dependency and traceability. By using test dependency information, one can obtain a test slicing algorithm to detect the scenarios that are affected and thus they are candidates for regression testing. By using traceability information, one can find affected components and their associated test scenarios and test cases for regression testing. With the same dependency and traceability information one can use the ripple effect analysis to identify all affected, including directly or indirectly, scenarios and thus the set of test cases can be selected for regression testing. This paper also provides several alternative test-case selection approaches and a hybrid approach to meet various requirements. A web-based tool has been developed to support these regression testing tasks.
Raymond A. Paul, Wei-Tek Tsai, Xiaoying Bai
COMPSAC4
2001 End-To-End Integration Testing Design
abstract
Integration testing has always been a challenge especially if the system under test is large with many subsystems and interfaces. This paper proposes an approach to design End-to-End (E2E) integration testing, including test scenario specification, test case generation and tool support. Test scenarios are specified as thin threads, each of which represents a single function from an end user's point of view. Thin threads can be organized hierarchically into a tree with each branch consisting of a set of related thin threads representing a set of related functionality. A test engineer can use thin-thread trees to generate test cases systematically, as well as carry out other related tasks such as risk analysis and assignment, regression testing, ripple effect analysis. A prototype tool has been developed to support E2E testing in a distributed environment on the J2EE platform.
Wei-Tek Tsai, Xiaoying Bai, Raymond A. Paul, Weiguang Shao, Vishal Agarwal
COMPSAC2
2001 Distributed End-to-End Testing Management
abstract
Testing is the primary means for quality assurance for enterprise systems, and integration testing is often the most time consuming and expensive part of testing. Recently Department of Defense proposed an End-to-End (E2E) integration testing process to address the challenge of testing large integrated information systems. The E2E testing activities include test thin-thread tree construction, condition tree specification, test configuration management, risk analysis, regression testing, ripple effect analysis, test scenario/case generation, rest result analysis, statistical analysis, and project management. A tool has been developed to support this E2E testing on J2EE using EJB and XML with Cloudscape relational database management system. This web-based tool also allows distributed collaboration and remote project management by the E2E testing participants including project managers, contractors, designers and testers.
Xiaoying Bai, Wei-Tek Tsai, Techeng Shen, Raymond A. Paul
EDOC1
2001 XML-based E2E Test Report Management
Raymond A. Paul, Wei-Tek Tsai, Xiaoying Bai
ER4