Ping Yu 0004

dblp:72/1358-4 · DBLP profile ↗
← Back
47ranked-venue papers
9as first author
14since 2021 · last 2026
0000-0002-7910-9396ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 18 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-authorSecurity and privacy · 2Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Optimising clinical information extraction: a comparative study of retrieval-augmented generation techniques in clinical notes
abstract
Extracting clinically meaningful information from free-text notes in specific clinical settings, such as Australian aged care facilities, remains challenging due to the heterogeneity of text documents, which lack standardised structure and terminology. Retrieval-augmented generation (RAG) can improve the precision and grounding of large language model (LLM) outputs; however, the choice of retrieval strategies remains understudied despite its critical importance for clinical information extraction (IE). Using real-world clinical notes from Australian aged care facilities, we systematically compare six retrieval methods within a unified RAG pipeline: sparse retrieval (BM25), dense retrieval (bi-encoder embeddings), dense retrieval with cross-encoder rerank (abbreviated as dense reranking), dynamic linear fusion of sparse and dense scores, reciprocal rank fusion (RRF), and hybrid coarse-to-fine reranking. We evaluate these strategies on two clinical named entity recognition tasks, extracting agitation symptoms in dementia (n = 208) and identifying malnutrition risk factors (n = 208) across five dimensions: context relevance, answer quality, source faithfulness, contextual diversity, and item-level accuracy. A repeated-measures ANOVA reveals that reranking and hybrid ensemble strategies significantly outperform both standalone sparse (BM25) and dense retrieval across both tasks. For agitation extraction, Dense reranking achieves the highest Answer F1 (0.946), Context Diversity (0.895) and Item-level Accuracy (0.963). For malnutrition, ensemble methods yield the best Item-level Accuracy (0.944), followed closely by dense reranking. Through error analysis, we identify three LLM error types in clinical named entity recognition: intrinsic hallucination, extrinsic hallucination, and false negatives, and elucidate how RAG mitigates each. These findings demonstrate that reranking-based retrieval substantially enhances the performance of RAG pipelines for clinical information extraction. It offers a practical approach for improving automated analysis of unstructured clinical text. Our four-stage experimental workflow - document indexing, context retrieval, LLM generation, and structured output formatting - provides a replicable framework for future clinical information extraction research and downstream predictive modelling.
Hengyi Zhang, Dinithi Vithanage, Chao Deng 0009, Ping Yu 0004
J. Biomed. Informatics4
2025 Empowering large language models for automated clinical assessment with generation-augmented retrieval and hierarchical chain-of-thought
abstract
BACKGROUND: Understanding and extracting valuable information from electronic health records (EHRs) is important for improving healthcare delivery and health outcomes. Large language models (LLMs) have demonstrated significant proficiency in natural language understanding and processing, offering promises for automating the typically labor-intensive and time-consuming analytical tasks with EHRs. Despite the active application of LLMs in the healthcare setting, many foundation models lack real-world healthcare relevance. Applying LLMs to EHRs is still in its early stage. To advance this field, in this study, we pioneer a generation-augmented prompting paradigm "GAPrompt" to empower generic LLMs for automated clinical assessment, in particular, quantitative stroke severity assessment, using data extracted from EHRs. METHODS: The GAPrompt paradigm comprises five components: (i) prompt-driven selection of LLMs, (ii) generation-augmented construction of a knowledge base, (iii) summary-based generation-augmented retrieval (SGAR); (iv) inferencing with a hierarchical chain-of-thought (HCoT), and (v) ensembling of multiple generations. RESULTS: GAPrompt addresses the limitations of generic LLMs in clinical applications in a progressive manner. It efficiently evaluates the applicability of LLMs in specific tasks through LLM selection prompting, enhances their understanding of task-specific knowledge from the constructed knowledge base, improves the accuracy of knowledge and demonstration retrieval via SGAR, elevates LLM inference precision through HCoT, enhances generation robustness, and reduces hallucinations of LLM via ensembling. Experiment results demonstrate the capability of our method to empower LLMs to automatically assess EHRs and generate quantitative clinical assessment results. CONCLUSION: Our study highlights the applicability of enhancing the capabilities of foundation LLMs in medical domain-specific tasks, i.e., automated quantitative analysis of EHRs, addressing the challenges of labor-intensive and often manually conducted quantitative assessment of stroke in clinical practice and research. This approach offers a practical and accessible GAPrompt paradigm for researchers and industry practitioners seeking to leverage the power of LLMs in domain-specific applications. Its utility extends beyond the medical domain, applicable to a wide range of fields.
Zhanzhong Gu, Wenjing Jia, Massimo Piccardi, Ping Yu 0004
Artif. Intell. Medicine4
2025 Aeromagnetic Compensation Based on Deep Transfer Learning Combined With Physical Constraints
Yuzhuo Zhao, Ping Yu 0004, Pengyu Lu
IEEE Trans. Geosci. Remote. Sens.3
2024 Automatic quantitative stroke severity assessment based on Chinese clinical named entity recognition with domain-adaptive pre-trained large language model
abstract
BACKGROUND: Stroke is a prevalent disease with a significant global impact. Effective assessment of stroke severity is vital for an accurate diagnosis, appropriate treatment, and optimal clinical outcomes. The National Institutes of Health Stroke Scale (NIHSS) is a widely used scale for quantitatively assessing stroke severity. However, the current manual scoring of NIHSS is labor-intensive, time-consuming, and sometimes unreliable. Applying artificial intelligence (AI) techniques to automate the quantitative assessment of stroke on vast amounts of electronic health records (EHRs) has attracted much interest. OBJECTIVE: This study aims to develop an automatic, quantitative stroke severity assessment framework through automating the entire NIHSS scoring process on Chinese clinical EHRs. METHODS: Our approach consists of two major parts: Chinese clinical named entity recognition (CNER) with a domain-adaptive pre-trained large language model (LLM) and automated NIHSS scoring. To build a high-performing CNER model, we first construct a stroke-specific, densely annotated dataset "Chinese Stroke Clinical Records" (CSCR) from EHRs provided by our partner hospital, based on a stroke ontology that defines semantically related entities for stroke assessment. We then pre-train a Chinese clinical LLM coined "CliRoberta" through domain-adaptive transfer learning and construct a deep learning-based CNER model that can accurately extract entities directly from Chinese EHRs. Finally, an automated, end-to-end NIHSS scoring pipeline is proposed by mapping the extracted entities to relevant NIHSS items and values, to quantitatively assess the stroke severity. RESULTS: Results obtained on a benchmark dataset CCKS2019 and our newly created CSCR dataset demonstrate the superior performance of our domain-adaptive pre-trained LLM and the CNER model, compared with the existing benchmark LLMs and CNER models. The high F1 score of 0.990 ensures the reliability of our model in accurately extracting the entities for the subsequent automatic NIHSS scoring. Subsequently, our automated, end-to-end NIHSS scoring approach achieved excellent inter-rater agreement (0.823) and intraclass consistency (0.986) with the ground truth and significantly reduced the processing time from minutes to a few seconds. CONCLUSION: Our proposed automatic and quantitative framework for assessing stroke severity demonstrates exceptional performance and reliability through directly scoring the NIHSS from diagnostic notes in Chinese clinical EHRs. Moreover, this study also contributes a new clinical dataset, a pre-trained clinical LLM, and an effective deep learning-based CNER model. The deployment of these advanced algorithms can improve the accuracy and efficiency of clinical assessment, and help improve the quality, affordability and productivity of healthcare services.
Zhanzhong Gu, Xiangjian He, Ping Yu 0004, Wenjing Jia, Xiguang Yang, Penghui Hu, Shiyan Chen, Yiguang Lin
Artif. Intell. Medicine3
2024 Unsupervised Signal Anomaly Transformer method: Achieving bearing life anomaly detection without the need for failure samples
Ping Yu 0004, Mengmeng Ping, Jialin Ma, Jie Cao 0014
Eng. Appl. Artif. Intell.1
2024 Applying generative AI with retrieval augmented generation to summarize and extract key clinical information from electronic health records
abstract
BACKGROUND: Malnutrition is a prevalent issue in aged care facilities (RACFs), leading to adverse health outcomes. The ability to efficiently extract key clinical information from a large volume of data in electronic health records (EHR) can improve understanding about the extent of the problem and developing effective interventions. This research aimed to test the efficacy of zero-shot prompt engineering applied to generative artificial intelligence (AI) models on their own and in combination with retrieval augmented generation (RAG), for the automating tasks of summarizing both structured and unstructured data in EHR and extracting important malnutrition information. METHODOLOGY: We utilized Llama 2 13B model with zero-shot prompting. The dataset comprises unstructured and structured EHRs related to malnutrition management in 40 Australian RACFs. We employed zero-shot learning to the model alone first, then combined it with RAG to accomplish two tasks: generate structured summaries about the nutritional status of a client and extract key information about malnutrition risk factors. We utilized 25 notes in the first task and 1,399 in the second task. We evaluated the model's output of each task manually against a gold standard dataset. RESULT: The evaluation outcomes indicated that zero-shot learning applied to generative AI model is highly effective in summarizing and extracting information about nutritional status of RACFs' clients. The generated summaries provided concise and accurate representation of the original data with an overall accuracy of 93.25%. The addition of RAG improved the summarization process, leading to a 6% increase and achieving an accuracy of 99.25%. The model also proved its capability in extracting risk factors with an accuracy of 90%. However, adding RAG did not further improve accuracy in this task. Overall, the model has shown a robust performance when information was explicitly stated in the notes; however, it could encounter hallucination limitations, particularly when details were not explicitly provided. CONCLUSION: This study demonstrates the high performance and limitations of applying zero-shot learning to generative AI models to automatic generation of structured summarization of EHRs data and extracting key clinical information. The inclusion of the RAG approach improved the model performance and mitigated the hallucination problem.
Mohammad Alkhalaf, Ping Yu 0004, Mengyang Yin, Chao Deng 0009
J. Biomed. Informatics2
2024 An Empirical Study on Automated Test Generation Tools for Java: Effectiveness and Challenges
Xiang-Jun Liu, Ping Yu 0004, Xiao-Xing Ma
J. Comput. Sci. Technol.2
2024 Inverseformer: Dual-Branch Network With Attention Enhancement for Density Interface Inversion
abstract
Density interface inversion directly reveals the density distribution of underground structures, playing a significant role in the field of geophysics. Deep learning, due to its ability to extract complex and nonlinear density structures, has been widely applied to density interface inversion. Current convolution-based neural network methods for inverting density interfaces are limited by the finite receptive field of the convolution paradigm, neglecting long-range dependencies and posing challenges for high-precision inversion of density interfaces. Additionally, when using deep learning for density interface inversion, different channels contain different information, necessitating the effective utilization of this information variation in the channel dimension. To address these issues, this letter proposes a novel dual-path Transformer network for density interface inversion, named Inverseformer. The Local-Global branch introduces Transformer to establish long-range density dependencies and, combined with the U-net network, jointly captures multiscale underground density interface features. The Channel branch introduces a channel attention mechanism for the model to operate attentively between channels, effectively capturing density interface structural changes in the channels. Synthetic model experiments show that the proposed dual-branch network for inverting density interfaces has smaller model errors and better data fitting, resulting in higher inversion accuracy and more stable inversion results. Subsequently, its effectiveness and value were validated in actual data from the Brittany region of France.
Ping Yu 0004, Longran Zhou, Pengyu Lu, Guan-Lin Huang, Fengyi Bi
IEEE Geosci. Remote. Sens. Lett.1
2024 Tensorial multi-view subspace clustering with side constraints for elevator security warning
Huangzhen Xu, Licheng Ruan, Yuzhou Ni, Hongwei Yin, Ping Yu 0004, Xinmin Cheng
Multim. Syst.5
2022 Randoop-TSR: Random-based Test Generator with Test Suite Reduction
abstract
Software testing plays a very important role in the software development process. Automated test generation tools increase the effectiveness and efficiency of software testing, and alleviate the problem of low efficiency caused by writing hand-crafted test cases. However, different test case generation methods vary in the size, code coverage, and fault detection capacity of the automatically-produced test suites. Automated test case generation tool based on random testing, Randoop as a representative, randomly and incrementally generates a large number of method sequences, which gives various possible combinations of calling methods, but the size of the test suite is not proportional to test quality. Therefore, there exists a lot of redundancy in the test cases. This paper proposes Randoop-TSR, an approach for identifying and eliminating redundant test cases on the basis of Randoop to improve the process of test generation. Our approach adopts three strategies to realize the removal of redundancy, namely: (i) similarity-based input sequence selection; (ii) redundant and duplicate assert statements elimination based on test smell detection; (iii) redundant test cases elimination without breaking test requirements (i.e., code coverage and mutation score). Randoop-TSR can eliminate redundancy effectively, and greatly reduce the size of test suites and execution time. Furthermore, our approach improves the efficiency and understandability of test cases while retaining code coverage and mutation score.
Ping Yu 0004
Internetware2
2022 AUGraft: Graft New API Usage into Old Code
abstract
Software libraries are essential in software development process. When a library evolves, client applications that rely on the library APIs are supposed to migrate their code to utilize new or updated features. Code migration is a challenging task, as developers need to spend time and effort searching for and understanding alternate API usage. Some tools collected code migration instances by mining software repositories and guided old code to evolve according to heuristic migration rules. Most of them focused on statement level modification and relied on a high degree of matching between deprecated API usage and migration instances, thus the applicable scenarios were limited. We propose a new approach, AUGraft, which automatically grafts new API usage into old code by matching and editing fine-grained abstract syntax tree elements. By searching 1,881 projects on FDroid, AUGraft collected 442 unique migration instances with Android SDK, of which 418 are correct. We applied AUGraft on 15 projects and got 367 migration results, of which 294 usages are useful, and 225 usages in them are migrated perfectly. We demonstrate AUGraft can improve the utilization of migration instances with high code migration accuracy.
Ping Yu 0004
Internetware2
2022 Real-Time Aeromagnetic Compensation With Compressed and Accelerated Neural Networks
abstract
As neural networks become an increasingly popular technique in the field of aeromagnetic compensation, there is an increasing demand for hardware systems with more computing power. Compared with the linear regression method, applying a neural network to the task of real-time compensation is difficult because of insufficient computing resources in the unmanned aerial vehicle (UAV) flight detection platform. To perform real-time compensation calculations with limited computing resources, we optimized back propagation neural network (OBPNN) through model compression and acceleration. In this study, we found that the most time-consuming part of network training is the iterative updating of the weights in the BPNN interference model. Using transfer learning, we replace the randomly initialized weights (RWs) with pretrained weights, thereby greatly reducing the number of iterations required. We also apply other model compression and acceleration algorithms. In a case study of our new technique, we implement the fast training of the OBPNN on a Raspberry Pi 4B system. This network processes approximately 316 samples per 0.1 s, which is fast enough to complete aeromagnetic compensation in real time.
Ping Yu 0004, Fengyi Bi
IEEE Geosci. Remote. Sens. Lett.2
2022 An Aeromagnetic Compensation Algorithm Based on a Deep Autoencoder
abstract
Magnetic compensation is a necessary step in the aeromagnetic data processing. While the aeromagnetic compensation model is a linear regression model, the multicollinearity of the variables in the model reduces the accuracy of the compensation model. To solve this problem, we propose a deep autoencoder (DAE) aeromagnetic compensation algorithm. The DAE searches the direction of maximum change in the data by using the gradient descent backpropagation algorithm. The special structure of the encoder can compress the representation of the coefficient matrix and extract data features, thereby weakening the correlation between the coefficient matrix variables. The features obtained after dimension reduction are used in the compensation calculation. We validate the DAE algorithm by applying it to data collected by an unmanned aerial vehicle and demonstrate that the DAE magnetic compensation results are better than those of the principal component analysis algorithm.
Ping Yu 0004
IEEE Geosci. Remote. Sens. Lett.1
2021 MOOC Student Dropout Rate Prediction via Separating and Conquering Micro and Macro Information
Jiayin Lin, Geng Sun 0002, Jun Shen 0001, David E. Pritchard, Ping Yu 0004, Tingru Cui, Li Li 0006, Ghassan Beydoun
ICONIP (6)5
2020 A novel parallel accelerated CRPF algorithm
Jinhua Wang 0005, Jie Cao 0014, Ping Yu 0004, Kaijie Huang
Appl. Intell.4
2020 Heterogeneity in clinical research data quality monitoring: A national survey
Lauren Houston, Ping Yu 0004, Allison Martin, Yasmine C. Probst
J. Biomed. Informatics2
2020 From ideal to reality: segmentation, annotation, and recommendation, the vital trajectory of intelligent micro learning
Jiayin Lin, Geng Sun 0002, Tingru Cui, Jun Shen 0001, Ghassan Beydoun, Ping Yu 0004, David E. Pritchard, Li Li 0006, Shiping Chen 0001
World Wide Web7
2019 ParaAim: Testing Android Applications Parallel at Activity Granularity
abstract
Widely used commercial Android applications (apps) turn to be of complex GUIs and hundreds of activities. Testing this kind of apps is challenging. Existing automated testing tools cannot complete the testing of these complex apps in a short time. However, scaling these tools for parallelism to accelerate the testing is not straightforward. In this paper, we borrow the basic concepts from parallel computing, and introduce the parallel testing platform ParaAim. ParaAim partitions an app testing job into a set of tasks at the activity granularity. Starting from a specific entrance activity, ParaAim explores the UI-states of the app and each newly discovered activity spawns a new task. ParaAim dispatches the new task to an idle device and sets it up to the entrance by replaying an event sequence. In this manner, the independent parts of the app are explored simultaneously. We focus on the efficiency of this parallel GUI exploration schema. ParaAim assigns the tasks that have more possibility to find new activities with high priority for better performance. Event sequence minimization is also studied to reduce the time cost of replay, and we design widget fuzzy match technique to handle the state inconsistency issue. We evaluated ParaAim with 20 popular commercial apps on different settings of Android device cluster. The results show that ParaAim scales well and evidently increases the average speed of exploration to nearly 2 times with two devices, and 3 times with four devices, than a single device.
Chun Cao, Ping Yu 0004, Zhiyong Duan, Xiaoxing Ma
COMPSAC (1)3
2019 A Survey of Segmentation, Annotation, and Recommendation Techniques in Micro Learning for Next Generation of OER
abstract
With the fast development of Internet technologies and mobile devices, space-time boundaries of people's daily activities become blurred. This makes people utilizing their fragmented time become possible. In recent years, micro learning, which aims to make good use of people's fragmented time and deliver micro-format of open education resource to learners, has drawn wider attention. The generation and delivery of micro learning materials are two essential steps for the micro learning service. This paper first discusses the characteristics of three significant stages of a sophisticated micro learning system: segmentation, annotation, and recommendation, for learning materials. Then various state-of-the-art techniques for different processing stages are reviewed in this survey. Different segmentation and annotation strategies based on different information sources (such as content and users' interaction) are demonstrated and analysed. Soft computing, transfer learning, reinforcement learning, and context-aware techniques are also compared and discussed for solving different difficulties in recommending scenarios. We contribute this paper as the first work focusing on the three-phased techniques in micro learning.
Jiayin Lin, Geng Sun 0002, Jun Shen 0001, Tingru Cui, Ping Yu 0004, Li Li 0006
CSCWD5
2019 Fast Robustness Prediction for Deep Neural Network
abstract
Deep neural networks (DNNs) have achieved impressive performance in many difficult tasks. However, DNN models are essentially uninterpretable to humans, and unfortunately prone to adversarial attacks, which hinders their adoption in security and safety-critical scenarios. The robustness of a DNN model, which measures its stableness against adversarial attacks, becomes an important topic in both the machine learning and the software engineering communities. Analytical evaluation of DNN robustness is difficult due to the high-dimensionality of inputs, the huge amount of parameters, and the nonlinear network structure. In practice, the degree of robustness of DNNs is empirically approximated with adversarial searching, which is computationally expensive and cannot be applied in resource constrained settings such as embedded computing. In this paper, we propose to predict the robustness of a DNN model for each input with another DNN model, which takes the output of neurons of the former model as input. We train a regression model to encode the connections between output of the penultimate layer of a DNN model and its robustness. With this trained model, the robustness for an input can be predicted instantaneously. Experiments with MNIST and CIFAR10 datasets and LeNet, VGG and ResNet DNN models were conducted to evaluate the efficacy of the proposed approach. The results indicated that our approach achieved 0.05-0.21 mean absolute errors and significantly outperformed confidence and surprise adequacy-based approaches.
Yue-Huan Wang, Zenan Li, Jingwei Xu 0001, Ping Yu 0004, Xiaoxing Ma
Internetware4
2018 Defining and Developing a Generic Framework for Monitoring Data Quality in Clinical Research
Lauren Houston, Ping Yu 0004, Allison Martin, Yasmine C. Probst
AMIA2
2018 Accelerating Automated Android GUI Exploration with Widgets Grouping
abstract
Ensuring the quality of mobile applications (apps) needs to explore the GUI thoroughly. In practice, exhaustively exploring every GUI widget is unscalable on large real-world apps since it usually suffers from the problem of widgets explosion. To mitigate the problem, many existing testing tools usually detect and group homogeneous widgets heuristicly with different level of model abstraction since these widgets behave the same. However, no heuristic always works well. Heterogeneous widgets with divergent behaviors can be mistakenly grouped, which largely limits the testing effectiveness. This paper proposes a technique to effective GUI testing of Android apps with dynamic feedback-directed widgets grouping. Initially, we group the widgets according to the structure of the GUI. During testing, we observe behaviors of widgets in a group and regroup improperly-grouped widgets dynamically. Then, we apply a feedback-directed strategy to effectively accelerate the GUI exploration. The proposed technique is implemented as a practical tool for Android apps, named WGDroid. We evaluated WGDroid on 17 widely-used Android apps and compared it with the state-of-the-art GUI testing tools, i.e., AimDroid, SAPIENZ, and Monkey on both emulators and real devices. WGDroid outperformed the three tools in all testing coverages and also detected the most unique crashes. In particular, WGDroid discovered 208 more activities on 12 large benchmark apps on real devices and 11 more activities on another 5 benchmark apps on emulators, than the best of the other tools. These results show that WGDroid can significantly accelerate the GUI exploration.
Chun Cao, Hongjun Ge, Tianxiao Gu, Ping Yu 0004, Jian Lu 0001
APSEC5
2018 ReScue: crafting regular expression DoS attacks
abstract
Regular expression (regex) with modern extensions is one of the most popular string processing tools. However, poorly-designed regexes can yield exponentially many matching steps, and lead to regex Denial-of-Service (ReDoS) attacks under well-conceived string inputs. This paper presents Rescue, a three-phase gray-box analytical technique, to automatically generate ReDoS strings to highlight vulnerabilities of given regexes. Rescue systematically seeds (by a genetic search), incubates (by another genetic search), and finally pumps (by a regex-dedicated algorithm) for generating strings with maximized search time. We implemenmted the Rescue tool and evaluated it against 29,088 practical regexes in real-world projects. The evaluation results show that Rescue found 49% more attack strings compared with the best existing technique, and applying Rescue to popular GitHub projects discovered ten previously unknown ReDoS vulnerabilities.
Yuju Shen, Yanyan Jiang 0001, Chang Xu 0001, Ping Yu 0004, Xiaoxing Ma, Jian Lu 0001
ASE4
2018 Mining API usage change rules for software framework evolution
Ping Yu 0004, Chun Cao, Hao Hu 0001, Xiaoxing Ma
Sci. China Inf. Sci.1
2017 Xdroid: Testing Android Apps with Dependency Injection
abstract
The applications ("apps") running on Android need to be adequately tested to avoid faults. Researchers have developed a number of test input generation tools for automated app testing and tried to improve test coverage to detect as many faults as possible. However, existing testing tools achieve very low coverage for some specific apps because they highly depend on external factors to run properly such as business logic, content providers and so on. In this paper, we present Xdroid to catch when and what kind of dependencies apps require and inject them correspondingly in a lightweight way. Working with a built-in tool Xmonkey which generates GUI events directly on Android devices, Xdroid implements an effective testing engine to get a high coverage. We evaluate Xdroid with diverse Android apps and demonstrate that it outperforms Monkey for 17%, Sapienz for 22% in coverage and meanwhile reveals more bugs than manual testing. Overall, it combines the benefits of both manual testing and random testing to improve test coverage and detect bugs effectively.
Chun Cao, Chenglin Meng, Hongjun Ge, Ping Yu 0004, Xiaoxing Ma
COMPSAC (1)4
2017 API Usage Change Rules Mining based on Fine-grained Call Dependency Analysis
abstract
Software frameworks are widely used in application development. But APIs of a framework may change when it evolves to accommodate new feature requests or to fix bugs. Those changes may break existing client programs of the framework, so client programs need to be migrated to the updated release when the framework evolves. Some technologies (e.g. call dependency analysis) have been proposed to find replacement APIs between the old and new framework releases. However, existing approaches based on call dependency analysis take whole method body as an analysis unit. The context in which a method is called is ignored. In this paper, we present a fine-grained approach named AUC-Miner to infer API usage change rules between two releases of the framework. To take method invocation context into consideration, we propose an approach to get more precise call relationship changes by code splitting. We also analyze indirect method invocations to re-fine call dependency analysis. After elaborating API usage change transactions, we adopt frequent item-set mining to generate API replacement rules. Text similarity and some heuristics to identify evolution of root methods are also applied in the mining progress. The evaluation of AUC-Miner on three popular frameworks shows that its precision is higher than basic call dependency analysis and another API replacement recommendation tool named AURA.
Ping Yu 0004, Chun Cao, Hao Hu 0001, Xiaoxing Ma
Internetware1
2017 How effectively can spreadsheet anomalies be detected: An empirical study
Ruiqing Zhang, Chang Xu 0001, Shing-Chi Cheung, Ping Yu 0004, Xiaoxing Ma, Jian Lu 0001
J. Syst. Softw.4
2016 Suppressing detection of inconsistency hazards with pattern learning
Chang Xu 0001, Wenhua Yang 0001, Xiaoxing Ma, Ping Yu 0004, Jian Lu 0001
Inf. Softw. Technol.5
2016 SIT: Sampling-based interactive testing for self-adaptive apps
Yi Qin 0002, Chang Xu 0001, Ping Yu 0004, Jian Lu 0001
J. Syst. Softw.3
2015 Versioning Distributed Transactions for Dynamic Component Reconfiguration
abstract
Dynamic reconfiguration enables components of a distributed system to be updated without restarting the whole system. One major challenge of this technology lies in how to preserve the system consistency while minimizing the disruption. Ma et al. proposed an approach to meet the challenge by managing runtime dependencies among components. However, this approach puts heavy burden on network communication as it requires components to send out every dependency changing event in each ongoing transaction, which is unnecessary in the concurrent environment. In this paper, we propose a transaction-versioning approach that identifies correct version of components for all ongoing transactions. The process can be completed within one single round of message changing among involved components. Comparing with Ma's approach, the communication overhead of our approach is linear with the system scale, which is greatly reduced especially when multiple transactions are running on the system, while the timeliness is not affected. Extensive experiments through simulation show the correctness and efficiency of our approach.
Chun Cao, Huating Liu, Ziling Lu, Ping Yu 0004
APSEC4
2015 ABC: Accelerated Building of C/C++ Projects
abstract
Software building is recurring and time-consuming. Based on the finding that a significant portion of compilations in incremental build is unnecessary, we propose by path compilation, an efficient build technique that avoids unnecessary recompilation with automated detection of redundant dependencies and unessential changes in source files. The technique is lightweight and transparent to software developers, and can be easily applied to existing build systems. We evaluated our approach on a set of real-world open source projects. The results show that 83% ~ 97% of the recompilations are unnecessary and our approach can accelerate the incremental build up to 44.20%.
Ying Zhang 0071, Yanyan Jiang 0001, Chang Xu 0001, Xiaoxing Ma, Ping Yu 0004
APSEC5
2015 ConRec: A Software Framework for Context-Aware Recommendation Based on Dynamic and Personalized Context
abstract
Contextual information is proven helpful to recommender system. And context-aware recommender system(CARS) has been applied in various applications. To improve the accuracy of context-aware recommendation and make recommender application development easier, we develop a lightweight software framework named ConRec, which introduces a dynamic context oriented approach to extend traditional reduction based recommender. This framework takes the dynamic nature of context into full consideration from different aspects to get better recommendation result. The dynamism of context exists in the process of context modeling, the computation of context weight and the handling of newly emergent context. In ConRec, context is dynamically modeled by clustering similar context values into one set automatically, rather than statically predefined by domain experts. Users' preferences to different types of context are explicitly measured through context weighting function based on real dataset. Moreover, ConRec supports incrementally adding new type of context to recommendation process, which reduces much cost of re-building the whole recommender model. Based on our improved reduction-based algorithm, ConRec is built as a highly scalable and reusable software framework for developing context-aware recommender applications. Finally, we evaluate our proposed approach on public datasets and get more accurate recommendation than traditional methods.
Ping Yu 0004, Chun Cao, Feng Xu 0007, Jian Lu 0001
COMPSAC2
2015 An Event-Based Formal Framework for Dynamic Software Update
abstract
Dynamic Software Update (DSU) is a technique to upgrade running programs without shutting them down. DSU can improve system availability and maintenance flexibility. However, its adoption in practice is still limited due to the risk of system misbehavior that careless DSU may bring. To reduce this risk we propose a formal framework for the specification and verification of DSU. Different from previous approaches where DSU is described from the viewpoint of program's internal state transitions, our framework focuses on program's external behavior and its effect on its environment. This more abstract view avoids over specification of DSU and allows for better DSU flexibility. Based on this framework, we also devise a mechanism that automatically synthesizes runtime monitors to improve DSU timeliness without compromising its safety.
Shengwei An, Xiaoxing Ma, Chun Cao, Ping Yu 0004, Chang Xu 0001
QRS4
2014 SHAP: Suppressing the Detection of Inconsistency Hazards by Pattern Learning
abstract
Context-aware applications rely on contexts derived from sensory data to adapt their behavior. However, contexts can be inconsistent and cause application anomaly or crash. One popular solution is to detect and resolve context inconsistencies at runtime. However, we observe that many detected inconsistencies do not indicate real context problems. Instead, they are caused by improper inconsistency detection. These inconsistencies are harmless, and their resolution is unnecessary or may even cause new problems. We name them inconsistency hazards. Inconsistency hazards should be suppressed, but their occurrences resemble normal inconsistencies. In this paper, we present a pattern-learning based approach SHAP to suppressing the detection of inconsistency hazards. Our key insight is that the detection of such hazards is subject to certain patterns of context changes. These patterns, although difficult to specify manually, can be learned effectively from historical inconsistency detection data. We evaluated our SHAP experimentally through three context-aware applications. The results reported that SHAP can automatically suppress the detection of over 90% inconsistency hazards, while preserving the detection of over 98% normal inconsistencies, with only negligible overhead.
Chang Xu 0001, Wenhua Yang 0001, Ping Yu 0004, Xiaoxing Ma, Jiang Lu
APSEC (1)4
2013 Toward a seamless adaptation platform for Internetware
Chun Cao, Ping Yu 0004, Hao Hu 0001, Jian Lu 0001
Sci. China Inf. Sci.2
2013 Application mobility in pervasive computing: A survey
Ping Yu 0004, Xiaoxing Ma, Jiannong Cao 0001, Jian Lu 0001
Pervasive Mob. Comput.1
2010 The Similarity of Video Based on the Association Graph Construction of Video Objects
Ping Yu 0004
ICCCI (2)1
2009 ARTEMIS: an open coordination middleware system
abstract
This demo displays the use of the prototypical ARTEMIS middleware system, which is developed at Nanjing University to support the construction, execution and evolution of applications in the open, dynamic and decentralized network environment of the Internet. To adapt to such a new environment, software application systems must be more flexible, more reactive, and more evolvable than before [1]. Built upon services from autonomous external sources, these application systems also have to explicitly consider the trustworthiness of the services. With these considerations, this version of ARTEMIS middleware is featured by its support for (1) multiple coordination modes based on various software architecture styles; (2) dynamic software architecture-based online reconfigurations, in reaction to the runtime changes in the environment and requirements; (3) trustworthiness evaluation at both the service level and the system level, which helps users to ensure and improve user's satisfaction on the system constructed.
Ping Yu 0004, Chun Cao, Xiaoxing Ma, Jian Lu 0001
Internetware1
2008 On environment-driven software model for Internetware
Jian Lu 0001, Xiaoxing Ma, XianPing Tao, Chun Cao, Yu Huang 0002, Ping Yu 0004
Sci. China Ser. F Inf. Sci.6
2008 Multi-mode interaction middleware for software services
XianPing Tao, Xiaoxing Ma, Jian Lu 0001, Ping Yu 0004, Yu Zhou 0010
Sci. China Ser. F Inf. Sci.4
2007 A Trust Evolution Model for P2P Networks
Ye Tao 0012, Ping Yu 0004, Feng Xu 0007, Jian Lu 0001
ATC3
2007 Constructing Self-Adaptive Systems with Polymorphic Software Architecture
Xiaoxing Ma, Yu Zhou 0010, Ping Yu 0004, Jian Lu 0001
SEKE4
2006 Mobile Agent Enabled Application Mobility for Pervasive Computing
Ping Yu 0004, Jiannong Cao 0001, Weidong Wen, Jian Lu 0001
UIC1
2005 A Cryptographic Solution for General Access Control
Yibing Kong, Jennifer Seberry, Janusz R. Getta, Ping Yu 0004
ISC4
2005 Similarity retrieval of videos by using 3D C-string knowledge representation
Anthony J. T. Lee, Han-Pang Chiu, Ping Yu 0004
J. Vis. Commun. Image Represent.3
2005 3D Z-string: A new knowledge structure to represent spatio-temporal relations between objects in a video
Anthony J. T. Lee, Ping Yu 0004, Han-Pang Chiu, Ruey-Wen Hong
Pattern Recognit. Lett.2
2002 3D C-string: a new spatio-temporal knowledge representation for video database systems
Anthony J. T. Lee, Han-Pang Chiu, Ping Yu 0004
Pattern Recognit.3