EDBT 2026 Demo / reviewers in the wild / expert
Xiaoli Lian
dblp:157/0413
· DBLP profile ↗
26ranked-venue papers
7as first author
18since 2021 · last 2026
0000-0002-6100-7068ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 22 · 7 first-author · 14 since 2021Artificial intelligence and machine learning · 2 · 1 first-authorSystems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Systematic Literature Review of Distributed Multi-Agent Task Allocation: Core Dimensions, Their Interrelationships, and a RepositoryabstractAs a critical and challenging research area within multi-agent systems (MAS), distributed multi-agent task allocation (D-MATA) has motivated extensive study and application across diverse domains. Although several systematic reviews exist on multi-agent task allocation (MATA), none provide an in-depth, systematic analysis of the core dimensions of D-MATA, namely applications, problems, methods, and metrics, nor explore their interrelationships. Moreover, no comprehensive repository capturing such foundational knowledge is currently available. These gaps collectively hinder the effective learning, adoption, and further development of D-MATA knowledge. To fill this gap, we conduct a systematic literature review (SLR) of 107 D-MATA studies, examining them from the perspectives of these core dimensions and discussing the potential interrelationships among these dimensions. To better support comprehensive evaluation and cross-study comparison of D-MATA methods, we propose a multi-dimensional evaluation framework based on the metrics used in the selected literature. Additionally, we provide an open knowledge repository comprising 107 problem-method-evaluation entries derived from the selected literature, supporting reproducible in-depth research. This work delivers clear guidance for both researchers and practitioners while building a systematic knowledge foundation for future investigations in the field. Zitian Yang, Li Zhang 0029, Xiaoli Lian |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2026 | Beyond Functional Correctness: Exploring Hallucinations in LLM-Generated Code
Fang Liu 0032, Yang Liu 0003, Lin Shi 0006, Zhen Yang 0022, Li Zhang 0029, Xiaoli Lian, Zhong-Qi Li, Yuchi Ma |
IEEE Trans. Software Eng. | 6 |
| 2025 | EfficientEdit: Accelerating Code Editing via Edit-Oriented Speculative DecodingabstractLarge Language Models (LLMs) have demonstrated remarkable capabilities in code editing, substantially enhancing software development productivity. However, the inherent complexity of code editing tasks forces existing approaches to rely on LLMs’ autoregressive end-to-end generation, where decoding speed plays a critical role in efficiency. While inference acceleration techniques like speculative decoding are applied to improve the decoding efficiency, these methods fail to account for the unique characteristics of code editing tasks, where changes are typically localized and existing code segments are reused. To address this limitation, we propose EfficientEdit, a novel method that improves LLM-based code editing efficiency through two key mechanisms based on speculative decoding: (1) effective reuse of original code segments while identifying potential edit locations, and (2) efficient generation of edit content via high-quality drafts from edit-oriented draft models and a dynamic verification mechanism that balances quality and acceleration. Experimental results show that EfficientEdit can achieve up to 10.38× and 13.09× speedup compared to standard autoregressive decoding in CanItEdit and CodeIF-Bench, respectively, outperforming state-of-the-art inference acceleration approaches by up to 90.6%. The code and data are available at https://github.com/zhu-zhu-ding/EfficientEdit. Peiding Wang, Li Zhang 0029, Fang Liu 0032, Yinghao Zhu, Lin Shi 0006, Xiaoli Lian, Minxiao Li, An Fu |
ASE | 7 |
| 2025 | AutoPLC: Generating Vendor-Aware Structured Text for Programmable Logic ControllersabstractAmong the programming languages for Programmable Logic Controllers (PLCs), Structured Text (ST) is widely adopted for industrial automation due to its expressiveness and flexibility. However, major vendors implement ST with proprietary extensions and hardware-specific libraries - Siemens’ SCL and CODESYS’ ST each differ in syntax and functionality. This fragmentation forces engineers to relearn implementation details across platforms, creating substantial productivity barriers. To address this challenge, we developed AutoPLC, a framework capable of automatically generating vendor-aware ST code directly from natural language requirements. Our solution begins by building two essential knowledge sources tailored to each vendor’s specifications: a structured API library containing platform-exclusive functions, and an annotated case database that captures real-world implementation experience. Building on these foundations, we created a four-stage generation process that combines step-wise planning (enhanced with a lightweight natural language state machine support for control logic), contextual case retrieval using LLM-based reranking, API recommendation guided by industrial data, and dynamic validation through direct interaction with vendor IDEs. Implemented for Siemens TIA Portal and the CODESYS platform, AutoPLC achieves 90%+ compilation success on our 914-task benchmark (covering general-purpose and process control functions), outperforming all selected baselines, at an average cost of only $0.13 per task. Experienced PLC engineers positively assessed the practical utility of the generated code, including cases that failed compilation. Donghao Yang, Aolang Wu, Li Zhang 0029, Xiaoli Lian, Fang Liu 0032, Yuming Ren, Jiaji Tian, Xiaoyin Che |
ASE | 5 |
| 2025 | FastCoder: Accelerating Repository-level Code Generation via Efficient Retrieval and VerificationabstractCode generation is a latency-sensitive task that demands high timeliness. However, with the growing interest and inherent difficulty in repository-level code generation, most existing code generation studies focus on improving the correctness of generated code while overlooking the inference efficiency, which is substantially affected by the overhead during LLM generation. Although there has been work on accelerating LLM inference, these approaches are not tailored to the specific characteristics of code generation; instead, they treat code the same as natural language sequences and ignore its unique syntax and semantic characteristics, which are also crucial for improving efficiency. Consequently, these approaches exhibit limited effectiveness in code generation tasks, particularly for repository-level scenarios with considerable complexity and difficulty. To alleviate this issue, following draft-verification paradigm, we propose FastCoder, a simple yet highly efficient inference acceleration approach specifically designed for code generation, without compromising the quality of the output. FastCoder constructs a multi-source datastore, providing access to both general and project-specific knowledge, facilitating the retrieval of high-quality draft sequences. Moreover, FastCoder reduces the retrieval cost by controlling retrieval timing, and enhances efficiency through parallel retrieval and a context- and LLM preference-aware cache. Experimental results show that FastCoder can reach up to 2.53× and 2.54× speedup compared to autoregressive decoding in repository-level and standalone code generation tasks, respectively, outperforming state-of-the-art inference acceleration approaches by up to 88%. FastCoder can also be integrated with existing correctness-focused code generation approaches to accelerate the LLM generation process, and reach a speedup exceeding 2.6×. Qianhui Zhao, Li Zhang 0029, Fang Liu 0032, Xiaoli Lian, Qiaoyuanhe Meng, Ziqian Jiao, Zetong Zhou, Jia Li 0012, Lin Shi 0006 |
ASE | 4 |
| 2025 | Deep learning-based software engineering: progress, challenges, and opportunitiesabstractAbstract Researchers have recently achieved significant advances in deep learning techniques, which in turn has substantially advanced other research disciplines, such as natural language processing, image processing, speech recognition, and software engineering. Various deep learning techniques have been successfully employed to facilitate software engineering tasks, including code generation, software refactoring, and fault localization. Many studies have also been presented in top conferences and journals, demonstrating the applications of deep learning techniques in resolving various software engineering tasks. However, although several surveys have provided overall pictures of the application of deep learning techniques in software engineering, they focus more on learning techniques, that is, what kind of deep learning techniques are employed and how deep models are trained or fine-tuned for software engineering tasks. We still lack surveys explaining the advances of subareas in software engineering driven by deep learning techniques, as well as challenges and opportunities in each subarea. To this end, in this study, we present the first task-oriented survey on deep learning-based software engineering. It covers twelve major software engineering subareas significantly impacted by deep learning techniques. Such subareas spread out through the whole lifecycle of software development and maintenance, including requirements engineering, software development, testing, maintenance, and developer collaboration. As we believe that deep learning may provide an opportunity to revolutionize the whole discipline of software engineering, providing one survey covering as many subareas as possible in software engineering can help future research push forward the frontier of deep learning-based software engineering more systematically. For each of the selected subareas, we highlight the major advances achieved by applying deep learning techniques with pointers to the available datasets in such a subarea. We also discuss the challenges and opportunities concerning each of the surveyed software engineering subareas. Xiangping Chen, Xing Hu 0008, Yuan Huang 0002, He Jiang 0001, Weixing Ji, Yanjie Jiang, Yanyan Jiang 0001, Bo Liu 0094, Hui Liu 0003, Xiaoli Lian, Guozhu Meng, Xin Peng 0001, Hailong Sun 0001, Lin Shi 0006, Bo Wang 0050, Chong Wang 0013, Jifeng Xuan, Xin Xia 0001, Yibiao Yang, Yixin Yang 0006, Li Zhang 0029, Yuming Zhou, Lu Zhang 0023 |
Sci. China Inf. Sci. | 11 |
| 2024 | Navigating the AI Frontier: A Critical Literature Review on Integrating Artificial Intelligence into Software Engineering EducationabstractThe swift development of Artificial Intelligence (AI), namely the introduction of Large Language Models (LLMs), is drastically altering various industries and necessitating a major change in the way software engineering is taught. To equip upcoming software engineers with the knowledge and abilities to function in this AI-powered environment, curriculum and pedagogical techniques must be critically reevaluated. To better understand the integration of AI and LLMs into software engineering education, this study gives a thorough and critical analysis of the literature, looking at existing models, pedagogical frameworks, and enduring issues. We explore various approaches utilized by educational establishments, including as specialized AI and LLM courses, incorporating modules into pre-existing curricula, and utilizing open-source LLM materials. Our analysis, which is based on case studies and research data, thoroughly assesses how well these strategies enable software engineers to comprehend, make use of, and ethically create AI and LLMs. Key obstacles to the successful integration of AI and LLM are also identified by our analysis, including the inexperienced status of LLM educators, resource limitations, potential biases in AI and LLM algorithms, and insufficient instructor knowledge. Building on these discoveries, we provide solid answers to these problems and suggest interesting avenues for further study to improve the integration of AI and LLM. In the end, this study advocates for a multimodal strategy to get future software engineers ready for the impending AI and LLM future and secure their place in this quickly changing field. Chandan Kumar Sah, Xiaoli Lian, Muhammad Mirajul Islam |
CSEE&T | 2 |
| 2024 | DRMiner: Extracting Latent Design Rationale from Jira Issue LogsabstractSoftware architectures are usually meticulously designed to address multiple quality concerns and support long-term maintenance. However, there may be a lack of motivation for developers to document design rationales (i.e., the design alternatives and the underlying arguments for making or rejecting decisions) when they will not gain immediate benefit, resulting in a lack of standard capture of these rationales. With the turnover of developers, the architecture inevitably becomes eroded. This issue has motivated a number of studies to extract design knowledge from open-source communities in recent years. Unfortunately, none of the existing research has successfully extracted solutions alone with their corresponding arguments due to challenges such as the intricate semantics of online discussions and the lack of benchmarks for design rationale extraction. Jiuang Zhao, Zitian Yang, Li Zhang 0029, Xiaoli Lian, Donghao Yang, Xin Tan 0003 |
ASE | 4 |
| 2024 | Enhancing Automated Program Repair with Solution DesignabstractAutomatic Program Repair (APR) endeavors to autonomously rectify issues within specific projects, which generally encompasses three categories of tasks: bug resolution, new feature development, and feature enhancement. Despite extensive research proposing various methodologies, their efficacy in addressing real issues remains unsatisfactory. It's worth noting that, typically, engineers have design rationales (DR) on solution--- planed solutions and a set of underlying reasons---before they start patching code. In open-source projects, these DRs are frequently captured in issue logs through project management tools like Jira. This raises a compelling question: How can we leverage DR scattered across the issue logs to efficiently enhance APR? Jiuang Zhao, Donghao Yang, Li Zhang 0029, Xiaoli Lian, Zitian Yang, Fang Liu 0032 |
ASE | 4 |
| 2024 | ReqCompletion: Domain-Enhanced Automatic Completion for Software RequirementsabstractSoftware requirements are the driving force behind software development. As the cornerstone of the entire software lifecycle, the efficiency of crafting requirement specifications and the quality of these requirements significantly influence the duration of software development. Despite massive research on requirements elicitation, the reality is that requirements are often painstakingly crafted manually, word by word. This manual process is not only time-consuming but also prone to issues such as the misuse of terminology. To address these challenges, we introduce ReqCompletion, an approach designed to recommend the next token in real-time for given prefix of requirements description. ReqCompletion comprises two primary components. First, we have devised and integrated a knowledge-injection module into GPT-2—which stands as the largest available GPT model that allows for fine-tuning on specialized downstream tasks. This injection imbues GPT-2 with richer domain-specific knowledge, thus improving the relevance of the suggested tokens. Additionally, we employ a pointer network to optimize the recommendation quality by utilizing completed requirements as contextual support. Empirical evaluations using two public datasets demonstrate that ReqCompletion surpasses all baselines in performance (Recall@7 gains up to 65.87% than the second-best model). Furthermore, the effectiveness of its two pivotal design elements has been substantiated through rigorous ablation studies. The utility of our work has been evaluated preliminarily through a small user study. Xiaoli Lian, Jieping Ma, Heyang Lv, Li Zhang 0029 |
RE | 1 |
| 2024 | WiAi-ID: Wi-Fi-Based Domain Adaptation for Appearance-Independent Passive Person IdentificationabstractWi-Fi signal-based person identification has become a hot research topic due to the widespread deployment of Wi-Fi devices and the fact that these approaches are noncontact, passive, and privacy-preserving. While the existing related methods and systems have achieved good performance for person identification, they also encounter many significant challenges in practical applications. Due to the propagation properties of Wi-Fi signals, the signal at the receiver will change significantly when the user’s appearance changes. This makes single-appearance trained models unusable for cross-appearance recognition tasks. To address this challenge, we propose a deep learning-based framework for appearance-independent identification using Wi-Fi signals (WiAi-ID), the core of which lies in the fact that the domain discriminator and feature extractor are trained together in an adversarial manner, thus forcing the model to extract identity-inherent features independent of human appearance, and introduces a multiscale CNN adaptation module to capture time-span-based features. We collected Wi-Fi signal data of pedestrians with different appearances. The experimental results show that WiAi-ID can effectively eliminate the impact on identification due to pedestrian appearance variations and accordingly outperforms the current state-of-the-art video and wireless signal-based recognition methods. Haobo Li 0004, Zhengqi Liu, Pengfei Xu 0003, Xiaoli Lian, Xiaojiang Chen |
IEEE Internet Things J. | 7 |
| 2024 | CDTC: Automatically establishing the trace links between class diagrams in design phase and source codeabstractAbstract Context The UML class diagram is commonly used to model functional structures and software code structures in both the preliminary and detailed design stages. And the abstraction level of UML class diagrams is usually higher than that of source code. Usually, there is a lack of trace links between these class diagrams and the source code, which may cause difficulties in understanding the source code, and affect the software evolution and maintenance. Objective The main goal of this article is to establish the trace links between highly abstracted UML class diagrams in the design phase and source code, and eventually help practitioners better understand source code. Method We propose an approach for the automated trace link establishment between UML class diagrams in the design phase and source code. To address the problem of abstraction level gap between them, we extend the UML class diagram by mining the synonymous phrases of class names and deducing the latent missing relationships between classes from multiple design documents. Then we build the trace links with a two‐phase approach including initial construction with fuzzy matching and further optimization by class relationship inference. Results Experiments on five open‐source projects show that the recalls of our approach are over 94%, and the F2‐scores are over 88%, with the gains of 30% to 60% than the four baselines. Conclusion Our work can be a reference for establishing the initial trace links between highly‐abstracted UML class diagrams and source code. Towards the higher abstraction of design diagrams, we extend UML class diagrams with the statistical analysis on multiple design documents. To guarantee the quality of trace links, we design a two‐phase approach by obtaining the “full but not good enough” trace links and filtering the “probably wrong” links. Experiments show that the main techniques of our approach behave as important role for tracing between high‐level class diagrams and source code. Fangwei Chen, Li Zhang 0029, Xiaoli Lian |
Softw. Pract. Exp. | 3 |
| 2024 | Usefulness of open domain model for identifying missing software requirements conceptsabstractSummary Detecting missing requirements during software development is crucial to avoid unexpected consequences. However, this task is challenging due to limited domain knowledge of requirements analysts and the dynamic nature of software requirements. Previous studies have shown that requirement‐oriented domain models can help identify omissions in requirements, but they are often incomplete for many domains. Meanwhile, domain models constructed from other artifacts are available online. This raises the question: Can these domain models be useful in identifying missing functional information in requirement specifications? To address this question, we conducted a study to measure the overlap between entities in domain models and requirements. We analyzed the occurrence of overlapped entities, considering four distribution characteristics: the type of entities in the domain model, the distribution of mapped entities in the domain model, the family belonging of the mapped entities in the domain model, and the distribution of mapped entities in the requirements. Based on our findings, we proposed recommendations for missing requirements. Additionally, we performed experiments, including the use of the proposed metric “ancestors of the highest level with the most mapped entities” (AHME). The results showed significant improvements with gains of 146% and 223% in the two domains, highlighting the benefits of these distribution characteristics. Li Zhang 0029, Xiaoli Lian |
Softw. Pract. Exp. | 3 |
| 2024 | DRIP: Segmenting individual requirements from software requirement documentsabstractAbstract Numerous academic research projects and industrial tasks related to software engineering require individual requirements as input. Unfortunately, according to our observation, several requirements may be packed in one paragraph without explicit boundaries in specification documents. To understand this problem's prevalence, we performed a preliminary study on the open requirement documents widely used in the academic community over the last 10 years, and found that 26% of them include this phenomenon. Several text segmentation approaches have been reported; however, they tend to identify topically coherent units which may contain more than one requirement. What is more, they do not take the constitutions of semantic units of requirements into consideration. Here we report a two‐phase learning‐based approach named DRIP to segment individual requirements from paragraphs. To be specific, we first propose a Requirement Segmentation Siamese framework, which models the similarity of sentences and their conjunction relations, and then detects the initial boundaries between individual requirements. Then, we optimize the boundaries heuristically based on the semantic completeness validation of the segments. Experiments with 1132 paragraphs and 6826 sentences show that DRIP outperforms the popular unsupervised and supervised text segmentation algorithms with respect to processing different documents (with accuracy gains of 57.65%–187.53%) and processing paragraphs of different complexity (with average accuracy gains of 54.46%–158.68%). We also show the importance of each component of DRIP to the segmentation. Li Zhang 0029, Xiaoli Lian, Heyang Lv |
Softw. Pract. Exp. | 3 |
| 2022 | Automatically recognizing the semantic elements from UML class diagram imagesabstractDesign models are essential for multiple tasks in software engineering, such as consistency checking, code generation, and design-to-code tracing. Almost all of these works need a semantically analyzable model to represent the software architecture design, e.g., a UML class diagram. Unfortunately, many design models are stored as images and embedded in text-based documentations, impeding the usage and evolution of these models. Thus, identifying the semantic elements of design models from images is important. However, there are lots of design models with different elements in diverse representations, which ask for different approaches for semantic elements extraction. In order to grasp an overview of the commonly used design model types, we conduct a survey on both open-source communities and industry. We find that design model diagrams are usually embedded in documents as pictures (73.72%), and UML class diagrams are the most used type (55.43%). Considering that there are limited studies on automatically recognizing the semantic elements from class diagram images, we propose an approach, which we call ReSECDI. ReSECDI includes our customized design for extracting UML class diagram elements based on image processing technologies. We design a rectangle clustering method for class recognition, to address the challenge that the presentation of classes may vary due to the UML constraints and tools’ styles. We design a polygonal line merging method and double-recognition-approximation method for relationship recognition to deal with the impact of low resolution on the detection. We evaluate the applicability of ReSECDI on 30 images drawn by three popular UML tools and 50 diagrams collected from the open-source communities, and get promising performances. ReSECDI can recognize all types of semantic elements commonly used. It has well applicability and can be used to process the images drawn by the mainstream tools and stored in different resolutions. Editor’s note: Open Science material was validated by the Journal of Systems and Software Open Science Board. Fangwei Chen, Li Zhang 0029, Xiaoli Lian, Nan Niu |
J. Syst. Softw. | 3 |
| 2021 | Putting software requirements under the microscope: automated extraction of their semantic elementsabstractThe relationships between software requirements work as the basis for several important software activities, such as change impact and developing cost analysis. Multiple types of relationships are mentioned in the RE literatures including normal (e.g., dependency) and abnormal ones (e.g., conflicts), and most of the existing work usually focus on the identification of one specific relationship. We collect and analyze the relations in the RE literatures, and find some common semantic elements of functional requirements are involved in the definition of multiple types of relations. Thus, to support automatically identifying diverse relationships, we propose our definition of the micro-level semantic constitution of functional requirement (M-FRDL), and one automatic approach for the element extraction, named by Micro-level Semantic elements Analyser of functional requirement (MISA). The experiments with three open requirement datasets show that our MISA can correctly identify about 94.93% elements of requirements on average. Weize Guo, Li Zhang 0029, Xiaoli Lian |
RE | 3 |
| 2021 | What can Open Domain Model Tell Us about the Missing Software Requirements: A Preliminary StudyabstractCompleteness is one of the most important attributes of software requirement specification. Unfortunately, incompleteness is one of the most difficult violations to detect. Some approaches have been proposed to detect missing requirements based on the requirement-oriented domain model. However, these kinds of models are actually lack for lots of domains. Fortunately, the domain models constructed for different purposes can usually be found online. This raises a question: whether or not these domain models are useful for finding the missing functional information in requirement specification? To explore this question, we design and conduct a preliminary study by computing the overlapping rate between the entities in domain models and the concepts of natural language software requirements, and then digging into four regularities of the occurrence of these entities(concepts) based on two example domains. The usefulness of these regularities, especially the one based our proposed metric AHME (with 54% and 70% of F2on the two domains), has been initially evaluated with an additional experiment. Li Zhang 0029, Xiaoli Lian |
RE | 3 |
| 2021 | A systematic gray literature review: The technologies and concerns of microservice application programming interfacesabstractAbstract The microservice application programming interface (API) becomes a growing concern in the IT industry, as a result of the increasing usage of microservice architecture style. There exist many successful practices among companies, communities, and so on. In contrast, the related academic research is still at an early stage, where lacks an overview of technologies for the design, implementation and operation of microservice APIs, as well as a general picture of concerns. In this article, we try to fill this gap by eliciting the technologies and concerns on microservice APIs and establishing a microservice API description model, with the intention of aiding researchers to gain an overview of this field and find possible research directions, and helping practitioners to better understand microservice APIs and be aware of the existing approaches for daily work. Twelve academic papers and 38 gray literatures are selected and analyzed following the systematic literature review approach. Besides, we give our observations from this study. For researchers, our findings show the most cared concerns of practitioners, and our description model can be used as a reference for new theories, experiments, and future research dimensions. For practitioners, our study can be used as a guideline for microservices experimentation and a starting point for practice. Fangwei Chen, Li Zhang 0029, Xiaoli Lian |
Softw. Pract. Exp. | 3 |
| 2020 | An improved mapping method for automated consistency check between software architecture and source codeabstractIn daily software development, inconsistencies between architecture and code inevitably occur with the continuous contribution, even under model-driven development which can trace between design and code. Many methods have been proposed for consistency checking, but most require huge human efforts on establishing the mappings between architectural and code elements. Besides, the multi-layered architecture and code increases the difficulties in inconsistency detection, while existing algorithms do not handle this well. Thus, we propose an improved mapping method for automated consistency check between software architecture and Java implementation, with the premises that initial tracing between architecture and code are established. To be specific, during software evolution, our method can automatically re-establish the mappings between architecture and code using initial tracing information. Then, with detailed inconsistency check rules, we detect the inconsistencies heuristically. Experiments with two projects show our method's high effectiveness with more than 98% of recall and 96% of precision. Fangwei Chen, Li Zhang 0029, Xiaoli Lian |
QRS | 3 |
| 2020 | Assisting engineers extracting requirements on components from domain documents
Xiaoli Lian, Wenchuang Liu, Li Zhang 0029 |
Inf. Softw. Technol. | 1 |
| 2018 | An approach for optimized feature selection in large-scale software product lines
Xiaoli Lian, Li Zhang 0029, Jing Jiang 0005, William Goss |
J. Syst. Softw. | 1 |
| 2017 | Mining Associations Between Quality Concerns and Functional RequirementsabstractThe cost and effort of developing software systems in a new technical area can be extensive. An organization must perform a domain analysis to discover competing products, analyze their architectures and features, and ultimately discover and specify product requirements. However, delivering high quality products, depends not only on gaining an understanding of functional requirements, but also of qualities such as performance, reliability, security, and usability. Discovering such concerns early in the requirements process drives architectural design decisions. This paper extends our prior work on mining functional requirements from large collections of domain documents, by proposing and evaluating a new technique for discovering and specifying quality concerns related to specific functional components. We evaluate our approach against three domains of Positive Train Control, Electronic Health Records, and Medical Infusion Pumps, and show that it significantly outperforms a basic information retrieval approach. Finally we classified the forms of retrieved information, discussed the utility of different types, and conducted a small study with an experienced engineer to investigate the quality of requirements produced using our approach. Xiaoli Lian, Jane Cleland-Huang, Li Zhang 0029 |
RE | 1 |
| 2016 | Mining Requirements Knowledge from Collections of Domain DocumentsabstractWhen organizations enter domains that are entirely new to them, they need to invest significant time and effort to acquire domain knowledge. This typically involves searching through a broad set of domain documents, retrieving relevant ones, and analyzing the textual content in order to discover and specify pertinent requirements. Depending on the nature of the domain and the availability of documentation, this task can be extremely time-consuming and may require non-trivial human effort. Furthermore, the task must often be performed repeatedly throughout early phases of the project. In this paper we first explore the effort needed to manually build a high-level domain model capturing the functional components. We then present MaRK (Mining Requirements Knowledge), which identifies and retrieves the documents containing descriptions of functional components in the domain model. Domain analysts can use this information to to specify requirements. We introduce and evaluate an algorithm which ranks domain documents according to their relevance to a component and then highlights sections of text which are likely to contain requirements-related information. We describe our process within the context of the Positive Train Control (PTC) domain with a repository of of 523 documents, representing 852MB of data. We empirically evaluate the MaRK relevance algorithm and its ability to retrieve relevant requirements knowledge for requirements related to PTC's On-Board Unit. Xiaoli Lian, Mona Rahimi, Jane Cleland-Huang, Li Zhang 0029, Remo Ferrai, Michael Smith 0025 |
RE | 1 |
| 2016 | Long-Term Active Integrator Prediction in the Evaluation of Code ContributionsabstractIn open source software (OSS) projects, integrators are given high-level access to repositories so that they could maintain and manage projects.Although integrators play a critical role in evaluating code changes for OSS projects, they may be short-term active.Long-term active integrators keep in evaluating code update submission and managing responses from contributors.In order to survive and succeed, OSS projects need to attract and retain long-term active integrators.To assist OSS projects to retain active integrators, we propose a method called LTAPredict to predict whether integrators will be longterm active in the evaluation of code contributions.LTAPredict collects activity data of integrators, extracts a rich set of features, and makes prediction via machine learning techniques.We perform experiments on 37 popular projects, containing a total of 1,073 integrators.Results show that based on the Decision Tree, LTAPredict achieves the accuracy as 0.829, the precision as 0.81, the recall as 0.827 and the F1 as 0.818.Meanwhile, we evaluate the feature importance to identify the most significant indicators of long-term active integrators.We observe that whether integrators becoming long-term active is associated with the number of active months and social distance with contributors in their first year as integrators.These findings assist OSS projects to identify potential long-term active integrators and adopt better strategies to retain them in the evaluation of code contributions. Jing Jiang 0005, Fuli Feng, Xiaoli Lian, Li Zhang 0029 |
SEKE | 3 |
| 2015 | Optimized feature selection towards functional and non-functional requirements in Software Product LinesabstractAs an important research issue in software product line, feature selection is extensively studied. Besides the basic functional requirements (FRs), the non-functional requirements (NFRs) are also critical during feature selection. Some NFRs have numerical constraints, while some have not. Without clear criteria, the latter are always expected to be the best possible. However, most existing selection methods ignore the combination of constrained and unconstrained NFRs and FRs. Meanwhile, the complex constraints and dependencies among features are perpetual challenges for feature selection. To this end, this paper proposes a multi-objective optimization algorithm IVEA to optimize the selection of features with NFRs and FRs by considering the relations among these features. Particularly, we first propose a two-dimensional fitness function. One dimension is to optimize the NFRs without quantitative constraints. The other one is to assure the selected features satisfy the FRs, and conform to the relations among features. Second, we propose a violation-dominance principle, which guides the optimization under FRs and the relations among features. We conducted comprehensive experiments on two feature models with different sizes to evaluate IVEA with state-of-the-art multi-objective optimization algorithms, including IBEAHD, IBEAε+, NSGA-II and SPEA2. The results showed that the IVEA significantly outperforms the above baselines in the NFRs optimization. Meanwhile, our algorithm needs less time to generate a solution that meets the FRs and the constraints on NFRs and fully conforms to the feature model. Xiaoli Lian, Li Zhang 0029 |
SANER | 1 |
| 2014 | An Evolutionary Methodology for Optimized Feature Selection in Software Product Lines
Xiaoli Lian, Li Zhang 0029 |
SEKE | 1 |