Jiasu Sun

dblp:76/2820 · DBLP profile ↗
← Back
33ranked-venue papers
0as first author
0since 2021 · last 2013
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 30Applied, interdisciplinary, general and emerging computing · 4Artificial intelligence and machine learning · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
11 papers
Software maintenance and evolution · 40% Program analysis · 28% Software testing · 15%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 18 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Software maintenance and evolution
software internationalization
0.542013
Locating Need-to-Externalize Constant Strings for Software Internationalization with Generalized String-Taint Analysis · IEEE Trans. Software Eng. 2013
Locating need-to-translate constant strings in web applications · SIGSOFT FSE 2010
TranStrL: An automatic need-to-translate string locator for software internationalization · ICSE 2009
Program analysis
static analysis
0.542013
Locating Need-to-Externalize Constant Strings for Software Internationalization with Generalized String-Taint Analysis · IEEE Trans. Software Eng. 2013
Locating need-to-translate constant strings in web applications · SIGSOFT FSE 2010
TranStrL: An automatic need-to-translate string locator for software internationalization · ICSE 2009
Software testing
model-based testing
0.212013
Effective Message-Sequence Generation for Testing BPEL Programs · IEEE Trans. Serv. Comput. 2013
Program analysis › static analysis
taint analysis
0.212013
Locating Need-to-Externalize Constant Strings for Software Internationalization with Generalized String-Taint Analysis · IEEE Trans. Software Eng. 2013
Software testing
test generation
0.212013
Effective Message-Sequence Generation for Testing BPEL Programs · IEEE Trans. Serv. Comput. 2013
Debugging and program repair
fault localization
0.122009
VIDA: Visual interactive debugging · ICSE 2009
A similarity-aware approach to testing based fault localization · ASE 2005
Software maintenance and evolution
feature location
0.122006
SNIAFL: Towards a static noninteractive approach to feature location · ACM Trans. Softw. Eng. Methodol. 2006
SNIAFL: Towards a Static Non-Interactive Approach to Feature Location · ICSE 2004
Debugging and program repair › software debugging
interactive debugging
0.112009
VIDA: Visual interactive debugging · ICSE 2009
Software maintenance and evolution › bug triage
bug report management
0.112008
An approach to detecting duplicate bug reports using natural language and execution information · ICSE 2008
Software maintenance and evolution › bug triage
duplicate bug report detection
0.112008
An approach to detecting duplicate bug reports using natural language and execution information · ICSE 2008
Software maintenance and evolution
program comprehension
0.112008
An objective-oriented approach to program comprehension using multiple information sources · Sci. China Ser. F Inf. Sci. 2008
Software maintenance and evolution › feature location
information retrieval-based feature location
0.112006
SNIAFL: Towards a static noninteractive approach to feature location · ACM Trans. Softw. Eng. Methodol. 2006
Services computing and microservices › service orchestration
BPEL
0.012013
Effective Message-Sequence Generation for Testing BPEL Programs · IEEE Trans. Serv. Comput. 2013
Services computing and microservices
service composition
0.012013
Effective Message-Sequence Generation for Testing BPEL Programs · IEEE Trans. Serv. Comput. 2013
Information retrieval
retrieval models
0.012004
SNIAFL: Towards a Static Non-Interactive Approach to Feature Location · ICSE 2004
Software maintenance and evolution
bug triage
0.012008
An approach to detecting duplicate bug reports using natural language and execution information · ICSE 2008
Program analysis › static analysis › interprocedural analysis
call graph analysis
0.012006
SNIAFL: Towards a static noninteractive approach to feature location · ACM Trans. Softw. Eng. Methodol. 2006
Software testing › regression testing
test suite reduction
0.012005
A similarity-aware approach to testing based fault localization · ASE 2005

Methods — techniques the papers use, named apart from their topics

information retrieval · 0.2string-taint analysis · 0.2API method analysis · 0.2message-sequence graph · 0.2generalized string-taint analysis · 0.2fault exposure · 0.2static analysis · 0.1flag propagation · 0.1natural language processing · 0.1execution trace analysis · 0.1BRCG · 0.0
YearPublicationVenuePosition
2013 Effective Message-Sequence Generation for Testing BPEL Programs
abstract
With the popularity of Web Services and Service-Oriented Architecture (SOA), quality assurance of SOA applications, such as testing, has become a research focus. Programs implemented by the Business Process Execution Language for Web Services (WS-BPEL), which can be used to compose partner Web Services into composite Web Services, are one popular kind of SOA applications. The unique features of WS-BPEL programs bring new challenges into testing. A test case for testing a WS-BPEL program is a sequence of messages that can be received by the WS-BPEL program under test. Previous research has not studied the challenges of message-sequence generation induced by unique features of WS-BPEL as a new language. In this paper, we present a novel methodology to generate effective message sequences for testing WS-BPEL programs. To capture the order relationship in a message sequence and the constraints on correlated messages imposed by WS-BPEL's routing mechanism, we model the WS-BPEL program under test as a message-sequence graph (MSG), and generate message sequences based on MSG. We performed experiments for our method and two other techniques with six WS-BPEL programs. The results show that the message sequences generated by using our method can effectively expose faults in the WS-BPEL programs.
Yitao Ni, Shan-Shan Hou, Lu Zhang 0023, Zhong Jie Li, Qian Lan, Hong Mei 0001, Jiasu Sun
IEEE Trans. Serv. Comput.8
2013 Locating Need-to-Externalize Constant Strings for Software Internationalization with Generalized String-Taint Analysis
abstract
Nowadays, a software product usually faces a global market. To meet the requirements of different local users, the software product must be internationalized. In an internationalized software product, user-visible hard-coded constant strings are externalized to resource files so that local versions can be generated by translating the resource files. In many cases, a software product is not internationalized at the beginning of the software development process. To internationalize an existing product, the developers must locate the user-visible constant strings that should be externalized. This locating process is tedious and error-prone due to 1) the large number of both user-visible and non-user-visible constant strings and 2) the complex data flows from constant strings to the Graphical User Interface (GUI). In this paper, we propose an automatic approach to locating need-to-externalize constant strings in the source code of a software product. Given a list of precollected API methods that output values of their string argument variables to the GUI and the source code of the software product under analysis, our approach traces from the invocation sites (within the source code) of these methods back to the need-to-externalize constant strings using generalized string-taint analysis. In our empirical evaluation, we used our approach to locate need-to-externalize constant strings in the uninternationalized versions of seven real-world open source software products. The results of our evaluation demonstrate that our approach is able to effectively locate need-to-externalize constant strings in uninternationalized software products. Furthermore, to help developers understand why a constant string requires translation and properly translate the need-to-externalize strings, we provide visual representation of the string dependencies related to the need-to-externalize strings.
Xiaoyin Wang, Lu Zhang 0023, Tao Xie 0001, Hong Mei 0001, Jiasu Sun
IEEE Trans. Software Eng.5
2010 Locating need-to-translate constant strings in web applications
abstract
Software internationalization aims to make software accessible and usable by users all over the world. For a Java application that does not consider internationalization at the beginning of its develop- ment stage, our previous work proposed an approach to locating need-to-translate constant strings in the Java code. However, when being applied on web applications, it can identify only constant strings that may go to the generated HTML texts, but cannot further distinguish constant strings visible at the browser side (need-to-translate) from other constant strings (not need-to-translate). In this paper, to address significant challenges in internationalizing web applications, we propose a novel approach to locating need-to-translate constant strings in web applications. Among those constant strings that may go to the generated HTML texts, our approach further distinguishes strings visible at the browser side from non-visible strings via a novel technique called flag propagation. We evaluated our approach on three real-world open source PHP-based web applications (in total near 17 KLOC): Squirrel Mail, Lime Survey, and Mrbs. The empirical results demonstrate that our approach accurately distinguishes visible strings from non-visible strings among all the constant strings that may go to the generated HTML texts, and is effective for locating need-to-translate constant strings in web applications.
Xiaoyin Wang, Lu Zhang 0023, Tao Xie 0001, Hong Mei 0001, Jiasu Sun
SIGSOFT FSE5
2010 Test input reduction for result inspection to facilitate fault localization
Dan Hao 0001, Tao Xie 0001, Lu Zhang 0023, Xiaoyin Wang, Jiasu Sun, Hong Mei 0001
Autom. Softw. Eng.5
2010 A biting-down approach to hierarchical decomposition of object-oriented systems based on structure analysis
abstract
Abstract System decomposition has been widely viewed as an effective means to facilitate the comprehension of complex software systems and/or capture potentially reusable components in them. In fact, various approaches to system decomposition have been intensively documented in the literature. However, during the process of system decomposition, only a few of them can also capture the target system's hierarchical organization structure, which is essential when the target system is very complex. In this paper, we present a biting‐down approach to hierarchical decomposition of object‐oriented systems. Compared with the previous hierarchical approaches, the distinct features of this approach are as follows. First, our approach does not rely on agglomeration, and thus can avoid some unnecessary calculations. Second, our approach does not require merging nodes when performing high‐level decomposition, and thus can avoid imprecision induced by the merging. To evaluate our approach, we conducted a case study and an experimental study on our approach. The results of these studies can confirm its effectiveness and its superiority over our previous approach. Copyright © 2009 John Wiley & Sons, Ltd.
Lu Zhang 0023, Jiasu Sun, Hong Mei 0001
J. Softw. Maintenance Res. Pract.4
2009 VIDA: Visual interactive debugging
abstract
Software debugging is time-consuming and effort-consuming. Although software debugging, especially fault-localization, has been studied for long, few practical debugging tools have been developed and used by the industry. In this paper we present VIDA, a visual interactive debugging tool, which has been integrated with the Eclipse Integrated Development Environment to support a programmer's debugging process. During the programmer's conventional debugging process, VIDA continuously recommends break-points for the programmer based on the analysis of execution information and the gathered feedback from the programmer. Moreover, VIDA provides a program outline to help the programmer choose breakpoints and visualizes the static dependency relation to help the programmer make estimation at breakpoints.
Dan Hao 0001, Lingming Zhang 0001, Lu Zhang 0023, Jiasu Sun, Hong Mei 0001
ICSE4
2009 Locating need-to-translate constant strings for software internationalization
abstract
Modern software applications require internationalization to be distributed to different regions of the world. In various situations, many software applications are not internationalized at early stages of development. To internationalize such an existing application, developers need to externalize some hard-coded constant strings to resource files, so that translators can easily translate the application into a local language without modifying its source code. Since not all the constant strings require externalization, locating those need-to-translate constant strings is a necessary task that developers must complete for internationalization. In this paper, we present an approach to automatically locating need-to-translate constant strings. Our approach first collects a list of API methods related to the graphical user interface (GUI), and then searches for need-to-translate strings from the invocations of these API methods based on string-taint analysis. We evaluated our approach on four real-world open source applications: RText, Risk, ArtOfIllusion, and Megamek. The results show that our approach effectively locates most of the need-to-translate constant strings in all the four applications.
Xiaoyin Wang, Lu Zhang 0023, Tao Xie 0001, Hong Mei 0001, Jiasu Sun
ICSE5
2009 TranStrL: An automatic need-to-translate string locator for software internationalization
abstract
Software internationalization is often necessary when distributing software applications to different regions around the world. In many cases, developers often do not internationalize a software application at the beginning of the development stage. To internationalize such an existing application, developers need to externalize some hard-coded constant strings to resource files, so that translators can easily translate the application to be in a local language without modifying its source code. Since not all the constant strings require externalization, locating those need-to-translate constant strings is a basic task that the developers must conduct. In this paper, we present TranStrL, an Eclipse plug-in tool that automatically locates need-to-translate constant strings in Java code. Our tool maintains a pre-collected list of API methods related to the Graphical User Interface (GUI), and then searches for need-to-translate strings in the source code starting from the invocations of these API methods using string-taint analysis.
Xiaoyin Wang, Lu Zhang 0023, Tao Xie 0001, Hong Mei 0001, Jiasu Sun
ICSE5
2009 Test-Data Generation Guided by Static Defect Detection
Dan Hao 0001, Lu Zhang 0023, Ming-Hao Liu 0001, Jiasu Sun
J. Comput. Sci. Technol.5
2009 Interactive Fault Localization Using Test Information
Dan Hao 0001, Lu Zhang 0023, Tao Xie 0001, Hong Mei 0001, Jiasu Sun
J. Comput. Sci. Technol.5
2009 Refining component description by leveraging user query logs
Yan Li 0067, Lu Zhang 0023, Jiasu Sun
J. Syst. Softw.4
2008 An approach to detecting duplicate bug reports using natural language and execution information
abstract
An open source project typically maintains an open bug repository so that bug reports from all over the world can be gathered. When a new bug report is submitted to the repository, a person, called a triager, examines whether it is a duplicate of an existing bug report. If it is, the triager marks it as DUPLICATE and the bug report is removed from consideration for further work. In the literature, there are approaches exploiting only natural language information to detect duplicate bug reports. In this paper we present a new approach that further involves execution information. In our approach, when a new bug report arrives, its natural language information and execution information are compared with those of the existing bug reports. Then, a small number of existing bug reports are suggested to the triager as the most similar bug reports to the new bug report. Finally, the triager examines the suggested bug reports to determine whether the new bug report duplicates an existing bug report. We calibrated our approach on a subset of the Eclipse bug repository and evaluated our approach on a subset of the Firefox bug repository. The experimental results show that our approach can detect 67%-93% of duplicate bug reports in the Firefox bug repository, compared to 43%-72% using natural language information alone.
Xiaoyin Wang, Lu Zhang 0023, Tao Xie 0001, John Anvik, Jiasu Sun
ICSE5
2008 Quota-constrained test-case prioritization for regression testing of service-centric systems
abstract
Test-case prioritization is a typical scenario of regression testing, which plays an important role in software maintenance. With the popularity of Web services, integrating Web services to build service-centric systems (SCSs) has attracted attention of many researchers and practitioners. During regression testing, as SCSs may use up constituent Web servicespsila request quotas (e.g., the upper limit of the number of requests that a user can send to a Web service during a certain time range), the quota constraint may delay fault exposure and the subsequent debugging. In this paper, we investigate quota-constrained test-case prioritization for SCSs, and propose quota-constrained strategies to maximize testing requirement coverage. We divide the testing time into time slots, and iteratively select and prioritize test cases for each time slot using integer linear programming (ILP). We performed an experimental study on our strategies together with three other strategies, and the results show that with the constraint of request quotas, our strategies can schedule test cases for execution in an order with higher effectiveness in exposing faults and achieving total and additional branch coverage.
Shan-Shan Hou, Lu Zhang 0023, Tao Xie 0001, Jiasu Sun
ICSM4
2008 Recommending Typical Usage Examples for Component Retrieval in Reuse Repositories
Yan Li 0067, Liangjie Zhang, Ge Li 0001, Jiasu Sun
ICSR5
2008 On similarity-awareness in testing-based fault localization
Dan Hao 0001, Lu Zhang 0023, Hong Mei 0001, Jiasu Sun
Autom. Softw. Eng.5
2008 An objective-oriented approach to program comprehension using multiple information sources
Wei Zhao 0006, Lu Zhang 0023, Jiasu Sun, Hong Mei 0001
Sci. China Ser. F Inf. Sci.3
2007 Mining User Query Logs to Refine Component Description
abstract
How to help reusers retrieve components efficiently and conveniently is critical to the success of the component-based software development (CBSD). In the literature, many research efforts have been put into the improvement of component retrieval mechanisms. Although many different retrieval methods have been proposed, nowadays the component description text is still an important information source for retrieval. Unfortunately, the descriptions of components often contain improper or even noisy information which could deteriorate the effectiveness of the retrieval mechanism. To alleviate the problem, in this paper, we propose an approach which can refine the component description by mining user query logs. The key idea of our approach is to extract proper information from the user query logs to amend the description of the component. Besides, logs from similar components are also taken into consideration. To evaluate the effectiveness of our approach, we performed an experimental study on the JDK. The experimental results demonstrated an improvement in terms of retrieval effectiveness.
Yan Li 0067, Shaobin Cheng, Lu Zhang 0023, Jiasu Sun
COMPSAC (1)5
2007 Applying Interface-Contract Mutation in Regression Testing of Component-Based Software
abstract
Regression testing, which plays an important role in software maintenance, usually relies on test adequacy criteria to select and prioritize test cases. However, with the wide use and reuse of black-box components, such as reusable class libraries and COTS components, it is challenging to establish test adequacy criteria for testing software systems built on components whose source code is not available. Without source code or detailed documents, the misunderstanding between the system integrators and component providers has become a main factor of causing faults in component-based software. In this paper, we apply mutation on interface contracts, which can describe the rights and obligations between component users and providers, to simulate the faults that may occur in this way of software development. The mutation adequacy score for killing the mutants of interface contracts can serve as a test adequacy criterion. We performed an experimental study on three subject systems to evaluate the proposed approach together with four other existing criteria. The experimental results show that our adequacy criterion is helpful for both selecting good-quality test cases and scheduling test cases in an order of exposing faults quickly in regression testing of component-based software.
Shan-Shan Hou, Lu Zhang 0023, Tao Xie 0001, Hong Mei 0001, Jiasu Sun
ICSM5
2007 An Exploratory Study of Web Services on the Internet
abstract
Web services technology has received much attention in the last few years, and a lot of research efforts have been devoted to utilizing services on the Internet to fulfill consumers' requirements. However, little research has been done on the current status of web services on the Internet, which has a great impact on current research. Enlightened by this situation, we made an exploratory study of the current status of web services on the Internet. Our study mainly focused on the investigation of four aspects, including the number, complexity, quality of description and the function diversity of available web services on the Internet. A web services investigation system is built up to harvest web services from the Internet and calculate the statistical results. The investigation results are reported in this paper, and, based on our study, the development trend of web services technology is also discussed in this paper.
Yan Li 0067, Liang-Jie Zhang, Ge Li 0001, Jiasu Sun
ICWS6
2006 Towards Interactive Fault Localization Using Test Information
abstract
Finding the location of a fault is a central task of debugging. Typically, a developer employs an interactive process for fault localization. To accelerate this task, several approaches have been proposed to automate fault localization. In practice, testing-based fault localization (TBFL), which uses test information to locate faults, has become a research focus. However, experimental results reported in the literature showed that current automation of fault localization can only serve as a means to confirming the search space and prioritizing search sequences, not a substitute of the interactive fault localization process. In this paper, we propose an approach based on test information to support the entire interactive fault localization process. During this process, the information gathered from previous interaction steps can be used to provide the ranking of suspicious statements for the current interaction step. As a feasibility study of our approach, we performed an experiment on applying our approach together with some other TBFL approaches on the Siemens programs, which have been used in the literature. Our experimental results show the effectiveness of our approach.
Dan Hao 0001, Lu Zhang 0023, Hong Mei 0001, Jiasu Sun
APSEC4
2006 An Approach to Test Data Generation for Killing Multiple Mutants
abstract
Software testing is an important technique for assurance of software quality. Mutation testing has been identified as a powerful fault-based technique for unit testing, and there has been some research on automatic generation of test data for mutation testing. However, existing approaches to this kind of test data generation usually generate test data according to one mutant at one time. Thus, more test data that are needed for achieving a given mutation score. In this paper, we propose a new approach to generating one test data according to multiple mutants that are mutated at the same location at one time. Thus, our approach can generate smaller test suite that can achieve the same mutation testing score. To evaluate our approach, we implemented a prototype tool based on our approach and carried out some preliminary experiments. The experimental results show that our approach is more cost-effective
Ming-Hao Liu 0001, You-Feng Gao, Jinhui Shan, Jiang-Hong Liu, Lu Zhang 0023, Jiasu Sun
ICSM6
2006 Ranking Component Retrieval Results by Leveraging User History Information
Yan Li 0067, Lu Zhang 0023, Jiasu Sun
SEKE5
2006 Identifying use cases in source code
Lu Zhang 0023, Zhiying Zhou, Dan Hao 0001, Jiasu Sun
J. Syst. Softw.5
2006 SNIAFL: Towards a static noninteractive approach to feature location
abstract
To facilitate software maintenance and evolution, a helpful step is to locate features concerned in a particular maintenance task. In the literature, both dynamic and interactive approaches have been proposed for feature location. In this article, we present a static and noninteractive method for achieving this objective. The main idea of our approach is to use information retrieval (IR) technology to reveal the basic connections between features and computational units in the source code. Due to the imprecision of retrieved connections, we use a static representation of the source code named BRCG (branch-reserving call graph) to further recover both relevant and specific computational units for each feature. A premise of our approach is that programmers should use meaningful names as identifiers. We also performed an experimental study based on two real-world software systems to evaluate our approach. According to experimental results, our approach is quite effective in acquiring the relevant and specific computational units for most features.
Wei Zhao 0006, Lu Zhang 0023, Jiasu Sun, Fuqing Yang
ACM Trans. Softw. Eng. Methodol.4
2005 Eliminating Harmful Redundancy for Testing-Based Fault Localization Using Test Suite Reduction: An Experimental Study
abstract
In the process of software maintenance, it is usually a time-consuming task to track down bugs. To reduce the cost on debugging, several approaches have been proposed to localize the fault(s) to facilitate debugging. Intuitively, testing-based fault localization (TBFL), such as dicing and TRANTULA, is quite promising as it can take the advantage of a large set of execution traces at the same time. However, redundant test cases may bias the distribution of the test suite and harm this kind of approaches. Therefore, we suggest that the test suite, which is the input of TBFL, should be reduced before used in TBFL. To evaluate whether and to what extent TBFL can benefit from test suite reduction, we performed an experimental study on two source programs. The experimental results show that, for test suites containing unevenly distributed redundant test cases, performing test suite reduction before applying TBFL may be more advantageous.
Dan Hao 0001, Lu Zhang 0023, Hao Zhong 0001, Hong Mei 0001, Jiasu Sun
ICSM5
2005 A Hierarchical Decomposition Method for Object-Oriented Systems Based on Identifying Omnipresent Clusters
abstract
System decomposition has become a research focus in software maintenance and evolution for a long time. In this paper, we propose a hierarchical decomposition method for object oriented systems based on identification of omnipresent clusters. The distinctive features of this approach are as follow: firstly, we addressed the problem of omnipresent classes for class diagram. Secondly, we proposed a hierarchical decomposition strategy which can deal with unbalanced hierarchical organization for software system. Thirdly, we presented a revised independency metric that can offset the impact of the measured clusters' size. The experimental results show that this method can be both effective and efficient.
Lu Zhang 0023, Jiasu Sun
ICSM3
2005 Requirements Guided Dynamic Software Clustering
abstract
In this paper, we propose a requirements guided dynamic approach to address software clustering -which aims at providing the logically meaningful and high-level decompositions of large and complex systems. In our approach, the hierarchical structure of functional requirements are constructed by a text document clustering technique named hierarchical agglomerative clustering (HAC) as a high-level skeleton to facilitate the further decomposition of source code through dynamic analysis. We also perform an experimental study based on a GNU system and present the quantitative and qualitative analysis of the experimental results.
Wei Zhao 0006, Lu Zhang 0023, Hong Mei 0001, Jiasu Sun
ICSM4
2005 A similarity-aware approach to testing based fault localization
abstract
Debugging is a time-consuming task in software development and maintenance. To accelerate this task, several approaches have been proposed to automate fault localization. In particular, testing based fault localization (TBFL), which utilizes the testing information to localize the faults, seem to be very promising. However, the similarity between test cases in the test suite has been ignored in the research on TBFL. In this paper, we investigate this similarity issue and propose a novel approach named similarity-aware fault localization (SAFL), which can calculate the suspicion probability of each statement with little impact by the similarity issue. To address and deal with the similarity between test cases, SAFL applies the theory of fuzzy sets to remove the uneven distribution of the test cases. We also performed an experimental study for two real-world programs at different size levels to evaluate SAFL together with another two approaches to TBFL. Experimental results show that SAFL is more effective than the other two approaches when the test suites contain injected redundancy, and SAFL can achieve a competitive result with normal test suites. SAFL can also be more effective than applying test suite reduction to current approaches to TBFL.
Dan Hao 0001, Lu Zhang 0023, Wei Zhao 0006, Hong Mei 0001, Jiasu Sun
ASE6
2004 SNIAFL: Towards a Static Non-Interactive Approach to Feature Location
abstract
To facilitate software maintenance and evolution, a helpful step is to locate features concerned in a particular maintenance task. In the literature, both dynamic and interactive approaches have been proposed for feature location. In this paper, we present a static and non-interactive method for achieving this objective. The main idea of our approach is to use the information retrieval (IR) technology to reveal the basic connections between features and computational units in source code. Due to the characteristics of the retrieved connections, we use a static representation of the source code named BRCG to further recover both the relevant and the specific computational units for each feature. Furthermore, we recover the relationships among the relevant units for each feature. A premise of our approach is that programmers should use meaningful names as identifiers. We perform an experimental study based on a GNU system to evaluate our approach. In the experimental study, we present the detailed quantitative experimental data and give the qualitative analytical results.
Wei Zhao 0006, Lu Zhang 0023, Jiasu Sun, Fuqing Yang
ICSE4
2004 An Experimental Study of Two Graph Analysis Based Component Capture Methods for Object-Oriented Systems
abstract
The problem of how to partition a software system and thus capture its overall architecture and its constituent components has become a research focus in the community of software engineering. In the literature, many methods have been proposed for solving this problem. For example, both top-down and bottom-up methods based on analyzing the graph representation of software systems have been proposed. We report an experimental study of a top-down method and a bottom-up method. In our study, we focus on the capability of component capture, the capability of architecture recovery and the time complexity for the two methods. According to our results on two real world systems, the studied bottom-up method is superior to the studied top-down method in both aspects, although the time complexity of the bottom-up method remains a big concern for large systems.
Renkuan Jiang, Lu Zhang 0023, Hong Mei 0001, Jiasu Sun
ICSM5
2004 Alternative Scalable Algorithms for Lattice-Based Feature Location
abstract
Considering the scalability of using formal concept analysis to locate features in source code, we present a set of alternative straightforward algorithms to achieve the same objectives. A preliminary experiment indicates that the alternative algorithms are more scalable to deal with the large numbers of data to some extent.
Wei Zhao 0006, Lu Zhang 0023, Dan Hao 0001, Hong Mei 0001, Jiasu Sun
ICSM5
2003 Discovering Use Cases from Source Code using the Branch-Reserving Call Graph
abstract
Understanding the behavior of a software system is an important problem in program comprehension. Use cases have been accepted as an effective means for describing behavioral requirements for a software system. We propose a novel approach for obtaining use cases from source code. The central idea of our approach is to use the branch-reserving call graph (BRCG) as the intermediate representation of a software program. We also provide strategies for pruning the BRCG to avoid generating too many fine-grained use cases. Use cases, which may just undergo some minor modifications from human experts, can be generated through traversing the pruned BRCG. The contributions of our approach are three-fold, i) This method represents a compromised approach, which differs from both the static and dynamic approaches for use case discovery, ii) This method takes into consideration the fact that it is the branch statements that separate one use case from another in source code. iii) This method can avoid intensive human involvement in determining the final set of use cases. We have also performed a case study for this method on a GNU system.
Lu Zhang 0023, Zhiying Zhou, Dan Hao 0001, Jiasu Sun
APSEC5
2003 Understanding How the Requirements Are Implemented in Source Code
abstract
For software maintenance and evolution, a common problem is to understand how each requirement is implemented in the source code. The basic solution of this problem is to find the fragment of source code that is corresponding to the implementation of each requirement. This can be viewed as a requirement-slicing problem - slicing the source code according to each individual requirement. We present an approach to find the set of functions that is corresponding to each requirement. The main idea of our method is to combine the information retrieval technology with the static analysis of source code structures. First, we retrieve the initial function sets through some information retrieval model using functional requirements as the queries and identifier information (such as function names, parameter names, variable names etc.) of functions in the source code as target documents. Then we complement each retrieved initial function set by analyzing the call graph extracted from the source code. A premise of our approach is that programmers should use meaningful names as identifiers. Furthermore, we perform an experimental study based on a GNU system. We use two basic metrics: precision and recall (which are the common practice in the information retrieval field), to evaluate our approach. We also compare the results directly acquired from information retrieval with those that are complemented through static source code structure analysis.
Wei Zhao 0006, Lu Zhang 0023, Jiasu Sun
APSEC5