Kunal Taneja

dblp:88/6573 · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
0since 2021 · last 2016
0000-0003-2099-9803ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 10 · 6 first-authorArtificial intelligence and machine learning · 1Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
9 papers
Software testing · 66% Software maintenance and evolution · 19% Program analysis · 10%
Network and information security
2 papers
Privacy and data protection · 54% Web and mobile security · 46%

Topics — the 16 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Software testing
test generation
0.442012
eXpress: guided path exploration for efficient regression test generation · ISSTA 2011
MODA: automated test generation for database applications via mock objects · ASE 2010
DiffGen: Automated Regression Unit-Test Generation · ASE 2008
Software maintenance and evolution
software evolution
0.322016
High-confidence software evolution · Sci. China Inf. Sci. 2016
Automated detection of api refactorings in libraries · ASE 2007
Program analysis › symbolic execution
dynamic symbolic execution
0.222011
DyTa: dynamic symbolic execution guided with static verification results · ICSE 2011
MODA: automated test generation for database applications via mock objects · ASE 2010
Software testing
regression testing
0.222011
eXpress: guided path exploration for efficient regression test generation · ISSTA 2011
DiffGen: Automated Regression Unit-Test Generation · ASE 2008
Software testing › test coverage › code coverage
statement coverage
0.112012
CarFast: achieving higher statement coverage faster · SIGSOFT FSE 2012
Software testing
test coverage
0.112012
CarFast: achieving higher statement coverage faster · SIGSOFT FSE 2012
Privacy and data protection
anonymization
0.112011
Testing software in age of data privacy: a balancing act · SIGSOFT FSE 2011
Software testing › automated testing
continuous testing
0.112011
eXpress: guided path exploration for efficient regression test generation · ISSTA 2011
Software testing › test generation
dynamic test generation
0.112011
DyTa: dynamic symbolic execution guided with static verification results · ICSE 2011
Program verification
static verification
0.112011
DyTa: dynamic symbolic execution guided with static verification results · ICSE 2011
Software testing › database testing
database application testing
0.112010
MODA: automated test generation for database applications via mock objects · ASE 2010
Software testing
test oracle
0.112010
MiTV: multiple-implementation testing of user-input validators for web applications · ASE 2010
Software testing › test generation
coverage-based test generation
0.112008
DiffGen: Automated Regression Unit-Test Generation · ASE 2008
Software testing
software reliability
0.112016
High-confidence software evolution · Sci. China Inf. Sci. 2016
Software maintenance and evolution › software evolution
library API evolution
0.112007
Automated detection of api refactorings in libraries · ASE 2007
Software maintenance and evolution
refactoring
0.112007
Automated detection of api refactorings in libraries · ASE 2007

Methods — techniques the papers use, named apart from their topics

dynamic symbolic execution · 0.4software evolution · 0.2program analysis · 0.2data privacy framework · 0.2mock objects · 0.2differential testing · 0.2search-based testing · 0.1static verification · 0.1guided path exploration · 0.1program instrumentation · 0.1
YearPublicationVenuePosition
2016 ICON: Inferring Temporal Constraints from Natural Language API Descriptions
abstract
Temporal constraints of an Application Programming Interface (API) are the allowed sequences of method invocations in the API governing the secure and robust operation of client software using the API. These constraints are typically described informally in natural language API documents, and therefore are not amenable to existing constraint-checking tools. Manually identifying and writing formal temporal constraints from API documents can be prohibitively time-consuming and error-prone. To address this issue, we propose ICON: an approach based on Machine Learning (ML) and Natural Language Processing (NLP) for identifying and inferring formal temporal constraints. To evaluate our approach, we use ICON to infer and formalize temporal constraints from the Amazon S3 REST API, the PayPal Payment REST API, and the java.io package in the JDK API. Our results indicate that ICON can effectively identify temporal constraint sentences (from over 4000 human annotated API sentences) with the average 79.0% precision and 60.0% recall. Furthermore, our evaluation demonstrates that ICON achieves an accuracy of 70% in inferring 77 formal temporal constraints from these APIs.
Rahul Pandita, Kunal Taneja, Laurie A. Williams, Teresa Tung
ICSME2
2016 High-confidence software evolution
Yingfei Xiong 0001, Dan Hao 0001, Xusheng Xiao, Kunal Taneja, Lu Zhang 0023, Tao Xie 0001
Sci. China Inf. Sci.6
2016 RUGRAT: Evaluating program analysis and testing tools and compilers with large generated random benchmark applications
abstract
Benchmarks are heavily used in different areas of computer science to evaluate algorithms and tools. In program analysis and testing, open-source and commercial programs are routinely used as benchmarks to evaluate different aspects of algorithms and tools. Unfortunately, many of these programs are written by programmers who introduce different biases, not to mention that it is very difficult to find programs that can serve as benchmarks with high reproducibility of results. We propose a novel approach for generating random benchmarks for evaluating program analysis and testing tools and compilers. Our approach uses stochastic parse trees, where language grammar production rules are assigned probabilities that specify the frequencies with which instantiations of these rules will appear in the generated programs. We implemented our tool for Java and applied it to generate a set of large benchmark programs of up to 5M lines of code each with which we evaluated different program analysis and testing tools and compilers. The generated benchmarks let us independently rediscover several issues in the evaluated tools. Copyright © 2014 John Wiley & Sons, Ltd.
Ishtiaque Hussain, Christoph Csallner, Mark Grechanik, Qing Xie 0003, Kunal Taneja, B. M. Mainul Hossain
Softw. Pract. Exp.6
2012 CarFast: achieving higher statement coverage faster
abstract
Test coverage is an important metric of software quality, since it indicates thoroughness of testing. In industry, test coverage is often measured as statement coverage. A fundamental problem of software testing is how to achieve higher statement coverage faster, and it is a difficult problem since it requires testers to cleverly find input data that can steer execution sooner toward sections of application code that contain more statements.
B. M. Mainul Hossain, Ishtiaque Hussain, Christoph Csallner, Mark Grechanik, Kunal Taneja, Qing Xie 0003
SIGSOFT FSE6
2011 DyTa: dynamic symbolic execution guided with static verification results
abstract
Software-defect detection is an increasingly important research topic in software engineering. To detect defects in a program, static verification and dynamic test generation are two important proposed techniques. However, both of these techniques face their respective issues. Static verification produces false positives, and on the other hand, dynamic test generation is often time consuming. To address the limitations of static verification and dynamic test generation, we present an automated defect-detection tool, called DyTa, that combines both static verification and dynamic test generation. DyTa consists of a static phase and a dynamic phase. The static phase detects potential defects with a static checker; the dynamic phase generates test inputs through dynamic symbolic execution to confirm these potential defects. DyTa reduces the number of false positives compared to static verification and performs more efficiently compared to dynamic test generation.
Xi Ge, Kunal Taneja, Tao Xie 0001, Nikolai Tillmann
ICSE2
2011 eXpress: guided path exploration for efficient regression test generation
abstract
Software programs evolve throughout their lifetime undergoing various changes. While making these changes, software developers may introduce regression faults. It is desirable to detect these faults as quickly as possible to reduce the cost involved in fixing them. One existing solution is continuous testing, which runs an existing test suite to quickly find regression faults as soon as code changes are saved. However, the effectiveness of continuous testing depends on the capability of the existing test suite for finding behavioral differences across versions.
Kunal Taneja, Tao Xie 0001, Nikolai Tillmann, Jonathan de Halleux
ISSTA1
2011 Testing software in age of data privacy: a balancing act
abstract
Database-centric applications (DCAs) are common in enterprise computing, and they use nontrivial databases. Testing of DCAs is increasingly outsourced to test centers in order to achieve lower cost and higher quality. When proprietary DCAs are released, their databases should also be made available to test engineers. However, different data privacy laws prevent organizations from sharing this data with test centers because databases contain sensitive information. Currently, testing is performed with anonymized data, which often leads to worse test coverage (such as code coverage) and fewer uncovered faults, thereby reducing the quality of DCAs and obliterating benefits of test outsourcing. To address this issue, we offer a novel approach that combines program analysis with a new data privacy framework that we design to address constraints of software testing. With our approach, organizations can balance the level of privacy with needs of testing. We have built a tool for our approach and applied it to nontrivial Java DCAs. Our results show that test coverage can be preserved at a higher level by anonymizing data based on their effect on corresponding DCAs.
Kunal Taneja, Mark Grechanik, Rayid Ghani, Tao Xie 0001
SIGSOFT FSE1
2010 MiTV: multiple-implementation testing of user-input validators for web applications
abstract
User-input validators play an essential role in guarding a web application against application-level attacks. Hence, the security of the web application can be compromised by defective validators. To detect defects in validators, testing is one of the most commonly used methodologies. Testing can be performed by manually writing test inputs and oracles, but this manual process is often labor-intensive and ineffective. On the other hand, automated test generators cannot generate test oracles in the absence of specifications, which are often not available in practice. To address this issue in testing validators, we propose a novel approach, called MiTV, that applies Multiple-implementation Testing for Validators, i.e., comparin gthe behavior of a validator under test with other validators of the same type. These other validators of the same type can be collected from either open or proprietary source code repositories. To show the effectiveness of MiTV, we applied MiTV on 53 different validators (of 6 common types) for web applications. Our results show that MiTV detected real defects in 70% of the validators.
Kunal Taneja, Madhuri R. Marri, Tao Xie 0001, Nikolai Tillmann
ASE1
2010 MODA: automated test generation for database applications via mock objects
abstract
Software testing has been commonly used in assuring the quality of database applications. It is often prohibitively expensive to manually write quality tests for complex database applications. Automated test generation techniques, such as Dynamic Symbolic Execution (DSE), have been proposed to reduce human efforts in testing database applications. However, such techniques have two major limitations: (1) they assume that the database that the application under test interacts with is accessible, which may not always be true; and (2) they usually cannot create necessary database states as a part of the generated tests.
Kunal Taneja, Yi Zhang 0051, Tao Xie 0001
ASE1
2008 Improving software reliability and productivity via mining program source code
abstract
A software system interacts with third-party libraries through various APIs. Insufficient documentation and constant refactorings of third-party libraries make API library reuse difficult and error prone. Using these library APIs often needs to follow certain usage patterns. These patterns aid developers in addressing commonly faced programming problems such as what checks should precede or follow API calls, how to use a given set of APIs for a given task, or what API method sequence should be used to obtain one object from another. Ordering rules (specifications) also exist between APIs, and these rules govern the secure and robust operation of the system using these APIs. These patterns and rules may not be well documented by the API developers. Furthermore, usage patterns and specifications might change with library refactorings, requiring changes in the software that reuse the library. To address these issues, we develop novel techniques (and their supporting tools) based on mining source code, assisting developers in productively reusing third party libraries to build reliable and secure software.
Tao Xie 0001, Mithun Acharya, Suresh Thummalapenta, Kunal Taneja
IPDPS4
2008 DiffGen: Automated Regression Unit-Test Generation
abstract
Software programs continue to evolve throughout their lifetime. Maintenance of such evolving programs, including regression testing, is one of the most expensive activities in software development. We present an approach and its implementation called DiffGen for automated regression unit-test generation and checking for Java programs. Given two versions of a Java class, our approach instruments the code by adding new branches such that if these branches can be covered by a test generation tool, behavioral differences between the two class versions are exposed. DiffGen then uses a coverage-based test generation tool to generate test inputs for covering the added branches to expose behavioral differences. We have evaluated DiffGen on finding behavioral differences between 21 classes and their versions. Experimental results show that our approach can effectively expose many behavioral differences that cannot be exposed by state-of-the-art techniques.
Kunal Taneja, Tao Xie 0001
ASE1
2008 Search-based inference of dialect grammars
Massimiliano Di Penta, Pierpaolo Lombardi, Kunal Taneja, Luigi Troiano
Soft Comput.3
2007 Automated detection of api refactorings in libraries
abstract
Software developers often do not build software from scratch but reuse software libraries. In theory, the APIs of a library should be stable, but in practice they do change and thus require changes in software that reuses the library. Our previous study of five reusable components shows that more than 80% of these API changes are caused by refactorings. If these refactorings could be automatically detected, they could be used to automatically upgrade applications.
Kunal Taneja, Danny Dig, Tao Xie 0001
ASE1