Shay Artzi

dblp:48/3637 · DBLP profile ↗
← Back
18ranked-venue papers
11as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 16 · 10 first-authorDatabases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
14 papers
Software testing · 44% Debugging and program repair · 24% Program analysis · 23%
Network and information security
4 papers
Systems and software security · 64% Web and mobile security · 36%

Topics — the 30 heaviest of 39, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Debugging and program repair
fault localization
0.542012
Fault Localization for Dynamic Web Applications · IEEE Trans. Software Eng. 2012
Directed test generation for effective fault localization · ISSTA 2010
Practical fault localization for dynamic web applications · ICSE (1) 2010
Software testing
test generation
0.442012
Fault Localization for Dynamic Web Applications · IEEE Trans. Software Eng. 2012
Finding Bugs in Web Applications Using Dynamic Test Generation and Explicit-State Model Checking · IEEE Trans. Software Eng. 2010
Directed test generation for effective fault localization · ISSTA 2010
Program analysis › constraint solving
string constraint solving
0.322012
HAMPI: A solver for word equations over strings, regular expressions, and context-free grammars · ACM Trans. Softw. Eng. Methodol. 2012
HAMPI: A String Solver for Testing, Analysis and Vulnerability Detection · CAV 2011
Software testing › test generation
dynamic test generation
0.222010
Finding Bugs in Web Applications Using Dynamic Test Generation and Explicit-State Model Checking · IEEE Trans. Software Eng. 2010
Finding bugs in dynamic web applications · ISSTA 2008
Program analysis
dynamic analysis
0.222010
Finding Bugs in Web Applications Using Dynamic Test Generation and Explicit-State Model Checking · IEEE Trans. Software Eng. 2010
Combined static and dynamic mutability analysis · ASE 2007
Debugging and program repair
automated program repair
0.112012
Automated repair of HTML generation errors in PHP applications using string constraint solving · ICSE 2012
Software testing
automated testing
0.112012
HAMPI: A solver for word equations over strings, regular expressions, and context-free grammars · ACM Trans. Softw. Eng. Methodol. 2012
Software testing › test input generation
concolic testing
0.112012
HAMPI: A solver for word equations over strings, regular expressions, and context-free grammars · ACM Trans. Softw. Eng. Methodol. 2012
Program analysis
constraint solving
0.112012
HAMPI: A solver for word equations over strings, regular expressions, and context-free grammars · ACM Trans. Softw. Eng. Methodol. 2012
Debugging and program repair › fault localization
statistical debugging
0.112012
Fault Localization for Dynamic Web Applications · IEEE Trans. Software Eng. 2012
Systems and software security › information flow tracking
taint analysis
0.112011
F4F: taint analysis of framework-based web applications · OOPSLA 2011
Systems and software security
vulnerability discovery
0.112011
F4F: taint analysis of framework-based web applications · OOPSLA 2011
Software testing › test generation
automated test generation
0.112011
A framework for automated testing of javascript web applications · ICSE 2011
Software testing › test generation › dynamic test generation
feedback-directed test generation
0.112011
A framework for automated testing of javascript web applications · ICSE 2011
Software testing › web application testing
javascript web application testing
0.112011
A framework for automated testing of javascript web applications · ICSE 2011
Software testing
test input generation
0.112011
HAMPI: A String Solver for Testing, Analysis and Vulnerability Detection · CAV 2011
Program analysis
static analysis
0.122011
Combined static and dynamic mutability analysis · ASE 2007
F4F: taint analysis of framework-based web applications · OOPSLA 2011
Software testing › test generation
directed test generation
0.112010
Directed test generation for effective fault localization · ISSTA 2010
Program analysis › symbolic execution
dynamic symbolic execution
0.112010
Directed test generation for effective fault localization · ISSTA 2010
Program verification › model checking
explicit-state model checking
0.112010
Finding Bugs in Web Applications Using Dynamic Test Generation and Explicit-State Model Checking · IEEE Trans. Software Eng. 2010
Program verification
model checking
0.112010
Finding Bugs in Web Applications Using Dynamic Test Generation and Explicit-State Model Checking · IEEE Trans. Software Eng. 2010
Debugging and program repair
bug reproduction
0.112009
ReCrashJ: a tool for capturing and reproducing program crashes in deployed applications · ESEC/SIGSOFT FSE 2009
Debugging and program repair › bug reproduction
crash reproduction
0.112009
ReCrashJ: a tool for capturing and reproducing program crashes in deployed applications · ESEC/SIGSOFT FSE 2009
Program analysis › static analysis
bug detection
0.112008
Finding bugs in dynamic web applications · ISSTA 2008
Software testing › fault detection
crash detection
0.112008
Finding bugs in dynamic web applications · ISSTA 2008
Software testing
web application testing
0.112008
Finding bugs in dynamic web applications · ISSTA 2008
Programming languages and type systems
type systems
0.112007
Object and reference immutability using java generics · ESEC/SIGSOFT FSE 2007
Software testing › test infrastructure
mock generation
0.112005
Automatic test factoring for java · ASE 2005
Web and mobile security › web security › web vulnerability detection
SQL injection detection
0.012012
HAMPI: A solver for word equations over strings, regular expressions, and context-free grammars · ACM Trans. Softw. Eng. Methodol. 2012
Debugging and program repair › automated program repair
test-based program repair
0.012012
Automated repair of HTML generation errors in PHP applications using string constraint solving · ICSE 2012

Methods — techniques the papers use, named apart from their topics

concolic execution · 0.4word equations · 0.3context-free grammar · 0.3string constraint solving · 0.3tarantula · 0.3dynamic analysis · 0.2static analysis · 0.1regular expressions · 0.1regular expression · 0.1ochiai · 0.1jaccard · 0.1taint analysis · 0.1reflective construct handling · 0.1HTML validation · 0.1
YearPublicationVenuePosition
2020 Subjective Search Intent Predictions using Customer Reviews
abstract
Query intent prediction is a component of information retrieval which improves result relevance through an understanding of latent user intents in addition to explicit query keywords. We target context-of-use intents, such as the activity for which a product is used and the target audience for a product, which are subjective and not usually indexed as product attributes in the catalog. We describe a method to predict latent query intents: we extract intents from product reviews on amazon.com and, using behavioral purchase signals that associate queries with the reviewed products, train query classifiers that label queries with the intents extracted from reviews. For example, we predict the activity "running" for the query "adidas mens pants." We show that our method can predict latent intents not indexed directly in the product catalog.
Adrian Boteanu, Emily Dutile, Adam Kiezun, Shay Artzi
CHIIR4
2012 Automated repair of HTML generation errors in PHP applications using string constraint solving
abstract
PHP web applications routinely generate invalid HTML. Modern browsers silently correct HTML errors, but sometimes malformed pages render inconsistently, cause browser crashes, or expose security vulnerabilities. Fixing errors in generated pages is usually straightforward, but repairing the generating PHP program can be much harder. We observe that malformed HTML is often produced by incorrect constant prints, i.e., statements that print string literals, and present two tools for automatically repairing such HTML generation errors. PHPQuickFix repairs simple bugs by statically analyzing individual prints. PHPRepair handles more general repairs using a dynamic approach. Based on a test suite, the property that all tests should produce their expected output is encoded as a string constraint over variables representing constant prints. Solving this constraint describes how constant prints must be modified to make all tests pass. Both tools were implemented as an Eclipse plugin and evaluated on PHP programs containing hundreds of HTML generation errors, most of which our tools were able to repair automatically.
Hesam Samimi, Max Schäfer, Shay Artzi, Todd D. Millstein, Frank Tip, Laurie J. Hendren
ICSE3
2012 HAMPI: A solver for word equations over strings, regular expressions, and context-free grammars
abstract
Many automatic testing, analysis, and verification techniques for programs can be effectively reduced to a constraint-generation phase followed by a constraint-solving phase. This separation of concerns often leads to more effective and maintainable software reliability tools. The increasing efficiency of off-the-shelf constraint solvers makes this approach even more compelling. However, there are few effective and sufficiently expressive off-the-shelf solvers for string constraints generated by analysis of string-manipulating programs, so researchers end up implementing their own ad-hoc solvers. To fulfill this need, we designed and implemented Hampi, a solver for string constraints over bounded string variables. Users of Hampi specify constraints using regular expressions, context-free grammars, equality between string terms, and typical string operations such as concatenation and substring extraction. Hampi then finds a string that satisfies all the constraints or reports that the constraints are unsatisfiable. We demonstrate Hampi's expressiveness and efficiency by applying it to program analysis and automated testing. We used Hampi in static and dynamic analyses for finding SQL injection vulnerabilities in Web applications with hundreds of thousands of lines of code. We also used Hampi in the context of automated bug finding in C programs using dynamic systematic testing (also known as concolic testing). We then compared Hampi with another string solver, CFGAnalyzer, and show that Hampi is several times faster. Hampi's source code, documentation, and experimental data are available at http://people.csail.mit.edu/akiezun/hampi 1
Adam Kiezun, Vijay Ganesh 0001, Shay Artzi, Philip J. Guo, Pieter Hooimeijer, Michael D. Ernst
ACM Trans. Softw. Eng. Methodol.3
2012 Fault Localization for Dynamic Web Applications
abstract
In recent years, there has been significant interest in fault-localization techniques that are based on statistical analysis of program constructs executed by passing and failing executions. This paper shows how the Tarantula, Ochiai, and Jaccard fault-localization algorithms can be enhanced to localize faults effectively in web applications written in PHP by using an extended domain for conditional and function-call statements and by using a source mapping. We also propose several novel test-generation strategies that are geared toward producing test suites that have maximal fault-localization effectiveness. We implemented various fault-localization techniques and test-generation strategies in Apollo, and evaluated them on several open-source PHP applications. Our results indicate that a variant of the Ochiai algorithm that includes all our enhancements localizes 87.8 percent of all faults to within 1 percent of all executed statements, compared to only 37.4 percent for the unenhanced Ochiai algorithm. We also found that all the test-generation strategies that we considered are capable of generating test suites with maximal fault-localization effectiveness when given an infinite time budget for test generation. However, on average, a directed strategy based on path-constraint similarity achieves this maximal effectiveness after generating only 6.5 tests, compared to 46.8 tests for an undirected test-generation strategy.
Shay Artzi, Julian Dolby, Frank Tip, Marco Pistoia
IEEE Trans. Software Eng.1
2011 HAMPI: A String Solver for Testing, Analysis and Vulnerability Detection
Vijay Ganesh 0001, Adam Kiezun, Shay Artzi, Philip J. Guo, Pieter Hooimeijer, Michael D. Ernst
CAV3
2011 A framework for automated testing of javascript web applications
abstract
Current practice in testing JavaScript web applications requires manual construction of test cases, which is difficult and tedious. We present a framework for feedback-directed automated test generation for JavaScript in which execution is monitored to collect information that directs the test generator towards inputs that yield increased coverage. We implemented several instantiations of the framework, corresponding to variations on feedback-directed random testing, in a tool called Artemis. Experiments on a suite of JavaScript applications demonstrate that a simple instantiation of the framework that uses event handler registrations as feedback information produces surprisingly good coverage if enough tests are generated. By also using coverage information and read-write sets as feedback information, a slightly better level of coverage can be achieved, and sometimes with many fewer tests. The generated tests can be used for detecting HTML validity problems and other programming errors.
Shay Artzi, Julian Dolby, Simon Holm Jensen, Anders Møller, Frank Tip
ICSE1
2011 F4F: taint analysis of framework-based web applications
abstract
This paper presents F4F (Framework For Frameworks), a system for effective taint analysis of framework-based web applications. Most modern web applications utilize one or more web frameworks, which provide useful abstractions for common functionality. Due to extensive use of reflective language constructs in framework implementations, existing static taint analyses are often ineffective when applied to framework-based applications. While previous work has included ad hoc support for certain framework constructs, adding support for a large number of frameworks in this manner does not scale from an engineering standpoint.
Manu Sridharan, Shay Artzi, Marco Pistoia, Salvatore Guarnieri, Omer Tripp, Ryan Berg
OOPSLA2
2010 Practical fault localization for dynamic web applications
abstract
We leverage combined concrete and symbolic execution and several fault-localization techniques to create a uniquely powerful tool for localizing faults in PHP applications. The tool automatically generates tests that expose failures, and then automatically localizes the faults responsible for those failures, thus overcoming the limitation of previous fault-localization techniques that a test suite be available upfront. The fault-localization techniques we employ combine variations on the Tarantula algorithm with a technique based on maintaining a mapping between statements and the fragments of output they produce. We implemented these techniques in a tool called Apollo, and evaluated them by localizing 75 randomly selected faults that were exposed by automatically generated tests in four PHP applications. Our findings indicate that, using our best technique, 87.7% of the faults under consideration are localized to within 1% of all executed statements, which constitutes an almost five-fold improvement over the Tarantula algorithm.
Shay Artzi, Julian Dolby, Frank Tip, Marco Pistoia
ICSE (1)1
2010 Directed test generation for effective fault localization
abstract
Fault-localization techniques that apply statistical analyses to execution data gathered from multiple tests are quite effective when a large test suite is available. However, if no test suite is available, what is the best approach to generate one? This paper investigates the fault-localization effectiveness of test suites generated according to several test-generation techniques based on combined concrete and symbolic (concolic) execution. We evaluate these techniques by applying the Ochiai fault-localization technique to generated test suites in order to localize 35 faults in four PHP Web applications. Our results show that the test-generation techniques under consideration produce test suites with similar high fault-localization effectiveness, when given a large time budget. However, a new, "directed" test-generation technique, which aims to maximize the similarity between the path constraints of the generated tests and those of faulty executions, reaches this level of effectiveness with much smaller test suites. On average, when compared to test generation based on standard concolic execution techniques that aims to maximize code coverage, the new directed technique preserves fault-localization effectiveness while reducing test-suite size by 86.1% and test-suite generation time by 88.6%.
Shay Artzi, Julian Dolby, Frank Tip, Marco Pistoia
ISSTA1
2010 Finding Bugs in Web Applications Using Dynamic Test Generation and Explicit-State Model Checking
abstract
Web script crashes and malformed dynamically generated webpages are common errors, and they seriously impact the usability of Web applications. Current tools for webpage validation cannot handle the dynamically generated pages that are ubiquitous on today's Internet. We present a dynamic test generation technique for the domain of dynamic Web applications. The technique utilizes both combined concrete and symbolic execution and explicit-state model checking. The technique generates tests automatically, runs the tests capturing logical constraints on inputs, and minimizes the conditions on the inputs to failing tests so that the resulting bug reports are small and useful in finding and fixing the underlying faults. Our tool Apollo implements the technique for the PHP programming language. Apollo generates test inputs for a Web application, monitors the application for crashes, and validates that the output conforms to the HTML specification. This paper presents Apollo's algorithms and implementation, and an experimental evaluation that revealed 673 faults in six PHP Web applications.
Shay Artzi, Adam Kiezun, Julian Dolby, Frank Tip, Danny Dig, Amit M. Paradkar, Michael D. Ernst
IEEE Trans. Software Eng.1
2009 ReCrashJ: a tool for capturing and reproducing program crashes in deployed applications
abstract
Many programs have latent bugs that cause the program to fail. In order to fix a failing program, is it crucial to be able to reproduce the failure consistently. However, reproducing a failure can be difficult and time-consuming, especially when the failure is discovered by a user in a deployed application.
Shay Artzi, Sunghun Kim 0001, Michael D. Ernst
ESEC/SIGSOFT FSE1
2009 Parameter reference immutability: formal definition, inference tool, and comparison
Shay Artzi, Adam Kiezun, Jaime Quinonez, Michael D. Ernst
Autom. Softw. Eng.1
2008 ReCrash: Making Software Failures Reproducible by Preserving Object States
Shay Artzi, Sunghun Kim 0001, Michael D. Ernst
ECOOP1
2008 Finding bugs in dynamic web applications
abstract
Web script crashes and malformed dynamically-generated Web pages are common errors, and they seriously impact usability of Web applications. Current tools for Web-page validation cannot handle the dynamically-generated pages that are ubiquitous on today's Internet. In this work, we apply a dynamic test generation technique, based on combined concrete and symbolic execution, to the domain of dynamic Web applications. The technique generates tests automatically, uses the tests to detect failures, and minimizes the conditions on the inputs exposing each failure, so that the resulting bug reports are small and useful in finding and fixing the underlying faults. Our tool Apollo implements the technique for PHP. Apollo generates test inputs for the Web application, monitors the application for crashes, and validates that the output conforms to the HTML specification. This paper presents Apollo's algorithms and implementation, and an experimental evaluation that revealed 214 faults in 4 PHP Web applications.
Shay Artzi, Adam Kiezun, Julian Dolby, Frank Tip, Danny Dig, Amit M. Paradkar, Michael D. Ernst
ISSTA1
2008 miRNAminer: A tool for homologous microRNA gene search
abstract
BACKGROUND: MicroRNAs (miRNAs), present in most metazoans, are small non-coding RNAs that control gene expression by negatively regulating translation through binding to the 3'UTR of mRNA transcripts. Previously, experimental and computational methods were used to construct miRNA gene repositories agreeing with careful submission guidelines. RESULTS: An algorithm we developed - miRNAminer - is used for homologous conserved miRNA gene search in several animal species. Given a search query, candidate homologs from different species are tested for their known miRNA properties, such as secondary structure, energy and alignment and conservation, in order to asses their fidelity. When applying miRNAminer on seven mammalian species we identified several hundreds of high-confidence homologous miRNAs increasing the total collection of (miRbase) miRNAs, in these species, by more than 50%. miRNAminer uses stringent criteria and exhibits high sensitivity and specificity. CONCLUSION: We present - miRNAminer - the first web-server for homologous miRNA gene search in animals. miRNAminer can be used to identify conserved homolog miRNA genes and can also be used prior to depositing miRNAs in public databases. miRNAminer is available at http://pag.csail.mit.edu/mirnaminer.
Shay Artzi, Adam Kiezun, Noam Shomron
BMC Bioinform.1
2007 Combined static and dynamic mutability analysis
abstract
Knowing which method parameters may be mutated during a method's executionis useful for many software engineering tasks. We present an approach todiscovering parameter reference immutability, in which several lightweight, scalable analyses are combined in stages, with each stage refining the overall result. The resulting analysis is scalable and combines the strengths of its component analyses. As one of the component analyses, we present a novel, dynamic mutability analysis and show how its results can be improved by random input generation. Experimental results on programs of up to 185 kLOC show that, compared to previous approaches, our approach increases both scalability and overall accuracy
Shay Artzi, Adam Kiezun, David Glasser, Michael D. Ernst
ASE1
2007 Object and reference immutability using java generics
abstract
A compiler-checked immutability guarantee provides useful documentation, facilitates reasoning, and enables optimizations. This paper presents Immutability Generic Java (IGJ), a novel language extension that expresses immutability without changing Java's syntax by building upon Java's generics and annotation mechanisms. In IGJ, each class has one additional type parameter that is Immutable, Mutable, or ReadOnly. IGJ guarantees both reference immutability (only mutable references can mutate an object) and object immutability (an immutable reference points to an immutable object). IGJ is the first proposal for enforcing object immutability within Java's syntax and type system, and its reference immutability is more expressive than previous work. IGJ also permits covariant changes of type parameters in a type-safe manner, e.g., a readonly list of integers is a subtype of a readonly list of numbers. IGJ extends Java's type system with a few simple rules. We formalize this type system and prove it sound. Our IGJ compiler works by type-erasure and generates byte-code that can be executed on any JVM without runtime penalty.
Yoav Zibin, Alex Potanin, Mahmood Ali, Shay Artzi, Adam Kiezun, Michael D. Ernst
ESEC/SIGSOFT FSE4
2005 Automatic test factoring for java
abstract
Test factoring creates fast, focused unit tests from slow system-widetests; each new unit test exercises only a subset of the functionalityexercised by the system test. Augmenting a test suite with factoredunit tests should catch errors earlier in a test run.One way to factor a test is to introduce 'mock' objects. If a testexercises a component T, which interacts with another component E (the'environment'), the implementation of E can be replaced by a mock.The mock checks that T's calls to E are as expected, and it simulatesE's behavior in response. We introduce an automatic technique fortest factoring. Given a system test for T and E, and a record of T'sand E's behavior when the system test is run, test factoring generatesunit tests for T in which E is mocked. The factored tests can isolatebugs in T from bugs in E and, if E is slow or expensive, improve testperformance or cost.We have built an implementation of automatic dynamic test factoring for theJava language. Our experimental data indicates that it can reduce therunning time of a system test suite by up to an order of magnitude.
David Saff, Shay Artzi, Jeff H. Perkins, Michael D. Ernst
ASE2