Erik Wittern

dblp:21/10311 · also John Erik Wittern · DBLP profile ↗
← Back
17ranked-venue papers
9as first author
1since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 16 · 8 first-author · 1 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
2 papers
Program analysis · 54% Services computing and microservices · 46%
Databases, data mining, and information retrieval
1 paper
Machine learning and data management · 100%

Topics — the 3 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Program analysis
static analysis
0.312017
Statically checking web API requests in JavaScript · ICSE 2017
Program analysis › data flow analysis › value analysis
string analysis
0.312017
Statically checking web API requests in JavaScript · ICSE 2017
Machine learning and data management
data management for machine learning
0.112021
Learning GraphQL Query Cost · ASE 2021

Methods — techniques the papers use, named apart from their topics

static query analysis · 1.0machine learning · 1.0inter-procedural string analysis · 0.3
YearPublicationVenuePosition
2021 Learning GraphQL Query Cost
abstract
GraphQL is a query language for APIs and a runtime for executing those queries, fetching the requested data from existing microservices, REST APIs, databases, or other sources. Its expressiveness and its flexibility have made it an attractive candidate for API providers in many industries, especially through the web. A major drawback to blindly servicing a client’s query in GraphQL is that the cost of a query can be unexpectedly large, creating computation and resource overload for the provider, and API rate-limit overages and infrastructure overload for the client. To mitigate these drawbacks, it is necessary to efficiently estimate the cost of a query before executing it. Estimating query cost is challenging, because GraphQL queries have a nested structure, GraphQL APIs follow different design conventions, and the underlying data sources are hidden. Estimates based on worst-case static query analysis have had limited success because they tend to grossly overestimate cost. We propose a machine-learning approach to efficiently and accurately estimate the query cost. We also demonstrate the power of this approach by testing it on query-response data from publicly available commercial APIs. Our framework is efficient and predicts query costs with high accuracy, consistently outperforming the static analysis by a large margin.
Georgios Mavroudeas, Guillaume Baudart, Alan Cha, Martin Hirzel, Jim Laredo, Malik Magdon-Ismail, Louis Mandel, Erik Wittern
ASE8
2020 Topology-Aware Continuous Experimentation in Microservice-Based Applications
Gerald Schermann, Fábio Oliveira, Erik Wittern, Philipp Leitner 0001
ICSOC3
2020 A principled approach to GraphQL query cost analysis
abstract
The landscape of web APIs is evolving to meet new client requirements and to facilitate how providers fulfill them. A recent web API model is GraphQL, which is both a query language and a runtime. Using GraphQL, client queries express the data they want to retrieve or mutate, and servers respond with exactly those data or changes. GraphQL’s expressiveness is risky for service providers because clients can succinctly request stupendous amounts of data, and responding to overly complex queries can be costly or disrupt service availability. Recent empirical work has shown that many service providers are at risk. Using traditional API management methods is not sufficient, and practitioners lack principled means of estimating and measuring the cost of the GraphQL queries they receive. In this work, we present a linear-time GraphQL query analysis that can measure the cost of a query without executing it. Our approach can be applied in a separate API management layer and used with arbitrary GraphQL backends. In contrast to existing static approaches, our analysis supports common GraphQL conventions that affect query cost, and our analysis is provably correct based on our formal specification of GraphQL semantics. We demonstrate the potential of our approach using a novel GraphQL query-response corpus for two commercial GraphQL APIs. Our query analysis consistently obtains upper cost bounds, tight enough relative to the true response sizes to be actionable for service providers. In contrast, existing static GraphQL query analyses exhibit over-estimates and under-estimates because they fail to support GraphQL conventions.
Alan Cha, Erik Wittern, Guillaume Baudart, James C. Davis 0001, Louis Mandel, Jim Laredo
ESEC/SIGSOFT FSE2
2019 An Empirical Study of GraphQL Schemas
Erik Wittern, Alan Cha, James C. Davis 0001, Guillaume Baudart, Louis Mandel
ICSOC1
2019 A mixed-method empirical study of Function-as-a-Service software development in industrial practice
Philipp Leitner 0001, Erik Wittern, Josef Spillner, Waldemar Hummer
J. Syst. Softw.2
2018 Generating GraphQL-Wrappers for REST(-like) APIs
Erik Wittern, Alan Cha, Jim Laredo
ICWE1
2018 Towards extracting web API specifications from documentation
abstract
Web API specifications are machine-readable descriptions of APIs. These specifications, in combination with related tooling, simplify and support the consumption of APIs. However, despite the increased distribution of web APIs, specifications are rare and their creation and maintenance heavily rely on manual efforts by third parties. In this paper, we propose an automatic approach and an associated tool called D2Spec for extracting significant parts of such specifications from web API documentation pages. Given a seed online documentation page of an API, D2Spec first crawls all documentation pages on the API, and then uses a set of machine-learning techniques to extract the base URL, path templates, and HTTP methods - collectively describing the endpoints of the API.
Jinqiu Yang 0001, Erik Wittern, Annie T. T. Ying, Julian Dolby, Lin Tan 0001
MSR2
2017 Statically checking web API requests in JavaScript
abstract
Many JavaScript applications perform HTTP requests to web APIs, relying on the request URL, HTTP method, and request data to be constructed correctly by string operations. Traditional compile-time error checking, such as calling a non-existent method in Java, are not available for checking whether such requests comply with the requirements of a web API. In this paper, we propose an approach to statically check web API requests in JavaScript. Our approach first extracts a request's URL string, HTTP method, and the corresponding request data using an inter-procedural string analysis, and then checks whether the request conforms to given web API specifications. We evaluated our approach by checking whether web API requests in JavaScript files mined from GitHub are consistent or inconsistent with publicly available API specifications. From the 6575 requests in scope, our approach determined whether the request's URL and HTTP method was consistent or inconsistent with web API specifications with a precision of 96.0%. Our approach also correctly determined whether extracted request data was consistent or inconsistent with the data requirements with a precision of 87.9% for payload data and 99.9% for query data. In a systematic analysis of the inconsistent cases, we found that many of them were due to errors in the client code. The here proposed checker can be integrated with code editors or with continuous integration tools to warn programmers about code containing potentially erroneous requests.
Erik Wittern, Annie T. T. Ying, Yunhui Zheng, Julian Dolby, Jim Laredo
ICSE1
2017 An empirical analysis of the docker container ecosystem on GitHub
abstract
Docker allows packaging an application with its dependencies into a standardized, self-contained unit (a so-called container), which can be used for software development and to run the application on any system. Dockerfiles are declarative definitions of an environment that aim to enable reproducible builds of the container. They can often be found in source code repositories and enable the hosted software to come to life in its execution environment. We conduct an exploratory empirical study with the goal of characterizing the Docker ecosystem, prevalent quality issues, and the evolution of Dockerfiles. We base our study on a data set of over 70000 Dockerfiles, and contrast this general population with samplings that contain the Top-100 and Top-1000 most popular Docker-using projects. We find that most quality issues (28.6%) arise from missing version pinning (i.e., specifying a concrete version for dependencies). Further, we were not able to build 34% of Dockerfiles from a representative sample of 560 projects. Integrating quality checks, e.g., to issue version pinning warnings, into the container build process could result into more reproducible builds. The most popular projects change more often than the rest of the Docker population, with 5.81 revisions per year and 5 lines of code changed on average. Most changes deal with dependencies, that are currently stored in a rather unstructured manner. We propose to introduce an abstraction that, for instance, could deal with the intricacies of different package managers and could improve migration to more light-weight images.
Jürgen Cito, Gerald Schermann, Erik Wittern, Philipp Leitner 0001, Sali Zumberi, Harald C. Gall
MSR3
2017 Who you gonna call?: analyzing web requests in Android applications
abstract
Relying on ubiquitous Internet connectivity, applications on mobile devices frequently perform web requests during their execution. They fetch data for users to interact with, invoke remote functionalities, or send user-generated content or meta-data. These requests collectively reveal common practices of mobile application development, like what external services are used and how, and they point to possible negative effects like security and privacy violations, or impacts on battery life. In this paper, we assess different ways to analyze what web requests Android applications make. We start by presenting dynamic data collected from running 20 randomly selected Android applications and observing their network activity. Next, we present a static analysis tool, Stringoid, that analyzes string concatenations in Android applications to estimate constructed URL strings. Using Stringoid, we extract URLs from 30, 000 Android applications, and compare the performance with a simpler constant extraction analysis. Finally, we present a discussion of the advantages and limitations of dynamic and static analyses when extracting URLs, as we compare the data extracted by Stringoid from the same 20 applications with the dynamically collected data.
Marianna Rapoport, Philippe Suter, Erik Wittern, Ondrej Lhoták, Julian Dolby
MSR3
2016 Benchmarking Web API Quality
David Bermbach, Erik Wittern
ICWE2
2016 A look at the dynamics of the JavaScript package ecosystem
abstract
The node package manager (npm) serves as the frontend to a large repository of JavaScript-based software packages, which foster the development of currently huge amounts of server-side Node. js and client-side JavaScript applications. In a span of 6 years since its inception, npm has grown to become one of the largest software ecosystems, hosting more than 230, 000 packages, with hundreds of millions of package installations every week. In this paper, we examine the npm ecosystem from two complementary perspectives: 1) we look at package descriptions, the dependencies among them, and download metrics, and 2) we look at the use of npm packages in publicly available applications hosted on GitHub. In both perspectives, we consider historical data, providing us with a unique view on the evolution of the ecosystem. We present analyses that provide insights into the ecosystem's growth and activity, into conflicting measures of package popularity, and into the adoption of package versions over time. These insights help understand the evolution of npm, design better package recommendation engines, and can help developers understand how their packages are being used.
Erik Wittern, Philippe Suter, Shriram Rajagopalan
MSR1
2016 Service feature modeling: modeling and participatory ranking of service design alternatives
Erik Wittern, Christian Zirpins
Softw. Syst. Model.1
2014 Feature-Based Configuration of Vendor-Independent Deployments on IaaS
abstract
Infrastructure as a Service (IaaS) is commonly used to deploy distributed systems like Web applications. For every component of a distributed system, IaaS consumers need to select, configure, and deploy a virtual machine (VM), the image to run on the VM, and software to install on the image. These tasks involve multiple decisions, require complex and error-prone manual effort, and may occur unexpectedly. To deal with these challenges, we present a feature-based configuration and vendor-independent deployment approach. An IaaS deployment model allows to describe and automatically perform the distributed system deployment vendor-independently. We use service feature modeling to support the decisions required for configuration. These modeling approaches are combined in a feature-based configuration and deployment process. It can be performed automatically and can be triggered in reaction to unexpected events. We present a proof-of-concept implementation that we use to perform a use case about deploying the Web application of the Barcoo service, showing the approach's applicability.
Erik Wittern, Alexander Lenk, Sebastian Bartenbach, Tobias Braeuer
EDOC1
2014 A Graph-Based Data Model for API Ecosystem Insights
abstract
APIs are increasingly important for companies to enable partners and consumers to access their services and resources. API ecosystems deal with related challenges like publication, promotion and provision of APIs by providers and identification, selection and consumption of APIs by consumers. To address these challenges, to match consumers with relevant APIs, and to support API providers and thus ultimately the ecosystem to evolve, API ecosystems rely on information about APIs, their usage and characteristics, and the social environment around them. We present an extensible, graph-based data model to capture the entities in an API ecosystem and their relations. The data model includes temporal information to capture the evolution of API ecosystems. Analysis operations on top of the data model provide insights for consumers, providers and the ecosystem provider to address the introduced challenges. We present a system implementing the conceptualized data model. We integrate this system with an API ecosystem used in the context of a hackathon event to continuously collect data. We furthermore show the data model's capabilities to represent a well-known dataset about ProgrammableWeb and to drive analysis operations on both datasets.
Erik Wittern, Jim Laredo, Maja Vukovic, Vinod Muthusamy, Aleksander Slominski
ICWS1
2012 Cloud Service Selection Based on Variability Modeling
Erik Wittern, Jörn Kuhlenkamp, Michael Menzel 0002
ICSOC1
2012 Participatory Service Design through Composed and Coordinated Service Feature Models
Erik Wittern, Nelly Schuster, Jörn Kuhlenkamp, Stefan Tai
ICSOC1