EDBT 2026 Demo / reviewers in the wild / expert
Asher Trockman
dblp:221/1541
· DBLP profile ↗
10ranked-venue papers
6as first author
4since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 6 · 3 first-authorArtificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Deep learning architectures and training · 64% Efficient and distributed learning · 12% Trustworthy machine learning · 12% | |
| Software engineering, system software, and programming languages
3 papers |
Empirical software engineering · 51% Software maintenance and evolution · 49% |
Topics — the 13 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training
convolutional neural network |
1.2 | 2 | 2023 | Understanding the Covariance Structure of Convolutional Filters · ICLR 2023 Orthogonalizing Convolutional Layers with the Cayley Transform · ICLR 2021 |
Empirical software engineering
mining software repositories |
1.1 | 3 | 2020 | Heard it through the Gitvine: an empirical study of tool diffusion across the npm ecosystem · ESEC/SIGSOFT FSE 2020 Tool choice matters: JavaScript quality assurance tools and usage outcomes in GitHub projects · ICSE 2019 Adding sparkle to social coding: an empirical study of repository badges in the npm ecosystem · ICSE 2018 |
Machine learning › Efficient and distributed learning › model compression › knowledge distillation
model distillation |
0.9 | 1 | 2025 | Antidistillation Sampling · NeurIPS 2025 |
Machine learning › Trustworthy machine learning
model security |
0.9 | 1 | 2025 | Antidistillation Sampling · NeurIPS 2025 |
Machine learning › Deep learning architectures and training › attention mechanism
self-attention |
0.7 | 1 | 2023 | Mimetic Initialization of Self-Attention Layers · ICML 2023 |
Machine learning › Deep learning architectures and training › neural network training
training on small datasets |
0.7 | 1 | 2023 | Mimetic Initialization of Self-Attention Layers · ICML 2023 |
Machine learning › Deep learning architectures and training
transformer |
0.7 | 1 | 2023 | Mimetic Initialization of Self-Attention Layers · ICML 2023 |
Machine learning › Deep learning architectures and training
weight initialization |
0.7 | 1 | 2023 | Mimetic Initialization of Self-Attention Layers · ICML 2023 |
Software maintenance and evolution
software ecosystems |
0.6 | 3 | 2020 | Tool choice matters: JavaScript quality assurance tools and usage outcomes in GitHub projects · ICSE 2019 Heard it through the Gitvine: an empirical study of tool diffusion across the npm ecosystem · ESEC/SIGSOFT FSE 2020 Adding sparkle to social coding: an empirical study of repository badges in the npm ecosystem · ICSE 2018 |
Machine learning › Deep learning architectures and training › convolutional neural network › convolution design
orthogonal convolution |
0.5 | 1 | 2021 | Orthogonalizing Convolutional Layers with the Cayley Transform · ICLR 2021 |
Software maintenance and evolution › software ecosystems
dependency management |
0.4 | 1 | 2019 | Tool choice matters: JavaScript quality assurance tools and usage outcomes in GitHub projects · ICSE 2019 |
Natural language and speech › Language models and text generation › large language model reasoning
reasoning traces |
0.3 | 1 | 2025 | Antidistillation Sampling · NeurIPS 2025 |
Software maintenance and evolution › software ecosystems
npm ecosystem |
0.1 | 1 | 2020 | Heard it through the Gitvine: an empirical study of tool diffusion across the npm ecosystem · ESEC/SIGSOFT FSE 2020 |
Methods — techniques the papers use, named apart from their topics
token probability distribution · 0.9sampling strategy · 0.9spectral analysis · 0.7mimetic initialization · 0.7covariance analysis · 0.7orthogonal parameterization · 0.5cayley transform · 0.5survival analysis · 0.4network science · 0.4regression analysis · 0.4mixed-methods · 0.4time series analysis · 0.3survey · 0.3statistical modeling · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Antidistillation SamplingabstractFrontier models that generate extended reasoning traces inadvertently produce token sequences that can facilitate model distillation. Recognizing this vulnerability, model owners may seek sampling strategies that limit the effectiveness of distillation without compromising model performance. *Antidistillation sampling* provides exactly this capability. By strategically modifying a model's next-token probability distribution, antidistillation sampling poisons reasoning traces, rendering them significantly less effective for distillation while preserving the model's utility. Yash Savani, Asher Trockman, Zhili Feng, Yixuan Even Xu, Avi Schwarzschild, Alexander Robey, Marc Finzi, J. Zico Kolter |
NeurIPS | 2 |
| 2023 | Understanding the Covariance Structure of Convolutional Filters
Asher Trockman, Devin Willmott, J. Zico Kolter |
ICLR | 1 |
| 2023 | Mimetic Initialization of Self-Attention LayersabstractIt is notoriously difficult to train Transformers on small datasets; typically, large pre-trained models are instead used as the starting point. We explore the weights of such pre-trained Transformers (particularly for vision) to attempt to find reasons for this discrepancy. Surprisingly, we find that simply initializing the weights of self-attention layers so that they "look" more like their pre-trained counterparts allows us to train vanilla Transformers faster and to higher final accuracies, particularly on vision tasks such as CIFAR-10 and ImageNet classification, where we see gains in accuracy of over 5% and 4%, respectively. Our initialization scheme is closed form, learning-free, and very simple: we set the product of the query and key weights to be approximately the identity, and the product of the value and projection weights to approximately the negative identity. As this mimics the patterns we saw in pre-trained Transformers, we call the technique "mimetic initialization". Asher Trockman, J. Zico Kolter |
ICML | 1 |
| 2021 | Orthogonalizing Convolutional Layers with the Cayley Transform
Asher Trockman, J. Zico Kolter |
ICLR | 1 |
| 2020 | Heard it through the Gitvine: an empirical study of tool diffusion across the npm ecosystemabstractAutomation tools like continuous integration services, code coverage reporters, style checkers, dependency managers, etc. are all known to provide significant improvements in developer productivity and software quality. Some of these tools are widespread, others are not. How do these automation "best practices" spread? And how might we facilitate the diffusion process for those that have seen slower adoption? In this paper, we rely on a recent innovation in transparency on code hosting platforms like GitHub---the use of repository badges---to track how automation tools spread in open-source ecosystems through different social and technical mechanisms over time. Using a large longitudinal data set, multivariate network science techniques, and survival analysis, we study which socio-technical factors can best explain the observed diffusion process of a number of popular automation tools. Our results show that factors such as social exposure, competition, and observability affect the adoption of tools significantly, and they provide a roadmap for software engineers and researchers seeking to propagate best practices and tools. Hemank Lamba, Asher Trockman, Daniel Armanios, Christian Kästner, Heather Miller, Bogdan Vasilescu |
ESEC/SIGSOFT FSE | 2 |
| 2019 | Tool choice matters: JavaScript quality assurance tools and usage outcomes in GitHub projectsabstractQuality assurance automation is essential in modern software development. In practice, this automation is supported by a multitude of tools that fit different needs and require developers to make decisions about which tool to choose in a given context. Data and analytics of the pros and cons can inform these decisions. Yet, in most cases, there is a dearth of empirical evidence on the effectiveness of existing practices and tool choices. We propose a general methodology to model the time-dependent effect of automation tool choice on four outcomes of interest: prevalence of issues, code churn, number of pull requests, and number of contributors, all with a multitude of controls. On a large data set of npm JavaScript projects, we extract the adoption events for popular tools in three task classes: linters, dependency managers, and coverage reporters. Using mixed methods approaches, we study the reasons for the adoptions and compare the adoption effects within each class, and sequential tool adoptions across classes. We find that some tools within each group are associated with more beneficial outcomes than others, providing an empirical perspective for the benefits of each. We also find that the order in which some tools are implemented is associated with varying outcomes. David Kavaler, Asher Trockman, Bogdan Vasilescu, Vladimir Filkov |
ICSE | 2 |
| 2019 | A panel data set of cryptocurrency development activity on GitHubabstractCryptocurrencies are a significant development in recent years, featuring in global news, the financial sector, and academic research. They also hold a significant presence in open source development, comprising some of the most popular repositories on GitHub. Their openly developed software artifacts thus present a unique and exclusive avenue to quantitatively observe human activity, effort, and software growth for cryptocurrencies. Our data set marks the first concentrated effort toward high-fidelity panel data of cryptocurrency development for a wide range of metrics. The data set is foremost a quantitative measure of developer activity for budding open source cryptocurrency development. We collect metrics like daily commits, contributors, lines of code changes, stars, forks, and subscribers. We also include financial data for each cryptocurrency: the daily price and market capitalization. The data set includes data for 236 cryptocurrencies for 380 days (roughly January 2018 to January 2019). We discuss particularly interesting research opportunities for this combination of data, and release new tooling to enable continuing data collection for future research opportunities as development and application of cryptocurrencies mature. Rijnard van Tonder, Asher Trockman, Claire Le Goues |
MSR | 2 |
| 2019 | Striking gold in software repositories?: an econometric study of cryptocurrencies on GitHubabstractCryptocurrencies have a significant open source development presence on GitHub. This presents a unique opportunity to observe their related developer effort and software growth. Individual cryptocurrency prices are partly driven by attractiveness, and we hypothesize that high-quality, actively-developed software is one of its influences. Thus, we report on a study of a panel data set containing nearly a year of daily observations of development activity, popularity, and market capitalization for over two hundred open source cryptocurrencies. We find that open source project popularity is associated with higher market capitalization, though development activity and quality assurance practices are insignificant variables in our models. Using Granger causality tests, we find no compelling evidence for a dynamic relation between market capitalization and metrics such as daily stars, forks, watchers, commits, contributors, and lines of code changed. Asher Trockman, Rijnard van Tonder, Bogdan Vasilescu |
MSR | 1 |
| 2018 | Adding sparkle to social coding: an empirical study of repository badges in the npm ecosystemabstractIn fast-paced, reuse-heavy, and distributed software development, the transparency provided by social coding platforms like GitHub is essential to decision making. Developers infer the quality of projects using visible cues, known as signals, collected from personal profile and repository pages. We report on a large-scale, mixed-methods empirical study of npm packages that explores the emerging phenomenon of repository badges, with which maintainers signal underlying qualities about their projects to contributors and users. We investigate which qualities maintainers intend to signal and how well badges correlate with those qualities. After surveying developers, mining 294,941 repositories, and applying statistical modeling and time-series analyses, we find that non-trivial badges, which display the build status, test coverage, and up-to-dateness of dependencies, are mostly reliable signals, correlating with more tests, better pull requests, and fresher dependencies. Displaying such badges correlates with best practices, but the effects do not always persist. Asher Trockman, Shurui Zhou, Christian Kästner, Bogdan Vasilescu |
ICSE | 1 |
| 2018 | "Automatically assessing code understandability" reanalyzed: combined metrics matterabstractPrevious research shows that developers spend most of their time understanding code. Despite the importance of code understandability for maintenance-related activities, an objective measure of it remains an elusive goal. Recently, Scalabrino et al. reported on an experiment with 46 Java developers designed to evaluate metrics for code understandability. The authors collected and analyzed data on more than a hundred features describing the code snippets, the developers' experience, and the developers' performance on a quiz designed to assess understanding. They concluded that none of the metrics considered can individually capture understandability. Expecting that understandability is better captured by a combination of multiple features, we present a reanalysis of the data from the Scalabrino et al. study, in which we use different statistical modeling techniques. Our models suggest that some computed features of code, such as those arising from syntactic structure and documentation, have a small but significant correlation with understandability. Further, we construct a binary classifier of understandability based on various interpretable code features, which has a small amount of discriminating power. Our encouraging results, based on a small data set, suggest that a useful metric of understandability could feasibly be created, but more data is needed. Asher Trockman, Keenen Cates, Mark Mozina, Christian Kästner, Bogdan Vasilescu |
MSR | 1 |