Jack Clark

dblp:127/1112 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
3since 2021 · last 2024
0000-0003-3886-7657ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Transaction processing and concurrency control · 75% Data mining · 25%
Software engineering, system software, and programming languages
2 papers
Compilers and program optimization · 55% Software testing · 37% Program verification · 8%
Artificial intelligence
2 papers
Transfer learning and domain adaptation · 41% Vision and language · 28% Language models and text generation · 19%

Topics — the 14 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining › anomaly detection › outlier detection
isolation-based anomaly detection
0.812024
Validating Database System Isolation Level Implementations with Version Certificate Recovery · EuroSys 2024
Transaction processing and concurrency control
isolation level verification
0.812024
Validating Database System Isolation Level Implementations with Version Certificate Recovery · EuroSys 2024
Transaction processing and concurrency control › serializability
serializability violation detection
0.812024
Validating Database System Isolation Level Implementations with Version Certificate Recovery · EuroSys 2024
Transaction processing and concurrency control
transaction isolation
0.812024
Validating Database System Isolation Level Implementations with Version Certificate Recovery · EuroSys 2024
Software testing › fuzzing › system software fuzzing
compiler fuzzing
0.712023
Taking Back Control in an Intermediate Representation for GPU Computing · Proc. ACM Program. Lang. 2023
Compilers and program optimization
intermediate representation
0.712023
Taking Back Control in an Intermediate Representation for GPU Computing · Proc. ACM Program. Lang. 2023
Compilers and program optimization
verified compilation
0.712023
Taking Back Control in an Intermediate Representation for GPU Computing · Proc. ACM Program. Lang. 2023
Computer vision › Vision and language › vision-language pretraining
contrastive vision-language pretraining
0.512021
Learning Transferable Visual Models From Natural Language Supervision · ICML 2021
Machine learning › Transfer learning and domain adaptation
zero-shot transfer
0.512021
Learning Transferable Visual Models From Natural Language Supervision · ICML 2021
Natural language and speech › Language models and text generation › neural language model
autoregressive language model
0.412020
Language Models are Few-Shot Learners · NeurIPS 2020
Machine learning › Transfer learning and domain adaptation
few-shot learning
0.412020
Language Models are Few-Shot Learners · NeurIPS 2020
Software testing › database testing
database system testing
0.212024
Validating Database System Isolation Level Implementations with Version Certificate Recovery · EuroSys 2024
Program verification
formal modeling
0.212023
Taking Back Control in an Intermediate Representation for GPU Computing · Proc. ACM Program. Lang. 2023
Computer vision › Vision and language › cross-modal supervision
natural language supervision
0.112021
Learning Transferable Visual Models From Natural Language Supervision · ICML 2021

Methods — techniques the papers use, named apart from their topics

white-box checking · 1.5version certificate recovery · 1.5expected serialization order · 1.5fuzzing · 0.7formal modeling · 0.7pre-training · 0.5contrastive learning · 0.5in-context learning · 0.4autoregressive language model · 0.4
YearPublicationVenuePosition
2024 Validating Database System Isolation Level Implementations with Version Certificate Recovery
abstract
Transactions are a key feature of database systems and isolation levels specify the behavior of concurrently executing transactions. Ensuring their correct behavior is crucial. Recently, many isolation anomalies have been found in production database systems. Checkers can be used to validate that a particular execution conforms to a desired isolation level. However, state-of-the-art checkers cannot handle predicate operations, which are both common in real-world workloads and essential for distinguishing between the repeatable read and serializable isolation levels. In this work, we address this issue by proposing an efficient white-box checker, Emme. Our key idea is to use information that is easily provided by database systems to efficiently check the isolation level of a given transaction history. We present version certificate recovery, a method of recovering the version order and each operation's version from the database system under test. For efficiency, we also propose the concept of an expected serialization order, which obviates the need to define and recover a version certificate for many serializable concurrency control protocols. We have implemented version certificate recovery for three widely used database systems---PostgreSQL, CockroachDB, and TiDB. We demonstrate that Emme is 1.2-3.6× faster than Elle, a state-of-the-art checker. Using the expected serialization order, we obtain a further speedup of 34-430× compared to Emme when checking histories containing predicate operations. We show that our approach can identify invalid histories that cannot be detected by Elle and also show that it can find realistic bugs purposely introduced by an engineer.
Jack Clark, Alastair F. Donaldson, John Wickerson, Manuel Rigger
EuroSys1
2023 Taking Back Control in an Intermediate Representation for GPU Computing
abstract
We describe our experiences successfully applying lightweight formal methods to substantially improve and reformulate an important part of Standard Portable Intermediate Representation SPIRV, an industry-standard language for GPU computing. The formal model that we present has allowed us to (1) identify several ambiguities and needless complexities in the way that structured control flow was defined in the SPIRV specification; (2) interact with the authors of the SPIRV specification to rectify these problems; (3) validate the developer tools and conformance test suites that support the SPIRV language by cross-checking them against our formal model, improving the tools, test suites, and our models in the process; and (4) develop a novel method for fuzzing SPIRV compilers to detect miscompilation bugs that leverages our formal model. The latest release of the SPIRV specification incorporates the revised set of control-flow definitions that have arisen from our work. Furthermore, our novel compiler-fuzzing technique has led to the discovery of twenty distinct, previously unknown bugs in SPIRV compilers from Google, the Khronos Group, Intel, and Mozilla. Our work showcases the practical impact that formal modelling and analysis techniques can have on the design and implementation of industry-standard programming languages.
Vasileios Klimis, Jack Clark, Alan Baker, David Neto, John Wickerson, Alastair F. Donaldson
Proc. ACM Program. Lang.2
2021 Learning Transferable Visual Models From Natural Language Supervision
abstract
State-of-the-art computer vision systems are trained to predict a fixed set of predetermined object categories. This restricted form of supervision limits their generality and usability since additional labeled data is needed to specify any other visual concept. Learning directly from raw text about images is a promising alternative which leverages a much broader source of supervision. We demonstrate that the simple pre-training task of predicting which caption goes with which image is an efficient and scalable way to learn SOTA image representations from scratch on a dataset of 400 million (image, text) pairs collected from the internet. After pre-training, natural language is used to reference learned visual concepts (or describe new ones) enabling zero-shot transfer of the model to downstream tasks. We study the performance of this approach by benchmarking on over 30 different existing computer vision datasets, spanning tasks such as OCR, action recognition in videos, geo-localization, and many types of fine-grained object classification. The model transfers non-trivially to most tasks and is often competitive with a fully supervised baseline without the need for any dataset specific training. For instance, we match the accuracy of the original ResNet-50 on ImageNet zero-shot without needing to use any of the 1.28 million training examples it was trained on.
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, Ilya Sutskever
ICML10
2020 Language Models are Few-Shot Learners
abstract
We demonstrate that scaling up language models greatly improves task-agnostic, few-shot performance, sometimes even becoming competitive with prior state-of-the-art fine-tuning approaches. Specifically, we train GPT-3, an autoregressive language model with 175 billion parameters, 10x more than any previous non-sparse language model, and test its performance in the few-shot setting. For all tasks, GPT-3 is applied without any gradient updates or fine-tuning, with tasks and few-shot demonstrations specified purely via text interaction with the model. GPT-3 achieves strong performance on many NLP datasets, including translation, question-answering, and cloze tasks. We also identify some datasets where GPT-3's few-shot learning still struggles, as well as some datasets where GPT-3 faces methodological issues related to training on large web corpora.
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Thomas Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeff Wu 0003, Clemens Winter, Christopher Hesse, Mark Chen 0003, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, Dario Amodei
NeurIPS26