Scott Gray

dblp:191/4641 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
3since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Generative modeling · 35% Deep learning architectures and training · 17% Transfer learning and domain adaptation · 15%
Computer architecture, parallel and distributed computing, and storage systems
3 papers
Cloud and datacenter computing · 43% Distributed systems · 33% Hardware accelerators and domain-specific architectures · 15%
Databases, data mining, and information retrieval
1 paper
Database system architecture and tuning · 50% Indexing and storage engines · 50%

Topics — the 13 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
autoregressive model
0.512021
Zero-Shot Text-to-Image Generation · ICML 2021
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.512021
Zero-Shot Text-to-Image Generation · ICML 2021
Natural language and speech › Language models and text generation › neural language model
autoregressive language model
0.412020
Language Models are Few-Shot Learners · NeurIPS 2020
Machine learning › Transfer learning and domain adaptation
few-shot learning
0.412020
Language Models are Few-Shot Learners · NeurIPS 2020
Indexing and storage engines
secondary index
0.412019
FoundationDB Record Layer: A Multi-Tenant Structured Datastore · SIGMOD Conference 2019
Cloud and datacenter computing › cloud storage
multi-tenant cloud storage
0.312018
CloudKit: Structured Storage for Mobile Applications · Proc. VLDB Endow. 2018
Machine learning › Deep learning architectures and training
convolutional neural network
0.212016
Fast Algorithms for Convolutional Neural Networks · CVPR 2016
Machine learning › Deep learning architectures and training › efficient deep learning
efficient convolution
0.212016
Fast Algorithms for Convolutional Neural Networks · CVPR 2016
Machine learning › Efficient and distributed learning
model compression
0.212016
Fast Algorithms for Convolutional Neural Networks · CVPR 2016
Cloud and datacenter computing
cloud storage
0.112018
CloudKit: Structured Storage for Mobile Applications · Proc. VLDB Endow. 2018
Storage systems › data management
petabyte-scale data management
0.112018
CloudKit: Structured Storage for Mobile Applications · Proc. VLDB Endow. 2018
Hardware accelerators and domain-specific architectures › machine learning accelerator
CNN accelerator
0.112016
Fast Algorithms for Convolutional Neural Networks · CVPR 2016
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.112016
Fast Algorithms for Convolutional Neural Networks · CVPR 2016

Methods — techniques the papers use, named apart from their topics

schema management · 0.8indexing · 0.8winograd minimal filtering · 0.5transformer · 0.5token stream modeling · 0.5FFT-based convolution · 0.5in-context learning · 0.4autoregressive language model · 0.4
YearPublicationVenuePosition
2023 Introducing Computational Thinking in Middle-Schools using a Culturally-responsive Game through a Researcher-Practitioner Partnership
abstract
There is a national need to increase the number of minority students entering STEM fields with essential computing skills. To increase minority students’ interest and engagement in computing, a researcher-practitioner partnership between the University of Texas at El Paso and the El Paso Independent School District, developed and implemented a culturally and linguistically responsive curriculum and pedagogy to introduce computational thinking (CT) in two middle schools across different subject areas in a borderland region. The curriculum leveraged the Sol y Agua game – a bilingual, culturally-responsive game designed to engage students of this region in CT. This paper describes the process and initial findings of this project. The quantitative data from in-game analyses show that students utilized the language change feature to switch from English to Spanish more frequently than the other way – highlighting the need for educational platforms relatable to students through language, environment, and cultural context. Analyses of the qualitative data indicate that while teachers/team members understood CT and translanguaging concepts and taught lesson units that provided opportunities to practice both, CT and translanguaging were largely implicit in the curriculum. In collaborative analyses of these patterns, teachers described additional supports that would help them to make CT instruction and translanguaging strategies more explicit in the content and pedagogy, highlighting the need for systematic, targeted integration of these concepts.
Ismael Villanueva-Miranda, Katherine Mortimer, Monika Akbar, Romelia Rodriguez Reyes, Cynthia Ontiveros, Scott Gray, Pedro Delgado, Victor Medrano, Melissa Anderson, Jacob Ramirez, Jesus Vazquez
IEEE Big Data6
2022 The Sol y Agua RPP: A Bilingual and Culturally Responsive Approach to Introduce Computational Thinking in Middle School
abstract
The Sol y Agua researcher-practitioner partnership (RPP) project introduces computational thinking (CT) in the middle school of the Paso del Norte region using a linguistically and culturally responsive approach. At the core of this RPP is the Sol y Agua game, a bilingual, culturally- and environmentally-relevant educational game developed at the University of Texas at El Paso to introduce computing and STEM topics in middle school. The Sol y Agua RPP includes some critical areas for a successful RPP, including partnership building and the focus on a linguistically and culturally-responsive pedagogy and content development. We describe our approach to build a sustainable RPP, incorporating bilingual pedagogy, and integrating CT through a culturally- and environmentally-relevant game as part of our RPP experience.
Monika Akbar, Katherine Mortimer, Grecia Navarrete, Stephanie Galvan, George Molina, Romelia Rodriguez Reyes, Cynthia Ontiveros, Scott Gray, Sarah Escandon, Monica Lyons, Pedro Delgado, Victor Medrano, Haleigh Kneedler, Patricia Benitez, Jacob Ramirez, Jesus Vazquez, Melissa Anderson
SIGCSE (2)8
2021 Zero-Shot Text-to-Image Generation
abstract
Text-to-image generation has traditionally focused on finding better modeling assumptions for training on a fixed dataset. These assumptions might involve complex architectures, auxiliary losses, or side information such as object part labels or segmentation masks supplied during training. We describe a simple approach for this task based on a transformer that autoregressively models the text and image tokens as a single stream of data. With sufficient data and scale, our approach is competitive with previous domain-specific models when evaluated in a zero-shot fashion.
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen 0003, Ilya Sutskever
ICML4
2020 Language Models are Few-Shot Learners
abstract
We demonstrate that scaling up language models greatly improves task-agnostic, few-shot performance, sometimes even becoming competitive with prior state-of-the-art fine-tuning approaches. Specifically, we train GPT-3, an autoregressive language model with 175 billion parameters, 10x more than any previous non-sparse language model, and test its performance in the few-shot setting. For all tasks, GPT-3 is applied without any gradient updates or fine-tuning, with tasks and few-shot demonstrations specified purely via text interaction with the model. GPT-3 achieves strong performance on many NLP datasets, including translation, question-answering, and cloze tasks. We also identify some datasets where GPT-3's few-shot learning still struggles, as well as some datasets where GPT-3 faces methodological issues related to training on large web corpora.
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Thomas Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeff Wu 0003, Clemens Winter, Christopher Hesse, Mark Chen 0003, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, Dario Amodei
NeurIPS24
2019 FoundationDB Record Layer: A Multi-Tenant Structured Datastore
abstract
The FoundationDB Record Layer is an open source library that provides a record-oriented data store with semantics similar to a relational database implemented on top of FoundationDB, an ordered, transactional key-value store. The Record Layer provides a lightweight, highly extensible way to store structured data. It offers schema management and a rich set of query and indexing facilities, some of which are not usually found in traditional relational databases, such as nested record types, indexes on commit versions, and indexes that span multiple record types. The Record Layer is stateless and built for massive multi-tenancy, encapsulating and isolating all of a tenant's state, including indexes, into a separate logical database. We demonstrate how the Record Layer is used by CloudKit, Apple's cloud backend service, to provide powerful abstractions to applications serving hundreds of millions of users. CloudKit uses the Record Layer to host billions of independent databases, many with a common schema. Features provided by the Record Layer enable CloudKit to provide richer APIs and stronger semantics with reduced maintenance overhead and improved scalability.
Christos Chrysafis, Ben Collins, Scott Dugas, Jay Dunkelberger, Moussa Ehsan, Scott Gray, Alec Grieser, Ori Herrnstadt, Kfir Lev-Ari, Mike McMahon, Nicholas Schiefer, Alexander Shraer
SIGMOD Conference6
2018 CloudKit: Structured Storage for Mobile Applications
abstract
CloudKit is Apple's cloud backend service and application development framework that provides strongly-consistent storage for structured data and makes it easy to synchronize data across user devices or share it among multiple users. Launched more than 3 years ago, CloudKit forms the foundation for more than 50 Apple apps, including many of our most important and popular applications such as Photos, iCloud Drive, Notes, Keynote, and News, as well as many third-party apps. To deliver this at large scale, CloudKit explicitly leverages multi-tenancy at the application level as well as at the user level to guide efficient data placement and distribution. By using CloudKit application developers are free to focus on delivering the application front-end and logic while relying on CloudKit for scale, consistency, durability and security. CloudKit manages petabytes of data and handles hundreds of millions of users around the world on a daily basis.
Alexander Shraer, Alexandre Aybes, Bryan Davis, Christos Chrysafis, Dave Browning, Eric Krugler, Eric Stone, Harrison Chandler, Jacob Farkas, Jonathan Ruben, Michael Ford, Mike McMahon, Nathan Williams, Nicolas Favre-Felix, Nihar Sharma, Ori Herrnstadt, Paul Seligman, Raghav Pisolkar, Scott Dugas, Scott Gray, Shirley Lu, Sytze Harkema, Valentin Kravtsov, Vanessa Hong, Wan Ling Yih, Yizuo Tian
Proc. VLDB Endow.21
2016 Fast Algorithms for Convolutional Neural Networks
abstract
Deep convolutional neural networks take GPU-days of computation to train on large data sets. Pedestrian detection for self driving cars requires very low latency. Image recognition for mobile phones is constrained by limited processing resources. The success of convolutional neural networks in these situations is limited by how fast we can compute them. Conventional FFT based convolution is fast for large filters, but state of the art convolutional neural networks use small, 3 3 filters. We introduce a new class of fast algorithms for convolutional neural networks using Winograd's minimal filtering algorithms. The algorithms compute minimal complexity convolution over small tiles, which makes them fast with small filters and small batch sizes. We benchmark a GPU implementation of our algorithm with the VGG network and show state of the art throughput at batch sizes from 1 to 64.
Andrew Lavin, Scott Gray
CVPR2