James Cook

dblp:58/3899 · DBLP profile ↗
← Back
15ranked-venue papers
10as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Theory of computation · 6 · 6 first-author · 4 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 2 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 first-author
YearPublicationVenuePosition
2026 Efficient Catalytic Graph Algorithms
abstract
We give fast, simple, and implementable catalytic logspace algorithms for two fundamental graph problems. First, a randomized catalytic algorithm for $s\to t$ connectivity running in $\widetilde{O}(nm)$ time, and a deterministic catalytic algorithm for the same running in $\widetilde{O}(n^3 m)$ time. The former algorithm is the first algorithmic use of randomization in $\mathsf{CL}$. The algorithm uses one register per vertex and repeatedly ``pushes'' values along the edges in the graph. Second, a deterministic catalytic algorithm for simulating random walks which in $\widetilde{O}( m T^2 / \varepsilon )$ time estimates the probability a $T$-step random walk ends at a given vertex within $\varepsilon$ additive error. The algorithm uses one register for each vertex and increments it at each visit to ensure repeated visits follow different outgoing edges. Prior catalytic algorithms for both problems did not have explicit runtime bounds beyond being polynomial in $n$.
James Cook, Edward Pyne
ITCS1
2025 The Structure of Catalytic Space: Capturing Randomness and Time via Compression
abstract
STOC ’25, Prague, Czechia
James Cook, Jiatu Li, Ian Mertz, Edward Pyne
STOC1
2024 Defending Against Social Engineering Attacks in the Age of LLMs
abstract
Lin Ai, Tharindu Sandaruwan Kumarage, Amrita Bhattacharjee, Zizhou Liu, Zheng Hui, Michael S. Davinroy, James Cook, Laura Cassani, Kirill Trapeznikov, Matthias Kirchner, Arslan Basharat, Anthony Hoogs, Joshua Garland, Huan Liu, Julia Hirschberg. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Lin Ai, Tharindu Kumarage, Amrita Bhattacharjee, Zizhou Liu, Zheng Hui, Michael Davinroy, James Cook, Laura Cassani, Kirill Trapeznikov, Matthias Kirchner, Arslan Basharat, Anthony Hoogs, Joshua Garland, Huan Liu 0001, Julia Hirschberg
EMNLP7
2024 Tree Evaluation Is in Space O(log n · log log n)
abstract
The Tree Evaluation Problem (TreeEval) (Cook et al. 2009) is a central candidate for separating polynomial time (P) from logarithmic space (L) via composition. While space lower bounds of Ω(log2 n) are known for multiple restricted models, it was recently shown by Cook and Mertz (2020) that TreeEval can be solved in space O(log2 n/loglogn). Thus its status as a candidate hard problem for L remains a mystery. Our main result is to improve the space complexity of TreeEval to O(logn · loglogn), thus greatly strengthening the case that Tree Evaluation is in fact in L. We show two consequences of these results. First, we show that the KRW conjecture (Karchmer, Raz, and Wigderson 1995) implies L ⊈NC1; this itself would have many implications, such as branching programs not being efficiently simulable by formulas. Our second consequence is to increase our understanding of amortized branching programs, also known as catalytic branching programs; we show that every function f on n bits can be computed by such a program of length Poly(n) and width 2O(n).
James Cook, Ian Mertz
STOC1
2023 Optical and Detector Design of the Ocean Color Instrument for the NASA Pace Mission
abstract
The Ocean Color Instrument (OCI) on NASA’s Plankton, Aerosol, Cloud, ocean Ecosystem mission is a hyperspectral imager with high SNR, precision and dynamic range, and with a very low striping artifact level in the 342-887 nm wavelength range with a spectral resolution of 5 nm in 2.5 nm steps, providing a significant technological advancement over previous ocean imagers. To achieve this, OCI is designed with specialized optical imaging and opto-electronic detection systems that push the boundaries of several state-of-the-art technologies. This paper provides an overview of these systems together with their achieved performances and discussions of their key design challenges.
Ulrik Gliese, David Kubalak, Zakk Rhodes, Craig R. Auletti, Sachidananda R. Babu, Branimir Blagojevic, Kasey Boggs, Robert Bousquet, Gregory Bredthauer, Gary L. Brown, Nga T. Cao, Thomas L. Capon, James Champagne, Leland H. Chemerys, Felix N. Chi, Brian L. Clemons, James Cook, William B. Cook, Nicholas P. Costen, Kevin R. Dahya, Paul V. Dizon, Roy Esplin, Robert Estep, Ali Feizi, Steven H. Feng, Eric T. Gorman, Jeffrey Guzek, O. A. Haddad, Claef F. Hakun, Locksley B. Haynes, Michael J. Hersh, Carrie S. Hill, David G. Holliday, Luis Ramos-Izquierdo, Kim S. Jepsen, Emily Kan, Bradford P. Kercheval, Saman Kholdebarin, Joseph J. Knuble, Anh T. La, Erik D. Laurila, Michael R. Lin, Albert J. Mariano, Lane A. Meier, Gerhard Meister, Bryan Monosmith, David Mott, Michael M. Mulloney, Quang V. Nguyen, Thomas J. Nolan, Matthew A. Owens, James Peterson, Manuel A. Quijada, Knute A. Ray, Kenneth Squire, Christopher P. Stull, Joe Thomes, Eugene Waluschka, Yiting Wen, Mark E. Wilson, Jeremy Werdell
IGARSS17
2023 Creating a Public Repository for Joining Private Data
abstract
How can one publish a dataset with sensitive attributes in a way that both preserves privacy and enables joins with other datasets on those same sensitive attributes? This problem arises in many contexts, e.g., a hospital and an airline may want to jointly determine whether people who take long-haul flights are more likely to catch respiratory infections. If they join their data by a common keyed user identifier such as email address, they can determine the answer, though it breaks privacy. This paper shows how the hospital can generate a private sketch and how the airline can privately join with the hospital's sketch by email address. The proposed solution satisfies pure differential privacy and gives approximate answers to linear queries and optimization problems over those joins. Whereas prior work such as secure function evaluation requires sender/receiver interaction, a distinguishing characteristic of the proposed approach is that it is non-interactive. Consequently, the sketch can be published to a repository for any organization to join with, facilitating data discovery. The accuracy of the method is demonstrated through both theoretical analysis and extensive empirical evidence.
James Cook, Milind Shyani, Nina Mishra
NeurIPS1
2022 Trading Time and Space in Catalytic Branching Programs
James Cook, Ian Mertz
CCC1
2021 ReHouSED: A novel measurement of Veteran housing stability using natural language processing
abstract
Housing stability is an important determinant of health. The US Department of Veterans Affairs (VA) administers several programs to assist Veterans experiencing unstable housing. Measuring long-term housing stability of Veterans who receive assistance from VA is difficult due to a lack of standardized structured documentation in the Electronic Health Record (EHR). However, the text of clinical notes often contains detailed information about Veterans' housing situations that may be extracted using natural language processing (NLP). We present a novel NLP-based measurement of Veteran housing stability: Relative Housing Stability in Electronic Documentation (ReHouSED). We first develop and evaluate a system for classifying documents containing information about Veterans' housing situations. Next, we aggregate information from multiple documents to derive a patient-level measurement of housing stability. Finally, we demonstrate this method's ability to differentiate between Veterans who are stably and unstably housed. Thus, ReHouSED provides an important methodological framework for the study of long-term housing stability among Veterans receiving housing assistance.
Alec B. Chapman, Audrey L. Jones, A. Taylor Kelley, Barbara E. Jones, Lori Gawron, Ann Elizabeth Montgomery, Thomas Byrne, Ying Suo, James Cook, Warren B. P. Pettey, Kelly S. Peterson, Makoto Jones, Richard Nelson
J. Biomed. Informatics9
2020 BusTr: Predicting Bus Travel Times from Real-Time Traffic
abstract
We present BusTr, a machine-learned model for translating road traffic forecasts into predictions of bus delays, used by Google Maps to serve the majority of the world's public transit systems where no official real-time bus tracking is provided. We demonstrate that our neural sequence model improves over DeepTTE, the state-of-the-art baseline, both in performance (-30% MAPE) and training stability. We also demonstrate significant generalization gains over simpler models, evaluated on longitudinal data to cope with a constantly evolving world.
Richard Barnes 0002, Senaka Buthpitiya, James Cook, Alex Fabrikant, Andrew Tomkins, Fangzhou Xu
KDD3
2020 Catalytic approaches to the tree evaluation problem
abstract
The study of branching programs for the Tree Evaluation Problem (TreeEval), introduced by S. Cook et al. (TOCT 2012), remains one of the most promising approaches to separating L from P. Given a label in [k] at each leaf of a complete binary tree and an explicit function in [k]2 → [k] for recursively computing the value of each internal node from its children, the problem is to compute the value at the root node. (While the original problem allows an arbitrary-degree tree, we focus on binary trees.) The problem is parameterized by the alphabet size k and the height h of the tree. A branching program implementing the straightforward recursive algorithm uses Θ((k + 1) h ) states, organized into 2 h −1 layers of width up to k h . Until now no better deterministic algorithm was known.
James Cook, Ian Mertz
STOC1
2019 Hard to Park?: Estimating Parking Difficulty at Scale
abstract
In this paper we consider the problem of estimating the difficulty of parking at a particular time and place; this problem is a critical sub-component for any system providing parking assistance to users. We describe an approach to this problem that is currently in production in Google Maps, providing inferences in cities across the world. We present a wide range of features intended to capture different aspects of parking difficulty and study their effectiveness both alone and in combination. We also evaluate various model architectures for the prediction problem. Finally, we present challenges faced in estimating parking difficulty in different regions of the world, and the approaches we have taken to address them.
Neha Arora 0001, James Cook, Ravi Kumar 0001, Yechen Li, Huai-Jen Liang, Andrew Tomkins, Iveel Tsogsuren
KDD2
2013 How to grow more pairs: suggesting review targets for comparison-friendly review ecosystems
abstract
We consider the algorithmic challenges behind a novel interface that simplifies consumer research of online reviews by surfacing relevant comparable review bundles: reviews for two or more of the items being researched, all generated in similar enough circumstances to provide for easy comparison. This can be reviews by the same reviewer, or by the same demographic category of reviewer, or reviews focusing on the same aspect of the items. But such an interface will work only if the review ecosystem often has comparable review bundles for common research tasks.
James Cook, Alex Fabrikant, Avinatan Hassidim
WWW1
2013 Group chats on Twitter
abstract
We report on a new kind of group conversation on Twitter that we call a group chat. These chats are periodic, synchronized group conversations focused on specific topics and they exist at a massive scale. The groups and the members of these groups are not explicitly known. Rather, members agree on a hashtag and a meeting time (e.g, 3pm Pacific Time every Wednesday) to discuss a subject of interest. Topics of these chats are numerous and varied. Some are support groups, for example, post-partum depression and mood disorder groups. Others are about a passionate interest: topics include skiing, photography, movies, wine and foodie communities. We develop a definition of a group that is inspired by how sociologists define groups and present an algorithm for discovering groups. We prove that our algorithms find all groups under certain assumptions. While these groups are of course known to the people who participate in the discussions, what we do not believe is known is the scale and variety of groups. We provide some insight into the nature of these groups based on over two years of tweets. Finally, we show that group chats are a growing phenomenon on Twitter and hope that reporting their existence propels their growth even further.
James Cook, Krishnaram Kenthapadi, Nina Mishra
WWW1
2012 Your two weeks of fame and your grandmother's
abstract
Did celebrity last longer in 1929, 1992 or 2009? We investigate the phenomenon of fame by mining a collection of news articles that spans the twentieth century, and also perform a side study on a collection of blog posts from the last 10 years. By analyzing mentions of personal names, we measure each person's time in the spotlight, and watch the distribution change from a century ago to a year ago. We expected to find a trend of decreasing durations of fame as news cycles accelerated and attention spans became shorter. Instead, we find a remarkable consistency through most of the period we study. Through a century of rapid technological and societal change, through the appearance of Twitter, communication satellites and the Internet, we do not observe a significant change in typical duration of celebrity. We also study the most famous of the famous, and find different results depending on our method for measuring duration of fame. With a method that may be thought of as measuring a spike of attention around a single narrow news story, we see the same result as before: stories last as long now as they did in 1930. A second method, which may be thought of as measuring the duration of public interest in a person, indicates that famous people's presence in the news is becoming longer rather than shorter, an effect most likely driven by the wider distribution and higher volume of media in modern times. Similar studies have been done with much shorter timescales specifically in the context of information spreading on Twitter and similar social networking site. However, to the best of our knowledge, this is the first massive scale study of this nature that spans over a century of archived data, thereby allowing us to track changes across decades.
James Cook, Atish Das Sarma, Alex Fabrikant, Andrew Tomkins
WWW1
2009 Goldreich's One-Way Function Candidate and Myopic Backtracking Algorithms
James Cook, Omid Etesami, Rachel Miller, Luca Trevisan 0001
TCC1