VLDB 2026 Research / reviewers in the wild / expert
Jerry Zhang
dblp:39/9312
· DBLP profile ↗
9ranked-venue papers
1as first author
6since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Artificial intelligence and machine learning · 2Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Understanding the Security Impact of CHERI on the Operating System KernelabstractCapability Hardware Enhanced RISC Instructions (CHERI) is a set of hardware extensions that allow enforcement of spatial and temporal safety for unsafe programming languages like C. CHERI utilizes the concept of hardware capabilities to enforce bounds checks on all memory accesses and a hardware-assisted revocation scheme to enforce temporal safety. In theory, CHERI offers a surprising mix of practicality and strong security guarantees for traditionally unsafe environments like operating system kernels: capability extensions block a range of safety-related vulnerabilities common to low-level systems code while requiring only modest engineering effort. Our work takes a deep look at the potential impact of CHERI on the security of a commodity operating system kernel. We analyze a total of 439 kernel vulnerabilities in Linux and FreeBSD kernels. Our analysis shows that CHERI can block 61 % of kernel vulnerabilities if temporal safety is implemented in the kernel (35% if capability revocation is off). Enabling CHERI requires a modest effort, e.g., porting the FreeBSD kernel to support pure-capability mode of execution took 7 months. Finally, compared to Rust, which is able to mitigate 84% of kernel exploits, CHERI achieves the rate of 70% (38% if revocation is off). While CHERI is less effective, enabling it in the kernel requires a much lower development effort. Zhaofeng Li 0004, Jerry Zhang, Joshua Tlatelpa-Agustin, Anton Burtsev |
ACSAC | 2 |
| 2025 | Cross-Batch Aggregation for Streaming Learning from Label Proportions in Industrial-Scale Recommendation Systems
Jonathan Valverde, Tiansheng Yao, Xiang Li 0124, Yin Zhang 0011, Andrew Evdokimov, Adam Kraft, Samuel Ieong, Jerry Zhang, Ed H. Chi, Zhiyuan Cheng 0002 |
RecSys | 9 |
| 2025 | Atmosphere: Practical Verified Kernels with Rust and VerusabstractRecent advances in programming languages and automated formal reasoning have changed the balance between the complexity and practicality of developing formally verified systems. Our work leverages Verus, a new verifier for Rust that combines ideas of linear types, permissioned reasoning, and automated verification based on satisfiability modulo theories (SMT), for the development of a formally verified microkernel, Atmosphere. Zhaofeng Li 0004, Jerry Zhang, Vikram Narayanan, Anton Burtsev |
SOSP | 3 |
| 2024 | Rust for Linux: Understanding the Security Impact of Rust in the Linux KernelabstractRust-for-Linux (RFL) is a new framework that allows development of Linux kernel extensions in Rust. At first glance, RFL is a huge step forward in terms of improving the security of the kernel: As a safe programming language, Rust can eliminate wide classes of low-level vulnerabilities. Yet, in practice, low-level driver code – complex driver interface, a combination of reference counting and manual memory management, arithmetic pointer and index operations, unsafe type casts, and numerous logical invariants about the data structures exchanged with the kernel might significantly limit the security impact of Rust.This work takes a careful look at how Rust can impact the security of driver code. Specifically, we ask the question: What classes (and what fraction) of vulnerabilities typically found in device driver code can be eliminated by reimplementing device drivers in Rust? We find that Rust can eliminate large classes of safety-related vulnerabilities, but naturally struggles to address protocol violations and semantic errors. Moreover, to be fully eliminated, many classes of flaws require careful programming discipline to avoid memory leaks and runtime panics (e.g., explicit checks for integer overflows and option types), careful implementation of Drop traits, as well as correct implementation of reference counting. Our analysis of 240 driver vulnerabilities that are present in device drivers in the last four years, shows that 82 could be automatically eliminated by Rust, 113 require specific programming idioms and developer’s involvement, and 45 remain unaffected by Rust. We hope that our work can improve the understanding of potential flaws in Rust drivers and result in more secure kernel code. Zhaofeng Li 0004, Vikram Narayanan, Jerry Zhang, Anton Burtsev |
ACSAC | 4 |
| 2024 | Self-Auxiliary Distillation for Sample Efficient Learning in Google-Scale RecommendersabstractIndustrial recommendation systems process billions of daily user feedback which are complex and noisy. Efficiently uncovering user preference from these signals becomes crucial for high-quality recommendation. We argue that those signals are not inherently equal in terms of their informative value and training ability, which is particularly salient in industrial applications with multi-stage processes (e.g., augmentation, retrieval, ranking). Considering that, in this work, we propose a novel self-auxiliary distillation framework that prioritizes training on high-quality labels, and improves the resolution of low-quality labels through distillation by adding a bilateral branch-based auxiliary task. This approach enables flexible learning from diverse labels without additional computational costs, making it highly scalable and effective for Google-scale recommenders. Our framework consistently improved both offline and online key business metrics across three Google major products. Notably, self-auxiliary distillation proves to be highly effective in addressing the severe signal loss challenge posed by changes such as Apple iOS policy. It further delivered significant improvements in both offline (+17% AUC) and online metrics for a Google Apps recommendation system. This highlights the opportunities of addressing real-world signal loss problems through self-auxiliary distillation techniques. Yin Zhang 0011, Xiang Li 0124, Tiansheng Yao, Andrew Evdokimov, Jonathan Valverde, Jerry Zhang, Evan Ettinger, Ed H. Chi, Zhiyuan Cheng 0002 |
RecSys | 8 |
| 2023 | TVA: A multi-party computation system for secure and expressive time series analytics
Muhammad Faisal 0001, Jerry Zhang, John Liagouris, Vasiliki Kalavri, Mayank Varia |
USENIX Security Symposium | 2 |
| 2014 | Immersive and collaborative data visualization using virtual reality platformsabstractEffective data visualization is a key part of the discovery process in the era of “big data”. It is the bridge between the quantitative content of the data and human intuition, and thus an essential component of the scientific path from data into knowledge and understanding. Visualization is also essential in the data mining process, directing the choice of the applicable algorithms, and in helping to identify and remove bad data from the analysis. However, a high complexity or a high dimensionality of modern data sets represents a critical obstacle. How do we visualize interesting structures and patterns that may exist in hyper-dimensional data spaces? A better understanding of how we can perceive and interact with multidimensional information poses some deep questions in the field of cognition technology and human-computer interaction. To this effect, we are exploring the use of immersive virtual reality platforms for scientific data visualization, both as software and inexpensive commodity hardware. These potentially powerful and innovative tools for multi-dimensional data visualization can also provide an easy and natural path to a collaborative data visualization and exploration, where scientists can interact with their data and their colleagues in the same visual space. Immersion provides benefits beyond the traditional “desktop” visualization tools: it leads to a demonstrably better perception of a datascape geometry, more intuitive data understanding, and a better retention of the perceived relationships in the data. Ciro Donalek, S. George Djorgovski, Alex Cioc, Anwell Wang, Jerry Zhang, Elizabeth Lawler, Stacy Yeh, Ashish Mahabal, Matthew J. Graham, Andrew J. Drake, Scott Davidoff, Jeffrey S. Norris, Giuseppe Longo |
IEEE BigData | 5 |
| 2011 | Location-based image retrieval for urban environmentsabstractImage based localization is an important problem with many applications. The basic idea is to match a user generated query image against a database of geo-tagged images with known 6 degrees of freedom poses. Once this retrieval problem is solved, it is possible to recover the pose of the query image. A challenging problem in image retrieval is performance degradation as the size of the image database grows. In this paper we describe an approach to large scale image retrieval for user localization in urban environment by taking advantage of coarse position estimates available, e.g. via cell tower triangulation, on many mobile devices today. The basic idea is to partition the large image database for a large region into a number of overlapping cells each with its own prebuilt search and retrieval structure. We demonstrate retrieval results over a ~12,000 image database covering a 1 km2area of downtown Berkeley. Jerry Zhang, Aaron Hallquist, Eric Liang, Avideh Zakhor |
ICIP | 1 |
| 2000 | HMM adaptation using vector taylor series for noisy speech recognitionabstractIn this paper we address the problem of robustness of speech recognition systems in noisy environments. The goal is to estimate the parameters of a HMM that is matched to a noisy environment, given a HMM trained with clean speech and knowledge of the acoustical environment. We propose a method based on truncated vector Taylor series that approximates the performance of a system trained with that corrupted speech. We also provide insight on the approximations used in the model of the environment and compare them with the lognormal approximation in PMC. 1. Alex Acero, Li Deng 0001, Trausti T. Kristjansson, Jerry Zhang |
INTERSPEECH | 4 |