Victor Zhang

dblp:35/9166 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
4since 2021 · last 2026
0009-0000-8077-1099ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 3 since 2021Systems, architecture and hardware · 2Theory of computation · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Theoretical computer science
2 papers
Graph algorithms and graph theory · 67% Algorithms and data structures · 33%
Databases, data mining, and information retrieval
2 papers
Graph data management · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Electronic design automation · 50% Reconfigurable computing and FPGAs · 50%

Topics — the 6 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Graph algorithms and graph theory › graph connectivity
connected components
1.322024
GraphZeppelin: How to Find Connected Components (Even When Graphs Are Dense, Dynamic, and Massive) · ACM Trans. Database Syst. 2024
GraphZeppelin: Storage-Friendly Sketching for Connected Components on Dynamic Graph Streams · SIGMOD Conference 2022
Algorithms and data structures › data streams
streaming algorithms
0.822024
GraphZeppelin: Storage-Friendly Sketching for Connected Components on Dynamic Graph Streams · SIGMOD Conference 2022
GraphZeppelin: How to Find Connected Components (Even When Graphs Are Dense, Dynamic, and Massive) · ACM Trans. Database Syst. 2024
Graph data management › graph processing
streaming graph processing
0.812024
GraphZeppelin: How to Find Connected Components (Even When Graphs Are Dense, Dynamic, and Massive) · ACM Trans. Database Syst. 2024
Graph algorithms and graph theory
graph connectivity
0.812024
GraphZeppelin: How to Find Connected Components (Even When Graphs Are Dense, Dynamic, and Massive) · ACM Trans. Database Syst. 2024
Algorithms and data structures › sketching
linear sketches
0.212024
GraphZeppelin: How to Find Connected Components (Even When Graphs Are Dense, Dynamic, and Massive) · ACM Trans. Database Syst. 2024
Electronic design automation
high-level synthesis
0.112011
LegUp: high-level synthesis for FPGA-based processor/accelerator systems · FPGA 2011

Methods — techniques the papers use, named apart from their topics

streaming algorithms · 1.5linear sketching data structures · 1.5sketching · 1.1c-to-hardware compilation · 0.1
YearPublicationVenuePosition
2026 OpenZL: Using Graphs to Compress Smaller and Faster
abstract
In the last few decades, research techniques have improved lossless compression ratios by significantly increasing processing time. However, these techniques have not gained popularity in industry because production systems require high throughput and low resource utilization. Instead, real world improvements in compression are increasingly realized by building application-specific compressors which can exploit knowledge about the structure and semantics of the data being compressed. Application-specific compressor systems outperform even the best generic compressors, but these techniques have severe drawbacks -- they are inherently limited in applicability, are hard to develop, and are difficult to maintain and deploy. In this work, we show that these challenges can be overcome with a new compression strategy. We propose the "graph model" of compression, a new theoretical framework for representing compression as a directed acyclic graph of modular codecs. OpenZL implements this framework and compresses data into a self-describing wire format, any configuration of which can be decompressed by a universal decoder. OpenZL's design enables rapid development of application-specific compressors with minimal code. Experimental results demonstrate that OpenZL achieves superior compression ratios and speeds compared to state-of-the-art general-purpose compressors on a variety of real-world datasets. Compared to ratio-focused deep-learning compressors, OpenZL is competitive on ratio while being many orders of magnitude faster. Internal deployments at Meta have also shown consistent improvements in size and/or speed, with development timelines reduced from months to days. OpenZL thus represents a significant advance in practical, scalable, and maintainable data compression for modern data-intensive applications.
Yann Collet, Nick Terrell, W. Felix P. Handte, Danielle Rozenblit, Victor Zhang, Yaelle Goldschlag, Jennifer Lee, Elliot Gorokhovsky, Yonatan Komornik, Daniel Riegel, Stan Angelov, Nadav Rotem
ICDE5
2025 Exploring the Landscape of Distributed Graph Sketching
abstract
Recent work has initiated the study of dense graph processing using graph sketching methods, which drastically reduce space costs by lossily compressing information about the input graph. In this paper, we explore the strange and surprising performance landscape of sketching algorithms. We highlight both their surprising advantages for processing dense graphs that were previously prohibitively expensive to study, as well as the current limitations of the technique. Most notably, we show how sketching can avoid bottlenecks that limit conventional graph processing methods.
David Tench, Evan West, Kenny Zhang, Michael A. Bender, Daniel DeLayo, Martin Farach-Colton, Gilvir Gill, Tyler Seip, Victor Zhang
ALENEX9
2024 GraphZeppelin: How to Find Connected Components (Even When Graphs Are Dense, Dynamic, and Massive)
abstract
Finding the connected components of a graph is a fundamental problem with uses throughout computer science and engineering. The task of computing connected components becomes more difficult when graphs are very large, or when they are dynamic, meaning the edge set changes over time subject to a stream of edge insertions and deletions. A natural approach to computing the connected components problem on a large, dynamic graph stream is to buy enough RAM to store the entire graph. However, the requirement that the graph fit in RAM is an inherent limitation of this approach and is prohibitive for very large graphs. Thus, there is an unmet need for systems that can process dense dynamic graphs, especially when those graphs are larger than available RAM. We present a new high-performance streaming graph-processing system for computing the connected components of a graph. This system, which we call GraphZeppelin , uses new linear sketching data structures ( CubeSketch ) to solve the streaming connected components problem and as a result requires space asymptotically smaller than the space required for a lossless representation of the graph. GraphZeppelin is optimized for massive dense graphs: GraphZeppelin can process millions of edge updates (both insertions and deletions) per second, even when the underlying graph is far too large to fit in available RAM. As a result GraphZeppelin vastly increases the scale of graphs that can be processed.
David Tench, Evan West, Victor Zhang, Michael A. Bender, Abiyaz Chowdhury, Daniel DeLayo, J. Ahmed Dellas, Martin Farach-Colton, Tyler Seip, Kenny Zhang
ACM Trans. Database Syst.3
2022 GraphZeppelin: Storage-Friendly Sketching for Connected Components on Dynamic Graph Streams
abstract
Finding the connected components of a graph is a fundamental problem with uses throughout computer science and engineering. The task of computing connected components becomes more difficult when graphs are very large, or when they are dynamic, meaning the edge set changes over time subject to a stream of edge insertions and deletions. A natural approach to computing the connected components on a large, dynamic graph stream is to buy enough RAM to store the entire graph. However, the requirement that the graph fit in RAM is prohibitive for very large graphs. Thus, there is an unmet need for systems that can process dense dynamic graphs, especially when those graphs are larger than available RAM.
David Tench, Evan West, Victor Zhang, Michael A. Bender, Abiyaz Chowdhury, J. Ahmed Dellas, Martin Farach-Colton, Tyler Seip, Kenny Zhang
SIGMOD Conference3
2013 LegUp: An open-source high-level synthesis tool for FPGA-based processor/accelerator systems
abstract
It is generally accepted that a custom hardware implementation of a set of computations will provide superior speed and energy efficiency relative to a software implementation. However, the cost and difficulty of hardware design is often prohibitive, and consequently, a software approach is used for most applications. In this article, we introduce a new high-level synthesis tool called LegUp that allows software techniques to be used for hardware design. LegUp accepts a standard C program as input and automatically compiles the program to a hybrid architecture containing an FPGA-based MIPS soft processor and custom hardware accelerators that communicate through a standard bus interface. In the hybrid processor/accelerator architecture, program segments that are unsuitable for hardware implementation can execute in software on the processor. LegUp can synthesize most of the C language to hardware, including fixed-sized multidimensional arrays, structs, global variables, and pointer arithmetic. Results show that the tool produces hardware solutions of comparable quality to a commercial high-level synthesis tool. We also give results demonstrating the ability of the tool to explore the hardware/software codesign space by varying the amount of a program that runs in software versus hardware. LegUp, along with a set of benchmark C programs, is open source and freely downloadable, providing a powerful platform that can be leveraged for new research on a wide range of high-level synthesis topics.
Andrew Canis, Jongsok Choi, Mark Aldham, Victor Zhang, Ahmed Kammoona, Tomasz S. Czajkowski, Stephen Brown 0003, Jason Helge Anderson
ACM Trans. Embed. Comput. Syst.4
2011 LegUp: high-level synthesis for FPGA-based processor/accelerator systems
abstract
In this paper, we introduce a new open source high-level synthesis tool called LegUp that allows software techniques to be used for hardware design. LegUp accepts a standard C program as input and automatically compiles the program to a hybrid architecture containing an FPGA-based MIPS soft processor and custom hardware accelerators that communicate through a standard bus interface. Results show that the tool produces hardware solutions of comparable quality to a commercial high-level synthesis tool.
Andrew Canis, Jongsok Choi, Mark Aldham, Victor Zhang, Ahmed Kammoona, Jason Helge Anderson, Stephen Brown 0003, Tomasz S. Czajkowski
FPGA4