Dong Zhou 0006

dblp:15/2101-6 · DBLP profile ↗
← Back
15ranked-venue papers
4as first author
3since 2021 · last 2023
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 1 first-author · 1 since 2021Computer networks · 4 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 4 · 2 since 2021Databases, data management, data science and information retrieval · 1Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2023 Mercury: Fast Transaction Broadcast in High Performance Blockchain Systems
abstract
Blockchain systems must be secure and offer high performance. These systems rely on transaction broadcast mechanisms to provide both of these features. Unfortunately, in today’s systems, the broadcast mechanisms are highly inefficient.We present Mercury, a new transaction broadcast protocol designed for high performance blockchains. Mercury shortens the transaction propagation delay using two techniques: a virtual coordinate system and an early outburst strategy. Simulation results show that Mercury outperforms prior propagation schemes and decreases overall propagation latency by up to 44%. When implemented in Conflux, an open-source high-throughput blockchain system, Mercury reduces transaction propagation latency by over 50% with less than 5% bandwidth overhead.
Mingxun Zhou, Liyi Zeng, Peilun Li, Fan Long, Dong Zhou 0006, Ivan Beschastnikh, Ming Wu 0007
INFOCOM6
2022 Utilizing Parallelism in Smart Contracts on Decentralized Blockchains by Taming Application-Inherent Conflicts
abstract
Traditional public blockchain systems typically had very limited transaction throughput because of the bottleneck of the consensus protocol itself. With recent advances in consensus technology, the performance limit has been greatly lifted, typically to thousands of transactions per second. With this, transaction execution has become a new performance bottleneck. Exploiting parallelism in transaction execution is a clear and direct way to address this and to further increase transaction throughput. Although some recent literature introduced concurrency control mechanisms to execute smart contract transactions in parallel, the reported speedup that they can achieve is far from ideal. The main reason is that the proposed parallel execution mechanisms cannot effectively deal with the conflicts inherent in many blockchain applications.
Péter Garamvölgyi, Dong Zhou 0006, Fan Long, Ming Wu 0007
ICSE3
2021 PipeZK: Accelerating Zero-Knowledge Proof with a Pipelined Architecture
abstract
Zero-knowledge proof (ZKP) is a promising cryptographic protocol for both computation integrity and privacy. It can be used in many privacy-preserving applications including verifiable cloud outsourcing and blockchains. The major obstacle of using ZKP in practice is its time-consuming step for proof generation, which consists of large-size polynomial computations and multi-scalar multiplications on elliptic curves. To efficiently and practically support ZKP in real-world applications, we propose PipeZK, a pipelined accelerator with two subsystems to handle the aforementioned two intensive compute tasks, respectively. The first subsystem uses a novel dataflow to decompose large kernels into smaller ones that execute on bandwidth-efficient hardware modules, with optimized off-chip memory accesses and on-chip compute resources. The second subsystem adopts a lightweight dynamic work dispatch mechanism to share the heavy processing units, with minimized resource underutilization and load imbalance. When evaluated in 28 nm, PipeZK can achieve 10x speedup on standard cryptographic benchmarks, and 5x on a widely-used cryptocurrency application, Zcash.
Ye Zhang 0042, Shuo Wang 0009, Xian Zhang 0001, Jiangbin Dong, Xingzhong Mao, Fan Long, Dong Zhou 0006, Mingyu Gao 0001, Guangyu Sun 0003
ISCA8
2020 Shrec: bandwidth-efficient transaction relay in high-throughput blockchain systems
abstract
The success of Bitcoin and Ethereum has attracted many efforts to build high-throughput blockchain systems. This paper focuses on transaction dissemination --- a rather overlooked issue in these systems. We argue that efficient transaction dissemination is the key for a blockchain system to sustain at high-throughput --- usually thousands of transactions per second --- and the existing solutions fell short at doing so.
Chenxing Li, Peilun Li, Ming Wu 0007, Dong Zhou 0006, Fan Long
SoCC5
2020 A Decentralized Blockchain with High Throughput and Fast Confirmation
Chenxing Li, Peilun Li, Dong Zhou 0006, Ming Wu 0007, Guang Yang 0020, Wei Xu 0005, Fan Long, Andrew Chi-Chih Yao
USENIX ATC3
2020 Fast Software Cache Design for Network Appliances
Dong Zhou 0006, Huacheng Yu, Michael Kaminsky, David G. Andersen
USENIX ATC1
2015 Raising the Bar for Using GPUs in Software Packet Processing
Anuj Kalia, Dong Zhou 0006, Michael Kaminsky, David G. Andersen
NSDI2
2015 Scaling Up Clustered Network Appliances with ScaleBricks
abstract
This paper presents ScaleBricks, a new design for building scalable, clustered network appliances that must "pin" flow state to a specific handling node without being able to choose which node that should be. ScaleBricks applies a new, compact lookup structure to route packets directly to the appropriate handling node, without incurring the cost of multiple hops across the internal interconnect. Its lookup structure is many times smaller than the alternative approach of fully replicating a forwarding table onto all nodes. As a result, ScaleBricks is able to improve throughput and latency while simultaneously increasing the total number of flows that can be handled by such a cluster. This architecture is effective in practice: Used to optimize packet forwarding in an existing commercial LTE-to-Internet gateway, it increases the throughput of a four-node cluster by 23%, reduces latency by up to 10%, saves memory, and stores up to 5.7x more entries in the forwarding table.
Dong Zhou 0006, Hyeontaek Lim, David G. Andersen, Michael Kaminsky, Michael Mitzenmacher, Ren Wang 0001, Ajaypal Singh
SIGCOMM1
2014 Rex: replication at the speed of multi-core
abstract
Standard state-machine replication involves consensus on a sequence of totally ordered requests through, for example, the Paxos protocol. Such a sequential execution model is becoming outdated on prevalent multi-core servers. Highly concurrent executions on multi-core architectures introduce non-determinism related to thread scheduling and lock contentions, and fundamentally break the assumption in state-machine replication. This tension between concurrency and consistency is not inherent because the total-ordering of requests is merely a simplifying convenience that is unnecessary for consistency. Concurrent executions of the application can be decoupled with a sequence of consensus decisions through consensus on partial-order traces, rather than on totally ordered requests, that capture the non-deterministic decisions in one replica execution and to be replayed with the same decisions on others. The result is a new multi-core friendly replicated state-machine framework that achieves strong consistency while preserving parallelism in multi-thread applications. On 12-core machines with hyper-threading, evaluations on typical applications show that we can scale with the number of cores, achieving up to 16 times the throughput of standard replicated state machines.
Chuntao Hong, Mao Yang 0004, Dong Zhou 0006, Lidong Zhou, Li Zhuang
EuroSys4
2013 Scalable, high performance ethernet forwarding with CuckooSwitch
abstract
Several emerging network trends and new architectural ideas are placing increasing demand on forwarding table sizes. From massive-scale datacenter networks running millions of virtual machines to flow-based software-defined networking, many intriguing design options require FIBs that can scale well beyond the thousands or tens of thousands possible using today's commodity switching chips.
Dong Zhou 0006, Hyeontaek Lim, Michael Kaminsky, David G. Andersen
CoNEXT1
2013 When Cycles Are Cheap, Some Tables Can Be Huge
Dong Zhou 0006, Hyeontaek Lim, Michael Kaminsky, David G. Andersen
HotOS2
2013 KuaFu: Closing the parallelism gap in database replication
abstract
Database systems are nowadays increasingly deployed on multi-core commodity servers, with replication to guard against failures. Database engine is best designed to scale with the number of cores to offer a high degree of parallelism on a modern multi-core architecture. On the other hand, replication traditionally resorts to a certain form of serialization for data consistency among replicas. In the widely used primary/backup replication with log shipping, concurrent executions on the primary and the serialized log replay on a backup creates a serious parallelism gap. Our experiment on MySQL with a 16-core configuration shows that the serial replay of a backup can sustain only less than one third of the throughput achievable on the primary under an OLTP workload. This paper proposes KuaFu to close the parallelism gap on replicated database systems by enabling concurrent replay of transactions on a backup. KuaFu maintains write consistency on backups by tracking transaction dependencies. Concurrent replay on a backup does introduce read inconsistency between the primary and backups. KuaFu further leverages multi-version concurrency control to produce snapshots in order to restore the consistency semantics. We have implemented KuaFu on MySQL; our evaluations show that KuaFu allows a backup to keep up with the primary while preserving replication consistency.
Chuntao Hong, Dong Zhou 0006, Mao Yang 0004, Carbo Kuo, Lidong Zhou
ICDE2
2013 Space-Efficient, High-Performance Rank and Select Structures on Uncompressed Bit Sequences
Dong Zhou 0006, David G. Andersen, Michael Kaminsky
SEA1
2011 Software fault isolation with API integrity and multi-principal modules
abstract
The security of many applications relies on the kernel being secure, but history suggests that kernel vulnerabilities are routinely discovered and exploited. In particular, exploitable vulnerabilities in kernel modules are common. This paper proposes LXFI, a system which isolates kernel modules from the core kernel so that vulnerabilities in kernel modules cannot lead to a privilege escalation attack. To safely give kernel modules access to complex kernel APIs, LXFI introduces the notion of API integrity, which captures the set of contracts assumed by an interface. To partition the privileges within a shared module, LXFI introduces module principals. Programmers specify principals and API integrity rules through capabilities and annotations. Using a compiler plugin, LXFI instruments the generated code to grant, check, and transfer capabilities between modules, according to the programmer's annotations. An evaluation with Linux shows that the annotations required on kernel functions to support a new module are moderate, and that LXFI is able to prevent three known privilege-escalation vulnerabilities. Stress tests of a network driver module also show that isolating this module using LXFI does not hurt TCP throughput but reduces UDP throughput by 35%, and increases CPU utilization by 2.2-3.7x.
Yandong Mao, Haogang Chen 0001, Dong Zhou 0006, Xi Wang 0005, Nickolai Zeldovich, M. Frans Kaashoek
SOSP3
2011 G2: A Graph Processing System for Diagnosing Distributed Systems
Dong Zhou 0006, Haoxiang Lin, Mao Yang 0004, Fan Long, Chaoqiang Deng, Changshu Liu, Lidong Zhou
USENIX ATC2