EDBT 2026 Demo / reviewers in the wild / expert
Chen Zou 0001
dblp:205/0395-1
· DBLP profile ↗
8ranked-venue papers
5as first author
4since 2021 · last 2024
0000-0003-0120-4032ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 4 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | VarVE: Bringing SIMD Performance to Variable-Width ValuesabstractProcessor datapaths grew from 4 to 512 bits via Single-Instruction-Multiple-Data (SIMD) parallelism. SIMD applies the same operation to multiple values, which increases performance and reduces the instruction count. However, these evolvements do not provide support for variable-width values, so programmers ‘pad’ values to align to outliers, wasting the upper bits with zeros in both registers and datapath. We propose VarVE, a vector instruction set extension built upon the state-of-the-art vector-length agnostic SIMD instruction set: ARM SVE. VarVE provides native support for variable-width values within a vector, avoiding padding waste, thus making better use of the SIMD datapath. VarVE's design enables a flexible strip mining model with a variety of optimizations. Evaluation of VarVE shows 60x speedup over ARM in kernels with element packing and unpacking, and 1.3x - 5.4x speedup over SVE for pure-compute filtering in TPC-H benchmarks. VarVE also achieves 2x speedup on a neural network inference task. All these results exemplify VarVE's general ability to improve datapath and memory system efficiency. Chen Zou 0001, Andrew A. Chien |
ICCD | 1 |
| 2022 | ASSASIN: Architecture Support for Stream Computing to Accelerate Computational StorageabstractComputational storage adds computing to storage devices, providing potential benefits in offload, data-reduction, and lower energy. Successful computational SSD architectures should match growing flash bandwidth, which in turn requires high SSD DRAM memory bandwidth. This creates a memory wall scaling problem, resulting from SSDs’ stringent power and cost constraints.A survey of recent computational SSD research shows that many computational storage offloads are suited to stream computing. To exploit this opportunity, we propose a novel general-purpose computational SSD and core architecture, called ASSASIN (Architecture Support for Stream computing to Accelerate computatIoNal Storage). ASSASIN provides a unified set of compute engines between SSD DRAM and the flash array. This eliminates the SSD DRAM bottleneck by enabling direct computing on flash data streams. ASSASIN further employs a crossbar to achieve performance even when flash data layout is uneven and preserve independence for page layout decisions in the flash translation layer. With stream buffers and scratchpad memories, ASSASIN core’s memory hierarchy and instruction set extensions provide superior low-latency access at low-power and effectively keep streaming flash data out of the in-SSD cache-DRAM memory hierarchy, thereby solving the memory wall.Evaluation shows that ASSASIN delivers 1.5x - 2.4x speedup for offloaded functions compared to state-of-the-art computational SSD architectures. Further, ASSASIN’s streaming approach yields 2.0x power efficiency and 3.2x area efficiency improvement. And these performance benefits at the level of computational SSDs translate to 1.1x - 1.5x end-to-end speedups on data analytics workloads. Chen Zou 0001, Andrew A. Chien |
MICRO | 1 |
| 2021 | Computational Storage to Increase the Analysis Capability of Tier-2 HEP Data SitesabstractLarge Hadron Collider (LHC) produces collision data at 100 PB/year which needs to be stored and analyzed for high energy physics (HEP) theories. We reconsider the design choices of HEP data centers and evaluate different upgrade options to improve their analysis capacity.Results show that computational storage to be the cost-effective and power-efficient upgrade option. Computational disks in the storage cluster deliver a 9.3-fold speedup for Higgs Boson analysis. This exceeds the speedup from all other upgrades considered (faster network: 100 to 1000 Gbps, upgrade from HDDs to SSDs). Chen Zou 0001, Andrew A. Chien, Robert W. Gardner, Ilija Vukotic |
CLUSTER | 1 |
| 2021 | PSACS: Highly-Parallel Shuffle Accelerator on Computational StorageabstractShuffle is an indispensable process in distributed online analytical processing systems to enable task-level parallelism exploitation via multiple nodes. As a data-intensive data reorganization process, shuffle implemented on general-purpose CPUs not only incurs data traffic back and forth between the computing and storage resources, but also pollutes the cache hierarchy with almost zero data reuse. As a result, shuffle can easily become the bottleneck of distributed analysis pipelines.Our PSACS approach attacks these bottlenecks with the rising computational storage paradigm. Shuffle is offloaded to the storage-side PSACS accelerator to avoid polluting computing node memory hierarchy and enjoy the latency, bandwidth and energy benefits of near-data computing. Further, the microarchitecture of PSACS exploits data-, subtask-, and task-level parallelism for high performance and a customized scratchpad for fast on-chip random access.PSACS achieves 4.6x—5.7x shuffle throughput at kernel-level and up to 1.3x overall shuffle throughput with only a twentieth of CPU utilization comparing to software baselines. These mount up to 23% end-to-end OLAP query speedup on average. Chen Zou 0001, Hui Zhang 0033, Andrew A. Chien, Yang-Seok Ki |
ICCD | 1 |
| 2020 | A Novel Heuristic Search Method for Two-Level Approximate Logic SynthesisabstractRecently, much attention has been paid to approximate computing, a novel design paradigm for error-tolerant applications. It can significantly reduce area, power, and delay of circuits by introducing an acceptable amount of error. In this paper, we propose a new heuristic method for two-level approximate logic synthesis. The problem is to identify an approximate sum-of-product (SOP) expression under a given error rate (ER) constraint so that it has the fewest literals. The basic idea of our method is to find an optimal set of input combinations for 0-to-1 output complement (SICC). For this purpose, we first identify all prime SICCs, which are fundamental SICCs in the sense that the optimal SICC is very likely to be a union of a subset of the prime SICCs. Then, we search among all subsets of the prime SICCs the optimal subset, which leads to a final good approximate SOP. We further propose four speed-up techniques. The experiments on benchmarks showed that our method is better than the previous state-of-the-art method and our speed-up techniques are effective. For an ER threshold of 0.8%, our method can reduce 15.8% literals on average. Sanbao Su, Chen Zou 0001, Weijiang Kong, Jie Han 0001, Weikang Qian |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2019 | Accelerating Raw Data Analysis with the ACCORDA Software and Hardware ArchitectureabstractThe data science revolution and growing popularity of data lakes make efficient processing of raw data increasingly important. To address this, we propose the ACCelerated Operators for Raw Data Analysis (ACCORDA) architecture. By extending the operator interface (subtype with encoding) and employing a uniform runtime worker model, ACCORDA integrates data transformation acceleration seamlessly, enabling a new class of encoding optimizations and robust high-performance raw data processing. Together, these key features preserve the software system architecture, empowering state-of-art heuristic optimizations to drive flexible data encoding for performance. ACCORDA derives performance from its software architecture, but depends critically on the acceleration of the Unstructured Data Processor (UDP) that is integrated into the memory-hierarchy, and accelerates data transformation tasks by 16x-21x (parsing, decompression) to as much as 160x (deserialization) compared to an x86 core. We evaluate ACCORDA using TPC-H queries on tabular data formats, exercising raw data properties such as parsing and data conversion. The ACCORDA system achieves 2.9x-13.2x speedups when compared to SparkSQL, reducing raw data processing overhead to a geomean of 1.2x (20%). In doing so, ACCORDA robustly matches or outperforms prior systems that depend on caching loaded data, while computing on raw, unloaded data. This performance benefit is robust across format complexity, query predicates, and selectivity (data statistics). ACCORDA's encoding-extended operator interface unlocks aggressive encoding-oriented optimizations that deliver 80% average performance increase over the 7 affected TPC-H queries. Yuanwei Fang, Chen Zou 0001, Andrew A. Chien |
Proc. VLDB Endow. | 2 |
| 2017 | UDP: a programmable accelerator for extract-transform-load workloads and moreabstractBig data analytic applications give rise to large-scale extract-transform-load (ETL) as a fundamental step to transform new data into a native representation. ETL workloads pose significant performance challenges on conventional architectures, so we propose the design of the unstructured data processor (UDP), a software programmable accelerator that includes multi-way dispatch, variable-size symbol support, Flexible-source dispatch (stream buffer and scalar registers), and memory addressing to accelerate ETL kernels both for current and novel future encoding and compression. Specifically, UDP excels at branch-intensive and symbol and pattern-oriented workloads, and can offload them from CPUs. Yuanwei Fang, Chen Zou 0001, Aaron J. Elmore, Andrew A. Chien |
MICRO | 2 |
| 2017 | Motion artifact removal based on periodical property for ECG monitoring with wearable systems
Chen Zou 0001, Yajie Qin, Chenglu Sun, Wei Li 0134, Wei Chen 0015 |
Pervasive Mob. Comput. | 1 |