Jiyuan Zhang 0002

dblp:93/7807-2 · DBLP profile ↗
← Back
2ranked-venue papers
2as first author
0since 2021 · last 2020
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Hardware accelerators and domain-specific architectures · 61% Parallel and multicore computing · 30% Embedded and real-time systems · 9%
Databases, data mining, and information retrieval
1 paper
Query processing and optimization · 100%
Theoretical computer science
1 paper
Computational complexity · 100%

Topics — the 4 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization › query execution › relational operators
set intersection
0.412020
FESIA: A Fast and SIMD-Efficient Set Intersection Approach on Modern CPUs · ICDE 2020
Computational complexity › communication complexity › two-party communication
set intersection
0.412020
FESIA: A Fast and SIMD-Efficient Set Intersection Approach on Modern CPUs · ICDE 2020
Hardware accelerators and domain-specific architectures › machine learning accelerator
direct convolution
0.312018
High Performance Zero-Memory Overhead Direct Convolutions · ICML 2018
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.312018
High Performance Zero-Memory Overhead Direct Convolutions · ICML 2018

Methods — techniques the papers use, named apart from their topics

bitmap filtering · 0.9SIMD · 0.9loop optimization · 0.3direct convolution · 0.3
YearPublicationVenuePosition
2020 FESIA: A Fast and SIMD-Efficient Set Intersection Approach on Modern CPUs
abstract
Set intersection is an important operation and widely used in both database and graph analytics applications. However, existing state-of-the-art set intersection methods only consider the size of input sets and fail to optimize for the case in which the intersection size is small. In real-world scenarios, the size of most intersections is usually orders of magnitude smaller than the size of the input sets, e.g., keyword search in databases and common neighbor search in graph analytics. In this paper, we present FESIA, a new set intersection approach on modern CPUs. The time complexity of our approach is O(n/√w + r), in which w is the SIMD width, and n and r are the size of input sets and intersection size, respectively. The key idea behind FESIA is that it first uses bitmaps to filter out unmatched elements from the input sets, and then selects suitable specialized kernels (i.e., small function blocks) at runtime to compute the final intersection on each pair of bitmap segments. In addition, all data structures in FESIA are designed to take advantage of SIMD instructions provided by vector ISAs with various SIMD widths, including SSE, AVX, and the latest AVX512. Our experiments on both real-world and synthetic datasets show that our intersection method achieves more than an order of magnitude better performance than conventional scalar implementations, and up to 4x better performance than state-of-the-art SIMD implementations.
Jiyuan Zhang 0002, Yi Lu 0010, Daniele G. Spampinato, Franz Franchetti
ICDE1
2018 High Performance Zero-Memory Overhead Direct Convolutions
abstract
The computation of convolution layers in deep neural networks typically rely on high performance routines that trade space for time by using additional memory (either for packing purposes or required as part of the algorithm) to improve performance. The problems with such an approach are two-fold. First, these routines incur additional memory overhead which reduces the overall size of the network that can fit on embedded devices with limited memory capacity. Second, these high performance routines were not optimized for performing convolution, which means that the performance obtained is usually less than conventionally expected. In this paper, we demonstrate that direct convolution, when implemented correctly, eliminates all memory overhead, and yields performance that is between 10% to 400% times better than existing high performance implementations of convolution layers on conventional and embedded CPU architectures. We also show that a high performance direct convolution exhibits better scaling performance, i.e. suffers less performance drop, when increasing the number of threads.
Jiyuan Zhang 0002, Franz Franchetti, Tze Meng Low
ICML1