VLDB 2026 Research / reviewers in the wild / expert
Aleksandar Zlateski
dblp:163/2100
· DBLP profile ↗
9ranked-venue papers
6as first author
0since 2021 · last 2019
0000-0003-0261-1967ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 5 first-authorArtificial intelligence and machine learning · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Hardware accelerators and domain-specific architectures · 72% Parallel and multicore computing · 28% | |
| Artificial intelligence
2 papers |
Segmentation and scene understanding · 30% Trustworthy machine learning · 30% 3D vision · 20% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 100% |
Topics — the 10 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning › Data-centric AI
label quality |
0.3 | 1 | 2018 | On the Importance of Label Quality for Semantic Segmentation · CVPR 2018 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.3 | 1 | 2018 | On the Importance of Label Quality for Semantic Segmentation · CVPR 2018 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
convolution optimization |
0.3 | 1 | 2018 | Optimizing N-dimensional, winograd-based convolution for manycore CPUs · PPoPP 2018 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › convolution optimization
winograd convolution |
0.3 | 1 | 2018 | Optimizing N-dimensional, winograd-based convolution for manycore CPUs · PPoPP 2018 |
Parallel and multicore computing › parallel algorithms
parallel algorithm design |
0.3 | 1 | 2017 | A Multicore Path to Connectomics-on-Demand · PPoPP 2017 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
CNN accelerator |
0.2 | 1 | 2016 | ZNNi: maximizing the inference throughput of 3D convolutional networks on CPUs and GPUs · SC 2016 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › DNN inference
CNN inference |
0.2 | 1 | 2016 | ZNNi: maximizing the inference throughput of 3D convolutional networks on CPUs and GPUs · SC 2016 |
Machine learning › Deep learning architectures and training
convolutional neural network |
0.2 | 1 | 2015 | Recursive Training of 2D-3D Convolutional Networks for Neuronal Boundary Prediction · NIPS 2015 |
Computer vision › 3D vision
volumetric image analysis |
0.2 | 1 | 2015 | Recursive Training of 2D-3D Convolutional Networks for Neuronal Boundary Prediction · NIPS 2015 |
Bioinformatics and computational biology › computational neuroscience
connectomics |
0.2 | 2 | 2017 | A Multicore Path to Connectomics-on-Demand · PPoPP 2017 Recursive Training of 2D-3D Convolutional Networks for Neuronal Boundary Prediction · NIPS 2015 |
Methods — techniques the papers use, named apart from their topics
recursive training · 0.4multicore CPU parallelism · 0.43d convolution · 0.4winograd algorithm · 0.3synthetic data generation · 0.3n-dimensional convolution · 0.3fine-tuning · 0.3convolutional network · 0.3padded and pruned FFTs · 0.2CPU-GPU co-execution · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | The anatomy of efficient FFT and winograd convolutions on modern CPUsabstractWinograd-based convolution has quickly gained traction as a preferred approach to implement convolutional neural networks (ConvNet) on various hardware platforms because it could require fewer floating point operations than FFT-based or direct convolutions. Aleksandar Zlateski, Zhen Jia 0001, Kai Li 0001, Frédo Durand |
ICS | 1 |
| 2018 | On the Importance of Label Quality for Semantic SegmentationabstractConvolutional networks (ConvNets) have become the dominant approach to semantic image segmentation. Producing accurate, pixel-level labels required for this task is a tedious and time consuming process; however, producing approximate, coarse labels could take only a fraction of the time and effort. We investigate the relationship between the quality of labels and the performance of ConvNets for semantic segmentation. We create a very large synthetic dataset with perfectly labeled street view scenes. From these perfect labels, we synthetically coarsen labels with different qualities and estimate human-hours required for producing them. We perform a series of experiments by training ConvNets with a varying number of training images and label quality. We found that the performance of ConvNets mostly depends on the time spent creating the training labels. That is, a larger coarsely-annotated dataset can yield the same performance as a smaller finely-annotated one. Furthermore, fine-tuning coarsely pre-trained ConvNets with few finely-annotated labels can yield comparable or superior performance to training it with a large amount of finely-annotated labels alone, at a fraction of the labeling cost. We demonstrate that our result is also valid for different network architectures, and various object classes in an urban scene. Aleksandar Zlateski, Ronnachai Jaroensri, Prafull Sharma, Frédo Durand |
CVPR | 1 |
| 2018 | Optimizing N-dimensional, winograd-based convolution for manycore CPUsabstractRecent work on Winograd-based convolution allows for a great reduction of computational complexity, but existing implementations are limited to 2D data and a single kernel size of 3 by 3. They can achieve only slightly better, and often worse performance than better optimized, direct convolution implementations. We propose and implement an algorithm for N-dimensional Winograd-based convolution that allows arbitrary kernel sizes and is optimized for manycore CPUs. Our algorithm achieves high hardware utilization through a series of optimizations. Our experiments show that on modern ConvNets, our optimized implementation, is on average more than 3 x, and sometimes 8 x faster than other state-of-the-art CPU implementations on an Intel Xeon Phi manycore processors. Moreover, our implementation on the Xeon Phi achieves competitive performance for 2D ConvNets and superior performance for 3D ConvNets, compared with the best GPU implementations. Zhen Jia 0001, Aleksandar Zlateski, Frédo Durand, Kai Li 0001 |
PPoPP | 2 |
| 2017 | Compile-time optimized and statically scheduled N-D convnet primitives for multi-core and many-core (Xeon Phi) CPUsabstractConvolutional networks (ConvNets), largely running on GPUs, have become the most popular approach to computer vision. Now that CPUs are closing the FLOPS gap with GPUs, efficient CPU algorithms are becoming more important. We propose a novel parallel and vectorized algorithm for N-D convolutional layers. Our goal is to achieve high utilization of available FLOPS, independent of ConvNet architecture and CPU properties (e.g. vector units, number of cores, cache sizes). Our approach is to rely on the compiler to optimize code, thereby removing the need for hand-tuning. We assume that the network architecture is known at compile-time. Our serial algorithm divides the computation into small sub-tasks designed to be easily optimized by the compiler for a specific CPU. Sub-tasks are executed in an order that maximizes cache reuse. We parallelize the algorithm by statically scheduling tasks to be executed by each core. Our novel compile-time recursive scheduling algorithm is capable of dividing the computation evenly between an arbitrary number of cores, regardless of ConvNet architecture. It introduces zero runtime overhead and minimal synchronization overhead. We demonstrate that our serial primitives efficiently utilize available FLOPS (75--95%), while our parallel algorithm attains 50--90% utilization on 64+ core machines. Our algorithm is competitive with the fastest CPU implementation to date (MKL2017) for 2D object recognition, and performs much better for image segmentation. For 3D ConvNets we demonstrate comparable performance to the latest GPU hardware and software even though the CPU is only capable of half the FLOPS of the GPU. Aleksandar Zlateski, H. Sebastian Seung |
ICS | 1 |
| 2017 | A Multicore Path to Connectomics-on-DemandabstractThe current design trend in large scale machine learning is to use distributed clusters of CPUs and GPUs with MapReduce-style programming. Some have been led to believe that this type of horizontal scaling can reduce or even eliminate the need for traditional algorithm development, careful parallelization, and performance engineering. This paper is a case study showing the contrary: that the benefits of algorithms, parallelization, and performance engineering, can sometimes be so vast that it is possible to solve "cluster-scale" problems on a single commodity multicore machine. Alexander Matveev, Yaron Meirovitch, Hayk Saribekyan, Wiktor Jakubiuk, Tim Kaler, Gergely Ódor, David M. Budden, Aleksandar Zlateski, Nir Shavit |
PPoPP | 8 |
| 2017 | Scalable training of 3D convolutional networks on multi- and many-cores
Aleksandar Zlateski, Kisuk Lee, H. Sebastian Seung |
J. Parallel Distributed Comput. | 1 |
| 2016 | ZNN - A Fast and Scalable Algorithm for Training 3D Convolutional Networks on Multi-core and Many-Core Shared Memory MachinesabstractConvolutional networks (ConvNets) have become a popular approach to computer vision. It is important to accelerate ConvNet training, which is computationally costly. We propose a novel parallel algorithm based on decomposition into a set of tasks, most of which are convolutions or FFTs. Applying Brent's theorem to the task dependency graph implies that linear speedup with the number of processors is attainable within the PRAM model of parallel computation, for wide network architectures. To attain such performance on real shared-memory machines, our algorithm computes convolutions converging on the same node of the network with temporal locality to reduce cache misses, and sums the convergent convolution outputs via an almost wait-free concurrent method to reduce time spent in critical sections. We implement the algorithm with a publicly available software package called ZNN. Benchmarking with multi-core CPUs shows that ZNN can attain speedup roughly equal to the number of physical cores. We also show that ZNN can attain over 90× speedup on a many-core CPU (Xeon Phi™ Knights Corner). These speedups are achieved for network architectures with widths that are in common use. The task parallelism of the ZNN algorithm is suited to CPUs, while the SIMD parallelism of previous algorithms is compatible with GPUs. Through examples, we show that ZNN can be either faster or slower than certain GPU implementations depending on specifics of the network architecture, kernel sizes, and density and size of the output patch. ZNN may be less costly to develop and maintain, due to the relative ease of general-purpose CPU programming. Aleksandar Zlateski, Kisuk Lee, H. Sebastian Seung |
IPDPS | 1 |
| 2016 | ZNNi: maximizing the inference throughput of 3D convolutional networks on CPUs and GPUsabstractSliding window convolutional networks (ConvNets) have become a popular approach to computer vision problems such as image segmentation and object detection and localization. Here we consider the parallelization of inference, i.e., the application of a previously trained ConvNet, with emphasis on 3D images. Our goal is to maximize throughput, defined as the number of output voxels computed per unit time. We propose CPU and GPU primitives for convolutional and pooling layers, which are combined to create CPU, GPU, and CPU-GPU inference algorithms. The primitives include convolution based on highly efficient padded and pruned FFTs. Our theoretical analyses and empirical tests reveal a number of interesting findings. For example, adding host RAM can be a more efficient way of increasing throughput than adding another GPU or more CPUs. Furthermore, our CPU-GPU algorithm can achieve greater throughput than the sum of CPU-only and GPU-only throughputs. Aleksandar Zlateski, Kisuk Lee, H. Sebastian Seung |
SC | 1 |
| 2015 | Recursive Training of 2D-3D Convolutional Networks for Neuronal Boundary PredictionabstractEfforts to automate the reconstruction of neural circuits from 3D electron microscopic (EM) brain images are critical for the field of connectomics. An important computation for reconstruction is the detection of neuronal boundaries. Images acquired by serial section EM, a leading 3D EM technique, are highly anisotropic, with inferior quality along the third dimension. For such images, the 2D max-pooling convolutional network has set the standard for performance at boundary detection. Here we achieve a substantial gain in accuracy through three innovations. Following the trend towards deeper networks for object recognition, we use a much deeper network than previously employed for boundary detection. Second, we incorporate 3D as well as 2D filters, to enable computations that use 3D context. Finally, we adopt a recursively trained architecture in which a first network generates a preliminary boundary map that is provided as input along with the original image to a second network that generates a final boundary map. Backpropagation training is accelerated by ZNN, a new implementation of 3D convolutional networks that uses multicore CPU parallelism for speed. Our hybrid 2D-3D architecture could be more generally applicable to other types of anisotropic 3D images, including video, and our recursive framework for any image labeling problem. Kisuk Lee, Aleksandar Zlateski, Ashwin Vishwanathan, H. Sebastian Seung |
NIPS | 2 |