VLDB 2026 Research / reviewers in the wild / expert
Igor Durdanovic
dblp:81/108
· DBLP profile ↗
7ranked-venue papers
0as first author
0since 2021 · last 2017
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4Systems, architecture and hardware · 2Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Efficient and distributed learning · 95% Optimization for machine learning · 5% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Hardware accelerators and domain-specific architectures · 68% Reconfigurable computing and FPGAs · 20% Parallel and multicore computing · 12% |
Topics — the 6 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning › model compression › pruning › structured pruning
channel pruning |
0.3 | 1 | 2017 | Pruning Filters for Efficient ConvNets · ICLR (Poster) 2017 |
Machine learning › Efficient and distributed learning
model compression |
0.3 | 1 | 2017 | Pruning Filters for Efficient ConvNets · ICLR (Poster) 2017 |
Machine learning › Efficient and distributed learning › model compression › pruning
structured pruning |
0.3 | 1 | 2017 | Pruning Filters for Efficient ConvNets · ICLR (Poster) 2017 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.1 | 1 | 2008 | A Massively Parallel Digital Learning Processor · NIPS 2008 |
Reconfigurable computing and FPGAs › reconfigurable computing
FPGA-based machine learning accelerators |
0.0 | 1 | 2008 | A Massively Parallel Digital Learning Processor · NIPS 2008 |
Parallel and multicore computing › parallel computing
parallel machine learning |
0.0 | 1 | 2004 | Parallel Support Vector Machines: The Cascade SVM · NIPS 2004 |
Methods — techniques the papers use, named apart from their topics
filter pruning · 0.3convolutional neural network · 0.3kernel matrix decomposition · 0.1cascade SVM · 0.1multiply-accumulate · 0.1SIMD · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2017 | Pruning Filters for Efficient ConvNets
Hao Li 0022, Asim Kadav, Igor Durdanovic, Hanan Samet, Hans Peter Graf |
ICLR (Poster) | 3 |
| 2009 | A Massively Parallel Coprocessor for Convolutional Neural NetworksabstractWe present a massively parallel coprocessor for accelerating Convolutional Neural Networks (CNNs), a class of important machine learning algorithms. The coprocessor functional units, consisting of parallel 2D convolution primitives and programmable units performing sub-sampling and non-linear functions specific to CNNs, implement a ldquometa-operatorrdquo to which a CNN may be compiled to. The coprocessor is serviced by distributed off-chip memory banks with large data bandwidth. As a key feature, we use low precision data and further increase the effective memory bandwidth by packing multiple words in every memory operation, and leverage the algorithmpsilas simple data access patterns to use off-chip memory as a scratchpad for intermediate data, critical for CNNs. A CNN is mapped to the coprocessor hardware primitives with instructions to transfer data between the memory and coprocessor. We have implemented a prototype of the CNN coprocessor on an off-the-shelf PCI FPGA card with a single Xilinx Virtex5 LX330T FPGA and 4 DDR2 memory banks totaling 1 GB. The coprocessor prototype can process at the rate of 3.4 billion multiply accumulates per second (GMACs) for CNN forward propagation, a speed that is 31x faster than a software implementation on a 2.2 GHz AMD Opteron processor. For a complete face recognition application with the CNN on the coprocessor and the rest of the image processing tasks on the host, the prototype is 6-10times faster, depending on the host-coprocessor bandwidth. Murugan Sankaradass, Venkata Jakkula, Srihari Cadambi, Srimat T. Chakradhar, Igor Durdanovic, Eric Cosatto, Hans Peter Graf |
ASAP | 5 |
| 2009 | A Massively Parallel FPGA-Based Coprocessor for Support Vector MachinesabstractWe present a massively parallel FPGA-based coprocessor for Support Vector Machines (SVMs), a machine learning algorithm whose applications include recognition tasks such as learning scenes, situations and concepts, and reasoning tasks such as analyzing the recognized scenes and semantics. The coprocessor architecture, targeted at both SVM training and classification, is based on clusters of vector processing elements (VPEs) operating in single-instruction multiple data (SIMD) mode to take advantage of large amounts of data parallelism in the application. We use the FPGA's DSP elements as parallel multiply-accumulators (MACs), a core computation in SVMs. A key feature of the architecture is that it is customized to low precision arithmetic which permits one DSP unit to perform two or more MACs in parallel. Low precision also reduces the required number of parallel off-chip memory accesses by packing multiple data words on the FPGA-memory bus. We have built a prototype using an off-the-shelf PCI-based FPGA card with a Xilinx Virtex 5 FPGA and 1 GB DDR2 memory. For SVM training, we observe application-level end-to-end computation speeds of over 9 billion multiply-accumulates per second (GMACs). For SVM classification, using data packing, the application speed increases to 14 GMACs. The FPGA-based system is about 20times faster than a dual Opteron 2.2 GHz processor CPU, and dissipates around 10 W of power. Srihari Cadambi, Igor Durdanovic, Venkata Jakkula, Murugan Sankaradass, Eric Cosatto, Srimat T. Chakradhar, Hans Peter Graf |
FCCM | 2 |
| 2008 | A Massively Parallel Digital Learning ProcessorabstractWe present a new, massively parallel architecture for accelerating machine learning algorithms, based on arrays of variable-resolution arithmetic vector processing elements (VPE). Groups of VPEs operate in SIMD (single instruction multiple data) mode, and each group is connected to an independent memory bank. In this way memory bandwidth scales with the number of VPE, and the main data flows are local, keeping power dissipation low. With 256 VPEs, implemented on two FPGA (field programmable gate array) chips, we obtain a sustained speed of 19 GMACS (billion multiply-accumulate per sec.) for SVM training, and 86 GMACS for SVM classification. This performance is more than an order of magnitude higher than that of any FPGA implementation reported so far. The speed on one FPGA is similar to the fastest speeds published on a Graphics Processor for the MNIST problem, despite a clock rate of the FPGA that is six times lower. High performance at low clock rates makes this massively parallel architecture particularly attractive for embedded applications, where low power dissipation is critical. Tests with Convolutional Neural Networks and other learning algorithms are under way now. Hans Peter Graf, Srihari Cadambi, Igor Durdanovic, Venkata Jakkula, Murugan Sankaradass, Eric Cosatto, Srimat T. Chakradhar |
NIPS | 3 |
| 2004 | Parallel Support Vector Machines: The Cascade SVMabstractWe describe an algorithm for support vector machines (SVM) that can be parallelized efficiently and scales to very large problems with hundreds of thousands of training vectors. Instead of analyzing the whole training set in one optimization step, the data are split into subsets and optimized separately with multiple SVMs. The partial results are combined and filtered again in a ‘Cascade’ of SVMs, until the global optimum is reached. The Cascade SVM can be spread over multiple processors with minimal communication overhead and requires far less memory, since the kernel matrices are much smaller than for a regular SVM. Convergence to the global optimum is guaranteed with multiple passes through the Cascade, but already a single pass provides good generalization. A single pass is 5x – 10x faster than a regular SVM for problems of 100,000 vectors when implemented on a single processor. Parallel implementations on a cluster of 16 processors were tested with over 1 million vectors (2-class problems), converging in a day or two, while a regular SVM never converged in over a week. Hans Peter Graf, Eric Cosatto, Léon Bottou, Igor Durdanovic, Vladimir Vapnik |
NIPS | 4 |
| 2001 | Evolution of Cooperative Problem Solving in an Artificial EconomyabstractWe address the problem of how to reinforce learning in ultracomplex environments, with huge state-spaces, where one must learn to exploit a compact structure of the problem domain. The approach we propose is to simulate the evolution of an artificial economy of computer programs. The economy is constructed based on two simple principles so as to assign credit to the individual programs for collaborating on problem solutions. We find empirically that starting from programs that are random computer code, we can develop systems that solve hard problems. In particular, our economy learned to solve almost all random Blocks World problems with goal stacks that are 200 blocks high. Competing methods solve such problems only up to goal stacks of at most 8 blocks. Our economy has also learned to unscramble about half a randomly scrambled Rubik's cube and to solve several commercially sold puzzles. Eric B. Baum, Igor Durdanovic |
Neural Comput. | 2 |
| 1996 | Correctness of Compiling Occam to Transputer CodeabstractThis paper contributes to the development of a rigorous mathematical framework for the study of provably correct compilation techniques. The proposed method is developed through an implementation of a real-life non-toy imperative programming language with non-determinism and parallelism, i.e. Occam, to a commercial machine, namely the Transputer. We provide a mathematical definition of the Transputer Instruction Set architecture for executing Occam together with a correctness proof for a general compilation schema of Occam programs into Transputer code. We start from the ground model, an abstract processor, running a high and a low priority queue of Occam processes, which formalizes the semantics of Occam at the abstraction level of atomic Occam instructions. We develop here increasingly more refined levels of Transputer semantics, proving correctness (and when possible also completeness) for each refinement step. Along the way we collect our proof assumptions, a set of natural conditions for a compiler to be correct, thus making our proof applicable to a large class of compilers. As a by-product our construction provides a challenging realistic case study for proof verification by theorem provers. Egon Börger, Igor Durdanovic |
Comput. J. | 2 |