EDBT 2026 Demo / reviewers in the wild / expert
Manjunath Kudlur
dblp:49/1513
· DBLP profile ↗
20ranked-venue papers
4as first author
0since 2021 · last 2018
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 3 first-authorSoftware engineering, systems software and programming languages · 7 · 2 first-authorArtificial intelligence and machine learning · 2Graphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
8 papers |
Distributed systems · 30% Parallel and multicore computing · 29% Hardware accelerators and domain-specific architectures · 10% | |
| Artificial intelligence
3 papers |
Efficient and distributed learning · 67% Representation and self-supervised learning · 33% | |
| Software engineering, system software, and programming languages
6 papers |
Compilers and program optimization · 46% Operating systems · 27% Program analysis · 27% | |
| Computer graphics and multimedia
1 paper |
Visual content generation and editing · 100% |
Topics — the 28 heaviest of 32, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
distributed training |
0.3 | 1 | 2018 | Dynamic control flow in large-scale machine learning · EuroSys 2018 |
Distributed systems
distributed machine learning |
0.3 | 1 | 2018 | Dynamic control flow in large-scale machine learning · EuroSys 2018 |
Machine learning › Representation and self-supervised learning › visual representation
style representation learning |
0.3 | 1 | 2017 | A Learned Representation For Artistic Style · ICLR (Poster) 2017 |
Visual content generation and editing › style transfer
artistic style transfer |
0.3 | 1 | 2017 | A Learned Representation For Artistic Style · ICLR (Poster) 2017 |
Visual content generation and editing
style transfer |
0.3 | 1 | 2017 | A Learned Representation For Artistic Style · ICLR (Poster) 2017 |
Machine learning › Efficient and distributed learning
large-scale learning |
0.2 | 1 | 2016 | TensorFlow: A System for Large-Scale Machine Learning · OSDI 2016 |
Distributed systems
large-scale machine learning systems |
0.2 | 1 | 2016 | TensorFlow: A System for Large-Scale Machine Learning · OSDI 2016 |
Parallel and multicore computing
parallel programming models |
0.2 | 2 | 2012 | Designing a unified programming model for heterogeneous machines · SC 2012 Orchestrating the execution of stream programs on multicore platforms · PLDI 2008 |
Operating systems › resource management
deadlock avoidance |
0.2 | 2 | 2009 | The theory of deadlock avoidance via discrete control · POPL 2009 Gadara: Dynamic Deadlock Avoidance for Multithreaded Programs · OSDI 2008 |
GPUs and heterogeneous computing
heterogeneous programming models |
0.1 | 1 | 2012 | Designing a unified programming model for heterogeneous machines · SC 2012 |
Parallel and multicore computing › parallel programming models
unified programming model |
0.1 | 1 | 2012 | Designing a unified programming model for heterogeneous machines · SC 2012 |
Compilers and program optimization › vectorization
SIMD vectorization |
0.1 | 1 | 2010 | MacroSS: macro-SIMDization of streaming applications · ASPLOS 2010 |
Parallel and multicore computing
data parallelism |
0.1 | 1 | 2010 | MacroSS: macro-SIMDization of streaming applications · ASPLOS 2010 |
Program analysis › program representation
control flow graph |
0.1 | 1 | 2009 | The theory of deadlock avoidance via discrete control · POPL 2009 |
Reconfigurable computing and FPGAs › FPGA compilation
compiler mapping |
0.1 | 1 | 2009 | Bridging the computation gap between programmable processors and hardwired accelerators · HPCA 2009 |
Hardware accelerators and domain-specific architectures › accelerator architecture
programmable accelerator |
0.1 | 1 | 2009 | Bridging the computation gap between programmable processors and hardwired accelerators · HPCA 2009 |
Compilers and program optimization
code generation |
0.1 | 1 | 2008 | Orchestrating the execution of stream programs on multicore platforms · PLDI 2008 |
Program analysis
dynamic analysis |
0.1 | 1 | 2008 | Gadara: Dynamic Deadlock Avoidance for Multithreaded Programs · OSDI 2008 |
Compilers and program optimization › instruction scheduling
software pipelining |
0.1 | 1 | 2008 | Orchestrating the execution of stream programs on multicore platforms · PLDI 2008 |
Parallel and multicore computing › parallel programming models
stream programming |
0.1 | 1 | 2008 | Orchestrating the execution of stream programs on multicore platforms · PLDI 2008 |
Electronic design automation
high-level synthesis |
0.1 | 1 | 2005 | Cost Sensitive Modulo Scheduling in a Loop Accelerator Synthesis System · MICRO 2005 |
Reconfigurable computing and FPGAs
modulo scheduling |
0.1 | 1 | 2005 | Cost Sensitive Modulo Scheduling in a Loop Accelerator Synthesis System · MICRO 2005 |
Electronic design automation › high-level synthesis
scheduling |
0.1 | 1 | 2005 | Cost Sensitive Modulo Scheduling in a Loop Accelerator Synthesis System · MICRO 2005 |
Processor architecture and microarchitecture › instruction set architecture › instruction set extension
application-specific instruction set extension |
0.0 | 1 | 2004 | Application-Specific Processing on a General-Purpose Core via Transparent Instruction Set Customization · MICRO 2004 |
Processor architecture and microarchitecture
instruction set architecture |
0.0 | 1 | 2004 | Application-Specific Processing on a General-Purpose Core via Transparent Instruction Set Customization · MICRO 2004 |
Processor architecture and microarchitecture › instruction set architecture
instruction set customization |
0.0 | 1 | 2004 | Application-Specific Processing on a General-Purpose Core via Transparent Instruction Set Customization · MICRO 2004 |
Processor architecture and microarchitecture
chip multiprocessor |
0.0 | 1 | 2008 | Orchestrating the execution of stream programs on multicore platforms · PLDI 2008 |
Compilers and program optimization
instruction scheduling |
0.0 | 1 | 2005 | Cost Sensitive Modulo Scheduling in a Loop Accelerator Synthesis System · MICRO 2005 |
Methods — techniques the papers use, named apart from their topics
data flow graphs · 0.6data flow graph · 0.6neural style transfer · 0.6integer linear programming · 0.3macro-SIMDization · 0.2GASNet runtime · 0.1decomposition · 0.1branch-and-bound · 0.1discrete control theory · 0.1compiler mapping · 0.1trace cache · 0.1subgraph identification · 0.1dynamic analysis · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2018 | Dynamic control flow in large-scale machine learningabstractMany recent machine learning models rely on fine-grained dynamic control flow for training and inference. In particular, models based on recurrent neural networks and on reinforcement learning depend on recurrence relations, data-dependent conditional execution, and other features that call for dynamic control flow. These applications benefit from the ability to make rapid control-flow decisions across a set of computing devices in a distributed system. For performance, scalability, and expressiveness, a machine learning system must support dynamic control flow in distributed and heterogeneous environments. Martín Abadi, Paul Barham 0001, Eugene Brevdo, Michael Burrows, Andy Davis, Jeffrey Dean, Sanjay Ghemawat, Tim Harley, Peter Hawkins, Michael Isard, Manjunath Kudlur, Rajat Monga, Derek Gordon Murray, Xiaoqiang Zheng |
EuroSys | 12 |
| 2017 | Exploring the structure of a real-time, arbitrary neural artistic stylization network
Golnaz Ghiasi, Honglak Lee, Manjunath Kudlur, Vincent Dumoulin, Jonathon Shlens |
BMVC | 3 |
| 2017 | A Learned Representation For Artistic Style
Vincent Dumoulin, Jonathon Shlens, Manjunath Kudlur |
ICLR (Poster) | 3 |
| 2016 | TensorFlow: A System for Large-Scale Machine Learning
Martín Abadi, Paul Barham 0001, Jianmin Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek Gordon Murray, Benoit Steiner, Paul A. Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Xiaoqiang Zheng |
OSDI | 11 |
| 2012 | Designing a unified programming model for heterogeneous machinesabstractWhile high-efficiency machines are increasingly embracing heterogeneous architectures and massive multithreading, contemporary mainstream programming languages reflect a mental model in which processing elements are homogeneous, concurrency is limited, and memory is a flat undifferentiated pool of storage. Moreover, the current state of the art in programming heterogeneous machines tends towards using separate programming models, such as OpenMP and CUDA, for different portions of the machine. Both of these factors make programming emerging heterogeneous machines unnecessarily difficult. We describe the design of the Phalanx programming model, which seeks to provide a unified programming model for heterogeneous machines. It provides constructs for bulk parallelism, synchronization, and data placement which operate across the entire machine. Our prototype implementation is able to launch and coordinate work on both CPU and GPU processors within a single node, and by leveraging the GASNet runtime, is able to run across all the nodes of a distributed-memory machine. Michael Garland, Manjunath Kudlur, Yili Zheng |
SC | 2 |
| 2010 | MacroSS: macro-SIMDization of streaming applications
Amir Hormati, Yoonseo Choi, Mark Woh, Manjunath Kudlur, Rodric M. Rabbah, Trevor N. Mudge, Scott A. Mahlke |
ASPLOS | 4 |
| 2009 | Flextream: Adaptive Compilation of Streaming Applications for Heterogeneous ArchitecturesabstractIncreasing demand for performance and efficiency has driven the computer industry toward multicore systems. These systems have become the industry standard in almost all segments of the computer market from high-end servers to handheld devices. In order to efficiently use these systems, an extensive amount of research and industry support has been devoted to developing explicitly parallel programming paradigms, such as streaming models, and new compiler techniques. One important challenge that arises in multicore systems is the ability to dynamically adapt a running application to a target architecture in the face of changes in resource availability (e.g., number of cores, available memory or bandwidth). In this paper, we focus on the increasingly important area of streaming computing and introduce Flextream as a flexible compilation framework that can dynamically adapt applications to the changing characteristics of the underlying architecture. We believe this is an important contribution as software developers grapple with the details of parallelism in a rapidly changing architecture landscape. Flextream achieves its goals through a combination of static compilation and dynamic adaptation techniques. Our results indicate that Flextreampsilas approach can achieve high-performance resource allocations that are within an average of 9% of the optimal solution with low overhead for a wide range of streaming applications. Amir Hormati, Yoonseo Choi, Manjunath Kudlur, Rodric M. Rabbah, Trevor N. Mudge, Scott A. Mahlke |
PACT | 3 |
| 2009 | Bridging the computation gap between programmable processors and hardwired acceleratorsabstractNew media and signal processing applications demand ever higher performance while operating within the tight power constraints of mobile devices. A range of hardware implementations is available to deliver computation with varying degrees of area and power efficiency, from general-purpose processors to application-specific integrated circuits (ASICs). The tradeoff of moving towards more efficient customized solutions such as ASICs is the lack of flexibility in terms of hardware reusability and programmability. In this paper, we propose a customized semi-programmable loop accelerator architecture that exploits the efficiency gains available through high levels of customization, while maintaining sufficient flexibility to execute multiple similar loops. A customized instance of the loop accelerator architecture is generated for a particular loop and then the data and control paths are proactively generalized in an efficient manner to increase flexibility. A compiler mapping phase is then able to map other loops onto the same hardware. The efficiency of the programmable accelerator is compared with non-programmable accelerators and with the OpenRISC 1200 general purpose processor. The programmable accelerator is able to achieve up to 34x better power efficiency and 30x better area efficiency than a simple general purpose processor, while trading off as little as 2x power and area efficiency to the non-programmable accelerator. Kevin Fan, Manjunath Kudlur, Ganesh S. Dasika, Scott A. Mahlke |
HPCA | 2 |
| 2009 | The theory of deadlock avoidance via discrete controlabstractDeadlock in multithreaded programs is an increasingly important problem as ubiquitous multicore architectures force parallelization upon an ever wider range of software. This paper presents a theoretical foundation for dynamic deadlock avoidance in concurrent programs that employ conventional mutual exclusion and synchronization primitives (e.g., multithreaded C/Pthreads programs). Beginning with control flow graphs extracted from program source code, we construct a formal model of the program and then apply Discrete Control Theory to automatically synthesize deadlock-avoidance control logic that is implemented by program instrumentation. At run time, the control logic avoids deadlocks by postponing lock acquisitions. Discrete Control Theory guarantees that the program instrumented with our synthesized control logic cannot deadlock. Our method furthermore guarantees that the control logic is maximally permissive: it postpones lock acquisitions only when necessary to prevent deadlocks, and therefore permits maximal runtime concurrency. Our prototype for C/Pthreads scales to real software including Apache, OpenLDAP, and two kinds of benchmarks, automatically avoiding both injected and naturally occurring deadlocks while imposing modest runtime overheads. Yin Wang 0001, Stéphane Lafortune, Terence Kelly, Manjunath Kudlur, Scott A. Mahlke |
POPL | 4 |
| 2008 | Optimus: efficient realization of streaming applications on FPGAsabstractIn this paper, we introduce Optimus: an optimizing synthesis compiler for streaming applications. Optimus compiles programs written in a high level streaming language to either software or hardware implementations. The compiler uses a hierarchical compilation strategy that separates concerns between macro- and micro-functional requirements. Macro-functional concerns address how components (modules) are assembled to implement larger more complex applications. Micro-functional issues deal with synthesis issues of the module internals. Optimus thus allows software developers who lack deep hardware design expertise to transparently leverage the advantages of hardware customization without crossing the semantic gap between high level languages and hardware description languages. Optimus generates streaming hardware that achieves on average 40x speedup over our baseline embedded processor for a fraction of the energy. Additionally, our results show that streaming-specific optimizations can further improve performance by 255% and reduce the area requirements by 16% in average. These designs are competitive with Handel-C implementations for some of the same benchmarks. Amir Hormati, Manjunath Kudlur, Scott A. Mahlke, David F. Bacon, Rodric M. Rabbah |
CASES | 2 |
| 2008 | Modulo scheduling for highly customized datapaths to increase hardware reusabilityabstractIn the embedded domain, custom hardware in the form of ASICs is often used to implement critical parts of applications when performance and energy efficiency goals cannot be met with software implementations on a general purpose processor or DSP. The downsides of using ASICs include high non-recurring engineering costs, inability to accommodate changes in the application after production, and inability to reuse hardware for new applications. However, by allowing a degree of post-programmability, the hardware can retain high performance and energy efficiency while increasing flexibility and reusability. The difficulty with programmable custom hardware lies in mapping new applications onto an existing datapath that is both sparse and irregular. This paper proposes a constraint-driven modulo scheduler that maps software-pipelineable loops onto programmable loop accelerator hardware. The scheduler is able to target accelerators with widely varying levels of datapath functional capability and connectivity, and thus, varying degrees of programmability. The paper investigates the ability of the scheduler to map new loops onto existing hardware, which depends on both the degree of programmability of the hardware as well as the similarity of the new loop to the original loop for which the hardware was designed. Kevin Fan, Hyunchul Park 0001, Manjunath Kudlur, Scott A. Mahlke |
CGO | 3 |
| 2008 | Gadara: Dynamic Deadlock Avoidance for Multithreaded Programs
Yin Wang 0001, Terence Kelly, Manjunath Kudlur, Stéphane Lafortune, Scott A. Mahlke |
OSDI | 3 |
| 2008 | Orchestrating the execution of stream programs on multicore platformsabstractWhile multicore hardware has become ubiquitous, explicitly parallel programming models and compiler techniques for exploiting parallelism on these systems have noticeably lagged behind. Stream programming is one model that has wide applicability in the multimedia, graphics, and signal processing domains. Streaming models execute as a set of independent actors that explicitly communicate data through channels. This paper presents a compiler technique for planning and orchestrating the execution of streaming applications on multicore platforms. An integrated unfolding and partitioning step based on integer linear programming is presented that unfolds data parallel actors as needed and maximally packs actors onto cores. Next, the actors are assigned to pipeline stages in such a way that all communication is maximally overlapped with computation on the cores. To facilitate experimentation, a generalized code generation template for mapping the software pipeline onto the Cell architecture is presented. For a range of streaming applications, a geometric mean speedup of 14.7x is achieved on a 16-core Cell platform compared to a single core. Manjunath Kudlur, Scott A. Mahlke |
PLDI | 1 |
| 2007 | Hierarchical coarse-grained stream compilation for software defined radioabstractSoftware Defined Radio (SDR) is an emerging embedded domain where the physical layer of wireless protocols is implemented in software rather than the traditional application specific hardware. The operation throughput requirements of current third-generation (3G) wireless protocols are an order of magnitude higher than the capabilities of modern DSP processors. Due to this steep performance requirement, heterogeneous multiprocessor system-on-chip designs have been proposed to support SDR. Given the difficulty in compiling traditional digital signal processors, these new multiprocessor architectures provide even greater challenges for the programmers and compilers. In this paper, we utilize a hierarchical dataflow programming model, referred to as SPIR, that is designed for modeling SDR applications. We then present a coarse-grained data ow compilation strategy that assigns a SDR protocol's DSP kernels onto multiple processors, allocates memory buffers, and determines an execution schedule that meets a prescribed throughput. Unlike traditional approaches, coarse-grained compilation exploits task-level parallelism by scheduling concurrent DSP kernels instead of instructions. Because of the streaming nature of SDR protocols, we adapted an existing instruction-level software pipelining technique, modulo scheduling, for coarse-grained compilation. Our compilation methodology is able to generate parallel code that achieves near linear speedup on a SDR multiprocessor system. Yuan Lin 0002, Manjunath Kudlur, Scott A. Mahlke, Trevor N. Mudge |
CASES | 2 |
| 2006 | Modulo graph embedding: mapping applications onto coarse-grained reconfigurable architecturesabstractCoarse-grained reconfigurable architectures (CGRAs) present an appealing hardware platform by providing the potential for high computation throughput, scalability, low cost and energy efficiency. CGRAs consist of an array of function units and register files generally organized as a two dimensional grid. The most difficult challenge with deploying CGRAs is compiler scheduling technology that can map software implementations of compute intensive loops onto the array. Traditional schedulers are not suitable because they do not take into account the explicit routing of operand values that is necessary. In essence, the problem of binding operations to time slots and resources is extended to also include explicit routing of operands from producers to consumers. To tackle this problem, this paper introduces a software pipelining technique for mapping loop bodies onto CGRAs, referred to as modulo graph embedding. We leverage graph embedding from graph theory, which is used to draw graphs onto a target space. The loop body is essentially drawn onto the CGRA mesh, subject to modulo resource usage constraints. Modulo graph embedding is effective because it can take into account the communication structure of the loop body during mapping. On average, a compute utilization of 56-68% is achieved for a set of loop kernels across three 4x4 CGRA designs. Hyunchul Park 0001, Kevin Fan, Manjunath Kudlur, Scott A. Mahlke |
CASES | 3 |
| 2005 | Cost Sensitive Modulo Scheduling in a Loop Accelerator Synthesis SystemabstractScheduling algorithms used in compilers traditionally focus on goals such as reducing schedule length and register pressure or producing compact code. In the context of a hardware synthesis system where the schedule is used to determine various components of the hardware, including datapath, storage, and interconnect, the goals of a scheduler change drastically. In addition to achieving the traditional goals, the scheduler must proactively make decisions to ensure efficient hardware is produced. This paper proposes two exact solutions for cost sensitive modulo scheduling, one based on an integer linear programming formulation and another based on branch-and-bound search. To achieve reasonable compilation times, decomposition techniques to break down the complex scheduling problem into phase ordered sub-problems are proposed. The decomposition techniques work either by partitioning the dataflow graph into smaller subgraphs and optimally scheduling the subgraphs, or by splitting the scheduling problem into two phases, time slot and resource assignment. The effectiveness of cost sensitive modulo scheduling in minimizing the costs of function units, register structures, and interconnection wires are evaluated within a fully automatic synthesis system for loop accelerators. The cost sensitive modulo scheduler increases the efficiency of the resulting hardware significantly compared to both traditional cost unaware and greedy cost aware modulo schedulers. Kevin Fan, Manjunath Kudlur, Hyunchul Park 0001, Scott A. Mahlke |
MICRO | 2 |
| 2004 | Automatic Synthesis of Customized Local Memories for Multicluster Application Accelerators
Manjunath Kudlur, Kevin Fan, Michael L. Chu, Scott A. Mahlke |
ASAP | 1 |
| 2004 | FLASH: Foresighted Latency-Aware Scheduling Heuristic for Processors with Customized DatapathsabstractApplication-specific instruction set processors (ASIPs) have the potential to meet the challenging cost, performance, and power goals of future embedded processors by customizing the hardware to suit an application. A central problem is creating compilers that are capable of dealing with the heterogeneous and nonuniform hardware created by the customization process. The processor datapath provides an effective area to customize, but specialized datapaths often have nonuniform connectivity between the function units, making the effective latency of a function unit dependent on the consuming operation. Traditional instruction schedulers break down in this environment due to their locally greedy nature of binding the best choice for a single operation even though that choice may be poor due to a lack of communication paths. To effectively schedule with nonuniform connectivity, we propose a foresighted latency-aware scheduling heuristic (FLASH) that performs lookahead across future scheduling steps to estimate the effects of a potential binding. FLASH combines a set of lookahead heuristics to achieve effective foresight with low compile-time overhead. Manjunath Kudlur, Kevin Fan, Michael L. Chu, Rajiv A. Ravindran, Nathan Clark, Scott A. Mahlke |
CGO | 1 |
| 2004 | Application-Specific Processing on a General-Purpose Core via Transparent Instruction Set CustomizationabstractApplication-specific instruction set extensions are an effective way of improving the performance of processors. Critical computation subgraphs can be accelerated by collapsing them into new instructions that are executed on specialized function units. Collapsing the subgraphs simultaneously reduces the length of computation as well as the number of intermediate results stored in the register file. The main problem with this approach is that a new processor must be generated for each application domain. While new instructions can be designed automatically, there is a substantial amount of engineering cost incurred to verify and to implement the final custom processor. In this work, we propose a strategy to transparent customization of the core computation capabilities of the processor without changing its instruction set. A congurable array of function units is added to the baseline processor that enables the acceleration of a wide range of data flow subgraphs. To exploit the array, the microarchitecture performs subgraph identification at run-time, replacing them with new microcode instructions to configure and utilize the array. We compare the effectiveness of replacing subgraphs in the fill unit of a trace cache versus using a translation table during decode, and evaluate the tradeoffs between static and dynamic identification of subgraphs for instruction set customization. Nathan Clark, Manjunath Kudlur, Hyunchul Park 0001, Scott A. Mahlke, Krisztián Flautner |
MICRO | 2 |
| 2004 | Performance analysis of methods that overcome false sharing effects in software DSMs
Manjunath Kudlur, R. Govindarajan |
J. Parallel Distributed Comput. | 1 |