EDBT 2026 Demo / reviewers in the wild / expert
Oscar Palomar
dblp:59/11209
· DBLP profile ↗
29ranked-venue papers
2as first author
2since 2021 · last 2023
0000-0001-6729-4187ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 27 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
9 papers |
Processor architecture and microarchitecture · 49% Hardware accelerators and domain-specific architectures · 12% Performance modeling and evaluation · 9% | |
| Artificial intelligence
1 paper |
Robot navigation and mapping · 100% | |
| Theoretical computer science
1 paper |
Algorithms and data structures · 100% |
Topics — the 30 heaviest of 36, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Processor architecture and microarchitecture
instruction set architecture |
0.9 | 2 | 2023 | Vitruvius+: An Area-Efficient RISC-V Decoupled Vector Coprocessor for High Performance Computing Applications · ACM Trans. Archit. Code Optim. 2023 Future Vector Microprocessor Extensions for Data Aggregations · ISCA 2016 |
Processor architecture and microarchitecture
vector processor |
0.7 | 2 | 2020 | A RISC-V Simulator and Benchmark Suite for Designing and Evaluating Vector Architectures · ACM Trans. Archit. Code Optim. 2020 An Integrated Vector-Scalar Design on an In-Order ARM Core · ACM Trans. Archit. Code Optim. 2017 |
Processor architecture and microarchitecture › instruction set architecture
RISC-V |
0.7 | 1 | 2023 | Vitruvius+: An Area-Efficient RISC-V Decoupled Vector Coprocessor for High Performance Computing Applications · ACM Trans. Archit. Code Optim. 2023 |
Processor architecture and microarchitecture › instruction set architecture
vector extension |
0.7 | 1 | 2023 | Vitruvius+: An Area-Efficient RISC-V Decoupled Vector Coprocessor for High Performance Computing Applications · ACM Trans. Archit. Code Optim. 2023 |
Performance modeling and evaluation › simulation
architectural simulation |
0.4 | 1 | 2020 | A RISC-V Simulator and Benchmark Suite for Designing and Evaluating Vector Architectures · ACM Trans. Archit. Code Optim. 2020 |
Performance modeling and evaluation
benchmarking |
0.4 | 1 | 2020 | A RISC-V Simulator and Benchmark Suite for Designing and Evaluating Vector Architectures · ACM Trans. Archit. Code Optim. 2020 |
Processor architecture and microarchitecture › instruction set architecture › vector extension
RISC-V vector extension |
0.4 | 1 | 2020 | A RISC-V Simulator and Benchmark Suite for Designing and Evaluating Vector Architectures · ACM Trans. Archit. Code Optim. 2020 |
Robotics › Robot navigation and mapping › SLAM
real-time SLAM |
0.3 | 1 | 2018 | Navigating the Landscape for Real-Time Localization and Mapping for Robotics and Virtual and Augmented Reality · Proc. IEEE 2018 |
Robotics › Robot navigation and mapping
SLAM |
0.3 | 1 | 2018 | Navigating the Landscape for Real-Time Localization and Mapping for Robotics and Virtual and Augmented Reality · Proc. IEEE 2018 |
Electronic design automation
design space exploration |
0.3 | 1 | 2018 | Navigating the Landscape for Real-Time Localization and Mapping for Robotics and Virtual and Augmented Reality · Proc. IEEE 2018 |
Energy-efficient computing › low-power design
low-power processor design |
0.3 | 1 | 2017 | An Integrated Vector-Scalar Design on an In-Order ARM Core · ACM Trans. Archit. Code Optim. 2017 |
Distributed systems
data aggregation |
0.2 | 1 | 2016 | Future Vector Microprocessor Extensions for Data Aggregations · ISCA 2016 |
Parallel and multicore computing
data parallelism |
0.2 | 1 | 2016 | Future Vector Microprocessor Extensions for Data Aggregations · ISCA 2016 |
Processor architecture and microarchitecture › SIMD
SIMD extensions |
0.2 | 1 | 2016 | Future Vector Microprocessor Extensions for Data Aggregations · ISCA 2016 |
High-performance computing › code optimization
vectorization |
0.2 | 1 | 2016 | Future Vector Microprocessor Extensions for Data Aggregations · ISCA 2016 |
Processor architecture and microarchitecture
SIMD |
0.2 | 1 | 2015 | VSR sort: A novel vectorised sorting algorithm & architecture extensions for future microprocessors · HPCA 2015 |
Algorithms and data structures › sequence algorithms › sorting › integer sorting
radix sort |
0.2 | 1 | 2015 | VSR sort: A novel vectorised sorting algorithm & architecture extensions for future microprocessors · HPCA 2015 |
Algorithms and data structures › sequence algorithms
sorting |
0.2 | 1 | 2015 | VSR sort: A novel vectorised sorting algorithm & architecture extensions for future microprocessors · HPCA 2015 |
Processor architecture and microarchitecture › instruction set architecture
instruction set extension |
0.2 | 2 | 2015 | Vector Extensions for Decision Support DBMS Acceleration · MICRO 2012 VSR sort: A novel vectorised sorting algorithm & architecture extensions for future microprocessors · HPCA 2015 |
Processor architecture and microarchitecture
out-of-order execution |
0.2 | 1 | 2023 | Vitruvius+: An Area-Efficient RISC-V Decoupled Vector Coprocessor for High Performance Computing Applications · ACM Trans. Archit. Code Optim. 2023 |
Processor architecture and microarchitecture › out-of-order execution
register renaming |
0.2 | 1 | 2023 | Vitruvius+: An Area-Efficient RISC-V Decoupled Vector Coprocessor for High Performance Computing Applications · ACM Trans. Archit. Code Optim. 2023 |
Memory systems › memory access patterns
irregular memory access |
0.2 | 1 | 2014 | APMC: advanced pattern based memory controller (abstract only) · FPGA 2014 |
Memory systems
memory access patterns |
0.2 | 1 | 2014 | APMC: advanced pattern based memory controller (abstract only) · FPGA 2014 |
Memory systems
memory controller |
0.2 | 1 | 2014 | APMC: advanced pattern based memory controller (abstract only) · FPGA 2014 |
Energy-efficient computing
power and energy modeling |
0.2 | 1 | 2014 | Power estimation tool for system on programmable chip based platforms (abstract only) · FPGA 2014 |
Electronic design automation
power estimation |
0.2 | 1 | 2014 | Power estimation tool for system on programmable chip based platforms (abstract only) · FPGA 2014 |
Hardware accelerators and domain-specific architectures
database accelerator |
0.1 | 1 | 2012 | Vector Extensions for Decision Support DBMS Acceleration · MICRO 2012 |
Processor architecture and microarchitecture › SIMD
SIMD instructions |
0.1 | 1 | 2015 | VSR sort: A novel vectorised sorting algorithm & architecture extensions for future microprocessors · HPCA 2015 |
Reconfigurable computing and FPGAs
FPGA-based memory controller |
0.1 | 1 | 2014 | APMC: advanced pattern based memory controller (abstract only) · FPGA 2014 |
Embedded and real-time systems › embedded hardware platform
multicore embedded systems |
0.1 | 1 | 2014 | Power estimation tool for system on programmable chip based platforms (abstract only) · FPGA 2014 |
Methods — techniques the papers use, named apart from their topics
machine learning · 1.0heterogeneous acceleration · 1.0JIT compilation · 1.0simulation · 0.4parameterizable vector model · 0.4gem5 simulation · 0.4chaining logic · 0.3block-based vector execution · 0.3simulation framework · 0.2descriptor-based prefetching · 0.2memory simulation · 0.1cycle-accurate microarchitectural simulation · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | eProcessor: European, Extendable, Energy-Efficient, Extreme-Scale, Extensible, Processor EcosystemabstractThe eProcessor project aims at creating a RISC-V full stack ecosystem. The eProcessor architecture combines a high-performance out-of-order core with energy-efficient accelerators for vector processing and artificial intelligence with reduced-precision functional units. The design of this architecture follows a hardware/software co-design approach with relevant application use cases from the high-performance computing, bioinformatics and artificial intelligence domains. Two eProcessor prototypes will be developed based on two fabricated eProcessor ASICs integrated into a computer-on-module. Lluc Alvarez, Abraham Ruiz, Arnau Bigas-Soldevilla, Pavel Kuroedov, Alberto González 0004, Hamsika Mahale, Noe Bustamante, Albert Aguilera, Francesco Minervini, Javier Salamero, Oscar Palomar, Vassilis Papaefstathiou, Antonis Psathakis, Nikolaos Dimou, Michalis Giaourtas, Iasonas Mastorakis, Giorgos Ieronymakis, Georgios-Michail Matzouranis, Vassilis Flouris, Nikolaos Kossifidis, Manolis Marazakis, Bhavishya Goel, Madhavan Manivannan, Ahsen Ejaz, Panagiotis Strikos, Mateo Vázquez, Ioannis Sourdis, Pedro Trancoso, Per Stenström, Jens Hagemeyer, Lennart Tigges, Nils Kucza, Jean-Marc Philippe, Ioannis Papaefstathiou |
CF | 11 |
| 2023 | Vitruvius+: An Area-Efficient RISC-V Decoupled Vector Coprocessor for High Performance Computing ApplicationsabstractThe maturity level of RISC-V and the availability of domain-specific instruction set extensions, like vector processing, make RISC-V a good candidate for supporting the integration of specialized hardware in processor cores for the High Performance Computing (HPC) application domain. In this article, 1 we present Vitruvius+, the vector processing acceleration engine that represents the core of vector instruction execution in the HPC challenge that comes within the EuroHPC initiative. It implements the RISC-V vector extension (RVV) 0.7.1 and can be easily connected to a scalar core using the Open Vector Interface standard. Vitruvius+ natively supports long vectors: 256 double precision floating-point elements in a single vector register. It is composed of a set of identical vector pipelines (lanes), each containing a slice of the Vector Register File and functional units (one integer, one floating point). The vector instruction execution scheme is hybrid in-order/out-of-order and is supported by register renaming and arithmetic/memory instruction decoupling. On a stand-alone synthesis, Vitruvius+ reaches a maximum frequency of 1.4 GHz in typical conditions (TT/0.80V/25°C) using GlobalFoundries 22FDX FD-SOI. The silicon implementation has a total area of 1.3 mm 2 and maximum estimated power of ∼920 mW for one instance of Vitruvius+ equipped with eight vector lanes. Francesco Minervini, Oscar Palomar, Osman S. Unsal, Enrico Reggiani, Josue V. Quiroga, Joan Marimon, Carlos Rojas 0001, Roger Figueras, Abraham Ruiz, Alberto González 0004, Jonnatan Mendoza, Iván Vargas 0001, César Hernández, Joan Cabre, Lina Khoirunisya, Mustapha Bouhali, Julian Pavon, Francesc Moll, Mauro Olivieri, Mario Kovac, Mate Kovac, Leon Dragic, Mateo Valero, Adrián Cristal |
ACM Trans. Archit. Code Optim. | 2 |
| 2020 | A RISC-V Simulator and Benchmark Suite for Designing and Evaluating Vector ArchitecturesabstractVector architectures lack tools for research. Consider the gem5 simulator, which is possibly the leading platform for computer-system architecture research. Unfortunately, gem5 does not have an available distribution that includes a flexible and customizable vector architecture model. In consequence, researchers have to develop their own simulation platform to test their ideas, which consume much research time. However, once the base simulator platform is developed, another question is the following: Which applications should be tested to perform the experiments? The lack of Vectorized Benchmark Suites is another limitation. To face these problems, this work presents a set of tools for designing and evaluating vector architectures. First, the gem5 simulator was extended to support the execution of RISC-V Vector instructions by adding a parameterizable Vector Architecture model for designers to evaluate different approaches according to the target they pursue. Second, a novel Vectorized Benchmark Suite is presented: a collection composed of seven data-parallel applications from different domains that can be classified according to the modules that are stressed in the vector architecture. Finally, a study of the Vectorized Benchmark Suite executing on the gem5-based Vector Architecture model is highlighted. This suite is the first in its category that covers the different possible usage scenarios that may occur within different vector architecture designs such as embedded systems, mainly focused on short vectors, or High-Performance-Computing (HPC), usually designed for large vectors. Cristóbal Ramírez, César-Alejandro Hernández-Calderón, Oscar Palomar, Osman S. Unsal, Marco A. Ramírez 0001, Adrián Cristal |
ACM Trans. Archit. Code Optim. | 3 |
| 2019 | SimAcc: A Configurable Cycle-Accurate Simulator for Customized Accelerators on CPU-FPGAs SoCsabstractThis paper describes a flexible infrastructure for fast computer architecture simulation and prototyping of accelerator IP. A trend for System-on-Chips is to include application specific accelerators on the die. However, there is still a key research problem that needs to be addressed: How do hardware accelerators interact with the processors of a system and what is the impact on overall performance? To solve this problem, we propose an infrastructure that can directly simulate unmodified application executables with FPGA hardware accelerators. Unmodified application binaries are dynamically instrumented to generate processor load/store and program counter events and any memory accesses generated by accelerators, that are sent to an FPGA-based out-of-order pipeline model. The key features of our infrastructure are the ability to code exclusively at the user level, to dynamically discover and use available hardware models at run time, to test and simultaneously optimize hardware accelerators in an heterogeneous system. In terms of evaluation, we present a comparison between our system and Gem5 to demonstrate accuracy and relative performance, using the SPEC CPU benchmarks; even though our system is implemented on Zynq XC7045 which integrates dual 667MHz Arm Cortex-A9s with substantial FPGA resources, it outperforms Gem5 running on a Xeon E3 3.2 GHz with 32GBs of RAM. We also evaluate our infrastructure in simulating the interaction of accelerators with processors using accelerators taken from the Mach Benchmark Suite and other custom accelerators from computer vision applications. Konstantinos Iordanou, Oscar Palomar, John Mawer, Cosmin Gorgovan, Andy Nisbet, Mikel Luján |
FCCM | 2 |
| 2018 | Navigating the Landscape for Real-Time Localization and Mapping for Robotics and Virtual and Augmented RealityabstractVisual understanding of 3-D environments in real time, at low power, is a huge computational challenge. Often referred to as simultaneous localization and mapping (SLAM), it is central to applications spanning domestic and industrial robotics, autonomous vehicles, and virtual and augmented reality. This paper describes the results of a major research effort to assemble the algorithms, architectures, tools, and systems software needed to enable delivery of SLAM, by supporting applications specialists in selecting and configuring the appropriate algorithm and the appropriate hardware, and compilation pathway, to meet their performance, accuracy, and energy consumption goals. The major contributions we present are: 1) tools and methodology for systematic quantitative evaluation of SLAM algorithms; 2) automated, machine-learning-guided exploration of the algorithmic and implementation design space with respect to multiple objectives; 3) end-to-end simulation tools to enable optimization of heterogeneous, accelerated architectures for the specific algorithmic requirements of the various SLAM algorithmic approaches; and 4) tools for delivering, where appropriate, accelerated, adaptive SLAM solutions in a managed, JIT-compiled, adaptive runtime context. Sajad Saeedi G., Bruno Bodin, Harry Wagstaff, Andy Nisbet, Luigi Nardi, John Mawer, Nicolas Melot, Oscar Palomar, Emanuele Vespa, Tom Spink, Cosmin Gorgovan, Andrew M. Webb 0002, James Clarkson, Erik Tomusk, Thomas Debrunner, Kuba Kaszyk, Pablo González de Aledo Marugán, Andrey Rodchenko, Graham D. Riley, Christos Kotselidis, Björn Franke, Michael F. P. O'Boyle, Andrew J. Davison, Paul H. J. Kelly, Mikel Luján, Steve Furber |
Proc. IEEE | 8 |
| 2018 | Vector Processing-Aware Advanced Clock-Gating Techniques for Low-Power Fused Multiply-AddabstractThe need for power efficiency is driving a rethink of design decisions in processor architectures. While vector processors succeeded in the high-performance market in the past, they need a retailoring for the mobile market that they are entering now. Floating-point (FP) fused multiply-add (FMA), being a functional unit with high power consumption, deserves special attention. Although clock gating is a well-known method to reduce switching power in synchronous designs, there are unexplored opportunities for its application to vector processors, especially when considering active operating mode. In this research, we comprehensively identify, propose, and evaluate the most suitable clock-gating techniques for vector FMA units (VFUs). These techniques ensure power savings without jeopardizing the timing. We evaluate the proposed techniques using both synthetic and “real-world” application-based benchmarking. Using vector masking and vector multilane-aware clock gating, we report power reductions of up to 52%, assuming active VFU operating at the peak performance. Among other findings, we observe that vector instruction-based clock-gating techniques achieve power savings for all vector FP instructions. Finally, when evaluating all techniques together, using “real-world” benchmarking, the power reductions are up to 80%. Additionally, in accordance with processor design trends, we perform this research in a fully parameterizable and automated fashion. Ivan Ratkovic, Oscar Palomar, Milan Stanic, Osman S. Unsal, Adrián Cristal, Mateo Valero |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2017 | The Potential of Dynamic Binary Modification and CPU-FPGA SoCs for SimulationabstractIn this paper we describe a flexible infrastructure that can directly interface unmodified application executables with FPGA hardware acceleration IP in order to 1), facilitate faster computer architecture simulation, and 2), to prototype microarchitecture or accelerator IP. Dynamic binary modification tool plugins are directly interfaced to the application under evaluation via flexible software interfaces provided by a userspace hardware control library that also manages access to a parameterised Bluespec IP library. We demonstrate the potential of our infrastructure with two use cases with unmodified application executables where, 1), an executable is dynamically instrumented to generate load/store and program counter events that are sent to FPGA hardware accelerated in-order microarchitecture pipeline, and memory hierarchy models, and 2), the design of a branch predictor is prototyped using an FPGA. The key features of our infrastructure are the ability to instrument at instruction level granularity, to code exclusively at the user level, and to dynamically discover and use available hardware models at run time, thus, we enable software developers to rapidly investigate and evaluate parameterised Bluespec microarchitecture and accelerator IP models. We present a comparison between our system and GEM5, the industry standard ARM architecture simulator, to demonstrate accuracy and relative performance, even though our system is implemented on an Xilinx Zynq 7000 FPGA board with tightly coupled FPGA and ARM Cortex A9 processors, it outperforms GEM5 running on a Xeon with 32GBs of RAM (400x vs 700x slowdown over native execution). John Mawer, Oscar Palomar, Cosmin Gorgovan, Andy Nisbet, William B. Toms, Mikel Luján |
FCCM | 2 |
| 2017 | An Integrated Vector-Scalar Design on an In-Order ARM CoreabstractIn the low-end mobile processor market, power, energy, and area budgets are significantly lower than in the server/desktop/laptop/high-end mobile markets. It has been shown that vector processors are a highly energy-efficient way to increase performance; however, adding support for them incurs area and power overheads that would not be acceptable for low-end mobile processors. In this work, we propose an integrated vector-scalar design for the ARM architecture that mostly reuses scalar hardware to support the execution of vector instructions. The key element of the design is our proposed block-based model of execution that groups vector computational instructions together to execute them in a coordinated manner. We implemented a classic vector unit and compare its results against our integrated design. Our integrated design improves the performance (more than 6×) and energy consumption (up to 5×) of a scalar in-order core with negligible area overhead (only 4.7% when using a vector register with 32 elements). In contrast, the area overhead of the classic vector unit can be significant (around 44%) if a dedicated vector floating-point unit is incorporated. Our block-based vector execution outperforms the classic vector unit for all kernels with floating-point data and also consumes less energy. We also complement the integrated design with three energy/performance-efficient techniques that further reduce power and increase performance. The first proposal covers the design and implementation of chaining logic that is optimized to work with the cache hierarchy through vector memory instructions, the second proposal reduces the number of reads/writes from/to the vector register file, and the third idea optimizes complex memory access patterns with the memory shape instruction and unified indexed vector load. Milan Stanic, Oscar Palomar, Timothy Hayes 0001, Ivan Ratkovic, Adrián Cristal, Osman S. Unsal, Mateo Valero |
ACM Trans. Archit. Code Optim. | 2 |
| 2016 | POSTER: An Integrated Vector-Scalar Design on an In-order ARM CoreabstractIn the low-end mobile processor market, power, energy and area budgets are significantly lower than in other markets (e.g. servers or high-end mobile markets). It has been shown that vector processors are a highly energy-efficient way to increase performance; however adding support for them incurs area and power overheads that would not be acceptable for low-end mobile processors. In this work, we propose an integrated vector-scalar design for the ARM architecture that mostly reuses scalar hardware to support the execution of vector instructions. The key element of the design is our proposed block-based model of execution that groups vector computational instructions together to execute them in a coordinated manner. Milan Stanic, Oscar Palomar, Timothy Hayes 0001, Ivan Ratkovic, Osman S. Unsal, Adrián Cristal, Mateo Valero |
PACT | 2 |
| 2016 | Energy minimization at all layers of the data center: The ParaDIME project
Oscar Palomar, Santhosh Kumar Rethinagiri, Gulay Yalcin, J. Rubén Titos Gil, Pablo Prieto, Emma Torrella, Osman S. Unsal, Adrián Cristal, Pascal Felber, Anita Sobe, Yaroslav Hayduk, Mascha Kurpicz, Christof Fetzer, Thomas Knauth, Malte Schneegaß, Jens Struckmeier, Dragomir Milojevic |
DATE | 1 |
| 2016 | Future Vector Microprocessor Extensions for Data AggregationsabstractAs the rate of annual data generation grows exponentially, there is a demand to aggregate and summarise vast amounts of information quickly. In the past, frequency scaling was relied upon to push application throughput. Today, Dennard scaling has ceased and further performance must come from exploiting parallelism. Single instruction-multiple data (SIMD) instruction sets offer a highly efficient and scalable way of exploiting data-level parallelism (DLP). While microprocessors originally offered very simple SIMD support targeted at multimedia applications, these extensions have been growing both in width and functionality. Observing this trend, we use a simulation framework to model future SIMD support and then propose and evaluate five different ways of vectorising data aggregation. We find that although data aggregation is abundant in DLP, it is often too irregular to be expressed efficiently using typical SIMD instructions. Based on this observation, we propose a set of novel algorithms and SIMD instructions to better capture this irregular DLP. Furthermore, we discover that the best algorithm is highly dependent on the characteristics of the input. Our proposed solution can dynamically choose the optimal algorithm in the majority of cases and achieves speedups between 2.7x and 7.6x over a scalar baseline. Timothy Hayes 0001, Oscar Palomar, Osman S. Unsal, Adrián Cristal, Mateo Valero |
ISCA | 2 |
| 2016 | A Fully Parameterizable Low Power Design of Vector Fused Multiply-Add Using Active Clock-Gating TechniquesabstractThe need for power-efficiency is driving a rethink of design decisions in processor architectures. While vector processors succeeded in the high-performance market in the past, they need a re-tailoring for the mobile market that they are entering now. Floating point fused multiply-add, being a power consuming functional unit, deserves special attention. Although clock-gating is a well-known method to reduce switching power in synchronous designs, there are unexplored opportunities for its application to vector processors, especially when considering active operating mode. In this research, we comprehensively identify, propose, and evaluate the most suitable clock-gating techniques for vector fused multiply-add units (VFU). These techniques ensure power savings without jeopardizing the timing. Using vector masking and vector multi-lane-aware clock-gating, we report power reductions of up to 52%, assuming active VFU operating at the peak performance. Among other findings, we observe that vector instruction-based clock-gating techniques achieve power savings for all vector floating-point instructions. We perform this research in a fully parameterizable and automated fashion using various tools at both architectural and circuit levels. Ivan Ratkovic, Oscar Palomar, Milan Stanic, Osman S. Unsal, Adrián Cristal, Mateo Valero |
ISLPED | 2 |
| 2016 | Exploring Energy Reduction in Future Technology Nodes via Voltage Scaling with Application to 10nmabstractDI-fusion, le Dépôt institutionnel numérique de l'ULB, est l'outil de référencementde la production scientifique de l'ULB.L'interface de recherche DI-fusion permet de consulter les publications des chercheurs de l'ULB et les thèses qui y ont été défendues. Gulay Yalcin, Santhosh Kumar Rethinagiri, Oscar Palomar, Osman S. Unsal, Adrián Cristal, Dragomir Milojevic |
PDP | 3 |
| 2016 | Architectural support for efficient message passing on shared memory multi-cores
J. Rubén Titos Gil, Oscar Palomar, Osman S. Unsal, Adrián Cristal |
J. Parallel Distributed Comput. | 2 |
| 2015 | Runtime-Aware Architectures
Marc Casas, Miquel Moretó, Lluc Alvarez, Emilio Castillo, Dimitrios Chasapis, Timothy Hayes 0001, Luc Jaulmes, Oscar Palomar, Osman S. Unsal, Adrián Cristal, Eduard Ayguadé, Jesús Labarta, Mateo Valero |
Euro-Par | 8 |
| 2015 | Heterogeneous Platform to Accelerate Compute Intensive ApplicationsabstractNowadays image processing applications are widely used in various industries such as traffic, safety, medical engineering, etc. In this paper, we propose a power and energy efficient heterogeneous platform to accelerate image processing applications. To achieve this efficiency, we propose a novel hybrid platform which consists of a Xilinx Zynq (ARM+FPGA) and an NVidias Jetson TK1 (ARM+GPU) coupled with PCIe card. In applications such face recognition, we optimized major tasks in detection and recognition in order to achieve a speedup of 69× when compared to sequential execution on the ARM core, 4.8× against Zynq platform (ARM+FPGA), 3.2× against NVidia platform (ARM+GPU) and 40% more energy efficient against sequential execution. Santhosh Kumar Rethinagiri, Oscar Palomar, Javier Arias Moreno, Osman S. Unsal, Adrián Cristal |
FCCM | 2 |
| 2015 | Trigeneous Platforms for Energy Efficient Computing of HPC ApplicationsabstractIn this paper, we present two novel real-time heterogeneous platforms with three kinds of devices (CPU, GPU, FPGA), i.e. trigeneous platforms, for efficiently accelerating computation intensive applications in both the high-performance computing and the embedded system domains. In the high-performance computing domain, the entire platform is implemented on a workstation which consists of an Intel Xeon E5 processor, a Nvidia Tesla GPU and a Xilinx Virtex 7 FPGA. The second platform is built for achieving high-performance in the real-time embedded system domain. For this platform, we use a Xilinx Zynq and Nvidia Jetson TK1 board. In these platforms, the communication is performed using PCIe Gen3 and PCIe Gen2 cards respectively. We conducted experiments using 5 real-time and high throughput computation-data-intensive applications, namely cone beam computed tomography, face recognition, HEVC UHD decoding, number plate recognition and motion tracking. All the applications are mapped to the devices of the proposed trigeneous platforms, based on the energy efficiency of the different tasks on each device but also minimizing data transfers and maximizing parallelism. With this trigeneous platform, we are able to achieve an average speed-up of 21x compared to a CPU-GPU platform, 24x compared to a CPU-FPGA platform and 70x compared to Quad-core CPU alone execution. The proposed trigeneous platforms save 43%, 56% and 64% of the energy when compared to CPU-GPU, CPU-FPGA and quad-core CPU platforms respectively. Furthermore, we also implemented these applications by using a single programming language (OpenCL) on the trigeneous platforms and achieved 6x of speed-up on average over the quad-core setup. Santhosh Kumar Rethinagiri, Oscar Palomar, Javier Arias Moreno, Osman S. Unsal, Adrián Cristal |
HiPC | 2 |
| 2015 | VSR sort: A novel vectorised sorting algorithm & architecture extensions for future microprocessorsabstractSorting is a widely studied problem in computer science and an elementary building block in many of its subfields. There are several known techniques to vectorise and accelerate a handful of sorting algorithms by using single instruction-multiple data (SIMD) instructions. It is expected that the widths and capabilities of SIMD support will improve dramatically in future microprocessor generations and it is not yet clear whether or not these sorting algorithms will be suitable or optimal when executed on them. This work extrapolates the level of SIMD support in future microprocessors and evaluates these algorithms using a simulation framework. The scalability, strengths and weaknesses of each algorithm are experimentally derived. We then propose VSR sort, our own novel vectorised non-comparative sorting algorithm based on radix sort. To facilitate the execution of this algorithm we define two new SIMD instructions and propose a complementary hardware structure for their execution. Our results show that VSR sort has maximum speedups between 14.9x and 20.6x over a scalar baseline and an average speedup of 3.4x over the next-best vectorised sorting algorithm. Timothy Hayes 0001, Oscar Palomar, Osman S. Unsal, Adrián Cristal, Mateo Valero |
HPCA | 2 |
| 2015 | VPM: Virtual power meter tool for low-power many-core/heterogeneous data center prototypesabstractPower and energy consumption of data centers are steadily increasing and the work performed by the data centers is not proportional to the power dissipated, where every μA is a revenue for the entity. On the one hand, the hardware community is proposing various methodologies to address this issue such as low-power processors, heterogeneity, etc. to reduce the power of the servers. On the other hand, the software community proposes mechanisms such as virtual machines (VMs), work-load scheduling, etc. to increase the utilization of the processor. In order to properly evaluate the impact of these mechanisms, we need an accurate power monitoring and estimation tool at the hardware host level, the VM level and the system-level. This paper proposes a novel power monitoring middleware on a low-power platform at the node level (ARM Big.LITTLE) and an estimation methodology by using a simulator for future data center prototypes at any given level of virtualization. First, we built an instrumentation framework to measure the power based on hardware counter activities and with respect to current fluctuation. This allows us to build power models for the corresponding platforms, which are fed into the middleware to estimate power on the fly. Furthermore, we used the same framework for future low-power processors such as ARM Cortex-A57 and -A53 based platforms, which are integrated into the architectural simulator by providing an API to estimate power with the power model. Second, we present a machine learning-based energy efficient scheduling of the VMs that leverages VPM. The results obtained with the power monitoring middleware differ less than 2% from real board measurements and 5% when using the simulation environment regardless of the number of virtual machines used. Furthermore, we reduced 40% of energy consumption on average when compared to default scheduling of the KVM hypervisor. Santhosh Kumar Rethinagiri, Oscar Palomar, Javier Arias Moreno, Osman S. Unsal, Adrián Cristal |
ICCD | 2 |
| 2015 | DiMP: Architectural Support for Direct Message Passing on Shared Memory Multi-coresabstractThanks to programming approaches like actor-based models, message passing is regaining popularity outside large-scale scientific computing for building scalable distributed applications in many-core processors. Unfortunately, the mismatch between message passing models and today's shared-memory hardware provided by commercial vendors results in suboptimal performance and loss of efficiency. This paper presents a set of architectural extensions to reduce the overheads incurred by message passing workloads running on shared memory multi-core architectures. It describes the instruction set extensions and the hardware implementation. In order to facilitate programmability, the proposed extensions are used by a message passing library, allowing programs to take advantage of them transparently. As a proof-of-concept, we use a modified MPICH library and MPI programs to evaluate the proposal. Experimental results show that, on average, our proposal spends 60% less cycles performing data transfers in MPI functions, and reduces the L1 data cache misses in said functions to a fourth. J. Rubén Titos Gil, Oscar Palomar, Osman S. Unsal, Adrián Cristal |
ICPP | 2 |
| 2014 | PVMC: Programmable Vector Memory ControllerabstractIn this work, we propose a Programmable Vector Memory Controller (PVMC), which boosts noncontiguous vector data accesses by integrating descriptors of memory patterns, a specialized local memory, a memory manager in hardware, and multiple DRAM controllers. We implemented and validated the proposed system on an Altera DE4 FPGA board. We compare the performance of our proposal with a vector system without PVMC as well as a scalar only system. When compared with a baseline vector system, the results show that the PVMC system transfers data sets up to 2.2× to 14.9× faster, achieves between 2.16× to 3.18× of speedup for 5 applications and consumes 2.56 to 4.04 times less energy. Tassadaq Hussain, Oscar Palomar, Osman S. Unsal, Adrián Cristal, Eduard Ayguadé, Mateo Valero |
ASAP | 2 |
| 2014 | EVX: Vector execution on low power EDGE coresabstractIn this paper, we present a vector execution model that provides the advantages of vector processors on low power, general purpose cores, with limited additional hardware. While accelerating data-level parallel (DLP) workloads, the vector model increases the efficiency and hardware resources utilization. We use a modest dual issue core based on an Explicit Data Graph Execution (EDGE) architecture to implement our approach, called EVX. Unlike most DLP accelerators which utilize additional hardware and increase the complexity of low power processors, EVX leverages the available resources of EDGE cores, and with minimal costs allows for specialization of the resources. EVX adds a control logic that increases the core area by 2.1%. We show that EVX yields an average speedup of 3x compared to a scalar baseline and outperforms multimedia SIMD extensions. Milovan Duric, Oscar Palomar, Aaron Smith, Osman S. Unsal, Adrián Cristal, Mateo Valero, Doug Burger |
DATE | 2 |
| 2014 | ParaDIME: Parallel Distributed Infrastructure for Minimization of EnergyabstractDramatic environmental and economic impact of the ever increasing power and energy consumption of modern computing devices in data centers is now a critical challenge. On one hand, designers use technology scaling as one of the methods to face the phenomenon called dark silicon (only segments of a chip function concurrently due to power restrictions). On the other hand, designers use extreme-scale systems such as teradevices to meet the performance needs of their applications which in turn increases the power consumption of the platform. In order to overcome these challenges, we need novel computing paradigms that address energy efficiency. One of the promising solutions is to incorporate parallel distributed methodologies at different abstraction levels. The FP7 project ParaDIME focuses on this objective to provide different distributed methodologies (software-hardware techniques) at different abstraction levels to attack the power-wall problem. In particular, the ParaDIME framework will utilize: circuit and architecture operation below safe voltage limits for drastic energy savings, specialized energy-aware computing accelerators, heterogeneous computing, energy-aware runtime, approximate computing and power-aware message passing. The major outcome of the project will be a processor architecture for a heterogeneous distributed system that utilizes future device characteristics for drastic energy savings. Wherever possible, ParaDIME will adopt multidisciplinary techniques, such as hardware support for message passing, runtime energy optimization utilizing new hardware energy performance counters, use of accelerators for error recovery from sub-safe voltage operation, and approximate computing through annotated code. Furthermore, we will establish and investigate the theoretical limits of energy savings at the device, circuit, architecture, runtime and programming model levels of the computing stack, as well as quantify the actual energy savings achieved by the ParaDIME approach for the complete computing stack with the real environment. Santhosh Kumar Rethinagiri, Oscar Palomar, Anita Sobe, Thomas Knauth, Wojciech M. Barczynski, Gulay Yalcin, Yaroslav Hayduk, Adrián Cristal, Osman S. Unsal, Pascal Felber, Christof Fetzer, Julien Ryckaert, Gina Alioto |
DSD | 2 |
| 2014 | APMC: advanced pattern based memory controller (abstract only)abstractIn this paper, we present APMC, the Advanced Pattern based Memory Controller, that uses descriptors to support both regular and irregular memory access patterns without using a master core. It keeps pattern descriptors in memory and prefetches the complex 1D/2D/3D data structure into its special scratchpad memory. Support for irregular Memory accesses are arranged in the pattern descriptors at program-time and APMC manages multiple patterns at run-time to reduce access latency. The proposed APMC system reduces the limitations faced by processors/accelerators due to irregular memory access patterns and low memory bandwidth. It gathers multiple memory read/write requests and maximizes the reuse of opened SDRAM banks to decrease the overhead of opening and closing rows. APMC manages data movement between main memory and the specialized scratchpad memory; data present in the specialized scratchpad is reused and/or updated when accessed by several patterns. The system is implemented and tested on a Xilinx ML505 FPGA board. The performance of the system is compared with a processor with a high performance memory controller. The results show that the APMC system transfers regular and irregular datasets up to 20.4x and 3.4x faster respectively than the baseline system. When compared to the baseline system, APMC consumes 17% less hardware resources, 32% less on-chip power and achieves between 3.5x to 52x and 1.4x to 2.9x of speedup for regular and irregular applications respectively. The APMC core consumes 50% less hardware resources than the baseline system's memory controller. In this paper, we present APMC, the Advanced Pattern based Memory Controller, an intelligent memory controller that uses descriptors to supports both regular and irregular memory access patterns. support of the master core. It keeps pattern descriptors in memory and prefetches the complex data structure into its special scratchpad memory. Memory accesses are arranged in the pattern descriptors at program-time and APMC manages multiple patterns at run-time to reduce access latency. The proposed APMC system reduces the limitations faced by processors/accelerators due to irregular memory access patterns and low memory bandwidth. The system is implemented and tested on a Xilinx ML505 FPGA board. The performance of the system is compared with a processor with a high performance memory controller. The results show that the APMC system transfers regular and irregular datasets up to 20.4x and 3.4x faster respectively than the baseline system. When compared to the baseline system, APMC consumes 17% less hardware resources, 32% less on-chip power and achieves between 3.5x to 52x and 1.4x to 2.9x of speedup for regular and irregular applications respectively. The APMC core consumes 50% less hardware resources than the baseline system's memory controller.memory accesses. Tassadaq Hussain, Oscar Palomar, Osman S. Unsal, Adrián Cristal, Eduard Ayguadé, Mateo Valero, Santhosh Kumar Rethinagiri |
FPGA | 2 |
| 2014 | Power estimation tool for system on programmable chip based platforms (abstract only)abstractThe ever increasing complexity of the applications result in the development of power hungry processors. There is a scarcity of standalone tools that have a good trade off between estimation speed and accuracy to estimate power/energy at an earlier phase of design flow. There are very few tools that addresses the design space exploration issue based on power and energy. In this paper, we propose a virtual platform based standalone power and energy estimation tool for System-on-Programmable Chip (SoPC) embedded platforms, which is independent of in-house tools. There are two steps involved in this tool development. The first step is power model generation. For the power model development, we used functional parameters to set up generic power models for the different parts of the system. This is a onetime activity. In the second step, a simulation based virtual platform framework is developed to evaluate accurately the activities used in the related power models developed in the first step. The combination of the two steps lead to a hybrid power estimation, which gives a better trade-off between accuracy and speed. The proposed tool has several benefits: it considers the power consumption of the embedded system in its entirety and leads to accurate estimates without a costly and complex material. The proposed tool is also scalable for exploring complex embedded multi-core architectures. Santhosh Kumar Rethinagiri, Oscar Palomar, Adrián Cristal, Osman S. Unsal |
FPGA | 2 |
| 2014 | MAPC: Memory access pattern based controllerabstractTraditionally, system designers have attempted to improve system performance by scheduling the processing cores and by exploring different memory system configurations and there is comparatively less work done scheduling the accesses at the memory system level and exploring data accesses on the memory systems. In this paper, we propose a memory access pattern based controller (MAPC). MAPC organizes data accesses in descriptors, prioritizes them with respect to the number and size of transfer requests. When compared to the baseline multicore system, the MAPC based system achieves between 2.41× to 5.34× of speedup for different applications, consumes 28% less hardware resources and 13% less dynamic power. Tassadaq Hussain, Oscar Palomar, Osman S. Unsal, Adrián Cristal, Eduard Ayguadé, Mateo Valero |
FPL | 2 |
| 2014 | AMMC: Advanced Multi-Core Memory ControllerabstractIn this work, we propose an efficient scheduler and intelligent memory manager known as AMMC (Advanced Multi-Core Memory Controller), which proficiently handles data movement and computational tasks. The proposed AMMC system improves performance by managing complex data transfers at run-time and scheduling multi-cores without the intervention of a control processor nor an operating system. AMMC has been coupled with a heterogeneous system that provides both general-purpose cores and application specific accelerators. The AMMC system is implemented and tested on a Xilinx ML505 evaluation FPGA board. The performance of the system is compared with a microprocessor based system that has been integrated with the Xilkernel operating system. Results show that the AMMC based multi-core system consumes 48% less hardware resources, 27.9% less on-chip power and achieves 6.8x of speed-up compared to the MicroBlaze-based multi-core system. Tassadaq Hussain, Oscar Palomar, Osman S. Unsal, Adrián Cristal, Eduard Ayguadé, Mateo Valero, Shakaib A. Gursal |
FPT | 2 |
| 2012 | Vector Extensions for Decision Support DBMS AccelerationabstractDatabase management systems (DBMS) have become an essential tool for industry and research and are often a significant component of data centres. As a result of this criticality, efficient execution of DBMS engines has become an important area of investigation. This work takes a top-down approach to accelerating decision support systems (DSS) on x86-64 microprocessors using vector ISA extensions. In the first step, a leading DSS DBMS is analysed for potential data-level parallelism. We discuss why the existing multimedia SIMD extensions (SSE/AVX) are not suitable for capturing this parallelism and propose a complementary instruction set reminiscent of classical vector architectures. The instruction set is implemented using unintrusive modifications to a modern x86-64 micro architecture tailored for DSS DBMS. The ISA and micro architecture are evaluated using a cycle-accurate x86-64 micro architectural simulator coupled with a highly-detailed memory simulator. We have found a single operator is responsible for 41% of total execution time for the TPC-H DSS benchmark. Our results show performance speedups between 1.94x and 4.56x for an implementation of this operator run with our proposed hardware modifications. Timothy Hayes 0001, Oscar Palomar, Osman S. Unsal, Adrián Cristal, Mateo Valero |
MICRO | 2 |
| 2009 | Reusing cached schedules in an out-of-order processor with in-order issue logicabstractThe complex and powerful out-of-order issue logic dismisses the repetitive nature of the code, unlike what caches or branch predictors do. We show that 90% of the cycles, the group of instructions selected by the issue logic belongs to just 13% of the total different groups issued: the issue logic of an out-of-order processor is constantly re-discovering what it has already found. To benefit from the repetitive nature of instruction issue, we move the scheduling logic after the commit stage, out of the critical path of execution. The schedules created there are cached and reused to feed a simple in-order issue logic, that could result in a higher frequency design. We present the complete design of our ReLaSch processor, that achieves the same average IPC than a conventional out-of-order processor, and a 1.56 speed-up over the IPC of an in-order processor. We actually surpass the out-of-order IPC in 23 out of 40 SPEC benchmarks, mainly because the broader vision of the code after the commit stage allows creating better schedules. Oscar Palomar, Toni Juan, Juan J. Navarro |
ICCD | 1 |