Grigoris Dimitroulakos

dblp:05/2913 · also Gregory Dimitroulakos · DBLP profile ↗
← Back
29ranked-venue papers
8as first author
1since 2021 · last 2023
0000-0001-7580-3914ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 25 · 8 first-authorSoftware engineering, systems software and programming languages · 4 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
2 papers
Software maintenance and evolution · 77% Programming languages and type systems · 16% Empirical software engineering · 7%

Topics — the 3 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Software maintenance and evolution › software maintenance
maintenance effort prediction
0.712023
Simulating Software Evolution to Evaluate the Reliability of Early Decision-making among Design Alternatives toward Maintainability · ACM Trans. Softw. Eng. Methodol. 2023
Programming languages and type systems
simulation
0.212023
Simulating Software Evolution to Evaluate the Reliability of Early Decision-making among Design Alternatives toward Maintainability · ACM Trans. Softw. Eng. Methodol. 2023
Empirical software engineering › software metrics
software quality metrics
0.112017
Early Evaluation of Implementation Alternatives of Composite Data Structures Toward Maintainability · ACM Trans. Softw. Eng. Methodol. 2017

Methods — techniques the papers use, named apart from their topics

time series analysis · 0.7statistical validation · 0.7simulation · 0.7formal model · 0.3comparison model · 0.3
YearPublicationVenuePosition
2023 Simulating Software Evolution to Evaluate the Reliability of Early Decision-making among Design Alternatives toward Maintainability
abstract
Critical decisions among design altern seventh atives with regards to maintainability arise early in the software design cycle. Existing comparison models relayed on the structural evolution of the used design patterns are suitable to support such decisions. However, their effectiveness on predicting maintenance effort is usually verified on a limited number of case studies under heterogeneous metrics. In this article, a multi-variable simulation model for validating the decision-making reliability of the derived formal comparison models for the significant designing problem of recursive hierarchies of part-whole aggregations, proposed in our prior work, is introduced. In the absence of a strict validation, the simulation model has been thoroughly calibrated concerning its decision-making precision based on empirical distributions from time-series analysis, approximating the highly uncertain nature of actual maintenance process. The decision reliability of the formal models has been statistically validated on a sample of 1,000 instances of design attributes representing the entire design space of the problem. Despite the limited accuracy of measurements, the results show that the models demonstrate an increasing reliability in a long-term perspective, even under assumptions of high variability. Thus, the modeling theory discussed in our prior work delivers reliable models that significantly reduce decision-risk and relevant maintenance cost.
Chris Karanikolas, Grigoris Dimitroulakos, Kostas Masselos
ACM Trans. Softw. Eng. Methodol.2
2020 A Retargetable MATLAB-to-C Compiler Exploiting Custom Instructions and Data Parallelism
abstract
This article presents a MATLAB-to-C compiler that exploits custom instructions present in state-of-the-art processor architectures and supports semi-automatic vectorization. A parameterized processor model is used to describe the target instruction set architecture to achieve user-friendly retargetability. Custom instructions are represented via specialized intrinsic functions in the generated code, which can then be used as input to any C/C++ compiler supporting the target processor. In addition, the compiler supports the generation of data parallel/vectorized code through the introduction of data packing/unpacking statements. The compiler has been used for code generation targeting ARM and x86 architectures for several benchmarks. The vectorized code generated by the compiler achieves an average speedup of 4.1× and 2.7× for packed fixed and floating point data, respectively, compared to scalarized code for ARM architecture and an average speedup of 3.1× and 1.5× for packed fixed and floating point data, respectively, for x86 architecture. Implementing data parallel instructions directly in the assembly code would have required a lot of design effort, and it would not been sustainable across evolving platform variants. Thus, the compiler can be employed to efficiently speed up critical sections of the target application. The compiler is therefore potentially employable to raise the design abstraction and reduce development time for both embedded and general-purpose applications.
Ioannis Latifis, Karthick Parashar, Grigoris Dimitroulakos, Hans Cappelle, Christakis Lezos, Kostas Masselos, Francky Catthoor
ACM Trans. Embed. Comput. Syst.3
2020 A Locality Optimizer for Loop-dominated Applications Based on Reuse Distance Analysis
abstract
Source code optimization can heavily improve software code implementation quality while still being complementary to conventional compilers’ optimizations. Source code analysis tools are very useful in supporting source code optimization. This article discusses MemAssist, a source-level optimization environment for semi-automatic locality optimization of loop-dominated code. MemAssist applies reuse distance analysis and a relevant optimization algorithm to explore the design space. It generates a set of suggestions for locality optimizing loop transformations that reduce data cache miss rate and execution time. MemAssist has been used to optimize a number of applications. Experimental results show that MemAssist leads to cache miss rate reduction at all cache layers, memory accesses reduction by up to 42%, and to a speedup of up to three times. Therefore, MemAssist can be used for efficient early-stage software optimization leading to development effort and time reduction.
Christakis Lezos, Grigoris Dimitroulakos, Ioannis Latifis, Kostas Masselos
ACM Trans. Design Autom. Electr. Syst.2
2017 A MATLAB Vectorizing Compiler Targeting Application-Specific Instruction Set Processors
abstract
This article discusses a MATLAB-to-C vectorizing compiler that exploits custom instructions, for example, for Single Instruction Multiple Data (SIMD) processing and instructions for complex arithmetic present in Application-Specific Instruction Set Processors (ASIPs). Custom instructions are represented via specialized intrinsic functions in the generated code, and the generated code can be used as input to any C/C++ compiler supporting the target processor. Furthermore, the specialized instruction set of the target processor is described in a parameterized way using a target processor-independent architecture description approach, thus allowing the support of any processor. The compiler has been used for the generation of application code for two different ASIPs for several benchmarks. The code generated by the compiler achieves a speedup between 2× --74× and 2× --97× compared to the code generated by the MathWorks MATLAB-to-C compiler. Experimental results also prove that the compiler efficiently exploits SIMD custom instructions achieving a 3.3 factor speedup compared to cases where no SIMD processing is used. Thus the compiler can be employed to reduce the development time/effort/cost and time to market through raising the abstraction of application design in an embedded systems/system-on-chip development context.
Ioannis Latifis, Karthick Parashar, Grigoris Dimitroulakos, Hans Cappelle, Christakis Lezos, Kostas Masselos, Francky Catthoor
ACM Trans. Design Autom. Electr. Syst.3
2017 Early Evaluation of Implementation Alternatives of Composite Data Structures Toward Maintainability
abstract
Selecting between different design options is a crucial decision for object-oriented software developers that affects code quality characteristics. Conventionally developers use their experience to make such decisions, which leads to suboptimal results regarding code quality. In this article, a formal model for providing early estimates of quality metrics of object-oriented software implementation alternatives is proposed. The model supports software developers in making fast decisions in a systematic way early during the design phase to achieve improved code characteristics. The approach employs a comparison model related to the application of the Visitor design pattern and inheritance-based implementation on structures following the Composite design pattern. The model captures maintainability as a metric of software quality and provides precise assessments of the quality of each implementation alternative. Furthermore, the model introduces the structural maintenance cost metric based on which the progressive analysis of the maintenance process is introduced. The proposed approach has been applied to several test cases for different relevant quality metrics. The results prove that the proposed model delivers accurate estimations. Thus, the proposed methodology can be used for comparing different implementation alternatives against various measures and quality factors before code development, leading to reduced effort and cost for software maintenance.
Chris Karanikolas, Grigoris Dimitroulakos, Kostas Masselos
ACM Trans. Softw. Eng. Methodol.2
2016 Matlab to C compilation targeting Application Specific Instruction Set Processors
Ioannis Latifis, Karthick Parashar, Grigoris Dimitroulakos, Hans Cappelle, Christakis Lezos, Kostas Masselos, Francky Catthoor
DATE3
2016 Compiler-Directed Data Locality Optimization in MATLAB
abstract
Array programming languages, such as MATLAB, are often used for algorithm development by scientists and engineers without taking into consideration implementation related issues and with limited emphasis on relevant optimizations. Application code optimization, especially in terms of data storage and transfer behavior, is still an important issue and heavily affects implementations' quality in terms of performance, power consumption etc. Efficient approaches for the optimization of high level application code are required to derive high quality implementations while still reducing development time and cost. This paper presents MemAssist, a software tool supporting application developers in detecting parts of the application code in MATLAB that do not exploit efficiently the targeted processor architecture and especially the memory hierarchy. Furthermore, the proposed tool guides application developers in applying code transformations in MATLAB for the optimization of the algorithm's temporal data locality. An image processing algorithm has been optimized using MemAssist as a practical usage scenario. Experimental results prove that the use of MemAssist can heavily reduce cache misses (up to 40%) and improve execution time (up to 30% speedup) on two different processor architectures. Thus, MemAssist can be used for optimized application code development that can lead to efficient implementations while still reducing development time and cost.
Christakis Lezos, Ioannis Latifis, Grigoris Dimitroulakos, Kostas Masselos
SCOPES3
2015 Reuse distance analysis for locality optimization in loop-dominated applications
Christakis Lezos, Grigoris Dimitroulakos, Kostas Masselos
DATE2
2009 Compiler assisted architectural exploration framework for coarse grained reconfigurable arrays
Grigoris Dimitroulakos, Nikos Kostaras, Michalis D. Galanis, Constantinos E. Goutis
J. Supercomput.1
2007 Compiler assisted architectural exploration for coarse grained reconfigurable arrays
abstract
A large number of factors influence the hardware cost and the mapping efficiency of applications on coarse grain reconfigurable architectures. This paper investigates for the first time in a unified way the four factors that are directly related with the efficiency of a coarse grain reconfigurable array architecture namely; the area the clock frequency, the scheduling efficiency and performance. An exploration framework has been build for estimating the values of the 4 a forementioned factors for different architecture alternatives. The exploration framework is composed of an existing retargetable compiler framework from which we estimate the mapping efficiency and the parametric realization of the coarse grained reconfigurable array architecture in hardware description language from which we estimate the clock frequency and the area of each architecture instance. The experiments refer to different architecture alternatives in terms of the processing elements' interconnection network, the register files' size, their number of input/output ports, and finally the available bandwidth. Totally 72 architecture scenarios have been studied revealing how each characteristic influences performance and area for efficiently make design decisions.
Grigoris Dimitroulakos, Nikos Kostaras, Michalis D. Galanis, Constantinos E. Goutis
ACM Great Lakes Symposium on VLSI1
2007 Improving performance and energy consumption in embedded microprocessor platforms with a flexible custom coprocessor data-path
abstract
The speedups and the energy reductions achieved in a generic single-chip microprocessor system by employing a high-performance data-path are presented. The data-path acts as a coprocessor that accelerates computational intensive kernel sections thereby increasing the overall performance. It is composed by Flexible Computational Components that can realize any two-level sequence of primitive operations. The automated coprocessor synthesis method from high-level software description and its integration to a design flow for executing applications on the system is presented. The estimated application speedups of eight real-life applications, relative to the software execution on the microprocessor range from 1.78 to 4.00, while the overhead in circuit area is small. The energy savings have an average value of 61%. A comparison with another high-performance data-path showed that the proposed coprocessor achieves smaller area-time products, by an average of 23%, for the synthesized data-paths.
Michalis D. Galanis, Grigoris Dimitroulakos, Constantinos E. Goutis
ACM Great Lakes Symposium on VLSI2
2007 Speedups and Energy Savings of Microprocessor Platforms with a Coarse-Grained Reconfigurable Data-Path
abstract
This paper presents the performance improvements and the energy reductions by coupling a high-performance coarse-grained reconfigurable data-path with a microprocessor in a generic platform. The datapath has been previously introduced by the authors. It is composed by computational units able to realize complex operations which aid in improving the performance of time critical application parts, called kernels. A design flow is proposed for mapping high-level software descriptions to the microprocessor system. Eight real-life applications are mapped on three different instances of the system. Significant overall application speedups, relative to a software-only solution, ranging from 1.74 to 3.94 are reported being close to theoretical speedup bounds. Average energy savings of 59% are achieved, while the reduction in the system energy-delay product ranges from 66% to 92%.
Michalis D. Galanis, Grigoris Dimitroulakos, Constantinos E. Goutis
IPDPS2
2007 Design space exploration of an optimized compiler approach for a generic reconfigurable array architecture
Grigoris Dimitroulakos, Michalis D. Galanis, Constantinos E. Goutis
J. Supercomput.1
2007 Exploring the speedups of embedded microprocessor systems utilizing a high-performance coprocessor data-path
Michalis D. Galanis, Grigoris Dimitroulakos, Constantinos E. Goutis
J. Supercomput.2
2007 Speedups in embedded systems with a high-performance coprocessor datapath
abstract
This article presents the speedups achieved in a generic single-chip microprocessor system by employing a high-performance datapath. The datapath acts as a coprocessor that accelerates computational-intensive kernel sections thereby increasing the overall performance. We have previously introduced the datapath which is composed of Flexible Computational Components (FCCs). These components can realize any two-level template of primitive operations. The automated coprocessor synthesis method from high-level software description and its integration to a design flow for executing applications on the system is presented. For evaluating the effectiveness of our coprocessor approach, analytical study in respect to the type of the custom datapath and to the microprocessor architecture is performed. The overall application speedups of several real-life applications relative to the software execution on the microprocessor are estimated using the design flow. These speedups range from 1.75 to 5.84, with an average value of 3.04, while the overhead in circuit area is small. The design flow achieved the acceleration of the applications near to theoretical speedup bounds. A comparison with another high-performance datapath showed that the proposed coprocessor achieves smaller area-time products by an average of 23% for the generated datapaths. Additionally, the FCC coprocessor achieves better performance in accelerating kernels relative to software-programmable DSP cores.
Michalis D. Galanis, Grigoris Dimitroulakos, Spyros Tragoudas, Constantinos E. Goutis
ACM Trans. Design Autom. Electr. Syst.2
2007 Speedups and Energy Reductions From Mapping DSP Applications on an Embedded Reconfigurable System
abstract
This paper presents performance improvements and energy savings from mapping real-world benchmarks on an embedded single-chip platform that includes coarse-grained reconfigurable logic with a microprocessor. The reconfigurable hardware is a 2-D array of processing elements connected with a mesh-like network. Analytical results derived from mapping seven real-life digital signal processing applications, with the aid of an automated design flow, on six different instances of the system architecture are presented. Significant overall application speedups relative to an all-software solution, ranging from 1.81 to 3.99 are reported being close to theoretical speedup bounds. Additionally, the energy savings range from 43% to 71%. Finally, a comparison with a system coupling a microprocessor with a very long instruction word core shows that the microprocessor/coarse-grained reconfigurable array platform is more efficient in terms of performance and energy consumption.
Michalis D. Galanis, Grigoris Dimitroulakos, Constantinos E. Goutis
IEEE Trans. Very Large Scale Integr. Syst.2
2006 Exploring the design space of an optimized compiler approach for mesh-like coarse-grained reconfigurable architectures
abstract
In this paper we study the performance improvements and trade-offs derived from an optimized mapping approach applied on a parametric coarse grained reconfigurable array architecture. The processing elements' local register files and the processing elements' interconnection network is exploited for caching memory data values with data reuse opportunities. The data reused values are transferred through the processing elements' interconnection network hence, relieving the bus from the burden of transferring these values. A novel mapping algorithm is also proposed that uses a modulo scheduling technique. This algorithm targets on a flexible architecture template which permits experimental exploration over different architecture alternatives. The experimental results showed that the operation parallelism was significantly improved by our mapping approach. Additionally, we have outlined the relation that exists between the performance improvements and the memory access latency, the interconnection network and the processing elements' register file size.
Grigoris Dimitroulakos, Michalis D. Galanis, Constantinos E. Goutis
IPDPS1
2006 Design flow for optimizing performance in processor systems with on-chip coarse-grain reconfigurable logic
abstract
A design flow for processor platforms with on-chip coarse-grain reconfigurable logic is presented. The reconfigurable logic is realized by a 2-dimensional array of processing elements. Performance is improved by accelerating critical software loops, called kernels, on the reconfigurable array. Basic steps of the design flow have been automated. A procedure for detecting critical loops in the input C code was developed, while a mapping technique for coarse grain reconfigurable arrays, based on software pipelining, was also devised. Analytical results derived from mapping five real-life DSP applications on eight different instances of a generic system architecture are presented. Large values of instructions per cycle were achieved on two reconfigurable arrays that resulted in high-performance kernel mapping. Additionally, by mapping critical code on the reconfigurable logic, speedups ranging from 1.27 to 3.18 relative to an all-processor execution were achieved
Michalis D. Galanis, Grigoris Dimitroulakos, Constantinos E. Goutis
IPDPS2
2006 Mapping DSP applications on processor systems with coarse-grain reconfigurable hardware
abstract
In this paper, we present performance results from mapping five real-world DSP applications on an embedded system-on-chip that incorporates coarse-grain reconfigurable logic with an instruction-set processor. The reconfigurable logic is realized by a 2-dimensional array of processing elements. A mapping flow for improving application's performance by accelerating critical software parts, called kernels, on the coarse-grain reconfigurable array is proposed. Profiling is performed for detecting critical kernel code. For mapping the detected kernels on the reconfigurable logic a priority-based mapping algorithm has been developed. The experiments for three different instances of a generic system show that the speedup from executing kernels on the reconfigurable array ranges from 9.9 to 151.1, with an average value of 54.1, relative to the kernels' execution on the processor. Important overall application speedups, due to the kernels' acceleration, have been reported for the five applications. These overall performance improvements range from 1.3 to 3.7, with an average value of 2.3, relative to an all-software execution.
Michalis D. Galanis, Grigoris Dimitroulakos, Constantinos E. Goutis
IPDPS2
2006 Resource constrained modulo scheduling for coarse-grained reconfigurable arrays
abstract
It is widely known that bandwidth limitations degrade parallel systems' performance. This paper presents a mapping methodology for coarse-grain reconfigurable arrays which alleviates the bandwidth bottleneck by exploiting the processing elements interconnection network for transferring values with data reuse opportunities. A novel mapping algorithm is also proposed that uses a resource-aware modulo scheduling technique. From the application of the proposed mapping approach, significant improvements in performance were achieved while we have also quantified these improvements in respect to crucial architecture parameters such as the memory latency and the register file size. For this reason, our methodology targets on a parametric architecture template which can model a large number of existing architectures of this kind
Grigoris Dimitroulakos, Michalis D. Galanis, Constantinos E. Goutis
ISCAS1
2006 Mapping DSP applications on processor/coarse-grain reconfigurable array architectures
abstract
Results from mapping five real-world DSP applications on a system-on-chip that incorporates coarse-grain reconfigurable hardware with an instruction-set processor is presented. The reconfigurable logic is realized by a 2-dimensional array of processing elements. A mapping method for improving application's performance by accelerating critical software parts, called kernels, on the coarse-grain reconfigurable array is proposed. For mapping the detected kernels on the reconfigurable logic a priority-based mapping algorithm has been developed. Important overall application speedups, due to the kernels' acceleration, have been reported for the five applications. These overall performance improvements range from 1.27 to 3.07, with an average value of 2.16, relative to an all-software execution
Michalis D. Galanis, Grigoris Dimitroulakos, Constantinos E. Goutis
ISCAS2
2006 Performance Improvements from Partitioning Applications to FPGA Hardware in Embedded SoCs
Michalis D. Galanis, Grigoris Dimitroulakos, Constantinos E. Goutis
J. Supercomput.2
2006 Partitioning Methodology for Heterogeneous Reconfigurable Functional Units
Michalis D. Galanis, Grigoris Dimitroulakos, Constantinos E. Goutis
J. Supercomput.2
2005 Alleviating the Data Memory Bandwidth Bottleneck in Coarse-Grained Reconfigurable Arrays
abstract
It is widely known that parallel operation execution in multiprocessor systems generates a respective increase in memory accesses. Since the memory and bus subsystems provide a limited access bandwidth, the applications performance cannot be that high as the multiprocessor system capabilities promise. This is the case for the 2D coarse-grained reconfigurable arrays for which a mapping methodology that aims in improving the mapped applications' performance by alleviating the data bandwidth bottleneck, is presented in this paper. This is achieved by exploiting the applications' data reuse opportunities both at the data dependence and source code level and the architecture's foreground memory. The methodology considers a realistic 2D coarse-grained reconfigurable architecture template, which can model the majority of the existing coarse-grained reconfigurable array architectures. The experimental results show a significant reduction in both execution time and memory accesses for two architecture scenarios that has been achieved by the application of the proposed methodology on a representative set of DSP applications.
Grigoris Dimitroulakos, Michalis D. Galanis, Constantinos E. Goutis
ASAP1
2005 Speedups from Partitioning Critical Software Parts to Coarse-Grain Reconfigurable Hardware
abstract
In this paper, we propose a hardware/software partitioning method for improving applications' performance in embedded systems. Critical software parts are accelerated on hardware of a single-chip generic system comprised by an embedded processor and coarse-grain reconfigurable hardware. The reconfigurable hardware is realized by a 2D array of processing elements. The partitioning flow utilizes an analysis procedure at the basic-block level for detecting kernels in software. A list-based mapping algorithm has been developed for estimating the execution cycles of kernels on coarse-grain reconfigurable arrays. The proposed partitioning flow has been largely automated for a program description in C language. Extensive hardware/software experiments on five real-life applications are presented. It is shown that the benchmarks spend an average of 69% of their instruction count in 11% on average of their code that correspond to the kernels' code. The results illustrate that by mapping critical code on coarse-grain reconfigurable hardware, speedups ranging from 1.2 to 3.7, with an average value of 2.2, are achieved.
Michalis D. Galanis, Grigoris Dimitroulakos, Constantinos E. Goutis
ASAP2
2005 Accelerating Applications by Mapping Critical Kernels on Coarse-Grain Reconfigurable Hardware in Hybrid Systems
abstract
In this paper, we propose a method for speeding-up applications by partitioning them between the reconfigurable hardware blocks of different granularity and mapping critical parts of applications on the coarse-grain reconfigurable hardware. The partitioning method consists of four steps; the intermediate representation creation, the kernel identification, the mapping onto coarse-grain reconfigurable blocks, and the mapping onto the FPGA hardware. The method is validated using five real-world applications, where the speedup relative to an all-FPGA solution ranges from 1.4 to 3.1.
Michalis D. Galanis, Grigoris Dimitroulakos, Constantinos E. Goutis
FCCM2
2005 Performance Improvements using Coarse-Grain Reconfigurable Logic in Embedded SoCs
abstract
A hardware/software partitioning methodology for improving applications' performance in embedded single-chip systems is presented. Critical software parts are accelerated on hardware of a system comprised by an embedded processor and coarse-grain reconfigurable hardware. The reconfigurable hardware is realized by a 2-dimensional array of processing elements. The partitioning method uses a basic-block level analysis procedure for detecting kernels in software. A mapping algorithm for coarse-grain reconfigurable arrays has been developed for estimating the execution time of kernels on the reconfigurable hardware. The proposed partitioning flow has been largely automated for a program description in C language. Analytical hardware/software experiments on five real-world applications are given. The results show that by mapping critical parts on coarse-grain reconfigurable hardware, speedups ranging from 1.2 to 3.7, with an average value of 2.3, are achieved.
Grigoris Dimitroulakos, Michalis D. Galanis, Constantinos E. Goutis
FPL1
2005 A high-throughput, memory efficient architecture for computing the tile-based 2D discrete wavelet transform for the JPEG2000
Grigoris Dimitroulakos, Michalis D. Galanis, Athanasios Milidonis, Constantinos E. Goutis
Integr.1
2004 An Automated C++ Code and Data Partitioning Framework for Data Management of Data-Intensive Applications
Athanasios Milidonis, Grigoris Dimitroulakos, Michalis D. Galanis, George Theodoridis, Constantinos E. Goutis, Francky Catthoor
SCOPES2