Paraskevas Evripidou

dblp:66/3873 · DBLP profile ↗
← Back
26ranked-venue papers
6as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 20 · 6 first-authorDatabases, data management, data science and information retrieval · 3Computer networks · 1Software engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
6 papers
Parallel and multicore computing · 87% Processor architecture and microarchitecture · 7% Reconfigurable computing and FPGAs · 3%
Databases, data mining, and information retrieval
1 paper
Data models and query languages · 50% Data integration and cleaning · 50%

Topics — the 22 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing › speculative parallelization
data-triggered threads
0.522017
Data-Driven Concurrency for High Performance Computing · ACM Trans. Archit. Code Optim. 2017
Architectural Support for Data-Driven Execution · ACM Trans. Archit. Code Optim. 2014
Parallel and multicore computing
parallel programming models
0.322017
Data-Driven Concurrency for High Performance Computing · ACM Trans. Archit. Code Optim. 2017
Block Scheduling of Iterative Algorithms and Graph-Level Priority Scheduling in a Simulated Data-Flow Multiprocessor · IEEE Trans. Parallel Distributed Syst. 1993
Parallel and multicore computing › parallel programming models
dataflow programming
0.312017
Data-Driven Concurrency for High Performance Computing · ACM Trans. Archit. Code Optim. 2017
Parallel and multicore computing › parallel computation models
data-driven execution
0.212014
Architectural Support for Data-Driven Execution · ACM Trans. Archit. Code Optim. 2014
Processor architecture and microarchitecture
multithreading
0.112006
Data-Driven Multithreading Using Conventional Microprocessors · IEEE Trans. Parallel Distributed Syst. 2006
Parallel and multicore computing › parallel scheduling
thread scheduling
0.112006
Data-Driven Multithreading Using Conventional Microprocessors · IEEE Trans. Parallel Distributed Syst. 2006
Processor architecture and microarchitecture
chip multiprocessor
0.112014
Architectural Support for Data-Driven Execution · ACM Trans. Archit. Code Optim. 2014
Reconfigurable computing and FPGAs
FPGA prototyping
0.112014
Architectural Support for Data-Driven Execution · ACM Trans. Archit. Code Optim. 2014
Parallel and multicore computing
synchronization
0.112014
Architectural Support for Data-Driven Execution · ACM Trans. Archit. Code Optim. 2014
Data integration and cleaning
heterogeneous data sources
0.112005
A Relationally Complete Visual Query Language for Heterogeneous Data Sources and Pervasive Querying · ICDE 2005
Data models and query languages › query language
visual query language
0.112005
A Relationally Complete Visual Query Language for Heterogeneous Data Sources and Pervasive Querying · ICDE 2005
Parallel and multicore computing
MPI
0.012001
Net-dbx: A Web-Based Debugger of MPI Programs Over Low-Bandwidth Lines · IEEE Trans. Parallel Distributed Syst. 2001
Parallel and multicore computing › parallel computing
parallel program debugging
0.012001
Net-dbx: A Web-Based Debugger of MPI Programs Over Low-Bandwidth Lines · IEEE Trans. Parallel Distributed Syst. 2001
Memory systems
cache management
0.012006
Data-Driven Multithreading Using Conventional Microprocessors · IEEE Trans. Parallel Distributed Syst. 2006
Memory systems › cache
cache miss reduction
0.012006
Data-Driven Multithreading Using Conventional Microprocessors · IEEE Trans. Parallel Distributed Syst. 2006
Compilers and program optimization
memory optimization
0.011995
Advanced Array Optimizations for High Performance Functional Languages · IEEE Trans. Parallel Distributed Syst. 1995
Parallel and multicore computing › dataflow computing
dataflow scheduling
0.011993
Block Scheduling of Iterative Algorithms and Graph-Level Priority Scheduling in a Simulated Data-Flow Multiprocessor · IEEE Trans. Parallel Distributed Syst. 1993
Parallel and multicore computing › task scheduling
scheduling for performance
0.011993
Block Scheduling of Iterative Algorithms and Graph-Level Priority Scheduling in a Simulated Data-Flow Multiprocessor · IEEE Trans. Parallel Distributed Syst. 1993
Debugging and program repair › concurrent program debugging
distributed debugging
0.012001
Net-dbx: A Web-Based Debugger of MPI Programs Over Low-Bandwidth Lines · IEEE Trans. Parallel Distributed Syst. 2001
Parallel and multicore computing › parallel programming models
message passing
0.012001
Net-dbx: A Web-Based Debugger of MPI Programs Over Low-Bandwidth Lines · IEEE Trans. Parallel Distributed Syst. 2001
High-performance computing › numerical linear algebra › linear solver
iterative linear solvers
0.011993
Block Scheduling of Iterative Algorithms and Graph-Level Priority Scheduling in a Simulated Data-Flow Multiprocessor · IEEE Trans. Parallel Distributed Syst. 1993
High-performance computing
scientific computing systems
0.011993
Block Scheduling of Iterative Algorithms and Graph-Level Priority Scheduling in a Simulated Data-Flow Multiprocessor · IEEE Trans. Parallel Distributed Syst. 1993

Methods — techniques the papers use, named apart from their topics

dynamic dataflow · 0.3data-driven multithreading · 0.3simulation · 0.1web-based interface · 0.1thread synchronization unit · 0.1java · 0.1cacheflow policy · 0.1tuple relational calculus · 0.1relational algebra · 0.1predictive storage preallocation · 0.0framework preconstruction · 0.0aggregate storage subsumption · 0.0
YearPublicationVenuePosition
2019 Toward data-driven architectural support in improving the performance of future HPC architectures
George Matheou, Vassos Soteriou, Paraskevas Evripidou
Parallel Comput.3
2017 Data-Driven Concurrency for High Performance Computing
abstract
In this work, we utilize dynamic dataflow/data-driven techniques to improve the performance of high performance computing (HPC) systems. The proposed techniques are implemented and evaluated through an efficient, portable, and robust programming framework that enables data-driven concurrency on HPC systems. The proposed framework is based on data-driven multithreading (DDM), a hybrid control-flow/dataflow model that schedules threads based on data availability on sequential processors. The proposed framework was evaluated using several benchmarks, with different characteristics, on two different systems: a 4-node AMD system with a total of 128 cores and a 64-node Intel HPC system with a total of 768 cores. The performance evaluation shows that the proposed framework scales well and tolerates scheduling overheads and memory latencies effectively. We also compare our framework to MPI, DDM-VM, and OmpSs@Cluster. The comparison results show that the proposed framework obtains comparable or better performance.
George Matheou, Paraskevas Evripidou
ACM Trans. Archit. Code Optim.2
2014 Architectural Support for Data-Driven Execution
abstract
The exponential growth of sequential processors has come to an end, and thus, parallel processing is probably the only way to achieve performance growth. We propose the development of parallel architectures based on data-driven scheduling. Data-driven scheduling enforces only a partial ordering as dictated by the true data dependencies, which is the minimum synchronization possible. This is very beneficial for parallel processing because it enables it to exploit the maximum possible parallelism. We provide architectural support for data-driven execution for the Data-Driven Multithreading (DDM) model. In the past, DDM has been evaluated mostly in the form of virtual machines. The main contribution of this work is the development of a highly efficient hardware support for data-driven execution and its integration into a multicore system with eight cores on a Virtex-6 FPGA. The DDM semantics make barriers and cache coherence unnecessary, which reduces the synchronization latencies significantly and makes the cache simpler. The performance evaluation has shown that the support for data-driven execution is very efficient with negligible overheads. Our prototype can support very small problem sizes (matrix 16×16) and ultra-lightweight threads (block of 4x4) that achieve speedups close to linear. Such results cannot be achieved by software-based systems.
George Matheou, Paraskevas Evripidou
ACM Trans. Archit. Code Optim.2
2013 The TERAFLUX Project: Exploiting the DataFlow Paradigm in Next Generation Teradevices
abstract
Thanks to the improvements in semiconductor technologies, extreme-scale systems such as teradevices (i.e., composed by 1000 billion of transistors) will enable systems with 1000+ general purpose cores per chip, probably by 2020. Three major challenges have been identified: programmability, manageable architecture design, and reliability. TERAFLUX is a Future and Emerging Technology (FET) large-scale project funded by the European Union, which addresses such challenges at once by leveraging the dataflow principles. This paper describes the project and provides an overview of the research carried out by the TERAFLUX consortium.
Marco Solinas, Rosa M. Badia, François Bodin, Albert Cohen 0001, Paraskevas Evripidou, Paolo Faraboschi, Bernhard Fechner, Guang R. Gao, Arne Garbade, Sylvain Girbal, Daniel Goodman 0001, Behram Khan, Souad Koliai, Feng Li 0016, Mikel Luján, Laurent Morin, Avi Mendelson, Nacho Navarro, Antoniu Pop, Pedro Trancoso, Theo Ungerer, Mateo Valero, Sebastian Weis, Ian Watson, Stéphane Zuckerman, Roberto Giorgi
DSD5
2011 DDM-VMc: the data-driven multithreading virtual machine for the cell processor
abstract
In this paper we present the Data-Driven Multithreading Virtual Machine for the Cell Processor (DDM-VMc). Data-Driven Multithreading is a non-blocking multithreading model that decouples the synchronization from the computation portions of a program allowing them to execute asynchronously in a data-flow manner. The core of the DDM model is the Thread Scheduling Unit (TSU), which schedules threads dynamically at runtime based on data availability. DDM-VMc implements the TSU as a software module running on the PPE core of the Cell, allowing the SPE cores to execute the program threads. DDM-VMc virtualizes the parallel resources of the Cell, handles the heterogeneity of the cores and manages the Cell memory hierarchy efficiently.
Samer Arandi, Paraskevas Evripidou
HiPEAC2
2009 Programming Abstractions and Toolchain for Dataflow Multithreading Architectures
abstract
The need to exploit multi-core systems for parallel processing has revived the concept of dataflow. In particular, the dataflow multithreading architectures have proven to be good candidates for these systems. In this work we propose an abstraction layer that enables compiling and running a program written for an abstract dataflow multithreading architecture on different implementations. More specifically, we present a set of compiler directives that provide the programmer with the means to express most types of dependencies between code segments. In addition, we present the corresponding toolchain that transforms this code into a form that can be compiled for different implementations of the model. As a case study for this work, we present the usage of the toolchain for the TFlux and DTA architectures.
Kyriakos Stavrou, Demos Pavlou, Marios Nikolaides, Panayiotis Petrides, Paraskevas Evripidou, Pedro Trancoso, Zdravko Popovic, Roberto Giorgi
ISPDC5
2008 TFlux: A Portable Platform for Data-Driven Multithreading on Commodity Multicore Systems
abstract
In this paper we present thread flux (TFlux), a complete system that supports the data-driven multithreading (DDM) model of execution. TFlux virtualizes any details of the underlying system therefore offering the same programming model independently of the architecture. To achieve this goal, TFlux has a runtime support that is built on top of a commodity operating system. Scheduling of threads is performed by the thread synchronization unit (TSU), which can be implemented either as a hardware or a software module. In addition, TFlux includes a preprocessor that, along with a set of simple compiler directives, allows the user to easily develop DDM programs. The preprocessor then automatically produces the TFlux code, which can be compiled using any commodity C compiler, therefore automatically producing code to any ISA. TFlux has been validated on three platforms. A Simics-based multicore system with a TSU hardware module (TFluxHard), a commodity 8-core Intel Core2 QuadCore-based system with a software TSU module (TFluxSoft), and a Cell/BE system with a software TSU module (TFluxCell). The experimental results show that the performance achieved is close to linear speedup, on average 21x for the 27 nodes TFluxHard, and 4.4x on a 6 nodes TFluxSoft and TFluxCell. Most importantly, the observed speedup is stable across the different platforms thus allowing the benefits of DDM to be exploited on different commodity systems.
Kyriakos Stavrou, Marios Nikolaides, Demos Pavlou, Samer Arandi, Paraskevas Evripidou, Pedro Trancoso
ICPP5
2007 A Pragmatic Methodology to Web Service Discovery
abstract
As in linguistics the semantics co-exist with pragmatics to provide complete and unambiguous meaning to utterances, in the same way the pragmatics should augment the semantics in intelligent applications. Service-oriented computing is an area where such intelligent applications are greatly facilitated. In this paper we propose a pragmatic methodology to Web service discovery which utilizes both the pragmatics and the semantics. This methodology aims to solve a very basic problem of existing semantic discovery approaches: the inability of selecting the most appropriate service among many semantically equivalent Web services.
Electra Tamani, Paraskevas Evripidou
ICWS2
2006 Applying Trust Mechanisms in an agent-based P2P Network of Service Providers and Requestors
Electra Tamani, Paraskevas Evripidou
CCGRID2
2006 Active Folders: A Metaphor for Developing and Interacting with Context-Aware Applications
abstract
Query by Browsing (QBB) is a relationally complete paradigm for creating Visual Query Language (VQL)s. It has been adapted successfully for use on handheld devices in the form of the Chiromancer interface, thus placing considerable search capabilities at the fingertips of mobile users. In this paper we describe an enhancement to the QBB paradigm to accommodate continuous queries through the use of active folders, i.e. folders whose contents are constantly updated. We thus demonstrate how such continuous queries coupled with a suitable context-describing database schema can be used to support a variety of context-aware application scenarios. A prototype architecture and context schema are described as are the extensions to the QBB paradigm and the additions to the Chiromancer interface needed to support these scenarios.
Stavros Polyviou, George Samaras, Paraskevas Evripidou
MDM3
2006 Data-Driven Multithreading Using Conventional Microprocessors
abstract
This paper describes the data-driven multithreading (DDM) model and how it may be implemented using off-the-shelf microprocessors. Data-driven multithreading is a nonblocking multithreading execution model that tolerates internode latency by scheduling threads for execution based on data availability. Scheduling based on data availability can be used to exploit cache management policies that reduce significantly cache misses. Such policies include firing a thread for execution only if its data is already placed in the cache. We call this cache management policy the CacheFlow policy. The core of the DDM implementation presented is a memory mapped hardware module that is attached directly to the processor's bus. This module is responsible for thread scheduling and is known as the thread synchronization unit (TSU). The evaluation of DDM was performed using simulation of the data-driven network of workstations (D2NOW). D2NOW is a DDM implementation built out of regular workstations augmented with the TSU. The simulation was performed for nine scientific applications, seven of which belong to the SPLASH-2 suite. The results show that DDM can tolerate well both the communication and synchronization latency. Overall, for 16 and 32-node D2NOW machines the speedup observed was 14.4 and 26.0, respectively
Costas Kyriacou, Paraskevas Evripidou, Pedro Trancoso
IEEE Trans. Parallel Distributed Syst.2
2005 CFP Taxonomy of the Approaches for Dynamic Web Content Acceleration
Stavros Papastavrou, George Samaras, Paraskevas Evripidou, Panos K. Chrysanthis
ADBIS3
2005 A Relationally Complete Visual Query Language for Heterogeneous Data Sources and Pervasive Querying
abstract
In this paper we introduce and formally define Query by Browsing (QBB), a scalable, relationally complete visual query language based on the desktop user interface paradigm and tuple relational calculus that allows the formulation of complex queries over relational, entity-relationship, object-oriented and XML data sources on a variety of handheld and desktop platforms. It is to our knowledge the first visual query language to combine the important characteristics of usability, scalability, expressive power and flexibility. We support these claims by demonstrating the similarity of the QBB paradigm to the popular desktop user interface paradigm, by relating it to relational calculus and relational algebra and by describing Chiromancer II, a Web-based implementation of the QBB paradigm for handheld devices. We also discuss ways in which non-relational sources can be represented and queried and compare QBB to related work in the area of visual query languages for a variety of data models. We finally offer conclusions and thoughts for future work.
Stavros Polyviou, George Samaras, Paraskevas Evripidou
ICDE3
2004 Net-dbx-G: a Web-based debugger of MPI programs over Grid environments
abstract
Net-dbx-G is a tool that utilizes Java and other World Wide Web tools as an interface to Grid services to help Grid application developers debug their MPI programs from anywhere in the Internet. Net-dbx-G is a source level debugger with the full power of gdb (the GNU Debugger), providing the user with full Grid functionality. The portability of the tool is of great importance as well because it enables the tool to be used on heterogeneous nodes that participate in an MPI enabled, Grid environment. The users of our system simply point their browser to the Net-dbx-G webpage and authenticate to the virtual organization that they are part of. Having at their disposal the shared resources participating with that VO, they start debugging by interacting with the provided GUI environment of the tool. The users can dynamically select which MPI processes to view/debug.
Panayiotis Neophytou, Neophytos Neophytou, Paraskevas Evripidou
CCGRID3
2004 CacheFlow: A Short-Term Optimal Cache Management Policy for Data Driven Multithreading
Costas Kyriacou, Paraskevas Evripidou, Pedro Trancoso
Euro-Par2
2004 Mobile Agents for Wireless Computing: The Convergence of Wireless Computational Models with Mobile-Agent Technologies
Constantinos Spyrou, George Samaras, Evaggelia Pitoura, Paraskevas Evripidou
Mob. Networks Appl.4
2001 The PaCMAn Metacomputer: parallel computing with Java mobile agents
Paraskevas Evripidou, George Samaras, Christoforos Panayiotou, Evaggelia Pitoura
Future Gener. Comput. Syst.1
2001 D3-Machine: A decoupled data-driven multithreaded architecture with variable resolution support
Paraskevas Evripidou
Parallel Comput.1
2001 Net-dbx: A Web-Based Debugger of MPI Programs Over Low-Bandwidth Lines
abstract
This paper describes Net-dbx, a tool that utilizes Java and other World Wide Web tools for the debugging of MPI programs from anywhere in the Internet. Net-dbx is a source-level interactive debugger with the full power of gdb (the GNU Debugger) augmented with the debug functionality of the public-domain MPI implementation environments. The main effort was on a low overhead, yet powerful, graphical interface supported by low-bandwidth connections. The portability of the tool is of great importance as well because it enables the tool to be used on heterogeneous nodes that participate in an MPI multicomputer. Both needs are satisfied a great deal by the use of WWW browsing tools and the Java programming language. The user of our system simply points his/her browser to the Net-dbx page, logs in to the destination system, and starts debugging by interacting with the tool, just as with any GUI environment. The user can dynamically select which MPI processes to view/debug. A special WWW-based environment has been designed and implemented to host the system prototype.
Neophytos Neophytou, Paraskevas Evripidou
IEEE Trans. Parallel Distributed Syst.2
1998 Exploiting Course Grain Parallelism from FORTRAN by Mapping it to IF1
Adrianos Lachanas, Paraskevas Evripidou
Euro-Par2
1998 Net-dbx: A Java Powered Tool for Interactive Debugging of MPI Programs Across the Internet
Neophytos Neophytou, Paraskevas Evripidou
Euro-Par2
1995 Incorporating Input/Output Operations Into Dynamic Data-Flow Graphs
Paraskevas Evripidou, Jean-Luc Gaudiot
Parallel Comput.1
1995 Advanced Array Optimizations for High Performance Functional Languages
abstract
We discuss and evaluate three optimizations for reducing memory management overhead and data copying costs in SISAL 1.2 programs that build arrays. The first, called framework preconstruction, eliminates superfluous allocate-deallocate sequences in cyclic computations. The second, called aggregate storage subsumption, reduces the management overhead for compound array components. The third, called predictive storage preallocation, eliminates superfluous data copying in filtered array constructions and simplifies their parallelization. We have added all three optimizations to the Optimizing SISAL Compiler with rewarding improvements in SISAL program performance on vector-parallel machines such as those built by Cray Computer Corporation, Convex, and Cray Research.>
David C. Cann, Paraskevas Evripidou
IEEE Trans. Parallel Distributed Syst.2
1993 Block Scheduling of Iterative Algorithms and Graph-Level Priority Scheduling in a Simulated Data-Flow Multiprocessor
abstract
Iterative methods for solving linear systems are discussed. Although these methods are inherently highly sequential, it is shown that much parallelism could be exploited in a data-flow system by scheduling the iterative part of the algorithms in blocks and by looking ahead across several iterations. This approach is general and will apply to other iterative and loop-based problems. It is also demonstrated by simulation that relying solely on data-driven scheduling of parallel and unrolled loops results in low resource utilization and poor performance. A graph-level priority scheduling mechanism has been developed that greatly improves resource utilization and yields higher performance.>
Paraskevas Evripidou, Jean-Luc Gaudiot
IEEE Trans. Parallel Distributed Syst.1
1990 A Decoupled Graph/Computation Data-Driven Architecture with Variable-Resolution Actors
Paraskevas Evripidou, Jean-Luc Gaudiot
ICPP (1)1
1988 Iterative Algorithms in a Data-Driven Environment
Paraskevas Evripidou, Jean-Luc Gaudiot
ICPP (1)1