Anastasios Psarras

dblp:144/4454 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
0since 2021 · last 2020
0000-0001-6151-9242ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 5 first-authorSoftware engineering, systems software and programming languages · 5 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Interconnection networks and networks-on-chip · 42% Cloud and datacenter computing · 26% Integrated circuit design · 15%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Integrated circuit design
digital circuit design
0.312017
A Dual-Clock Multiple-Queue Shared Buffer · IEEE Trans. Computers 2017
Interconnection networks and networks-on-chip › router architecture
shared buffer design
0.312017
A Dual-Clock Multiple-Queue Shared Buffer · IEEE Trans. Computers 2017
Interconnection networks and networks-on-chip › router architecture
network-on-chip router
0.212016
ShortPath: A Network-on-Chip Router with Fine-Grained Pipeline Bypassing · IEEE Trans. Computers 2016
Interconnection networks and networks-on-chip › router architecture
network-on-chip router microarchitecture
0.212016
ShortPath: A Network-on-Chip Router with Fine-Grained Pipeline Bypassing · IEEE Trans. Computers 2016
Cloud and datacenter computing
quality of service
0.212016
PhaseNoC: Versatile Network Traffic Isolation Through TDM-Scheduled Virtual Channels · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2016
Cloud and datacenter computing
traffic isolation
0.212016
PhaseNoC: Versatile Network Traffic Isolation Through TDM-Scheduled Virtual Channels · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2016
Processor architecture and microarchitecture
chip multiprocessor
0.112016
PhaseNoC: Versatile Network Traffic Isolation Through TDM-Scheduled Virtual Channels · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2016
Processor architecture and microarchitecture
many-core architecture
0.112016
PhaseNoC: Versatile Network Traffic Isolation Through TDM-Scheduled Virtual Channels · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2016
Performance modeling and evaluation › simulation › communication system simulation
network simulation
0.112016
ShortPath: A Network-on-Chip Router with Fine-Grained Pipeline Bypassing · IEEE Trans. Computers 2016

Methods — techniques the papers use, named apart from their topics

linked-list organization · 0.3dual-clock design · 0.3standard-cell-based synthesis · 0.2placed-and-routed layout · 0.2opportunistic bandwidth stealing · 0.2TDM scheduling · 0.2
YearPublicationVenuePosition
2020 The Mesochronous Dual-Clock FIFO Buffer
abstract
To increase system composability and facilitate timing closure, fully synchronous clocking is replaced by more relaxed clocking schemes, such as mesochronous clocking. Under this regime, the modules at the two ends of a mesochronous interface receive the same clock signal, thus operating under the same clock frequency, but the edges of the arriving clock signals may exhibit an unknown phase relationship. In such cases, clock synchronization is needed when sending data across modules. In this brief, we present a novel mesochronous dual-clock first-input-first-output (FIFO) buffer that can handle both clock synchronization and temporary data storage, by synchronizing data implicitly through the explicit synchronization of only the flow-control signals. The proposed design can operate correctly even when the transmitter and the receiver are separated by a long link whose delay cannot fit within the target operating frequency. In such scenarios, the proposed mesochronous FIFO can be extended to support multicycle link delays in a modular manner and with minimal modifications to the baseline architecture. When compared with the other state-of-the-art dual-clock mesochronous FIFO designs, the new architecture is demonstrated to yield a substantially lower cost implementation.
Dimitris Konstantinou, Anastasios Psarras, Chrysostomos Nicopoulos, Giorgos Dimitrakopoulos
IEEE Trans. Very Large Scale Integr. Syst.2
2017 A Dual-Clock Multiple-Queue Shared Buffer
abstract
Multiple parallel queues are versatile hardware data structures that are extensively used in modern digital systems. To achieve maximum scalability, the multiple queues are built on top of a dynamically-allocated shared buffer that allocates the buffer space to the various active queues, based on a linked-list organization. This work focuses on dynamically-allocated multiple-queue shared buffers that allow their read and write ports to operate in different clock domains. The proposed dual-clock shared buffer follows a tightly-coupled organization that merges the tasks of signal synchronization across asynchronous clock domains and queueing (buffering), in a common hardware module. When compared to other state-of-the-art dual-clock multiple-queue designs, the new architecture is demonstrated to yield a substantially lower-cost implementation. Specifically, hardware area savings of up to 55 percent are achieved, while still supporting full-throughput operation.
Anastasios Psarras, Michalis Paschou, Chrysostomos Nicopoulos, Giorgos Dimitrakopoulos
IEEE Trans. Computers1
2016 CrossOver: Clock domain crossing under virtual-channel flow control
Michalis Paschou, Anastasios Psarras, Chrysostomos Nicopoulos, Giorgos Dimitrakopoulos
DATE2
2016 A Low-Power Network-on-Chip Architecture for Tile-based Chip Multi-Processors
abstract
Technology scaling of tiled-based CMPs reduces the physical size of each tile and increases the number of tiles per die. This trend directly impacts the on-chip interconnect; even though the tile population increases, the inter-tile link distances scale down proportionally to the tile dimensions. The decreasing inter-tile wire lengths can be exploited to enable swift link traversal between neighboring tiles, after appropriate wire engineering. Building on this premise, we propose a technique to rapidly transfer its between adjacent routers in half a clock cycle, by utilizing both edges of the clock during the sending and receiving operations. Half-cycle link traversal enables, for the first time, substantial reductions in (a) link power, irrespective of the data switching profile, and (b) buffer power (through buffer-size reduction), without incurring any latency/throughput loss. In fact, the proposed architecture also yields some latency improvements over a baseline NoC. Detailed hardware analysis using placed-and-routed designs, and cycle-accurate full-system simulations corroborate the significant power and latency improvements.
Anastasios Psarras, Junghee Lee 0004, Pavlos M. Mattheakis, Chrysostomos Nicopoulos, Giorgos Dimitrakopoulos
ACM Great Lakes Symposium on VLSI1
2016 ShortPath: A Network-on-Chip Router with Fine-Grained Pipeline Bypassing
abstract
Scalable Network-on-Chip (NoC) architectures should achieve high-throughput and low-latency operation without exceeding the stringent area/energy constraints of modern Systems-on-Chip (SoC), even when operating under a high clock frequency. Such requirements directly impact the NoC routers and interfaces comprising the NoC architecture. This paper focuses on the micro-architecture of NoC routers and presents ShortPath, a pipelined router architecture that can achieve high-speed implementations by parallelizing as much as possible - and without resorting to speculation - the allocation steps involved in the operation of a VC-based router. Most importantly, ShortPath is augmented with a fine-grained pipeline bypassing mechanism, which skips all stages without contention and “fast-forwards” the flits to the first point of contention. Pipeline bypassing in ShortPath is always productive, and even if a flit loses in arbitration, it does not repeat any of the stages already bypassed. Extensive network simulations and hardware analysis - using standard-cell-based synthesis and placed-and-routed layout - corroborate the efficiency of ShortPath, in terms of both network performance and hardware complexity, as compared to the most relevant current state-of-the-art architecture.
Anastasios Psarras, Ioannis Seitanidis, Chrysostomos Nicopoulos, Giorgos Dimitrakopoulos
IEEE Trans. Computers1
2016 PhaseNoC: Versatile Network Traffic Isolation Through TDM-Scheduled Virtual Channels
abstract
As multi/many-core architectures evolve, the demands on the network-on-chip (NoC) are amplified. In addition to high performance and physical scalability, the NoC is increasingly required to also provide specialized functionality, such as network virtualization, flow isolation, and quality-of-service. Although traditional architectures supporting virtual channels (VCs) offer the resources for flow partitioning and isolation, an adversarial workload can still interfere and degrade the performance of other workloads that are active in a different set of VCs. In this paper, we present PhaseNoC, a truly noninterfering VC-based architecture that adopts time-division multiplexing at the VC level. Distinct flows, or application domains, mapped to disjoint sets of VCs are isolated, both inside the router's pipeline and at the network level. Any latency overhead is minimized by appropriate scheduling of flows in separate phases of operation, irrespective of the chosen topology. When strict isolation is not required, the proposed architecture can employ opportunistic bandwidth stealing. This novel mechanism works synergistically with the baseline PhaseNoC techniques to improve the overall latency/throughput characteristics of the NoC, while still preserving performance isolation. Experimental results corroborate that-with lower cost than state-of-the-art NoC architectures, and with minimum latency overhead-PhaseNoC removes any flow interference and allows for efficient network traffic isolation.
Anastasios Psarras, Junghee Lee 0004, Ioannis Seitanidis, Chrysostomos Nicopoulos, Giorgos Dimitrakopoulos
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2015 PhaseNoC: TDM scheduling at the virtual-channel level for efficient network traffic isolation
Anastasios Psarras, Ioannis Seitanidis, Chrysostomos Nicopoulos, Giorgos Dimitrakopoulos
DATE1
2015 Timing-resilient Network-on-Chip architectures
abstract
Networks-on-Chip (NoC) have been established as the de facto standard for on-chip communication in multi-/many-core systems, due to their innate scalability properties pertaining to performance and physical implementation. Spanning the entire chip, the NoC suffers from both inter-die and intra-die variations. In addition to static variability, the NoC is also afflected by dynamic variations, such as fast VDD droops, temperature, and aging effects. Such variations cause unpredictable behavior in the timing characteristics of the NoC components, which need significant timing margins to ensure always-correct operation. Consequently, the potential for high-frequency operation is impeded. In this paper, we propose a timing-error-resilient mechanism called TRNoC, which allows the NoC to operate at higher frequencies, at the expense of sporadically experiencing run-time timing errors, which are handled by an error recovery strategy. TRNoC includes a lightweight timing-error detection mechanism, together with a distributed error recovery mechanism, which allow for lossless operation, while leaving the error-free parts of the network unaffected. Hardware implementation results demonstrate the efficiency of TRNoC and its potential as a scalable timing-error-resilient NoC architecture.
Alexandros Panteloukas, Anastasios Psarras, Chrysostomos Nicopoulos, Giorgos Dimitrakopoulos
IOLTS2
2015 ElastiStore: Flexible Elastic Buffering for Virtual-Channel-Based Networks on Chip
abstract
As multicore systems transition to the many-core realm, the pressure on the interconnection network is substantially elevated. The network on chip (NoC) is expected to undertake the expanding demands of the ever-increasing numbers of processing elements, while its area/power footprint remains severely constrained. Hence, low-cost NoC designs that achieve high-throughput and low-latency operation are imperative for future scalability. While the buffers of the NoC routers are key enablers of high performance, they are also major consumers of area and power. In this paper, we extend elastic buffer (EB) architectures to support multiple virtual channels (VCs), and we derive ElastiStore, a novel lightweight EB architecture that minimizes buffering requirements without sacrificing performance. ElastiStore uses just one register per VC and a shared buffer sized large enough to merely cover the round-trip time that appears either on the NoC links or due to the internal pipeline of the NoC routers. The integration of the proposed EB scheme in the NoC router enables the design of efficient architectures, which offer the same performance as baseline VC-based routers, albeit at a significantly lower cost. Cycle-accurate network simulations including both synthetic traffic patterns and real application workloads running in a full-system simulation framework verify the efficacy of the proposed architecture. Moreover, the hardware implementation results using a 45-nm standard-cell library demonstrate ElastiStore's efficiency.
Ioannis Seitanidis, Anastasios Psarras, Kypros Chrysanthou, Chrysostomos Nicopoulos, Giorgos Dimitrakopoulos
IEEE Trans. Very Large Scale Integr. Syst.2
2014 Hardware primitives for the synthesis of multithreaded elastic systems
abstract
Elastic systems operate in a dataflow-like mode using a distributed scalable control and tolerating variable-latency computations. At the same time, multithreading increases the utilization of processing units and hides the latency of each operation by time-multiplexing operations of different threads in the datapath. This paper proposes a model to unify multithreading and elasticity. A new multithreaded elastic control protocol is introduced supported by low-cost elastic buffers that minimize the storage requirements without sacrificing performance. To enable the synthesis of multithreaded elastic architectures, new hardware primitives are proposed and utilized in two circuit examples to prove the applicability of the proposed approach.
Giorgos Dimitrakopoulos, Ioannis Seitanidis, Anastasios Psarras, K. Tsiouris, Pavlos M. Mattheakis, Jordi Cortadella
DATE3
2014 ElastiStore: An elastic buffer architecture for Network-on-Chip routers
abstract
The design of scalable Network-on-Chip (NoC) architectures calls for new implementations that achieve high-throughput and low-latency operation, without exceeding the stringent area-energy constraints of modern Systems-on-Chip (SoC). The router's buffer architecture is a critical design aspect that affects both network-wide performance and implementation characteristics. In this paper, we extend Elastic Buffer (EB) architectures to support multiple Virtual Channels (VC) and we derive ElastiStore, a novel lightweight elastic buffer architecture that minimizes buffering requirements, without sacrificing performance. The integration of the proposed elastic buffering scheme in the NoC router enables the design of new router architectures - both single-cycle and two-stage pipelined - which offer the same performance as baseline VC-based routers, albeit at a significantly lower area/power cost.
Ioannis Seitanidis, Anastasios Psarras, Giorgos Dimitrakopoulos, Chrysostomos Nicopoulos
DATE2
2014 ElastiNoC: A self-testable distributed VC-based Network-on-Chip architecture
abstract
Network-on-Chip (NoC) design tries to keep a balance between network performance and physical implementation flexibility. The adoption of Virtual Channels (VC) holds promise for scalable NoC design. VCs allow for traffic separation and isolation, enable deadlock avoidance and improve network performance. In this paper, we present ElastiNoC, a novel distributed VC-based router architecture that enjoys all the benefits offered by VCs and leads to efficient silicon-aware implementations. The proposed architecture utilizes an efficient buffering strategy and allows for modular pipelined organizations that increase the clock frequency. Moreover, it offers maximum freedom in terms of physical placement, by allowing the NoC components to be physically spread throughout the chip, irrespective of the network topology. The combined effect of all supported features enables significant delay reductions under equal performance, when compared to state-of-the-art VC-based NoC implementations. Moreover, the careful addition of self-test structures allows ElastiNoC to enjoy fully distributed Built-In Self Testability (BIST), where testing unfolds in phases and reaches high fault coverage with small test application time.
Ioannis Seitanidis, Anastasios Psarras, Emmanouil Kalligeros, Chrysostomos Nicopoulos, Giorgos Dimitrakopoulos
NOCS2