Spiridon F. Beldianu

dblp:42/9378 · DBLP profile ↗
← Back
6ranked-venue papers
5as first author
0since 2021 · last 2015
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 first-authorComputer networks · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Energy-efficient computing · 61% Hardware accelerators and domain-specific architectures · 30% Processor architecture and microarchitecture · 9%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Energy-efficient computing
power gating
0.212015
Performance-Energy Optimizations for Shared Vector Accelerators in Multicores · IEEE Trans. Computers 2015
Energy-efficient computing
power management
0.212015
Performance-Energy Optimizations for Shared Vector Accelerators in Multicores · IEEE Trans. Computers 2015
Hardware accelerators and domain-specific architectures › data-parallel accelerator
vector accelerator
0.212015
Performance-Energy Optimizations for Shared Vector Accelerators in Multicores · IEEE Trans. Computers 2015
Processor architecture and microarchitecture
chip multiprocessor
0.112015
Performance-Energy Optimizations for Shared Vector Accelerators in Multicores · IEEE Trans. Computers 2015

Methods — techniques the papers use, named apart from their topics

performance-energy measurement · 0.2FPGA prototyping · 0.2
YearPublicationVenuePosition
2015 Performance-Energy Optimizations for Shared Vector Accelerators in Multicores
abstract
For multicore processors with a private vector coprocessor (VP) per core, VP resources may not be highly utilized due to limited data-level parallelism (DLP) in applications. Also, under low VP utilization static power dominates the total energy consumption. We enhance here our previously proposed VP sharing framework for multicores in order to increase VP utilization while reducing the static energy. We describe two power-gating (PG) techniques to dynamically control the VP's width based on utilization figures. Floating-point results on an FPGA prototype show that the PG techniques reduce the energy needs by 30-35 percent with negligible performance reduction as compared to a multicore with the same amount of hardware resources where, however, each core is attached to a private VP.
Spiridon F. Beldianu, Sotirios G. Ziavras
IEEE Trans. Computers1
2014 ASIC Design of Shared Vector Accelerators for Multicore Processors
abstract
Vector coprocessor (VP) resources are often underutilized due to the lack of sustained DLP (data-level parallelism) or the presence of vector-length variations in application code. Our work is motivated by: a) the omnipresence of vector operations in high-performance scientific and embedded applications, b) the need for performance and energy efficiency, and c) applications that must often handle various vector sizes. Our design for VP sharing in multicores enhances performance while maintaining low area and energy costs. Our 40nm ASIC design yields 16.66 GFLOPs/Watt. Also, a detailed clock and power gating analysis further proves the viability of our approach.
Spiridon F. Beldianu, Sotirios G. Ziavras
SBAC-PAD1
2013 FPGA and ASIC square root designs for high performance and power efficiency
abstract
Floating-point square root is a fundamental operation in signal processing and various HPC applications. Since this is an expensive operation in resource and energy consumption, its efficient implementation should be of priority in future multicores that will face dark silicon issues. This paper presents a low-cost, low-power consumption design to calculate the square root using the IEEE754 single-precision floating-point format. Two versions of the design are investigated with and without clock gating (CG), respectively. Evaluation involves FPGA and ASIC technologies at 40 and 65 nm. Substantial performance growth and reduced power consumption are gained as compared to a popular iterative solution. The ASIC design demonstrates much lower power consumption, which at 40 nm is lower than that at 65 nm by about a threefold. At 40 nm, CG for the ASIC realization is justified primarily for low activity rates.
Shashank Suresh, Spiridon F. Beldianu, Sotirios G. Ziavras
ASAP2
2013 Multicore-based vector coprocessor sharing for performance and energy gains
abstract
For most of the applications that make use of a dedicated vector coprocessor, its resources are not highly utilized due to the lack of sustained data parallelism which often occurs due to vector-length variations in dynamic environments. The motivation of our work stems from: (a) the mandate for multicore designs to make efficient use of on-chip resources for low power and high performance; (b) the omnipresence of vector operations in high-performance scientific and emerging embedded applications; (c) the need to often handle a variety of vector sizes; and (d) vector kernels in application suites may have diverse computation needs. We present a robust design framework for vector coprocessor sharing in multicore environments that maximizes vector unit utilization and performance at substantially reduced energy costs. For our adaptive vector unit, which is attached to multiple cores, we propose three basic shared working policies that enforce coarse-grain, fine-grain, and vector-lane sharing. We benchmark these vector coprocessor sharing policies for a dual-core system and evaluate them using the floating-point performance, resource utilization, and power/energy consumption metrics. Benchmarking for FIR filtering, FFT, matrix multiplication, and LU factorization shows that these coprocessor sharing policies yield high utilization and performance with low energy costs. The proposed policies provide 1.2--2 speedups and reduce the energy needs by about 50% as compared to a system having a single core with an attached vector coprocessor. With the performance expressed in clock cycles, the sharing policies demonstrate 3.62--7.92 speedups compared to optimized Xeon runs. We also introduce performance and empirical power models that can be used by the runtime system to estimate the effectiveness of each policy in a hybrid system that can simultaneously implement this suite of shared coprocessor policies.
Spiridon F. Beldianu, Sotirios G. Ziavras
ACM Trans. Embed. Comput. Syst.1
2011 On-chip Vector Coprocessor Sharing for Multicores
abstract
For most of the applications that make use of a vector coprocessor, the resources are not highly utilized due to the lack of sustained data parallelism, which sometimes occurs due to vector-length changes in dynamic environments. The motivation of our work stems from (a) the mandate for multicore designs to make efficient use of the on-chip resources, (b) the frequent presence of vector operations in high-performance scientific and embedded applications, (c) the increased probability that different cores may deal with different vector lengths at various times, and (d) different vector kernels in the same or different application suites may have diverse computation needs. Our objective is to provide a versatile design framework that can facilitate vector coprocessor sharing among multiple cores in a manner that maximizes resource utilization while also yielding very high performance at reduced cost. We propose three basic shared vector coprocessor architectures for multicores based on coarse-grain, fine-grain and vector lane sharing. We benchmark these distinct vector architectures for a dual-core system using the floating-point performance and resource utilization metrics. Our analysis shows that vector lane sharing, where the number of vector lanes assigned to a core can be controlled dynamically, provides the greatest flexibility and generally yields very good results. Since, however, each of the three design choices has its own performance advantages under certain vector-load conditions, we ultimately suggest a hybrid vector coprocessor design that can support all three architectural choices as per the core and application collective needs.
Spiridon F. Beldianu, Sotirios G. Ziavras
PDP1
2009 Re-Configurable Parallel Match Evaluators Applied to Scheduling Schemes for Input-Queued Packet Switches
abstract
The performance of matching schemes for input- queued (IQ) packet switches is mainly defined by the selection policy adopted. This policy can be aimed to produce a large weight sum for matched input-output pairs, where each input- output pair is assigned a weight, or to produce a large match size in the number of matched pairs, giving place to maximum weight matching or maximum size matching, respectively. However, schedulers can only provide a single match in function of the selection (of candidate ports) policy adopted and of the backlogged traffic at the input queues. A parallel match evaluator was recently proposed to provide not one but several match options at the same time. This approach evaluates several predefined and fixed matches and picks the match with the largest size. However, the fixed permutations of the evaluated matches may produce low performance under traffic with nonuniform distributions because of the limited number of choices. This paper proposes to make the parallel match evaluator configurable and two schemes to provide diverse and changeable matches such that the matches (and therefore, the evaluator) become adaptable to the traffic pattern. The proposed schemes were tested under uniform and nonuniform traffic patterns and the results show that these schemes provide high performance, even when scheduling is performed between periods of multiple time slots, or framed intervals. The proposed approach can be used for configuring slow micro-electro-mechanical (MEM) optical switch fabrics.
Spiridon F. Beldianu, Roberto Rojas-Cessa, Eiji Oki, Sotirios G. Ziavras
ICCCN1