Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Toni Juan

dblp:48/3704 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
0since 2021 · last 2009
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 4 first-authorSoftware engineering, systems software and programming languages · 4 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Processor architecture and microarchitecture · 48% Parallel and multicore computing · 18% GPUs and heterogeneous computing · 17%

Topics — the 12 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture
many-core architecture
0.112008
Larrabee: a many-core x86 architecture for visual computing · ACM Trans. Graph. 2008
Processor architecture and microarchitecture
branch prediction
0.021998
Dataflow Analysis of Branch Mispredictions and Its Application to Early Resolution of Branch Outcomes · MICRO 1998
Dynamic History-length Fitting: A Third Level of Adaptivity for Branch Prediction · ISCA 1998
Processor architecture and microarchitecture › instruction set architecture
vector extension
0.012002
Tarantula: A Vector Extension to the Alpha Architecture · ISCA 2002
Processor architecture and microarchitecture
vector processor
0.012002
Tarantula: A Vector Extension to the Alpha Architecture · ISCA 2002
Processor architecture and microarchitecture › branch prediction
two-level adaptive branch prediction
0.011998
Dynamic History-length Fitting: A Third Level of Adaptivity for Branch Prediction · ISCA 1998
Processor architecture and microarchitecture
value prediction
0.011998
Dataflow Analysis of Branch Mispredictions and Its Application to Early Resolution of Branch Outcomes · MICRO 1998
Memory systems
cache
0.011996
The Difference-bit Cache · ISCA 1996
Memory systems › memory access latency
cache access latency
0.011996
The Difference-bit Cache · ISCA 1996
Memory systems › cache
cache organization
0.011996
The Difference-bit Cache · ISCA 1996
Memory systems › cache › cache organization
set-associative cache
0.011996
The Difference-bit Cache · ISCA 1996
High-performance computing
scientific computing systems
0.012002
Tarantula: A Vector Extension to the Alpha Architecture · ISCA 2002
Parallel and multicore computing
dataflow computing
0.011998
Dataflow Analysis of Branch Mispredictions and Its Application to Early Resolution of Branch Outcomes · MICRO 1998

Methods — techniques the papers use, named apart from their topics

vector processing · 0.1software task scheduling · 0.1binning · 0.1simulation · 0.0trace-driven simulation · 0.0hardware mechanism design · 0.0dataflow analysis · 0.0access time modeling · 0.0
YearPublicationVenuePosition
2009 Reusing cached schedules in an out-of-order processor with in-order issue logic
abstract
The complex and powerful out-of-order issue logic dismisses the repetitive nature of the code, unlike what caches or branch predictors do. We show that 90% of the cycles, the group of instructions selected by the issue logic belongs to just 13% of the total different groups issued: the issue logic of an out-of-order processor is constantly re-discovering what it has already found. To benefit from the repetitive nature of instruction issue, we move the scheduling logic after the commit stage, out of the critical path of execution. The schedules created there are cached and reused to feed a simple in-order issue logic, that could result in a higher frequency design. We present the complete design of our ReLaSch processor, that achieves the same average IPC than a conventional out-of-order processor, and a 1.56 speed-up over the IPC of an in-order processor. We actually surpass the out-of-order IPC in 23 out of 40 SPEC benchmarks, mainly because the broader vision of the code after the commit stage allows creating better schedules.
Oscar Palomar, Toni Juan, Juan J. Navarro
ICCD2
2008 Larrabee: a many-core x86 architecture for visual computing
abstract
This paper presents a many-core visual computing architecture code named Larrabee, a new software rendering pipeline, a manycore programming model, and performance analysis for several applications. Larrabee uses multiple in-order x86 CPU cores that are augmented by a wide vector processor unit, as well as some fixed function logic blocks. This provides dramatically higher performance per watt and per unit of area than out-of-order CPUs on highly parallel workloads. It also greatly increases the flexibility and programmability of the architecture as compared to standard GPUs. A coherent on-die 2 nd level cache allows efficient inter-processor communication and high-bandwidth local data access by CPU cores. Task scheduling is performed entirely with software in Larrabee, rather than in fixed function logic. The customizable software graphics rendering pipeline for this architecture uses binning in order to reduce required memory bandwidth, minimize lock contention, and increase opportunities for parallelism relative to standard GPUs. The Larrabee native programming model supports a variety of highly parallel applications that use irregular data structures. Performance analysis on those applications demonstrates Larrabee's potential for a broad range of parallel computation.
Larry Seiler, Doug Carmean, Eric Sprangle, Tom Forsyth, Michael Abrash, Pradeep Dubey, Stephen Junkins, Adam T. Lake, Jeremy Sugerman, Robert Cavin, Roger Espasa, Ed Grochowski, Toni Juan, Pat Hanrahan
ACM Trans. Graph.13
2002 Tarantula: A Vector Extension to the Alpha Architecture
abstract
Tarantula is an aggressive floating point machine targeted at technical, scientific and bioinformatics workloads, originally planned as a follow-on candidate to the EV8 processor. Tarantula adds to the EV8 core a vector unit capable of 32 double-precision flops per cycle. The vector unit fetches data directly from a 16 MByte second level cache with a peak bandwidth of sixty four 64-bit values per cycle. The whole chip is backed by a memory controller capable of delivering over 64 GBytes/s of raw bandwidth. Tarantula extends the Alpha ISA with new vector instructions that operate on new architectural state. Salient features of the architecture and implementation are: (1) it fully integrates into a virtual-memory cache-coherent system without changes to its coherency protocol, (2) provides high bandwidth for non-unit stride memory accesses, (3) supports gather/scatter instructions efficiently, (4) fully integrates with the EV8 core with a narrow, streamlined interface, rather than acting as a co-processor (5) can achieve a peak of 104 operations per cycle, and (6) achieves excellent "real-computation" per transistor and per watt ratios. Our detailed simulations show that Tarantula achieves an average speedup of 5X over EV8, out of a peak speedup in terms of flops of 8X. Furthermore, performance on gather/scatter intensive benchmarks such as Radix Sort is also remarkable: a speedup of almost 3X over EV8 and 15 sustained operations per cycle. Several benchmarks exceed 20 operations per cycle.
Roger Espasa, Federico Ardanaz, Julio Gago, Roger Gramunt, Isaac Hernandez, Toni Juan, Joel S. Emer, Stephen Felix, P. Geoffrey Lowney, Matthew Mattina, André Seznec
ISCA6
2001 How to compare the performance of two SMT microarchitectures
abstract
In this paper we discuss methods and metrics for comparing the performance of two simultaneous multithreading microarchitectures. We identify conditions under which the instructions-per-cycle metric may be misleading for comparing two simultaneous multithreading microarchitectures for the same amount of work. Part of the problem is isolated to the definition of what is same work. When simulating a mix of independent programs under the same initial conditions on two different simultaneous multithreading microarchitectures there are two approaches to ensure the work of the two runs is same: constant-work-per-thread or variablework-per-thread. For both approaches the total number of instructions in the run is constant, however, for the first, the instructions from each thread is also constant, whereas for the second is not. We claim that: (a) when simulating two microarchitectures with the constant-work-per-thread approach, the instructions-percycle is sufficient to compare them to establish the microarchitecture with the best performance, (b) when variable-work-per-thread approach is used the instruction-per-cycle may be inadequate for comparing performance. We attribute this to the inability of the instructions-per-cycle metric to account for differences in the load-balance of the two runs. A new performance metric,SMT-speedup, is proposed that enables accurate comparison of the performance of two simultaneous multithreading microarchitectures for runs with different load-balance. The new metric considers the loadbalance in terms of the size and performance of each thread. In light of the insight gain in this paper we contend that a simultaneous multithreading microarchitecture may need to trade-off throughput and load-balance to achieve the best performance.
Yiannakis Sazeides, Toni Juan
ISPASS2
1998 Dynamic History-length Fitting: A Third Level of Adaptivity for Branch Prediction
abstract
Accurate branch prediction is essential for obtaining high performance in pipelined superscalar processors that execute instructions speculatively. Some of the best current predictors combine a part of the branch address with a fixed amount of global history of branch outcomes in order to make a prediction. These predictors cannot perform uniformly well across all workloads because the best amount of history to be used depends on the code, the input data and the frequency of context switches. Consequently, all predictors that use a fixed history length are therefore unable to perform up to their maximum potential. We introduce a method-called DHLF-that dynamically determines the optimum history length during execution, adapting to the specific requirements of any code, input data and system workload. Our proposal adds an extra level of adaptivity to two-level adaptive branch predictors. The DHLF method can be applied to any one of the predictors that combine global branch history with the branch address. We apply the DHLF method to gshare (dhlf-gshare) and obtain near-optimal results for all SPECint95 benchmarks, with and without context switches. Some results are also presented for gskewed (dhlf-gskewed), confirming that other predictors can benefit from our proposal.
Toni Juan, Kana Sanjeevan, Juan J. Navarro
ISCA1
1998 Dataflow Analysis of Branch Mispredictions and Its Application to Early Resolution of Branch Outcomes
abstract
The goal of this study is twofold: to analyze in detail the nature of conditional branch mispredictions in correlation based branch predictors, and, based on this analysis, to reduce the impact of branch mispredictions on processor performance by decreasing the branch resolution delay instead of improving the branch prediction accuracy. We classify conditional branches with the highest number of mispredictions according to the nature of their branch condition analytical expression. Based on these expressions, we can analyze and even precisely explain the origin of mispredictions in many cases. Moreover we find that many such branches belong to small sets of blocks inside loops, and within such sets we find that some of the branch expressions have regularity properties. We show how to exploit this regularity property by anticipating the branch outcome, where anticipation is a combination of value prediction and normal dataflow execution. We investigate a hardware mechanism to implement the concept of branch outcome anticipation. This mechanism relies on the separate execution of the normal program flow and a branch flow, which is a subset of the program flow corresponding to copies of the instructions needed to compute branch outcomes. The branch flow uses the regularity properties of branch condition expressions to get ahead of the normal program flow whenever possible. Currently, the mechanism can only target a subset of the conditional branches, but with these branches we experimentally show that the anticipation mechanism successfully reduces the average branch misprediction latency by 60%.
Alexandre Farcy, Olivier Temam, Roger Espasa, Toni Juan
MICRO4
1997 Data Caches for Superscalar Processors
abstract
As the number of instructions executed in parallel increases, superscalar processors will require higher bandwidth from data caches.Because of the high cost of true multi-ported caches, alternative cache designs must be evaluated.The purpose of this study is to examine the data cache bandwidth requirements of high-degree superscalar processors, and investigate alternative solutions.The designs studied range from classic solutions like multi-banked caches to more complex solutions recently proposed in the literature.The performance tradeoffs of these different cache designs are examined in details.Then, using a chip area cost model, all solutions are compared with respect to both cost and performance.While many cache designs seem capable of achieving high cache bandwidth, the best cost/performance t,radeoff varies significantly depending on the dedicated area cost, ranging from multi-banked cache designs to hybrid multi-banked/multi-ported caches or even true multi-ported caches.For instance, we find that an 8-bank cache with minor optimizations perform 10% better than a true a-port cache at half the cost, or that a 4-bank 2 ports per bank cache performs better than a true 4-port cache and uses 45% less chip area.
Toni Juan, Juan J. Navarro, Olivier Temam
International Conference on Supercomputing1
1997 Reducing TLB power requirements
abstract
Translation look-aside buffers (TLBs) are small caches to speed-up address translation in processors with virtual memory.This paper considers two issues: (1) a comparison of the power consumption of fully-associative, set-associative, and direct mapped TLBs for the same miss rate and (2) the proposal of modifications of the basic cells and of the structure of set-associative TLBs to reduce the power.The power evaluation is done using a model and the miss rates are obtained from simulations of the SPEC92 benchmark.With respect to ( 1) we conclude that for small TLBs (high miss rates) fully-associative TLBs consume less power but for larger TLBs (low miss rates) set-associative TLBs are better.Moreover, the proposed modifications produce significant reductions in power consumption.Our evaluations show a reduction of 40 to 60% compared to the best traditional TLB.The proposed TLB implementation produces an increase in delay and in area.However, these increases are tolerable because the cycle time is determined by the slower cache and because the TLB area corresponds to only a small portion of the chip area.196 __ __, ~-~-,-, 'II -F 'r-pT-yy-,~';ri '-I:.* : ,< 7,
Toni Juan, Tomás Lang, Juan J. Navarro
ISLPED1
1996 Block Algorithms for Sparse Matrix Computations on High Performance Workstations
abstract
In this paper we analyze the use of Blocklng (tiling), Data Precopying and Software Pipelining to improve the performance of sparse matrix computations on superscalar workstations.In particular, we analyze the case of the Sparse Matrix by dense Matrix operation.The analysis focusses on the practical aspects that can be observed when programming such problem on present workstations with several memory levels.The problem is studied on the Alpha 21064 based workstation DEC 3000/800.Simulations of the memory hierarchyare also used to understand the behaviour of the algorithms.The results obtained show that there is a clear difference between the dense case and the sparse case in terms of the compromises to be adopted to optimize the algorithms.The analysis can be of interest to numerical library and compiler designers.1
Juan J. Navarro, Elena García-Diego, Josep Lluís Larriba-Pey, Toni Juan
International Conference on Supercomputing4
1996 The Difference-bit Cache
abstract
The difference-bit cache is a two-way set-associative cache with an access time that is smaller than that of a conventional one and close or equal to that of a direct-mapped cache. This is achieved by noticing that the two tags for a set have to differ at least by one bit and by using this bit to select the way. In contrast with previous approaches that predict the way and have two types of hits (primary of one cycle and secondary of two to four cycles), all hits of the difference-bit cache are of one cycle. The evaluation of the access time of our cache organization has been performed using a recently proposed on-chip cache access model.
Toni Juan, Tomás Lang, Juan J. Navarro
ISCA1
1994 MOB forms: a class of multilevel block algorithms for dense linear algebra operations
abstract
Multilevel block algorithms exploit the data locality in linear algebra operations when executed in machines with several levels in the memory hierarchy. It is shown that the family we call Multilevel Orthogonal Block (MOB) algorithms is optimal and easy to design and that using the multilevel approach produces significant performance improvements. The effect of interference in the cache, of the TLB misses, and of page faults are also considered. The multilevel block algorithms are evaluated analytically for an ideal memory system with M cache levels without interferences. Moreover, experimental results of the MOB forms in some present high performance workstations are presented.
Juan J. Navarro, Toni Juan, Tomás Lang
International Conference on Supercomputing2