F. Jesús Sánchez

dblp:43/6635 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
0since 2021 · last 2011
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 15 · 5 first-authorSoftware engineering, systems software and programming languages · 4

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
9 papers
Processor architecture and microarchitecture · 70% Memory systems · 16% Reconfigurable computing and FPGAs · 10%
Software engineering, system software, and programming languages
8 papers
Compilers and program optimization · 100%

Topics — the 17 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture
instruction scheduling
0.232009
AGAMOS: A Graph-Based Approach to Modulo Scheduling for Clustered Microarchitectures · IEEE Trans. Computers 2009
Effective instruction scheduling techniques for an interleaved cache clustered VLIW processor · MICRO 2002
Graph-partitioning based instruction scheduling for clustered processors · MICRO 2001
Processor architecture and microarchitecture › instruction-level parallelism › VLIW
clustered VLIW
0.242005
Distributed Data Cache Designs for Clustered VLIW Processors · IEEE Trans. Computers 2005
Flexible Compiler-Managed L0 Buffers for Clustered VLIW Processors · MICRO 2003
Effective instruction scheduling techniques for an interleaved cache clustered VLIW processor · MICRO 2002
Processor architecture and microarchitecture › instruction-level parallelism › VLIW
VLIW processor
0.132005
Distributed Data Cache Designs for Clustered VLIW Processors · IEEE Trans. Computers 2005
Flexible Compiler-Managed L0 Buffers for Clustered VLIW Processors · MICRO 2003
Effective instruction scheduling techniques for an interleaved cache clustered VLIW processor · MICRO 2002
Processor architecture and microarchitecture › clustered architecture
clustered microarchitecture
0.122009
AGAMOS: A Graph-Based Approach to Modulo Scheduling for Clustered Microarchitectures · IEEE Trans. Computers 2009
Graph-partitioning based instruction scheduling for clustered processors · MICRO 2001
Reconfigurable computing and FPGAs
modulo scheduling
0.122009
AGAMOS: A Graph-Based Approach to Modulo Scheduling for Clustered Microarchitectures · IEEE Trans. Computers 2009
Modulo scheduling for a fully-distributed clustered VLIW architecture · MICRO 2000
Processor architecture and microarchitecture › multithreading
speculative multithreading
0.112008
Mitosis: A Speculative Multithreaded Processor Based on Precomputation Slices · IEEE Trans. Parallel Distributed Syst. 2008
Compilers and program optimization
instruction scheduling
0.142003
Graph-partitioning based instruction scheduling for clustered processors · MICRO 2001
Flexible Compiler-Managed L0 Buffers for Clustered VLIW Processors · MICRO 2003
Effective instruction scheduling techniques for an interleaved cache clustered VLIW processor · MICRO 2002
Compilers and program optimization › parallelization
thread-level speculation
0.112005
Mitosis compiler: an infrastructure for speculative threading based on pre-computation slices · PLDI 2005
Memory systems
cache coherence
0.112005
Distributed Data Cache Designs for Clustered VLIW Processors · IEEE Trans. Computers 2005
Memory systems
cache design
0.112005
Distributed Data Cache Designs for Clustered VLIW Processors · IEEE Trans. Computers 2005
Distributed systems
distributed caching
0.112005
Distributed Data Cache Designs for Clustered VLIW Processors · IEEE Trans. Computers 2005
Memory systems › cache coherence › cache coherence protocol
snoopy coherence
0.112005
Distributed Data Cache Designs for Clustered VLIW Processors · IEEE Trans. Computers 2005
Processor architecture and microarchitecture
speculative execution
0.112005
Mitosis compiler: an infrastructure for speculative threading based on pre-computation slices · PLDI 2005
Processor architecture and microarchitecture › instruction-level parallelism
VLIW
0.022000
Modulo scheduling for a fully-distributed clustered VLIW architecture · MICRO 2000
Cache Sensitive Modulo Scheduling · MICRO 1997
Compilers and program optimization › loop transformation
loop scheduling
0.012009
AGAMOS: A Graph-Based Approach to Modulo Scheduling for Clustered Microarchitectures · IEEE Trans. Computers 2009
Memory systems
cache
0.022003
Flexible Compiler-Managed L0 Buffers for Clustered VLIW Processors · MICRO 2003
Effective instruction scheduling techniques for an interleaved cache clustered VLIW processor · MICRO 2002
Compilers and program optimization › instruction scheduling
software pipelining
0.011997
Cache Sensitive Modulo Scheduling · MICRO 1997

Methods — techniques the papers use, named apart from their topics

pseudoschedules · 0.2multilevel graph partitioning · 0.2instruction replication · 0.2speculation · 0.2hardware-software co-design · 0.2instruction scheduling · 0.1pre-computation slices · 0.1modulo scheduling · 0.1compiler-managed memory · 0.1attraction buffers · 0.1
YearPublicationVenuePosition
2011 Global productiveness propagation: a code optimization technique to speculatively prune useless narrow computations
abstract
This paper proposes a unique hardware-software collaborative strategy to remove useless work at 16-bit data-width granularity. The underlying motivation is to design a low power execution platform by exploiting 'narrow' computations. The proposal uses a strictly narrow bit-wide microarchitecture (16-bit integer datapath), which realizes the goal of a low cost, low hardware complexity, low power execution engine. Software dynamically maps the 64-bit computations by translating them into an equivalent 16-bit instruction stream and optimizing them.
Indu Bhagat, Enric Gibert, F. Jesús Sánchez, Antonio González 0001
LCTES3
2009 AGAMOS: A Graph-Based Approach to Modulo Scheduling for Clustered Microarchitectures
abstract
This paper presents AGAMOS, a technique to modulo schedule loops on clustered microarchitectures. The proposed scheme uses a multilevel graph partitioning strategy to distribute the workload among clusters and reduces the number of intercluster communications at the same time. Partitioning is guided by approximate schedules (i.e., pseudoschedules), which take into account all of the constraints that influence the final schedule. To further reduce the number of intercluster communications, heuristics for instruction replication are included. The proposed scheme is evaluated using the SPECfp95 programs. The described scheme outperforms a state-of-the-art scheduler for all programs and different cluster configurations. For some configurations, the speedup obtained when using this new scheme is greater than 40 percent, and for selected programs, performance can be more than doubled.
Alex Aletà, Josep M. Codina, F. Jesús Sánchez, Antonio González 0001, David R. Kaeli
IEEE Trans. Computers3
2008 Mitosis: A Speculative Multithreaded Processor Based on Precomputation Slices
abstract
This paper presents the Mitosis framework, which is a combined hardware-software approach to speculative multithreading, even in the presence of frequent dependences among threads. Speculative multithreading increases single-threaded application performance by exploiting thread-level parallelism speculatively - that is, executing code in parallel even when the compiler or runtime system cannot guarantee the parallelism exists. The proposed approach is based on predicting/computing thread input values via software, through a piece of code that is added at the beginning of each thread (the pre-computation slice). A pre-computation slice is expected to compute the correct thread input values most of the time, but not necessarily always. This allows aggressive optimization techniques to be applied to the slice to make it very short. This paper focuses on the microarchitecture that supports this execution model. The primary novelty of the microarchitecture is the hardware support for the execution and validation of pre-computation slices. Additionally, this paper presents new architectures for the register file and the cache memory in order to support multiple versions of each variable and allow for efficient roll-back in case of misspeculation. We show that the proposed microarchitecture, together with the compiler support, achieves an average speedup of 2.2 for applications that conventional non-speculative approaches are not able to parallelize at all.
Carlos Madriles, Carlos García Quiñones, F. Jesús Sánchez, Pedro Marcuello, Antonio González 0001, Dean M. Tullsen, Hong Wang 0003, John Paul Shen
IEEE Trans. Parallel Distributed Syst.3
2007 Virtual Cluster Scheduling Through the Scheduling Graph
abstract
This paper presents an instruction scheduling and cluster assignment approach for clustered processors. The proposed technique makes use of a novel representation named the scheduling graph which describes all possible schedules. A powerful deduction process is applied to this graph, reducing at each step the set of possible schedules. In contrast to traditional list scheduling techniques, the proposed scheme tries to establish relations among instructions rather than assigning each instruction to a particular cycle. The main advantage is that wrong or poor schedules can be anticipated and discarded earlier. In addition, cluster assignment of instructions is performed using another novel concept called virtual clusters, which define sets of instructions that must execute in the same cluster. These clusters are managed during the deduction process to identify incompatibilities among instructions. The mapping of virtual to physical clusters is postponed until the scheduling of the instructions has finalized. The advantages this novel approach features include: (1) accurate scheduling information when assigning, and, (2) accurate information of the cluster assignment constraints imposed by scheduling decisions. We have implemented and evaluated the proposed scheme with superblocks extracted from Speclnt95 and MediaBench. The results show that this approach produces better schedules than the previous state-of-the-art. Speed-ups are up to 15%, with average speed-ups ranging from 2.5% (2-Clusters) to 9.5% (4-Clusters)
Josep M. Codina, F. Jesús Sánchez, Antonio González 0001
CGO2
2006 Instruction scheduling for a clustered VLIW processor with a word-interleaved cache
abstract
Abstract Clustering is a common technique to overcome the wire delay problem incurred by the evolution of technology. Fully distributed architectures, where the register file, the functional units and the data cache are partitioned, are particularly effective to deal with these constraints and moreover they are very scalable. In this paper, effective instruction scheduling techniques for a word‐interleaved cache clustered VLIW processor are presented. Such scheduling techniques rely on (i) loop unrolling and variable alignment to increase the fraction of local accesses, (ii) a latency assignment process to schedule memory instructions with an appropriate latency, and (iii) different heuristics to assign memory instructions to clusters. Memory consistency is guaranteed by constraining the assignment of memory instructions to clusters. In addition, the use of Attraction Buffers is also introduced. An Attraction Buffer is a hardware mechanism that allows some data replication in order to increase the number of local accesses and, in consequence, reduces stall time. Performance results for the Mediabench benchmark suite demonstrate the effectiveness of the presented techniques and mechanisms. The number of local accesses is increased by more than 25% by using the mentioned scheduling techniques, while stall time is reduced by more than 30% when Attraction Buffers are used. Finally, IPC results for such an architecture are 10% and 5% better compared to those of a clustered VLIW processor with a centralized/unified data cache depending on the scheduling heuristic, respectively. Copyright © 2006 John Wiley & Sons, Ltd.
Enric Gibert, F. Jesús Sánchez, Antonio González 0001
Concurr. Comput. Pract. Exp.2
2005 Mitosis compiler: an infrastructure for speculative threading based on pre-computation slices
Carlos García Quiñones, Carlos Madriles, F. Jesús Sánchez, Pedro Marcuello, Antonio González 0001, Dean M. Tullsen
PLDI3
2005 Distributed Data Cache Designs for Clustered VLIW Processors
abstract
Wire delays are a major concern for current and forthcoming processors. One approach to deal with this problem is to divide the processor into semi-independent units referred to as clusters. A cluster usually consists of a local register file and a subset of the functional units, while the L1 data cache typically remains centralized in What we call partially distributed architectures. However, as technology evolves, the relative latency of such a centralized cache will increase, leading to an important impact on performance. In this paper, we propose partitioning the L1 data cache among clusters for clustered VLIW processors. We refer to this kind of design as fully distributed processors. In particular; we propose and evaluate three different configurations: a snoop-based cache coherence scheme, a word-interleaved cache, and flexible LO-buffers managed by the compiler. For each alternative, instruction scheduling techniques targeted to cyclic code are developed. Results for the Mediabench suite'show that the performance of such fully distributed architectures is always better than the performance of a partially distributed one with the same amount of resources. In addition, the key aspects of each fully distributed configuration are explored.
Enric Gibert, F. Jesús Sánchez, Antonio González 0001
IEEE Trans. Computers2
2003 Local Scheduling Techniques for Memory Coherence in a Clustered VLIW Processor with a Distributed Data Cache
abstract
Clustering is a common technique to deal with wire delays. Fully-distributed architectures, where the register file, the functional units and the cache memory are partitioned, are particularly effective to deal with these constraints and besides they are very scalable. However the distribution of the data cache introduces a new problem: memory instructions may reach the cache in an order different to the sequential program order, thus possibly violating its contents. In this paper two local scheduling mechanisms that guarantee the serialization of aliased memory instructions are proposed and evaluated: the construction of memory dependent chains (MDC solution), and two transformations (store replication and load-store synchronization) applied to the original data dependence graph (DDGT solution). These solutions do not require any extra hardware. The proposed scheduling techniques are evaluated for a word-interleaved cache clustered VLIW processor (although these techniques can also be used for any other distributed cache configuration). Results for the Mediabench benchmark suite demonstrate the effectiveness of such techniques. In particular, the DDGT solution increases the proportion of local accesses by 16% compared to MDC, and stall time is reduced by 32% since load instructions can be freely scheduled in any cluster However the MDC solution reduces compute time and it often outperforms the former. Finally the impact of both techniques on an architecture with attraction buffers is studied and evaluated.
Enric Gibert, F. Jesús Sánchez, Antonio González 0001
CGO2
2003 Flexible Compiler-Managed L0 Buffers for Clustered VLIW Processors
abstract
Wire delays are a major concern for current and forthcoming processors. One approach to attack this problem is to divide the processor into semi-independent units referred to as clusters. A cluster usually consists of a local register file and a subset of the functional units, while the data cache remains centralized. However, as technology evolves, the latency of such a centralized cache increase leading to an important performance impact. In this paper, we propose to include flexible low-latency buffers in each cluster in order to reduce the performance impact of higher cache latencies. The reduced number of entries in each buffer permits the design of flexible ways to map data from L1 to these buffers. The proposed L0 buffers are managed by the compiler, which is responsible to decide which memory instructions make us of them. Effective instruction scheduling techniques are proposed to generate code that exploits these buffers. Results for the Mediabench benchmark suite show that the performance of a clustered VLIW processor with a unified L1 data cache is improved by 16% when such buffers are used. In addition, the proposed architecture also shows significant advantages over both MultiVLIW processors and clustered processors with a word-interleaved cache, two state-of-the-art designs with a distributed L1 data cache.
Enric Gibert, F. Jesús Sánchez, Antonio González 0001
MICRO2
2002 An interleaved cache clustered VLIW processor
abstract
Clustered microarchitectures are becoming a common organization due to their potential to reduce the penalties caused by wire delays and power consumption. Fully-distributed architectures are particularly effective to deal with these constraints, and besides they are very scalable. However, the distribution of the data cache memory poses a significant challenge and may be critical for performance. In this work, a distributed data cache VLIW architecture based on an interleaved cache organization along with cyclic scheduling techniques are proposed. Moreover, the use of Attraction Buffers for such an architecture is introduced. Attraction Buffers are a novel hardware mechanism to increase the percentage of local accesses. The idea is to allow the movement of some data towards the clusters that need it.Performance results for 9 Mediabench benchmarks show that our scheduling techniques are able to hide the increased memory latency when accessing data mapped in a remote cluster. In addition, the local hit ratio is increased by 15% and stall time is reduced by 30% when using the same scheduling techniques with an interleaved cache clustered processor with Attraction Buffers. Finally, the proposed architecture is compared with a state-of-the-art distributed architecture such as the multiVLIW. Results show that the performance of an interleaved cache clustered VLIW processor with Attraction Buffers is similar to that of the multiVLIW architecture, whereas the former has a lower hardware complexity.
Enric Gibert, F. Jesús Sánchez, Antonio González 0001
ICS2
2002 Effective instruction scheduling techniques for an interleaved cache clustered VLIW processor
abstract
Clustering is a common technique to overcome the wire delay problem incurred by the evolution of technology. Fully-distributed architectures, where the register file, the functional units and the data cache are partitioned, are particularly effective to deal with these constraints and besides they are very scalable. In this paper effective instruction scheduling techniques for a clustered VLIW processor with a word-interleaved cache are proposed Such scheduling techniques rely on: (i) loop unrolling and variable alignment to increase the percentage of local accesses, (ii) a latency assignment process to schedule memory operations with an appropriate latency and (iii) different heuristics to assign instructions to clusters. In particular, the number of local accesses is increased by more than 25% if these techniques are used and the ratio of stall time over compute time is small. Next, the main source of remote accesses and stall time is investigated. Stall time is mainly due to remote hits, and Attraction Buffers are used to increase local accesses and reduce stall time. Stall time is reduced by 29% and 34% depending on the scheduling heuristic. IPC results for a word-interleaved cache clustered VLIW processor are similar to those of the multiVLIW (a cache-coherent clustered processor with a more complex hardware design), and are 10% and 5% better (depending on the scheduling heuristic) than the IPC for a clustered processor with a unified cache.
Enric Gibert, F. Jesús Sánchez, Antonio González 0001
MICRO2
2001 Graph-partitioning based instruction scheduling for clustered processors
abstract
This paper presents a novel scheme to schedule loops for clustered microarchitectures. The scheme is based on a preliminary cluster assignment phase implemented through graph partitioning techniques followed by a scheduling phase that integrates register allocation and spill code generation. The graph partitioning scheme is shown to be very effective due to its global view of the whole code while the partition is generated. Results show a significant speedup when compared with previously proposed techniques. For some processor configuration the average speedup for the SPECfp95 is 23% with respect to the published scheme with the best performance. Besides, the proposed scheme is much faster (between 2-7 times, depending on the configuration).
Alex Aletà, Josep M. Codina, F. Jesús Sánchez, Antonio González 0001
MICRO3
2000 The Effectiveness of Loop Unrolling for Modulo Scheduling in Clustered VLIW Architectures
abstract
Clustered organizations are becoming a common trend in the design of VLIW architectures. In this work we propose a novel modulo scheduling approach for such architectures. The proposed technique performs the cluster assignment and the instruction scheduling in a single pass, which is shown to be more effective than doing first the assignment and later the scheduling. We also show that loop unrolling significantly enhances the performance of the proposed scheduler especially when the communication channel among clusters is the main performance bottleneck. By selectively unrolling some loops, we can obtain the best performance with the minimum increase in code size. Performance evaluation for the SPECfp95 shows that the clustered architecture achieves about the same IPC (Instructions Per Cycle) as a unified architecture with the same resources. Moreover when the cycle time is taken into account, a 4-cluster configurations is 3.6 times faster than the unified architecture.
F. Jesús Sánchez, Antonio González 0001
ICPP1
2000 Modulo scheduling for a fully-distributed clustered VLIW architecture
abstract
Clustering is an approach that many microprocessors are adopting in recent times in order to mitigate the increasing penalties of wire delays. We propose a novel clustered VLIW architecture which has all its resources partitioned among clusters, including the cache memory. A modulo scheduling scheme for this architecture is also proposed. This algorithm takes into account both register and memory inter-cluster communications so that the final schedule results in a cluster assignment that favors cluster locality in cache references and register accesses. It has been evaluated for both 2- and 4-cluster configurations and for differing numbers and latencies of inter-cluster buses. The proposed algorithm produces schedules with very low communication requirements and outperforms previous cluster-oriented schedulers.
F. Jesús Sánchez, Antonio González 0001
MICRO1
1999 A locality sensitive multi-module cache with explicit management
abstract
Article Free Access Share on A locality sensitive multi-module cache with explicit management Authors: Jesús Sánchez Department of Computer Architecture, Universitat Politècnica de Catalunya, Barcelona, Spain Department of Computer Architecture, Universitat Politècnica de Catalunya, Barcelona, SpainView Profile , Antonio González Department of Computer Architecture, Universitat Politècnica de Catalunya, Barcelona, Spain Department of Computer Architecture, Universitat Politècnica de Catalunya, Barcelona, SpainView Profile Authors Info & Claims ICS '99: Proceedings of the 13th international conference on SupercomputingJune 1999 Pages 51–59https://doi.org/10.1145/305138.305158Published:01 May 1999Publication History 20citation310DownloadsMetricsTotal Citations20Total Downloads310Last 12 Months18Last 6 weeks2 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
F. Jesús Sánchez, Antonio González 0001
International Conference on Supercomputing1
1999 Software Data Prefetching for Software Pipelined Loops
F. Jesús Sánchez, Antonio González 0001
J. Parallel Distributed Comput.1
1997 Cache Sensitive Modulo Scheduling
abstract
This paper focuses on the interaction between software prefetching (both binding and nonbinding) and software pipelining for VLIW machines. First, it is shown that evaluating software pipelined schedules without considering memory effects can be rather inaccurate due to stalls caused by dependences with memory instructions (even if a lockup-free cache is considered). It is also shown that the penalty of the stalls is in general higher than the effect of spill code. Second, we show that in general binding-schemes are more powerful than nonbinding ones for software pipelined schedules. Finally, the main contribution of this paper is an heuristic scheme that schedules some memory operations according to the locality estimated at compile time and other attributes of the dependence graph. The proposed scheme is shown to outperform other heuristic approaches since it achieves a better trade-off between compute and stall time than the others.
F. Jesús Sánchez, Antonio González 0001
MICRO1