EDBT 2026 Demo / reviewers in the wild / expert
Marc Lupon
dblp:68/7750
· DBLP profile ↗
6ranked-venue papers
3as first author
1since 2021 · last 2021
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 3 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Memory systems · 34% Processor architecture and microarchitecture · 31% Parallel and multicore computing · 16% | |
| Software engineering, system software, and programming languages
2 papers |
Concurrent programming · 69% Runtime systems and virtual machines · 31% |
Topics — the 13 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Processor architecture and microarchitecture › computer arithmetic › floating-point arithmetic
fused multiply-add |
0.2 | 1 | 2014 | Speculative hardware/software co-designed floating-point multiply-add fusion · ASPLOS 2014 |
Electronic design automation
hardware/software co-design |
0.2 | 1 | 2014 | Speculative hardware/software co-designed floating-point multiply-add fusion · ASPLOS 2014 |
Processor architecture and microarchitecture
instruction set architecture |
0.2 | 1 | 2014 | Speculative hardware/software co-designed floating-point multiply-add fusion · ASPLOS 2014 |
Concurrent programming
transactional memory |
0.1 | 1 | 2011 | Safe and efficient supervised memory systems · HPCA 2011 |
Memory systems › memory consistency
memory consistency model |
0.1 | 1 | 2011 | Safe and efficient supervised memory systems · HPCA 2011 |
Memory systems › memory consistency › memory consistency model
total store order |
0.1 | 1 | 2011 | Safe and efficient supervised memory systems · HPCA 2011 |
Memory systems
cache coherence |
0.1 | 1 | 2010 | A Dynamically Adaptable Hardware Transactional Memory · MICRO 2010 |
Memory systems › cache coherence
conflict management |
0.1 | 1 | 2010 | A Dynamically Adaptable Hardware Transactional Memory · MICRO 2010 |
Parallel and multicore computing › transactional memory
hardware transactional memory |
0.1 | 1 | 2010 | A Dynamically Adaptable Hardware Transactional Memory · MICRO 2010 |
Parallel and multicore computing
transactional memory |
0.1 | 1 | 2010 | A Dynamically Adaptable Hardware Transactional Memory · MICRO 2010 |
Runtime systems and virtual machines › binary translation
dynamic binary translation |
0.1 | 1 | 2014 | Speculative hardware/software co-designed floating-point multiply-add fusion · ASPLOS 2014 |
Distributed systems
concurrency control |
0.0 | 1 | 2010 | A Dynamically Adaptable Hardware Transactional Memory · MICRO 2010 |
Storage systems › file systems
versioning |
0.0 | 1 | 2010 | A Dynamically Adaptable Hardware Transactional Memory · MICRO 2010 |
Methods — techniques the papers use, named apart from their topics
speculative instruction-fusion optimization · 0.4cycle-accurate simulation · 0.4formal specification · 0.2RTL implementation · 0.2eager and lazy versioning · 0.1conflict prediction · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Accelerating ML Recommendation with over a Thousand RISC-V/Tensor Processors on Esperanto's ET-SoC-1 ChipabstractThe ET-SoC-1 has over a thousand RISC-V processors on a single TSMC 7nm chip, including: • 1088 energy-efficient ET-Minion 64-bit RISC-V in-order cores each with a vector/tensor unit • 4 high-performance ET-Maxion 64-bit RISC-V out-of-order cores • >160 million bytes of on-chip SRAM • Interfaces for large external memory with low-power LPDDR4x DRAM and eMMC FLASH • PCIe x8 Gen4 and other common I/O interfaces • Innovative low-power architecture and circuit techniques allows entire chip to • Compute at peak rates of 100 to 200 TOPS • Operate using under 20 watts for ML recommendation workloads David R. Ditzel, Roger Espasa, Nivard Aymerich, Allen Baum, Tom Berg, Jim Burr, Eric Hao, Jayesh Iyer, Miquel Izquierdo, Shankar Jayaratnam, Darren Jones, Chris Klingner, Stephen Lee, Marc Lupon, Grigorios Magklis, Bojan Maric, Rajib Nath, Mike Neilly, J. Duane Northcutt, Bill Orner, Jose Renau, Gerard Reves, Xavier Reves, Tom Riordan, Pedro Sanchez, Sridhar Samudrala, Guillem Sole, Raymond Tang, Tommy Thorn, Sebastia Tortella, Daniel Yau |
HCS | 15 |
| 2015 | Reactive clocks with variability-tracking jitterabstractThe growing variability in nanoelectronic devices, due to uncertainties from the manufacturing process and environmental conditions (power supply, temperature, aging), requires increasing design guardbands, forcing circuits to work with conservative clock frequencies. Various schemes for clock generation based on ring oscillators and adaptive clocks have been proposed with the goal to mitigate the power and performance losses attributable to variability. However, there has been no systematic analysis to quantify the benefits of such schemes and no sign-off method has been proposed for timing correctness. This paper presents and analyzes a Reactive Clocking scheme with Variability-Tracking Jitter (RClk) that uses variability as an opportunity to reduce power by continuously adjusting the clock frequency to the varying environmental conditions, and thus, reduces guardband margins significantly. Power can be reduced between 20% and 40% at iso-performance and performance can be boosted by similar amounts at iso-power. Additionally, energy savings can be translated to substantial advantages in terms of reliability and thermal management. More importantly, the technology can be adopted with minimal modifications to conventional EDA flows. Jordi Cortadella, Luciano Lavagno, Pedro Lopez, Marc Lupon, Alberto Moreno-Conde, Antoni Roca 0001, Sachin S. Sapatnekar |
ICCD | 4 |
| 2014 | Speculative hardware/software co-designed floating-point multiply-add fusionabstractA Fused Multiply-Add (FMA) instruction is currently available in many general-purpose processors. It increases performance by reducing latency of dependent operations and increases precision by computing the result as an indivisible operation with no intermediate rounding. However, since the arithmetic behavior of a single-rounding FMA operation is different than independent FP multiply followed by FP add instructions, some algorithms require significant revalidation and rewriting efforts to work as expected when they are compiled to operate with FMA--a cost that developers may not be willing to pay. Because of that, abundant legacy applications are not able to utilize FMA instructions. In this paper we propose a novel HW/SW collaborative technique that is able to efficiently execute workloads with increased utilization of FMA, by adding the option to get the same numerical result as separate FP multiply and FP add pairs. In particular, we extended the host ISA of a HW/SW co-designed processor with a new Combined Multiply-Add (CMA) instruction that performs an FMA operation with an intermediate rounding. This new instruction is used by a transparent dynamic translation software layer that uses a speculative instruction-fusion optimization to transform FP multiply and FP add sequences into CMA instructions. The FMA unit has been slightly modified to support both single-rounding and double-rounding fused instructions without increasing their latency and to provide a conservative fall-back path in case of mispeculation. Evaluation on a cycle-accurate timing simulator showed that CMA improved SPECfp performance by 6.3% and reduced executed instructions by 4.7%. Marc Lupon, Enric Gibert, Grigorios Magklis, Sridhar Samudrala, Raúl Martínez, Kyriakos Stavrou, David R. Ditzel |
ASPLOS | 1 |
| 2011 | Safe and efficient supervised memory systemsabstractSupervised Memory systems use out-of-band metabits to control and monitor accesses to normal data memory for such purposes as transactional memory and memory typestate trackers. Previous proposals demonstrate the value of supervised memory systems, but have typically (1) assumed sequential consistency (while most deployed systems use weaker models), and (2) used ad hoc, informal memory specifications (that can be ambiguous and/or incorrect). This paper seeks to make many previous proposals more practical. This paper builds a foundation for future supervised memory systems which (1) operate with the TSO and ×86 memory models, and (2) are formally specified using two supervised memory models. The simpler TSOallmodel requires all metadata and data accesses to obey TSO, but precludes using store buffers for supervised accesses. The more complex TSOdatamodel relaxes some ordering constraints (allowing store buffer use) but makes programmer reasoning more difficult. To get the benefits of both models, we propose Safe Supervision, which asks programmers to avoid using metabits from one location to order accesses to another. Programmers that obey safe supervision can reason with the simpler semantics of TSOallwhile obtaining the higher performance of TSOdata. Our approach is similar to how data-race-free programs can run on relaxed systems and yet appear sequentially consistent. Finally, we show that TSOdatacan (a) provide significant performance benefit (up to 22%) over TSOalland (b) can be incorporated correctly and with low overhead into the RTL of an industrial multi-core chip design (OpenSPARC T2). Jayaram Bobba, Marc Lupon, Mark D. Hill, David A. Wood 0001 |
HPCA | 2 |
| 2010 | A Dynamically Adaptable Hardware Transactional MemoryabstractMost Hardware Transactional Memory (HTM) implementations choose fixed version and conflict management policies at design time. While eager HTM systems store transactional state in-place in memory and resolve conflicts when they are produced, lazy HTM systems buffer the transactional state in specialized hardware and defer the resolution of conflicts until commit time. Each scheme has its strengths and weaknesses, but, unfortunately, both approaches are too inflexible in the way they manage data versioning and transactional contention. Thus, fixed HTM systems may result in a significant performance opportunity loss when they execute complex transactional applications. In this paper, we present DynTM (Dynamically Adaptable HTM), the first fully-flexible HTM system that permits the simultaneous execution of transactions using complementary version and conflict management strategies. In the heart of DynTM is a novel coherence protocol that allows tracking conflicts among eager and lazy transactions. Both the eager and the lazy execution modes of DynTM exhibit very high performance compared to modern HTM systems. For example, the DynTM lazy execution mode implements local commits to improve on previous proposals. In addition, lazy transactions share the majority of hardware support with eager transactions, reducing substantially the hardware cost compared to other lazy HTM systems. By utilizing a simple predictor to decide the best execution mode for each transaction at runtime, DynTM obtains an average speedup of 34% over HTM systems that employ fixed version and conflict management policies. Marc Lupon, Grigorios Magklis, Antonio González 0001 |
MICRO | 1 |
| 2009 | FASTM: A Log-based Hardware Transactional Memory with Fast Abort RecoveryabstractVersion management, one of the key design dimensions of hardware transactional memory (HTM) systems, defines where and how transactional modifications are stored. Current HTM systems use either eager or lazy version management. Eager systems that keep new values in-place while they hold old values in a software log, suffer long delays when aborts are frequent because the pre-transactional state is recovered by software. Lazy systems that buffer new values in specialized hardware offer complex and inefficient solutions to handle hardware overflows, which are common in applications with coarse-grain transactions. In this paper, we present FASTM, an eager log-based HTM that takes advantage of the processorpsilas cache hierarchy to provide fast abort recovery. FASTM uses a novel coherence protocol to buffer the transactional modifications in the first level cache and to keep the non-speculative values in the higher levels of the memory hierarchy. This mechanism allows fast abort recovery of transactions that do not overflow the first level cache resources. Contrary to lazy HTM systems, committing transactions do not have to perform any actions in order to make their results visible to the rest of the system. FASTM keeps the pre-transactional state in a software-managed log as well, which permits the eviction of speculative values and enables transparent execution even in the case of cache overflow. This approach simplifies eviction policies without degrading performance, because it only falls back to a software abort recovery for transactions whose modified state has overflowed the cache. Simulation results show that FASTM achieves a speed-up of 43% compared to LogTM-SE, improving the scalability of applications with coarse-grain transactions and obtaining similar performance to an ideal eager HTM with zero-cost abort recovery. Marc Lupon, Grigorios Magklis, Antonio González 0001 |
PACT | 1 |