VLDB 2026 Research / reviewers in the wild / expert
Mohamed A. Gomaa
dblp:15/82
· DBLP profile ↗
3ranked-venue papers
3as first author
0since 2021 · last 2005
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 first-authorSoftware engineering, systems software and programming languages · 3 · 3 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Hardware reliability and fault tolerance · 45% Processor architecture and microarchitecture · 40% Energy-efficient computing · 11% |
Topics — the 8 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Processor architecture and microarchitecture › multithreading
simultaneous multithreading |
0.1 | 2 | 2004 | Heat-and-run: leveraging SMT and CMP to manage power density through the operating system · ASPLOS 2004 Transient-Fault Recovery for Chip Multiprocessors · ISCA 2003 |
Processor architecture and microarchitecture › dynamic optimization
instruction reuse |
0.1 | 1 | 2005 | Opportunistic Transient-Fault Detection · ISCA 2005 |
Hardware reliability and fault tolerance › soft errors
soft error rate reduction |
0.1 | 1 | 2005 | Opportunistic Transient-Fault Detection · ISCA 2005 |
Hardware reliability and fault tolerance › error detection
transient fault detection |
0.1 | 1 | 2005 | Opportunistic Transient-Fault Detection · ISCA 2005 |
Processor architecture and microarchitecture
chip multiprocessor |
0.0 | 1 | 2003 | Transient-Fault Recovery for Chip Multiprocessors · ISCA 2003 |
Hardware reliability and fault tolerance › soft errors
transient fault |
0.0 | 1 | 2003 | Transient-Fault Recovery for Chip Multiprocessors · ISCA 2003 |
Hardware reliability and fault tolerance › error recovery
transient fault recovery |
0.0 | 1 | 2003 | Transient-Fault Recovery for Chip Multiprocessors · ISCA 2003 |
Processor architecture and microarchitecture › out-of-order execution
issue queue |
0.0 | 1 | 2005 | Opportunistic Transient-Fault Detection · ISCA 2005 |
Methods — techniques the papers use, named apart from their topics
trace-driven simulation · 0.1thread migration · 0.0thread co-scheduling · 0.0simulation · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2005 | Opportunistic Transient-Fault DetectionabstractCMOS scaling increases susceptibility of microprocessors to transient faults. Most current proposals for transient-fault detection use full redundancy to achieve perfect coverage while incurring significant performance degradation. However, most commodity systems do not need or provide perfect coverage. A recent paper explores this leniency to reduce the soft-error rate of the issue queue during L2 misses while incurring minimal performance degradation. Whereas the previous paper reduces soft-error rate without using any redundancy, we target better coverage while incurring similarly-minimal performance degradation by opportunistically using redundancy. We propose two semi-complementary techniques, called partial explicit redundancy (PER) and implicit redundancy through reuse (IRTR), to explore the trade-off between soft-error rate and performance. PER opportunistically exploits low-ILP phases and L2 misses to introduce explicit redundancy with minimal performance degradation. Because PER covers the entire pipeline and exploits not only L2 misses but all low-ILP phases, PER achieves better coverage than the previous work. To achieve coverage in high-ILP phases as well, we propose implicit redundancy through reuse (IRTR). Previous work exploits the phenomenon of instruction reuse to avoid redundant execution while falling back on redundant execution when there is no reuse. IRTR takes reuse to the extreme of performance-coverage trade-off and completely avoids explicit redundancy by exploiting reuse's implicit redundancy within the main thread for fault detection with virtually no performance degradation. Using simulations with SPEC2000, we show that PER and IRTR achieve better tradeoff between soft-error rate and performance degradation than the previous schemes. Mohamed A. Gomaa, T. N. Vijaykumar |
ISCA | 1 |
| 2004 | Heat-and-run: leveraging SMT and CMP to manage power density through the operating systemabstractPower density in high-performance processors continues to increase with technology generations as scaling of current, clock speed, and device density outpaces the downscaling of supply voltage and thermal ability of packages to dissipate heat. Power density is characterized by localized chip hot spots that can reach critical temperatures and cause failure. Previous architectural approaches to power density have used global clock gating, fetch toggling, dynamic frequency scaling, or resource duplication to either prevent heating or relieve overheated resources in a superscalar processor. Previous approaches also evaluate design technologies where power density is not a major problem and most applications do not overheat the processor. Future processors, however, are likely to be chip multiprocessors (CMPs) with simultaneously-multithreaded (SMT) cores. SMT CMPs pose unique challenges and opportunities for power density. SMT and CMP increase throughput and thus on-chip heat, but also provide natural granularities for managing power-density. This paper is the first work to leverage SMT and CMP to address power density. We propose heat-and-run SMT thread assignment to increase processor-resource utilization before cooling becomes necessary by co-scheduling threads that use complimentary resources. We propose heat-and-run CMP thread migration to migrate threads away from overheated cores and assign them to free SMT contexts on alternate cores, leveraging availability of SMT contexts on alternate CMP cores to maintain throughput while allowing overheated cores to cool. We show that our proposal has an average of 9% and up to 34% higher throughput than a previous superscalar technique running the same number of threads. Mohamed A. Gomaa, Michael D. Powell, T. N. Vijaykumar |
ASPLOS | 1 |
| 2003 | Transient-Fault Recovery for Chip MultiprocessorsabstractTo address the increasing susceptibility of commodity chip multiprocessors (CMPs) to transient faults, we propose Chiplevel Redundantly Threaded multiprocessor with Recovery (CRTR). CRTR extends the previously-proposed CRT for transient-fault detection in CMPs, and the previously-proposed SRTR for transient-fault recovery in SMT. All these schemes achieve fault tolerance by executing and comparing two copies, called leading and trailing threads, of a given application. Previous recovery schemes for SMT do not perform well on CMPs. In a CMP, the leading and trailing threads execute on different processors to achieve load balancing and reduce the probability of a fault corrupting both threads; whereas in an SMT, both threads execute on the same processor. The inter-processor communication required to compare the threads introduces latency and bandwidth problems not present in an SMT.To hide inter-processor latency, CRTR executes the leading thread ahead of the trailing thread by maintaining a long slack, enabled by asymmetric commit. CRTR commits the leading thread before checking and the trailing thread after checking, so that the trailing thread state may be used for recovery. Previous recovery schemes commit both threads after checking, making a long slack suboptimal. To tackle inter-processor bandwidth, CRTR not only increases the bandwidth supply by pipelining the communication paths, but also reduces the bandwidth demand. By reasoning that faults propagate through dependences, previously-proposed Dependence-Based Checking Elision (DBCE) exploits (true) register dependence chains so that only the value of the last instruction in a chain is checked. However, instructions that mask operand bits may mask faults and limit the use of dependence chains. We propose Death- and Dependence-Based Checking Elision (DDBCE), which chains a masking instruction only if the source operand of the instruction dies after the instruction. Register deaths ensure that masked faults do not corrupt later computation. Using SPEC2000, we show that CRTR incurs negligible performance loss compared to CRT for inter-processor (one-way) latency as high as 30 cycles, and that the bandwidth requirements of CRT and CRTR with DDBCE are 5.2 and 7.1 bytes/cycle, respectively. Mohamed A. Gomaa, Chad Scarbrough, Irith Pomeranz, T. N. Vijaykumar |
ISCA | 1 |