Amit Golander

dblp:63/3690 · DBLP profile ↗
← Back
14ranked-venue papers
7as first author
5since 2021 · last 2025
0009-0000-6798-6183ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 6 first-author · 5 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2025 Accelerating AI through Novel Compute Architecture and Digital Photonic Gates
abstract
The growing computational demands and performance limits of AI models are often constrained by their underlying mathematical processing units. This issue spans various model sizes and hardware---from laptop CPUs to large GPU clusters. Current units, like a CPU's Floating-Point Unit (FPU), use multi-cycle pipelines. While this design supports complex computations, it also adds latency and slows floating-point performance. Spending multiple cycles on an operation is slow, but preferable to lowering the global clock frequency to support these intricate units.
Fahim Bellan, Amit Golander
SYSTOR2
2025 Compute-based Fault Tolerance for DNN
abstract
Deep neural network (DNN) systems use many GPUs, which can fail---making fault tolerance (FT) essential to avoid cluster restarts. Traditional FT relies on frequent checkpointing, incurring high bandwidth and memory costs. We propose an alternative strategy using GPU redundancy, introducing uniform and heterogeneous encoding approaches. We analyze their costs and recommend usage scenarios, especially for emerging rack-scale AI computers
Adi Molkho, Amit Golander, Oded Schwartz
SYSTOR2
2024 Cheap & Fast File-aaS for AI by Combining Scale-out Virtiofs, Block layouts and Delegations
abstract
File service Supply & Demand forces are changing: on the demand side, AI has significantly increased single-tenant File-aaS performance requirements. On the supply side, economic forces are moving the clients from traditional servers to rack-scale computers. Rack-scale computers encourage heterogeneous compute and are equipped with DPUs (Data-Processing Unit, also known as Smart NIC). DPUs offer higher efficiency, but may be over utilized at times. For this reason we would like our software to be flexible and run on general purpose compute when DPU utilization is high.
Sagi Manole, Amit Golander
SYSTOR2
2023 Near-Memory Processing Offload to Remote (Persistent) Memory
abstract
Traditional Von Neumann computing architectures are struggling to keep up with the rapidly growing demand for scale, performance, power-efficiency and memory capacity. One promising approach to this challenge is Remote Memory, in which the memory is over RDMA fabric [1]. We enhance the remote memory architecture with Near Memory Processing (NMP), a capability that offloads particular compute tasks from the client to the server side as illustrated in Figure 1. Similar motivation drove IBM to offload object processing to their remote KV storage [2].
Roei Kisous, Amit Golander, Yigal Korman, Tim Gubner, Rune Humborstad, Manyi Lu
SYSTOR2
2022 pmAddr: a persistent memory centric computing architecture
abstract
Traditional Processor-centric computing architectures do not scale-out well, because servers do not share their local main memories. To bypass this architectural limitation, programmers place their shared state on shared storage. But since Storage is slow (many hundreds of microseconds), they speed performance by duplicating the shared state to the compute nodes and have a complex coherent protocols to try keep all copies in sync. In recent years, Memory-centric architectures were proposed as an alternative.
Amit Golander, Shai Taharlev, Yigal Korman
SYSTOR1
2018 Accelerating Unmodified Databases using Persistent Memory and Flash Storage Tiers
abstract
Recent breakthroughs in Storage Class Memory (SCM) technologies have driven Persistent Memory (PM) devices to become commodity off-the-shelf components in 2018. PM devices are byte addressable, plug into the memory interconnect, and run at near memory speeds, densities and price points. PM availability is led by Fast PM, comprised from backed-DRAM devices such as NVDIMM-N, and will follow soon with Slow PM, comprised of new SCM materials, such as Intel 3D XPoint NVDIMM. Fast and Slow PM devices vary in speeds, densities and cost, but both are orders of magnitude faster than Flash devices and an order of magnitude more expensive per GB.
Amit Golander, Netanel Katzburg, Omer Zilberberg
SYSTOR1
2017 Persistent memory over fabric (PMoF)
abstract
Persistent Memory (PM) is an emerging family of technologies that are: persistent; byte addressable; and respond in near-memory speeds. PM devices, also referred to as non-volatile DIMMs or NVDIMMs, connect to the low-latency CPU memory interconnect. PM-based solutions can achieve local persistency within a micro second, which is two orders-of-magnitude faster compared to modern Flash solutions [1].
Amit Golander, Sagi Manole, Yigal Korman
SYSTOR1
2014 Protein Sequence Pattern Matching: Leveraging Application Specific Hardware Accelerators
abstract
Digitalization has brought a tremendous momentum to health care research. Recognition of patterns in proteins is crucial for identifying possible functions of newly discovered proteins, as well as analysis of known proteins for previously undetermined activity. In this paper, the workload consists of locating patterns from the PROSITE database in protein sequences. We optimize the pattern search task by using a new breed of processors that merge network and server attributes. We leverage massive multithreading and regular-expression (RegX) hardware accelerators; the latter were designed and built for an entirely different application - high-bandwidth deep-packet inspection. Our multithreading optimization achieves 18x improvement, but by harnessing a RegX accelerator we were able to further demonstrate a significant 392x improvement relative to software pattern matching. Moreover, performance per area and power consumption are improved by multiple orders of magnitude as well.
Sagi Manole, Amit Golander, Shlomo Weiss
IEEE Trans. Computers2
2014 L1-L2 Interconnect Design Methodology and Arbitration in 3-D IC Multicore Compute Clusters
abstract
We introduce a novel 3-D implementation of the interconnect between cores and shared L2 cache banks for multicore clusters. The 3-D structure extends cluster sizes that can be supported with tolerable wire delays. As a result of the shorter connections achieved by splitting existing 2-D design into four layers, performance is improved and area and power are reduced. The splitting enables implementation of a better arbitration scheme, which leads to additional performance improvement.
Alexei Jolondz, Shlomo Weiss, Amit Golander
IEEE Trans. Very Large Scale Integr. Syst.3
2013 High Compression Rate and Ratio Using Predefined Huffman Dictionaries
abstract
Current Huffman coding modes are optimal for a single metric: compression ratio (quality) or rate (performance). We recognize that real life data can usually be classified to families of data types and thus the Huffman dictionary can be reused instead of recalculated. In this paper, we show how to balance the trade-off between compression ratio and rate, without modifying existing standards and legacy decompression implementations.
Amit Golander, Shai Tahar, Lior Glass, Giora Biran, Sagi Manole
DCC1
2013 Leveraging predefined huffman dictionaries for high compression rate and ratio
abstract
The explosion of data, both in motion and in rest, along with the popularity of cloud computing, have resulted in the need for better in-line compression solutions. Current Huffman coding modes are optimal for a single metric: compression ratio (quality) or rate (performance). The ratio-focused mode, for example, further compresses the data by 15% as compared to the rate-focused mode, but takes 20% longer to execute.
Amit Golander, Shai Tahar, Lior Glass, Giora Biran, Sagi Manole
SYSTOR1
2009 Checkpoint allocation and release
abstract
Out-of-order speculative processors need a bookkeeping method to recover from incorrect speculation. In recent years, several microarchitectures that employ checkpoints have been proposed, either extending the reorder buffer or entirely replacing it. This work presents an in-dept-study of checkpointing in checkpoint-based microarchitectures, from the desired content of a checkpoint, via implementation trade-offs, and to checkpoint allocation and release policies. A major contribution of the article is a novel adaptive checkpoint allocation policy that outperforms known policies. The adaptive policy controls checkpoint allocation according to dynamic events, such as second-level cache misses and rollback history. It achieves 6.8% and 2.2% speedup for the integer and floating point benchmarks, respectively, and does not require a branch confidence estimator. The results show that the proposed adaptive policy achieves most of the potential of an oracle policy whose performance improvement is 9.8% and 3.9% for the integer and floating point benchmarks, respectively. We exploit known techniques for saving leakage power by adapting and applying them to checkpoint-based microarchitectures. The proposed applications combine to reduce the leakage power of the register file to about one half of its original value.
Amit Golander, Shlomo Weiss
ACM Trans. Archit. Code Optim.1
2008 Stateful hardware decompression in networking environment
abstract
Compression and Decompression can significantly lower the network bandwidth requirements for common internet traffic. Driven by the demands of an enterprise network intrusion system, this paper defines and examines the requirements of popular dictionary-based decompression in the real-time network processing scenario. In particular, a "stateful" decompression is required that arises out of the packet oriented nature of current networks, where the decompression of the data of a packet depends on the decompressed contents of its preceeding packets composing the same data stream. We propose an effective hardware decompression acceleration engine, which fetches the history data into the accelerator's fast memory on-demand and hides the associated latency by exploring the parallelism of the dictionary-based decompression process. We specify and evaluate various design and implementation options of the fetch-on-demand mechanism, i.e. prefetch most frequently used history, on-accelerator history buffer management, and reuse of fetched history data. Through simulation-based performance study, we show the effectiveness of the proposed mechanism on hiding the overhead of stateful decompression. We further show the effects of the design options and the impact on the overall performance of the network service stack of an intrusion prevension system.
Hao Yu 0008, Hubertus Franke, Giora Biran, Amit Golander, Terry Nelms, Brian M. Bass
ANCS4
2008 Hiding the misprediction penalty of a resource-efficient high-performance processor
abstract
Misprediction is a major obstacle for increasing speculative out-of-order processors performance. Performance degradation depends on both the number of misprediction events and the recovery time associated with each one of them. In recent years a few checkpoint based microarchitectures have been proposed. In comparison with ROB-based processors, checkpoint processors are scalable and highly resource efficient. Unfortunately, in these proposals the misprediction recovery time is proportional to the instruction queue size. In this paper we analyze methods to reduce the misprediction recovery time. We propose a new register file management scheme and techniques to selectively flush the instruction queue and the load store queue, and to isolate deeply pipelined execution units. The result is a novel checkpoint processor with Constant misprediction RollBack time (CRB). We further present a streamlined, cost-efficient solution, which saves complexity at the price of slightly lower performance.
Amit Golander, Shlomo Weiss
ACM Trans. Archit. Code Optim.1