VLDB 2026 Research / reviewers in the wild / expert
Shmuel Wimer
dblp:49/3159
· DBLP profile ↗
30ranked-venue papers
11as first author
3since 2021 · last 2023
0000-0002-5728-0061ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 26 · 10 first-author · 2 since 2021Theory of computation · 4 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | A Self-Refreshable Bit-Cell for Single-Cycle Refreshing of Embedded MemoriesabstractPower supply voltage reduction is a primary enabler for sustaining the increasing demand for ultra-low power processors. On-die memories, which are traditionally implemented by SRAM, stop functioning properly when the supply voltage is scaled down aggressively; hence, embedded DRAM (eDRAM) bit-cells are used instead. These bit-cells leak their data strongly in one direction, whereas the leakage in the opposite direction is considerably lower. Due to their intrinsic limited Data Retention Time (DRT), these memories require power-hungry refreshing, which degrades performance. In an attempt to extend the DRT of a bit-cell, theoretically to infinity, compounds of various types of storage nodes in a single bit-cell, storing the datum and its complement, were examined here. A rigorous proof shows that under realistic leakage models, there is an inherent incompleteness preventing the proper readout and decision of the stored value after a certain time. Adopting the idea of dual-polarity complementary storage nodes, a new eDRAM self-refreshable bit-cell is proposed that yields a considerably extended DRT. The dual-polarity property enables the refreshing of an entire memory array in a single clock cycle, thus almost nullifying the unavoidable performance loss occurred by row-by-row ordinary power-hungry refreshing. Binyamin Frankel, Shmuel Wimer |
IEEE Trans. Computers | 2 |
| 2023 | Energy Efficiency of Opportunistic Refreshing for Gain-Cell Embedded DRAMabstractOn-die memories, which are traditionally implemented by SRAM, stop functioning properly when the supply voltage is scaled down aggressively; hence, embedded DRAMs (eDRAMs) are used instead. Opportunistic refreshing was proved to eliminate the performance loss incurred by the eDRAM refreshing must. We show here that Gain-Cell eDRAM (GCeDRAM) supplemented with opportunistic refreshing consumes significantly smaller power and energy than SRAM. Analysis supported by hardware simulations demonstrate that the same design point achieves maximum performance and minimum energy. Replacement of the data memory in the ultra-low power processor PULPino from SRAM to opportunistically refreshed GCeDRAM yielded 30% energy savings in the memory, which translated into 7% savings in the entire processor. Binyamin Frankel, Eyal Sarfati, Davide Rossi 0001, Shmuel Wimer |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2021 | Resource allocation in rooted trees subject to sum constraints and nonlinear cost functions
Nir Halman, Shmuel Wimer |
Inf. Process. Lett. | 2 |
| 2018 | Queuing-Based eDRAM Refreshing for Ultra-Low Power ProcessorsabstractUltra-low power processors designed to work at very low voltage are the enablers of the internet of things (IoT) era. Their internal memories, which are usually implemented by a static random access memory (SRAM) technology, stop functioning properly at low voltage. Some recent commercial products have replaced SRAM with embedded memory (eDRAM), in which stored data are destroyed overtime, thus requiring periodic refreshing that causes performance loss. This article presents a queuing-based opportunistic refreshing algorithm that eliminates most if not all of the performance loss and is shown to be optimal. The queues used for refreshing miss refreshing opportunities not only when they are saturated but also when they are empty, hence increasing the probability of performance loss. We examine the optimal policy for handling a saturated and empty queue, and the ways in which system performance depends on queue capacity and memory size. This analysis results in a closed-form performance expression capturing read/write probabilities, memory size and queue capacity leading to CPU-internal memory architecture optimization. Binyamin Frankel, Roi Herman, Shmuel Wimer |
IEEE Trans. Computers | 3 |
| 2017 | Probability-Driven Multibit Flip-Flop Integration With Clock GatingabstractData-driven clock gated (DDCG) and multibit flip-flops (MBFFs) are two low-power design techniques that are usually treated separately. Combining these techniques into a single grouping algorithm and design flow enables further power savings. We study MBFF multiplicity and its synergy with FF data-to-clock toggling probabilities. A probabilistic model is implemented to maximize the expected energy savings by grouping FFs in increasing order of their data-to-clock toggling probabilities. We present a front-end design flow, guided by physical layout considerations for a 65-nm 32-bit MIPS and a 28-nm industrial network processor. It is shown to achieve the power savings of 23% and 17%, respectively, compared with designs with ordinary FFs. About half of the savings was due to integrating the DDCG into the MBFFs. Doron Gluzer, Shmuel Wimer |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2016 | Energy efficient deeply fused dot-product multiplication architectureabstractDot-product is frequently used in many signal processing applications, and for this reason digital signal processors (DSPs) commonly include hardware units to speed up its calculation. In this paper, we propose a new circuit architecture for dot-product hardware units. This design fuses all the steps of the dot-product computation into a unit called a fused dot-product multiplier (FDPM), resulting in enhancements in performance, power and area vs. designs such as multiply-accumulate (MAC) units. We designed and implemented the FDPM in 28-nanometer and 65-nanometer technologies. Comparison of the fused unit to an implementation by an EDA design compiler showed that our architecture consumed 38% less power and 30% less area for the same clock cycle. Comparison to a previously proposed FDPM architecture showed a per-bit 1.9X - 4.7X energy reduction. Analysis of the dependence of the per-bit power consumption on vector length and word-width indicated a good match between theory and practice. Shmuel Wimer, Israel Koren |
ASAP | 1 |
| 2016 | Energy efficient computing by multi-mode addition
Amir Albeck, Shmuel Wimer |
Integr. | 2 |
| 2015 | Timing-constrained power minimization in VLSI circuits by simultaneous multilayer wire spacing
Konstantin Moiseev, Shmuel Wimer, Avinoam Kolodny |
Integr. | 2 |
| 2015 | Energy efficient hybrid adder architecture
Shmuel Wimer, Amnon Stanislavsky |
Integr. | 1 |
| 2014 | Cell-based interconnect migration by hierarchical optimization
Eugene Shaphir, Ron Y. Pinter, Shmuel Wimer |
Integr. | 3 |
| 2014 | Planar CMOS to multi-gate layout conversion for maximal fin utilization
Shmuel Wimer |
Integr. | 1 |
| 2014 | Design Flow for Flip-Flop Grouping in Data-Driven Clock GatingabstractClock gating is a predominant technique used for power saving. It is observed that the commonly used synthesis-based gating still leaves a large amount of redundant clock pulses. Data-driven gating aims to disable these. To reduce the hardware overhead involved, flip-flops (FFs) are grouped so that they share a common clock enabling signal. The question of what is the group size maximizing the power savings is answered in a previous paper. Here we answer the question of which FFs should be placed in a group to maximize the power reduction. We propose a practical solution based on the toggling activity correlations of FFs and their physical position proximity constraints in the layout. Our data-driven clock gating is integrated into an Electronic Design Automation (EDA) commercial backend design flow, achieving total power reduction of 15%-20% for various types of large-scale state-of-the-art industrial and academic designs in 40 and 65 manometer process technologies. These savings are achieved on top of the sClock gating is a predominant technique used for power saving. It is observed that the commonly used synthesis-based gating still leaves a large amount of redundant clock pulses. Data-driven gating aims to disable these. To reduce the hardware overhead involved, flip-flops (FFs) are grouped so that they share a common clock enabling signal. The question of what is the group size maximizing the power savings is answered in a previous paper. Here we answer the question of which FFs should be placed in a group to maximize the power reduction. We propose a practical solution based on the toggling activity correlations of FFs and their physical position proximity constraints in the layout. Our data-driven clock gating is integrated into an Electronic Design Automation (EDA) commercial backend design flow, achieving total power reduction of 15%-20% for various types of large-scale state-of-the-art industrial and academic designs in 40 and 65 manometer process technologies. These savings are achieved on top of the savings obtained by clock gating synthesis performed by commercial EDA tools, and gating manually inserted into the register transfer level design.avings obtained by clock gating synthesis performed by commercial EDA tools, and gating manually inserted into the register transfer level design. Shmuel Wimer, Israel Koren |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2012 | Using well-solvable quadratic assignment problems for VLSI interconnect applications
Ben Emanuel, Shmuel Wimer, Gershon Wolansky |
Discret. Appl. Math. | 2 |
| 2012 | The Optimal Fan-Out of Clock Network for Power Minimization by Adaptive GatingabstractGating of the clock signal in VLSI chips is nowadays a mainstream design methodology for reducing switching power consumption. In this paper we develop a probabilistic model of the clock gating network that allows us to quantify the expected power savings and the implied overhead. Expressions for the power savings in a gated clock tree are presented and the optimal gater fan-out is derived, based on flip-flops toggling probabilities and process technology parameters. The resulting clock gating methodology achieves 10% savings of the total clock tree switching power. The timing implications of the proposed gating scheme are discussed. The grouping of FFs for a joint clocked gating is also discussed. The analysis and the results match the experimental data obtained for a 3-D graphics processor and a 16-bit microcontroller, both designed at 65-nanometer technology. Shmuel Wimer, Israel Koren |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2011 | A Cost Effective Centralized Adaptive Routing for Networks-on-ChipabstractAs the number of applications and programmable units in CMPs and MPSoCs increases, the Network-on-Chip (NoC) encounters unpredictable, heterogeneous and time dependent traffic loads. This motivates the introduction of adaptive routing mechanisms that balance the NoC's loads and achieve higher throughput compared with traditional oblivious routing schemes. An effective adaptive routing scheme should be based on a global view of the network state. However, most current adaptive routing schemes, following off-chip networks, are based on distributed reactions to local congestion. In this paper we leverage the unique on-chip capabilities and introduce a novel paradigm of NoC centralized adaptive routing. Our scheme continuously monitors the global traffic load in the network and modifies the routing of packets to improve load balancing accordingly. We present a specific design for the case of mesh topology, where XY or YX routes are adaptively selected for each source-destination pair. We show that while our implementation is lightweight and scalable in hardware costs, it outperforms oblivious and distributed adaptive routing schemes in terms of load balancing and average packet delay. Ran Manevich, Israel Cidon, Avinoam Kolodny, Isask'har Walter, Shmuel Wimer |
DSD | 5 |
| 2010 | Interconnect power and delay optimization by dynamic programming in gridded design rulesabstractThe lithography used for 32 nanometers and smaller VLSI process technologies restricts the admissible interconnect widths and spaces to a small set of discrete values with some interdependencies, so that traditional interconnect sizing by continuous-variable optimization techniques becomes impossible. We present a dynamic programming (DP) algorithm for simultaneous sizing and spacing of all wires in interconnect bundles (or bus structures), yielding the optimal power-delay tradeoff curve. It sets the width and spacing of all interconnects simultaneously, thus finding the global optimum. The DP algorithm is generic and can handle a variety of power-delay objectives, such as total power or delay, or weighted sum of both, power-delay product, max delay and alike. The algorithm consistently yields more than 10% dynamic power and 5% delay reduction for interconnect channels in industrial microprocessor blocks designed in 32 nanometer process technology, when applied as a post-layout optimization step to redistribute wires within interconnect channels of fixed width, without changing the area of the original layout. Konstantin Moiseev, Avinoam Kolodny, Shmuel Wimer |
ISPD | 3 |
| 2010 | Interconnect Bundle Sizing Under Discrete Design RulesabstractThe lithography used for 32 nm and smaller very large scale integrated process technologies restricts the admissible interconnect widths and spaces to a small set of discrete values with some interdependencies, so that traditional interconnect sizing by continuous-variable optimization techniques becomes impossible. We present a dynamic programming (DP) algorithm for simultaneous sizing and spacing of all wires in interconnect bundles, yielding the optimal power-delay tradeoff curve. DP algorithm sets the width and spacing of all interconnects simultaneously, thus finding the global optimum. The DP algorithm is generic and can handle a variety of power-delay objectives, such as total power or delay, or weighted sum of both, power-delay product, max delay, and alike. The algorithm consistently yields 6% dynamic power and 5% delay reduction for interconnect channels in industrial microprocessor blocks designed in 32 nm process technology, when applied as a post-layout optimization step to redistribute wires within interconnect channels of fixed width, without changing the area of the original layout. Konstantin Moiseev, Avinoam Kolodny, Shmuel Wimer |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2009 | Power-delay optimization in VLSI microprocessors by wire spacingabstractThe problem of optimal space allocation among interconnect wires in a VLSI layout, in order to minimize the switching power consumption and the average signal delay, is addressed in this article. We define a Weighted Power-Delay Sum (WPDS) objective function and derive necessary and sufficient conditions for the existence of optimal interwire space allocation, based on the notion of capacitance density. At the optimum, every wire must be in equilibrium of its line-to-line weighted capacitance density on its two opposite sides, and the WPDS of the whole circuit is minimal if and only if capacitance density is uniformly distributed across the entire layout. This condition is shown to be equivalent to all paths of the layout cross-capacitance graph having the same length and all cuts having the same flow. An implementation which has been used in the design of a recent commercial high-end microprocessor and yielded 17% power reduction and 9% delay reduction in top-level interconnects is presented. Konstantin Moiseev, Avinoam Kolodny, Shmuel Wimer |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2008 | On optimal ordering of signals in parallel wire bundles
Konstantin Moiseev, Shmuel Wimer, Avinoam Kolodny |
Integr. | 2 |
| 2008 | Timing-aware power-optimal ordering of signalsabstractA computationally efficient technique for reducing interconnect active power in VLSI systems is presented. Power reduction is accomplished by simultaneous wire spacing and net ordering, such that cross-capacitances between wires are optimally shared. The existence of a unique power-optimal wire order within a bundle is proven, and a method to construct this order is derived. The optimal order of wires depends only on the activity factors of the underlying signals; hence, it can be performed prior to spacing optimization. By using this order of wires, optimality of the combined solution is guaranteed (as compared with any other ordering and spacing of the wires). Timing-aware power optimization is enabled by simultaneously considering timing criticality weights and activity factors for the signals. The proposed algorithm has been applied to various interconnect layouts, including wire bundles from high-end microprocessor circuits in 65 nm technology. Interconnect power reduction of 17% on average has been observed in such bundles. Konstantin Moiseev, Avinoam Kolodny, Shmuel Wimer |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2006 | Timing optimization of interconnect by simultaneous net-ordering, wire sizing and spacingabstractThis paper addresses the problem of ordering and sizing parallel wires in a single metal layer within an interconnect channel of a given width, such that cross-capacitances are optimally shared for circuit timing optimization. Using an Elmore delay model including cross capacitances for a bundle of wires, we show that an optimal wire ordering is uniquely determined, such that best timing can be obtained by proper allocation of wire widths and inter-wire spaces. The optimal order, called BMI (balanced monotonic interleaved) depends only on the size of drivers for a wide range of cases. Heuristics are presented for simultaneous ordering, sizing and spacing of wires. Examples for 90-nanometer technology are analyzed and discussed Konstantin Moiseev, Shmuel Wimer, Avinoam Kolodny |
ISCAS | 2 |
| 1993 | On Paths with the Shortest Average Arc Length in Weighted Graphs
Shmuel Wimer, Israel Koren, Israel Cederbaum |
Discret. Appl. Math. | 1 |
| 1993 | An efficient algorithm for some multirow layout problemsabstractThree multirow layout problems are presented: transistor orientation, contact positioning and symbolic-to-shape translation. It is shown that these multirow problems have a common property, called quantitative dependency. Using this property, an optimization technique which is based on a penalty-delay strategy is presented. It is proved that the penalty-delay strategy assures optimality, and that the optimal solution can be obtained in linear time. The algorithmic approach is based on the observation that optimal layout decisions in any region within a cell or a macro depend only on quantitative measures of the decisions in other regions, rather than on their details. This suggests a departure from the traditional approach of handling the different regions separately and combining them afterward into a single unit, an approach that may degrade the quality of the final layout. Instead, the entire macro can be processed at once, taking into account the mutual quantitative dependency between distinguished regions.> Jack A. Feldman, Israel A. Wagner, Shmuel Wimer |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 1992 | Balanced Block Spacing for VLSI Layout
Israel Cederbaum, Israel Koren, Shmuel Wimer |
Discret. Appl. Math. | 3 |
| 1989 | Depth-first-search and dynamic programming algorithms for efficient CMOS cell generationabstractAn algorithmic framework is presented for mapping CMOS circuit diagrams into area-efficient, high-performance layouts in the style of one-dimensional transistor arrays. Using efficient search techniques and accurate evaluation methods, the huge solution space that is typical to such problems is transversed extremely fast, yielding designs of hand-layout quality. In addition to generating circuits that meet prespecified layout constraints in the context of a fixed target image, on-the-fly optimizations are performed to meet secondary optimization criteria. A practical dynamic programming routing algorithm is utilized to accommodate the special conditions that arise in this context. This algorithm has been implemented and is currently used at IBM for cell-library generation.> Reuven Bar-Yehuda, Jack A. Feldman, Ron Y. Pinter, Shmuel Wimer |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 1989 | Optimal aspect ratios of building blocks in VLSIabstractThe building blocks in a given floorplan have several possible physical implementations yielding different layouts. A discussion is presented of the problem of selecting an optimal implementation for each building block so that the area of the final layout is minimized. A polynomial algorithm that solves this problem for slicing floorplans was presented elsewhere, and it has been proved that for general (nonslicing) floorplans the problem is NP-complete. The authors suggest a branch-and-bound algorithm which proves to be very efficient and can handle successfully large general nonslicing floorplans. The high efficiency of the algorithm stems from the branching strategy and the bounding function used in the search procedure. The branch-and-bound algorithm is supplemented by a heuristic minimization procedure which further prunes the search, is computationally efficient, and does not prevent achieving a global minimum. Finally, the authors show how the nonslicing and the slicing algorithms can be combined to handle efficiently very large general floorplans.> Shmuel Wimer, Israel Koren, Israel Cederbaum |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1988 | Optimal Aspect Ratios of Building Blocks in VLSI
Shmuel Wimer, Israel Koren, Israel Cederbaum |
DAC | 1 |
| 1988 | Analysis of strategies for constructive general block placementabstractThe problem of general block placement in VLSI is considered, using the constructive approach in which blocks are selected and located one at a time. Some well-known strategies are presented for the selection of the next block to be located, novel ones are proposed, and a methodology to evaluate them is established. It is then shown that the optimization problem arising in constructive placement can be reduced to several much simpler sub problems. Objective functions for locating the selected block to achieve a good layout are presented for three different metrics: the squared Euclidean, rectilinear, and Euclidean. Appropriate optimization problems are obtained and solved analytically, using efficient computation schemes. These solutions have been implemented and are used in a real VLSI chip design environment. It is shown that the squared Euclidean and the rectilinear metrics are preferable to the Euclidean one.> Shmuel Wimer, Israel Koren |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1987 | Optimal Chaining of CMOS Transistors in a Functional CellabstractWe describe an algorithm that maps a CMOS circuit diagram into an area-efficient, high-performance layout in the style of a transistor chain. It is superior to other published algorithms of this kind in terms of the class of input circuits it accepts, its efficiency, and the quality of the results it produces. This algorithm is intended for the automatic generation of basic cells in a custom or semicustom design environment, thereby removing the burden of arduous mask definition from the designer. We show how our method was used to compose cells in a row into a functional slice (e.g. an adder) that can be used in, say, a data path. Shmuel Wimer, Ron Y. Pinter, Jack A. Feldman |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1983 | HOPLA-PLA optimization and synthesis
Shmuel Wimer, N. Sharfman |
DAC | 1 |