EDBT 2026 Demo / reviewers in the wild / expert
Bülent Abali
dblp:32/3092
· DBLP profile ↗
26ranked-venue papers
6as first author
1since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 24 · 6 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
11 papers |
Memory systems · 58% Storage systems · 19% Hardware accelerators and domain-specific architectures · 8% | |
| Software engineering, system software, and programming languages
1 paper |
Operating systems · 100% |
Topics — the 30 heaviest of 38, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems › memory compression
cache compression |
1.1 | 2 | 2024 | Enterprise-Class Cache Compression Design · HPCA 2024 Touché: Towards Ideal and Efficient Cache Compression By Mitigating Tag Area Overheads · MICRO 2019 |
Memory systems
cache |
0.8 | 1 | 2024 | Enterprise-Class Cache Compression Design · HPCA 2024 |
Hardware accelerators and domain-specific architectures › domain-specific accelerator
compression accelerator |
0.4 | 1 | 2020 | Data Compression Accelerator on IBM POWER9 and z15 Processors : Industrial Product · ISCA 2020 |
Storage systems
data compression |
0.4 | 1 | 2020 | Data Compression Accelerator on IBM POWER9 and z15 Processors : Industrial Product · ISCA 2020 |
Storage systems › data compression
lossless compression |
0.4 | 1 | 2020 | Data Compression Accelerator on IBM POWER9 and z15 Processors : Industrial Product · ISCA 2020 |
Memory systems
memory compression |
0.4 | 2 | 2018 | Attaché: Towards Ideal Memory Compression by Mitigating Metadata Bandwidth Overheads · MICRO 2018 Hardware Compressed Main Memory: Operating System Support and Performance Evaluation · IEEE Trans. Computers 2001 |
Memory systems › memory hierarchy
cache hierarchy |
0.2 | 1 | 2024 | Enterprise-Class Cache Compression Design · HPCA 2024 |
Memory systems
cache management |
0.1 | 1 | 2019 | Touché: Towards Ideal and Efficient Cache Compression By Mitigating Tag Area Overheads · MICRO 2019 |
Memory systems
memory bandwidth |
0.1 | 1 | 2018 | Attaché: Towards Ideal Memory Compression by Mitigating Metadata Bandwidth Overheads · MICRO 2018 |
Memory systems
non-volatile memory |
0.1 | 1 | 2009 | Enhancing lifetime and security of PCM-based main memory with start-gap wear leveling · MICRO 2009 |
Memory systems › non-volatile memory
phase change memory |
0.1 | 1 | 2009 | Enhancing lifetime and security of PCM-based main memory with start-gap wear leveling · MICRO 2009 |
Storage systems › flash and SSD › flash memory management
wear leveling |
0.1 | 1 | 2009 | Enhancing lifetime and security of PCM-based main memory with start-gap wear leveling · MICRO 2009 |
Embedded and real-time systems › real-time scheduling › multicore scheduling
cache-aware scheduling |
0.1 | 1 | 2008 | PAM: a novel performance/power aware meta-scheduler for multi-core systems · SC 2008 |
Energy-efficient computing
energy-aware scheduling |
0.1 | 1 | 2008 | PAM: a novel performance/power aware meta-scheduler for multi-core systems · SC 2008 |
Parallel and multicore computing › parallel scheduling
resource-aware scheduling |
0.1 | 1 | 2008 | PAM: a novel performance/power aware meta-scheduler for multi-core systems · SC 2008 |
Electronic design automation › high-level synthesis
scheduling |
0.1 | 1 | 2008 | PAM: a novel performance/power aware meta-scheduler for multi-core systems · SC 2008 |
Cloud and datacenter computing › virtualization
i/o virtualization |
0.1 | 1 | 2006 | High Performance VMM-Bypass I/O in Virtual Machines · USENIX ATC, General Track 2006 |
Cloud and datacenter computing
virtualization |
0.1 | 1 | 2006 | High Performance VMM-Bypass I/O in Virtual Machines · USENIX ATC, General Track 2006 |
Performance modeling and evaluation
benchmarking |
0.0 | 2 | 2001 | Performance of Hardware Compressed Main Memory · HPCA 2001 Hardware Compressed Main Memory: Operating System Support and Performance Evaluation · IEEE Trans. Computers 2001 |
Memory systems › memory compression
hardware compressed memory |
0.0 | 1 | 2001 | Performance of Hardware Compressed Main Memory · HPCA 2001 |
Memory systems
main memory |
0.0 | 1 | 2001 | Hardware Compressed Main Memory: Operating System Support and Performance Evaluation · IEEE Trans. Computers 2001 |
Memory systems › memory compression
main memory compression |
0.0 | 1 | 2001 | Performance of Hardware Compressed Main Memory · HPCA 2001 |
Performance modeling and evaluation
cache performance modeling |
0.0 | 1 | 2008 | PAM: a novel performance/power aware meta-scheduler for multi-core systems · SC 2008 |
Interconnection networks and networks-on-chip › routing algorithms
adaptive routing |
0.0 | 1 | 1999 | A New Switch Chip for IBM RS/6000 SP Systems · SC 1999 |
Interconnection networks and networks-on-chip
switch architecture |
0.0 | 1 | 1999 | A New Switch Chip for IBM RS/6000 SP Systems · SC 1999 |
Cloud and datacenter computing › virtualization
virtual machine monitor |
0.0 | 1 | 2006 | High Performance VMM-Bypass I/O in Virtual Machines · USENIX ATC, General Track 2006 |
Parallel and multicore computing › parallel algorithms › parallel combinatorial algorithms
parallel selection |
0.0 | 1 | 1993 | Balanced Parallel Sort on Hypercube Multiprocessors · IEEE Trans. Parallel Distributed Syst. 1993 |
Parallel and multicore computing › parallel algorithms › sorting
parallel sorting |
0.0 | 1 | 1993 | Balanced Parallel Sort on Hypercube Multiprocessors · IEEE Trans. Parallel Distributed Syst. 1993 |
Operating systems › resource management
memory management |
0.0 | 1 | 2001 | Hardware Compressed Main Memory: Operating System Support and Performance Evaluation · IEEE Trans. Computers 2001 |
Operating systems › resource management › memory management
virtual memory |
0.0 | 1 | 2001 | Hardware Compressed Main Memory: Operating System Support and Performance Evaluation · IEEE Trans. Computers 2001 |
Methods — techniques the papers use, named apart from their topics
prediction-assisted adaptive compression · 0.8on-chip integration · 0.4compression · 0.4sub-ranking · 0.3hardware prediction · 0.3wear leveling · 0.1address indirection · 0.1performance prediction · 0.1cache modeling · 0.1device pass-through · 0.1queueing analysis · 0.0hardware compression · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Enterprise-Class Cache Compression DesignabstractLarger cache sizes closer to processor cores increase processing efficiency, but physical limitations restrict cache sizes at a given latency. Effective cache capacity can be expanded via the inline compression of data as it enters a lower level cache. Using the IBM Telum®processor cache hierarchy as a comparative baseline, this paper presents a custom compression scheme designed for small, line-sized data blocks, examines op-timal compressor/decompressor placement, solutions to common compression drawbacks, and proposes a tiered design blueprint to facilitate product integration. The impact of compression and prediction-assisted adaptive compression on effective cache capacity, hit rate and access latency across several typical industry workloads is explored. Alper Buyuktosunoglu, David Trilla, Bülent Abali, Deanna Postles Dunn Berger, Craig R. Walters, Jang-Soo Lee |
HPCA | 3 |
| 2020 | Data Compression Accelerator on IBM POWER9 and z15 Processors : Industrial ProductabstractLossless data compression is highly desirable in enterprise and cloud environments for storage and memory cost savings and improved utilization I/O and network. While the value provided by compression is recognized, its application in practice is often limited because it's a processor intensive operation resulting low throughput and high elapsed time for compression intense workloads.The IBM POWER9 and IBM z15 systems overcome the shortcomings of existing approaches by including a novel on-chip integrated data compression accelerator. The accelerator reduces processor cycles, I/O traffic, memory and storage footprint of many applications practically with zero hardware cost. The accelerator also eliminates the cost and I/O slots that would have been necessary with FPGA/ASIC based compression adapters. On the POWER9 chip, a single accelerator uses less than 0.5% of the processor chip area, but provides a 388x speedup factor over the zlib compression software running on a general-purpose core and provides a 13x speedup factor over the entire chip of cores. On a POWER9 system, the accelerators provide an end-to-end 23% speedup to Apache Spark TPC-DS workload compared to the software baseline. The z15 chip doubles the compression rate of POWER9 resulting in even much higher speedup factors over the compression software running on general-purpose cores. On a maximally configured z15 system topology, on-chip compression accelerators provide up to 280 GB/s data compression rate, the highest in the industry. Overall, the on-chip accelerators significantly advance the state of the art in terms of area, throughput, latency, compression ratio, reduced processor utilization, power/energy efficiency, and integration into the system stack.This paper describes the architecture, and novel elements of the POWER9 and z15 compression/decompression accelerators with emphasis on trade-offs that made the on-chip implementation possible. Bülent Abali, Bart Blaner, John J. Reilly, Matthias Klein, Craig B. Agricola, Bedri Sendir, Alper Buyuktosunoglu, Christian Jacobi 0002, William J. Starke, Haren Myneni |
ISCA | 1 |
| 2019 | Touché: Towards Ideal and Efficient Cache Compression By Mitigating Tag Area OverheadsabstractCompression is seen as a simple technique to increase the effective cache capacity. Unfortunately, compression techniques either incur tag area overheads or restrict cache block placement to only include neighboring addresses. Ideally, we should be able to place compressed cache blocks without any restrictions or overheads. Seokin Hong, Bülent Abali, Alper Buyuktosunoglu, Michael B. Healy, Prashant J. Nair |
MICRO | 2 |
| 2018 | Towards a Composable Computer SystemabstractThe recent advancement of technology in both software and hardware enables us to revisit the concept of the composable architecture in the system design. The composable system design provides flexibility to serve a variety of workloads. The system offers a dynamic co-design platform that allows experiments and measurements in a controlled environment. This speeds up the system design and software evolution. It also decouples the lifecycles of components. The design consideration includes adopting available technology with the understanding of application characteristics. With the flexibility, we show the design has the potential to be the infrastructure of both cloud computing and HPC architecture serving a variety of workloads. I-Hsin Chung, Bülent Abali, Paul Crumley |
HPC Asia | 2 |
| 2018 | Attaché: Towards Ideal Memory Compression by Mitigating Metadata Bandwidth OverheadsabstractMemory systems are becoming bandwidth constrained and data compression is seen as a simple technique to increase their effective bandwidth. However, data compressionrequires accessing Metadata which incurs additional bandwidth overheads. Even after using a Metadata-Cache, the bandwidth overheads of Metadata can reduce the benefits of compression. This paper proposes Attaché, a framework that reduces the overheads of Metadata accesses. The Attaché framework consists of two components. The first component, called the Blended Metadata Engine (BLEM), enables data and its Metadata to be accessed together. BLEM incurs additional Metadata accesses only 0.003% times and removes almost all Metadata bandwidth overheads. The second component, called theCompression Pre-dictor(COPR), predicts if the memory block is compressed. TheCOPR predictor uses a fine-grained line-level predictor, a coarse-grained page-level predictor, and a global indicator. This enables Attaché to predict the compressibility of the memory block before sending a memory read request. We implement Attaché on a memory system that uses Sub-Ranking. On average, Attaché achieves 15.3% speedup (ideal 17%) and saves 22% energy consumption (ideal 23%) when compared to a baseline system that does not employ data compression. Attaché is completely hardware-based and uses only 368KB of SRAM. Seokin Hong, Prashant J. Nair, Bülent Abali, Alper Buyuktosunoglu, Kyu-Hyoun Kim, Michael B. Healy |
MICRO | 3 |
| 2018 | Towards a Single-Host Many-GPU SystemabstractAs computation-intensive tasks such as deep learning and big data analysis take advantage of GPU based accelerators, the interconnection links may become a bottleneck. In this paper, we investigate the upcoming performance bottleneck of multi-accelerator systems, as the number of accelerators equipped with single host grows. We instrumented the host PCIe fabric to measure the data transfer and compared it with the measurements from the software tool. It shows how the data transfer (P2P) helps to avoid the bottleneck on the interconnection links, but multi-GPU performance does not scale up as expected due to the control messages. We quantify the impact of host control messages with suggestions to remedy scalability bottlenecks. We also implement the proposed strategy on Lulesh to validate the concept. The result shows our strategy can save 59.86% time cost of the kernel and 13.32% PCIe H2D payload. Ming-Hung Chen, I-Hsin Chung, Bülent Abali, Paul Crumley |
SBAC-PAD | 3 |
| 2017 | Composable architecture for rack scale big data computing
Chung-Sheng Li, Hubertus Franke, Colin Parris, Bülent Abali, Mukil Kesavan, Victor Chang 0001 |
Future Gener. Comput. Syst. | 4 |
| 2011 | High-Throughput, Lossless Data Compresion on FPGAsabstractLoss less compression is often used before writing data to a storage medium or transmitting across a transmission medium. Compression aids by saving storage space or transmission bandwidth, a decompression operation is performed when the data is subsequently read. Though this scheme has clear benefits, the execution time of compression and decompression is critical to its application in real-time systems. Software compression utilities are often slow, leading to degraded system performance. Hardware-based solutions, on the other hand, often drive large resource requirements and are not amenable to supporting future algorithmic changes. In the current article, we present a high-throughput, streaming, loss less compression algorithm and its efficient implementation on FPGAs. The proposed solution provides a peak throughput of 1GB/sec per engine, with a sustained overall measured throughput of 2.66GB/sec on a PCIe-based FPGA board with two compression and two decompression engines. This result represents an overall speedup of 13.6× over reference software implementation. The proposed design is very lean, and, with multiple engines running in parallel, provides a path to potential speedups of up to two orders of magnitude. In the current implementation, the achievable overall throughput is limited only by the available PCIe bus bandwidth. Bharat Sukhwani, Bülent Abali, Bernard Brezzo, Sameh W. Asaad |
FCCM | 2 |
| 2009 | Virtualization polling engine (VPE): using dedicated CPU cores to accelerate I/O virtualizationabstractVirtual machine (VM) technologies are making rapid progress and VM performance is approaching that of native hardware in many aspects. Achieving high performance for I/O virtualization remains a challenge, however, especially for high speed networking devices such as 10 Gigabit Ethernet 10 GbE) NICs. Traditional software-based approaches to I/O virtualization usually suffer significant performance degradation compared with native hardware. Hardware-based approaches that allow direct device accessin VMs can achieve good performance, albeit at the expense of increased hardware cost and increased complexity in achieving tasks such as VM checkpointing, migration, and record/reply.Recently, the trend in microprocessor design has shifted from achieving higher CPU frequencies to putting more cores in a single chip, thus the cost of each core is rapidly decreasing. In this paper, we propose a new I/O virtualization approach called the Virtualization Polling Engine (VPE). VPE introduces a concept called virtualization onload, which takes advantage of dedicated CPU cores to help with the virtualization of I/O devices by using an event-driven execution model with dedicated polling threads. It can significantly reduce virtualization overhead and achieve performance close to the hardware-based approaches without requiring special hardware support.Using our VPE approach, we developed a prototype called KVM-VPE to provide Ethernet virtualization support for KVM. Our experiments in a 10GbE testbed showed that VPE significantly outperformed the original KVM. In Netperf TCP tests our prototype achieved over 5 times the bandwidth for transmitting (Tx) and over 3 times the bandwidth for receiving (Rx) compared with the original KVM. KVM-VPE also supports direct user application access to the virtual Ethernet interfaces and achieved 7.4 μs end-to-end latency between two VMs on different machines in our testbed. Overall, our research demonstrated that VPE is a promising approach to high performance I/O virtualization in the coming multicore era. Jiuxing Liu, Bülent Abali |
ICS | 2 |
| 2009 | Evaluating high performance communication: a power perspectiveabstractRecently, high speed interconnects capable of remote direct memory access (RDMA) such as InfiniBand and iWARP have gained considerable popularity due to their superb latency and bandwidth. Most existing studies about RDMA have focused mainly on its performance aspect. However, as power management has become essential for high-end systems such as enterprise servers and high performance computing nodes which are often equipped with RDMA capable network adapters, it is very important for us to take a fresh look at the benefits of RDMA from the power perspective. Jiuxing Liu, Dan E. Poff, Bülent Abali |
ICS | 3 |
| 2009 | Enhancing lifetime and security of PCM-based main memory with start-gap wear levelingabstractPhase Change Memory (PCM) is an emerging memory technology that can increase main memory capacity in a cost-effective and power-efficient manner. However, PCM cells can endure only a maximum of 107 - 108 writes, making a PCM based system have a lifetime of only a few years under ideal conditions. Furthermore, we show that non-uniformity in writes to different cells reduces the achievable lifetime of PCM system by 20x. Writes to PCM cells can be made uniform with Wear-Leveling. Unfortunately, existing wear-leveling techniques require large storage tables and indirection, resulting in significant area and latency overheads. Moinuddin K. Qureshi, John P. Karidis, Michele Franceschini, Vijayalakshmi Srinivasan, Luis A. Lastras, Bülent Abali |
MICRO | 6 |
| 2008 | Sysman: A Virtual File System for Managing Clusters
Mohammad Banikazemi, David Daly, Bülent Abali |
LISA | 3 |
| 2008 | PAM: a novel performance/power aware meta-scheduler for multi-core systemsabstractSharing resources such as caches and main memory bandwidth in multi-core systems requires a more sophisticated scheduling scheme. PAM is a low-overhead, user-level meta-scheduler which does not require any hardware or software changes. In particular, it operates by detecting resource congestions and providing guidelines to the standard system scheduler by limiting the assignment of processes to subsets of available cores. PAM contains a cache model that it uses to predict the impact of new schedules. PAM can be used to improve the system along three dimensions: performance, power, and energy consumption (and any combination of these three). On our prototype, we show individual benchmarks can improve by up to 33% and the overall system performance can be improved by as much as 14%. Mohammad Banikazemi, Dan E. Poff, Bülent Abali |
SC | 3 |
| 2007 | Nomad: migrating OS-bypass networks in virtual machinesabstractVirtual machine (VM) technology is experiencing a resurgence due to various benefits including ease of management, security and resource consolidation. Live migration of virtual machines allows transparent movement of OS instances and hosted applications across physical machines. It is one of the most useful features of VM technology because it provides a powerful tool for effective administration of modern cluster environments. Migrating network resources is one of the key problems that need to be addressed in the VM migration process. Existing studies of VM migration have focused on traditional I/O interfaces such as Ethernet. However, modern high-speed interconnects with intelligent NICs pose significantly more challenges as they have additional features including hardware level reliable services and direct I/O accesses. In this paper we present Nomad, a design for migrating modern interconnects with the aforementioned features, focusing on cluster environments running VMs. We introduce a thin namespace virtualization layer to efficiently address location dependent resource handles and a handshake protocol which transparently maintains reliable service semantics during migration. We demonstrate our design by implementing a prototype based on the Xen virtual machine monitor and InfiniBand. Our performance analysis shows that Nomad can achieve efficient migration of network resources, even in environments with stringent communication performance requirements. Wei Huang 0003, Jiuxing Liu, Matthew J. Koop, Bülent Abali, Dhabaleswar K. Panda 0001 |
VEE | 4 |
| 2006 | A case for high performance computing with virtual machinesabstractVirtual machine (VM) technologies are experiencing a resurgence in both industry and research communities. VMs offer many desirable features such as security, ease of management, OS customization, performance isolation, check-pointing, and migration, which can be very beneficial to the performance and the manageability of high performance computing (HPC) applications. However, very few HPC applications are currently running in a virtualized environment due to the performance overhead of virtualization. Further, using VMs for HPC also introduces additional challenges such as management and distribution of OS images.In this paper we present a case for HPC with virtual machines by introducing a framework which addresses the performance and management overhead associated with VM-based computing. Two key ideas in our design are: Virtual Machine Monitor (VMM) bypass I/O and scalable VM image management. VMM-bypass I/O achieves high communication performance for VMs by exploiting the OS-bypass feature of modern high speed interconnects such as Infini-Band. Scalable VM image management significantly reduces the overhead of distributing and managing VMs in large scale clusters. Our current implementation is based on the Xen VM environment and InfiniBand. However, many of our ideas are readily applicable to other VM environments and high speed interconnects.We carry out detailed analysis on the performance and management overhead of our VM-based HPC framework. Our evaluation shows that HPC applications can achieve almost the same performance as those running in a native, non-virtualized environment. Therefore, our approach holds promise to bring the benefits of VMs to HPC applications with very little degradation in performance. Wei Huang 0003, Jiuxing Liu, Bülent Abali, Dhabaleswar K. Panda 0001 |
ICS | 3 |
| 2006 | High Performance VMM-Bypass I/O in Virtual Machines
Jiuxing Liu, Wei Huang 0003, Bülent Abali, Dhabaleswar K. Panda 0001 |
USENIX ATC, General Track | 3 |
| 2005 | Storage-Based Intrusion Detection for Storage Area Networks (SANs)abstractStorage systems are the next frontier for providing protection against intrusion. Since storage systems see changes to persistent data, several types of intrusions can be detected by storage systems. Intrusion detection (ID) techniques can be deployed in various storage systems. In this paper, we study how intrusions can be detected at the block storage level and in SAN environments. We propose novel approaches for storage-based intrusion detection and discuss how features of state-of-the-art block storage systems can be used for intrusion detection and recovery of compromised data. In particular we present two prototype systems. First we present a real time intrusion detection system (IDS), which has been integrated within a storage management and virtualization system. In this system incoming requests for storage blocks are examined for signs of intrusions in real time. We then discuss how intrusion detection schemes can be deployed as an appliance loosely coupled with a SAN storage system. The major advantage of this approach is that it does not require any modification and enhancement to the storage system software. In this approach, we use the space and time efficient point-in-time copy operation provided by SAN storage devices. We also present performance results showing that the impact of ID on the overall storage system performance is negligible. Recovering data in compromised systems is also discussed. Mohammad Banikazemi, Dan E. Poff, Bülent Abali |
MSST | 3 |
| 2001 | Performance of Hardware Compressed Main MemoryabstractA new memory subsystem called Memory Expansion Technology (MXT) has been built for compressing main memory contents. MXT effectively doubles the physically available memory. This paper provides an analysis of the performance impact of memory compression using the SPEC2000 benchmarks and a database benchmark. Results show that the hardware compression of memory has a negligible performance penalty compared to a standard memory. We also show that many applications' memory contents can be compressed usually by a factor of two to one. We demonstrate this using industry benchmarks, web server benchmarks, and contents of popular web sites. Bülent Abali, Hubertus Franke, Dan E. Poff, T. Basil Smith |
HPCA | 1 |
| 2001 | Adaptive Routing on the New Switch Chip for IBM SP Systems
Bülent Abali, Craig B. Stunkel, Jay Herring, Mohammad Banikazemi, Dhabaleswar K. Panda 0001, Cevdet Aykanat, Yucel Aydogan |
J. Parallel Distributed Comput. | 1 |
| 2001 | Design Alternatives for Virtual Interface Architecture and an Implementation on IBM Netfinity NT Cluster
Mohammad Banikazemi, Bülent Abali, Lorraine M. Herger, Dhabaleswar K. Panda 0001 |
J. Parallel Distributed Comput. | 2 |
| 2001 | Hardware Compressed Main Memory: Operating System Support and Performance EvaluationabstractA new memory subsystem, called Memory Xpansion Technology (MXT), has been built for compressing main memory contents. MXT effectively doubles the physically available memory transparently to the CPUs, input/output devices, device drivers, and application software. An average compression ratio of two or greater has been observed for many applications. Since compressibility of memory contents varies dynamically, the size of the memory managed by the operating system is not fixed. In this paper, we describe operating system techniques that can deal with such dynamically changing memory sizes. We also demonstrate the performance impact of memory compression using the SPEC CPU2000 and SPECweb99 benchmarks. Results show that the hardware compression of memory has a negligible performance penalty compared to a standard memory for many applications. For memory starved applications and benchmarks such as SPECweb99, memory compression improves the performance significantly. Results also show that the memory contents of many applications can be compressed, usually by a factor of two to one. Bülent Abali, Mohammad Banikazemi, Hubertus Franke, Dan E. Poff, T. Basil Smith |
IEEE Trans. Computers | 1 |
| 2000 | Efficient Virtual Interface Architecture (VIA) Support for the IBM SP Switch-Connected NT ClustersabstractThe IBM SP Switch-Connected NT cluster is one the newest clustering platforms available. In this paper, we discuss an experimental implementation of the Virtual Interface Architecture for this platform. We discuss different design issues involved in this implementation. In particular, we explain how the virtual-to-physical address translation can be implemented efficiently with a minimum Network Interface Card (NIC) memory requirement. We show how caching the VIA descriptors on the NIC can reduce the communication latency. We also present an efficient scheme for implementing the VIA door bells without any hardware support. A comprehensive performance evaluation study of the implementation is provided. The performance of the implemented VIA surpasses that of other existing software implementations of the VIA and is comparable to that of a hardware VIA implementation. The peak measured bandwidth for our system is observed to be 101.4 MBytes/s and the one-way latency for short messages is 18.2 microseconds. It is to be noted that the VIA implementation presented in this paper is not a part of any IBM product and no assumptions should be made regarding its availability as a product in the future. Mohammad Banikazemi, Vijay Moorthy, Dhabaleswar K. Panda 0001, Lorraine M. Herger, Bülent Abali |
IPDPS | 5 |
| 2000 | Adaptive Routing in RS/6000 SP-Like Bidirectional Multistage Interconnection NetworksabstractThe IBM RS/6000 SP is one of the most successful commercially available multicomputers. SP owes its success partially to the scalable, high bandwidth, low latency network. In this paper, we present the adaptive routing scheme used in the new SP network switch chip called the Switch2. We show that the adaptive routing methods outperform the oblivious routing methods on SP like multistage networks. It is shown that the adaptive routing increases the network throughout by up to 229% over oblivious routing in some cases. We also study the effect of output selection, functions on the network performance. We present six different output selection functions and study their performance for different system parameters and communication patterns through extensive simulation. The results show that three of these selection functions perform similar to each other and outperform the other selection functions consistently. Therefore, unlike previous research findings for meshes and tori, there is no need to use multiple (or hybrid) selection functions to obtain the best performance for the SP like bidirectional multistage interconnection networks. We also provide an analysis of the cost-effectiveness of different selection functions with respect to the complexity of their hardware implementations. Mohammad Banikazemi, Dhabaleswar K. Panda 0001, Craig B. Stunkel, Bülent Abali |
IPDPS | 4 |
| 1999 | A New Switch Chip for IBM RS/6000 SP SystemsabstractThis paper describes the architecture of a third-generation switching element which may appear in future IBM RS/6000 SP interconnection networks.In this paper this ASIC will be referred as the Switch3 switch chip.Like its predecessors, Switch3 is an 8-port device implementing output-queuing using the high-utilization central-buffering technique.However, Switch3 offers significant enhancements over these existing SP switch chips by incorporating advances in both VLSI technology and in recent interconnection network research.Switch3 introduces a new form of adaptive routing with the potential to significantly improve network bandwidth.It also offers support for collective communication via a powerful hardware multicast replication capability.The technology advances allow link bandwidth to be improved to 500 MB/s per direction per link, and allow the central buffer size to be doubled compared to the current SP switch.Furthermore, the larger Switch3 input buffers are capable of supporting link lengths of up to 100 meters, enabling richly-connected, scalable topologies with a high aggregate bandwidth.Finally, Switch3 offers a number of other significant enhancements including limited support for high-priority traffic and detailed performance monitoring information.1. Note that Switch3 is simply a place-holder name for this paper.It is not an internal or external IBM product name. Craig B. Stunkel, Jay Herring, Bülent Abali, Rajeev Sivaram |
SC | 3 |
| 1997 | Clock Synchronization on a Multicomputer
Bülent Abali, Craig B. Stunkel, Caroline Benveniste |
J. Parallel Distributed Comput. | 1 |
| 1993 | Balanced Parallel Sort on Hypercube MultiprocessorsabstractA parallel sorting algorithm for sorting n elements evenly distributed over 2/sup d/ p nodes of a d-dimensional hypercube is presented. The average running time of the algorithm is O((n log n)/p+p log 2n). The algorithm maintains a perfect load balance in the nodes by determining the (kn/p)th elements (k1,. . ., (p-1)) of the final sorted list in advance. These p-1 keys are used to partition the sorted sublists in each node to redistribute data to the nodes to be merged in parallel. The nodes finish the sort with an equal number of elements (n/p) regardless of the data distribution. A parallel selection algorithm for determining the balanced partition keys in O(p log2n) time is presented. The speed of the sorting algorithm is further enhanced by the distance-d communication capability of the iPSC/2 hypercube computer and a novel conflict-free routing algorithm. Experimental results on a 16-node hypercube computer show that the sorting algorithm is competitive with the previous algorithms and faster for skewed data distributions.> Bülent Abali, Füsun Özgüner, Abdulla Bataineh |
IEEE Trans. Parallel Distributed Syst. | 1 |