EDBT 2026 Demo / reviewers in the wild / expert
Lotfi Mhamdi
dblp:89/5339
· DBLP profile ↗
32ranked-venue papers
12as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 23 · 8 first-author · 8 since 2021Systems, architecture and hardware · 4 · 3 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient Early Network Intrusion Detection based on Sub-Flow Segmentation
Chonghao Pei, Lotfi Mhamdi, Ali F. Almutairi, Rami Langar |
ICC | 2 |
| 2025 | Bandwidth Fair Share Using Reinforcement Learning for IoT NetworksabstractThe exponential growth of the Internet of Things (IoT) is driven by advancements in wireless communication, data analytics, and sensor technology, leading to a proliferation of smart devices and applications. To keep pace with the massive data volumes generated by IoT networks, dynamic features and resource constraints should be carefully considered compared with traditional networks, especially for traffic Congestion Control (CC). Reinforcement Learning (RL) based congestion control has shown high potential in dealing with network congestion. However, RL-based schemes mandate expensive resource requirements (at the sender side) and impaired fairness performance problems. This paper addresses this shortcoming by proposing a lightweight RL-based Congestion Control scheme that overcomes this problem. In particular, we propose an intelligent, RL-based, “Share by Allocating switch Buffer” CC scheme capable of overcoming congestion by minimizing intermediate switch buffer overflow. Our scheme termed SAB-RL has been extensively tested and shown to tackle congestion at its root, thereby outperforming other CC proposals. Qihang Jiao, Lotfi Mhamdi |
ICC | 2 |
| 2024 | Using Siamese Networks and Autoencoders for Feature Reduction in IoT SecurityabstractIn the development of the Internet of Things (IoT), machine learning has become a significant solution to cyberattack challenges. However, the vast and heterogeneous IoT network traffic poses challenges for resource-constrained devices. To address this concern, this paper introduces the SiaAE approach, integrating the Siamese network with autoencoders, aimed at reducing the feature dimensionality of IoT network traffic data. This design enables the autoencoder to differentiate between detailed information among samples of the same or different categories while reducing dimensionality. We evaluated SiaAE on three public datasets: CICIDS2017, N-BaIoT, and TON-IoT. The results demonstrate that SiaAE outperforms traditional Principal Component Analysis (PCA) and standard autoencoders (AE) in feature reduction and class separation, thereby enhancing the accuracy of classification tasks. Specifically, SiaAE achieved accuracy rates of 99.94%, 99.97%, and 97.99% on the three datasets, respectively. Its effectiveness is proven in improving the performance of machine learning models for network security. Chonghao Pei, Lotfi Mhamdi |
GLOBECOM | 2 |
| 2024 | Securing SDN: Hybrid autoencoder-random forest for intrusion detection and attack mitigationabstractSoftware Defined Networking (SDN) has revolutionized network administration by providing centralized management through software, enabling traffic adjustment independent of the data plane. Despite the benefits, SDN networks are prone to security threats from external sources, thus necessitating the implementation of security measures. Unfortunately, most existing efforts have been just a simple mapping of earlier solutions into the SDN environments. This paper addresses the problem of SDN security based on deep learning in a purely native SDN environment, where a Deep Learning intrusion detection module is tailored to a native SDN environment. In particular, we propose a hybrid Deep AutoEncoder with a Random Forest classifier model (DAERF) to enhance intrusion detection performance in a native SDN environment. The proposed model is incorporated into a novel adaptive framework for attack mitigation in SDN environments. The proposed framework consists of a three-layer protection mechanism for detecting and preventing attacks. It is based on entropy-based detection, hybrid machine learning in the control layer and proactive services monitoring in the application layer. Experimental results have shown that our DEARF proposed autoencoder model achieved anomaly detection rates in excess of 98% in stand-alone mode as well as when incorporated within the framework, making it highly solution for next generation SDN networks. Lotfi Mhamdi, Mohd Mat Isa |
J. Netw. Comput. Appl. | 1 |
| 2023 | Hybrid Learning Blockchain assisted approach to Secure Software Defined NetworksabstractSDN has become a great self-sufficient management and network configuration paradigm. In contrast to traditional networks, SDN separates the control plane from the data plane, reducing the number of operations and the time required to add new services. Despite these benefits, modern technology also introduces risks and weaknesses. As a result, a key component of SDN design is the creation of high-performance intrusion detection systems (IDSs) to categorize hostile activity. Recently, many works have introduced Machine L/DL based solutions for intrusion detection, but the problem with Machine Learning techniques is that they can produce a relatively high rate of false positives, which is a major concern for modern IDS. While deep learning techniques have not yet demonstrated a high efficiency in intrusion detection, blockchain technology can help learning Machine learning /Deep learning models to be reused and shared with confidence as well as reducing the false positive rate generated by ML algorithms, thereby improving the efficiency of DL models in the IDS field. This article presents a successful integration of ML/DL learning with blockchain to improve the utilisation of machine learning techniques to detect intrusions in SDN networks by improving the construction of prediction models, and improving the robustness of learning-based solutions. Hédi Hamdi, Lotfi Mhamdi, Mahmood A. Mahmood |
GLOBECOM | 2 |
| 2023 | Network Intrusion Detection in Software-Defined Network Using Deep and Machine LearningabstractSoftware Defined Network (SDN) has been known for the great potential to become the development direction of a new generation network architecture. For the large-scale deployment of SDN in the future, security issues are a big concern. Accordingly, this paper applies several Machine Learning (ML)and deep learning (DL) models for Network Intrusion Detection System (NIDS), respectively, aiming at improving the accuracy performance on NIDS. The benchmark dataset, NSL-KDD, is used for evaluating the performance of the algorithms. Extensive experiments show that the F-measure rate can reach up to 87.72 % for multiclass label on NSL- KDD data set with twenty-two features using KNN algorithm. Furthermore, compared to other ML models, the k-Nearest Neighbour model has better performance under multiclass classification through numerous experiments. Lotfi Mhamdi, Hédi Hamdi, Mahmood A. Mahmood |
GLOBECOM | 1 |
| 2022 | Light-Weight Congestion Control in Constrained IoT NetworksabstractThe Internet of Things (IoT) is a growing technology that remotely connects multiple devices (ranging across many fields and applications) over the Internet. The scalability of an IoT network mandates a reliable transport infrastructure. Traditional TCP control protocol is unsuitable for such domain, mainly due to energy and power consumption reasons. A lighter version of TCP, Light Weight IP (lwIP) provides a promising solution for current and projected future scalable IoT infrastructures. However, the original lwlP is just a simple mapping of the protocol, without insight into the IoT specific requirements. This paper examines the 1wIp congestion control mechanism and addresses its shortcomings. In particular, a detailed examination is devoted to the various metrics such as Retransmission Time-Outs (RTOs) and its back-off epochs, the congestion window behaviour and progress in the absence (and presence) of congestion. In particular, we propose a set of novel algorithms to address both the IoT constraints nature (light-weight) as well as keeping up with scalability in IoT network size and performance. A detailed simulation study has been conducted to endorse the viability of our proposed set of algorithms for next-generation IoT networks. Hussam Abdul Khalek, Lotfi Mhamdi |
GLOBECOM | 2 |
| 2022 | Hybrid Deep Autoencoder with Random Forest in Native SDN Intrusion Detection EnvironmentabstractThis paper introduces a hybrid deep autoencoder with a random forest classifier model to enhance intrusion detection performance in a native SDN environment. A deep learning architecture combining a deep autoencoder with random forest learning feature representation of traffic flows natively collected from the SDN environment. Publicly available packet Capture (PCAP) files of recorded traffic flows were used in the SDN network for flow feature extraction and real-time implementation. The results show very high and consistent performance metrics, with an average of 0.9 receiver-operating characteristics area under curve (ROC AUC) recorded. Furthermore, we compared the performance achieved using the original dataset with previous research to investigate the performance achieved using the same model developed. The flow-based intrusion model presented outperforms other publicly available methods, with a traffic anomaly detection rate of 98% accuracy and precision. Mohd Mat Isa, Lotfi Mhamdi |
ICC | 2 |
| 2019 | F-DCTCP: Fair Congestion Control for SDN-Based Data Center NetworksabstractNetwork congestion control and management has long been a major issue in providing low-latency, high throughput and high link-utilisation. In particular, Data Center Network (DCN) environments, often dominated by partition-aggregate workloads, suffer performance collapse due to TCP Incast caused by inadequate congestion control parameters. This paper proposes a step in solving this problem. In particular, we describe a novel framework to overcome and control congestion in DCNs based on Software Defined Networking (SDN). We propose a native SDN-based congestion control mechanism, termed Fair Data Center TCP (F-DCTCP). As we shall see, the experimental results show that F-DCTCP outperforms all previous proposals by providing the best combined overall performance in terms of throughout, fairness and flow completion times, making it highly attractive for next generation networks. Jonathan Aina, Lotfi Mhamdi, Hédi Hamdi |
ISNCC | 2 |
| 2018 | A New Paradigm to Build Scalable Packet-Switches for Data Center NetworksabstractThis paper presents the design, implementation, and evaluation of a class of packet-switching fabric architectures. Based on the well-investigated three-stage Clos-network, we propose a variety of packet-switches that are constructed by adding the most beneficial Network-on-Chip (NoC) paradigm which offers many distinct and practical advantages. Compared to the conventional crossbar switches, the NoC-based architectures provide better path-diversity, simple packet scheduling and speedup. A gradual design method is adopted to enhance the performance of the NoC switch, and several related issues such as the congestion avoidance, micro level load-balancing, and costeffectiveness are addressed. The NoC switches exhibit a high scalability potential in - both - the port count and traffic volume, making them a good candidate for the next-generation Data Center Networks. Fadoua Hassen, Lotfi Mhamdi |
GLOBECOM | 2 |
| 2017 | High-radix packet-switching architecture for Data Center NetworksabstractWe propose a highly scalable packet-switching architecture that suits for demanding Data Center Networks (DCNs). The design falls into the category of buffered multistage switches. It affiliates the three-stage Clos-network and the Networks-on-Chip (NoC) paradigm. We also suggest a congestion-aware routing algorithm that shares the traffic load among the switch's central modules via interleaved connecting links. Unlike conventional switches, the current proposal provides better path diversity, simple scheduling, speedup and robustness to load variation. Simulation results show that the switch scales well with the port-count and traffic fluctuation and that it outperforms different switches under many traffic patterns. Fadoua Hassen, Lotfi Mhamdi |
HPSR | 2 |
| 2017 | High-capacity clos-network switch for data center networksabstractScaling-up Data Center Networks (DCNs) should be done at the network level as well as the switching elements level. The glaring reason for this, is that switches/routers deployed in the DCN can bound the network capacity and affect its performance if improperly chosen. Many multistage switching architectures have been proposed to fit for the next-generation networking needs. However all of them are either performance limited or too complex to be implemented. Targeting scalability and performance, we propose the design of a large-capacity switch in which we affiliate a multistage design with a Networks-on-Chip (NoC) design. The proposal falls into the category of buffered multistage switches. Still, it has a different architectural aspect and scheduling process. Dissimilar to common point-to-point crossbars, NoCs used at the heart of the three-stage Closnetwork allow multiple packets simultaneously in the modules where they can be adaptively transported using a pipelined scheduling scheme. Our simulations show that the switch scales well with the load and size variation. It outperforms a variety of architectures under a range of traffic arrivals. Fadoua Hassen, Lotfi Mhamdi |
ICC | 2 |
| 2017 | A scalable packet-switch architecture based on OQ NoCs for data center networks
Fadoua Hassen, Lotfi Mhamdi |
Comput. Networks | 2 |
| 2016 | Diagnosis of hybrid systems through Observers and Timed AutomataabstractThis paper deals the diagnosis of hybrid systems. The objective is the integration of these two modeling tools: Observers and Timed Automata for the design of a global diagnostic model. The latter allows to detect and locate faults, minimize repair time and provide a reliable diagnosis and easily readable despite the complexity of the equipment. As regards methodology, the work consists in automate model generation procedures and failure indicators in formal form and interchangeable to integrate them into the diagnostic system. On the implementation plan, the results were applied to a hydraulic system with two tanks. Lotfi Mhamdi, B. Maaref, Hedi Dhouibi, Hassani Messaoud, Zineb Simeu-Abazi |
CoDIT | 1 |
| 2016 | Congestion-Aware Multistage Packet-Switch Architecture for Data Center NetworksabstractData Center Networks (DCNs) have gone through major evolutionary changes over the past decades. Yet, it is still difficult to predict loads fluctuation and congestion spikes in the network switching fabric. Conventional multistage switches/routers used in data center fabrics barely deal with load balancing. Congestion management is often processed at the edge modules. However, neither the architecture of switches/routers, nor their inner routing algorithms tend to consider traffic balancing and congestion management. In this paper, we propose a flexible design of a scalable multistage switch with crossconnected UniDirectional Network-on-Chip based central blocs (UDNs). We also introduce a congestion-aware routing to forward packets adaptively. We compare the current switch architecture to the state-of-the art previous multistage switches under different traffic types. Simulations of various switch settings have shown that the proposed architecture maintains high throughput and low latency performance. Fadoua Hassen, Lotfi Mhamdi |
GLOBECOM | 2 |
| 2016 | Providing performance guarantees in data center network switching fabricsabstractThis paper proposes a novel and highly scalable multistage packet-switch design based on Networks-on-Chip (NoC). In particular, we describe a three-stage Clos packet-switch fabric with a Round-Robin packets dispatching scheme where each central stage module is an Output-Queued Unidirectional NoC (OQ-UDN), instead of the conventional single-hop crossbar. We test the switch performance under different traffic profiles. In addition to experimental results, we present an analytical approximation for the theoretical throughput of the switch under Bernoulli i.i.d arrivals. We also provide an upper-bound estimation of the end-to-end blocking probability in the proposed switch to help predict performance and to optimize the design. Fadoua Hassen, Lotfi Mhamdi |
HPSR | 2 |
| 2016 | A scalable packet-switch based on output-queued NoCs for data centre networksabstractThe switch fabric in a Data-Center Network (DCN) handles constantly variable loads. This is stressing the need for high-performance packet switches able to keep pace with climbing throughput while maintaining resiliency and scalability. Conventional multistage switches with their space-memory variants proved to be performance limited as they do not scale well with the proliferating DC requirements. Most proposals are either too complex to implement or not cost effective. In this paper, we present a highly scalable multistage switching architecture for DC switching fabrics. We describe a three-stage Clos packet-switch fabric with Output-Queued Unidirectional NoC (OQ-UDN) modules and Round-Robin packets dispatching scheme. The proposed OQ Clos-UDN architecture avoids the need for complex and costly input modules and simplifies the scheduling process. Thanks to a dynamic packets dispatching and the multi-hop nature of the UDN modules, the switch provides load balancing and path-diversity. We compared our proposed architecture to state-of-the art previous architectures under extensive uniform and non-uniform DC traffic types. Simulations of various switch settings have shown that the proposed OQ Clos-UDN outperforms previous proposals and maintains high throughput and latency performance. Fadoua Hassen, Lotfi Mhamdi |
ICC | 2 |
| 2016 | Deep learning approach for Network Intrusion Detection in Software Defined NetworkingabstractSoftware Defined Networking (SDN) has recently emerged to become one of the promising solutions for the future Internet. With the logical centralization of controllers and a global network overview, SDN brings us a chance to strengthen our network security. However, SDN also brings us a dangerous increase in potential threats. In this paper, we apply a deep learning approach for flow-based anomaly detection in an SDN environment. We build a Deep Neural Network (DNN) model for an intrusion detection system and train the model with the NSL-KDD Dataset. In this work, we just use six basic features (that can be easily obtained in an SDN environment) taken from the forty-one features of NSL-KDD Dataset. Through experiments, we confirm that the deep learning approach shows strong potential to be used for flow-based anomaly detection in SDN environments. Tuan A. Tang, Lotfi Mhamdi, Desmond C. McLernon, Syed Ali Raza Zaidi, Mounir Ghogho |
WINCOM | 2 |
| 2014 | A survey on architectures and energy efficiency in Data Center Networks
Ali Hammadi, Lotfi Mhamdi |
Comput. Commun. | 2 |
| 2013 | Scheduling multicast traffic in partially buffered crossbar switchesabstractThe tremendous increase of multicast traffic applications on the Internet has led to intensive studies for efficient multicast traffic support in Internet routers and switches. Various solutions have been proposed for the unbuffered (IQ) and buffered (CICQ) crossbar switches. However, they were either too complex to run at high speed or too expensive to be readily implemented in hardware. The Partially Buffered Crossbar (PBC) has been proposed and shown to achieve the performance of buffered crossbars while having a low cost close to that of unbuffered crossbars under unicast traffic. This paper conducts one of the first studies on multicast traffic support in PBC switches and proposes a novel round-robin based scheduling algorithm termed MSRR. Based on extensive simulations, we show that with the proposed algorithm, a PBC switch outperforms its unbuffered and buffered counterparts with a small number of internal buffers, making it highly attractive for next generation networks. Cao Di, Lotfi Mhamdi |
ISCC | 2 |
| 2010 | Performance Guarantees in Partially Buffered Crossbar SwitchesabstractMost of today's high capacity switches and Internet routers do not provide performance guarantees. This is attributed to their underlaying interconnection topology (i.e. the crossbar) and/or to their impractically complex scheduling algorithms. This paper derives a study for a Partially Buffered Crossbar (PBC) switch to practically provide throughput and fairness guarantees. We show how a PBC switch with a two-cell internal buffering per output and a speed-up of two can provide guaranteed performance on throughput under any admissible i.i.d traffic scheme. We propose a scheduling algorithm, named AF-DROP- PR, that enables us to derive bounds on the average cell latency and on the guaranteed fairness among the competing flows. AF- DROP-PR maximizes the total throughput while guaranteeing service fairness among all flows. We explore how the allocation bandwidth of individual flows is traded-off with the overall fairness index. We conjecture that the PBC switch, with its pipelined and distributed AF-DROP-PR scheduling, is attractive not only for providing performance guarantees, but also for its high capacity and low cost. Nikolaos Skalis, Lotfi Mhamdi |
GLOBECOM | 2 |
| 2009 | Internet-Router Buffered Crossbars Based on Networks on ChipabstractThe scalability and performance of the Internet depends critically on the performance of its packet switches. Current packet switches are based on single-hop crossbar fabrics, with line cards that use virtual output-queueing to reduce head-of-line blocking. In this paper we propose to use a multi-hop network on a chip (NOC) as the crossbar fabric, with FIFO-queued line cards. The use of a multi-hop crossbar fabric has several advantages. (1) Speed-up, i.e. the crossbar fabric can operate faster because NOC inter-router wires are shorter than those in a single-hop crossbar, and because arbitration is distributed instead of centralised. (2) Load balancing because paths from different input-output port pairs share the same router buffers, unlike the internal buffers of buffered crossbar fabric that are dedicated to a single inputoutput pair. (3) Path diversity allows traffic from an input port to follow different paths to its destination output port. This results in further load balancing, especially for non-uniform traffic patterns. (4) Simpler line-card design: the use of FIFOs on the line cards simplifies both the line cards and the (interchip) flow control between the crossbar fabric and line cards, reducing the number of (expensive) chip pins required for flow control. (5) Scalability, in the sense that the crossbar speed is independent of the number of ports, which is not the case for single-hop crossbar fabrics. We analyzed the performance of our architecture both analytically and by simulation, and show that it performs well for a wide range of traffic conditions and switch sizes. Additionally we prototyped a 32 × 32 NOC-based crossbar fabric in a 65 nm CMOS technology. The unoptimised implementation operates at 413 MHz, achieving an aggregate throughput in excess of 1010ATM cells per second. Kees Goossens, Lotfi Mhamdi, Iria Varela Senin |
DSD | 2 |
| 2009 | Efficient Multicast Support in Buffered Crossbars using Networks on ChipabstractThe Internet growth coupled with the variety of its services is creating an increasing need for multicast traffic support by backbone routers and packet switches. Recently, buffered crossbar (CICQ) switches have shown high potential in efficiently handling multicast traffic. However, they were unable to deliver optimal performance despite their expensive and complex crossbar fabric. This paper proposes an enhanced CICQ switching architecture suitable for multicast traffic. Instead of a dedicated internal crosspoint buffer for every input-output pair of ports, the crossbar is designed as a multi-hop Network on Chip (NoC). Designing the crossbar as a NoC offers several advantages such as low latency, internal fabric load balancing and path diversity. It also obviates the requirement of the virtual output queuing by allowing simple FIFO structure without performance degradation. We designed appropriate routing for the NoC as well as on-chip router scheduling and tested its performance under a wide range of input multicast traffic. Simulations results showed that our proposal outperforms the CICQ architecture and offers a viable architectural alternative. We also studied the effect of various parameters such as the depth of the NoC as well as the speedup requirement for high-bandwidth multicast switching. Iria Varela Senin, Lotfi Mhamdi, Kees Goossens |
GLOBECOM | 2 |
| 2009 | Distributed parallel scheduling algorithms for high-speed virtual output queuing switchesabstractThis paper presents a novel scalable switching architecture for input queued switches with its proper arbitration algorithms. In contrast to traditional switching architectures where the scheduler is implemented by one single centralized scheduling device, the proposed architecture connects several single scheduling devices in series and a distributed scheduling algorithm is run sequentially on them, whereby the inputs of each single scheduling device build connections to a group of outputs, considering both their local transmission requests as well as global outputs availability information. We show that a pipeline pattern can be used to increase the efficiency of the scheduling scheme with scheduling algorithms running in parallel on all the separate scheduling devices. We first introduce a distributed parallel round robin scheduling algorithm (DPRR) for the proposed architecture. Through the analysis of simulation results on various admissible traffics, it is shown that the performance of DPRR is much better than, or very close to the performance of, other round robin scheduling algorithms. We also prove that under Bernoulli i.i.d. uniform traffic DPRR achieves 100% throughput. Secondly, we introduce a distributed parallel round robin scheduling algorithm with memory (DPRRM) as an improved version of DPRR to make it stable under any admissible traffic. Lotfi Mhamdi, Mounir Hamdi |
ISCC | 1 |
| 2009 | PBC: A Partially Buffered Crossbar Packet SwitchabstractThe crossbar fabric is widely used as the interconnect of high-performance packet switches due to its low cost and scalability. There are two main variants of the crossbar fabric: unbuffered and internally buffered. On one hand, unbuffered crossbar fabric switches exhibit the advantage of using no internal buffers. However, they require a complex scheduler to solve input and output ports contention. Internally, buffered crossbar fabric switches, on the other hand, overcome the scheduling complexity by means of distributed schedulers. However, they require expensive internal buffers-one per crosspoint. In this paper, we propose a novel architecture, namely, the partially buffered crossbar (PBC) switching architecture, where a small number of separate internal buffers are maintained per output. Our goal is to design a PBC switch having the performance of buffered crossbars and a cost comparable to that of unbuffered crossbars. We propose a class of round robin scheduling algorithms for the PBC architecture. Simulations results show that using as few as eight buffers per fabric column and irrespective of the number N of input ports of the switch, we can achieve similar performance to buffered crossbars that use N buffers per fabric output. Lotfi Mhamdi |
IEEE Trans. Computers | 1 |
| 2009 | On the Integration of Unicast and Multicast Cell Scheduling in Buffered Crossbar SwitchesabstractInternet traffic is a mixture of unicast and multicast flows. Integrated schedulers capable of dealing with both traffic types have been designed mainly for Input Queued (IQ) buffer-less crossbar switches. Combined Input and crossbar queued (CICQ) switches, on the other hand, are known to have better performance than their buffer-less predecessors due to their potential in simplifying the scheduling and improving the switching performance. The design of integrated schedulers in CICQ switches has thus far been neglected. In this paper, we propose a novel CICQ architecture that supports both unicast and multicast traffic along with its appropriate scheduling. In particular, we propose an integrated round-robin-based scheduler that efficiently services both unicast and multicast traffic simultaneously. Our scheme, named multicast and unicast round robin scheduling (MURS), has been shown to outperform all existing schemes under various traffic patterns. Simulation results suggested that we can trade the size of the internal buffers for the number of input multicast queues. We further propose a hardware implementation of our algorithm for a 16 times 16 buffered crossbar switch. The implementation results suggest that MURS can run at 20 Gbps line rate and a clock cycle time of 2.8 ns, reaching an aggregate switching bandwidth of 320 Gbps. Lotfi Mhamdi |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2008 | A Partially Buffered Crossbar packet switching architecture and its schedulingabstractThe crossbar fabric is widely used as the interconnect of high-performance packet switches due to its low cost and scalability. There are two main variants of the crossbar fabric: unbuffered and internally buffered. On one hand, unbuffered crossbar fabric switches exhibit the advantage of using no internal buffers. However, they require a centralized and complex scheduler. Internally buffered crossbar fabric switches, on the other hand, overcome the scheduling complexity by means of distributed schedulers. However, they require expensive internal buffers-one per crosspoint. In this paper we propose a novel architecture, namely the Partially Buffered Crossbar (PBC) switching architecture, where a small number of separate internal buffers are maintained per output. Our goal is to design a PBC switch having the performance of buffered crossbars and a cost comparable to that of unbuffered crossbars. We propose a class of round-robin scheduling algorithms for the PBC switch. Simulations results show that using as few as 8 buffers per output port and irrespective of the number,N, of input ports of the switch, we can achieve even better performance than buffered crossbars that use N buffers per output port. Lotfi Mhamdi |
ISCC | 1 |
| 2006 | A reconfigurable hardware based embedded scheduler for buffered crossbar switchesabstractIn this paper, we propose a new internally buffered crossbar (IBC) switching architecture where the input and output distributed schedulers are embedded inside the crossbar fabric chip. As opposed to previous designs, where these schedulers are spread across input and output line cards, our design allows the schedulers to have cheap and fast access to the internal buffers, optimizes the flow control mechanism and makes the IBC more scalable. We employed the Xilinx Virtex-4FX platform to show the feasibility of our proposal and implemented a reconfigurable hardware based IBC switch with the maximum port count that we could fit on a single chip. The experiments suggest that a 24x24 IBC switch running a 10 Gbps port speed and a clock cycle time of 6.4 ns can be implemented. Lotfi Mhamdi, Christoforos Kachris, Stamatis Vassiliadis |
FPGA | 1 |
| 2006 | High-performance switching based on buffered crossbar fabrics
Lotfi Mhamdi, Mounir Hamdi, Christoforos Kachris, Stephan Wong, Stamatis Vassiliadis |
Comput. Networks | 1 |
| 2004 | Scheduling multicast traffic in internally buffered crossbar switchesabstractScheduling multicast traffic has been an active research topic due to the tremendous growth of multicast traffic (audio, video, teleconferencing, etc.) on the Internet. Considerable research work has been done on input queued (IQ) switches to handle the multicast traffic. Unfortunately, all the proposed solutions were of no practical value because they either lack performance or were simply not practical. Internally buffered crossbar (IBC) switches, on the other hand, have been considered as a robust alternative to buffer-less crossbar switches to improve the switching performance. However, no work has ever been done on multicasting in IBC switches. In this paper, we fill this gap and study, for the first time, the multicasting problem in IBC switches. In particular, we propose a simple scheduling scheme named multicast cross-points round robin (MXRR) for the IBC switch architecture. Our scheme was shown to handle multicast traffic more efficiently and far better than all previous schemes. Yet, MXRR is both practical and achieves high performance. Lotfi Mhamdi, Mounir Hamdi |
ICC | 1 |
| 2003 | Output queued switch emulation by a one-cell-internally buffered crossbar switchabstractThe output queued (OQ) switching architecture shows optimal performance amongst all queuing approaches. However, OQ switches lack scalability due to high memory-bandwidth constraints. An OQ switch can be exactly emulated by a more scalable crossbar switch (i.e., input-queued - IQ - switch) and a small speedup (Chuang, S. et al., 1998). Unfortunately, this result was not of practical use due to the high complexity of the proposed scheduling scheme. A similar result was shown by B. Magill et al. (see Conf. on Commun. Control and Computing, 2002) and was based on the internally buffered crossbar (IBC) switching architecture. While the latter result seems to overcome the complexity issue, the scheduling scheme presented is costly. We extend our previous work (Mhamdi and Hamdi, IEEE ICC'03, vol.3, p.1659-63, 2003) and prove the same result as Magill et al., but with lower hardware requirements.. In particular, we propose a simple scheduling scheme, named modified current arrival first-lowest TTL (time-to-live) first (MCAF-LTF), that does not require a costly time stamping mechanism. Based on the MCAF-LTF, we prove that, with a speedup of just 2, a one-cell-internally buffered crossbar switch can exactly emulate an OQ switch. The reduced complexity of our proposed scheme makes it of high practical value and allows it to be readily implemented in ultra-high capacity networks. Lotfi Mhamdi, Mounir Hamdi |
GLOBECOM | 1 |
| 2003 | Practical scheduling algorithms for high-performance packet switchesabstractAs buffer-less scheduling algorithms reach their practical limitations due to higher port numbers and data rates, buffered crossbars gained a lot of interest recently because of the great potential they have in solving the complexity and scalability issues faced by their buffer-less predecessors. In particular, the internally buffered switching architecture was shown, through distributed scheduling algorithms, to be able to sustain the current and expected increases in Internet throughput rates. In this paper, we propose a class of distributed scheduling algorithms for the internally buffered crossbar switching architecture. As will be shown, the distributed nature of these algorithms makes them of high practical value. That is, they can be implemented in real-time for high-speed input traffic. In addition, we will demonstrate, through simulation, that these scheduling algorithms outperform state-of-the-art related algorithms in this area. Lotfi Mhamdi, Mounir Hamdi |
ICC | 1 |