VLDB 2026 Research / reviewers in the wild / expert
Roberto Rojas-Cessa
dblp:50/3423
· DBLP profile ↗
65ranked-venue papers
17as first author
11since 2021 · last 2026
0000-0001-6075-9869ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 44 · 14 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GATE: Optimal Profit vs Maximum Application Throughput in Data Centers
Jorge Medina, Chuan-Bi Lin, Roberto Rojas-Cessa, H. Jonathan Chao |
HPSR | 3 |
| 2026 | Threshold-based Wavelet Decomposition for 3D Point Clouds in Digital Environments
Mikhail I. Smirnov, Ziqian Dong, Roberto Rojas-Cessa |
ICC | 3 |
| 2025 | Partitioning Prompts for Higher Efficacy in Network Design with Large Language ModelabstractIn this paper, we propose deliverable partitioning in prompt design to assist Large Language Models (LLMs) in improving response correctness for network design and configuration. While recent research has explored the use of LLMs to enhance network management efficiency, their responses often remain inconsistent, incomplete, or inaccurate. Often, LLM-generated configurations contain missing or erroneous configuration commands, which can lead to operational failures. Our proposed partitioning methodology aims to mitigate these issues by decomposing complex network configuration tasks into simplified and focused tasks. To evaluate the effectiveness of this approach, we introduce a scoring policy and conduct extensive experiments across three levels of network complexity and varying degrees of design choice ambiguity. We also compare the performance of leading LLMs, including ChatGPT, Copilot, and DeepSeek. Our findings indicate that partitioning the inquiry process leads to more accurate and consistent responses than non-partitioned approaches, especially in scenarios where design parameters are explicitly defined and leave some but small room, as ambiguity, for inference. Vishnu Komanduri, Scott Alessio, Sebastian Estropia, Gokhan Yerdelen, Tyler Ferreira, Murali Gunti, Ziqian Dong, Roberto Rojas-Cessa |
HPSR | 8 |
| 2025 | eFlight: RL Scheme for Autonomous Drones to Efficiently Fly through ObstaclesabstractThe flight time of uncrewed autonomous vehicles (UAVs) is constrained by its battery capacity, restricting its application in long-duration missions. To address this challenge, we propose eFlight, a hybrid scheme that uses a reinforcement-learning heuristic to augment A* for path finding. eFlight reduces both node expansions and computation time while finding energy-efficient paths in obstacle-dense 3D airspace. We compare eFlight with conventional path-planning algorithms for point-to-point flights on areas of various dimensions and with various obstacle densities. The results show that eFlight achieves a dual advantage: finding low-energy paths with short computation times. In high-density obstacle environment, eFlight identifies the lowest energy consumption path in 89.5% of the trials. Compared to the baseline scheme, eFlight reduces computation time by 90.6% ± 26.6% and energy by 7.13% ± 9.96%. Yihan Xu 0003, Chuan-Bi Lin, Cong Wang 0015, Ziqian Dong, Roberto Rojas-Cessa |
SEC | 6 |
| 2024 | Comparing Link Sharing and Flow Completion Time in Traditional and Learning-based TCPabstractCongestion control design in TCP has primarily focused on maximizing throughput, reducing delay, or minimizing packet loss. Such has been the case in the surge of TCP approaches using machine and deep learning. However, flow completion time, average throughput, and fairness index are the key performance indicators more noticeable to users and used for applications, and thus must be evaluated. We theorize, that an ideal congestion control scheme would have a small average flow completion time, and high average throughput and high fairness index. We aim to analyze the performance of a wide-variety of congestion control schemes to determine the importance of these metrics in designing a congestion control scheme. With this objective, we propose a modified reinforcement learning version of TCP; RL-TCP+, to demonstrate how flow completion time can be minimized and to evaluate it’s impact on bandwidth sharing. Through extensive experimentation, we show that greater link-sharing and fairness do not always result in lower flow completion time, and that flow-prioritization could prove beneficial in certain scenarios. Vishnu Komanduri, Cong Wang 0015, Roberto Rojas-Cessa |
HPSR | 3 |
| 2023 | PEAK: Policy Event Assessment of COVID-19 Cases at the Start of the Pandemic in New York CityabstractThe impact of events and associated public health announcements on COVID-19 incidence remains an interesting and open question for future response and prevention. To address this issue, we propose a policy event impact assessment framework (PEAK) that quantifies the impact of policies and events on COVID-19 incidence in this paper. PEAK uses timeseries change point detection to estimate how health policies and events affected COVID-19 incidence during the most difficult period of the pandemic experienced in New York City and uses the long short-term memory for impact analysis at each change point. We analyze 26 public announcements on COVID19 and events that occurred in New York City from March 2020 to February 2021. The results show the top 10 largest-impact change points identified by PEAK and the events that caused such impacts. Amit Hiremath, Ziqian Dong, Roberto Rojas-Cessa |
ICTAI | 3 |
| 2023 | DICE: Data Imputation for Cost Estimates from Multiple Sources to Model User Decision-MakingabstractUnderstanding key factors that affect users’ commute mode choice is essential to design policies that promote sustainable transportation. However, the reliance on survey data for these studies often faces incomplete data challenges. One of the regional transportation surveys obtained for the study on commute mode decision-making misses 97% of the parking cost data, an important factor in people’s decision-making. To tackle the problem, we propose the data imputation for cost estimates (DICE) scheme to synthesize data from multiple sources to infer the missing data. DICE linearly maps imputed values to missing entries based on the assumption that higher-income users can spend more on their commute. In the absence of ground truth data, we propose to use the accuracy of the regression model trained with the imputed data as a metric to evaluate DICE. We train the regression model with 75% of the imputed data, test it with the remainder, and evaluate it with the complete cases. The prediction accuracy of the test data and the evaluation data are 0.89 and 0.77, respectively. The results indicate that the imputed data and complete cases share similar distributions and the model trained with the imputed data can perform classification. We tested DICE using a 1995 transportation survey and a 2021 housing survey data sets where cost is considered a key feature in decision-making. In both cases, the regression model achieves higher than 0.7 prediction accuracy, which proves the applicability of DICE on different data sets. Hailun Wu, Ziqian Dong, Roberto Rojas-Cessa |
ICTAI | 3 |
| 2023 | Effect of the incident angle of a transmitting laser light on the coverage of a NLOS-FSO network
Paa Kwesi Esubonteng, Roberto Rojas-Cessa |
Comput. Networks | 2 |
| 2022 | A Machine Learning Approach to Estimating Queuing Delay on a Router over a Single-Hop PathabstractQueuing delay is a dynamic network parameter that plays an important role in defining the performance of Internet applications over an end-to-end path. However, measurement of queuing delay is challenging because it requires a large infrastructural support from the path under test. In this paper, we propose an active scheme to measure queuing delay on a router using a probe-gap model. The scheme uses a popular data-clustering algorithm to process its data samples; therefore, its measurement efficacy is not dependent on the issues related to infrastructural access, certain variations (e.g., compression) in the probe gaps, and the number of clusters in the data processing. Here, we present a detailed evaluation of the scheme against the current state-of-the-art on a single-hop path through ns-3 simulation. Our results show that the proposed scheme is robust, consistent, quick, and highly accurate under different traffic conditions. Travis Ricker, Khondaker Musfakus Salehin, Alex Chen, Eiji Oki, Roberto Rojas-Cessa |
ICC | 6 |
| 2022 | Multidepot Drone Path Planning With Collision AvoidanceabstractIntersections of flight paths in multidrone missions are indications of a high likelihood of in-flight drone collisions. This likelihood can be proactively minimized during path planning. This article proposes two offline collision-avoidance multidrone path-planning algorithms: 1) DETACH and 2) STEER. Large drone tasks can be divided into smaller ones and carried out by multiple drones. Each drone follows a planned flight path that is optimized to efficiently perform the task. The path planning of the set of drones can then be optimized to complete the task in a short time, with minimum energy expenditure, or with maximum waypoint coverage. Here, we focus on maximizing waypoint coverage. Different from existing schemes, our proposed offline path-planning algorithms detect and remove possible in-flight collisions. They are based on a constrained nearest-neighbor search algorithm that aims to cover a large number of waypoints per flight path. DETACH and STEER perform vector intersection check for flight path analysis, but each at different stages of path planning. We evaluate the waypoint coverage of the proposed algorithms through a novel profit model and compare their performance on a work area with different waypoint densities. Our results show that STEER covers 40% more waypoints and generates 20% more profit than DETACH in high-density waypoint scenarios. Kun Shen, Rutuja Shivgan, Jorge Medina, Ziqian Dong, Roberto Rojas-Cessa |
IEEE Internet Things J. | 5 |
| 2021 | Indirect Line-Of-Sight Free-Space Optical Communications Using Diffuse ReflectionabstractFree-space optical communications (FSOC) has been proved to achieve the highest data rate in wireless communications, and yet, its long-time adoption remains reserved for a very selective set of applications. A culprit of the limited adoption of FSOC is its point-to-point link setup used by this technology as a result of the required a) the direct Line-of-Sight (LOS) between transmitter and receiver, and b) the needed fine alignment of the transceivers of the communicating stations. Indirect LOS FSOC (ID-FSOC) is an alternative to FSOC where transmitter and receiver might not be in LOS of each other but to a diffuse reflector (DR). Such a reflector not only helps to cover out-of-sight areas but also converts the optical link onto a broadcast channel where one station may communicate with all other stations that may have LOS to the reflector. Here, we discuss some of the properties of this novel paradigm and argue that it may represent a new facet of free space of optical communications that provides a great range of applications for high-speed data communications. Roberto Rojas-Cessa |
HPSR | 1 |
| 2020 | PEQ: Scheduling Time-Sensitive Data-Center Flows using Weighted Flow Sizes and DeadlinesabstractWe propose a deadline-aware flow scheduling scheme, called Preemptive Efficient Queuing (PEQ), which takes both the deadlines and sizes of the flows into account for efficient flow scheduling of flows in a data-center network in this paper. PEQ benefits from evaluating the deadlines and flow sizes of flows to improve application throughput. Our results show that PEQ outperforms the state-of-the-art deadline-aware transport schemes in terms of the application throughput and average flow completion time, under traffic with different deadline distributions. Results also show that PEQ yields a near-optimal application throughput in cases where short flows have long deadlines and long flows have strict deadlines. In other words, PEQ attains higher performance than the counterpart schemes on traffic with broader distribution of deadlines and flow sizes. Vinay K. Gopalakrishna, Yagiz Kaymak, Chuan-Bi Lin, Roberto Rojas-Cessa |
HPSR | 4 |
| 2020 | Effectiveness of Many-to-Many GRASP-Based Routing Algorithms for Power DistributionabstractIn this paper we propose three modified versions of the Greedy SmAlleSt-cost Path first (GRASP) algorithm for minimizing the total transmission cost in a digital microgrid (DMG). The Simple Dynamic GRASP (SDG) is a cached-version of GRASP that dynamically updates the available capacity of links. The total path Transmission cost Dynamic GRASP (TDG) is a cached version of GRASP that uses the path's transmission cost as the criteria to select the smallest-cost paths. The Brute Force TDG (BFT) is a cached version of GRASP that explores all paths between loads and sources and selects the smallest-cost paths by evaluating the path's transmission costs. Our results show that SDG achieves the smallest total transmission cost and the fewest unsatisfied loads. Although TDG and BFT may achieve similar total transmission costs to those of SDG and GRASP, our results show that using the path's transmission costs in the selection of the smallest-cost path is not as effective as using the path's link costs. However, in more complex power networks, we show that the cached TDG and BFT yield fewer unsatisfied loads than that of a GRASP implementation without a cache. Jorge Medina, Zhengqi Jiang, Roberto Rojas-Cessa |
HPSR | 3 |
| 2019 | TRIDENT: A Load-Balancing Clos-Network Packet Switch With Queues Between Input and Central Stages and In-Order ForwardingabstractWe propose a three-stage load balancing packet switch and its configuration scheme. The input- and central-stage switches are bufferless crossbars, and the output-stage switches are buffered crossbars. We call this switch ThRee-stage Clos-network swItch with queues at the middle stage and DEtermiNisTic scheduling (TRIDENT), and the switch is cell based. The proposed configuration scheme uses predetermined and periodic interconnection patterns in the input and central modules to load-balance and route traffic, therefore, it has low configuration complexity. The operation of the switch includes a mechanism applied at input and output modules to forward cells in sequence. TRIDENT achieves 100% throughput under uniform and nonuniform admissible traffic with independent and identical distributions (i.i.d.). The switch achieves this high performance using a low-complexity architecture while performing in-sequence forwarding and no central-stage expansion or memory speedup. We analyze the operations the configuration mechanisms perform on the traffic traversing the switch. We use this analysis to prove that the switch achieves 100% through under i.i.d. traffic. We also show that the switch forward cells in-sequence. We present a simulation analysis as a practical demonstration of the switch performance under uniform and nonuniform i.i.d. traffic. Oladele Theophilus Sule, Roberto Rojas-Cessa |
IEEE Trans. Commun. | 2 |
| 2019 | A Split-Central-Buffered Load-Balancing Clos-Network Switch With In-Order ForwardingabstractWe propose a configuration scheme for a load-balancing Clos-network (LBC) packet switch that has split central modules and buffers in between the split modules. Our split-central-buffered LBC switch is cell-based. The switch has four stages, namely input, central-input, central-output, and output stages. The proposed configuration scheme uses a pre-determined and periodic interconnection pattern in the input and split central modules to load-balance and route traffic. The LBC switch has low configuration complexity. The operation of the switch includes a mechanism applied at input and split-central modules to forward cells in sequence. The switch achieves 100% throughput under uniform and nonuniform admissible traffic with independent and identical distributions (i.i.d.). The switch uses no speedup nor memory expansion. We demonstrate the properties of the switch through traffic and timing analysis. Oladele Theophilus Sule, Roberto Rojas-Cessa, Ziqian Dong, Chuan-Bi Lin |
IEEE/ACM Trans. Netw. | 2 |
| 2018 | Optimal Positioning of Ground Base Stations in Free-Space Optical Communications for High-Speed TrainsabstractIn this paper, we propose two different free-space-optics (FSO) coverage models for next-generation high-speed-train communications. To the best of our knowledge, these are the first coverage models proposed for FSO seamless handover. The models provide different coverage areas for performing seamless signal handover and uninterrupted ground-to-train communication. The first model uses two different wavelengths in adjacent covered areas and the second one uses a single wavelength. We find the optimal distance from the train track to a ground base station and the distance between base stations to provide seamless connectivity and handover while minimizing the number of base stations along the track. We base our estimations on a realistic model of an FSO system and provide numerical evaluations demonstrating the performance of the proposed coverage models. We show the different amounts of received power on ground-to-train communications as a function of the location of ground base stations. We also consider the effect of fog on the FSO link as the most attenuating condition for FSO communications. Our results show that communication rates of 1 Gpbs and higher may be achieved with the proposed station positioning and coverage models. Sina Fathi Kazerooni, Yagiz Kaymak, Roberto Rojas-Cessa, Jianghua Feng, Nirwan Ansari, MengChu Zhou, Tairan Zhang |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2017 | SRA: Slot reservation announcement scheme for medium access control of IEEE 802.11 crowded networks in emergency scenariosabstractIn this paper we propose the Slot Reservation Announcement (SRA) access scheme to minimize the occurrence of channel-access collisions in infrastructure IEEE 802.11 networks under emergency and crowded scenarios. SRA is based on slot reservation in lieu of a randomized backoff approach to increase the channel-access success ratio. The proposed scheme rids IEEE 802.11-like access networks of throughput and utilization collapse under crowded scenarios. In turn, it increases channel utilization. This enabled access is critical for crowded networks under emergency scenarios where many stations suddenly contend for the channel. This scheme may be used as fallback mechanism to avoid throughput collapse in such scenarios and to enable critical communications. Our simulation results show maximum bandwidth utilization and increased throughput as compared to an IEEE 802.11 network and a recent reservation-based access scheme under crowded conditions. Sina Fathi Kazerooni, Roberto Rojas-Cessa |
ICC | 2 |
| 2017 | Reducing Frequency of Request Communications with Pro-Active and Aggregated Power Management for the Controlled Delivery Power GridabstractWe present a feasibility analysis of the controlled delivery power grid (CDG) that uses aggregated power request by users to reduce communications overhead. The CDG, as an approach to the power grid, uses a data network to communicate requests and grants of power in the distribution of electrical power. These requests and grants allow the energy supplier know the power demand in advance and to designate the loads and the time when power is supplied to them. Each load is assigned a power-network address that is used for communication of requests and grants with the energy supplier. With addressed loads, power is only delivered to selected loads. However, issuing a request for power before delivery takes place requires knowing the demand of power the load consumes during the operation interval. However, it is a general concern that having issuing requests in a time-slot basis may risk request losses and therefore, generate intermittent supply. Therefore, we propose request aggregation to minimize the number of requests issued. We show by simulation that the CDG with request aggregation attains high performance, in terms of satisfaction ratio and waiting time for power supply. Haard Shah, Matthew Petrula, Roberto Rojas-Cessa, Haim Grebel |
MASS | 3 |
| 2016 | Disjoint Superposition for Reduction of Conjoined Prefixes in IP Lookup for Actual IPv6 Forwarding TablesabstractHelix is a recently-proposed scheme that performs IP lookup in a single memory access. Helix uses parallel prefix matching at the different prefix lengths and the position of prefixes in a binary tree for reducing the amount of memory used. The scheme enables fast table updates as prefixes are kept in their original form. In Helix, a large number of prefixes is stored in a very small amount of memory and route updates, as the lookup process, is performed in a single memory access. Most IPv6 testings of IP lookup schemes are performed on forwarding tables generated synthetically from IPv4 tables as IPv6 tables have a small prefix count. However, the prefix distribution of address blocks in actual tables may be correlated and that may result in using large amounts of memory to represent them. Here, we proposed the application of disjoint superpositions in Helix to further reduce the amount of memory used to represent these forwarding tables. We show that under IPv6 forwarding tables, Helix prevails in performing lookup operations in a single memory access time. Roberto Rojas-Cessa, Taweesak Kijkanjanarat, Wara Wangchai, Krutika Patil, Narathip Thirapittayatakul |
GLOBECOM | 1 |
| 2016 | Method for measuring the packet processing time of Internet workstations with the detection of interrupt coalescenceabstractThe packet processing time (PPT) of an end host (i.e., workstation) is the time elapsed between the arrival of a packet at the data-link layer and the time the packet is processed at the application layer (RFCs 2679). A recent work presented an active scheme to measure PPT over Internet paths using a packet-pair based probing structure considering its importance as a network parameter in the Internet. However, the existing scheme does not consider the effect of interrupt coalescence (IC) in network interface cards (NICs), which are configured with an IC under high transmission speeds. In this paper, we propose an enhancement to the existing scheme for measuring PPT when there is an IC available in the NIC of the workstation under test. The enhanced scheme first detects IC and then measures PPT using the measured gaps of the probing structure. We evaluated the enhanced scheme through testbed experiments using two different workstations and under different transmission speeds, i.e., 10, 100, and 1000 Mb/s. Our results show that the proposed scheme consistently measures PPT with high efficacy. Khondaker Musfakus Salehin, Roberto Rojas-Cessa |
HPSR | 2 |
| 2015 | Helix: IP lookup scheme based on helicoidal properties of binary trees
Roberto Rojas-Cessa, Taweesak Kijkanjanarat, Wara Wangchai, Krutika Patil, Narathip Thirapittayatakul |
Comput. Networks | 1 |
| 2015 | Scheme to Measure Packet Processing Time of a Remote Host through Estimation of End-Link CapacityabstractAs transmission speeds increase faster than processing speeds, the packet processing time (PPT) of a host is becoming more significant in the measurement of different network parameters in which packet processing by the host is involved. The PPT of a host is the time elapsed between the arrival of a packet at the data-link layer and the time the packet is processed at the application layer (RFCs 2679 and 2681). To measure the PPT of a host, stamping the times when these two events occur is needed. However, time stamping at the data-link layer may require placing a specialized packet-capture card and the host under test in the same local network. This makes it complex to measure the PPT of remote end hosts. In this paper, we propose a scheme to measure the PPT of an end host connected over a single- or multiple-hop path and without requiring time stamping at the data-link layer. The proposed scheme is based on measuring the capacity of the link connected to the host under test. The scheme was tested on an experimental testbed and in the Internet, over a U.S. inter-state path and an international path between Taiwan and the U.S. We show that the proposed scheme consistently measures PPT of a host. Khondaker Musfakus Salehin, Roberto Rojas-Cessa, Chuan-Bi Lin, Ziqian Dong, Taweesak Kijkanjanarat |
IEEE Trans. Computers | 2 |
| 2014 | Containing sybil attacks on trust management schemes for peer-to-peer networksabstractIn this paper, we introduce a framework to detect possible sybil attacks against a trust management scheme of peer-to-peer (P2P) networks used for limiting the proliferation of malware. Sybil attacks may underscore the effectivity of such schemes as malicious peers may use bogus identities to artificially manipulate the reputation, and therefore, the levels of trust of several legitimate and honest peers. The framework includes a k-means clustering scheme, a method to verify the transactions reported by peers, and identification of possible collaborations between peers. We prove that as the amount of public information on peers increases, the effectivity of sybil attacks may decrease. We study the performance of each of these mechanisms, in terms of the number of infected peers in a P2P network, using computer simulation. We show the effect of each mechanism and their combinations. We show that the combination of these schemes is effective and efficient. Lin Cai 0003, Roberto Rojas-Cessa |
ICC | 2 |
| 2014 | DAQ: Deadline-Aware Queue scheme for scheduling service flows in data centersabstractWe propose a scheme to schedule the transmission of data center traffic to guarantee a transmission rate for long flows without affecting the rapid transmission required by short flows. We call the proposed scheme Deadline-Aware Queue (DAQ). The traffic of a data center can be broadly classified into long and short flows, where the terms long and short refer to the amount of data to be transmitted. In a data center, the long flows require modest transmission rates to keep maintenance, data updates, and functional operation. Short flows require either fast service or be serviced within a tight deadline. Satisfaction of both classes of bandwidth demands is needed. DAQ uses per-class queues at supporting switches, keeps minimum flow state information, and uses a simple but effective flow control. The credit-based flow control, employed between switch and data sources, ensures lossless transmissions. We study the performance of DAQ and compare it to those of other existing schemes. The results show that the proposed scheme improves the achievable throughput for long flows up to 37% and the application throughput for short flows up to 33% when compared to other schemes. DAQ guarantees a minimum throughput for long flows despite the presence of heavy loads of short flows. Cong Ding 0006, Roberto Rojas-Cessa |
ICC | 2 |
| 2013 | Hybrid optoelectronic packet switch with multiple wavelength conversion through an electronic packet switchabstractIn this paper, we propose an optoelectronic switch that resolves contention in an optical switch through an electronic switch. The optoelectronic switch switches packets in the optical domain, and uses an electronic switch to store and forward packets that lose contention through other wavelengths. We investigate two modalities of the electronic switch, namely, local and global, and compare the performance of these modalities, in terms of packet loss rate. The simulation results show that the global modality achieves the lower packet loss rate. We also show that a small number of ports for wavelength conversion suffices to achieve a low packet loss rate with global modality. Ziqian Dong, Roberto Rojas-Cessa |
HPSR | 2 |
| 2013 | Minimizing scheduling complexity with a Clos-network space-space-memory (SSM) packet switchabstractIn this paper we propose a three-stage space-space-memory (SSM) Clos-network switch that uses crosspoint buffers in the third-stage modules to eliminate the need for performing multiple iterations in for port matching. We show that the proposed switch not only reduces the configuration complexity of space-space-space (S3) switches but also improves switching performance and relaxes configuration timing. We demonstrate these advantages by comparing the performance of the proposed switch using the weighted module-first no-port (WFM-NP) matching scheme to that of a S3switch using the original scheduling scheme (with port matching). For higher utilization of the SSM switch, we propose the weighted central-module-link matching (WCMM) scheme. The WCMM scheme rescinds multiple iterations for module matching and yet, it achieves higher performance than the WFM-NP scheme. The advantages of the SSM switch are achieved without memory speedup. The memory addition is a small cost to trade for complexity reduction and performance improvement. Chuan-Bi Lin, Roberto Rojas-Cessa |
HPSR | 2 |
| 2013 | Measurement of packet processing time of an Internet host using asynchronous packet capture at the data-link layerabstractAs transmission speeds increase faster than processing speeds, the packet processing time (PPT) of a host (i.e., workstation) is becoming more significant in the measurement of different network parameters, in which packet processing by the host is involved. The PPT of a host is the time elapsed between the arrival of a packet at the data-link layer and the time the packet is processed at the application layer of the TCP/IP protocol stack (RFCs 2679 and 2681). In this paper, we propose a methodology to measure the PPT of a host using Internet Control Message Protocol (ICMP) packet and a specialized packet-capture card that does not require synchronization between the host under test and the packet-capture card. We tested the proposed methodology on two hosts with different specifications. The experimental results show that the proposed methodology consistently measures PPT. Khondaker Musfakus Salehin, Roberto Rojas-Cessa |
ICC | 2 |
| 2012 | Task and Server Assignment for Reduction of Energy Consumption in DatacentersabstractEnergy consumption of cloud data centers accounts for a major operational cost. This paper presents an optimization model for task scheduling to minimize task processing time and energy consumption in data centers for cloud computing. We formulate an integer programming optimization problem to minimize the expected energy consumption of homogenous tasks in a data center with a large number of servers and propose the most-efficient-server first greedy task scheduling algorithm to minimize energy expenditure. We show that the proposed task scheduling can minimize the energy expenditure while bounding the average task waiting time. We present a simulation of the proposed task scheduling scheme to show an optimum number of servers to achieve small task processing times and to minimize energy consumption. Ziqian Dong, Roberto Rojas-Cessa |
NCA | 3 |
| 2012 | Throughput analysis of shared-memory crosspoint buffered packet switchesabstractThis study presents a theoretical throughput analysis of two buffered-crossbar switches, called shared-memory crosspoint buffered (SMCB) switches, in which crosspoint buffers are shared by two or more inputs. In one of the switches, the shared-crosspoint buffers are dynamically partitioned and assigned to the sharing inputs, and memory is sped up. In the other switch, inputs are arbitrated to determine which of them accesses the shared-crosspoint buffers, and memory speedup is avoided. SMCB switches have been shown to achieve a throughput comparable to that of a combined input-crosspoint buffered (CICB) switch with dedicated crosspoint buffers to each input but, with less memory than a CICB switch. The two analysed SMCB switches use random selection as the arbitration scheme. The authors modelled the states of the shared-crosspoint buffers of the two switches using a Markov-modulated process and prove that the throughput of the proposed switches approaches 100% under independent and identically distributed uniform traffic. In addition, the authors provide numerical evaluations of the derived formulas to show how the throughput approaches asymptotically to 100%. Ziqian Dong, Roberto Rojas-Cessa |
IET Commun. | 2 |
| 2011 | Memory-memory-memory Clos-network packet switches with in-sequence serviceabstractOut-of-sequence is a problem faced by multi-stage buffered Clos-network switches. This paper proposes two buffered three-stage Clos-network packet switches that service packets in sequence and provide high switching performance. The proposed switches require short configuration times as compared to existing bufferless or partially buffered Clos-network switches. The proposed switches use time stamps assigned at the input modules to identify the order of packets in the switch. The switches use time-stamp monitoring mechanisms either at the input modules in a switch called the MMM-IM switch, or at the output modules in a switch called the MMM-OM switch to keep packets in sequence. Synchronization among different switch modules is not required in the proposed switches. The switching performance study presented in this paper shows that in-sequence monitoring at the IM provides higher performance and larger scalability than in-sequence monitoring at the output. Furthermore, the throughput of the MMM-IM switch is comparable to that of a switch that may service packets out of sequence. Ziqian Dong, Roberto Rojas-Cessa, Eiji Oki |
HPSR | 2 |
| 2011 | Avoiding Speedup from Bandwidth Overhead in a Practical Output-Queued Packet SwitchabstractWe study a time-slotted and synchronous output queued switch that uses packet concatenation to reduce band width overhead, or CS-OQ switch. The CS-OQ switch avoids memory speedup through the use of dedicated optical lines connecting inputs to outputs. The dedicated lines are combined with virtual input queues (VIQs) at the outputs and segmentation queues at the inputs. The segmentation queues are used for packet concatenation. We show that the performance of the CS-OQ switch approaches that of an ideal output queued switch, or asynchronous OQ (A-OQ) switch by selecting a suitable segmentation length in combination with packet concatenation. Lin Cai 0003, Roberto Rojas-Cessa, Taweesak Kijkanjanarat |
ICC | 2 |
| 2011 | Experimental performance evaluation of a virtual software routerabstractSoftware routers (SRs) are an alternative low-cost and moderate-performance router solutions implemented with general-purpose workstations able to host multiple network interface cards (NICs). Workstations can be programmed to forward packets between different NICs and to participate in routing functions. Virtualization can be used to model new protocols or hardware systems in software and without modifying the host's kernel. However virtualized routers are expected to suffer from performance degradation because of software execution overhead. In this paper, we investigate the performance impact of a virtual software router (VSR) in comparison to that of a SR. We present the performance of VSRs hosted by different workstations - with different number of processing cores. Roberto Rojas-Cessa, Khondaker Musfakus Salehin, Komlan Egoh |
LANMAN | 1 |
| 2011 | Load-Balanced Combined Input-Crosspoint Buffered Packet SwitchesabstractCombined input-crosspoint buffered (CICB) switches can achieve high switching performance without speedup. However, the dedicated crosspoint buffers in a CICB switch may not be efficiently used, and throughput degradation may occur. This throughput degradation is especially observable under flows with high data rates and long distances between the line cards and the buffered crossbar. This paper introduces two load-balanced CICB switches: the load-balancing CICB switch with full access (LB-CICB-FA) and the load-balancing CICB switch with single access (LB-CICB-SA). The proposed switches use the crosspoint buffers efficiently and support long distances between the line cards and buffered crossbar with crosspoint buffers smaller than those in a CICB switch by a factor of N, where N is the number of ports. It is proven that the LB-CICB-FA switch with random selection of the configuration of the load-balancing stage, input queues, and crosspoint queues is weakly stable under admissible independent and identical distributed (i.i.d.) traffic. Additional simulation results support the correctness of the theoretical analysis. Furthermore, it is shown that the throughput of the LB-CICB-SA switch with the longest-queue first (LQF) and first-come first-served (FCFS) as input and output arbitrations, respectively, is 100% under admissible i.i.d. traffic. The proposed switches keep cells in sequence and use no speedup. The low implementation complexity of the load-balancing stage is discussed and shown to be small. Roberto Rojas-Cessa, Ziqian Dong |
IEEE Trans. Commun. | 1 |
| 2010 | Ternary-Search-Based Scheme to Measure Link Available-Bandwidth in Wired NetworksabstractAccurate measurement of available bandwidth (ABW) is an important parameter to analyze network performance. Active measurement is an attractive approach as it has the advantage of controllability and flexibility for performing network measurement. However, it can affect both the data traffic and the measurement process itself if a significant amount of probe traffic is injected into the network. Furthermore, measurement must be completed in short time to effectively monitor the network state. In this paper, we propose a fast ABW measurement scheme that generates a small amount of probe traffic to achieve an acceptable measurement accuracy. The proposed scheme achieves an accuracy comparable to that of popular existing schemes. We present a performance study of the proposed scheme through ns2 simulation under different traffic conditions. Khondaker Musfakus Salehin, Roberto Rojas-Cessa |
GLOBECOM | 2 |
| 2010 | Scheme to measure One-Way Delay Variation with detection and removal of clock skewabstractOne-Way Delay Variation (OWDV) has become increasingly of interest to evaluate network state and service quality, especially for real-time and streaming services such as VoIP and video. Measurement of these parameters needs to be performed with the layout infrastructure. Many schemes for OWD measurements require clock synchronization at the source and destination through Global-Positioning System (GPS) or the Network Time Protocol (NTP). In clock-synchronized approaches, the accuracy of the measurement of OWDV depends on the achieved accuracy of clock synchronization. GPS provides high-accuracy clock synchronization. However, the deployment of GPS on legacy network equipment might be slow and costly. This paper proposes a method for measuring OWDV without recurring to clock synchronization. However, clock skew may affect the measurement of OWDV. The proposed approach is based on the measurement of Inter-Packet Delay (IPD) and Accumulated OWDV (AOWDV). This paper shows the performance of the proposed scheme via simulation and through experimentation in a VoIP network. The presented simulation and experimental results indicate that clock skew can be efficiently measured and removed and that OWDV can be measured without requiring clock synchronization. Makoto Aoki, Eiji Oki, Roberto Rojas-Cessa |
HPSR | 3 |
| 2010 | Three-Dimensional Based Trust Management Scheme for Virus Control in P2P NetworksabstractPeer-to-peer (P2P) networking is widely used to exchange, contribute, or obtain files from any participating user. In these networks, worms, viruses and intruding files find an open door to the downloading host, creating a convenient environment for successful proliferation throughout the network. Trust management is a promising proactive mechanism to prevent virus dissemination. Current trust models use peer reputation for this purpose. However, when viruses have infectious properties, peer reputation may not be enough to limit their proliferation. In this paper, we show that peer reputation alone cannot bound epidemics in an infectious environment. Therefore, this paper introduces a trust management scheme that uses the combination of trust values of peers and infection values of both peers and content. Moreover, to improve the efficiency on the calculation of trust values of ratio-based normalization models, we propose a model for trust value calculation using a three dimensional (3D) normalization to represent peer activity with high accuracy. We show that the proposed trust management scheme can bound virus proliferation to a small number of peers, without inhibiting file-downloading activity. Lin Cai 0003, Roberto Rojas-Cessa |
ICC | 2 |
| 2010 | Distributed Diffusion-Based Mesh Algorithm for Distributed Mesh Construction in Wireless Ad Hoc and Sensor NetworksabstractReliable mesh communications in dense wireless ad hoc networks require the creation of both self organizing mesh structures and mesh routing protocols to accomplish efficient and reliable communications with the added infrastructure redundancy. To date, much of the research in the area has focused on communication protocol design. The investigations often are based on a mesh network structure already fully formed and some times fixed to the underlying physical node topology. Therefore, there is a need for a platform to build mesh networks with structural flexibility and to provide management functions to network- and application-level protocols. In this paper, we propose the distributed diffusion-based mesh (DDM) algorithm for distributed mesh construction that instructs distributed nodes on how to make the desired connections with their neighbors. We accomplish this by introducing the concept of connection rule, which defines allowed connections at each mesh node, combined with a token signal that initiates and controls the structure and boundaries of the resulting mesh. We argue that slight changes in mesh network structure greatly affect network performance and show how the combined use of rule and token signal offers control over the resulting mesh structure. This methodology can be used for cross-layer optimization to achieve a network topology suitable for different network applications. As compared with existing protocols, our algorithm also provides a large reduction in communication overhead. Komlan Egoh, Roberto Rojas-Cessa, Nirwan Ansari |
ICC | 2 |
| 2010 | Performance of an Optical Packet Switch with Parametric Wavelength ConvertersabstractIn an optical packet switch (OPS), input fibers carry multiple wavelengths, which carry packets to one or more output fibers. As several wavelengths from different inputs could be destined to the same output fiber, one wavelength can be connected and the others remain disconnected, losing the carried packets. Because of the multiple wavelengths available at an output fiber, wavelength conversion in the OPS of the unconnected wavelengths into those available can increase the number of connections. A parametric wavelength converter (PWC) provides multi-channel wavelength conversion where wavelengths can be converted to another. A PWC uses a pump wavelength that can be flexibly chosen to define which wavelengths can be converted, defining the so-called wavelength conversion pairs. However, it is unknown which set of pump wavelengths, and therefore the set of connection pairs, should be selected to improve the OPS performance while minimizing the number of PWCs in the OPS. Therefore, this paper proposes a pump wavelength selection policy for an OPS that uses different pump wavelengths, one for each PWC, within an arbitrarily selected interval. This policy is called variety rich (VR) policy. This paper also introduces a non-wavelength blocking OPS (NWB-OPS) to make full use of PWCs. The switch performance is evaluated through computer simulation. The results show that the proposed policy with different pump wavelengths achieves the highest performance when compared to another of similar complexity. Furthermore, the performance study shows that small sizes of the interval to select a pump wavelength are more beneficial than larger ones. Nattapong Kitsuwan, Roberto Rojas-Cessa, Motoharu Matsuura, Eiji Oki |
ICC | 2 |
| 2010 | Maximal weight matching scheme with frame occupancy-based for input-queued packet switchesabstractVirtual output queues (VOQs) is widely used by input-queued (IQ) switches to eliminate the head-of-line (HOL) blocking phenomenon that limits matching and switching performance. It has been shown that IQ switches can provide 100% throughput under admissible traffic when using maximum-weight matching schemes or iterative maximal-weight matching schemes with a speedup of two or more. These different approaches require either a high computation complexity or large resolution times for high-speed switches. Therefore, there is a need for low-complexity and fast matching schemes that provide high throughput under several admissible traffic patterns, including those with nonuniform distributions, without recurring to speedup nor multiple iterations. We proposed earlier weightless matching schemes based on the captured-frame concept for IQ switches to provide high throughput under uniform and nonuniform traffic patterns, when using a single iteration and no speedup. In this paper, we proposed to use the capture-frame concept in weighted matching schemes and show that performance improvement can be achieved. We apply this concept to the longest-queue first (LQF) selection to create the unlimited LQF (uFLQF) schemes. We show via simulation that uFLQF achieves high throughput under uniform and nonuniform traffic patterns with a single iteration. Chuan-Bi Lin, Roberto Rojas-Cessa |
LANMAN | 2 |
| 2010 | Schemes to measure available bandwidth and link capacity with ternary search and compound probe for packet networksabstractAccurate measurement of network parameters such as available bandwidth (ABW) and link capacity are needed for analyzing network performance. Active measurement is an attractive approach as it has the advantage of controllability and flexibility for performing network measurement. However, it can affect both the data traffic and the measurement process itself, affecting the accuracy of the measurement if significant amount of probe traffic is injected into the network. Furthermore, measurement must be completed in short time to effectively monitor the network state. In this paper, we present two new measurement schemes: one for measuring ABW, and the other for measuring per-hop link capacities of an end-to-end path. The ABW measurement scheme performs measurement in short period of time and with small amount of probe traffic, and it achieves accuracy comparable to that of IGI and Pathload. The link-capacity measurement scheme provides immunity to cross traffic. We present ns-2 simulation results of the ABW and link-capacity measurement schemes to show their performance. Khondaker Musfakus Salehin, Roberto Rojas-Cessa |
LANMAN | 2 |
| 2010 | Task-execution scheduling schemes for network measurement and monitoring
Roberto Rojas-Cessa, Nirwan Ansari |
Comput. Commun. | 2 |
| 2010 | Combined methodology for measurement of available bandwidth and link capacity in wired packet networksabstractAccurate measurement of network parameters such as available bandwidth (ABW), link capacity, delay, packet loss and jitter are used to support and monitor several network functions, for example traffic engineering, quality-of-service (QoS) routing, end-to-end transport performance optimisation and link capacity planning. However, proactive network measurement schemes can impact both the data traffic and the measurement process itself, affecting the accuracy of the estimation if a significant amount of probe traffic is injected into the network. In this work, the authors propose two measurement schemes, one for measuring ABW and the other for measuring link capacity, both of them use a combination of data probe packets and Internet control messaging protocol (ICMP) packets. Our schemes perform ABW and link-capacity measurements in a short time and with a small amount of probe traffic. The authors show a performance study of our measurement schemes and compare their accuracy to those of other existing measurement schemes and also show that the proposed schemes achieve shorter convergence time than other existing schemes and high accuracy. Khondaker Musfakus Salehin, Roberto Rojas-Cessa |
IET Commun. | 2 |
| 2009 | Re-Configurable Parallel Match Evaluators Applied to Scheduling Schemes for Input-Queued Packet SwitchesabstractThe performance of matching schemes for input- queued (IQ) packet switches is mainly defined by the selection policy adopted. This policy can be aimed to produce a large weight sum for matched input-output pairs, where each input- output pair is assigned a weight, or to produce a large match size in the number of matched pairs, giving place to maximum weight matching or maximum size matching, respectively. However, schedulers can only provide a single match in function of the selection (of candidate ports) policy adopted and of the backlogged traffic at the input queues. A parallel match evaluator was recently proposed to provide not one but several match options at the same time. This approach evaluates several predefined and fixed matches and picks the match with the largest size. However, the fixed permutations of the evaluated matches may produce low performance under traffic with nonuniform distributions because of the limited number of choices. This paper proposes to make the parallel match evaluator configurable and two schemes to provide diverse and changeable matches such that the matches (and therefore, the evaluator) become adaptable to the traffic pattern. The proposed schemes were tested under uniform and nonuniform traffic patterns and the results show that these schemes provide high performance, even when scheduling is performed between periods of multiple time slots, or framed intervals. The proposed approach can be used for configuring slow micro-electro-mechanical (MEM) optical switch fabrics. Spiridon F. Beldianu, Roberto Rojas-Cessa, Eiji Oki, Sotirios G. Ziavras |
ICCCN | 2 |
| 2009 | Analysis of Space-Space-Space Clos-Network Packet SwitchabstractThe throughput of a packet switch is a major switch property, and therefore, of major interest to analyze it. An approximation of the throughput of a staged random selection algorithm with a single iteration under uniform for a three-stage Clos-network packet switch, also called a Space- Space-Space (S3) Clos-network packet switch, has been recently presented. However, the difference between this approximation and the actual throughput of the staged random selection algorithm is significant. To address this issue, this paper presents a theoretical throughput analysis of the staged random selection algorithm with a single iteration for a S3Clos-network switch and show that the throughput is higher than that estimated by the existing approximation. Second, the paper extends the analysis to calculate the throughput of the staged random selection algorithm with multiple iterations by considering the analysis of the parallel iterative matching scheme, which is a random-based matching scheme for single-stage switches. The introduced derivation carefully considers the behavior of the selection algorithm at the switching modules in all three stages of the switch. The probability that a request reaches the third-stage modules is affected by the matching results at the second-stage modules. Numerical evaluations of the analytical formulas are performed. The results show that the staged random selection algorithm with multiple iterations for a S3Clos-network switch without internal expansion can achieve 100% throughput under uniform traffic. Eiji Oki, Nattapong Kitsuwan, Roberto Rojas-Cessa |
ICCCN | 3 |
| 2009 | Proactive Routing for Congestion Avoidance in Network Recovery under Single-Link FailuresabstractRouting schemes for network recovery aim to find connection paths under the scenarios of link or node failures that disrupt end-to-end connectivity. Several recovery schemes consider a proactive search for alternative links (or paths) that can restore the lost connectivity after the failure. However, the effect that the modification of network traffic after the switch over the alternative paths brings can undermine the objective of these schemes as links can become overloaded. This effect can cause temporary instability, requiring a new selection of paths and interrupt communications. This new search would defeat the purpose of pro-actively disposing of alternative paths. To address this issue, this paper proposes an on-demand and proactive routing approach, where the demand is based on providing new paths for existing network traffic while avoiding link congestion. Because some links may get congested in an attempt to accommodate re-routed flows in a recovery tree, on-demand routing is based on the selection of paths according to the existing traffic load in the failed link and the available bandwidth in the alternative links. This approach produces one or multiple-paths from source nodes to destination nodes to satisfy network load requirements and to avoid interruptions produce by route oscillations. The proposed scheme can create as many trees as needed to satisfy traffic flowing. Roberto Rojas-Cessa |
ICTAI | 1 |
| 2008 | Input- and Output-Based Shared-Memory Crosspoint-Buffered Packet Switches for Multicast Traffic Switching and ReplicationabstractThe incorporation of broadcast and multimedia- on-demand services are expected to increase multicast traffic in packet networks, and therefore in switches and routers. Combined input-crosspoint buffered (CICB) switches can provide high performance under uniform multicast traffic, however, at the expense of N2crosspoint buffers. In this paper, we introduce an output-based shared-memory crosspoint-buffered (O-SMCB) packet switch where the crosspoint buffers are shared by two outputs and use no speedup. The proposed switch provides high performance under admissible uniform and nonuniform multicast traffic models while using 50% of the memory used in CICB switches. Furthermore, the O-SMCB switch provides higher throughput than an SMCB switch with buffers shared by inputs, or I-SMCB, previously proposed, despite the strong similarities between the architectures of these two switches. In this paper, we study the performance of the O-SMCB switch under uniform and nonuniform multicast traffic models and compare it to the I-SMCB switch. Ziqian Dong, Roberto Rojas-Cessa |
ICC | 2 |
| 2008 | Module-First Matching Schemes for Scalable Input-Queued Space-Space-Space Clos-Network Packet SwitchesabstractClos-network switches were proposed as a scalable architecture for the implementation of large-capacity circuit switches. In packet switching, the three-stage Clos-network architecture uses small switches as modules to assemble a switch with large number of ports or aggregated ports with high data rates. Current schemes for configuration of input-queued three- stage Clos-network (IQC) switches involve port matching and path routing assignment, in that order. The implementation of a scheduler capable of matching thousands of ports in large-size switches is complex because of the large port count. To decrease the scheduler complexity for such switches (e.g., 1024 ports or more), we propose a configuration scheme for IQC switches that hierarchizes the matching process. In a practical scenario our scheme performs routing first and port matching thereafter. This approach applies the reduction concept of Clos networks to the matching process. The application of this approach results in a feasible size of schedulers for up to Exabit-capacity switches, an independent configuration of the middle stage modules from port matches, a reduction of the matching communication overhead between different stages, and a release of the switching function to the last-stage modules in a 3-stage switch. We show that the switching performance of the proposed approach using weight- based and weightless selection schemes is high under uniform and nonuniform traffic. Chuan-Bi Lin, Roberto Rojas-Cessa |
ICC | 2 |
| 2007 | Parallel Search Trie-Based Scheme for Fast IP LookupabstractAs data rates in the Internet increase, the Internet Protocol (IP) address lookup is required to be resolved in shorter resolution times. IP address lookup involves finding the longest matching prefix from a database of prefixes that better matches the destination address of a packet. The fastest IP-address lookup solutions are based on ternary content addressable memories (TCAMs), which can resolve the IP lookup in one memory-access time. However, TCAMs have a high power consumption and large complexity that may limit their scalability and storage capacity. An alternative is to use random access memory (RAM) that stores a forwarding table in a trie form. Proposed trie-based solutions for IP lookup require three or more memory-access times in the worst-case scenario. This makes them unattractive despite their reduced power consumption. In this paper, we propose a flexible and fast trie-based IP-lookup algorithm where parallel searching is performed. This algorithm performs lookup in two memory- access times whith a feasible amount of memory or three memory access times with reduced memory. Roberto Rojas-Cessa, Lakshmi Ramesh, Ziqian Dong, Lin Cai 0003, Nirwan Ansari |
GLOBECOM | 1 |
| 2007 | Captured-frame matching schemes for scalable input-queued packet switches
Roberto Rojas-Cessa, Chuan-Bi Lin |
Comput. Commun. | 1 |
| 2006 | Shared-Memory Combined Input-Crosspoint Buffered Packet Switch for Differentiated ServicesabstractCombined input-crosspoint buffered (CICB) packet switches with dedicated crosspoint buffers require a minimum amount of memory in the buffered crossbar of N2ldr k ldr L bytes, where N is the number of ports and k is the crosspoint buffer size, which is defined by the distance between the line cards and the buffered crossbar, and L is the cell (packet) size in bytes, to avoid buffer underflow under high-speed data flows. To support P traffic classes with different priorities, CICB switches requires N2ldrkldrLldrP bytes to avoid blocking of high priority cells. In this paper, we study a shared-memory crosspoint buffered packet switch that uses small crosspoint buffers and no speedup to support differentiated services and long distances between the line cards and the buffered crossbar in practical implementations. The proposed switch requires 1/m of memory amount in a CICB switch to achieve similar throughput performance. Ziqian Dong, Roberto Rojas-Cessa |
GLOBECOM | 2 |
| 2006 | Improving the Accuracy of EEAC-SV with Smart Packet MarkingabstractExplicit endpoint admission control with service vector (EEAC-SV) is a solution that allows a multi-hop connection in a Diffserv network to utilize different service classes in different routers along the path, thus enabling the network to provide QoS with fine granularity and improved network utilization. EEAC-SV relies on the QoS information in the pre-marked probing packet to select the optimal service vector. However, owing to the overhead constraint, the QoS information in the probing packet can only be conveyed in a quantized manner. The mathematical model indicates that with the same amount of overhead, the performance of EEAC-SV can be optimized with the proper packet marking strategy. In this paper, we propose the smart packet marking strategy to improve the EEAC-SV performance in terms of reducing the false routing probability. Smart packet marking is a stochastic solution which, based on the knowledge of the history of network QoS performances and user behaviors, defines the boundaries of each quantization level in such a way to improve the routing accuracy. Through extensive simulations, we demonstrate that smart packet marking outperforms the packet marking strategy that does not take stochastic information into consideration. Nirwan Ansari, Roberto Rojas-Cessa |
GLOBECOM | 3 |
| 2006 | Framed Round-Robin Arbitration with Explicit Feedback Control for Combined Input-Crosspoint Buffered Packet SwitchesabstractThis paper introduces a frame-based round-robin arbitration scheme with explicit feedback control (FRE) for combined input-crosspoint buffered packet switches. We consider a CICB switch with fixed-length packets, called cells. The proposed scheme dynamically sets the frame size according to the amount of cell accumulation at the input queues. We study FRE when applied to a continuous system. We transport the FRE concept into a discrete system as an arbitration scheme for a packet switch, and study the switching performance. We combine FRE, which is used as the input arbitration scheme, with other round-robin based schemes, used as output arbitration schemes. The resulting combined schemes provide high throughput under several admissible traffic patterns when used in a switch with one-cell crosspoint buffers and no speedup. Roberto Rojas-Cessa |
ICC | 2 |
| 2005 | Load-balanced CICB packet switch with support for long round-trip timesabstractCombined input-crosspoint buffered (CICB) packet switches relax arbitration timing and provide high-performance switching. However, the amount of memory in buffered crossbars required to achieve 100% throughput under flows with high data rates is proportional to the number of ports, N, and the crosspoint buffer size k, which is defined by the distance between the line cards and the buffered crossbar. Long distances between the line cards and the buffered crossbar can make a CICB switch costly to implement. In this paper, we propose a load-balanced CICB packet switch to support long distances between the buffered crossbar and the line cards using crosspoint buffers of small size. The proposed switch reduces the required crosspoint buffer size by a factor of N and keeps the cells in sequence. Roberto Rojas-Cessa, Ziqian Dong, Sotirios G. Ziavras |
GLOBECOM | 1 |
| 2005 | Matching schemes with captured-frame eligibility for input-queued packet switchesabstractVirtual output queues (VOQs) are widely used by input-queued (IQ) switches to eliminate the head-of-line (HOL) blocking phenomena, which limits switching performance. An effective matching scheme must provide high throughput under several admissible traffic patterns and keep implementation complexity low. A variety of matching schemes for IQ switches that deliver high throughput under uniform traffic have been proposed. However, there is a need of matching schemes that provide high throughput under several admissible traffic patterns, including those with nonuniform distributions. In this paper, we introduce the captured frame-size concept for matching schemes in IQ switches. We use the captured-frame eligibility concept in a round-robin based scheme, uFORM, and in a random-base scheme, uFPIM, to improve switching performance under nonuniform traffic patterns. The uFPIM scheme is based in the parallel iterative matching (PIM) scheme and shows the throughput improvement achieved with the captured frame concept. The uFORM scheme provides high performance under nonuniform traffic while keeping the high performance that round-robin schemes are known to have under uniform traffic. Roberto Rojas-Cessa, Chuan-Bi Lin |
ICC | 1 |
| 2005 | On the combined input-crosspoint buffered switch with round-robin arbitrationabstractInput-buffered switches have been widely considered for implementing feasible packet switches. However, their matching process may not be time-efficient for switches with high-speed ports. Buffered crossbars (BXs) are an alternative to relax timing for packet switches with high-speed ports and to provide high-performance switching. BX switches were originally considered expensive, as the memory amount required in the crosspoints (XPs) is proportional to the square of the number of ports (O(N/sup 2/)). This limitation is now less stringent with the advances on chip-fabrication techniques, and when considering small crosspoint (XP) buffer sizes. In this paper, we study a combined input-crosspoint buffered packet switch, named CIXB, with virtual output queues (VOQs) at the inputs, and arbitration based on round-robin selection. We show that the CIXB switch achieves 100% throughput under uniform traffic, and high performance under nonuniform traffic, using one-cell XP buffer size and no speedup. Roberto Rojas-Cessa, Eiji Oki, H. Jonathan Chao |
IEEE Trans. Commun. | 1 |
| 2004 | Frame occupancy-based round-robin matching scheme for input-queued packet switchesabstractThe use of virtual output queues (VOQ) in input-queued (IQ) switches can eliminate the head-of-line (HOL) blocking phenomenon, which limits switching performance. An effective matching scheme for IQ switches with VOQ must provide high throughput under admissible traffic patterns while keeping the implementation feasible. This paper proposes a matching scheme for IQ switches that provides high throughput under uniform and a nonuniform traffic pattern, called unbalanced. The proposed matching scheme, FORM, is primarily based on round-robin selection and the captured-frame concept. We show via simulation that this scheme delivers over 99% throughput under unbalanced traffic and retains the high performance under uniform traffic that round-robin matching schemes are known to offer. Roberto Rojas-Cessa, Chuan-Bi Lin |
GLOBECOM | 1 |
| 2004 | Round-robin with adaptable-size-frame arbitration for input-crosspoint buffered switchesabstractCombined input-crosspoint buffered switches relax arbitration timing and provide high-performance switching for packet switches with high-speed ports. It has been shown that these switches, with one-cell crosspoint buffer and round-robin arbitration at input and output ports, provide 100% throughput under uniform traffic. However, under admissible traffic patterns with nonuniform distributions, only weight-based selection schemes are reported to provide high throughput. This paper proposes a round-robin based arbitration scheme for a combined input-crosspoint buffered packet switch. The presented scheme uses adaptable-size frames, where the frame size is determined by the received service. The resulting switch provides nearly 100% throughput for several admissible traffic patterns, including uniform and unbalanced traffic, using one-cell crosspoint buffers. Roberto Rojas-Cessa |
ICC | 1 |
| 2004 | Maximum weight matching dispatching scheme in buffered Clos-network packet switchesabstractThe scalability of Clos-network switches makes them an alternative to single-stages switches for implementing large-size packet switches. This paper introduces a cell dispatching scheme, called Maximum Weight Matching Dispatching (MWMD) scheme, for buffered Clos-network switches. The MWMD scheme is based on a maximum weight matching algorithm for input-buffered switches. This paper shows that, with request queues in the buffered Clos-network architecture, the MWMD scheme is able to achieve a 100% throughput for independent admissible traffic, without allocating any buffers in the second stage and without expanding the internal bandwidth. As a practical scheme, a maximal oldest-cell-first matching dispatching (MOMD) scheme is also introduced. MOMD shows that using a finite number of iterations in the dispatching scheme, the throughout under unbalanced traffic pattern can be high. Roberto Rojas-Cessa, Eiji Oki, H. Jonathan Chao |
ICC | 1 |
| 2003 | Concurrent fault detection for a multiple-plane packet switchabstractIn high-speed and high-capacity packet switches, system reliability is critical to avoid loss of huge amounts of information and retransmission of traffic. We propose a series of concurrent fault-detection mechanisms for a multiple-plane crossbar-based packet switch. Our switch model, called the m+z model, has m active planes and z spare planes. This switch has distributed arbiters on each plane. The spare planes, used for substitution of faulty active ones, are also used in the fault-detection mechanism, thus providing fault detection and fault location for all switching planes. Our detection schemes are able to detect a single fault quickly without increasing transmission overhead. The proposed schemes can be used for switches with different numbers of active planes and a small number of spare planes. Roberto Rojas-Cessa, Eiji Oki, H. Jonathan Chao |
IEEE/ACM Trans. Netw. | 1 |
| 2002 | PCRRD: a pipeline-based concurrent round-robin dispatching scheme for Clos-network switchesabstractThis paper proposes a pipeline-based concurrent round-robin dispatching scheme, called PCRRD, for Clos-network switches. Our previously proposed concurrent round-robin dispatching (CRRD) scheme provides 100% throughput under uniform traffic by using simple round-robin arbiters, but it has the strict timing constraint that the dispatching scheduling has to be completed within one cell time slot. This is a bottleneck in building high-performance switching systems. To relax the strict timing constraint of CRRD, we propose to use more than one scheduler engine, up to P, so called subschedulers. Each subscheduler is allowed to take more than one time slot for dispatching. Every time slot, one out of P subschedulers provides the dispatching result. The subschedulers adopt our original CRRD algorithm. We show that PCRRD preserves 100% throughput under uniform traffic of our original CRRD algorithm, while ensuring the cell-sequence order. Since the constraint of the scheduling timing is dramatically relaxed, it is suitable for high-performance switching systems even when the switch size increases and port speed is high (e.g., 40 Gbit/s). Eiji Oki, Roberto Rojas-Cessa, H. Jonathan Chao |
ICC | 2 |
| 2002 | Concurrent round-robin-based dispatching schemes for Clos-network switchesabstractA Clos-network switch architecture is attractive because of its scalability. Previously proposed implementable dispatching schemes from the first stage to the second stage, such as random dispatching (RD), are not able to achieve high throughput unless the internal bandwidth is expanded. This paper presents two round-robin-based dispatching schemes to overcome the throughput limitation of the RD scheme. First, we introduce a concurrent round-robin dispatching (CRRD) scheme for the Clos-network switch. The CRRD scheme provides high switch throughput without expanding internal bandwidth. CRRD implementation is very simple because only simple round-robin arbiters are adopted. We show via simulation that CRRD achieves 100% throughput under uniform traffic. When the offered load reaches 1.0, the pointers of round-robin arbiters at the first- and second-stage modules are completely desynchronized and contention is avoided. Second, we introduce a concurrent master-slave round-robin dispatching (CMSD) scheme as an improved version of CRRD to make it more scalable. CMSD uses hierarchical round-robin arbitration. We show that CMSD preserves the advantages of CRRD, reduces the scheduling time by 30% or more when arbitration time is significant and has a dramatically reduced number of crosspoints of the interconnection wires between round-robin arbiters in the dispatching scheduler with a ratio of 1//spl radic/N, where N is the switch size. This makes CMSD easier to implement than CRRD when the switch size becomes large. Eiji Oki, Zhigang Jing, Roberto Rojas-Cessa, H. Jonathan Chao |
IEEE/ACM Trans. Netw. | 3 |
| 2001 | PMM: a pipelined maximal-sized matching scheduling approach for input-buffered switchesabstractThis paper proposes an innovative pipeline-based maximal-sized matching scheduling approach, called PMM, for input-buffered switches. It dramatically relaxes the timing constraint for arbitration with a maximal matching scheme. In the PMM approach, arbitration operates in a pipelined manner, where K subschedulers are used. Each subscheduler is allowed to take more than one time slot for its matching. Every time slot, one of them provides the matching result. The subscheduler can adopt a pre-existing efficient maximal matching algorithm such as iSLIP and DRRM. PMM maximizes the efficiency of the adopted arbitration scheme by allowing sufficient time for a number of iterations. We show that PMM preserves 100% throughput under uniform traffic and fairness for best-effort traffic of the pre-existing algorithm. Eiji Oki, Roberto Rojas-Cessa, H. Jonathan Chao |
GLOBECOM | 2 |
| 2001 | Fast fault detection for a multiple-plane packet switchabstractIn high-speed and high-capacity packet switches, system reliability is critical to avoid the loss of a huge amount of information and to avoid re-transmission of traffic. We propose a series of concurrent fault-detection mechanisms for a multiple-plane crossbar-based packet switch. Our switch model, called the m + z model, has m active planes and z spare planes. This switch has distributed arbiters on each plane. The spare planes, used for substitution of faulty active ones, are also used in the fault detection mechanism, thus providing sufficient data redundancy for fault detection and location. Our detection scheme is able to detect a single fault in one time slot without increasing transmission overhead. The proposed schemes can be used for switches with different numbers of active planes and the number of spare planes needed for fault detection is small. Roberto Rojas-Cessa, Eiji Oki, H. Jonathan Chao |
GLOBECOM | 1 |
| 2001 | CIXOB-k: combined input-crosspoint-output buffered packet switchabstractWe propose a novel architecture, a combined input-crosspoint-output buffered (CIXOB-k, where k is the size of the crosspoint buffer) Switch. CIXOB-k architecture provides 100% throughput under uniform and unbalanced traffic. It also provides timing relaxation and scalability. CIXOB-k is based on a switch with combined input-crosspoint buffering (CIXB-k) and round-robin arbitration. CIXB-k has a better performance than a non-buffered crossbar that uses iSLIP arbitration scheme. CIXOB-k uses a small speedup to provide 100% throughput under unbalanced traffic. We analyze the effect of the crosspoint buffer size and the switch size under uniform and unbalanced traffic for CIXB-k. We also describe solutions for relaxing the crosspoint memory amount and scalability for a CIXOB-k switch with a large number of ports. Roberto Rojas-Cessa, Eiji Oki, H. Jonathan Chao |
GLOBECOM | 1 |
| 2001 | Concurrent round-robin dispatching scheme in a clos-network switchabstractA Clos-network switch architecture is attractive because of its scalability. Previously proposed implementable dispatching schemes from the first stage to the second stage, such as random dispatching, are not able to achieve a high throughput unless the internal bandwidth is expanded. This paper proposes a concurrent round-robin dispatching (CRRD) scheme for a Clos-network switch, to overcome the throughput limitation of the random dispatching scheme. The CRRD scheme provides high switch throughput without expanding internal bandwidth. CRRD implementation is very simple because only simple roundrobin arbiters are adopted. In CRRD, the round-robin arbiters concurrently perform the matching between requesting cells and output links in each first-stage module to dispatch the cells to available second-stage modules. We show that CRRD achieves 100% throughput under uniform traffic. When the offered load reaches 1.0, the pointers of roundrobin arbiters at the first-stage and second-stage modules are effectively desynchronized and contention is avoided. key words: Packet switch, Clos-network switch, dispatching, arbitration, throughput Eiji Oki, Zhigang Jing, Roberto Rojas-Cessa, H. Jonathan Chao |
ICC | 3 |