Mikel Jimenez Fernandez

dblp:326/4377 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 4 · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer networks
4 papers
Network performance modeling · 40% Datacenter networks · 15% Internet architecture and protocols · 13%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Distributed systems · 31% High-performance computing · 31% Interconnection networks and networks-on-chip · 31%

Topics — the 10 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Network performance modeling › network simulation
large-scale network simulation
1.012026
Enabling AI Network Cross-Layer Design and Operations with Arcadia: A Simulation Platform at Scale · NSDI 2026
Network performance modeling
network simulation
1.012026
Enabling AI Network Cross-Layer Design and Operations with Arcadia: A Simulation Platform at Scale · NSDI 2026
Datacenter networks › RDMA
RDMA over Converged Ethernet
0.812024
RDMA over Ethernet for Distributed Training at Meta Scale · SIGCOMM 2024
High-performance computing
collective communication
0.812024
RDMA over Ethernet for Distributed Training at Meta Scale · SIGCOMM 2024
Distributed systems › distributed machine learning
distributed training
0.812024
RDMA over Ethernet for Distributed Training at Meta Scale · SIGCOMM 2024
Interconnection networks and networks-on-chip › remote direct memory access
RDMA networks
0.812024
RDMA over Ethernet for Distributed Training at Meta Scale · SIGCOMM 2024
Internet architecture and protocols
wide area network
0.712023
EBB: Reliable and Evolvable Express Backbone Network in Meta · SIGCOMM 2023
Network optimization and economics › resource sharing
network sharing
0.612022
Network entitlement: contract-based network sharing with agility and SLO guarantees · SIGCOMM 2022
Routing and switching › traffic engineering
MPLS traffic engineering
0.212023
EBB: Reliable and Evolvable Express Backbone Network in Meta · SIGCOMM 2023
Routing and switching
traffic engineering
0.212023
EBB: Reliable and Evolvable Express Backbone Network in Meta · SIGCOMM 2023

Methods — techniques the papers use, named apart from their topics

simulation · 2.1distributed enforcement · 1.1distributed control agents · 0.7centralized traffic engineering · 0.7
YearPublicationVenuePosition
2026 Enabling AI Network Cross-Layer Design and Operations with Arcadia: A Simulation Platform at Scale
Zhaodong Wang, Satyajeet Ahuja, Mohammad Noormohammadpour, Gregory R. Steinbrecher, Thomas Fuller, Kevin Quirk, Mikel Jimenez Fernandez, Abhinav Triguna, Yan Cai 0018, Steve Politis, Petr Lapukhov, Naader Hasani, Ying Zhang 0022
NSDI9
2024 RDMA over Ethernet for Distributed Training at Meta Scale
abstract
The rapid growth in both computational density and scale in AI models in recent years motivates the construction of an efficient and reliable dedicated network infrastructure. This paper presents the design, implementation, and operation of Meta's Remote Direct Memory Access over Converged Ethernet (RoCE) networks for distributed AI training.
Adithya Gangidi, Rui Miao 0001, Shengbao Zheng, Sai Jayesh Bondu, Guilherme Loch Waltrick Goes, Hany Morsy, Rohit Puri, Mohammad Riftadi, Ashmitha Jeevaraj Shetty, Shuqiang Zhang, Mikel Jimenez Fernandez, Shashidhar Gandham, Hongyi Zeng
SIGCOMM12
2023 EBB: Reliable and Evolvable Express Backbone Network in Meta
abstract
We present the design, implementation, evaluation, deployment and production experiences of EBB (Express BackBone), a private WAN (Wide Area Network) connecting Meta's global data centers (DCs). Initiated in 2015, EBB now carries 100% of DC-DC traffic, witnessing remarkable growth over the years. A key design aspect of EBB is its multi-plane architecture, facilitating seamless deployment of a new control plane while ensuring operational simplicity. This architecture allows for efficient failure mitigation, standard maintenance, and capacity expansion by draining one or two planes without impacting service level objectives (SLOs). Another critical design decision is the hybrid model, combining distributed control agents and a central controller. EBB's centralized traffic engineering utilizes an MPLS-TE based solution to allocate paths periodically for different traffic classes based on service requirements, while its distributed control agents enable fast local failure recovery by pre-installing pre-computed backup paths in the data plane. We delve into our eight-year production experience, highlighting the successful deployment of multiple generations of EBB.
Marek Denis, Yuanjun Yao, Ashley Hatch, Chiunlin Lim, Shuqiang Zhang, Kyle Sugrue, Henry Kwok, Mikel Jimenez Fernandez, Petr Lapukhov, Sandeep Hebbani, Gaya Nagarajan, Omar Baldonado, Lixin Gao 0001, Ying Zhang 0022
SIGCOMM9
2022 Network entitlement: contract-based network sharing with agility and SLO guarantees
abstract
This paper presents Meta's Production Wide Area Network (WAN) Entitlement solution used by thousands of Meta's services to share the network safely and efficiently. We first introduce the Network Entitlement problem, i.e., how to share WAN bandwidth across services with flexibility and SLO guarantees. We present a new abstraction entitlement contract, which is stable, simple, and operationally friendly. The contract defines services' network quota and is set up between the network team and services teams to govern their obligations. Our framework includes two key parts: (1) an entitlement granting system that establishes an agile contract while achieving network efficiency and meeting long-term SLO guarantees, and (2) a large-scale distributed run-time enforcement system that enforces the contract on the production traffic. We demonstrate its effectiveness through extensive simulations and real-world end-to-end tests. The system has been deployed and operated for over two years in production. We hope that our years of experience provide a new angle to viewing WAN network sharing in production and will inspire follow-up research.
Satyajeet Ahuja, Vinayak Dangui, Kirtesh Patil, Manikandan Somasundaram, Mario A. Sánchez, Guanqing Yan, Mohammad Noormohammadpour, Alaleh Razmjoo, Grace Smith, Abhinav Triguna, Soshant Bali, Yuxiang Xiang, Prabhakaran Ganesan, Mikel Jimenez Fernandez, Petr Lapukhov, Guyue Liu, Ying Zhang 0022
SIGCOMM17