Fastpass: A Centralized “Zero-Queue” Datacenter Network
Name
sigc110-perry.pdf
Description
Paper
Size
1.77 MB
Format
Adobe PDF
Checksum (MD5)
b0836e0eea0fa2a4b866e297e342717b
Author(s) • • • •
Perry, Jonathan
Ousterhout, Amy Elizabeth
Balakrishnan, Hari
Shah, Devavrat
Fugal, Hans
Date Issued
August 2014
Journal
SIGCOMM 2014 proceedings
Publisher
Association for Computing Machinery
Citation
Perry, Jonathan, Amy Ousterhout, Hari Balakrishnan, Devavrat Shah, and Hans Fugal. "Fastpass: A Centralized “Zero-Queue” Datacenter Network." ACM SIGCOMM 2014, Chicago, Illinois, August 17-22, 2014, pp.307-318.
Version
Author's final manuscript
Abstract
An ideal datacenter network should provide several properties, including low median and tail latency, high utilization (throughput), fair allocation of network resources between users or applications, deadline-aware scheduling, and congestion (loss) avoidance. Current datacenter networks inherit the principles that went into the design of the Internet, where packet transmission and path selection decisions are distributed among the endpoints and routers. Instead, we propose that each sender should delegate control—to a centralized arbiter—of when each packet should be transmitted and what path it should follow. This paper describes Fastpass, a datacenter network architecture built using this principle. Fastpass incorporates two fast algorithms: the first determines the time at which each packet should be transmitted, while the second determines the path to use for that packet. In addition, Fastpass uses an efficient protocol between the endpoints and the arbiter and an arbiter replication strategy for fault-tolerant failover. We deployed and evaluated Fastpass in a portion of Facebook’s datacenter network. Our results show that Fastpass achieves high throughput comparable to current networks at a 240 reduction is queue lengths (4.35 Mbytes reducing to 18 Kbytes), achieves much fairer and consistent flow throughputs than the baseline TCP (5200 reduction in the standard deviation of per-flow throughput with five concurrent connections), scalability from 1 to 8 cores in the arbiter implementation with the ability to schedule 2.21 Terabits/s of traffic in software on eight cores, and a 2.5 reduction in the number of TCP retransmissions in a latency-sensitive service at Facebook.
MIT Department
Massachusetts Institute of Technology. Computer Science and Artificial Intelligence Laboratory
Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science
Terms of Use
Creative Commons Attribution-Noncommercial-Share Alike
Persistent DSpace Link
DOI of Published Version
http//dx.doi.org/10.1145/2619239.2626309
https://doi.org/10.1145/2619239.2626309