Failure Latency Budgets: Engineering Systems That Are Allowed to Slow Down
DOI:
https://doi.org/10.64137/3107-9458/ICACSIS-107Keywords:
Failure Latency Budgets, Graceful Degradation, Distributed Systems, Reliability Engineering, Service Level Objectives (Slos), Tail Latency, Fault Tolerance, Resilience Engineering, Performance Optimization, Cloud-Native ArchitectureAbstract
Conventionally, availability and performance engineering have been mainly based on a binary viewpoint: systems are either operational or not, healthy or failed, meeting SLA or violating it. This outlook has contributed to the development of redundancy, fault tolerance, and high-availability architectures. Nevertheless, as distributed systems become more complex and interconnected, the disadvantages of this up/down model are getting more and more evident. A lot of recent failures do not show up as complete outages but rather as slow, cascading degradations, raised latency, unavailability of a part of the features, and resource contention that gradually, and even strongly, decays user experience well before a formal outage has been declared. This paper presents Failure Latency Budgets as an idea, a technical and operational framework that clearly defines and distributes an acceptable slowdown under stress, considering controlled degradation as a normal and equal reliability goal. The main point is that sturdier systems should not just be designed to be failure-free but also to fail slowly, predictably, and transparently. We put forward a method that combines latency limits, dependency-aware service prioritization, adaptive load shedding, and user-focused service tiers into the scheme for planning reliability. Our results indicate that by legitimizing slowdown as a permitted and engineered condition, there is an enhancement of both technical resilience and organizational decision-making during incidents. Changing the way reliability is thought of by basing it on controlled performance decay rather than binary uptime, this paper adds a practical supplement to site reliability engineering (SRE) tactics and provides a more refined model for the design of distributed systems that operate dependably even in the presence of real-world contingencies.
References
[1] Thekkath, Chandramohan A., and Henry M. Levy. "Limits to low-latency communication on high-speed networks." ACM Transactions on Computer Systems (TOCS) 11.2 (1993): 179-203.
[2] Chachere, John, John Kunz, and Raymond Levitt. "The role of reduced latency in integrated concurrent engineering." CIFE, WP 116 (2009).
[3] So, Kelvin CW, and Emin Gün Sirer. "Latency and bandwidth-minimizing failure detectors." Proceedings of the 2Nd ACM SIGOPS/EuroSys European Conference on Computer Systems 2007. 2007.
[4] Rumble, Stephen M., et al. "It's time for low latency." 13th Workshop on Hot Topics in Operating Systems (HotOS XIII). 2011.
[5] Aqeel, Waqar. The Latency Budget: How to Save and What to Buy. Diss. Duke University, 2021.
[6] Hu, Biao, et al. "On-the-fly fast overrun budgeting for mixed-criticality systems." Proceedings of the 13th International Conference on Embedded Software. 2016.
[7] Wozniak, Ernest, et al. "Assigning time budgets to component functions in the design of time-critical automotive systems." Proceedings of the 29th ACM/IEEE international conference on Automated software engineering. 2014.
[8] Bennis, Mehdi, Mérouane Debbah, and H. Vincent Poor. "Ultrareliable and low-latency wireless communication: Tail, risk, and scale." Proceedings of the IEEE 106.10 (2018): 1834-1853.
[9] Das, Shidhartha, et al. "A self-tuning DVS processor using delay-error detection and correction." IEEE Journal of Solid-State Circuits 41.4 (2006): 792-804.
[10] Barber, Patrick, et al. "Quality failure costs in civil engineering projects." International Journal of Quality & Reliability Management 17.4/5 (2000): 479-492.
[11] Beyer, Betsy, et al. Site reliability engineering: how Google runs production systems. " O'Reilly Media, Inc.", 2016.
[12] Abdul-Rahman, H., et al. "Delay mitigation in the Malaysian construction industry." Journal of construction engineering and management 132.2 (2006): 125-133.
[13] Rajendran, Jeyavijayan, et al. "Fault analysis-based logic encryption." IEEE Transactions on computers 64.2 (2013): 410-424.
[14] Elbamby, Mohammed S., et al. "Toward low-latency and ultra-reliable virtual reality." IEEE network 32.2 (2018): 78-84.
[15] Parvez, Imtiaz, et al. "A survey on low latency towards 5G: RAN, core network and caching solutions." IEEE Communications Surveys & Tutorials 20.4 (2018): 3098-3130.


