Bare MetalSRE & DevOps · TermPhysical computing hardware used directly without a general-purpose virtualization layer between the operating system and the machine.Read more
CloudSRE & DevOps · ConceptA model for accessing computing storage networking and managed technology services through remotely operated infrastructure.Read more
Error BudgetSRE & DevOps · MetricThe amount of unreliability permitted while still meeting a service level objective.Read more
Graceful DegradationSRE & DevOps · ConceptA system's ability to preserve its most important functions when part of it is impaired.Read more
IncidentSRE & DevOps · TermAn unplanned event that degrades or interrupts a service or business operation.Read more
InfrastructureSRE & DevOps · TermThe foundational compute network storage runtime and supporting services on which software operates.Read more
InstanceSRE & DevOps · TermA particular running or allocated occurrence of a service application virtual machine container or other resource.Read more
LoggingSRE & DevOps · PracticeThe recording of structured events and diagnostic information produced by software and infrastructure.Read more
MTBFMean Time Between Failures · SRE & DevOpsThe average operating time between one failure and the next.Read more
MTTRMean Time to Recovery · SRE & DevOpsThe average time needed to restore service after a failure.Read more
ObservabilitySRE & DevOps · ConceptThe ability to understand a system's internal state from the signals it produces.Read more
On-CallSRE & DevOps · PracticeA rotation where designated people respond to urgent service issues outside normal planned work.Read more
On-PremisesSRE & DevOps · TermTechnology deployed and operated within infrastructure controlled at an organization's own facilities or private locations.Read more
PlaybookSRE & DevOps · TermA reusable strategy and set of actions for handling a broader situation.Read more
PostmortemSRE & DevOps · PracticeA retrospective document or meeting that records an incident's impact causes response and follow-up actions.Read more
RCARoot Cause Analysis · SRE & DevOpsA structured investigation into the underlying conditions that produced an unwanted outcome.Read more
RPORecovery Point Objective · SRE & DevOpsThe maximum acceptable amount of data loss measured backward in time.Read more
RTORecovery Time Objective · SRE & DevOpsThe target maximum time for restoring a service after disruption.Read more
RunbookSRE & DevOps · TermA documented procedure for diagnosing or handling a recurring operational situation.Read more
SLAService Level Agreement · SRE & DevOpsA formal agreement defining service commitments and consequences when they are missed.Read more
SLIService Level Indicator · SRE & DevOpsA measurement used to evaluate a specific aspect of service performance.Read more
SLOService Level Objective · SRE & DevOpsA target level of service performance measured over a defined period.Read more