DevOps
SRE

Rescue engineering for DevOps teams: how to triage, stabilise and recover failing infrastructure

Rescue engineering: what collapse-response engineers teach us about infrastructure resilience

A field playbook for incident response, blast radius control, observability and disaster recovery — borrowed from the engineers who stabilise collapsed buildings.

Rescue engineering is the discipline that takes over after a structural collapse: engineers who triage damaged buildings in minutes, shore up what is still standing and monitor it continuously so rescue teams can work without becoming casualties themselves. It is engineering in a broken system, with incomplete information and a hard clock.

That is also an accurate description of a serious production incident. At Gart Solutions we spend a large part of our time in exactly that regime — a Kubernetes cluster that will not schedule, a migration that stalled halfway, a database failing over at 3 a.m., a cloud bill that tripled overnight. Normal engineering assumptions have stopped applying, load paths are not what the diagram says, and every change you make could make things worse.

This article is a DevOps playbook written around that idea. Below you’ll find how we triage incidents, how we stabilise systems before investigating them, how we keep blast radius small, what to instrument, and how to build disaster recovery that has actually been tested. The rescue engineering parallels are there because they are genuinely useful models, not decoration.

TL;DR

  • Triage first: decide severity, scope and ownership in minutes — not root cause.
  • Stabilise before you debug. Roll back, fail over or shed load, then investigate with evidence preserved.
  • Keep headroom for the “lateral loads” of production: retry storms, traffic spikes, failing dependencies.
  • Control blast radius with canary releases, cells, feature flags and least privilege.
  • Treat observability like structural monitoring — alert on symptoms and SLO burn, not on every cause.
  • Run more than one recovery path, and test restores on a schedule.

What rescue engineering actually is

Rescue engineering is a branch of civil, structural and geotechnical engineering focused on assessing and stabilising severely damaged structures so search and rescue teams can operate. In the United States it sits inside FEMA’s Urban Search and Rescue system; internationally it is coordinated through INSARAG.

Three details matter for our purposes.

Triage is time-boxed. A Structures Specialist assesses a building in no more than 15 minutes, with a whole sector triaged within two hours. There is no time for modelling; decisions come from visible failure patterns. Structures too unstable to work on are marked “No Go” until proper equipment arrives.

Stabilisation comes before entry. In a trench rescue, walers and pneumatic struts are placed against the walls from outside the hazard zone before any rescuer climbs in. Passive protection is not enough — the shoring has to actively hold the ground in place.

Monitoring runs throughout the operation. Tiltmeters and crack meters watch the structure while people work inside it, triggering evacuation alarms if movement passes a threshold. Triage is redone after every aftershock, because the building is not the same structure it was an hour ago.

Time-boxed triage, stabilise before you enter, monitor continuously, re-assess after every shock. That is also a description of a mature incident management practice.

Rescue engineering practiceDevOps equivalentGart Solutions service
15-minute triage, “No Go” structuresSeverity classification, change freeze on fragile systemsIncidents Management
Shore the walls before enteringRoll back, fail over or shed load before debuggingInfrastructure Reliability Service
Design factor and lateral load allowanceCapacity headroom, autoscaling limits, load testingScaling & Performance Optimization
Tiltmeters and alarm thresholdsSLOs, golden signals, burn-rate alertsMonitoring and Observability
The L-zone around the trenchBlast radius control: canaries, cells, feature flagsDevOps Services, CI/CD
Parallel drilling Plans A, B and CBackups, replicas, multi-region failover, tested restoresBackup and Disaster Recovery (DRaaS)
Standardised hazard markingShared incident taxonomy, runbooks, service ownershipTechnical Support, DevOps Consulting

Triage: the first ten minutes of an incident

The most expensive mistake in incident response is starting with “why”. Rescue teams answer three different questions first: how bad is it, who is affected, and what do we have to work with.

Assign roles immediately. Even a three-person team benefits from separating the incident commander (decides, does not type), the operations lead (makes changes) and the communications lead (updates stakeholders and the status page). Without that split, the person with the deepest knowledge ends up writing Slack updates instead of fixing the system.

Classify severity with written criteria. Vague severity levels produce inconsistent responses. Ours are defined by user impact and reversibility, not by which team owns the component:

  • SEV1 — core user journeys are broken or data is at risk. Wake people up, open a call.
  • SEV2 — significant degradation or a broken path with a workaround. Immediate response, business hours escalation.
  • SEV3 — contained issue, no meaningful user impact yet. Ticket and schedule.

Declare your “No Go” systems. In a collapse, some structures are off-limits until the right equipment arrives. In production, the equivalent is an explicit decision that certain actions are prohibited during the incident — no schema migrations, no restarting a primary with replication lag, no manual edits to state that Terraform manages. Write them down before the incident, not during it.

Time-box the triage itself. If ten minutes of investigation has not produced a working hypothesis, stop investigating and stabilise instead. That single rule cuts mean time to recovery more than most tooling changes.

Stabilise before you debug

Nobody enters an unshored trench. Yet engineers routinely debug a live system that is still moving — tailing logs while the error rate climbs, attaching profilers to a pod that is being OOM-killed, reading a query plan while the connection pool saturates.

Stabilisation options, roughly in order of how often they work:

  1. Roll back the last change. Most incidents follow a deployment or config change. kubectl rollout undo, revert the Helm release, re-apply the previous Terraform state. If your last deploy is not trivially reversible, that is the finding of your next postmortem.
  2. Fail over. Promote a replica, shift traffic to a healthy region or availability zone, switch DNS or the load balancer to a warm standby.
  3. Shed load. Rate-limit, drop non-critical traffic, disable expensive endpoints, pause async consumers and batch jobs. A degraded service is better than one that is down.
  4. Turn features off. Feature flags let you disable the one code path causing the problem without a deploy — the cheapest stabilisation available if you have invested in them beforehand.
  5. Scale out, carefully. Adding capacity helps a resource shortage and makes a thundering-herd problem worse. Know which one you have before you scale.

Preserve the evidence before you restart anything. Rescue engineers document a structure before they change it. Take the snapshot, capture pod logs and heap dumps, export the metrics window, note the exact timestamps. Restarting is often the right move, and it also destroys the state you need to explain the incident afterwards. <!– promo:🛟 –>

Most outages run long for procedural reasons, not technical ones. Gart Solutions helps teams build incident processes, runbooks and reliability practices that shorten recovery — and reduce how often incidents happen at all. See our SRE and managed cloud operations →

Design factors: the headroom you need for the loads you didn’t plan

Timber rescue shoring runs on a thin safety margin — far thinner than a permanent structure — so engineers compensate by adding lateral bracing for forces they cannot predict. Every vertical shore is designed to resist a sideways push equal to a percentage of the weight it carries, and that requirement is raised where aftershocks are expected.

Production has the same lateral loads: retry storms, traffic spikes, a slow dependency, a noisy neighbour, a cold cache after a restart. Most platforms have no explicit allowance for any of them.

Practical headroom rules we apply on client platforms:

  • Know your real ceiling. Load test to failure at least once, in a production-like environment, so the limit is a measured number rather than a guess.
  • Size requests and limits deliberately. In Kubernetes, missing or wrong resource requests are the most common cause of cascading node pressure. Set requests from observed usage, keep limits realistic and use Pod Disruption Budgets for anything with a quorum.
  • Account for autoscaling latency. HPA reacts in tens of seconds; the cluster autoscaler plus node boot is minutes. Your headroom has to cover the gap, or you need pre-warmed capacity for known peaks.
  • Make retries safe. Exponential backoff with jitter, capped attempts, circuit breakers and idempotent handlers. Naive retries turn a brief blip into a self-inflicted denial of service.
  • Watch the quiet quotas. Cloud API rate limits, connection pool sizes, file descriptors, IP address space in a subnet and NAT gateway ports cause more incidents than CPU does.

Blast radius control: the L-zone of your platform

Around every trench there is a zone extending from the edge to a distance equal to the trench depth. No heavy equipment, no traffic, no spoil piles — because a surcharge near the edge is what triggers the second collapse.

The DevOps version is deciding, in advance, how far a single failure or a single mistake can travel.

  • Deploy progressively. Canary or blue-green releases with automated rollback on error-rate or latency thresholds. If a bad release reaches 100% of users, the deployment pipeline is the problem.
  • Partition the system. Separate clusters or namespaces per environment, cell-based or per-tenant isolation for large customers, and separate cloud accounts or subscriptions for production and non-production.
  • Restrict who can push on the edge. No standing human access to production, changes through IaC and pipelines, review on Terraform plans and approvals for destructive operations. Most severe incidents we are called into after the fact involve a manual change nobody reviewed.
  • Freeze changes when the ground is unstable. During a live incident and immediately after it, unrelated deployments are additional load on a structure that is already moving.
  • Isolate the network. Network policies, private endpoints and least-privilege IAM stop a compromised or misbehaving component from taking the rest of the platform with it.

Observability: tiltmeters on the load-bearing columns

During a long extrication, sensors on the structure report movement to the millimetre and sound an alarm before something falls. They monitor a handful of things that matter, with thresholds set in advance, and the alarm has one meaning: get out.

Most monitoring setups we inherit do the opposite — thousands of metrics, hundreds of alerts, no agreement about which ones mean “stop”.

What we put in place instead:

  • Service level objectives. Define what “working” means for each critical journey — availability, latency, error rate — and alert on error budget burn rate rather than on isolated threshold crossings.
  • The golden signals first. Latency, traffic, errors and saturation per service, with everything else as supporting detail during investigation.
  • Alert on symptoms, page on impact. High CPU is information; checkout failing is an alert. Anything that pages a human at night must be both urgent and actionable, with a linked runbook.
  • Traces and structured logs. Distributed tracing turns “the app is slow” into “this call to the payment provider takes 4 seconds”, which is the difference between an hour of guessing and a five-minute fix.
  • One place to look. Prometheus and Grafana for metrics, ELK, Fluentd or Graylog for logs, correlated by request ID and consistent labels. Fragmented tooling costs you minutes exactly when minutes are expensive.

One more borrowed habit: re-triage after aftershocks. After recovery a system is not back to normal — caches are cold, queues are backed up, a hotfix is in place, scaling limits were raised. Keep heightened monitoring for the rest of the day and re-check assumptions after each follow-up change.

Disaster recovery: drill more than one hole

When 33 miners were trapped roughly 700 m underground in Chile in 2010, rescuers ran three independent drilling plans in parallel — a raise-borer, an air-rotary rig and an oil platform rig. Plan B broke through first. The other two were not wasted effort; they were the reason a single failure could not end the rescue.

Most disaster recovery plans we audit are Plan A only: one backup job, one region, one restore procedure, one person who knows how it works.

What an actual recovery capability requires:

  • Written RTO and RPO per system. Different workloads deserve different answers. Without agreed numbers, “we have backups” is a feeling, not a plan.
  • Independent recovery paths. Backups that are not in the same account or subscription as production, with immutability or object lock so ransomware and a bad script cannot delete them. Cross-region replicas for the systems that justify the cost.
  • Infrastructure as code as a recovery mechanism. If your environment exists only as clicked-together resources, rebuilding it is archaeology. Terraform, Pulumi, Bicep or CDK make re-provisioning a pipeline run.
  • Kubernetes-specific plans. Cluster state, etcd or managed control plane recovery, persistent volume snapshots, and a tool such as Velero for namespace-level restores. A cluster is not a backup of itself.
  • Scheduled restore drills. An untested backup is an assumption. Restore into an isolated environment on a schedule, measure how long it takes and fix what you find. Nearly every first drill we run with a new client uncovers something — a missing secret, an undocumented dependency, a restore that takes six hours against a two-hour RTO.
  • A rollback in every change. The Chilean rescue capsule had a bottom escape hatch in case it jammed mid-shaft. Every deployment deserves the same assumption: a reverse path that has actually been exercised.

<!– promo:📡 –>

Backups are not a recovery plan until someone restores them. Gart Solutions designs and tests Backup and Disaster Recovery (DRaaS) setups with explicit RTO and RPO targets, plus the monitoring to know when you need them. Explore our DevOps and platform engineering services →

Shared language: the INSARAG marking of your platform

Rescue teams from different countries can hand over a worksite because the marking system is standardised: a painted box with the site ID, the team, the assessment level and the date, hazards written above it, the triage category below and an arrow pointing at the access route. No conversation required.

Engineering organisations lose time in incidents for the opposite reason — nobody is sure who owns a service, where its runbook lives or what its dependencies are. Fixes that cost little and pay off in the first incident:

  • A service catalogue with a named owner and an on-call rotation for every production component.
  • Runbooks next to the alerts that trigger them, kept short and current.
  • One place where incident state lives — severity, commander, timeline, current hypothesis — so anyone joining can read in rather than ask.
  • Blameless postmortems with a small number of committed, assigned actions, and a review of whether last month’s actions actually shipped.

How Gart Solutions works on failing infrastructure

Gart Solutions is a Kyiv-based team focused on DevOps, cloud solutions and infrastructure. Our engineers average 8.2 years of experience, around 70% are senior-level certified professionals, and we have delivered more than 50 successful projects. <!– stats –>

8.250+10+70%
Years of average experienceSuccessful projectsSenior and middle specialistsSenior-level certified engineers

Our engagement model follows the same sequence as a structured deployment: a free consultation to identify where the pressure is, a technical audit and architecture vision before anything is touched, alignment on targets, implementation, documentation and reporting for the client’s own team, and ongoing maintenance and support after delivery.

The services most relevant to reliability work: incidents management, monitoring and observability, infrastructure reliability, backup and disaster recovery (DRaaS), scaling and performance optimisation, and technical support — supported by DevOps as a Service, CI/CD, Kubernetes cluster and container management, cloud migration and cost optimisation, and managed AWS, Azure and GCP.

Day to day that means Terraform, Pulumi, AWS CDK and Azure Bicep for infrastructure as code; Jenkins, GitLab CI, GitHub Actions and Azure DevOps for pipelines; Kubernetes, OpenShift, Rancher and Nomad for orchestration; Prometheus, Grafana, ELK, Fluentd, Graylog and New Relic for observability; and ISO 27001, NIST and CIS practices with HashiCorp Vault, NeuVector and SonarQube on the security side.

Case study: infrastructure for a national healthcare platform

A healthcare client needed CI/CD and infrastructure for development and production environments serving a common base of medical records, prescriptions and insurance data across the country’s medical centres. We built it on hardware from a local provider, GiGa Cloud, using VMware ESXi and Terraform, connected it to the government E-Health platform and applied data masking in dev and test to satisfy GDPR and HIPAA requirements. Infrastructure creation was automated, environments became scalable on demand, and delivery moved to a self-managed pipeline on Jenkins, Docker and Kubernetes.

Case study: finding the real load path in Azure

For another client, most of the Azure bill came from a load balancer and network traffic. We configured the file share with a private endpoint in the same VNet as the AKS nodes so traffic stayed internal rather than crossing the load balancer. Network costs dropped by 90% — up to $400 a day — and we handed over recommendations for performance, security and reliability alongside the savings. The lesson generalises: find the actual load path before you reinforce anything.

Conclusion

Rescue engineering is a useful model for DevOps because it is honest about working in a broken system. It assumes information is incomplete, time is short and the structure may move again — and it responds with process: triage quickly, stabilise before you enter, monitor continuously, keep more than one way out.

Teams that adopt those habits do not have fewer hard nights because their systems never fail. They have fewer because failure stops being improvisation.

Fedir Kompaniiets

Fedir Kompaniiets

Co-founder & CEO, Gart Solutions · Cloud Architect & DevOps Consultant

Fedir is a technology enthusiast with over a decade of diverse industry experience. He co-founded Gart Solutions to address complex tech challenges related to Digital Transformation, helping businesses focus on what matters most — scaling. Fedir is committed to driving sustainable IT transformation, helping SMBs innovate, plan future growth, and navigate the “tech madness” through expert DevOps and Cloud managed services. Connect on LinkedIn.

Not sure how your infrastructure would hold up in a collapse?

Gart Solutions helps engineering teams stabilise, monitor and recover cloud infrastructure — from incident management and observability to tested disaster recovery. Book a free consultation.

Talk to our architects →

Let’s work together!

See how we can help to overcome your challenges

FAQ

What is rescue engineering, and why apply it to DevOps?

Rescue engineering is the structural and geotechnical discipline that stabilises collapsed buildings, trenches and mines so rescue teams can work safely. It is useful to DevOps because its constraints match those of a production incident: incomplete information, an unstable system, severe time pressure and real consequences for acting carelessly. The practices it has developed — time-boxed triage, stabilisation before investigation, continuous monitoring, parallel recovery plans — translate almost directly into incident management and disaster recovery practice.

What should a DevOps team do in the first ten minutes of an outage?

Assign roles (incident commander, operations, communications), classify severity against written criteria, and establish scope — which users and journeys are affected. Then decide whether to stabilise or investigate. If ten minutes of investigation has not produced a working hypothesis, stabilise first: roll back the last change, fail over, or shed load. Capture snapshots and logs before restarting anything so the evidence survives. Root cause analysis belongs in the postmortem, not in the first ten minutes.

How do you reduce blast radius in cloud infrastructure?

Deploy progressively with canary or blue-green releases and automated rollback thresholds. Separate environments into different clusters, accounts or subscriptions. Use cell-based or per-tenant isolation for large customers. Remove standing human access to production and route changes through infrastructure as code and reviewed pipelines. Apply network policies, private endpoints and least-privilege IAM so one compromised component cannot reach everything. Finally, freeze unrelated changes during and immediately after incidents.

How often should disaster recovery be tested?

At minimum annually for full recovery scenarios, and quarterly for restore drills of critical data stores — more often for systems with aggressive RTOs or frequent architectural change. The test that matters is a real restore into an isolated environment, timed against your stated RTO and RPO. Most first drills reveal a gap: a missing secret, an undocumented dependency or a restore that is far slower than assumed. Gart Solutions runs these drills as part of its DRaaS and DevOps engagements.

What's the difference between incident management and disaster recovery?

Incident management handles live disruptions — detection, severity classification, coordination, stabilisation and follow-up learning. Disaster recovery covers the larger failures where data, a cluster or an entire region is lost and must be rebuilt within agreed RTO and RPO targets. In rescue terms, incident management is triage and shoring; disaster recovery is the engineered extraction route prepared in advance. Mature teams need both, and the second is far harder to improvise.

What is the difference between disaster recovery and incident management?

Incident management is about responding to live disruptions: detecting them, classifying severity, coordinating responders, stabilising service and learning afterwards. Disaster recovery covers the larger failures where a system, data centre or region is lost and must be restored from backups or replicas within agreed recovery time and recovery point objectives. In rescue terms, incident management is triage and shoring; disaster recovery is the planned extraction route. Gart Solutions offers both, including Backup and Disaster Recovery as a Service.

How do we know whether our monitoring is good enough?

Ask three questions. Would you learn about a major outage from your monitoring before a customer tells you? Does every alert that pages someone at night have a clear action and a linked runbook? Can you tell, for each critical user journey, whether it is currently meeting its objective? If any answer is no, the gap is usually not more metrics but fewer, better-defined ones: service level objectives, the golden signals per service, and burn-rate alerting instead of dozens of static thresholds.

How do I know if my infrastructure needs an external DevOps engagement?

Warning signs include repeat outages with similar causes, recovery that takes hours, deployments the team is afraid to run, cloud bills nobody can explain, backups that have never been test-restored, and critical knowledge held by one or two people. A technical audit that maps how load and failure actually move through your system is usually the fastest way to find the real weak points. Gart Solutions begins every engagement with a free consultation and an audit before recommending any changes — you can see our services here.
arrow arrow

Thank you
for contacting us!

Please, check your email

arrow arrow

Thank you

You've been subscribed

We use cookies to enhance your browsing experience. By clicking "Accept," you consent to the use of cookies. To learn more, read our Privacy Policy