Resilience is usually assumed, not designed
Most multi-site WANs are resilient in theory and fragile in a specific, unexamined place. The architecture looks sound on a diagram, but nobody has traced what actually happens when a particular link, device, or data centre fails. Resilience was assumed during the build and never tested since. The result is an estate that runs well until the one untested failure occurs, usually at the worst time.
Building a truly resilient WAN is less about buying more hardware and more about finding the assumptions and removing the single points of failure hiding behind them.
Where the single points of failure hide
Even in well-run environments, the same weak points recur:
- A single circuit per site. One transport means one failure takes the site offline, however good the equipment behind it.
- Redundant links on the same physical path. Two circuits that share a duct or a provider backhaul are not two circuits. a single cable cut takes both.
- A single device. Dual circuits into one router simply move the single point of failure from the line to the box.
- Centralised dependencies. If every site backhauls through one data centre for internet or security, that data centre is a shared point of failure for the entire estate.
- Failover that has never been tested. A backup path that has not been exercised is a hope, not a control.
Trace one failure end to end
Pick a site and ask precisely what happens if its primary circuit fails right now. Who notices, how fast does traffic move, what depends on that path, and has the failover been tested this year. Following that single question end to end tends to reveal more than any architecture diagram.
The layers of WAN resilience
Resilience is not one control. It is several, applied where the risk justifies them.
Transport diversity
Genuine redundancy means more than a second circuit. It means diverse transports (for example fibre plus a separate broadband or cellular path) that do not share physical infrastructure, so no single event can take them both down.
Device and path redundancy
Removing the single-device failure point where it matters, and ensuring that traffic has an alternative path through the network, not just an alternative line into the building.
Intelligent, application-aware routing
This is where SD-WAN earns its place. Using multiple transports simultaneously with sub-second failover turns the loss of a link into a non-event rather than an outage, and lets the network keep the most important applications on the healthiest path.
Distributed rather than single-point services
Moving internet breakout and security enforcement closer to each site, rather than backhauling everything through one location, removes a shared dependency and improves performance at the same time. Extending this into a secure multi-site design applies the same standard everywhere.
Failover that actually works
The most common resilience failure is not the absence of a backup path. it is a backup path that does not work when called upon. Failover has to be fast enough that applications survive it, automatic enough that it does not wait on a human, and, above all, tested. A resilience design that is never exercised degrades quietly: circuits change, configurations drift, and the untested path stops being viable without anyone noticing. Regular testing is what keeps resilience real.
"A backup circuit you have never failed over to is not resilience. It is a theory. The only resilient WAN is one whose failure paths have actually been exercised."
Edge7 Networks, Networking PracticeResilience includes security
Availability is one dimension of resilience. security is another, and they should be designed together. A distributed WAN that improves availability while weakening security enforcement has traded one form of fragility for another. Consistent policy and inspection across every site, so that a resilient estate is also a defensible one, is part of the same design rather than a separate project.
A structured way forward
Improving WAN resilience does not require rebuilding the network. It starts with an assessment that maps the current topology, identifies the single points of failure, and prioritises them by the impact of the site or service they affect. From there, resilience is added where it matters most first, rather than everywhere at once. This keeps the work proportionate and lets the highest-risk gaps close early.
Edge7 Networks designs and runs resilient multi-site networks for organisations across Ireland and the UK, and has worked on WAN and SD-WAN since 2018. If you are not certain your WAN would survive the failure you have not tested, an assessment is the right place to begin.