Critical Communications · Explainer

Designing for the Bad Day

When the network you depend on is the first thing to fail

Designing for the Bad Day
Figure 1 — Redundancy, fallback and direct mode when the network is gone.

Critical communications networks are not built for the average Tuesday. They are built for the flood, the building collapse, the chemical plant fire — the moments when every other system is under stress and the people who need to talk most urgently are the ones whose calls cannot fail. That design philosophy runs all the way from the power supplies in the basement of a base station to the way a radio in a firefighter's pocket behaves when it can no longer hear the network at all.

01Redundancy: building failure in from the start

The first principle is that every single point of failure is a design flaw. A critical network treats redundancy not as an option to add later but as a structural requirement from the start. That means duplicate transmission paths between sites and the core — typically a primary fibre route and a microwave or satellite backup — so that a single cable cut does not island a site. It means duplicate power, with mains, generator and battery arranged so that the transition between them is seamless rather than a brief outage. And it means geographic redundancy in the core network itself: key switching and dispatch functions replicated across physically separate sites so that no single fire, flood or power failure takes down command.

TETRA, the digital trunked standard developed under ETSI and widely used by public safety agencies in Europe and beyond, was designed with this logic at its heart. A TETRA infrastructure can be configured so that individual base stations — called TETRA base stations (TBS) in the standard's terminology — retain local switching capability and continue to serve their coverage area even when they lose contact with the rest of the network. This is fallback mode, and it is a deliberate architectural feature rather than a workaround.

3GPP's mission-critical standards for LTE and 5G NR, developed under the MCPTT and MCData specifications, follow a parallel logic. Isolated E-UTRAN operation for public safety, known as IOPS, defines how an LTE base station can continue providing local connectivity to devices even when the link to the evolved packet core is severed. The network shrinks, but it does not disappear.

This is fallback mode, and it is a deliberate architectural feature rather than a workaround.

From this piece
radio-frequency circuit board detail, macro

02Fallback and graceful degradation

Redundancy buys time and narrows the failure surface, but no design eliminates risk entirely. The next layer of protection is graceful degradation — ensuring that as parts of the system fail, the remainder keeps delivering something useful rather than collapsing entirely.

For critical networks, this often means pre-positioning priority rules in the network itself. When capacity is constrained, emergency communications take precedence over routine traffic automatically, without requiring a human to intervene at a control room. TETRA achieves this through its priority access mechanism; 3GPP's mission-critical standards include access class barring and quality-of-service profiles that reserve headroom for users with the highest priority. A paramedic's voice call over a congested network does not compete equally with a data download — the system is tuned to make that guarantee hold under load.

Graceful degradation also means designing the human interface to reflect the state of the network honestly. A radio that silently fails — appearing to work while actually transmitting into nothing — is more dangerous than one that clearly indicates it is operating in a degraded mode. Indicators of coverage, network registration status and group affiliation are not cosmetic: they tell the user what assurances the system is actually providing in that moment.

03Direct mode: the last resort that must always work

The deepest fallback is direct mode operation, or DMO: radio-to-radio communication with no network involvement at all. When a firefighter descends into a basement beyond the reach of any base station, or when a disaster has damaged the infrastructure entirely, DMO allows terminals to communicate directly on a designated simplex channel.

A vehicle-mounted radio at the surface can relay calls between personnel underground and the command structure above.

From this piece

TETRA defines direct mode as a standard capability, and most public safety terminals support it. The push-to-talk discipline — press, wait, speak — remains identical whether the call is crossing a national trunked network or hopping directly between two handsets thirty metres apart. That consistency is intentional. Under stress, users should not have to remember which mode they are in and adjust their behaviour accordingly.

DMO range is limited — typically a few hundred metres in built environments — but it can be extended through gateway and repeater functions. A terminal that is within network coverage can act as a gateway, bridging DMO traffic back into the trunked network. A vehicle-mounted radio at the surface can relay calls between personnel underground and the command structure above. These mechanisms are also defined in the standard, not improvised in the field.

The architecture of a critical network is, in a sense, a set of nested fallback positions. The full networked system is first preference. Degraded-but-connected operation is second. Isolated local switching is third. Radio-to-radio direct mode is last. Each position is worse than the one before, but each is designed, tested and trained against — because the bad day does not announce itself, and the people caught in it need their communications to work anyway.