Telstra’s Critical Infrastructure Failure was Easy to Avoid

In July 2026 a fault within the Telstra telecommunications network caused a broad infrastructure outage to phones, emergency calls, payments, trains, and many other services. A time server restarted badly, told the network the wrong time, and the network then started refusing normal connections because its security checks no longer matched.

The press coverage highlighted a software patch that was not applied, and an undocumented change. Telstra executives have said that their procedures were followed, and that these procedures were not good enough. They also said a good deal about the complexity of the technology and that they could not promise that outages won’t happen in the future.

What has not been widely discussed is how poor the maintenance procedures appear to be, nor how basic the steps to avoid this infrastructure failure were.

With any piece of infrastructure undergoing maintenance there is a testing regime before returning the infrastructure to service. Sometimes this testing is quite simple, such as checking a road is clear after removing a fallen tree, or checking that a clock has the right time before putting it in charge of a telecommunications network.

And such a basic step is apparently what was not done; or was done in a way that did not help.

Now the maintenance that was done on the time server was not expected to impact its ability to tell the time; but any electrician will tell you to test your meter before you use it on live equipment. This ensures that an unexpected fault in the meter does not injure or kill the electrician or others. Testing the time server would have ensure people could call for an ambulance, police, fire brigade…

This begs the question – are basic tests not being done?

A well-managed time time service has safeguards. One is to automatically compare the time of all time servers against each other, and either isolate those that differ significantly from the majority, or flag the variation as a critical fault, allowing humans to rectify the problem rapidly. Another is that the various systems that use time servers compare the time of different servers, and ignore any that significantly disagree with the others.

In Telstra’s case it seems that the underlying design assumption is that none of the top-level time servers will ever be wrong, as the faulty time server was:

  • not automatically taken offline
  • not monitored and flagged as an urgent critical failure
  • implicitly trusted by every attached system

Monitoring is not complicated. Nor is it expensive to set up.

This begs the question – are there simple design flaws within Telstra’s critical infrastructure?

The takeaway for all infrastructure managers is this:

Always measure and test infrastructure components before returning them to service.

Leave a Reply

Your email address will not be published. Required fields are marked *