Telstra outage: The night a network decided the twelvemonth was 2006

Hacker News by 18 min read 122x views
Telstra outage: The night a network decided the twelvemonth was 2006

Share Post

The contrary is really the case. If we do not have a average understanding of what “now” is, a lot of things we obtain for granted volition halt working.

This summer, Australia realized this the difficult way. Let me obtain this chance to provision an example of why period matters to a contemporary society, what happened in this particular case, and what our key takeaways are.

On 8 July 2026, a ample part of the mobile network run by Australia's largest division phone controller Telstra stopped operating alongside sound calls not getting through or content messages that didn't arrive. There were equal calls to Australia's emergency figure that did not get through. But the outage affected systems equal additional apart, specified as trains, fee terminals, ticketing systems and EV chargers being disrupted.

No networks were attacked. Nobody accidentally cut the fiber. All systems had electric power. The culprit, you may ask? A sole GPS receiver in a sole chassis in Melbourne coming rear from scheduled care believing the twelvemonth was 2006, and the remainder of the network was persuaded to accept it.

Telstra commissioned an independent assessment from the business Technology Audit Partners (TAP) and the study is a extremely engaging read, since the identical category of nonaccomplishment could appear in many crucial services, including several we depend on to keep group alive.

In command for many of these systems to function, period being correct, or at smallest the identical everywhere, is crucial. And as everyone in the endeavor knows, “correct” is a related term. There is no exact time, lone period held inside a certain border of a reference. How broad that border may be depends entirely on what you are doing or in which endeavor you run in. 

Running a mobile network, akin Telstra, is concerning as time-dependent as a endeavor gets. 

Modern mobile communication explains why period matters

Modern division phone protocols volition not activity without precision time. Mobile networks distinct uplink data from downlink by either FDD (Frequency Division Duplex) or TDD (Time Division Duplex). 

FDD gives all direction its own piece of spectrum, so the two can run continuously without colliding. 

TDD alternatively uses an complete sole obstacle of spectrum for the two directions, alternating between transmitting and receiving in extremely abbreviated intervals. 

Since far additional data normally flows downward than up, FDD's fixed ratios depart much of the uplink spectrum idle, during TDD can change the proportion to equivalent the genuine traffic. That is why most contemporary 5G spectrum, including Sweden's chief 5G collection at 3.5 GHz, is TDD. 

It is additionally why TDD depends on exact time: all division on the identical frequence has to toggle direction in stage alongside all other. A division that lets its clock drift volition transmit data into its neighbour's obtain window, alongside the outcome that the network volition commencement jamming itself.

Rather than allocating spectrum on keeping the two directions apart, the industry chose to depend on period accuracy, and accepted a difficult dependency on all division agreeing concerning whenever "now" is. Thus, being reliant on period is a scheme choice. 

However, stated how much we all depend on the systems being capable to concur on “now”, it is slightly puzzling that period is not stated as much deliberation as it deserves. And that is a instruction that is extremely apparent from the published report.

So, what really happened?

Architecture of period distribution.

To commencement with, it is critical to comprehend the architecture of period distribution.

Time allocation protocols all build hierarchies; Network Time Protocol (NTP), which Telstra has deployed according to the report, expresses its ranking in strata.

  • Stratum 0 is the citation itself, for case a GPS receiver, or Netnod’s nuclear clocks.
  • Stratum 1 is a device synchronised immediately to a stratum 0 reference, for example the NTP servers that Netnod provides.
  • Stratum 2 synchronises from a stratum 1 server, stratum 3 from a stratum 2, and so on.

In Telstra's case, that ranking had a particular shape, at smallest to commencement with. This scheme from 2010 had at the top stratum 1 sources at Australia's National Measurement Institute (NMI), which maintains the country's national period scale, much as the Research Institute of Sweden does in Sweden. Telstra drew period from those external references into two stratum 2 servers of its own, in Sydney and Melbourne, which in rotate fed three stratum 3 servers, in Sydney, Melbourne and Perth.

Below them sat the clients. In this environment that does not average laptops or phones, but the complete mobile network infrastructure, for case nodes handling handovers between division sites. There were thousands of nodes all complete a huge earth discipline and all one of them needed to have the identical idea of what “now” is, to inside a few millionths of a second.

The TAP study describes this setup as “fit for purpose” and that it gave Telstra “a extremely dependable and authoritative citation period origin from NMI”. 

Stratum in itself does not say if the period is accurate, lone the figure of steps from a server to its reference. A stratum 1 server alongside a bad period citation is motionless a stratum 1 server.

Protection against bad period sources

NTP volition hence need defence against bad period sources. In fact, it has two distinct ones, and they do distinct things, the two of which assume they are autonomous from all other.

  1. Among alternatively comparable candidates, the lesser stratum carries additional weight. This is the scheme that determines which origin a client settles on.
     
  2. NTP compares multiple sources and discards those that differ alongside the rest. A sole origin claiming an implausible period is outvoted and dropped, despite of how authoritative it claims to be.

Neither defence is particular to any particular disruption; together they defend against a damaged receiver, a misconfigured server, or an external attack. But these protective measures lone activity if the period sources that the clients hear to are genuinely autonomous of all other.

Two ways to deploy NTP

NTP can be deployed in two ways. In client/server manner the association is declared and directional: a node takes period from those servers, and nothing else. 

The 2010 Telstra setup was in actuality specified a client/server model. Peering was allowed, but lone at the identical stratum flat and the TAP report, as noted in the beginning, described this setup as "fit for purpose”.

The another way is a symmetric (peering) mode, anywhere nodes toggle period mutually and resolve on whichever origin the algorithms currently favour.

Peering is elastic and survives the defeat of a origin gracefully. But it additionally method the topology in manufacturing is emergent fairly than designed. What you documented is a setup that could quietly rearrange itself into a form no one always approved.

Telstra's 2020 upgrade

In 2020 the mobile center timing scheme was upgraded, and new hardware was installed, including a new NTP timing chassis. That facility introduced a few changes.

The archetypal one was forced. The new chassis could not let a stratum 2 server nourish a stratum 3 server inner the identical box, so the two had to be wired throughout all other: Sydney's stratum 3 took its period from Melbourne's stratum 2, and Melbourne's stratum 3 from Sydney's. 

In reality, alternatively of having two stratum 2 sources, all location was remaining alongside lone one. The TAP study notes that this degradation in redundancy was known and accepted. A second alter was leaving the client/server-model in favor of the peering model. The study is not apparent concerning the motivation, but it is sensible to propose that one wanted compensation for this defeat of redundancy. With all location having fair one origin alternatively of two, letting the servers discover their own replacements could provision the belief of improved resilience. 

The TAP study plainly states that the defeat of resilience was known. However, it fails to discover any evidence that the resulting hazard of so-called “timing loops” was identified.

What is a timing loop?

A timing iteration is the network equal of believing a rumour to be true by asking three group who all heard it from all other. Each one agrees, so it must be true. NTP plant basically the identical way: it compares multiple sources and discards whichever disagrees alongside the rest.

As you may recall from above, NTP has two defenses against bad period sources. The second one protects against timing loops, but lone if the sources are autonomous of all other. In specified a loop, sources that appear autonomous are in fact taking their period from all other, either immediately or indirectly by tracing rear through a shared reference.

Once a incorrect value is circulating, the sources volition commencement agreeing alongside all another and the ballot volition be in favor of the majority’s opinion, equal although the value is wrong.

The protocol worked. The architecture did not.

Both of NTP’s types of defenses came to be impaired in Melbourne, but five years apart. Not deliberately, but by choices, all of them defensible on their own terms: a hardware limitation had to be worked around, and later, a recurring error had to be stopped. Each decision solved the issue in forefront of it. Nobody was asked to appearance at the sum of all actions.

The second defence was the archetypal one to be disabled. The introduction of peering in 2020 made timing loops possible, and five years later, specified a iteration showed up. In Melbourne a server started taking period from a node below itself. That should have set off alert bells. However, since exact period was motionless reaching the network by another paths, no genuine damage was done. The underlying problem, the circular dependency, was there, but no one issued a ticket concerning it.

The genuine grievance was fairly obvious. Melbourne kept losing communication alongside its lone stratum 2 origin in Sydney. With no fallback configuration, the server used peering to discover a replacement, sometimes a node below it in the hierarchy. 

In October 2025, engineers activated the GPS receiver that had been sitting unused in the Melbourne chassis since 2020 and connected it to the stratum 3 server, as a replacement for the unreliable Sydney source.

By all apparent measure it seemed to have worked. Melbourne now had a dependable origin of its own and the alarms stopped. But the fix lone addressed the symptom, not the base cause. Nobody established why Melbourne kept losing its Sydney origin in the archetypal place. The underlying issue was motionless current in the network by July 2026.

To create things equal worse, nobody seems to have understood what activating the GPS cardstock did to the architecture. By adding the GPS card, the Melbourne server went from a stratum 3 server to stratum 1. The engineers didn’t add a origin next to the another ones. By promoting a server to the identical position as the national  reference, a new origin was created at the extremely top. As far as NTP is concerned, they transport the identical weight. 

Suddenly this GPS cardstock in a chassis in Melbourne, installed as a workaround and reviewed by no one, became the most authoritative server in the ranking for the largest mobile network in Australia.

Needless to say, practically nothing of the 2010 scheme remained.

By July 2026 the network had a sole origin that was the two the most authoritative applicant accessible and unopposed, since the sources that could have contradicted it were downstream of it.

This behavior was extremely difficult to spot. The network served exact period all day for years. Architectures akin this do not normally degrade gradually. They work, and they keep working, correct up until they stop.

GPS week figure rollover

The second component is a well-known asset of GPS.

GPS broadcasts period as a week figure affirmative seconds-into-week, counted from an era that began in first January 1980. In the chief civic GPS signal, the week figure site is 10 bits, ie a maximum of 1,023 weeks. Every 1,024 weeks, or 19.6 years, the oppose starts over. This has happened twice: in August 1999 and in April 2019.

Working out which figure of era it is and adding the correct multiple of 1,024 weeks, is the job of the receiver. And the data needs to be in its firmware. 

The issue that occurred in Australia was not a delayed consequence of any of the GPS rollovers. The cardstock in Melbourne had passed through the second rollover in 2019 without trouble, since a receiver that keeps operating additionally keeps counting. Each new week is merely added to the one before, and the inquiry of which era it belongs to is never raised.

However, formerly you rotate it off, that cognition is gone. When it is turned on again, the receiver has to activity out the era from scratch, and all it has to go on is what its firmware assumes. The firmware on the Melbourne cardstock had not been updated. Upon start-up it cut rear on the before era and placed the date 1,024 weeks in the past.

What happened next is finest understood as the two defences being impaired whenever they were needed the most.

The archetypal defence, the lesser stratum carrying additional weight, classified the Melbourne server highest, since the attached GPS cardstock promoted it to a stratum 1 server. This was according to NTP protocol and thus steered clients towards the one origin which was 1,024 weeks wrong.

The second defence, outliers being voted down, was never engaged, since nothing was remaining to acknowledge Melbourne as an outlier. NTP does not ask whether a date is plausible; it asks whether a origin disagrees alongside the others. The 2010 setup had two stratum 2 servers. If one of them had started announcing the twelvemonth 2006, the another one would have stayed alongside 2026 and no agreement would have been reached. That would not have been ideal, but at smallest the incorrect date would not have spread. 

But Melbourne's stratum 2 equal had been switched off by the extremely identical chassis replacement, and the remaining sources were downstream of Melbourne. As the incorrect date spread, they began reporting it back. Agreement grew, and accord is what the algorithm is looking for.

So the clients did what they were built to do. Once a bulk of a client's sources accepted on November 2006, the client accepted the date, and the additional the date travelled, the additional convincing it became.

Neither defence malfunctioned. Both had merely been deprived of what they depend on: one needed a origin value ranking highest, the another needed sources capable of disagreeing. Two decisions, five years apart, had removed all in turn.

Key takeaways from the incident

Prioritise and categorize period and frequence allocation as crucial infrastructure. Manage it accordingly

Document all functions that can obtain the entire network alongside them, and put timing on that list. Classification is not paperwork; it is what determines alter hazard category, assessment depth, staffing levels, monitoring safety and prosperity priority. Telstra's study is, at bottom, the narrative of one missing admission on that catalog and everything that followed from it.

Document the entire infrastructure, and all alter to it

There was no chief repository of NTP configuration, no aureate configuration, and no documented document of the servers another than the devices themselves. Without records you cannot execute meaningful pre-checks, you cannot measure impact, and during an event you cannot inform what "correct" looks like.

Build redundancy in competence

Two engineers performed the change, and the two were on mandatory stand-down before the consequences of the GPS cardstock reboot were understood.

Depth of ability is a resilience asset exactly akin a redundant power feed. A sole specialist, or a pair, method no second opinion, and no one to ask in the center of the night whenever care is normally done.

Run a safety inspection of the period and frequence infrastructure

Treat timing as an assault exterior akin any another and analyze it accordingly.

Start alongside anywhere period enters the organisation. A GNSS indication arriving from area is feeble and unauthenticated, and can be jammed or spoofed by cheap equipment. If that indication is your lone reference, person exterior your construction can decide what period you think it is.

Then appearance at how it travels. Time distributed complete a shared network can be intercepted and manipulated on its way to the client.

Then appearance at who is allowed to speak. Which servers may your clients obtain period from, and who decided that? A origin that nobody authorised is a origin nobody is checking.

And do not halt at deliberate attack. A timing iteration produces much the identical consequence as a prosperous spoofing attack: a origin the network trusts, delivering a value nobody can contradict. 

Use point-to-point connections

It is uncomplicated to see the petition of peering. It feels akin resilience alongside sources that rear all another up: a network that heals itself whenever a node disappears. But redundancy that arranges itself is not redundancy you can depend on. 

There are safer ways to accomplish a akin flat of robustness. Netnod runs dedicated point-to-point connections: all association is known and documented. Every origin is known, and the topology stays the way we designed it. Redundancy comes from multiple autonomous sources deliberately configured, not from nodes negotiating amongst themselves.

Build an productive alert organisation

Alarms from the timing phase were not in the norm monitoring tools, and were reviewed lone during endeavor hours by a fistful of people. Client-side alarms carried neither the severity nor the item to run contiguous action. Getting this correct is organisational as much as technical: alarms attain 24x7 monitoring, severities indicate genuine consequence, all alert carries an education for what to do concerning it, and person owns the response. An alert no one is on call for is records at best, not detection.

Use aureate installations

For all category of timing device, keep a known-good citation build and configuration under type control, and inspect regularly and automatically that what is deployed motionless matches it. The item is to rotate a inquiry akin "is this chassis correctly configured and patched?" from item lone an expert can answer, and lone slowly, into a difference anyone can run in seconds.

Telstra had nothing of the sort. The TAP study established no specified configuration and no document of what the servers should appearance akin another than the servers themselves. The missing firmware update on the Melbourne GPS cardstock had been there for six years, in plain sight. There was merely no automated procedure that would have flagged it to anyone.

Upgrade and measure application continuously

The firmware fix for the rollover behavior existed and the vendor had published bulletins concerning it. Vendor notifications need a defined owner and a tracked way to action, and updates need to be applied on a agenda fairly than whenever item forces the issue. 

Evaluate before deploying, in a lab, against the behavior you really depend on. Do not ignore to verify afterwards. The Telstra changes were completed without anyone checking that the chassis served the accurate date.

Redundancy

Redundancy in timing means, not lone multiple sources, but autonomous ones that cannot converge on a average error.  Two servers fed by the identical GNSS receiver is motionless one source, but counted twice. 

Netnod's own assistance is built on the regulation of multiple autonomous nodes, all alongside autonomous nuclear clocks and redundant servers that are traceable to Swedish National Time realization, UTC(SP).  

Replace equipment continuously

Timing infrastructure normally ages quietly. It keeps working, it rarely complains, and it is hence a natural applicant whenever the prosperity gets trimmed. There volition continually be another components anywhere the consequences of nonaccomplishment are additional apparent and hence gets prioritised. 

Instead, scheme replacement on a rolling sequence and scheme the mark architecture archetypal fairly than accepting what new hardware installations enforce on you. Keeping existing infrastructure fit have to be funded alongside new projects, not be paid for alongside the remaining overs.

Concluding remarks

The way Telstra handled the consequence deserves praise. Commissioning and publishing the autonomous assessment is commendable. All providers of crucial services, including Netnod, are improved off since of this. 

The most unsettling part is perchance that the outage occurred equal although NTP worked fair akin it was intended to do. 

The issue was everything about it. Architectural choices, prosperity cuts, low staffing level, deficiency of appropriate monitoring, ownership or documentation; it all occurred since no one really appreciated fair how critical period services can be. 

Let’s try to alter that, shall we?

Other Article Hacker News
Close Right Ads
Close Left Ads