CREATE
  • Technology
    • BIOTECH
    • COMMUNICATIONS
    • COMPUTING
    • IMAGING
    • MATERIALS
    • ROBOTICS
    • SOFTWARE
  • Industry
    • DEFENCE
    • INFRASTRUCTURE
    • INNOVATION
    • MANUFACTURING
    • POLICY
    • PROJECTS
    • TRANSPORT
  • Sustainability
    • ENERGY
    • ENVIRONMENT
    • RESOURCES
  • Community
    • CULTURE
    • PEOPLE
  • Career
    • EDUCATION
    • INSPIRATION
    • LEADERSHIP
    • TRENDS
  • About
    • CONTACT
    • SUBSCRIBE
No Result
View All Result
CREATE
  • Technology
    • BIOTECH
    • COMMUNICATIONS
    • COMPUTING
    • IMAGING
    • MATERIALS
    • ROBOTICS
    • SOFTWARE
  • Industry
    • DEFENCE
    • INFRASTRUCTURE
    • INNOVATION
    • MANUFACTURING
    • POLICY
    • PROJECTS
    • TRANSPORT
  • Sustainability
    • ENERGY
    • ENVIRONMENT
    • RESOURCES
  • Community
    • CULTURE
    • PEOPLE
  • Career
    • EDUCATION
    • INSPIRATION
    • LEADERSHIP
    • TRENDS
  • About
    • CONTACT
    • SUBSCRIBE
No Result
View All Result
CREATE
No Result
View All Result
Home Sustainability Energy

How do engineers respond when fail-safes trigger total failure?

Elle Hardy by Elle Hardy
16 April 2026
in Energy, Features
Reading Time: 5 mins read
0
How do engineers respond when fail-safes trigger total failure?

Image: Getty

Engineers are increasingly being asked not just to solve discrete problems, but to manage and design for hyper-complex, interconnected systems.

In November 2023, a routine software upgrade brought Australia’s second-largest telecommunications network to its knees.

Ten months later in July 2024, a faulty security update crashed 8.5 million Windows computers worldwide.

And in April 2025, the entire Iberian Peninsula went dark in just 90 seconds.

Three different continents, three different sectors, one common thread: single points of failure that cascaded through integrated networks, bringing critical infrastructure to a standstill.

For Australian engineers designing and managing complex systems, these incidents offer vital lessons about a paradox at the heart of modern infrastructure. On the one hand, we are reaping the benefits of networks becoming more technologically, socially and economically interconnected.

Yet this same interconnection can transform localised errors into catastrophic failures that can bring entire countries to their knees. A test of engineering resilience, the challenge isn’t to eliminate integration, but ensure our hyper-complex systems are designed to contain failures and prevent contagion.

Optus outage

The Optus outage exemplifies how protective systems mixed with human error can worsen outcomes. A software upgrade at a Singapore exchange triggered an avalanche of network routing changes, including more than 940,000 Border Gateway Protocol announcements in an hour, compared to a normal rate of fewer than 3000.

The resulting flood of updates shut down 90 edge routers nationwide.

The Optus outage lasted more than
0 hours
The number of customers affected was more than
0 million

The routers did exactly what they were designed to do: disconnect themselves to avoid overload. But this ‘fail-safe’ response created a worse outcome – total network failure. With the IP Core network down, restoration required physical reconnection or manual rebooting of routers at sites across the country.

The outage lasted more than 12 hours and affected more than 10 million customers. Hospitals lost communications, banks couldn’t process transactions, Melbourne’s entire train network came to a halt, and even emergency services calls were affected.

When protective systems all responded identically to the same trigger, there was no fallback. According to Mark A. Gregory FIEAust, from RMIT’s School of Engineering, “an entire telecommunications network going offline is unusual. The network should be designed in such a way that redundancy (backups) and resiliency are built in from the outset.”

Mark A. Gregory FIEAust

In his submission to the Senate inquiry on the incident, Gregory added that there should be “‘more transparency and improved reporting to the regulator … on network design, management practices, redundancy and resiliency”. The inquiry agreed, recommending that telecommunications companies be bound by stricter security requirements, as with other critical infrastructure providers.

As for engineering lessons, mapping and segmenting, with diversity of technology and locations to avoid a single point of failure, are critical. Equally, having the right tools in place to detect catastrophic events and better investment in engineering are key components of avoiding a future event of this magnitude.

CrowdStrike incident

The July 2024 CrowdStrike incident exposed the significant vulnerabilities that arise when software operates outside established boundaries and deep within system architecture.

A faulty update to CrowdStrike’s Falcon Sensor security software – embedded at the Windows kernel level, below user-mode processes – triggered a catastrophic failure. The sensor functions by presenting itself as a device driver, and while its execution is subject to Microsoft’s WHQL certification, this process is lengthy.

To avoid recertification for every release, CrowdStrike evolved its update model to adjust configuration parameters rather than modify code, enabling rapid, “agile” updates, but shifting greater risk onto the operational environment.

When an untested configuration-parameter update was distributed, it triggered immediate system crashes worldwide. The malformed kernel-level parametric update forced affected systems into continuous reboot-crash cycles, preventing the operating system from loading into user mode. Although CrowdStrike withdrew the update within 78 minutes, the scale of impact was unprecedented: an estimated 8.5 million computers became unusable, each presenting the well-known “blue screen of death”.

“CrowdStrike will stand out as a turning point because it touches on three fundamental pillars of systems engineering.”
Jawahar Bhalla FIEAust CPEng

Restoring functionality required manual deletion of corrupted configuration files on individual machines. The cascading consequences were profound – grounded aircraft, postponed hospital procedures, degraded emergency services and widespread operational disruption across multiple sectors.

Jawahar Bhalla FIEAust CPEng, immediate past-president of the Systems Engineering Society of Australia (SESA), noted that while incidents such as Optus’s outage were more an emergent protective behaviour in a complex system of systems, the CrowdStrike failure was more fundamental, compromising the foundational reference architecture of the operating system. “CrowdStrike will stand out as a turning point because it touches on three fundamental pillars of systems engineering: architectural frameworks, transformative engineering methodologies and governance,” he said.

Jawahar Bhalla FIEAust

According to Bhalla, granting kernel-level access introduced a system-level hazard because it effectively “broke the architecture”. The incident illustrated that as digital infrastructure becomes more interconnected and dynamic, reference architectures must evolve, and engineering boundaries must extend beyond the system of interest to encompass its enabling systems. He argues that CrowdStrike may ultimately be viewed as a watershed moment in our understanding of complex systems – but only if its lessons are taken seriously. 

To that end, Bhalla identifies four pillars essential for resilient systems engineering: monitoring environmental changes that affect our configuration items; maintaining evolutionary integrity between conceptual design and fielded systems; establishing digital twin environments that faithfully mirror reality to enable synthetic testing and progressive deployment; and ensuring the availability of enabling resources.

“Leading through a crisis is a true test of a leader’s mettle – just as risk mitigation and crisis avoidance create a lasting legacy.”
Kathryn Guarini

“The CrowdStrike failure violated most of these principles – the use of the application as an ‘architectural extension’ within Windows and itself, the lack of holistic testing, the deployment of a breaking change directly to production, and the absence of enabling capabilities to roll back the change once released.”

Kathryn Guarini, a Senior Fellow of Electrical and Computer Engineering at Yale, said the Crowdstrike failure also highlighted the need to improve design robustness through chaos engineering, “where engineers create intentional failures in order to understand their impact, solve problems proactively and avoid large-scale service disruptions”.

Kathryn Guarini

In a crisis, “leadership matters,” she added. “Leading through a crisis is a true test of a leader’s mettle – just as risk mitigation and crisis avoidance create a lasting legacy.”

Learning lessons

Australia’s size and geographic isolation create unique challenges. Our infrastructure often lacks international redundancy options, making domestic resilience even more critical. Therefore, the goal isn’t just bouncing back from failures, but engineering systems that learn and improve from disruptions. 

As our infrastructure becomes more complex and interconnected, the question isn’t whether failures will occur, but whether we’ve designed systems that fail safely.

The examples of Optus and CrowdStrike demonstrate what happens when we don’t.

A version of this story was originally published in the February 2026 edition of create magazine with the headline “Connection failures”.

The risk of a complete system shutdown is greater than ever. This webinar explores the causes and impacts of total grid failure.

Tags: systems engineeringtelecommunicationspower network
Previous Post

Australia has a maths problem that training more engineers can help fix

Next Post

By the numbers: The latest Australian construction activity data

Elle Hardy

Elle Hardy

Elle is a freelance journalist. She has written for industry publications including the Australian Water Association's Current magazine, Mercer Magazine and BPay Banter.

Related Posts

How do you engineer empathy?
Features

How do you engineer empathy?

10 September 2026
Australian construction quarterly update, by the numbers
Features

Australian construction quarterly update, by the numbers

10 September 2026
Inside Australia’s new AI framework: what new regulations mean for engineers
AI framework

Inside Australia’s new AI framework: what new regulations mean for engineers

10 September 2026
Next Post
By the numbers: The latest Australian construction activity data

By the numbers: The latest Australian construction activity data

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

create is brought to you by Engineers Australia, Australia's national body for engineers and the voice of more than 120,000 members. Backing today's problem-solvers so they can shape a better tomorrow.
  • ABOUT US
  • CONTACT US
  • SITEMAP
  • PRIVACY POLICY
  • TERMS
  • SUBSCRIBE

© 2024 Engineers Australia

No Result
View All Result
  • Technology
    • BIOTECH
    • COMMUNICATIONS
    • COMPUTING
    • IMAGING
    • MATERIALS
    • ROBOTICS
    • SOFTWARE
  • Industry
    • DEFENCE
    • INFRASTRUCTURE
    • INNOVATION
    • MANUFACTURING
    • POLICY
    • PROJECTS
    • TRANSPORT
  • Sustainability
    • ENERGY
    • ENVIRONMENT
    • RESOURCES
  • Community
    • CULTURE
    • PEOPLE
  • Career
    • EDUCATION
    • INSPIRATION
    • LEADERSHIP
    • TRENDS
  • About
    • CONTACT
    • SUBSCRIBE