Topic 4.2 Notes – Fault Tolerance
1. What Fault Tolerance Is
A system is fault-tolerant if it can continue functioning even when one or more components fail.
That matters because in complex systems, failure is normal. Routers crash. Power goes out. Cables get damaged. Servers overload. Sometimes multiple things fail at once, like during a natural disaster.
If users expect the Internet to work 24/7, the system has to survive those failures.
Common causes of failure:
- Hardware malfunctions (broken routers, damaged cables)
- Power outages
- Natural disasters
- Cyberattacks
- Overloaded systems
Fault tolerance doesn’t mean “nothing ever breaks.” It means the system keeps operating anyway.
On a quiz, you might be asked to:
- Describe a benefit of fault tolerance
- Explain how a system remains functional after failure
- Identify where a system is still vulnerable
Keep the definition tight in your head:
Fault tolerance = continues to function despite failure.
2. Redundancy and Multiple Paths
The main reason the Internet is fault-tolerant is redundancy.
Redundancy means including extra components that can be used if others fail.
That could mean:
- Extra routers
- Extra cables
- Extra servers
- Multiple connections between locations
Redundancy uses more resources. That’s the tradeoff. More equipment, more cost, more maintenance. But it increases reliability.
Multiple Paths Between Devices
One key feature of the Internet is that there are usually multiple possible paths between two devices.
Remember from Topic 4.1:
- Data is broken into packets
- Each packet can travel independently
Here’s what that looks like conceptually. Notice how different packets from the same source can travel along different routes through the network before reaching their destination.

Packets traveling along multiple network paths
If:
- A router goes down
- A cable is cut
- A connection becomes unavailable
Then packets can be sent along a different route, if one exists.
This avoids a single point of failure, which is one component whose failure would shut down the entire system.
Important exam idea:
More possible routes between two points = greater reliability.
3. How Fault Tolerance Works on the Internet
The Internet uses abstractions for routing and transmitting data. Devices don’t need to know the full path from sender to receiver. Routing systems handle that.
When something fails, this is what happens:
- A device or connection becomes unavailable.
- The network detects the failure.
- Routing protocols update available paths.
- New packets are sent along a different route.
- The receiving device reassembles packets in order.
Because packets are independent:
- Some may take different paths.
- They may arrive out of order.
- The system reorders them correctly.
Communication might slow down, but it doesn’t completely stop.
That design is why localized damage doesn’t usually “break the Internet.”
If you see a question describing packets being rerouted after a router fails, they’re testing your understanding of dynamic routing + redundancy.
4. Benefits of Fault Tolerance
Increased Reliability
- The system keeps operating even when parts fail.
- Reduces total shutdowns.
- Users experience fewer disruptions.
Reliability increases as redundancy increases.
Reduced Impact of Cyberattacks
Consider a DDoS attack, where a server is flooded with traffic.
If there are:
- Multiple servers
- Multiple network connections
Traffic can be redirected. The system may slow down, but it won’t necessarily collapse.
Fault tolerance doesn’t stop the attack. It limits the damage.
Improved Scalability
Scalability means a system can handle growth.
Redundant routing options:
- Allow more packets to move simultaneously
- Support more connected devices
- Increase total network capacity
More routes between two points makes the system both more reliable and better able to handle growth.
5. Tradeoffs and Vulnerabilities
Cost of Redundancy
Redundancy requires:
- Extra infrastructure
- More hardware
- Ongoing maintenance
This increases cost and complexity.
A common AP-style question gives you two network designs and asks which is more reliable and why. The more fault-tolerant one will almost always have more alternative paths.
Remaining Vulnerabilities
Even fault-tolerant systems can fail if:
- Too many components fail at once
- There aren’t enough alternative paths
- A critical component is not redundant
- The system is overwhelmed beyond capacity
Be ready to spot:
- A single point of failure
- Where adding redundancy would improve reliability
- Why a system might still be vulnerable
No system is perfectly fault-tolerant.