thinkberg.com
TECHNICAL MEMORANDUM                                          TM-26-001

SUBJECT:      Cost of nine minutes of downtime at national scale
DATE:         2026-07-07


SUMMARY. On 9 July 2021 the backend issuing Germany's digital
vaccination certificates was unreachable for nine minutes. Cost:
roughly 30,000 certificates not issued. The redundancy was planned;
the failure modes were not.

1. BACKGROUND. The day before, a power block in one of our servers
failed - not quietly; it exploded. The VM on that machine was not on
the SAN, so there was nothing to fail over to: the first outage. The
surge had also damaged one of the network switches. Its replacement
was scheduled for the next day.

2. FINDINGS.

   2.1 During the switch replacement the next day, its twin did not
   take over - the redundancy that should have carried the traffic.

   2.2 The backend was down for nine minutes.

   2.3 In those nine minutes, roughly 30,000 certificates were not
   issued - people standing in pharmacies and vaccination centers,
   waiting. By that day the system had issued more than 60 million.

   2.4 Travel during the pandemic was, in many cases, only possible
   with a current test or vaccination certificate. Every German
   certificate came through this backend - there was no other path.

3. ASSUMPTIONS. If just 1% of those 30,000 were about to travel
abroad - a plane to board, a border to check - that is 300 people.
At a moderate 400–1,000 EUR per ticket, nine minutes put roughly
120,000–300,000 EUR of other people's travel at risk.

4. CONCLUSIONS. "How much redundancy is enough?" is not an
architecture question. It is arithmetic: what does one minute of
downtime cost, in your units - certificates, orders, patients,
plane tickets? Someone has to do that arithmetic before the
hardware does it for you.

5. ACTION. There was no technical fix on our side to make. An
exploding power block is bad luck. A VM with no SAN copy and a
failover twin that does not engage are not. Both sat in the
infrastructure beneath our software, so what remained was the
supplier relationship: we asked our data-center operator for better
due diligence on the services provided. Some incidents end in an
architecture change. This one ended in a hard conversation - and in
this memo.