On 22 July, at 10.44 in the morning, the hospital systems across the Waikato stopped answering. Waikato, Thames, Te Kūiti, Tokoroa and Taumarunui, five hospitals back on pen and paper for about 11 hours, with 68 appointments and 15 theatre cases deferred at Waikato Hospital alone. The cause was connectivity failure between its network and main data center1, which means the systems themselves were running perfectly well and nobody could reach them, and for some time afterwards Health NZ could say which part had failed without being able to say why.
I live in Hamilton, so this one happened close to home, and it took me straight back to Inland Revenue and the 33 priority one incidents we logged in 2013 on infrastructure IBM had stopped supporting. I’ve written about what the whiteboards did to that number, 33 to 7 to 3, in Two Kinds of Visibility, and about the 2am reboots we scheduled off our own data in Arcane by Necessity, and I promised in the first of those to come back one day to how we sweated those assets. This is that piece, and it turns on a tool I asked for and didn’t get.
The question I put to GKC
In 2014, our second year, we took a Splunk subscription. The New Zealand supplier was GKC, and David Anso and Glen Patrick spent a good deal of time with us, setting up the 45 inch monitoring TV our Director for IT Operations had bought, getting us through the training, teaching us to scrape and analyse the logs coming off the estate. What came out of it was a dashboard on that TV showing the WebSphere servers, the load, and whatever else we could feed it. Good piece of work, and genuinely useful.
The limit was data ingestion. Ingestion cost money, the budget didn’t stretch, so we were capped on how much we could take in at any one time, and when we hit the cap we culled what we were feeding in to protect the trends we cared about most. We were rationing our own visibility, which tells you something about the era we were operating in.
Somewhere during that training I asked the question that has stayed with me for 12 years. Why can’t we use this data to forecast the time to impact? The logs already hold the behaviour, we already know the shape of a server heading for trouble, so why can’t the system tell us what will fail, and when, before it does?
The answer was polite and it was correct. Great idea, they said, and it’s possible, but the technology and the algorithms are limited right now, and you’d have to supply the algorithms yourselves. Not one of us had a major in maths or statistics. Nobody in that team could have written them. So the idea died there, filed under too early, and we went back to doing the job by hand: benchmarks for when each server needed its cache cleared, thresholds watched by people, restarts scheduled for the hours the country was asleep.
Teru, my Senior IT Consultant, carried a spreadsheet that did the analytical half of this. Everything the tool would have done automatically, he did in columns.
Give 2014 the tool it asked for
The what if sets itself up. Drop today’s capability into that room. AI ingesting the logs without a cap, taking over Teru’s spreadsheet, doing the maths none of us could do, producing the sequence of failures instead of a picture of the present.
Splunk ships this now. IT Service Intelligence runs machine learning over historical service data to predict a service’s health score, and the horizon it offers is roughly 30 minutes, given a service with more than 5 good indicators and more than a week of history. The thing I sketched out loud in 2014 is a product feature in 2026. Hold onto that 30 minutes, because it matters more than it looks.
So: if we’d held that tool at Inland Revenue in 2014, would we have kept sweating those assets, and would the $1.5 billion transformation still have gone ahead?
We’d have sweated them longer
My answer surprised me, because I wanted the tool badly at the time and I assumed I’d say it would have made us braver about replacing things. I think the opposite.
Look at where we sat. A government department spends taxpayers’ money, and the obligation running underneath everything, the one my boss quoted whenever things went wrong, was the integrity of the tax system under the Tax Administration Act. Nowhere in that obligation does it say the servers have to be new. If the estate keeps running and the operating cost stays low, the return to the taxpayer improves, and a forecast that tells you precisely when each machine needs attention makes that arithmetic better. Sweating stops being a nervous gamble and becomes a scheduled, costed decision. We were already running a manual, low resolution version of exactly that, and holding 99.9 percent availability through tax peak season while we did it.
Which is worth putting next to what the transformation eventually reported: 7 of 10 measures achieved, digital uptake at 99 percent, and system availability of 99.9 percent2. The same availability number a team of 11 to 12 was already holding by hand, on kit two versions behind, three years before the programme started. Availability was never the argument for spending $1.5 billion. If forecasting had been the deciding variable, the case for replacement would have got weaker.
Better forecasting, longer sweating. That’s the half of this that’s true.
And the world wouldn’t have let us
Here’s the other half, and it took me some time to see it properly.
A forecast tells you when the thing you have will fail. It’s silent on when you have to replace it, because that clock sits outside your system, set by people who aren’t consulting your dashboard.
New Zealanders understand this through houses better than through servers. In Japan a house has an expected lifespan, and at the end of it you demolish and rebuild, so the clock is explicit. Here you maintain instead, and the maintenance isn’t optional, because the exterior paint goes, the piping goes, the wiring ages, the timber can rot, moisture gets in, appliances die, and the building code keeps moving underneath all of it, so what counted as a sound house 20 years ago needs work today to still count as one. The weather, the wear and the standard set that timing, and the house has no say in it.
Our estate had the same problem, and I could see it while we were congratulating ourselves on the reboot schedule. WebSphere was two versions behind and out of vendor support. We had an Enterprise Content Management System that had to be upgraded whether it was healthy or not, because without it the documentation and the correspondence Inland Revenue generates would eventually stop working with everything around it. Demand kept climbing, the agents’ bots hammered the portal harder every peak, and the expectations of what a tax administration should do online kept moving. None of that shows up on a failure curve. A perfect time-to-impact forecast would have told us the servers were fine until Thursday, and nothing at all about the debt we were building out of knowledge that lived only in the heads of the people running it.
That last part is the expensive bit. Sweat something long enough and the replacement bill is mostly the reconstruction of everything the old system knew and nobody wrote down. Transformation programmes are partly that invoice arriving.
The other failure mode
There’s a mirror version of this in the news here, and both belong in the same frame, because I don’t want to be read as saying that spending is the problem.
MBIE’s Biometric Capability Upgrade ran for seven years and was stopped in December 2025 having delivered nothing usable. By September 2026 the ministry had confirmed $50.7 million lost across it and a related project3, and Parliament’s Privileges Committee had found that the ministry deliberately misled a select committee about it, a contempt of Parliament. Ministers in two governments were, on the Minister’s account, managed rather than informed, and that detail is the tell that this isn’t a story about one party or one minister.
The root of it looks structural to me. A programme of work that never fixes its answer to two questions, who is this for and what is it for, will drift for as long as the money lasts. Business process improvement, cost reduction, something the public actually needs: pick one and hold it. If the specification moves every time the market ships something newer, there’s no outcome to measure against, only a target that keeps walking away, and years pass, and everyone involved can say honestly that they worked hard the whole time.
So one organisation deferred replacement for a decade under an absolute obligation, spent $1.5 billion, and delivered most of what it promised. Another spent tens of millions chasing a capability and delivered nothing. Neither outcome was decided by how well anybody could forecast a failure.
What the 30 minutes tells you
Which brings me back to Splunk’s horizon.
30 minutes is an operational number and a good one. It’s enough to move traffic, warn the service desk, get somebody to a keyboard, take a controlled restart instead of wearing an uncontrolled outage. In 2014 it would have made us better operators, saved Teru some late nights, and possibly taken 3 priority ones down to 1.
The same 30 minutes is silent on whether to spend $1.5 billion, on when the vendor stops patching, on when the regulation changes, and on when the last person who understands the batch schedule retires. Nobody reading that dashboard learns anything about the decision that actually mattered.
I asked for the tool in 2014 because I thought better visibility would improve the decision. Twelve years on, I think it improves the operation, and the decision was always being made on a different set of facts, most of which sit outside the building.
What I do with my own
I should admit that I sweat my own assets thoroughly. I’m running a 10th generation i5 laptop well past the 3 years the industry likes to quote, and when I replace it I’ll be looking at a refurbished 13th generation i7 rather than a new machine, because the value is better and the dollar stretches further. There’s a 2012 Mac Mini in the house doing duty as a NAS under Linux, more than 10 years old and not complaining. I keep an old laptop purely because it has a DVD writer and I haven’t digitised the discs yet. Linux on hardware Windows has given up on buys years, and for the work I do the performance difference is nothing I can feel.
I can do all of that because when my laptop dies, the consequence is my own inconvenience. No tax return, no surgery, no statutory obligation. The moment other people depend on it, the calculation changes, and the calculation was never about the hardware.
Everything has a best before date and an expiry date, and the gap between them is where the sweating happens.
I’d still have taken the tool in 2014.
- Waikato hospitals’ outage of 22 July 2026, its duration and the appointments and theatre cases deferred: RNZ, 23 July 2026, RNZ. ↩︎
- Inland Revenue’s final report on its business transformation outcomes: 7 of 10 measures achieved, digital uptake 99%, availability 99.9%: Inland Revenue Annual Report 2023–24, Reseller News. ↩︎
- Correction, 27 September 2026: the paragraph on MBIE has been corrected. The first version reported $33 million spent and a further $6 million surfacing afterwards. The cost is later confirmed by MBIE’s chief executive on 23 September. MBIE’s Biometric Capability Upgrade, the Privileges Committee’s finding of contempt, and the cost confirmed in September 2026: Privileges Committee report, 25 August 2026, MBIE, RNZ. ↩︎
