There is a new kid on the block - on every block - and in fact it is rebuilding the blocks as we speak and populating them with fictitious workers you won’t be able to tell from real workers – and yes, sooner or later it is coming for your job.
Testers, too, have absorbed this narrative: Judgement Day approaches, the AIs will be taking over, not quite like in the Terminator films, but the metal foot will shortly be stepping on our pay packets
It’s late-2026 now, and we still assume that, any day now, this worst-case scenario is coming. The metal foot is about to descend; the boss is about to call you in for a chat. But in fact, it’s much likelier the boss has called you in to ask you how AI is helping you to do your job better. As testers (and their bosses) are starting to notice, there is something very dystopian about the world where code checks code, machine scans machine, everything falsely passes, and LLMs surrender the earth to swarms of bugs in their software.
Why should we suppose that machine code would struggle to look inside itself and find a bug?
Well: how does it even know what a bug is? How can it distinguish one from the functioning software? This is a problem shared with code that has been around on this planet for far longer than the noughts and ones of binary: DNA.
If only we were allowed to report all the fitness-for-purpose bugs in our own DNA, which has been cheerfully replicating for billions of years without anyone to tell it what its purpose is. Could we have better eyesight, better hearing, could we stop craving the very foods that are bad for our health?
Without any human in the loop, we could end up with a world of badly designed (which is to say completely undesigned) machine applications, unaware of their flaws and excruciatingly slow to correct them.
It’s not just a theoretical problem: in the summer of 2025 an AI coding agent at Replit, left to work unsupervised on a live company database, was placed under an explicit freeze – no changes without express permission – and deleted the entire thing, wiping the records of more than a thousand companies and the executives who ran them. It then reported that all was well: it fabricated thousands of imaginary users to fill the hole it had made, produced status messages insisting nothing was wrong, and assured everyone the deletion could not be undone.
In fact, it could all be undone, but it wasn't until a human had trawled through the apparent wreckage that this was known.
Perhaps the ultimate legacy of AI then, will not be the many and various ways in which it accelerates testing and improves accuracy, but the many and various ways in which it confirms what a precious thing manual human checks were in the first place.
We must, after all, keep human hands on the reins. As this ingenious new tech takes to the air and loops the loop, a person should still be steering.
By Simon Royall, Test Analyst at Mando Group