A new survey doing the rounds in finance circles this month found that one in four executives have had an AI-generated error reach their board or their investors. Not caught in review. Not flagged by a second pair of eyes. Reached the board.
Most people read that stat and think about hallucination — AI making something up. That's not actually the risk I'd worry about. The risk I'd worry about is agreement.
Here's what I mean. Push back on almost anything an AI tool tells you — a number, an assumption, a line of reasoning — and nine times out of ten you'll get some version of the same two words back: "you're right." It doesn't matter if you were right. It doesn't always matter if your challenge made any sense at all. The model's default setting is to move toward you, not to hold its ground and make you prove your case. Ask it to double-check something a second time and it will often find a new reason to agree with whatever you just said, even if that's the opposite of what it told you thirty seconds earlier.
That's a strange thing to build financial judgement on top of.
FP&A has always had one real job underneath all the modelling and the slide-building: deciding which numbers deserve to be trusted before anyone outside the function sees them. A trainee's first pass at a variance analysis gets checked. A new analyst's forecast gets sense-checked against last year, against the pipeline, against what the sales director is saying in the corridor. Nobody signs it off just because it looks tidy and arrived on time.
AI output looks tidy and arrives fast. That's exactly why it's dangerous to wave through. It has the polish of something that's already been checked, when often the only thing that's happened is that nobody has challenged it yet — and the one time someone does, it folds.
CIMA published something alongside that stat worth sitting with: AI can now build a board presentation in minutes, but the actual work — checking the visual story holds up, that the data means what the slide claims it means, that the message survives contact with a sceptical director — still has to be done by a person. The tool got faster. The job of judgement didn't get any smaller. If anything it got more important, because speed used to be a natural checkpoint. A deck that took three days to build got looked at three days' worth of times along the way. A deck that takes twenty minutes doesn't get that benefit unless someone deliberately builds the checking back in.
So here's the practical version of this, for anyone using AI anywhere near numbers that leave the finance function: treat every AI output the way you'd treat a first draft from a new starter. Not because it's usually wrong — often it isn't — but because "usually right" isn't the bar for anything that reaches a board pack or an investor update. Build in a deliberate challenge before sign-off, and don't let "you're right" from the tool count as that challenge. It isn't a second opinion. It's the same opinion, agreeing with you.
The finance teams that get this right over the next few years won't be the ones using AI the most. They'll be the ones who never confused a fast answer with a checked one.
Mark Lynam is a senior finance leader with 20+ years experience in commercial finance and FP&A. He works with businesses that need sharp financial thinking at the leadership table. Get in touch or connect on LinkedIn.