Talk to us

The real test of an AI advice tool isn’t the demo

By
Team We Complement

AI is moving into financial advice at some pace. Research reported by FT Adviser in July found that the proportion of UK financial advisers using AI had risen from 43% to 74% in just twelve months, based on an Intelliflo survey of 209 advice professionals. That is a fairly significant shift in a relatively short period of time.

It probably does not come as much of a surprise. There are more tools available, they are getting better and firms are naturally interested in where they can make processes easier, remove repetitive work or help people get to the information they need more quickly.

Most providers will give you some kind of demonstration before you buy, and increasingly there is the option of a free trial as well. Both are useful. Seeing a system in action is important, and having a few weeks to put it through its paces is even better.

The problem comes if the test stops at whether the demo looked good or whether the first few outputs were impressive.

The more useful question is: what happens when you put your own work through it?

 

Real files are rarely as tidy as a demo

Anyone involved in technical planning will know that a live advice case does not always arrive as one neat bundle of perfectly consistent information.

A meeting note might refer to one level of income while a fact-find shows another. There may be two versions of an illustration in the file and it is not immediately obvious which one informed the final recommendation. A client’s objective may have changed during the advice process without every document being updated to reflect it. Sometimes the information is all there, but it is spread across provider correspondence, meeting notes, research and emails rather than sitting neatly in one place.

None of those things necessarily mean there is a problem with the advice. They are simply examples of the reality of working with information that has been gathered from different places, by different people and at different points in time.

There is also another type of case our technical planners regularly have to deal with: one where there is no obvious black-and-white answer. Two options may both be perfectly reasonable, but one works better because of something specific about that client. The important bit then is not simply whether the system can find the facts. It is whether the evidence and reasoning are clear enough for a professional to make, and later explain, that judgement.

That is where an AI tool starts to get a much more meaningful test.

 

Use the awkward cases

If you have a month to trial a system, it makes sense to use that month properly.

There will always be a temptation to start with straightforward cases because they are easier to compare and you want to understand how the system works. There is nothing wrong with that. But once you are comfortable with it, the awkward cases are probably going to tell you much more.

Put through something where documents conflict. Try a case with a piece of missing evidence. Use one where an older version of a document sits alongside the current one. Include a case where you already know there was a genuine judgement call to make.

Then look at what actually happens.

Did the tool notice the inconsistency? Did it rely on the right document? Did it challenge something that was already adequately evidenced? If it raised a concern, could the person reviewing the case see why? And just as importantly, if there was nothing wrong, did it leave the case alone rather than manufacturing a problem because it felt obliged to find one?

That is very different from asking whether an overall output looked good.

Tony explores this in the latest paper in our Advice, Evidence & Judgement series. Rather than beginning with a headline accuracy percentage, the paper suggests starting with a varied sample of cases and looking carefully at the failures: material gaps that were missed, unnecessary challenges, whether findings were grounded in the correct evidence and whether a reviewer could see what happened when professional judgement came into play.

 

A polished answer can still be wrong

This is one of the things that makes generative AI particularly interesting in advice.

The output can be extremely convincing.

A system can take a large amount of information and turn it into something clear, well structured and professional very quickly. That can be enormously useful. But it can also make a weak answer look much stronger than it actually is.

A beautifully written paragraph does not tell you whether the system has picked up the correct version of the illustration. A confident explanation does not tell you whether it has missed a contradictory meeting note. A sensible-sounding challenge does not tell you whether the evidence it says is missing was actually sitting elsewhere in the file.

Our technical planners would not accept something as correct simply because it was well written, and firms should probably apply the same thinking when evaluating technology.

The question becomes less about how good the output looks and more about whether you can follow it back to the evidence that supports it.

 

Testing AI is not about being cautious for the sake of it

There can sometimes be a sense that conversations about testing and controls are somehow anti-innovation. I do not think they are.

The FCA itself is actively supporting firms experimenting with AI through its AI Lab. Its AI Live Testing programme provides firms with a place to test AI systems in real-world conditions, and participants are expected to have considered both pre-deployment testing and how systems will be monitored after deployment. The FCA describes its wider aim as enabling the safe and responsible use of AI while supporting growth and innovation.

That feels like a sensible distinction.

The question is not whether AI should be used. It already is, and the numbers suggest adoption is moving quickly.

The question is how firms become confident that a tool is doing what they think it is doing once it becomes part of a real advice process.

A free trial gives firms a really useful opportunity to start answering that question, but only if the testing reflects the work the system will eventually be expected to deal with.

 

What would you want to know before relying on it?

Tony’s paper includes ten questions firms can use when looking at an AI tool, covering things such as what cases an accuracy claim was measured against, how disagreements were handled, whether findings can be traced back to the underlying evidence and whether firms can test representative cases of their own.

For me, though, the principle underneath all of them is quite simple.

The point of testing an AI tool on your own files is to get a much clearer picture of how it actually behaves in your advice process, not just how well it performs in a demonstration.

As AI becomes a bigger part of how advice is prepared, researched, reviewed and evidenced, being able to understand that difference is going to matter.

Read Tony’s full paper: Advice, Evidence & Judgement | We Complement

External reading: FCA AI Lab

Source referenced in the introduction: FT Adviser article on AI adoption among UK advisers

ISO/IEC 27001:2022 certified
UKAS-accredited information security management system
You can verify the validity of our ISO certificate via the UKAS register.

ISO/IEC 27001:2022 certified

Affiliate of

Consumer Duty Alliance

Proud to work with

Paradigm ValidPath

Contact

Old Brewery Business Centre
Castle Eden
Co. Durham
TS27 4SU

Tel: +44 (0)1472 728 030
Email: hello@wecomplement.co.uk

© 2026 We Complement | Privacy Policy
We Complement Limited registered in England & Wales under company number 13689379, ICO number ZB427271. Registered address: Old Brewery Business Centre, Castle Eden, Co. Durham, TS27 4SU.