Chatbots fail 57% of personal finance answers, UK test finds
A UK fintech tested 18 AI models and more than 10,000 responses, finding mainstream chatbots failed 57% of personal finance answers and 88% on multi-step queries.
Saturn, a UK financial technology firm, ran a large-scale test and found mainstream AI chatbots failed 57% of personal finance answers in a dataset of more than 10,000 responses. Failure rates rose to 88% on the most complex, multi-step questions.
The company submitted 121 personal finance questions to 18 free and paid AI models, repeating each query up to five times to capture variability. Models tested included versions of ChatGPT, Gemini, Claude and Copilot. Saturn counted an answer as a failure if it contained a factual error, left out a material point or omitted a required warning.
Across all models the failure rate was 57%. On the subset of harder, multi-step items the failure rate climbed to 88%. Free model outputs failed 63% of the time, compared with a 49% failure rate for paid versions. On the most difficult questions free models failed 93% of the time. Claude Opus 5 in a reasoning mode recorded the best performance but still failed 39% of answers.
Errors recorded in the study included miscalculations, missed changes to tax rules and invented regulations. Saturn gave an example of a pension tax response that, if followed, could have exposed a saver to an extra £17,500 in charges from HM Revenue and Customs.
Use of AI for financial decisions is growing. An EY survey of 18,000 consumers found 49% had used AI to support savings and investment decisions. The UK financial regulator reported in August that four in five less experienced investors had used AI for investing help. The EY research also found 56% of respondents trust AI tools, while 44% incorrectly believe AI-generated financial information is regulated.
A separate PensionBee poll of 1,000 US adults found nearly six in ten would act on financial guidance from a chatbot without independently checking it. About one in four respondents said a chatbot had already provided wrong information about their finances.
Amal Jolly, Saturn’s chief executive, warned millions rely on AI for money advice and urged the Financial Conduct Authority to move quickly because AI financial advice is unregulated and consumers lack the compensation rights they would have with a human adviser.
The report noted paid models generally produced fewer errors but still produced risky outputs. Testers recommended users exercise caution when relying on chatbots for taxes, pensions and investments, and called on regulators to set rules that clarify responsibilities and consumer protections when AI is used for money guidance.
The material on GNcrypto is intended solely for informational use and must not be regarded as financial advice. We make every effort to keep the content accurate and current, but we cannot warrant its precision, completeness, or reliability. GNcrypto does not take responsibility for any mistakes, omissions, or financial losses resulting from reliance on this information. Any actions you take based on this content are done at your own risk. Always conduct independent research and seek guidance from a qualified specialist. For further details, please review our Terms, Privacy Policy and Disclaimers.






