A new benchmark from BehindLogin reveals that most banking chatbots are failing to deliver meaningful customer outcomes, offering conversational interfaces without the underlying capability to resolve real user needs.
The BehindLogin Chatbot Experience Benchmark 2026, based on observed customer behaviour inside live banking apps, shows that while chatbots can handle basic queries, most struggle with natural language, fail to use customer data, and rarely support end-to-end task completion.
All Talk, No Action
At a basic level, chatbots are designed to reduce friction at moments where users need fast, clear support. That requires two things: accurate interpretation of intent and the ability to act on it. The benchmark shows consistent gaps across both.
Only 7 out of 20 providers can reliably interpret natural language queries. For the majority, performance is dependent on structured inputs, with a large proportion of conversations breaking down.
Even where intent is understood, execution remains limited, with 9 out of 20 providers surface customer data in a meaningful way. Beyond this, 6 out of 20 providers can complete account-based tasks within the chatbot.
In most cases, chatbots guide users rather than resolve issues, redirecting them back into app journeys or support flows they were trying to avoid. This creates a disconnect between interface and capability. The chatbot presents itself as a conversational layer, but lacks the depth to deliver outcomes.
“Banks have the advantage with customer data and execution, but aren’t using it. As LLMs raise expectations, that gap is becoming harder to ignore,” states BehindLogin CEO, Ollie Lane.
Awards
According to the report, a small group of providers demonstrate stronger execution across core areas, particularly in combining clearer interaction design with more reliable task guidance.
Best Overall: Revolut
Runner-Up: Bunq
Personalisation: Bank of America
Intelligent Action: Starling Bank
Conversational UX: Monzo
Leaderboard
Revolut leads the benchmark with a more consistent experience across intent recognition, structured responses and task progression.
Wider scores show a market in transition:
- Language and intent understanding remains inconsistent
- Data is available but under-leveraged
- Action is fragmented or offloaded to other channels
Experiences may feel responsive at surface level but lack depth under real usage conditions.

The BehindLogin Chatbot Experience Benchmark 2026 makes clear that the gap in banking chatbot performance is not simply a technology issue, but an execution issue exposed through real customer use.
Recommended
- NCSC Warns of Severe AI-enabled Chatbot Risks
- 8 Out of 10 AI Chatbots Can Assist in Planning Violent Attacks
- Comment: AI and chatbots | Key issues and risks for employers
By observing verified banking customers as they used chatbots inside their own apps, the benchmark moves beyond demos, marketing claims, and controlled test environments. This methodology matters because it shows how these tools behave in the moments that count: when customers ask natural questions, expect relevant use of their account data, and need an issue resolved without being pushed into another journey.
The findings suggest that many providers have built conversational interfaces without matching them with the operational capability required to deliver meaningful outcomes. While some banks and fintechs are beginning to show stronger performance, particularly around intent recognition, personalisation, and task completion, the wider market remains inconsistent.





