Why AI Governance in Banking Starts With Controlled Experimentation
How can a bank use AI to remove routine work without training its people to approve credible sounding output they have not properly examined?
A Copilot presentation can now arrive in the correct corporate design before the author has had time to make a coffee. The colours are right. The logo is in place. The layout looks polished. And yet the presentation can still be weak. That was my experience when I prepared a demonstratiåon of Copilot Agent Builder. I asked Copilot in PowerPoint to create a presentation that explained how the tool works and led users through building an agent. It produced an attractive first draft. But the text was sparse. The argument was generic. Some visuals actively obscured the intended message. The deck looked finished long before it was ready.
Banks do not face an AI adoption problem. They face a judgement problem: people are becoming faster at producing answers than at challenging them.
That small experience captures the next AI challenge for banks. The question is no longer whether employees can generate content, analyse a file or receive a summary. The question is whether they can recognise when an answer that looks complete is incomplete, unsupported or wrong in a way that matters. This is not a technical issue. It is a question of professional judgement.
An agent that knows its limits
Alongside the presentation, I conducted a test run for a contact note agent. The idea was simple. After a client conversation, a relationship manager often has a set of rough bullet points. The agent should turn those notes into a comprehensive, structured contact note. But it should do something more important than write well. It should show what is missing. If the relationship manager noted that a client discussed a potential investment, but did not record the objective, relevant constraints or agreed next steps, the agent should not fill the gaps with plausible language. It should flag the missing elements and ask the user to provide them. Under current internal rules, the agent is also not allowed to generate images or documents.
Those restrictions are not an inconvenience. They are the point. A useful contact note agent is not a substitute for the customer conversation. It is a controlled drafting and quality support mechanism. It turns fragmented notes into a reviewable record, makes omissions visible and leaves the relationship manager accountable for validating what is eventually stored. The process could then continue. Once the relationship manager has reviewed the note, another specialised agent could identify potential tasks and follow ups. One helper structures the conversation record. The next identifies commitments that may need action. The human decides whether the record is correct, whether the task exists and what should happen next.
This is a more realistic model for banking than the fantasy of fully autonomous work. Agents should act as a team of assistants around a professional, not as an invisible replacement for professional responsibility.
The dangerous comfort of “good enough”
The contact note example also exposes the risk. A well written output creates comfort. It has complete sentences, a professional tone and a logical structure. That makes it easy to read. It also makes it easy to accept without asking the most important question: where did this statement come from? The usual answer is “human in the loop.” But that phrase has become too vague to be useful. A person who clicks approve after skimming a polished answer is not exercising meaningful control. They are performing a ritual of accountability. The real risk is not that employees will trust obvious nonsense. Most people will catch that. The greater risk is that a fluent, almost correct output lowers the threshold for scrutiny. It feels good enough.
For contact notes, the review task is relatively concrete. The relationship manager should compare the draft with the original bullet points. Was anything added that was not actually said? Is an important detail missing? Does the final note accurately reflect the conversation?
KYC is a much more difficult case. A KYC reviewer does not just need to check whether a draft corresponds to an input. They need to assess whether an entire customer story is coherent. Does a customer’s declared income align with their accumulated wealth? Do documents, customer statements, transactions and risk indicators support one another? Does a claim sound credible but require corroboration from a trusted source?
That is where a capable agent could become genuinely valuable. It could gather and structure available evidence, identify inconsistencies, point to relevant passages in supporting documents and show what information remains unverified. It could make the reasoning path visible.
But it should not make the final judgement. The reviewer’s role is not to confirm that the agent produced a tidy narrative. It is to decide whether the narrative makes economic and factual sense. The agent can reduce preparation effort. It cannot inherit accountability for the conclusion.
Experimentation is a control
This is why banks should stop treating experimentation and governance as opposites. In a regulated institution, controlled experimentation is part of the control environment. Employees who never use AI cannot develop an informed understanding of its limitations. They may reject useful applications because the technology feels unfamiliar. Or they may accept weak answers because they do not recognise the typical ways in which AI can omit context, infer too much or present uncertainty with unjustified confidence. Both responses create risk. People develop judgement by applying AI to real but bounded tasks. They learn when a prompt lacks crucial context. They learn that a confident result can rest on a weak premise. They learn when refining and validating a task takes longer than doing it themselves. They learn which work benefits from assistance and which work should remain largely untouched.
I use Copilot Researcher for online research on a specific topic. I use Copilot Analyst to examine data, for example in Excel. Both can accelerate early exploration. But neither can understand the real business question without help. To obtain useful results, I need to explain what the data represents, what I want to know and what would count as a meaningful conclusion. That is not a weakness of the tool. It is the division of labour. For the same reason, I rarely ask AI to summarise my e-mails. In many cases, the detail is the context. It allows me to understand why an issue exists, how it developed and what response is appropriate. I also do not delegate my work prioritisation. My inbox is deliberately lean, and my own system already works efficiently.
The point is not to use AI everywhere. It is to know where it improves work and where it removes the very context needed for sound judgement.
Global evidence suggests that many organisations are still caught between experimentation and operational scale. McKinsey’s 2025 survey found widespread use of AI, but most organisations remained in experimentation or pilot phases rather than deploying it broadly across the enterprise. The difficult step is not obtaining access to a model. It is redesigning the routines through which people frame work, check outputs and make decisions [mckinsey]
Regulation points in the same direction. The EU AI Act defines AI literacy as the skills, knowledge and understanding needed to use AI systems in an informed way and interpret their outputs appropriately. Its Article 4 requires providers and deployers to take measures to ensure sufficient AI literacy among staff and other people acting on their behalf. [europa]
For Swiss banks, FINMA’s Guidance 08/2024 makes the practical stakes clear. The authority identifies risks related to models, robustness, explainability, bias, data quality, cyber security and reliance on third parties. It expects institutions to address governance and risk management when using AI. [finma]
A policy that tells employees to be careful is not enough. Carefulness is a capability. And capability is built through practice.
From shallow use to capability
The strongest objection is obvious. Encouraging frequent AI use could create shallow activity. Employees may generate prompts to signal adoption, ask AI to perform trivial tasks or produce outputs that do not improve customer outcomes, operational quality or risk management.
That risk is real. But shallow experimentation is often the starting point, not the end state. People rarely become confident users of a new tool because they first receive the perfect use case. They become confident because they try simple tasks, see what the tool can do, discover where it fails and gradually lose their reluctance to experiment. An employee who has never tested AI on a low risk task is unlikely to recognise either its potential or its limitations when a more consequential use case appears.
A usage target can therefore be useful. It signals that AI is not an optional curiosity reserved for enthusiasts. It creates repetition, familiarity and permission to try. But a number alone is not enough. It can easily turn into performative prompting: activity without learning, or usage without value.
The responsibility of the organisation is to turn initial experimentation into better judgement. That requires practical guidance, short learning sessions, examples drawn from real work and simple methods that help people identify suitable use cases. Most employees do not lack tasks. They lack a way to recognise which tasks are worth giving to AI, what context the tool needs and what degree of review the output requires. The objective should therefore not be to maximise prompts. It should be to move employees from asking, “What could AI do for me?” to asking, “Which part of this task should AI prepare, what must I still verify and where would it create no value at all?” That is how basic use becomes professional capability.
What banks should measure
The wrong metric is the number of prompts, licences or active users. That would merely create performative adoption. The useful measure is whether a real process has become more trustworthy. Can a relationship manager see where a contact note came from and what needs confirmation? Can a KYC reviewer trace a conclusion back to documents, transaction data and credible external sources? Can a team explain why an agent was allowed to make a suggestion, but not to execute an action? Can an employee identify when the output is not good enough?
The future advantage of a bank will not come simply from owning access to the strongest model. Models will become more available. Trust will not.
Trust will belong to institutions that can demonstrate where an output came from, what evidence supports it, what was deliberately excluded, who challenged it and who remains responsible when it is wrong. The most dangerous sentence in banking is therefore not: “The machine made the decision.” - It is: “Looks good to me.”
Consequences for banks
- Build traceability into customer documentation. A contact note agent should separate original input, generated text and missing information. It should make it easy for a relationship manager to identify whether a statement is grounded in the source notes or merely inferred.
- Turn KYC review into evidence review. A KYC agent should not produce a final verdict. It should surface inconsistencies, point to supporting documents, identify unsupported claims and present findings in a way that enables meaningful human validation. That aligns with FINMA’s focus on governance, risks and controls.[finma]
- Treat experimentation as a governed capability. Leadership should set an expectation that teams regularly test relevant, low risk use cases. It should complement that expectation with practical examples, guidance for selecting use cases and a mechanism to capture learning about limitations, risks and controls.
Quellen
- AI Mindset: Rethinking How We Work in 2026, Digital Age, 2026. The original argument that widespread AI use does not automatically redesign work.[digital-age]
- The State of AI in 2025: Agents, innovation, and transformation, McKinsey, 2025. Evidence on widespread AI use and the gap between experimentation and scaled deployment.[mckinsey]
- Regulation (EU) 2024/1689, European Union, 13 June 2024. Definition of AI literacy and Article 4 obligations.[europa]
- FINMA Guidance on governance and risk management when using artificial intelligence, FINMA, 18 December 2024. Swiss supervisory expectations regarding AI risks, governance and controls.[finma]





