Let’s consider a scenario: a junior banker uses AI to pore over hundreds of documents, create financial models and answer a buyer’s diligence question.
Turns out one of the quarterly EBITDA adjustments was wrong. The error carried downstream. The deal still went through.
Two years later, the acquisition is in dispute. The model has changed, the prompt is gone, and the reviewer has moved on.
Who signs the explanation?
That hypothetical is becoming an operating model. On August 11, Allvue announced an AI agent system that connects a loan transaction from the arrival of a loan notice through data consolidation in the general ledger. Software extracts the borrower, facility terms, and payment amounts. Cash gets reconciled. Approved data moves toward the official books.
Allvue’s product explanation highlights some of the machinery underneath: AI reads unstructured notices, deterministic rules match cash, and people validate exceptions before posting. The intelligence is spread across models, rules, integrations, and necessarily, human judgment.
Once AI output reaches a buyer, a ledger, or cash, a bad answer becomes the firm’s problem.
Firms will need the evidence chain behind the answer.
AI has entered the transaction
To be clear, the “AI banker” still has a human payroll behind it, for now. The title is shorthand for “AI-enabled banker” but the liability remains literal.
And the shift is already visible in M&A. The Wall Street Journal reported that CVC used an AI data portal and chatbot analyst during the sale of Greek ecommerce company Skroutz. Buyers received a portal that functioned like an investment memo. The chatbot answered financial and diligence questions, escalating selected requests to management.
Eilla AI offers another operating model. During the sale of CreateX and Native, its platform handled buyer identification and document-heavy execution while employed advisers directed the process. An offer arrived 15 days after outreach began.
The signature line sits in a different place in each model. What remains consistent is that every answer, exception, and posting still needs an authorized owner.
Everyone is building half a record
AI audit trails are already a product and research category. Collibra describes a governed record spanning inputs, outputs, actions, data access, people, and policies. A January paper proposes tamper-evident LLM audit trails that connect technical provenance with approvals and attestations. Financial institutions clearly understand the need to preserve AI activity for compliance.
The harder problem begins when one financial decision crosses several record systems. Anthropic provides two examples. Its Compliance API exposes organizational activity across chats, files, projects, users, and configuration changes. Future Claude models will also place a statistical watermark inside generated text. User identity, organization, chat history, ownership, and legal responsibility sit outside that watermark.
The problem is fragmentation. Anthropic knows what happened inside Claude. Allvue knows what happened inside the loan workflow. The approval system knows who signed off. The accounting system knows what posted.
But each system captures only one slice of the decision while the firm owns the whole outcome.

Those slices serve three different jobs:
- Artifact provenance: Who or what produced the output?
- System observability: What happened inside the model and workflow?
- Decision accountability: Which evidence, controls, and authority permitted the action?
A watermark covers part of the first. Compliance APIs and traces cover part of the second. The decision ledger connects all three across system boundaries.
The reconstructability test
Let’s take a sample buyer question: “How much of next year’s EBITDA depends on the top three customers?”
An agent retrieves contracts, forecasts, and the operating model, calculates the concentration, drafts a response, and routes it for release. The saved answer shows what went out. But it rarely reconstructs a missing contract, an overwritten spreadsheet, a tool failure, the reviewer’s edits, or the approval rule.
A file named FINAL_v12_reallyfinal.pdf is still just a file.
The institution needs to reconstruct five links:
Evidence —> Computation —> Judgment —> Authority —> Action
What did the system know? What did it calculate? Which checks and human judgments changed the answer? Who had authority to release it? What happened after release?

Ordinary logs record events inside one system. A decision ledger assembles the causal record across systems and ties the final action back to evidence and authority. When something fails, the firm can locate the break across source data, retrieval, calculation, model interpretation, or human review. Law and contract decide liability. The ledger gives every party a fighting chance of proving what happened.
The existing rulebook follows the activity
Regulators are already observing this transition. In July, 21 organizations gained access to Claude inside the FCA’s second Supercharged Sandbox cohort for use cases including agent-led payments, AI governance, and compliance automation.
FINRA’s Regulatory Notice 24-09 says its technology-neutral rules continue to apply when member firms use generative AI, including proprietary systems, third-party tools, and AI embedded inside existing products. Firms remain responsible for evaluating the system and supervising its use.
Recordkeeping history supplies the precedent by analogy. In August 2024, the SEC charged 26 firms over failures to preserve required electronic communications, imposing $392.75 million in combined penalties. Regulated work had moved into channels the firms could not adequately capture or supervise. AI expands that reconstruction problem across retrieved documents, models, tools, and agent vs. human review.
Current agents still need an evidence layer
BankerToolBench, built with input from 502 investment bankers, asks agents to navigate data rooms and produce linked financial deliverables. Its strongest tested model failed nearly half the criteria. Bankers rated zero percent of its outputs as client-ready. Zero.
The result captures the tension: agents are capable enough to enter professional workflows and unreliable enough to require evidence around every material release.
A July paper titled “Benchmarks Are Not Validation” argues that production approval requires system-level evidence. In production, an eval becomes part of the firm’s record that its standard was applied.
Human approval needs substance
Financial firms already have model-risk programs, vendor contracts, and legal holds. The decision ledger carries those controls across systems that retrieve, route, and act.

Human review remains central, provided the reviewer can actually review. A person with thirty seconds, a green approval button, and no access to the sources supplies ceremony.
Moreover, while AI multiplies deliverable generation, human review capacity still remains fixed. Without structural guardrails, review may inevitably decay into rubber-stamping.
Vendor diligence should ask whether the firm can export sources, agent traces, evals, model history, and approvals during a contract’s lifecycle. The regulator, client, or investment committee will still look to the institution for the answer.
A signature as infrastructure
Financial AI agents can answer a buyer, extract a loan notice, build a model, and prepare a journal entry. Every step can split the evidence across a provider, the institution, and the person who pressed approve.
Signing off is not the new part. Maker-checker and approval matrices have run finance for decades, and they worked because the maker's work papers were generally retained in systems the firm controlled. The checker could pull the model, the memo, or the email chain.
However, the maker is now a powerfully distributed system, and its work papers are scattered across a model provider, a data/workflow vendor, the web, and the firm’s own tools. The signature still transfers full accountability to the institution. The evidence behind it does not automatically arrive with it.
Finance already has systems of record for transactions, communications, and accounting. Agentic AI needs one for how evidence became an authorized action.
Sources:
- Allvue: Intelligent Loan Operations announcement
- Allvue: The Future of Investment Accounting Operations
- WSJ: AI Replaced Bankers on a CVC Sale Process
- Raconteur: Eilla AI Completes Europe’s First AI-Powered M&A Deal
- Eilla: CreateX and Native Case Study
- FINRA Regulatory Notice 24-09
- SEC: 2024 Electronic Communications Recordkeeping Settlements
- BankerToolBench
- Benchmarks Are Not Validation
- Anthropic: How Claude’s Text Watermark Works
- Anthropic: Claude Platform Compliance API
- FCA: Anthropic to Support the Supercharged Sandbox
- Collibra: AI Audit Trails
- Audit Trails for Accountability in Large Language Models
