Iter Advisors

LLMs in finance: compare ChatGPT, Claude and Gemini

Cases and sources reviewedBy ·

The right LLM for a finance team passes a reproducible business test in an authorised environment. Compare answer quality, traceability of figures and correction time using the same documents and instructions.

Dashboards and financial data analysis

English version published on 2 October 2026. Sources rechecked on the same date.

Which criteria should you compare across assistants?

ChatGPT, Claude and Gemini differ in more than writing style. Available functions depend on the plan, administrator settings and connected tools. This framework is a selection method, not a benchmark carried out by Iter.

Framework to complete for each tested environment
CriterionTestEvidence to retain
CalculationRecalculate a table containing credit notes, duplicates and a zero budgetFormulas or code and results reconciled to the answer key
DocumentsFind five clauses in a document setFile name, page and exact passage
ContextExplain a variance whose cause is not suppliedDistinction between fact and hypothesis
IntegrationRead only authorised foldersVerification of permissions and possible actions
CostMeasure test and correction timeLicences, consumption and review time
Three environments: documented functions and test questions
EnvironmentDocumented starting pointQuestion to verify
ChatGPTFile analysis and code-based calculationsCan you verify operations, exclusions and output files?
Claude inside a business applicationDocument analysis in the Hebbia caseDoes the chosen application preserve references to supporting documents?
Gemini in Data StudioAssistance with questions, calculated fields and presentationsAre these features available in your plan and data context?

These examples involve different products and integrations. They do not rank model quality. Compare the configurations actually available to your team.

ChatGPT documents file analysis using calculation tools. Ask every other candidate environment to demonstrate the same task on authorised company data, without assuming equivalent quality.

How do you run a useful comparative test?

  1. Prepare an authorised file set: variance table, closing procedure and supporting documents.
  2. Write an answer key with expected figures, sources and unanswered questions.
  3. Give every tool the same instructions and retain its complete answers.
  4. Record errors, fabricated references, omissions and correction minutes.
  5. Repeat sensitive cases: one correct answer does not demonstrate reliability.

An accurate result without a verifiable reference is insufficient for an audit note. Reject an environment if the test requires excessive access rights. The ChatGPT guide exercises offer a reproducible starting point.

What does the Hebbia case show about financial documents?

Anthropic describes Hebbia’s use of Claude to analyse financial and legal documents. Claude runs inside a platform that organises documents and tasks; the application goes beyond an isolated conversation.

For an SME, the lesson to test is documentary traceability. Request a supporting file and passage for every statement. The publisher’s cache-related time reduction must not be interpreted as a finance team’s productivity gain. Read the cases and their scope.

Which decisions must remain subject to human approval?

Budget assumptions, provisions, payment approvals and investor communications commit the company. An assistant can prepare a proposal; retain a named owner and a control trail.

For financial due diligence (FR), use the model to identify missing documents or prepare questions. A model failing to flag a risk does not prove that none exists. The final review also covers what was not supplied.

What should you check before connecting financial data?

Document who can read what, which actions are possible, where data is retained and how access can be revoked. Read OpenAI’s business-data commitments and Anthropic’s training policy for the actual plan selected. Do not transfer a condition from one product to another.

For the pilot, prefer read-only access and a restricted folder. A Fractional CFO can define the business need with IT and data owners. The roadmap connects selection to deployment stages.

Sébastien Doat

Your finance contact

Sébastien Doat

Founding partner and Fractional CFO

Frequently asked questions

Is there a best LLM for every finance department?

No. Selection depends on the file set, available tools, access, expected quality and verification cost. Compare environments on identical tasks.

Can an LLM perform calculations?

An assistant can use code, a spreadsheet or other calculation tools. Verify executed operations and assumptions; a number generated in prose is not evidence of calculation.

How do you measure financial-summary quality?

Compare against an answer key: accurate figures, retrievable sources, no invented causes, material points covered and correction time. Also test information missing from the file set.

Scope your project with a CFO

Share your needs, tools and difficulties. We can clarify the scope, deliverables and controls for support.

Discuss my project

Explore our Fractional CFO service

Sources & references

  1. Data analysis with ChatGPT — OpenAI Help Center, source rechecked on 2 October 2026.
  2. Hebbia: financial-document analysis with Claude — Anthropic / Claude, source rechecked on 2 October 2026.
  3. Data Studio Pro features (formerly Looker Studio) — Google Cloud, source rechecked on 2 October 2026.
  4. Business data protection — OpenAI, source rechecked on 2 October 2026.
  5. Data use for model training — Anthropic Privacy Center, source rechecked on 2 October 2026.