Claude vs. Copilot: Live Financial Model Build-Off — Who Wins?

events upcoming events May 08, 2026

 The idea was straightforward. Take two of the most capable AI systems currently in use, give them the same financial modeling task, and observe how each would approach it in real time.

There was no template to complete and no guided structure to follow. Each system had to interpret the business scenario, define its assumptions, and build a functioning financial model from scratch.

That was the premise behind a live session hosted by David Brown, where Claude and ChatGPT were put head-to-head under real conditions. From the outset, the expectation was clear. This was not about whether AI could produce output. It was about whether that output could stand up to the same standard applied to human financial modelers.

As David framed it at the beginning, both systems would be required to build a full three-statement model and effectively demonstrate whether they could pass the level of quality expected in professional financial modeling assessments.

Alongside him, Gemini was introduced as a co-judge, setting the tone for how the models would be assessed. The standard was deliberately high:

“Today we are putting two of the world’s most powerful models head-to-head in a brutal, no-holds-barred test of financial logic, accounting integrity, and Excel craftsmanship. Building a model is easy to fake on the surface, but the real test begins when you stress the logic behind it.”

With that, the exercise moved from setup into execution.

 

Two Systems, Two Approaches

Once the timer started, the difference between both systems became obvious almost immediately.

Claude paused. It outlined its thinking, broke the work into steps, defined its structure, and then began building with that structure in place. Every stage was visible. You could follow how it intended to move from assumptions to schedules to outputs.

ChatGPT approached the task differently. It engaged quickly, attempted to navigate the environment, hesitated briefly around accessing instructions, and then moved toward building. The output began to appear earlier, but the reasoning behind it was less clearly exposed.

Watching both unfold in real time, David called out the contrast directly. While Claude was structuring its workflow before execution, ChatGPT appeared to be interacting more aggressively with the system itself.

At one point, he noted that it seemed to be “trying to go into his system and pull information instead of just following the instructions,” introducing an unexpected dynamic to the exercise.

It became obvious that both systems were capable, but they were solving the same problem in fundamentally different ways.

 

When the Models Began to Look Convincing

As the session progressed, both models began to resemble something usable.

Income statements appeared. Cash flows were constructed. Balance sheets were produced. The structure looked familiar enough that, at first glance, either output could have passed as a completed piece of professional work.

This is where the session deliberately slowed down.

Because, as Gemini had already made clear, appearance was never the standard. The real question was whether the models could hold under pressure.

“On the surface, an AI can generate something that looks right. But if someone actually tests the logic, if they trace the relationships and push the model, that is where the quality truly shows.”

 

What Started to Break Beneath the Surface

Closer inspection revealed weaknesses in both models, although they appeared in different forms.

Claude’s model was structured and transparent, but not immune to error. Certain calculations were placed directly within the assumptions sheet instead of being separated properly into schedules. Some relationships were not implemented in a way that preserved consistency across periods. Circular references appeared in critical sections, and while the framework was clear, the execution had gaps.

What made this more revealing was what happened when it was challenged.

When David prompted it to confirm whether it was truly finished, Claude responded by reviewing its own work and identifying issues, pointing out areas where assumptions had not been applied correctly or where calculations needed to be adjusted. The implication was clear. Producing a model and validating that model are not the same task.

On the other hand, ChatGPT’s issues appeared differently. After an uncertain start, it produced a complete output quickly, but part of that output had been constructed out of view, hidden within a background sheet. From a speed perspective, this was efficient. From a usability perspective, it raised concerns.

The reaction during the session captured this clearly. If a model cannot be easily followed, then it cannot be easily trusted. What is hidden may be functional, but it becomes difficult to validate.

 

Judging Under Real Conditions

Gemini applied a structured rubric that covered coding discipline, financial statement integrity, model structure, presentation, and overall usability. The expectation was strict. A model could not simply look complete. It had to function correctly across all its components.

As the judging framework made clear: “The model must be transparent, logically consistent, and transferable. If a human takes over the model, they must be able to use it without breaking the underlying structure.”

Even with that structure in place, interpretation required more than automated scoring.

Throughout the evaluation, David stepped in to test specific areas, walking through relationships and checking whether results held when examined more closely. This part of the process reflected how financial models are actually reviewed in practice. Outputs are interrogated and not just accepted at face value.

The final scores favoured ChatGPT. Its ability to recover from an uneven start and produce a complete model within the time constraints carried weight in the overall assessment. The structure met enough of the core requirements to secure the win under the defined conditions.

However, the outcome did not close the broader question. If anything, it highlighted it more clearly.

 

 Watch the Replay!

 

What the Experiment Revealed

The most important insight from the session sits beyond the result itself.

Both systems demonstrated that financial models can now be generated quickly. Structure can be assembled. Outputs can be produced in a way that resembles professional work closely enough to pass an initial review.

The limitation appears when the focus shifts from generation to reliability.

A financial model is not valuable because it exists. It is valuable because it can be trusted. That trust depends on the ability to test assumptions, trace relationships, and ensure that outputs behave consistently under different conditions.

Speed does not remove this requirement. It amplifies the consequences of getting it wrong.

Errors can now scale faster. Weak logic can persist inside outputs that look complete. The difference between a working model and a reliable model becomes more important, not less.

 

Where This Leaves the Modern Analyst

The lesson from this session is not about AI capability. That part is no longer in question. The real distinction going forward is between those who can generate output and those who can evaluate it.

The role of the financial modeler is evolving from execution into interpretation. It is no longer enough to construct a model. The expectation is to understand it, test it, and ensure that it reflects reality.

This shift is already creating a divide. Some professionals can generate output quickly. Others can generate and validate it confidently. That difference becomes visible in the quality of work and the level of trust placed in it.

Understanding this shift is one thing. Developing the capability to operate within it is another.

The Financial Modeling Academy scholarship program is designed to bridge that gap. The program focuses not just on building models, but on understanding their structure, testing their integrity, and applying them in real-world situations where decisions depend on accuracy.

To learn more about the program, what to expect and how and when to register for the next cohort, get started here: https://bit.ly/FMARoutes

Be the professional with the skills to validate reality, challenge the data, and confidently drive strategic decisions.