Last week I opened a live GA4 reporting sheet, put the ChatGPT sidebar next to it, and switched on computer control. This wasn’t a sandbox and it wasn’t sample data. It was a reporting workbook I actually use, and I gave the model permission to click, type and build in it while I sat and watched.The whole thing runs a little over two minutes and you can watch the recording here. The part I keep thinking about isn’t where it built the report.

It’s the bit near the end where I stopped, clicked into a cell, and looked at what was actually sitting in there.


First, what “computer use” by ChatGPT means

This isn’t a chatbot writing formulas into a text box for you to copy. OpenAI’s documentation describes the capability as the model being able to “see and operate graphical user interfaces on macOS or Windows”. You grant screen recording and accessibility permissions, approve the specific app, and the model drives.

Now, there are two guardrails matter. OpenAI states that “You can stop the task or take over your computer at any time,” and that on Windows the model takes foreground control while it runs. Watching isn’t an optional nicety, it’s how the thing was designed, which is great.

The capability is real and it’s measurable. On SpreadsheetBench, a benchmark built from real-world spreadsheet editing tasks, OpenAI reports ChatGPT agent scoring 45.54% with direct .xlsx editing against Copilot in Excel’s 20.0%. That’s more than double Copilot, and it’s still under half the tasks.

Hold both of those numbers in your head at once and you’ve about the right posture for this thing.

What it was fast at, and what it stalled on

It’s always amazing seeing what these new models can do. They can now build full spreadsheets just like any of us can – formatted with good-looking tabs. But for those of us that have bled in the trenches of spreadsheets for years, we see where they eventually break down.

Formatting flew. Column widths, headers, number formats, all of it gone in seconds. That tracks with the general shape of these tools: they’re excellent at the mechanical work you’d otherwise do with your hands.

Other steps stalled. The sheet sat still long enough that I started narrating my own doubt on camera. The model was almost certainly working somewhere I couldn’t see. From the outside, “thinking” and “broken” look identical.

That gap is the honest state of the art. On OSWorld, a benchmark of 369 real computer tasks across web and desktop applications, humans complete 72.36% while the best model at publication managed 12.24%. That’s the figure from the benchmark’s 2024 publication and agents have improved a great deal since, but nobody should be surprised by a pause.

The uneven performance has a name. Ethan Mollick calls it the jagged frontier.

Everything inside the wall can be done by the AI, everything outside is hard for the AI to do.

Ethan Mollick, “Centaurs and Cyborgs on the Jagged Frontier”

The frontier is invisible and it’s not smooth. In the Boston Consulting Group field experiment Mollick co-authored, consultants working inside AI capability finished 12.2% more tasks, 25.1% faster, at more than 40% higher quality. On a task deliberately placed just outside it, accuracy dropped from 84% unaided to somewhere between 60% and 70% with AI assistance.

Same tool, same people, opposite result, and nothing on the screen tells you which side of the line you’re standing on.

Formula, or pasted value?

I asked it to build an NLP report. The content appeared directly in the spreadsheet. It looked right.

So I stopped and clicked into a cell.

Did it write a spreadsheet formula, or did it run Python somewhere else and paste the answer in as a value? Both put a number on the screen. Only one of them gives you something you can hand to a client.

This isn’t a purist objection. Microsoft published the same warning about its own product. The documentation for Excel’s COPILOT function, which Microsoft is retiring on 14 September 2026 about a year after shipping it, is unusually blunt:

COPILOT uses AI and can give incorrect responses. Formula results may change over time, even with the same arguments.

Microsoft, COPILOT function documentation

The same page tells users not to use the function for numerical calculations, not to use it for workbook lookups, and to avoid it for financial reporting, legal documents and other high-stakes scenarios. When the vendor shipping a feature tells you where not to point it, believe them. That it’s now being withdrawn makes the warning more instructive, not less.

Anthropic went the other direction and made traceability the product. Claude for Excel is built to track and explain its changes and to let users navigate directly to the cells it references. Different companies, same conclusion: the work has to be visible in the workbook.

The two outcomes, side by side

Logic lives in the cellLogic ran somewhere else
AuditabilityClick any number, read the formulaThe number is right, or it is not
RefreshChange an input, the sheet updatesThe sheet quietly goes stale
HandoffA client or an auditor can retrace itOnly the person who ran it knows
Failure modeVisible and fixableSilent

Spreadsheets fail silently. That’s their defining characteristic, and it predates AI by decades. In 2020 Public Health England under-reported 15,841 positive COVID-19 tests because the data sat in the legacy .XLS format, which caps out at 65,536 rows. Contact tracing stalled for thousands of people because of a limit nobody had inspected.

Nor is this rare. Summarizing the inspection studies in his field, University of Hawaii researcher Ray Panko reports that auditors found errors in 94% of the spreadsheets they examined closely. That’s a small sample, 85 workbooks, drawn on purpose from cases where somebody looked hard. Which is rather the point. Nobody looks hard at most spreadsheets.

The error rate was already bad. What’s new is that the cells now fill faster than anybody reads them.

The business writers got here first

Almost everything worth knowing about working alongside a fast, capable, occasionally wrong assistant was written down before any of us had one.

On trust

Ronald Reagan spent the back half of the Cold War repeating a Russian proverb until Gorbachev started teasing him about it. At the INF Treaty signing in December 1987:

The maxim is: Dovorey no provorey. Trust, but verify.

Ronald Reagan, Remarks on Signing the Intermediate-Range Nuclear Forces Treaty, 1987

Gorbachev told him he repeated it at every meeting. Reagan replied that he liked it. It’s the entire operating manual for AI in a client deliverable, in three words.

On fooling yourself

Richard Feynman, in his 1974 Caltech commencement address on what he called cargo cult science:

The first principle is that you must not fool yourself — and you are the easiest person to fool.

Richard P. Feynman, “Cargo Cult Science,” Caltech, 1974

There’s now a controlled experiment pointing exactly this way. METR ran a randomized trial with 16 experienced open-source developers on real tasks in their own repositories. The developers using AI tools were 19% slower. Afterwards, they estimated that AI had made them 20% faster.

They weren’t lying. It really did feel faster. Perceived productivity and measured productivity came apart completely, and the people best positioned to notice didn’t notice. Time your own runs before you trust the feeling. (The study used early-2025 models and METR has since revised its experimental design, so treat it as a caution about self-assessment rather than a fixed verdict on the tools.)

On doing the wrong thing efficiently

Peter Drucker, writing in Harvard Business Review in 1963, long before anyone could automate a bad process at scale:

There is surely nothing quite so useless as doing with great efficiency what should not be done at all.

Peter F. Drucker, “Managing for Business Effectiveness,” HBR, 1963

He drew the distinction sharply again in 1974: “Efficiency is concerned with doing things right. Effectiveness is doing the right things.” A model that builds the wrong report in seconds has not helped you. It has just made the wrong report cheaper to produce, which is worse, because now you will make more of them.

The line usually attributed to Bill Gates makes the same point about automation specifically: automation applied to an efficient operation magnifies the efficiency, and automation applied to an inefficient operation magnifies the inefficiency. (Widely credited to Business @ the Speed of Thought, 1999, though I couldn’t verify it against the book itself.)

On what your job just became

The moment you hand over the keyboard you stop being a producer and start being a manager. Andy Grove defined the job in High Output Management:

The output of a manager is the output of the organizational units under his or her supervision or influence.

Andrew S. Grove, High Output Management, 1983

Your output is now whatever the model produced, and you own all of it. Grove also had the best available warning about measurement, which reads uncannily well as advice about watching an agent work:

Indicators tend to direct your attention toward what they are monitoring. It is like riding a bicycle: you will probably steer it where you are looking.

Andrew S. Grove, High Output Management, 1983

If the only indicator you watch is “did the report appear,” that’s the only thing you will optimize for.

On going to look for yourself

Taiichi Ohno built the Toyota Production System on the principle that you don’t accept a report about a problem, you go and look at the problem. In Workplace Management:

I think it is important to get in the habit of verifying failures with one’s own eyes.

Taiichi Ohno, Workplace Management, 1982 (English edition 1988)

Toyota calls the broader practice genchi genbutsu, go and see for yourself. Here that means clicking the cell.

On knowing what you are doing

Warren Buffett, talking to Columbia business students in 1993:

Risk comes from not knowing what you’re doing.

Warren Buffett, Columbia Business School, 1993

People ask whether they can use this if they’re not great with spreadsheets. They can. The trouble is that running it and checking it are different skills, and only one of them is optional.

Simon Willison draws the same line for AI-written code, and it transfers directly:

If you’re going to put your name to it you need to be confident that you understand how and why it works.

Simon Willison, “Will the future of software development run on vibes?” 2025

One quote to stop using

While we are here: “what gets measured gets managed” is not Peter Drucker. The Drucker Institute confirmed as much to researcher Danny Buerkli, who went looking for the source and couldn’t find one. The nearest scholarly ancestor is V. F. Ridgway’s 1956 paper on the dysfunctional consequences of performance measurement, and Ridgway’s argument was the opposite of how the phrase gets used. He was warning that measuring things badly makes organizations worse.

Which is, as it happens, exactly the risk with an AI-generated report that looks authoritative and can’t be traced.

How to run this test yourself in ten minutes

If you want to know whether this fits your reporting workflow, stop reading threads about it and run it. Judge the output the way you’d judge a junior analyst’s first draft.

  1. Start with a sheet you already understand. Pick a report you’ve built by hand. You can’t audit an output you couldn’t have produced yourself, and the first run is about calibration, not speed.
  2. Connect real data before you prompt. Wire in the live source first. I use Adformatic’s Google Sheets reporting add-ons, which lean toward PPC and advertising data; Supermetrics covers a wider source list including GA4. Sample data teaches you nothing about the workflow.
  3. Ask for structure, not conclusions. Headers, layout, a first pass of formulas. Keep each generation small so it stays cheap to check.
  4. Click three cells at random. Formula or pasted value? If it’s a value, ask the model to rebuild it as a formula. This single habit is the difference between a report and a screenshot.
  5. Break one input on purpose. Change a date range or a channel filter. A sheet with real logic moves. A sheet full of pasted answers sits there looking confident.

Where I have landed

I’m keeping it. The throughput is real, and for anyone sitting on a backlog of spreadsheets they will never find an afternoon to build, that’s the whole argument.

But the throughput isn’t the story. Jim Collins spent five years studying what separated great companies from good ones and gave technology its own chapter in Good to Great, “Technology Accelerators.” Writing the finding up for Newsweek, he put it plainly: technology is “an accelerator of greatness already in place, never the principal cause of greatness or decline.”

That’s the frame I keep coming back to. Computer control in a spreadsheet accelerates whatever process you already have. If your reporting is well structured and you know what a good number looks like, it will make you noticeably faster. If your reporting is held together with pasted values and hope, it will help you produce more of that, faster, with a cleaner font. That is a process problem, and no model fixes it for you.

Karpathy has the cleanest framing of it:

It’s less Iron Man robots and more Iron Man suits that you want to build.

Andrej Karpathy, “Software Is Changing (Again),” 2025

I’m not fully settled on this, and I’d rather say so than pretend otherwise. Ask me again in three months, after I’ve either caught something ugly or quietly stopped checking as carefully as I did on day one. The second one is the failure mode I actually worry about.

For now: the suit does the heavy lifting. I’m still the one flying it.

Watch the two-minute test, then go click a cell in your own sheet. I’d like to know what you find.


This is the first entry in Testing AI, where we run real tools against real client work and publish what actually happened. More on how we think about AI in the reporting stack over on AI content analysis. If you want reporting you can hand to a client and defend line by line, talk to us.

Joe Robison

Founder & Consultant
Joe Robison is the founder of Green Flag Digital. He founded the agency in 2015 and has been heads-down scaling content marketing and SEO services for clients ever since. He is an occasional surfer, fledgling yogi, and sucker for organized travel tours.