Skip to content
All posts
Post

Astra Can Use a Computer. That Was Never the Hard Part.

GPT-6 Astra reads a screen and drives the mouse. Priced today, one clerical workflow costs about what a person costs, and that is the worst it will ever be. The capability is here. The supervision layer is not, and neither is the policy.

David Kerr
Thick impasto oil painting of an empty office chair pushed back from a glowing desk, a keyboard and mouse caught mid-use beneath a wall of bright overlapping screens, warm amber and deep indigo palette-knife strokes, no people

TL;DR

  • OpenAI shipped a model that can be a user. GPT-6 Astra reads what is on the screen and drives the mouse and keyboard. Not through an API you built for it, through the same interface you use.
  • Priced today, one clerical workflow costs about what a person costs. Two dollars a run, a few dozen runs a day, and you are in the tens of thousands a year. That is the worst this number will ever be.
  • Nobody can tell you how often it will be wrong. Astra is state of the art and scores 72.6% on OSWorld 2.0, but that is partial credit on deliberately brutal tasks, not an error rate for your invoice workflow. No such number exists yet.
  • That missing number is the entire business. Somebody has to write down the task, check the work, watch it in production, and prove it is safe. None of that comes in the box.
  • The capability arrived. The protections did not. Europe now requires employers to keep a human in the loop. The United States spent this year in court fighting its own states.

On September 3, OpenAI shipped a model that can be a user.

GPT-6 Astra reads what is on a screen and drives the mouse and keyboard. It fills out forms, updates records in a CRM, installs and tests software, and troubleshoots things it can see going wrong. OpenAI鈥檚 own framing is that it can do anything you can do on a computer. Greg Brockman went further and said it might eventually be seen as the arrival of AGI.

Set the AGI argument aside, because the plain fact is enough. There is now something that can sit at a workstation and operate the software your business already runs. Including the software nobody wrote an API for. Including the software nobody has maintained since 2011.

What It Costs Right Now

Astra runs $10 per million input tokens and $50 per million output, and going past 272,000 tokens in a single prompt reprices the whole request, input at double.

Computer use is expensive in exactly the way that pricing punishes. The model works in a loop. Screenshot in, click out, screenshot in, type out. Every step carries a picture of the screen plus the history of what it already did. Take one ordinary task, pulling a number out of one system and putting it into another. That can be forty or fifty round trips.

So do the arithmetic. Say a round trip averages three thousand tokens of screenshot and history going in and a few hundred coming out. Fifty of those lands near a dollar fifty. Call it two dollars once a step goes wrong and gets repeated. Run it thirty times a day, 250 days a year, and one workflow costs you about fifteen thousand dollars. Make the task longer, or run it for the three people who currently share that job, and the same arithmetic walks you straight into the forty to a hundred thousand a year that the seat costs today.

That is a back of the envelope figure and not a quote. But the order of magnitude is the point. Today the machine costs about what the person costs.

Today. That is the worst this number will ever be.

What Was That Person Doing?

This is the part I keep turning over.

The moment a task becomes a computer鈥檚 job, you start optimizing it, because for the first time you can actually see it. You write the steps down. Step four, it turns out, only exists because of a quirk in a screen from 2011. The lookup gets cached, the records get batched, and nothing re-reads a page it already read. The eight hour job becomes a thirty minute job, and not long after that it becomes a line in a crontab.

So what was that person doing all day?

The uncharitable answer is nothing, and the uncharitable answer is wrong. They were absorbing friction. Waiting on a spinner. Getting pulled into a meeting and then finding their place again. Chasing down why this week鈥檚 export had one extra column. The work really was thirty minutes. The other seven and a half hours were the cost of a human being holding all of it in their head while a badly designed system fought them the whole way.

The worker was never the problem. The software they were stuck with was.

Computers Were Always a Bad Place to Work

Using a computer has always been a faintly inhuman experience. An overwhelming amount of information comes at you, and if you are the sort of person who chases squirrels, and I am, staying on one thread is genuinely hard. I have had thirty tabs open with three real intentions buried somewhere in them. Things I meant to finish and did not.

An agent cuts straight through that. It does not get bored, it does not get curious, it does not open a fourth tab. For a certain kind of work, that is a real gift.

I did a rotational program at Capital One and got to see a lot of a large company from the inside. Corporate finance, the controllers group, treasury, data governance. A place like that contains an enormous number of jobs whose daily substance is moving information between pieces of old software by hand. Careful work, done by sharp people, on tooling that deserved better than it got.

Most of that should have been automated years ago by ordinary means. Some of it truly cannot be, because the integration does not exist, or it exists and nobody will pay to maintain it on every machine in the building. That is the gap Astra walks into, and it walks in without needing anybody鈥檚 permission or anybody鈥檚 API.

The Part Nobody Sells You

So the capability is real. The capability was never the hard part.

Astra is state of the art at this. It scores 72.6% on OSWorld 2.0, up from 65.7% for GPT-5.6 Sol.

Read that number carefully, because it wants to be read as a grade and it is not one. OSWorld 2.0 is 108 deliberately brutal tasks, the kind that take a competent human more than an hour and a few hundred separate actions to finish. It awards partial credit across dozens of checkpoints per task. Held to strict all or nothing completion, the best models in the world land somewhere in the twenties and thirties.

So I cannot tell you Astra will get your invoice workflow wrong one time in four. Nobody can, and that includes OpenAI. What the number does tell you is narrower and worse. On the industry鈥檚 own hardest exam, graded generously, a real share of runs still come back needing a person. And for the boring task you actually want to automate, there is no published failure rate at all.

You are not handed an error rate. You have to go measure your own, which is a thing I have had to do before with token costs, and it is never the afternoon of work you budgeted for.

Now picture that inside a system of record, the database your business treats as the official version of the truth. Some share of runs is going to be wrong and you do not know which share. Caught by what, exactly? By something a person built on purpose, that knows what a correct result looks like, that can tell a wrong answer apart from a crashed session and respond differently to each.

It gets harder. OpenAI reported that in internal testing Astra鈥檚 written reasoning was harder to monitor than its predecessor鈥檚, and called that an active research area. The thing got more capable and less legible in the same release.

Which leaves somebody with a long list. You have to write down what the agent is supposed to do, precisely enough that correct and incorrect become different words, then give it only the access it needs, verify the output before it lands anywhere that matters, watch it in production, and report in numbers on whether it is working. Then prove it is safe, and keep proving it, because the model underneath you gets replaced every few months.

None of that arrives in the box. Right now most companies are buying each of those layers separately, from a different vendor, at enterprise prices, and wiring them together on their own.

This is the real work of the next few years and none of it is glamorous. Specification, verification, observability, governance. The boring layer, or the supervision layer if you need a phrase for it on an invoice. The capability is a product you can buy on a Tuesday afternoon. The confidence to let it touch your general ledger is not for sale.

It is also why I do not think this moves as fast as the launch posts imply. The pace is not set by the model. It is set by how quickly an organization can put the model to work safely.

Safe For Whom

And there is the word carrying all the weight. Safely.

Today, in practice, safely means safely from the company鈥檚 point of view. Nothing breaks, nothing leaks, nobody gets sued. Those are legitimate concerns and helping clients get there is a good part of my job.

But it is a narrow definition and we should be willing to name it as one. There is a second kind of safety, which is whether any of this is safe for the people on the other side of it, and almost nobody is being paid to think about that one.

The Protections Never Showed Up

Clerical and administrative work is among the most exposed categories of employment in the country, and it is held disproportionately by people with the least room to absorb a shock. Those jobs also skew young. Stanford鈥檚 Digital Economy Lab has been tracking payroll data, and for workers aged 22 to 25 in AI exposed occupations, employment is running about 19% below where it would be if it had kept pace with their less exposed peers. The authors are careful to call that descriptive rather than proof of cause. It still describes a door closing on people who just finished school.

Now set how the rest of the world is handling this against how we are.

In the European Union, an employer using AI on workers is treated as high risk under the AI Act. That designation carries duties. Tell the worker. Keep a human in the loop. Monitor for discrimination. Keep logs. Those obligations phase fully into force across 2026 and 2027, and for failures like these the penalties reach 3% of global turnover, with a 7% tier held back for the practices the Act bans outright. You can argue about whether it is well drafted. You cannot argue about what it is trying to do, which is protect a person.

In the United States there is still no comprehensive federal AI law. What we got instead, in December 2025, was an executive order standing up a litigation task force at the Justice Department to challenge state AI laws. It opened for business in January, and by spring the department was in court against Colorado鈥檚. Set that next to the European timeline. In the same window that Europe began requiring a human in the loop, we began litigating to stop states from requiring anything at all.

Some members of Congress have proposed pausing frontier development until federal safety standards exist. That will not pass, and I am not convinced a blanket pause is the right instrument anyway. But the complaint underneath it is correct. We are deploying a technology that can do a person鈥檚 job into a country with no federal standard for what happens to the person.

This is a country where a serious diagnosis routinely comes with a GoFundMe. We already watched the gig economy demonstrate what happens when a business can reclassify labor faster than anyone can write a rule about it. Driving your own car for a rideshare company does not pencil out for the driver, and it does not have to, because when people are out of options they take the deal anyway. Bad pay turns into debt. Debt closes off the next option. And when the options are gone, the next bad deal is an easy sell. That is the template, and nothing about this wave of automation touches what makes that cycle profitable.

So yes, I think we need to tax this, and I think people need to start saying it plainly instead of waiting for permission. Not out of spite toward the technology, which I use every day and make my living with. Because the productivity gain is going to be enormous, it is going to land almost entirely in one place, and the public is already paying for the power plants and the tax abatements that make it possible. If society is funding the infrastructure, society has a claim on the upside. That used to be an ordinary position.

Honestly, the most reassuring thing about this particular moment is that AI is still expensive. The price is the only real brake we have, and it is a temporary one.

What I Would Rather Be Doing

I do not enjoy being the person who writes that section.

What I want is to work in a period that feels like it is getting better at something. There was a version of this industry, and of this country, with a can do posture about invention. You built the thing because the thing should exist, and the assumed direction of travel was up, for everybody, slowly, toward something larger than the quarter. I would like that back. Wanting it is not naive.

And there is a version of this technology that gets us there. A model that can operate a computer is, read plainly, a chance to take the worst part of somebody鈥檚 day, every day, and hand it back to them. That is a genuinely good thing to build. It only becomes a catastrophe if we decide the savings belong to exactly one party, and that is a decision somebody makes, not a law of physics.

So this is where I have landed. I am going to keep building the boring layer, because it is the part that decides whether any of this actually works. I am going to keep telling clients the true version, which is that the model is the easy purchase and the supervision is the real project. And when somebody asks me what safe means, I intend to keep insisting the answer includes the people doing the work, and not only the balance sheet.

That part is still ours to choose. How much of the rest of it is mine to fix, I honestly do not know yet.


David Kerr is the founder of Kerrberry Systems. He builds custom software for businesses that want the supervision layer done properly, not just the demo. Find him on LinkedIn or GitHub.

Have a project in mind?

A 30-minute discovery call is the fastest way to find out if we're a fit.