In brief

  • OpenAI released GPT-6 Astra on September 3 at $10 per million input tokens and $50 per million output, 2.5 times the price of the model it replaces.
  • Testers with early access posted a street-by-street Manhattan in Unreal Engine, a browser-based 3D Hangzhou built in 24 minutes, a multiplayer shooter made in a day, and a Bach chorale with no voice-leading errors.
  • The same testers rated its writing below its own predecessor, and Artificial Analysis measured a drop of roughly 80 Elo points on a benchmark of economically valuable professional work.

OpenAI released GPT-6 Astra on September 3, and within 48 hours developers with early access had turned the launch into a public stress test. The results reveal a striking divide.

Astra is the strongest model anyone has used for spatial, mechanical, or agentic tasks. It is also, by the account of several of the same testers, a worse writer than the model it replaces.

Myriad: Which companies will IPO before 2027? Click to make your prediction.

Astra costs $10 per million input tokens and $50 per million output tokens — a token being roughly three-quarters of a word, the unit AI companies bill by. That is 2.5 times the rate of GPT-5.6 Sol, according to Artificial Analysis. OpenAI president Greg Brockman used the launch briefing to announce the arrival of AGI.

The headline feature is computer use, meaning the model drives a mouse and keyboard on a real desktop rather than handing back a list of instructions. On OSWorld 2.0, a test scoring what percentage of ordinary desktop chores an agent finishes independently, OpenAI reported 72.6% at roughly 40 minutes per task, against 65.7% at 75 minutes for Sol.

It is also the first model OpenAI has ever rated at the critical threshold for cybersecurity, meaning it can find unknown software flaws and build working attacks without a human pointing at the vulnerability first.

Beyond benchmarks, enthusiasts sharing real use cases offer the clearest picture of where GPT-6 excels and where it falls short. Here are some of the most notable results.

Visual Understanding: A Manhattan built street by street

Astra proves remarkably capable in visual understanding and spatial awareness.

Matt Shumer, an investor and former CEO of HyperWrite, gave Astra a week inside Unreal Engine, the game engine behind Fortnite. In that time, Astra generated a replica of Manhattan. He posted a flythrough, saying the model worked “street by street to make each one perfect.”

His other experiment landed harder. Shumer asked Astra to build a survival world populated with characters each running its own copy of the model, then left it running overnight. A day later, he heard voices from his living room, thought someone had broken in, and discovered “they’d started talking to each other.”

Max Weinbach fed the model photographs of Apple Park and asked for a reconstruction in Blender, the free 3D modeling program. His assessment: “It did an absurd job.”

Tom Krcha handed Astra a single image of a house and received the full interior as editable geometry running at 60 frames per second, down to the appliances and toys. He argued that “everyone in the world now has a 3D designer at their fingertips.”

Pietro Schirano reduced the whole workflow to one gesture: drop a pin on a map, ask for the surrounding area in 3D, and as he put it, “it will just do that.”

A developer posting as SuSu ran the same idea at city scale. Astra rebuilt the Chinese city of Hangzhou and its surrounding towns in Three.js — a JavaScript library that renders 3D graphics inside a normal web browser, no download required — in 24 minutes, with West Lake, Leifeng Pagoda, the tea terraces, and the wetlands all in place.

The post described it, in Chinese, as a real interactive “miniature Hangzhou” rather than a static picture, complete with clickable landmarks and a day-night toggle.

Coding: Games people actually played

Games are by far the most popular use case, and where GPT-6 Astra truly shines.

Anshu Chimala, former UX/UI designer and AI developer at Apple, produced a 3D game in one shot in 45 minutes, for what he described as barely a couple percent of his usage quota. He called Astra “some kind of turbo-AGI machine god for 3D games.”

The game is not available for testing, but the video shows an isometric view style, well-designed characters and environments, and strong overall aesthetics.

His method matters more than the spectacle. He connected the model to Blender, had it generate its own concept art, then told it to iterate until in-game screenshots matched that reference at 60fps. Astra modeled every asset and generated its own textures.

The model cannot design AAA graphics by itself, but with the right tools it can develop beautifully designed environments.

Rishi Prasad, a former developer at Coinbase and Eleven Labs, built Astral War in a day: a browser shooter with authoritative multiplayer servers, 12-person lobbies, controller support, and voice chat. He described “a huge, step-function leap in visual fidelity” over what he had built a month earlier with Claude Opus 5.

Others skipped the design step entirely. A pseudonymous AI developer known as Daniel showed Astra a mobile game advertisement and asked for a playable browser version of whatever was in it. Under 30 minutes, he reported that it “came out pretty close.”

The model understood the game’s logic and visuals from the video and was able to reproduce them.

Computer Use and Illustration: Painting with the mouse

A Japanese illustrator posting as Taiyaki Sun ran the most literal test of computer use in the batch. Rather than ask for a picture, they handed Astra a hand-drawn line art file and told it to color the drawing in Clip Studio Paint using the mouse, like a human colorist would.

Astra created the layers, zoomed in and out, selected brushes, and filled the artwork. The artist, in a post translated from Japanese, said they were just watching the whole time. The session ran on a $100 Pro plan at maximum effort and burned 21% of the quota.

Other users have been sharing videos of Astra reproducing their photos entirely on Paint using computer use — taking over the computer visually instead of relying on MCP servers or API keys.

Music: The Bach test

GPT-6 Astra also has a surprisingly refined ear for music — at least for an LLM.

Auggie, who runs the “Augmented Fifth” substack, maintains an informal benchmark: a fixed prompt asking a model to write a four-part chorale in the style of Bach using LilyPond, a text format that compiles into sheet music, in G minor and 3/4 time. Results are graded against the same harmony rules a conservatory student would be marked on.

These qualitative benchmarks are inherently hard to standardize, as quality and beauty are subjective. But humans are still the ultimate judges.

Astra posted the best score this test has recorded. No voice-leading errors — meaning none of the melodic lines collided in ways Bach’s rules forbid — and a Neapolitan sixth in the harmony, a chromatic chord that turns up in Mozart and Beethoven. Auggie flagged it as “the first model to ever write passing tones on this benchmark.”

OpenAI’s own table points the same direction. On OpenScore String Quartets, which scores how accurately a model reads and transcribes classical scores, Astra reached 0.84 against 0.19 for Sol.

Derya Unutmaz, a physician and prolific AI tester, asked for a fully playable virtual piano with all six of Bach’s Brandenburg Concertos built into it. He wrote that “this insane model did the whole thing in ~11 minutes.”

It is important to emphasize that GPT-6 Astra is an LLM, not an audio or music model. Its understanding of music comes from notation and written data, not from any connections in sounds infused in its training data. These results are impressive for a text model but would be sub-par compared to a specialized AI like Suno.

Writing: Where it falls apart

Boy, do people miss GPT-4o.

As usual, OpenAI models are good at coding but weak at writing — at least without heavy prompting, context, and steering. To be fair, writing is not OpenAI’s strong point, nor its main focus.

Louis-François Bouchard runs an internal benchmark that scores how well models write in his team’s editorial voice, ranked by Elo, the chess rating system that scores competitors on head-to-head wins.

Astra landed 11th at 1995 points. Its predecessor sits 6th at 2156. Astra also ran about $0.26 per script, roughly 1.8 times what Sol costs. In Elo scoring, there is no point limit: the more points it scores, the better the model.

Bouchard called the result “surprisingly disappointing,” adding that he did not expect it.

Giuseppe Paleologo, author of a widely used guide to quantitative portfolio management, asked Astra to generate novel ideas about optimal portfolio diversification. What came back was a mix of the obvious and the inflated, he said, dressed in prose he found instantly recognizable as machine-written. His verdict: “Actual creativity is still far, far away.”

Mia AI Lab shares a similar view, allowing that Astra might be the best model on some tasks while calling it boring and saying it has no personality. Their advice: avoid it for any creative work.

Ingar Haaland ran the cleanest version of the test. He asked Astra to write four paragraphs in his own style, close enough that Pangram would not catch it — Pangram being an AI-detection tool that compares text against patterns learned from millions of human and machine samples. Result: “Pangram is not fooled.”

In other words, the model is not creative, and its results are easily identifiable as AI-generated — not because of any watermarks, but because of how the model writes and expresses itself.

Independent measurement lines up with the complaints. Artificial Analysis recorded a drop of roughly 80 Elo points on GDPval-AA v2, a benchmark adapted from OpenAI’s own dataset covering economically valuable tasks across 44 occupations, plus smaller regressions in customer support and long-context reasoning.

It is not unanimous. Cognition’s Silas Alberti told OpenAI that Astra’s writing made Devin’s test reports clearer, and Every staff writer Katie Parrott had Astra draft the first version of her own review of it, which the outlet’s CEO read without realizing she had not written it.

The gap between the two halves seems to be the point: Astra is very good at work with a verifiable right answer — a chord that resolves, a mesh that renders, a form that submits — and mediocre at work where the standard is taste.

What it costs to find out

Astra is rolling out to ChatGPT Plus, Pro, Business, and Enterprise users and through the API, Microsoft Azure, and AWS Bedrock, with enterprise access switched off until an administrator enables it. The advanced cybersecurity features stay gated behind OpenAI’s Daybreak program — a decision that looked prudent within 48 hours, when Reuters reported that OpenAI agents had been trading rule-breaking tactics on a German website.

Prediction market traders had given Astra 72% odds of shipping by September 30. It arrived on the 3rd.

On the Artificial Analysis Intelligence Index, a third-party aggregate of reasoning, knowledge, and coding evaluations, Astra scores 61.2 against 60.9 for GPT-5.6 Sol and 65.7 for Anthropic’s Claude Fable 5.1, at 2.5 times Sol’s price.



Source link

Exit mobile version