When you need frontier intelligence?

Two small tests, three new models, and the question of how much better an answer needs to be before the extra cost and waiting feel worthwhile.
I've been quiet for a few weeks. We finally moved into our new house, and it was about as exhausting as you can imagine. Packing, unpacking, selling things, throwing things away, and somehow buying more things. An endless succession of chores.
AI made some of them less painful. No army of humanoid robots carried our boxes. My wife and I used it to research repairs, compare contractors' quotes, find furniture, and work out everywhere we needed to change our address. That last one alone was useful enough to make me wonder how we managed before.
During the move, we mostly relied on Gemini Flash models. We had no complaints. Then the new Fable and Astra models arrived, and I caught myself wondering how much of that frontier intelligence I would actually use.
I admire progress at the frontier. Unfortunately, my everyday life contains considerably more furniture shopping than unsolved scientific problems. Was I missing something by reaching for the fast model?
I tried two comparisons: shopping for a coffee table and building a browser game. Astra gave me the shopping answer I liked best, but Flash was already useful. The more complex coding task made a stronger case for the frontier models. That difference is what interests me: how much better does an answer need to be before the extra cost and waiting feel worthwhile?
Three flagships in seventy-two hours
Anthropic introduced Claude Fable 5.1 on September 1, Google released Gemini 3.8 Flash on September 2, and OpenAI announced GPT-6 Astra on September 3. It was a busy week to have an opinion about AI.
For this comparison, I'm using "frontier" to refer to Fable and Astra, the two more expensive models I tested. Flash is my everyday workhorse. Those labels describe their role in this experiment; they don't settle which one will do a particular job better.
| Model | Input / 1M | Output / 1M |
|---|---|---|
| Claude Fable 5.1 | $10.00 | $50.00 |
| GPT-6 Astra | $10.00 | $50.00 |
| Gemini 3.8 Flash | $0.75 | $3.75 |
These are base API prices per million tokens, checked on September 13, 2026. Google's Flash rate is introductory pricing through December 31, 2026. Cached inputs can cost less, and these rates are separate from consumer subscription prices. Sources: Anthropic, Google, and OpenAI.
The price difference is substantial. Whether it buys something useful is what I wanted to find out.
Two small tests
Benchmarks help establish what models can do, but they don't tell me how much I'll value the difference while choosing furniture or working on a prototype. I wanted to try tasks I would actually ask an assistant to do.
This is a small, informal comparison: one sequence of prompts per model for each task. It reflects the models, tools, and settings I used, including High effort for Fable. It doesn't isolate model intelligence from the surrounding software, and repeated runs could produce different results. I'm treating it as a way to examine my own usage, not a general ranking.
I looked for two things: whether an answer respected the request, and whether the result was useful enough to justify the time spent getting it. For the game, that also meant separating how it looked from how it played.
Shopping for a coffee table
Furniture research consumed a surprising amount of my life during the move. Product names tell you very little, descriptions bury the materials, and "oak" can mean several different things. It seemed like a good test of an assistant's ability to turn a messy search into a useful shortlist.
I used these two prompts:
Find me five solid oak coffee tables in Scandinavian style, 40 to 52 inches wide, under $800, available in the US. For each one give the retailer, the dimensions, whether it's solid wood or veneer, and the price. Don't include anything you can't verify.
Of those five, which will hold up best over ten years with kids in the house? Explain what you're basing that on.
Three useful answers
Flash returned five tables with dimensions, prices, and descriptions of their construction. It followed up with a detailed durability comparison and a recommendation for the Article Baarlo. The whole exchange took seconds.
It wasn't a perfect match to the brief. The tables included veneer, which Flash disclosed, despite my request for solid oak. That would matter if solid wood were an absolute purchasing requirement. In this exercise, I still found the result useful: it gave me plausible alternatives and enough material information to understand the compromise. I would check the retailer's specifications before buying any of them.
Fable was more explicit about the difficulty of finding five exact matches. It separated potentially suitable options from mixed-material tables, over-budget choices, and unavailable products. Its recommendation favored the VarbrosHome Lake within the intended budget, subject to the size and final US price, with the Ethnicraft Nordic as a more expensive alternative. It also distinguished general furniture-construction reasoning from long-term test data on these particular tables.
Astra returned three tables it identified as qualifying: the Magnus from Icon By Design, the Nico from Castlery, and the Amoeba from Article. It explained why it had stopped at three, disclosed a style caveat, and recommended the Nico with moderate confidence. Its reasoning made the evidence and uncertainty easy to follow.
The three models ended up recommending different tables. Apparently, even the world of coffee tables leaves plenty of room for judgment. None of these answers could establish which table would actually survive ten years in my house, but all three helped me think through the purchase.
Astra was better but Flash was already useful
Of the three, I preferred Astra's answer. It was concise, easy to scan, and clear about what supported the recommendation. Models have become awfully chatty. When I ask a simple question, I don't always want a small book in return.
That preference doesn't mean I was unhappy with Flash. For the kind of shopping research I was doing during the move, a useful shortlist with disclosed compromises could be enough to keep me moving. Astra gave me a better answer; the question was whether I needed that improvement on every search.
The difference in waiting was substantial. Across both prompts, my recorded times were 5.6 seconds for Flash, about 3 minutes 50 seconds for Fable, and 12 minutes 3 seconds for Astra. Astra's first answer alone took nine minutes.
The apps reported estimated costs of about $0.01 for Flash, $2.60 for Fable, and $8.30 for Astra across the two prompts. These are the reported figures for my runs; caching and usage accounting can differ between apps. The gap was substantial.
I find it funny that the shortest answer came at the greatest cost. That doesn't mean brevity caused the cost: I was paying for the work leading up to the answer, not just the words I read.
If the extra work saves a long research session or helps me avoid a purchase I regret, it may be worth it. But for an ordinary furniture search, I wasn't convinced that the improvement justified the difference. I would still happily start with Flash and ask for another pass if the material requirement became decisive.
| Model | Prompt | Elapsed time | Input tokens | Output tokens | Estimated cost (USD) |
|---|---|---|---|---|---|
| Gemini 3.8 Flash | #1 | 3.2s | 2,700 | 450 | $0.00270 |
| Gemini 3.8 Flash | #2 | 2.4s | 3,700 | 480 | $0.00329 |
| Gemini 3.8 Flash | Total | 5.6s | 6,400 | 930 | ~$0.01 |
| Fable 5.1 High | #1 | ~3 minutes | ~1.5M | ~8,000 | $2.50 (with caching) |
| Fable 5.1 High | #2 | 50s | ~150k | ~900 | $0.10 (with caching) |
| Fable 5.1 High | Total | 3m50s | ~1.65M | ~8,900 | $2.60 |
| ChatGPT Astra | #1 | 9 minutes | 4,048,473 | 4,015 | $6.14 |
| ChatGPT Astra | #2 | 3m03s | 1,254,389 | 3,706 | $2.16 |
| ChatGPT Astra | Total | 12m03s | 5,302,862 | 7,721 | $8.30 |
Building a game in four steps
For the second test, I asked each model to build a browser game and then made the request progressively harder. I started with movement, added a visual direction, introduced gameplay systems, and finally asked for architectural changes and verification.
The screenshots show each stage. Flash is on the left, Fable in the middle, and Astra on the right.
Step one gets everyone moving
The first prompt asked for a single HTML file using Three.js from a CDN, with a ground plane, a capsule-shaped player, WASD movement, a following camera, mouse look, and jumping with gravity.
All three produced broadly similar starting points. The visual treatment differed, but this step gave me little reason to insist on a frontier model. A fast model was enough to get the basic idea onto the screen.

Step two separates the visual approaches
Next I asked for a forest, a humanoid high elf with a sword, and "hyper realistic, AAA quality" graphics.
That last part was deliberately ambitious and imprecise. None of these screenshots looks like a finished AAA game. The useful comparison is how each model interpreted the direction within a browser prototype.
Astra produced the richest-looking environment to my eye, with a stronger sense of atmosphere. Fable and Flash took simpler visual approaches. At this stage, Astra gave me the most compelling image, but appearance alone couldn't tell me how much I would enjoy playing it.

Step three makes the systems interact
The third prompt added three enemy types, each with idle, patrol, chase, attack, and flee states. I also asked for melee combat, damage numbers, an inventory with weapons that changed damage, and save/load that would survive refreshing the page during a fight.
This is where the task became more than adding visible features. An enemy needs to react to the player and to damage. Changing weapons needs to affect combat. Saving needs to preserve the relevant state so that loading resumes a coherent game.
Across the more developed versions, I found Fable's game more immediately playable. Astra continued to impress me visually, but its animation and gameplay felt awkward. Those are different kinds of quality, and my preference depended on whether I was looking at the scene or trying to play it.

Step four asks for more than a convincing demo
The final prompt asked for an entity-component-system architecture, quests with prerequisites and mutually exclusive outcomes, backward-compatible saves, 60fps with 200 enemies, and deterministic replay with a test showing that the final state matched.
Flash regressed at this stage. Fable and Astra held up better in my overall assessment. The gap was much clearer than it had been in the first, simple prototype.
I would be careful about calling any of these finished games. Screenshots and play impressions don't independently verify the architecture, save compatibility, or deterministic replay. A proper performance test also needs defined hardware and a repeatable workload. My conclusion here is narrower: as the prototype became more ambitious, the frontier models delivered results I valued more.

What the improvement cost
Adding up the four recorded prompts, I got these totals:
| Model | Prompt | Elapsed time | Input tokens | Output tokens | Estimated cost (USD) |
|---|---|---|---|---|---|
| Gemini 3.8 Flash | #1 | 1m 33s | ~158,800 | ~16,000 | $0.022 |
| Gemini 3.8 Flash | #2 | 1m 32s | ~175,900 | ~22,100 | $0.026 |
| Gemini 3.8 Flash | #3 | 1m 39s | ~307,500 | ~23,500 | $0.040 |
| Gemini 3.8 Flash | #4 | 4m 01s | ~3,181,300 | ~48,300 | $0.337 |
| Gemini 3.8 Flash | Total | 8m 45s | ~3,823,500 | ~109,900 | $0.425 |
| Fable 5.1 High | #1 | 1m 51s | 540,602 | 6,338 | $0.92 |
| Fable 5.1 High | #2 | 15m 41s | 1,921,759 | 71,246 | $5.60 |
| Fable 5.1 High | #3 | 13m 30s | 2,695,263 | 65,320 | $5.33 |
| Fable 5.1 High | #4 | 25m 23s | 7,498,533 | 129,403 | $11.06 |
| Fable 5.1 High | Total | 45m 24s | 12,656,157 | 272,307 | $22.91 |
| ChatGPT Astra | #1 | 3m 35s | 282,949 | 6,430 | $1.08 |
| ChatGPT Astra | #2 | 16m 08s | 1,468,121 | 26,931 | $3.19 |
| ChatGPT Astra | #3 | 11m 50s | 1,152,576 | 21,160 | $2.57 |
| ChatGPT Astra | #4 | 32m 42s | 3,307,000 | 49,916 | $6.76 |
| ChatGPT Astra | Total | 1h 04m 15s | 6,210,646 | 104,437 | $13.61 |
For the four-step game, the apps reported estimated costs of roughly $0.43 for Flash, $22.91 for Fable, and $13.61 for Astra. The frontier models asked much more of my budget as well as my patience.
Those totals describe this sequence of runs, rather than a general measure of model speed. They also leave out the time a person might spend fixing the result afterward.
That second number could change the decision. If I want to explore a rough idea, Flash's short turnaround is valuable. If I need to keep extending the game without losing existing behavior, spending longer on a stronger model may save me repair work. This experiment didn't measure that repair time, so I can't calculate the tradeoff precisely.
For this game, Fable gave me the more playable result, while Astra gave me the stronger visual direction. For a simple starting point, Flash was enough. The models became easier to distinguish as more parts of the game had to work together.
What makes the extra intelligence worth it
A subscription can make the cost of an individual request less visible. It doesn't make waiting disappear. If I'm sitting in front of the screen trying to make a quick decision, seconds and minutes feel very different. If I've handed over a substantial piece of work and gone off to do something else, a longer run may be perfectly reasonable.
The API price table is also only a starting point. Actual cost depends on uncached and cached inputs, output tokens, and any tool charges. A model that costs more per token may need fewer attempts to reach a result I'm happy with. These small tests don't measure every follow-up or repair I might need, so they can't establish the cheapest way to finish every task.
That's why I keep coming back to the result. For shopping, I liked Astra more, but Flash had already done enough to help. For the later game prompts, the difference affected whether I liked the environment and wanted to keep playing. I was asking for something more demanding, and I could see more value in the extra capability.
When I would reach for the frontier
I started this comparison wondering whether I needed frontier intelligence in everyday life. My answer is: sometimes, but I don't feel a need to use it for everything.
Flash remains a very good place to start for the ordinary tasks that fill my day. It helped during the move, gave me useful furniture research, and got a simple game prototype running. I don't need every one of those interactions to produce the best answer a model can possibly give.
I'd reach for Fable or Astra sooner when I want a more ambitious result, when several parts of a project have to work together, or when an improvement in quality is worth a longer wait. I'd also escalate when a fast model's answer leaves me doing too much of the work myself. The point at which that happens will depend on the task and on what I consider good enough.
Even then, the choice isn't obvious. In this game, Fable gave me the more immediately playable result, while Astra created the environment I liked looking at most. Trying both taught me more than a single overall ranking would have.
I'm glad these models exist, and I'll keep finding reasons to use them. I just don't want admiration for the frontier to become a habit of sending every small task there. During a move, a useful answer in a few seconds can be exactly what I need. When I want to build something more ambitious, I'm much more willing to pay and wait for a better one.
ps. If you're curious about all the prompts I used and model responses head to my substack: https://marcintreder.substack.com/ – I'm sharing more data with my premium subscribers.