🔍 Read the full analysis: An AI Stack With Distinct Roles: My September 2026 Setup on ThorstenMeyerAI.com
Get tech for your team delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
With six frontier AI models clustered within about 20 index points on the Artificial Analysis Intelligence Index but differing roughly 100-fold in cost per task, developer Thorsten Meyer outlines a September 2026 stack that assigns distinct roles: Claude Opus 5.5 as main builder and newly released GPT-6.1 Sol as low-cost detail and review model.
Developer Thorsten Meyer has published his working AI stack for 29 September 2026, built around a frontier-model field where six leading models sit within roughly 20 index points of each other on capability while their cost per task differs by about 100×. His setup assigns Claude Opus 5.5 as the main builder and the just-released GPT-6.1 Sol as a low-cost model for detailed work and code review, a division of labor driven less by raw intelligence than by price-performance at the task level.
According to Meyer’s write-up, the core question has shifted from “which model is smartest?” to “which model clears my quality bar at the lowest cost per task?” He cites data from the Artificial Analysis Intelligence Index v4.3.x, which he describes as “a map of general capability, not a verdict on your workload,” urging readers to shadow-test before switching models.
Three findings stand out in his data. First, Opus 5.5 (released 22 September, index 58 at max) outscores its more expensive sibling Fable 5.1 (53) by 5 points while costing less per task — $5.98 against $7.63. Second, Sonnet 5.5 at max effort costs $7.60 per task for 56 points, more than Opus at max for 2 fewer points, making it hard to justify at that setting. Third, GPT-6.1 Sol at xhigh costs $0.39 per task — roughly one-eighth of Astra’s $3.26 and one-twentieth of Fable’s $7.63 — for a score only 1 to 2 points lower.
Meyer also reports that the effort setting is a bigger cost lever than model choice. On Opus 5.5, moving from xhigh to max adds 2 index points and 73% more cost per task; from medium to max, cost rises 4.46× for 7 points. He runs Opus at high (54 points, $1.82 per task) for development and reserves xhigh for architecture, migrations, and trust boundaries.
Opus builds. Sol reviews. Jev decides.
One price tape, six models
Score against cost, at every effort setting
The effort dial moves the bill more than the model
Claude Opus 5.5
Claude Sonnet 5.5
GPT-6.1 Sol: near-Astra scores at a fraction of the price
Three published settings
| Setting | Index | Cost per task | Output tokens | First token |
|---|---|---|---|---|
| medium | 48 | $0.21 | 15M | 5.3 s |
| high | 50 | $0.32 | 25M | 57 s |
| xhigh | 51 | $0.39 | 36M | 69 s |
Same score band, very different bill
My stack: who builds, who reviews
Cheaper tokens are not cheaper work
Read the numbers with four warnings
Part 2: Jev, the model that decides instead of writing
One call in, typed answers out
Three question types
Confidence is the superpower
Three uses running in my publishing operation
The fit test, then the shadow test
- Replay 300 to 500 past decisions
- Compare overall and per confidence band
- Read 20 disagreements, decide who was right
- High band at 95% or better?
- Own flag, off by default
- Canary on 5 to 10 units
- Roll out in the confident band only
24 use cases, sorted by how well they fit
Proven in production
- 1Relevance gate
- 2Language check
- 3Classifier fallback
Publishing and content
- 4Thin-source detector
- 5Same-event dedupe
- 6Product fits roundup
- 7Disclosure present
- 8Headline quality
- 9Comment moderation
Commerce and support
- 10Support-ticket routing
- 11Return-reason coding
- 12Review to feature complaints
- 13Catalogue taxonomy
- 14Order-fraud pre-triage
Software and AI systems
- 15LLM guardrail
- 16RAG passage filter
- 17Citation check
- 18Tool and intent routing
- 19Log-line triage
- 20PR risk triage
Business ops and home
- 21Inbox triage
- 22Expense categorisation
- 23Lead qualification
- 24Smart-home intent
Limits, cost and one hard rule
Why Price-Performance Now Beats Model Loyalty
The stack illustrates a broader shift in how practitioners choose AI tools: as capability gaps between frontier models narrow to single index points, cost per task becomes the deciding factor. A review pass at $0.32 to $0.39 per task is cheap enough to run routinely on every meaningful change, which changes quality-assurance economics for development teams.
Meyer argues the review seat matters most: a different model family reviewing Opus’s output is a better check than Opus reviewing itself. He pairs that with four working rules, including “effort is not capability,” “a different model is not an independent review if both read the same flawed spec,” and “passing tests are not approval to ship.”
He also cautions that cheaper tokens are not cheaper work: halving model price saves only 12.5% of real cost, and one extra minute of human review can erase the saving — an example he flags as illustrative, not measured.
A Compressed Month of Frontier Releases
The stack caps an unusually crowded September 2026. According to Meyer’s timeline, Claude Fable 5.1 launched 1 September, GPT-6 Astra on 3 September, Claude Opus 5.5 and GPT-6 Luna on 22 September, Claude Sonnet 5.5 on 28 September, and GPT-6.1 Sol on 29 September — the same day as publication.
GPT-6.1 Sol launched at the same $2/$10 per 1M tokens (input/output) as its week-old predecessor. Meyer notes that even its medium setting matches the earlier GPT-6 Sol’s score of 48 at one-fifth the cost per task ($0.21 versus $1.06). Published per-token prices: Opus 5.5 at $4/$20 with cache reads at $0.20; Fable and Astra at $10/$50; Sol at $2/$10; Luna at $0.10/$0.50.
Meyer also deploys Jev, a decision model he describes as unable to write a sentence, for high-volume yes/no and routing judgements, alongside GPT-6 Luna for classification, extraction, and routing at $0.07 per task.
“In four weeks, the AI frontier stopped being a leaderboard and became a price curve.”
— Thorsten Meyer
Open Questions Around GPT-6.1 Sol
Several limits remain. Meyer reports that Artificial Analysis has not yet published low or max settings for GPT-6.1 Sol, and notes that a one-index-point difference is inside measurement noise. Sol’s high and xhigh settings take 57 to 69 seconds to produce a first token, which he says rules it out as an interactive model at those settings.
All scores come from a single benchmark family (Artificial Analysis Intelligence Index v4.3.x), and Meyer’s cost-per-task figures are specific to that index’s task set — real workloads may produce different ratios. His claim that model price is a small fraction of total cost is labeled illustrative, not measured. Whether Sol’s extremely concise output (25M tokens on the index against a median of 82M for comparable models) translates to shorter answers in production use is not established.
Watching Sol’s Benchmarks and Stack Evolution
Meyer expects to keep refining the stack as benchmarks mature. Immediate watch items include Artificial Analysis publishing low and max effort settings for GPT-6.1 Sol, which will show whether the price-performance gap widens or closes at the extremes.
He plans continued shadow-testing before any model becomes a default, and says Astra and Fable stay in the stack only as tie-breakers when Opus and Sol disagree — a role whose cost-effectiveness will be re-evaluated as cheaper models catch up. The larger open question for practitioners is whether the September compression of scores holds, which would push more decisions toward routing models like Jev and Luna rather than a single premium default.
Key Questions
What is the main idea behind this September 2026 AI stack?
With six frontier models within about 20 index points of each other but roughly 100× apart in cost per task, the setup assigns each model a role based on price-performance: Opus 5.5 builds, GPT-6.1 Sol handles details and review, and cheaper models like Luna and Jev handle classification and routing.
Why is GPT-6.1 Sol used for review instead of building?
According to Meyer, Sol scores 1 to 2 points below Astra and Fable but costs $0.39 per task instead of $3.26 or $7.63. That makes a review pass cheap enough to run routinely, while Opus 5.5 retains a 5-point lead at xhigh for demanding build work.
What are GPT-6.1 Sol’s main drawbacks?
Its high and xhigh settings take 57 to 69 seconds to first token, making it unsuitable for interactive use, and Artificial Analysis has not yet published low or max effort settings. Its benchmark scores also trail Opus 5.5 by about 5 points at comparable settings.
Does raising the effort setting make a model smarter?
No, per Meyer’s rule that “effort is not capability.” On Opus 5.5, going from xhigh to max adds only 2 index points for 73% more cost per task, and medium-to-max raises cost 4.46× for 7 points.
Where do the benchmark scores in the stack come from?
All scores come from the Artificial Analysis Intelligence Index v4.3.x, which Meyer describes as a general-capability map rather than a verdict on any specific workload. He recommends shadow-testing before switching models.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
