🔍 Read the full analysis: Opus, Sol, Jev: Three Roles In My September 2026 AI Stack on ThorstenMeyerAI.com
Get business pricing on monitors, keyboards and dev gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
Thorsten Meyer’s 29 September 2026 stack review assigns Claude Opus 5.5 as the main building model and newly launched GPT-6.1 Sol as a low-cost detail and review model. Six frontier models now score within roughly 20 index points while task costs differ by about 100x, shifting the decision from capability to price per task.
Independent AI commentator Thorsten Meyer published a full reorganisation of his working AI stack on 29 September 2026, the day GPT-6.1 Sol launched. His conclusion: with six frontier models clustered within about 20 index points on the Artificial Analysis Intelligence Index v4.3.x while their cost per task differs by roughly 100x, the practical question is no longer which model is smartest, but which one clears a quality bar at the lowest cost per task.
Meyer’s stack assigns four distinct roles. Claude Opus 5.5 (released 22 September, index score 58 at max, $5.98 per task) is the main builder, run at high or xhigh effort rather than max. GPT-6.1 Sol (released today, index 51 at xhigh, $0.39 per task) handles detail work and code review. GPT-6 Astra, Claude Fable 5.1 and Sonnet 5.5 serve as situational alternates, while GPT-6 Luna (index 37, $0.07 per task, 1,429 tasks per $100) covers classification and routing.
Three findings drive the arrangement, according to Meyer’s read of the Artificial Analysis data. First, Opus 5.5 outscores its more expensive sibling Fable 5.1 by 5 points while costing less per task. Second, Sonnet 5.5 at max effort costs more per task than Opus at max for 2 fewer points, which Meyer says makes it hard to justify at that setting. Third, GPT-6.1 Sol costs roughly one-eighth of Astra and one-twentieth of Fable per task for a score only 1 to 2 points lower.
The effort setting is identified as the largest cost lever. On Opus 5.5, moving from xhigh to max adds 2 index points and 73% more cost per task; from medium to max, cost rises 4.46x for 7 points. Meyer runs Opus at high (54 points, $1.82 per task) for development and reserves xhigh for hard problems such as architecture and migrations.
Opus builds. Sol reviews. Jev decides.
One price tape, six models
Score against cost, at every effort setting
The effort dial moves the bill more than the model
Claude Opus 5.5
Claude Sonnet 5.5
GPT-6.1 Sol: near-Astra scores at a fraction of the price
Three published settings
| Setting | Index | Cost per task | Output tokens | First token |
|---|---|---|---|---|
| medium | 48 | $0.21 | 15M | 5.3 s |
| high | 50 | $0.32 | 25M | 57 s |
| xhigh | 51 | $0.39 | 36M | 69 s |
Same score band, very different bill
My stack: who builds, who reviews
Cheaper tokens are not cheaper work
Read the numbers with four warnings
Part 2: Jev, the model that decides instead of writing
One call in, typed answers out
Three question types
Confidence is the superpower
Three uses running in my publishing operation
The fit test, then the shadow test
- Replay 300 to 500 past decisions
- Compare overall and per confidence band
- Read 20 disagreements, decide who was right
- High band at 95% or better?
- Own flag, off by default
- Canary on 5 to 10 units
- Roll out in the confident band only
24 use cases, sorted by how well they fit
Proven in production
- 1Relevance gate
- 2Language check
- 3Classifier fallback
Publishing and content
- 4Thin-source detector
- 5Same-event dedupe
- 6Product fits roundup
- 7Disclosure present
- 8Headline quality
- 9Comment moderation
Commerce and support
- 10Support-ticket routing
- 11Return-reason coding
- 12Review to feature complaints
- 13Catalogue taxonomy
- 14Order-fraud pre-triage
Software and AI systems
- 15LLM guardrail
- 16RAG passage filter
- 17Citation check
- 18Tool and intent routing
- 19Log-line triage
- 20PR risk triage
Business ops and home
- 21Inbox triage
- 22Expense categorisation
- 23Lead qualification
- 24Smart-home intent
Limits, cost and one hard rule
Why Cost Per Task Now Beats Raw Benchmarks
The piece matters because it documents a shift many practitioners are making: when top models sit within a few index points of each other, benchmark ranking alone stops driving decisions. At $0.39 per task, a review pass with Sol is cheap enough to run routinely on every meaningful change, which Meyer argues changes engineering hygiene rather than just budgets.
He also contends that a different model family reviewing Opus’s output is a better check than Opus reviewing itself, and that the low cost of Sol makes cross-model review practical. His working rules frame the limits: effort is not capability, a second model reading the same flawed spec is not independent review, passing tests are not approval to ship, and cheaper tokens do not reduce total work — he estimates halving model price saves 12.5% of real cost, which one extra minute of human review erases.
Sol’s Launch Numbers and the September Field
GPT-6.1 Sol launched at the same published prices as its week-old predecessor — $2 / $10 per 1M tokens — and Artificial Analysis already lists three effort levels: medium (index 48, $0.21 per task), high (50, $0.32) and xhigh (51, $0.39). Even the medium setting matches the earlier GPT-6 Sol’s score of 48 at one-fifth of its $1.06 per task, per Meyer’s reading.
Other published prices per 1M tokens: Opus 5.5 at $4 / $20 (cache reads $0.20), Fable and Astra at $10 / $50, and Luna at $0.10 / $0.50. Sol is also concise — the high setting used 25M output tokens on the index against a median of 82M for comparable models. Sonnet 5.5 at max wrote about 193k output tokens per task, which Meyer says is the most Artificial Analysis has measured.
“In four weeks, the AI frontier stopped being a leaderboard and became a price curve.”
— Thorsten Meyer
Caveats Behind the Stack Choices
Several limits are explicit in the source. Artificial Analysis has not yet published low or max effort settings for GPT-6.1 Sol, and Meyer notes that one index point is inside the noise. Sol’s high and xhigh settings take 57 to 69 seconds to produce a first token, making it unsuitable for interactive use at those levels.
Opus 5.5 still leads Sol by 5 points at xhigh (56 against 51), so the cheaper model is not a replacement at the top tier. All scores come from a single index version (v4.3.x), and Meyer stresses the rankings reflect his workload, not a universal verdict — the 12.5% cost-savings figure is described as illustrative rather than measured.
Watching for Sol’s Full Effort Curve
The next data points are the missing low and max effort settings for GPT-6.1 Sol once Artificial Analysis publishes them, which could shift where the model sits on the price-quality curve. Continued releases in October may compress the top of the field further, and Meyer’s own guidance is that any stack change should follow shadow testing on real workloads rather than index movements alone.
Key Questions
What is GPT-6.1 Sol, and when was it released?
GPT-6.1 Sol is a frontier model released on 29 September 2026 at $2 / $10 per 1M tokens. In Meyer’s assessment it scores 51 on the Artificial Analysis index at xhigh effort while costing $0.39 per task.
Why does Meyer use Opus 5.5 instead of Sol for building?
Opus 5.5 scores 56 at xhigh against Sol’s 51, a 5-point gap that matters for complex development work. Meyer runs Opus at high effort (54 points, $1.82 per task) for most building and xhigh for architecture and migrations.
What are GPT-6.1 Sol’s main drawbacks?
Its high and xhigh settings take 57 to 69 seconds to the first token, so it is not interactive at those levels, and Opus 5.5 still outscores it by 5 points. Low and max effort settings have not yet been published.
What does the Artificial Analysis index actually measure?
The Intelligence Index v4.3.x maps general capability across models. Meyer cautions that it is not a verdict on any specific workload and recommends shadow testing before switching models.
Is the cheapest model always the best value?
No. Meyer argues that effort settings move cost more than model choice, that benchmark gains of 1 to 2 points sit inside the noise, and that cheaper tokens do not reduce total work cost — human review time dominates savings.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
