🔍 Read the full analysis: How Opus, Sol, And Jev Shape My September 2026 AI Stack on ThorstenMeyerAI.com
Get movie nights delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
A Sept. 29, 2026 article describes a personal AI workflow built around Opus 5.5 for development, GPT-6.1 Sol for detailed review and Jev for high-volume decisions. Its benchmark data puts Sol far below several alternatives in reported task cost, but the figures come from one index and the author says teams should test models on their own workloads.
Meyer bases the comparison largely on the Artificial Analysis Intelligence Index v4.3.x, which he describes as a general capability measure rather than a verdict on any particular workload. In its listed top settings, Opus 5.5 scores 58 and costs $5.98 per task; GPT-6.1 Sol at xhigh scores 51 and costs $0.39. The article lists GPT-6 Luna at 37 and $0.07 per task, with other models between those results.
The workflow assigns Opus 5.5 at high effort to feature work, APIs, refactors and multi-file changes. Meyer says high delivers an index score of 54 at $1.82 per task, while xhigh is reserved for more difficult architecture, migration and trust-boundary work. He reports that xhigh scores 56 at $3.46 per task, and max scores 58 at $5.98.
GPT-6.1 Sol, released Sept. 29 according to the article, is used for examining specific files or diffs and reviewing Opus’s output. At high, the index lists Sol at 50 and $0.32 per task; xhigh scores 51 at $0.39. Meyer assigns Sonnet 5.5 and Luna to scoped side work and classification. He says Jev, which he describes as a decision model that cannot write sentences, handles high-volume routing and yes-or-no judgments; the source provides no benchmark figures for Jev.
Opus builds. Sol reviews. Jev decides.
One price tape, six models
Score against cost, at every effort setting
The effort dial moves the bill more than the model
Claude Opus 5.5
Claude Sonnet 5.5
GPT-6.1 Sol: near-Astra scores at a fraction of the price
Three published settings
| Setting | Index | Cost per task | Output tokens | First token |
|---|---|---|---|---|
| medium | 48 | $0.21 | 15M | 5.3 s |
| high | 50 | $0.32 | 25M | 57 s |
| xhigh | 51 | $0.39 | 36M | 69 s |
Same score band, very different bill
My stack: who builds, who reviews
Cheaper tokens are not cheaper work
Read the numbers with four warnings
Part 2: Jev, the model that decides instead of writing
One call in, typed answers out
Three question types
Confidence is the superpower
Three uses running in my publishing operation
The fit test, then the shadow test
- Replay 300 to 500 past decisions
- Compare overall and per confidence band
- Read 20 disagreements, decide who was right
- High band at 95% or better?
- Own flag, off by default
- Canary on 5 to 10 units
- Roll out in the confident band only
24 use cases, sorted by how well they fit
Proven in production
- 1Relevance gate
- 2Language check
- 3Classifier fallback
Publishing and content
- 4Thin-source detector
- 5Same-event dedupe
- 6Product fits roundup
- 7Disclosure present
- 8Headline quality
- 9Comment moderation
Commerce and support
- 10Support-ticket routing
- 11Return-reason coding
- 12Review to feature complaints
- 13Catalogue taxonomy
- 14Order-fraud pre-triage
Software and AI systems
- 15LLM guardrail
- 16RAG passage filter
- 17Citation check
- 18Tool and intent routing
- 19Log-line triage
- 20PR risk triage
Business ops and home
- 21Inbox triage
- 22Expense categorisation
- 23Lead qualification
- 24Smart-home intent
Limits, cost and one hard rule
A Workflow Built Around Task Cost
The account illustrates a practical shift in model selection: a small benchmark-score gap may matter less than the price of repeatedly running a model. Meyer says Sol’s review cost is low enough for him to use it on every meaningful change, while keeping Opus on work where he values its higher score. That is his reported practice, not evidence that the same division will suit every team.
His figures also draw attention to the price of effort settings. The article reports that Opus moves from $1.34 per task at medium to $5.98 at max, while its index score rises from 51 to 58. Meyer argues that human review time can erase savings from cheaper model calls, but the supplied source cuts off during its example and gives no measured result for that claim. Readers should treat it as a caution, not a quantified finding.
Benchmarks Behind the Model Split
The source says six models sit within roughly 20 points on the cited index while their listed per-task costs differ substantially. It presents Opus 5.5 as the top scorer among those models, and Sol and Luna as lower-cost options for selected work. Those comparisons depend on the index’s methods and settings; the article cautions readers to shadow-test models before switching.
The author also compares effort levels, which change both scores and costs. For Sonnet 5.5, the source lists high at 47 and $1.08 per task, versus max at 56 and $7.60. It says max produced about 193,000 output tokens per task on the index. These are reported index measurements, not guaranteed costs or performance for other prompts, tools or workloads.
“The index is a map of general capability, not a verdict on your workload, so shadow-test before you switch anything.”
— Thorsten Meyer, in the Sept. 29 article
Limits of the Published Comparison
The source does not specify how the index’s per-task costs map to an individual organization’s actual bills, including its prompts, usage patterns or human review time. It also says the index had not yet published GPT-6.1 Sol’s low or max settings. Meyer notes that a one-point score gap is within the noise, so small differences should not be read as definitive.
Jev is not independently benchmarked in the supplied material, and no test results are given to show how accurately it routes decisions. The source text also ends partway through an illustrative comparison about model prices and human review. It does not provide enough information to establish that example’s outcome.
Test Models Against Your Work
Meyer’s stated next step for readers is to shadow-test models on their own tasks before changing a workflow. That means comparing outputs and costs under the same requirements, then checking where human review time affects the total. The article does not announce a formal follow-up test or a date for updated benchmark figures.
Further comparison will depend on new index results for settings not yet listed and on workload-specific tests. Until those are available, the article’s stack remains one practitioner’s Sept. 29 snapshot, with Jev’s routing role described but not quantified.
Key Questions
What AI stack does Thorsten Meyer describe?
He says he uses Opus 5.5 to build, GPT-6.1 Sol for detailed work and review, and Jev for high-volume yes-or-no and routing decisions. He names Sonnet 5.5, Luna, Astra and Fable as alternatives for selected tasks.
Why does Meyer use GPT-6.1 Sol for review?
He cites its reported cost of $0.32 per task at high and $0.39 at xhigh, which he says makes routine review affordable. Those figures are from the Artificial Analysis index cited in the article.
Does the article establish that Opus 5.5 is best for every developer?
No. It reports Meyer’s own workflow and index comparisons. The article says the index measures general capability and advises readers to test models on their own workloads.
What remains unknown about Jev?
The article describes Jev as a decision model for routing and yes-or-no judgments, but supplies no benchmark scores, task costs or accuracy results for it.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
