🔍 Read the full analysis: Half-Price GPT‑6 Sol And Luna But Benchmark Scores Are Unaffected on ThorstenMeyerAI.com
Get business pricing on monitors, keyboards and dev gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
OpenAI has released GPT-6 Sol and Luna at 50% lower prices than GPT‑5.6, with benchmark scores largely unaffected. Cost savings could significantly expand AI use cases across industries.
OpenAI has introduced two new models, GPT‑6 Sol and Luna, priced at approximately 50% less than their GPT‑5.6 predecessors, without any notable change in benchmark scores. This move aims to make AI more accessible for a broader range of applications by drastically reducing operational costs, a shift that could influence how businesses deploy large language models.
The new models, GPT‑6 Sol and Luna, arrived on September 22, 2026, with prices cut by half: GPT‑6 Sol now costs $2.00 per 1 million input tokens and $10.00 per 1 million output tokens, down from $4 and $20 respectively. GPT‑6 Luna is priced at $0.10 for input and $0.50 for output tokens, compared to previous prices of $0.20 and $1.20. These reductions are enabled by improvements in caching and inference efficiency, with OpenAI passing the savings directly to users.
Despite the lower prices, independent analysis by Artificial Analysis indicates that benchmark scores for these models remain roughly the same as their predecessors. GPT‑6 Sol scores 48 on the Artificial Analysis Intelligence Index, well above the median of 25 for comparable models, with a 872,000-token context window. Luna scores 37, also above the median, with a 1 million-token context window. In coding tasks, Sol scores 57 on the Coding Agent Index, slightly above its predecessor, while Luna scores 41, slightly below. The models show improvements in hallucination reduction, with Sol decreasing hallucination rates from 92% to 60%, and Luna from 93% to 77%, primarily by declining to answer more often, which reduces errors but also lowers overall accuracy.
However, some evaluations reveal regressions in knowledge-based tasks, with both models losing Elo points in economic and knowledge work assessments, attributed to reduced presentation quality and omitted details in their outputs. This suggests a trade-off between cost, response length, and completeness of information, which could impact workflows requiring detailed, well-structured outputs.
GPT‑6 Sol and Luna: half the price, about the same intelligence
OpenAI’s September 22, 2026 release doesn’t raise the ceiling. It lowers the cost of everything below it, which changes what’s worth automating.
Per 1M input / output tokens. Cached input reads keep the 90% discount.
Cost per task, halved
Measured by Artificial Analysis as the weighted cost of one Intelligence Index task, at max effort.
The effort dial moves cost more than the model choice
| Model and effort | Intelligence Index | Cost per task |
|---|---|---|
| GPT‑6 Sol (max) | 48 | $1.06 |
| GPT‑6 Sol (low) | 34 | $0.13 |
| GPT‑6 Luna (max) | 37 | $0.07 |
| GPT‑6 Luna (low) | 21 | $0.0045 |
| GPT‑6 Luna (non‑reasoning) | 18 | $0.01 |
Sol at low effort keeps about 70% of its max score for roughly an eighth of the cost, because it writes far fewer reasoning tokens. For reference, Claude Opus 5.5 leads the same index at 58.
What got better, and what got worse
Better
- Hallucination rate on AA‑Omniscience: Sol 92% → 60%, Luna 93% → 77%
- Coding Agent Index: Sol 57, up 2 points, at ~50% lower cost per task
- OpenAI reports about half as many factual mistakes for Sol as its predecessor
- Higher cache hit rates; GitHub reports over 50% fewer prompt tokens needing fresh processing
Sol gets there partly by declining more: it attempts 83% of questions vs 99%, and accuracy falls 59% → 54%.
Worse
- GDPval‑AA v2.1: Sol down ~100 Elo, Luna down ~75
- AA‑Briefcase v1.1: Luna down ~45 Elo
- Coding Agent Index: Luna 41, down 2 points
- Both models write more output tokens per task than their predecessors
Reviewers attribute the drops to weaker presentation and deliverables that omit required elements.
What to do about it
Impact of Cost Reduction on AI Deployment Strategies
The release of GPT‑6 Sol and Luna at half the previous cost could significantly lower barriers for integrating large language models into business operations, research, and customer service. The unchanged benchmark scores imply that organizations can now access models with similar capabilities but at a fraction of the cost, potentially enabling new applications and scaling existing ones more affordably. This shift may accelerate AI adoption across industries, especially for tasks where cost was previously prohibitive.
As an affiliate, we earn on qualifying purchases.
Background on GPT Model Pricing and Performance
OpenAI’s previous models, including GPT‑5.6, were priced higher, limiting their use to enterprise-level applications or high-budget projects. The company has emphasized that improvements in caching and inference efficiency have allowed for substantial cost reductions without sacrificing performance. The September 2026 release expands on this approach by offering models that maintain competitive benchmark scores while dramatically lowering operational expenses, aligning with industry trends toward more accessible AI tools.
Independent evaluations, such as those by Artificial Analysis, have been instrumental in confirming that these models’ performance remains stable despite the price cuts. While some assessments noted regressions in specific knowledge tasks, the overall picture suggests a strategic focus on reducing costs while maintaining core capabilities.
large language model development tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Remaining Questions About Model Capabilities and Use Cases
It is not yet clear how these models will perform in real-world, high-stakes applications beyond benchmark tests, especially in complex reasoning or nuanced tasks. The impact of reduced presentation quality on enterprise workflows that rely on detailed outputs remains to be fully assessed. Additionally, the long-term effects of lower hallucination rates and potential regressions in knowledge tasks are still being studied, and further independent evaluations are awaited.
As an affiliate, we earn on qualifying purchases.
Next Steps for Adoption and Performance Verification
OpenAI is expected to continue refining these models and releasing updates that address the identified regressions while maintaining cost efficiency. Organizations should conduct their own testing before fully deploying GPT‑6 Sol and Luna in critical workflows. Industry analysts anticipate that further independent benchmarks and user feedback will clarify how these models perform in diverse applications, shaping future AI deployment strategies.
As an affiliate, we earn on qualifying purchases.
Key Questions
How much cheaper are GPT‑6 Sol and Luna compared to GPT‑5.6?
GPT‑6 Sol is approximately 50% less expensive per 1 million tokens, with input costs dropping from $4 to $2 and output from $20 to $10. Luna’s costs are halved as well, from $0.20 to $0.10 for input and from $1.20 to $0.50 for output tokens.
Do the new models perform worse on benchmark tests?
According to independent evaluations, benchmark scores remain roughly the same, with some metrics showing slight improvements and others indicating regressions, particularly in knowledge-related tasks.
What are the main improvements in GPT‑6 Sol and Luna?
They exhibit significantly reduced hallucination rates, better cost efficiency, and comparable or slightly improved coding performance, mainly achieved through caching and inference optimizations.
Are there any downsides to the new models?
Some assessments report regressions in the quality and completeness of outputs, especially in detailed knowledge tasks, which could affect workflows requiring comprehensive deliverables.
When will more real-world testing results be available?
Further independent evaluations and user feedback are expected in the coming months, which will help clarify the models’ suitability for various enterprise and research applications.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
