Openai

OpenAI launches GPT-6 Sol and Luna, halves API prices

Two lower-cost GPT‑6 tiers roll out Sept. 22 with a lasting 50% API cut for mid/low tiers

Two lower-cost GPT‑6 tiers roll out Sept. 22 with a lasting 50% API cut for mid/low tiers

OpenAI announced two new GPT‑6 models, GPT‑6 Sol and GPT‑6 Luna, on Sept. 22, 2026 and cut API prices for those tiers roughly in half compared with GPT‑5.6 promotional rates. The company said the move is meant to make advanced models cheaper to use at scale.

Sol and Luna are positioned below GPT‑6 Astra: they reuse many of Astra’s training advances but are tuned for lower cost and faster response times. OpenAI said improvements in caching and inference efficiency let it serve these models more cheaply.

OpenAI’s published price table shows the headline numbers: GPT‑6 Sol at $2 input / $10 output per 1 million tokens and GPT‑6 Luna at $0.10 input / $0.50 output per 1 million tokens — about 50% cheaper than the matching GPT‑5.6 promotional prices. The company lists these as standard per‑token rates and shows cache and long‑context rules that affect bills.

OpenAI and its executives framed the cut as a permanent re‑baseline rather than a short promotion. A company post and a senior OpenAI tweet described the 50% reduction as a lasting change intended to expand practical use cases. That language has shaped early reactions from builders and competitors.

The new models are available in the API and are rolling into ChatGPT Work and Codex today; OpenAI said Free and Go users can try Luna in the desktop app. The company also updated model guidance and the API changelog to include the GPT‑6 Sol and Luna model IDs.

For teams building agents and high‑volume inference workloads, the change alters unit economics. A 50% token price drop reduces direct per‑call costs and raises the break‑even point for running always‑on or high‑frequency agents. Analysts and early users say that effect, not small accuracy gains, will drive much of the near‑term demand shift.

OpenAI and several independent writeups published benchmark snippets and task‑cost comparisons with the launch. Community benchmarks posted alongside the release showed Luna pulling strong cost‑per‑task results on automation workloads, while Sol offers a middle ground of quality and price. Reviewers cautioned that task design, tokenization, and orchestration patterns still determine real costs.

The timing also amplifies competitive pressure. Anthropic shipped Claude Opus 5.5 the same week with its own lower prices and larger context windows; observers described the back‑and‑forth as a near‑term price war across major providers. That rivalry is already prompting re‑tests of cost‑sensitive architectures.

Practically, builders face two linked choices: move orchestration and high‑token tasks to cheaper models like Luna, or keep mission‑critical reasoning on Astra and use Sol as an intermediate orchestrator. Early community reports show some teams experimenting with hybrid stacks that mix Astra for verification and Luna for high‑volume subtasks. Those patterns affect both latency and monthly bills.

A few wrinkles matter for budgeting. OpenAI’s pricing page flags context‑length surcharges, cached‑input and cache‑write rates, and different processing tiers (Standard, Fast, etc.) that change final bills. Some independent analyses note the comparison baseline is GPT‑5.6 promotional pricing, which OpenAI had previously marked as available at least through Nov. 21, 2026 — so short‑term comparisons should be read with that caveat.

Investors, cloud partners and third‑party vendors are watching how the price cuts affect demand for on‑prem or third‑party inference alternatives. Lower API fees make hosted, pay‑as‑you‑go models more attractive for scale tasks, but high‑end models like Astra remain the premium choice for the hardest reasoning jobs. The market response over the next quarter will show whether lower prices mainly expand usage or simply shift workloads between vendors.

OpenAI’s release is part product update and part market signal: the company says efficiency gains underpin the cuts, and competitors immediately countered with lower rates of their own. For agent builders and high‑volume customers, the practical effect is immediate — cheaper per‑token bills and new options for architecture design — but the longer‑term landscape will be decided by performance tradeoffs, real task costs, and how providers price advanced features.