{"id":2728,"date":"2026-07-31T06:09:45","date_gmt":"2026-07-31T06:09:45","guid":{"rendered":"https:\/\/dmarketertayeeb.com\/blog\/openai-gpt-5-6-luna-terra-price-cuts\/"},"modified":"2026-07-31T06:09:45","modified_gmt":"2026-07-31T06:09:45","slug":"openai-gpt-5-6-luna-terra-price-cuts","status":"publish","type":"post","link":"https:\/\/dmarketertayeeb.com\/blog\/openai-gpt-5-6-luna-terra-price-cuts\/","title":{"rendered":"GPT-5.6 Price Cuts: New Luna, Terra and Sol Fast Mode Costs"},"content":{"rendered":"\n<p><strong>Updated 31 July 2026:<\/strong> OpenAI has changed the economics of the GPT-5.6 family. Luna API prices are down 80%, Terra prices are down 20%, paid Codex and ChatGPT Work usage goes further when those models are used, and GPT-5.6 Sol now has a faster API option. The headline is not simply \u201cAI got cheaper.\u201d The useful question is which work should move to Luna, which still belongs on Terra, and when paying twice the standard rate for Sol Fast mode actually makes business sense.<\/p>\n\n\n\n<p>The update first surfaced in <a href=\"https:\/\/x.com\/thsottiaux\/status\/2082883636177916306\" rel=\"noopener noreferrer\">Tibo\u2019s July 30 post<\/a> and an <a href=\"https:\/\/x.com\/OpenAI\/status\/2082878156483219672\" rel=\"noopener noreferrer\">official OpenAI announcement<\/a>. OpenAI then published a detailed first-party explanation of the new rates and the engineering behind them. I checked the numbers against the company\u2019s live model documentation rather than treating a viral post as the whole story.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">The GPT-5.6 price cuts in one table<\/h2>\n\n\n\n<p>Starting 30 July, the standard API rates per one million tokens are:<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table>\n  <thead>\n    <tr><th>Model<\/th><th>Previous input<\/th><th>New input<\/th><th>Previous output<\/th><th>New output<\/th><th>Change<\/th><\/tr>\n  <\/thead>\n  <tbody>\n    <tr><td>GPT-5.6 Luna<\/td><td>$1.00<\/td><td>$0.20<\/td><td>$6.00<\/td><td>$1.20<\/td><td>80% lower<\/td><\/tr>\n    <tr><td>GPT-5.6 Terra<\/td><td>$2.50<\/td><td>$2.00<\/td><td>$15.00<\/td><td>$12.00<\/td><td>20% lower<\/td><\/tr>\n    <tr><td>GPT-5.6 Sol<\/td><td>$5.00<\/td><td>$5.00<\/td><td>$30.00<\/td><td>$30.00<\/td><td>No standard-price change<\/td><\/tr>\n  <\/tbody>\n<\/table><\/figure>\n\n\n\n<p>Those current rates are confirmed in OpenAI\u2019s live documentation for <a href=\"https:\/\/developers.openai.com\/api\/docs\/models\/gpt-5.6-luna\" rel=\"noopener noreferrer\">Luna<\/a>, <a href=\"https:\/\/developers.openai.com\/api\/docs\/models\/gpt-5.6-terra\" rel=\"noopener noreferrer\">Terra<\/a>, and <a href=\"https:\/\/developers.openai.com\/api\/docs\/models\/gpt-5.6-sol\" rel=\"noopener noreferrer\">Sol<\/a>. Cached input is currently $0.02 for Luna, $0.20 for Terra, and $0.50 for Sol per million tokens. The model pages also retain the long-context rule: prompts above 272,000 input tokens are priced at twice the input rate and 1.5 times the output rate for the full request.<\/p>\n\n\n\n<p>That last detail matters. A cheaper model does not eliminate the need to control context. The cost curve is more forgiving, but a workflow that repeatedly sends a huge history, duplicated files and verbose tool output can still waste money. My <a href=\"https:\/\/dmarketertayeeb.com\/blog\/gpt-5-6-sol-codex-usage-272k-context-pricing\">evidence-based guide to GPT-5.6 Codex usage limits<\/a> explains that context and allowance problem in depth; this article focuses on the new July 30 economics.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">What changed for Codex and ChatGPT Work subscriptions<\/h2>\n\n\n\n<p>The API price cut is only half of the announcement. OpenAI says the lower Luna and Terra prices are reflected in how usage is counted against paid subscriptions in Codex and ChatGPT Work. Subscription prices and quota budgets did not change. Instead, work completed with Luna and Terra should consume fewer credits, allowing the same paid plan to go further.<\/p>\n\n\n\n<p>That distinction prevents two common misreadings:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n  <li><strong>This is not a cheaper ChatGPT or Codex subscription.<\/strong> The monthly plan price remains the same.<\/li>\n  <li><strong>This is not a blanket quota increase for every model.<\/strong> The benefit comes from lower usage accounting when Luna or Terra handles the work.<\/li>\n<\/ul>\n\n\n\n<p>OpenAI has not published a universal \u201cmessages per model\u201d conversion table. Usage depends on context, reasoning, tools, retrieval and caching. Do not convert an 80% API price cut into a promise of exactly five times as many subscription messages. The directional conclusion is sound\u2014Luna usage should go materially further\u2014but the meter is not a simple token invoice.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">How much do representative API jobs cost now?<\/h2>\n\n\n\n<p>Simple arithmetic makes the model-routing opportunity clearer. These examples use standard rates, stay below the long-context threshold, exclude tool-specific fees, and do not assume cached input.<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table>\n  <thead>\n    <tr><th>Illustrative job<\/th><th>Luna<\/th><th>Terra<\/th><th>Sol<\/th><\/tr>\n  <\/thead>\n  <tbody>\n    <tr><td>100K input + 10K output<\/td><td>$0.032<\/td><td>$0.32<\/td><td>$0.80<\/td><\/tr>\n    <tr><td>1M input + 100K output<\/td><td>$0.32<\/td><td>$3.20<\/td><td>$8.00<\/td><\/tr>\n    <tr><td>10M input + 1M output<\/td><td>$3.20<\/td><td>$32.00<\/td><td>$80.00<\/td><\/tr>\n  <\/tbody>\n<\/table><\/figure>\n\n\n\n<p>The gap is large enough to change architecture. At the new rate, Luna can process the same token volume for one-tenth of Terra\u2019s token cost and one twenty-fifth of Sol\u2019s in these examples. That does not mean the three models produce identical outcomes. It means teams can afford to test Luna on repeatable steps that were previously routed upward by habit.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">A cost-first deployment matrix<\/h2>\n\n\n\n<p>A lower token rate matters only when a workload has a measurable quality threshold. OpenAI\u2019s <a href=\"https:\/\/openai.com\/index\/advancing-the-price-performance-frontier-with-gpt-5-6\/\" rel=\"noopener noreferrer\">price-performance announcement<\/a> recommends matching intelligence, urgency, scale and the cost of error to the outcome. The practical way to apply that principle is to put every recurring job into a cost-and-acceptance test.<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table>\n  <thead>\n    <tr><th>Workload<\/th><th>Starting model<\/th><th>Why<\/th><th>Required control<\/th><\/tr>\n  <\/thead>\n  <tbody>\n    <tr><td>Lead or search-query classification<\/td><td>Luna<\/td><td>High volume, repeatable labels, low marginal cost<\/td><td>Sample accuracy and monitor drift<\/td><\/tr>\n    <tr><td>Metadata extraction and content inventory<\/td><td>Luna<\/td><td>Structured, bounded transformation<\/td><td>Schema validation and exception queue<\/td><\/tr>\n    <tr><td>Ad-copy or headline variant generation<\/td><td>Luna or Terra<\/td><td>Many candidates benefit from inexpensive breadth<\/td><td>Brand, policy and human review<\/td><\/tr>\n    <tr><td>SEO brief synthesis across several sources<\/td><td>Terra<\/td><td>Needs more judgment and evidence integration<\/td><td>Source links and factual review<\/td><\/tr>\n    <tr><td>Campaign diagnosis with conflicting signals<\/td><td>Sol<\/td><td>Ambiguity and causal reasoning matter<\/td><td>Verified exports and explicit assumptions<\/td><\/tr>\n    <tr><td>High-stakes strategy or migration plan<\/td><td>Sol, then Luna for execution steps<\/td><td>Use deep reasoning once, cheap execution repeatedly<\/td><td>Milestones, rollback and outcome checks<\/td><\/tr>\n  <\/tbody>\n<\/table><\/figure>\n\n\n\n<p>For evergreen guidance on the strengths of each tier and how to instruct them, use DMT\u2019s <a href=\"https:\/\/dmarketertayeeb.com\/blog\/gpt-5-6-sol-terra-luna-marketers-guide\">GPT-5.6 Sol, Terra and Luna guide<\/a>. The decision here is narrower: re-price the job, run a labelled evaluation, and move it only if the cheaper configuration maintains the required acceptance rate.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Sol Fast mode: 2.5\u00d7 speed for 2\u00d7 the standard price<\/h2>\n\n\n\n<p>OpenAI is replacing Priority Processing with Fast mode in the API. For GPT-5.6 Sol, the company says Fast mode can deliver up to 2.5 times Standard processing speed at twice the Standard price, without changing model intelligence. Existing requests tagged <code>priority<\/code> remain compatible and will use Fast mode.<\/p>\n\n\n\n<p>At current Sol rates, that implies $10 per million input tokens, $1 per million cached input tokens, and $60 per million output tokens for the faster tier. The premium can be rational when latency has measurable business value:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n  <li>A live analyst or developer is blocked on an urgent investigation.<\/li>\n  <li>An interactive product would otherwise feel unresponsive.<\/li>\n  <li>A time-sensitive campaign, outage or incident has a high delay cost.<\/li>\n  <li>Shorter turnaround lets a scarce expert review more completed work.<\/li>\n<\/ul>\n\n\n\n<p>Fast mode is usually poor value for overnight research, background enrichment, bulk audits or queues with no human waiting. The model is not smarter in Fast mode. Paying the premium for a batch job merely buys earlier completion. Measure the value of saved minutes rather than assuming faster is automatically better.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Auto-review is moving to Luna\u2014and may cost about 10\u00d7 less<\/h2>\n\n\n\n<p>In a follow-up to the announcement, OpenAI said Auto-review in the ChatGPT app and Codex CLI is moving from GPT-5.4 to GPT-5.6 Luna. Combined with Luna\u2019s new price, OpenAI expects Auto-review to cost about ten times less. Tibo described this as the \u201creview for me\u201d flow, which can help prevent many high-risk actions from being handled only by the main agent.<\/p>\n\n\n\n<p>The operational lesson is bigger than one feature. Agentic systems often waste their strongest model on every step: planning, scanning, implementation, review and final explanation. A better architecture separates roles. Sol can resolve ambiguity and define a plan; Luna can inspect routine actions against policies, run checks, classify exceptions and escalate only the uncertain cases. This is the same harness principle discussed in my <a href=\"https:\/\/dmarketertayeeb.com\/blog\/ai-agent-harness-context-compaction\">analysis of GPT-5.6 agent harnesses and context compaction<\/a>.<\/p>\n\n\n\n<p>Do not treat Auto-review as a substitute for approval on consequential actions. A cheaper model can strengthen coverage because it becomes economical to review more steps, but review quality still depends on instructions, context, permissions, tool boundaries and escalation rules. Financial changes, deletion, publishing, external messages and access-control decisions still need explicit safeguards.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Why OpenAI says these price cuts were possible<\/h2>\n\n\n\n<p>OpenAI attributes the change to improvements across the model, inference stack and agent harness. The company says GPT-5.6 Sol helped rewrite and optimize production GPU kernels inside a human-led process, reducing end-to-end model-serving cost by 20%. It also reports more than 15% improvement in token-generation efficiency from speculative-decoding experiments.<\/p>\n\n\n\n<p>Those are OpenAI\u2019s own production claims, not an independent audit of every workload. They still matter because the company is passing the gains into published customer prices rather than presenting efficiency only as a benchmark. The economic signal is observable: Luna moved from $1\/$6 to $0.20\/$1.20, and Terra from $2.50\/$15 to $2\/$12 per million input\/output tokens.<\/p>\n\n\n\n<p>For marketers, the interesting trend is recursive improvement. A capable model helps engineers make inference cheaper; lower inference cost makes more agent steps viable; more usage produces more opportunities to evaluate and improve the system. The durable advantage will not come from generating five times more generic content. It will come from using the extra budget for better research coverage, QA, experimentation and measurement.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">A seven-step migration checklist<\/h2>\n\n\n\n<ol class=\"wp-block-list\">\n  <li><strong>Inventory model assignments.<\/strong> List every workflow step currently using Sol, Terra, Luna or an older model, along with volume and failure cost.<\/li>\n  <li><strong>Move repeatable volume into an evaluation.<\/strong> Test Luna on classification, extraction, monitoring and bounded implementation using a representative labelled sample.<\/li>\n  <li><strong>Preserve an escalation path.<\/strong> Route low-confidence, policy-sensitive or conflicting cases to Terra, Sol or a human reviewer.<\/li>\n  <li><strong>Update budget calculators.<\/strong> Replace the superseded Terra and Luna rates, include cached input, and preserve the >272K long-context multiplier.<\/li>\n  <li><strong>Check latency separately from intelligence.<\/strong> Benchmark Standard and Fast mode on the same Sol tasks; pay the premium only when elapsed time changes the outcome.<\/li>\n  <li><strong>Audit auto-review boundaries.<\/strong> Confirm what it reviews, what it can block, and which actions still require explicit human confirmation.<\/li>\n  <li><strong>Measure cost per accepted result.<\/strong> A model that is ten times cheaper per token is not cheaper if its work is rejected twenty times more often.<\/li>\n<\/ol>\n\n\n\n<p>Teams building broader automations should pair this checklist with a governed <a href=\"https:\/\/dmarketertayeeb.com\/blog\/openai-codex-marketers-plugins-sites-workflows\">Codex and ChatGPT Work workflow<\/a>. Keep sources, approvals, receipts and rollback paths visible. Lower token prices increase what is affordable; they do not expand the permission to act.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">What this means for SEO and digital marketing teams<\/h2>\n\n\n\n<p>The biggest opportunity is not unlimited AI content. Search performance depends on usefulness, evidence, originality and editorial judgment. Cheap Luna runs can improve the supporting work around an article: crawl and inventory pages, classify queries, detect missing metadata, check links, extract claims for verification, compare structured fields and flag inconsistencies. Terra can synthesize the evidence, and Sol can handle a genuinely difficult strategic decision.<\/p>\n\n\n\n<p>The same approach applies to paid media and analytics. Luna can label thousands of search terms or creative variants, Terra can diagnose common themes, and Sol can evaluate a complex budget tradeoff where attribution, lag and business constraints conflict. My guide to <a href=\"https:\/\/dmarketertayeeb.com\/blog\/agentic-ai-in-marketing-2026\">agentic AI in marketing<\/a> covers how to design those multi-step systems without confusing automation volume with business value.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Frequently asked questions<\/h2>\n\n\n\n<h3 class=\"wp-block-heading\">What are the new GPT-5.6 Luna prices?<\/h3>\n\n\n\n<p>Standard API pricing is $0.20 per million input tokens, $0.02 per million cached input tokens, and $1.20 per million output tokens. OpenAI says this is an 80% reduction from the previous rates.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">What are the new GPT-5.6 Terra prices?<\/h3>\n\n\n\n<p>Standard API pricing is $2 per million input tokens, $0.20 per million cached input tokens, and $12 per million output tokens. OpenAI says this is a 20% reduction.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Did GPT-5.6 Sol become cheaper?<\/h3>\n\n\n\n<p>No. Standard Sol pricing remains $5 input, $0.50 cached input, and $30 output per million tokens. The new option is Fast mode, which costs twice the Standard price for up to 2.5 times the speed.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Did Codex subscription prices go down?<\/h3>\n\n\n\n<p>No. OpenAI says ChatGPT and Codex subscription prices and quota budgets are unchanged. Luna and Terra now consume fewer credits, so usage on those models should go further.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Does an 80% Luna price cut guarantee five times more Codex messages?<\/h3>\n\n\n\n<p>No. API rates and subscription usage meters are related but not identical. Codex usage varies with model, context, reasoning, tools, retrieval and caching. OpenAI confirms more favourable usage accounting for Luna and Terra, but not a fixed message multiplier.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Should every workflow move to Luna?<\/h3>\n\n\n\n<p>No. Move repeatable, measurable, lower-risk steps first. Keep Terra or Sol where ambiguity, quality requirements or error costs justify more intelligence, and retain human approval for consequential actions.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Bottom line<\/h2>\n\n\n\n<p>OpenAI\u2019s July 30 update creates a real price-performance opportunity. Luna is now cheap enough to make high-volume agent work practical, Terra is more economical for everyday implementation and analysis, Sol Fast mode offers a clear latency premium, and Auto-review becomes a more affordable control layer. The winners will not be the teams that simply generate more output. They will be the teams that route each step to the least expensive model that can meet a tested quality standard.<\/p>\n\n\n<aside>\n  \n\n<h2 class=\"wp-block-heading\">About the author<\/h2>\n\n\n  \n\n<p>Tayeeb Khan writes Digital Marketer Tayeeb\u2019s practitioner-focused coverage of AI, SEO and digital marketing workflows. This analysis separates confirmed OpenAI product facts from implementation guidance and uses current first-party documentation for pricing claims.<\/p>\n\n\n<\/aside>\n\n\n<p><em>Editorial methodology: pricing, availability and performance claims were checked on 31 July 2026 against OpenAI\u2019s official announcement, official X thread and current model documentation. Cost examples are arithmetic illustrations using published standard token rates; they exclude tool fees, regional variations and long-context multipliers unless stated.<\/em><\/p>\n","protected":false},"excerpt":{"rendered":"<p>OpenAI\u2019s July 30 GPT-5.6 update makes Luna and Terra cheaper, adds Sol Fast mode and improves paid Codex usage economics. Here is what changes in practice.<\/p>\n","protected":false},"author":1,"featured_media":2727,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[183,180,178],"tags":[314,319,320,299,296],"class_list":["post-2728","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai-in-marketing","category-ai-news","category-artificial-intelligence","tag-ai-pricing","tag-ai-workflows","tag-codex","tag-gpt-5-6","tag-openai","has-featured-image"],"_links":{"self":[{"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/posts\/2728","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/comments?post=2728"}],"version-history":[{"count":0,"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/posts\/2728\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/media\/2727"}],"wp:attachment":[{"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/media?parent=2728"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/categories?post=2728"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/tags?post=2728"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}