Model launches

AI-generated text

Practical guidance for using the GPT‑6 model family efficiently

This guide outlines practical recommendations for selecting and operating GPT‑6 models—Astra, GPT‑6.1 Sol, and GPT‑6 Luna—so teams can balance capability, latency, and cost.

Practical guidance for using the GPT‑6 model family efficiently

This practical guide explains how to choose and operate GPT‑6 family models—balancing capability, latency, and cost—while handling prompt design, long contexts, long‑running workflows, and pre‑deployment checks.

Which model to pick

The GPT‑6 family offers different models for different needs:

  • GPT‑6 Astra: for the hardest reasoning problems requiring maximum intelligence.
  • GPT‑6.1 Sol: for complex coding, research, and computer interaction tasks.
  • GPT‑6 Luna: for focused, repeatable work at scale (e.g., extracting invoice fields, classifying requests, producing structured summaries).

When selecting a model, compare each model’s pricing and consider the appropriate reasoning effort level to balance intelligence and cost.

Reasoning effort and speed settings

In the API you can set how much effort the model expends on a task:

  • Low: routine tasks such as fact extraction or small edits.
  • Medium: judgment‑requiring work like feature planning or comparing options.
  • High: difficult debugging, deeper analysis, careful review.
  • Extra high / Max: test when High is insufficient; keep only if the improvement justifies the extra time and cost.

You can change reasoning effort mid‑conversation without breaking cache. Use Fast mode when response time matters (e.g., chat apps or coding tools); it gives faster, more consistent responses at a higher per‑token cost. Use Ultrafast when the fastest token generation is worth the premium (available for GPT‑6 Astra), for example for rapid coding iterations.

Prompts, skills, and decision boundaries

Start with a clear assignment: the desired result, the audience, relevant context and constraints, and what counts as done.

Recommendations:

  • Keep skill descriptions short and explicit about when they should run.
  • Load supporting details only when needed and replace rigid recipes with guidance suited to your models.
  • Update AGENTS.md to explain when documents and tests are relevant and explicitly authorize safe routine workflows (e.g., local tests with disposable data and no production access).
  • Define decision boundaries: state which actions can proceed independently and which require approval.
  • Be prescriptive about persistence: specify what “done” includes (implementing the change, running it, inspecting results, fixing failures) and which decisions require review.

Eric Provencher of OpenAI’s Developer Experience team notes that models better understand nuance and ambiguity, so overly specific guidance can sometimes hinder results.

Managing long‑running work

GPT‑6 can handle tasks spanning hours or days. To keep such workflows on track, use the API features below:

  • Steering: mid‑turn updates can be sent via the Responses WebSocket API; updates are queued and do not cancel running tools or undo completed actions.
  • Asynchronous tools: let the model continue independent work while your app runs slower tasks (e.g., tests); the app returns results when ready.
  • Delegation and multi‑agent workflows: GPT‑6.1 Sol supports assigning independent subtasks to subagents and combining their findings into a final response (multi‑agent currently in beta).

Further practices:

  • Allow the model to ask clarifying questions that affect the next step; specify which work can continue while you answer.
  • If you’ll be away, tell Codex which tasks can proceed and when to pause for your input.
  • Redirect active work when requirements change by steering the task with clear instructions on what to change and what to keep.

Computer use and tool selection

Computer use enables GPT‑6 Astra, GPT‑6.1 Sol, and GPT‑6 Luna to interact directly with websites and desktop apps, including applications without an API. For example, the model can investigate a bug, fix code, and open the product in a browser to verify the fix.

Choose the simplest reliable way to do each step:

  • Use an API or connected tool when it can perform the job directly.
  • Use computer use when the model needs to read a screen, click buttons, or fill forms.
  • When building computer use into your app, provide a tool that can run code to control a browser or desktop (e.g., Playwright for browsers, PyAutoGUI for desktop apps).

Efficiency, caching, and compaction

To manage cost and context size:

  • Cut context that the task doesn’t need while keeping necessary evidence.
  • Run independent tasks together where supported so one slow step doesn’t block unrelated work.
  • Reuse shared context with prompt caching for recurring work; cached input tokens can cost up to 95% less than uncached input tokens, depending on the model.
  • Put stable instructions and reference material before changing task details and keep tool definitions consistent.
  • Use the caching dashboard and diagnostics guide to spot where reuse breaks down. Include cache writes and any long‑context rates when estimating end‑to‑end workflow cost.

For longer conversations, compaction reduces context size while preserving the state needed to continue.

Monitoring, testing, and pre‑deployment checks

Before deploying, run representative tasks and measure task success, latency, and cost per successful task. Decide how you will monitor model behavior and review the data controls for your application. Use the API deployment checklist and diagnostics guides to prepare for production.

Summary

The GPT‑6 family supports complex, multi‑step, long‑running workflows, but getting the best results requires deliberate choices: pick the right model and reasoning level, design concise and consistent prompts and skills, apply caching and compaction, and use steering, async tools, and delegation for long tasks. Measure success, latency, and cost before deployment and put monitoring and data controls in place.