Launch status, checked July 24, 2026: OpenAI launched ChatGPT Voice in Work and Codex on July 23. In the ChatGPT desktop app, eligible users can now speak to an agent, interrupt it naturally, start or redirect tasks, ask for progress, and coordinate work across multiple agents. The conversation is powered by GPT-Live, while the actual task continues to use the tools, permissions, model access, and usage pool available to Work or Codex.
This is more than speech-to-text, but it is not a magic override for your computer. Voice is a live control layer over the selected agentic experience. It does not bypass approvals, invent access to unavailable apps, or make Codex selectable as a standalone product on the web or mobile. OpenAI currently documents the full experience for the desktop app on macOS and Windows, with paired iOS remote access.
Early public discussion has mixed genuine use cases with understandable confusion about rollout timing, plan eligibility, mobile support, API access, and cost. This guide separates OpenAI’s confirmed product facts from user reports and my practical recommendations. Community posts are used only to identify questions people are asking, not to establish how the product works.
Codex Voice mode in one table
| Question | Confirmed answer as of July 24, 2026 |
|---|---|
| What launched? | ChatGPT Voice inside Work and Codex in the ChatGPT desktop app. |
| What can it do? | Start, prioritize, interrupt, or redirect tasks; coordinate multiple agents; resume supported project work; and provide spoken or on-screen progress updates. |
| Which desktop platforms? | macOS and Windows. |
| Is it available on the web? | No. Voice in Work and Codex is not a standalone web experience. |
| Is it available on mobile? | Not as a standalone Work or Codex Voice experience. OpenAI documents paired iOS remote access to supported desktop work. |
| Is everyone getting it at once? | No. Availability can depend on plan, region, workspace settings, app version, and gradual rollout. |
| Does it use a separate meter? | Connected Voice time is metered separately where flexible pricing applies. Delegated agent work still draws from the shared Work and Codex usage pool. |
| Is there an API? | The July 23 launch is a ChatGPT desktop product feature. OpenAI still describes GPT-Live API access as planned and offers a notification form. |
What OpenAI actually launched
OpenAI’s ChatGPT release notes describe the change simply: start a new task in Voice, speak naturally, interrupt, and ask Voice to start or coordinate work using the tools and permissions available to Work or Codex. The more detailed ChatGPT Voice documentation expands that into five confirmed functions:
- Start, prioritize, interrupt, or redirect tasks while work continues in the background.
- Coordinate multiple agents across active conversations and projects.
- Resume existing work using available project context and supported connected tools.
- Receive spoken or on-screen updates when work is progressing, blocked, or complete.
- See when ChatGPT is listening and mute or stop the microphone.
The key phrase is using the tools and permissions available to the selected experience. If Codex has access to a repository and terminal but not a calendar, Voice does not manufacture calendar access. If a workspace requires approval before a tool can send, publish, delete, or change production, speaking the instruction does not remove that approval. The interface changed; the permission model did not disappear.
This product release builds on OpenAI’s July 8 GPT-Live launch. GPT-Live uses a full-duplex architecture, which means it can listen and speak at the same time instead of waiting for rigid turns. It can decide repeatedly whether to keep listening, respond, pause, interrupt, or invoke a tool. For deeper work, it can delegate to a frontier model while preserving the live conversation.

The demo is useful for understanding the interaction pattern: the user explains an outcome conversationally, the agent continues working in the desktop app, and the user can interrupt or redirect it. The video stays on OpenAI’s official account; DMT hosts only the attributed launch poster so the page remains compatible with its security policy. Treat the footage as a product demonstration, not a promise that every shown tool, permission, model, or workspace configuration is available on every account.
This article originally focused on GPT-Live’s model architecture and marketing implications. The July 23 update is the practical next step: the voice layer is no longer limited to conversation. It can now direct longer-running work in Work and Codex. The sections below preserve the useful architecture and strategy context while adding current setup, pricing, troubleshooting, privacy, and safety guidance.
Codex Voice mode versus dictation, Chat Voice, and remote control
Much of the launch confusion comes from using the word “voice” for four different input patterns. They are related, but they are not interchangeable.
| Experience | Primary job | What happens after you speak | Best use |
|---|---|---|---|
| Dictation | Convert speech into an editable prompt | You review or edit the transcript, then send it as text | Precise prompts that benefit from speaking but need a final written check |
| Voice in Chat | Have a live conversation | ChatGPT responds conversationally using the features available in Chat | Questions, brainstorming, explanation, language practice, and everyday help |
| Voice in Work or Codex | Direct agentic work through live conversation | Work or Codex can start and coordinate tasks using its existing context, tools, and permissions | Longer tasks, progress checks, redirection, multi-agent coordination, and hands-free supervision |
| Paired iOS remote access | Stay connected to supported desktop work | The desktop host remains the execution environment while the phone provides a remote control surface | Checking progress or steering work away from the computer |
Use dictation when the wording itself is the deliverable. Use Voice when the conversation is part of the workflow. If you need to specify a filename, exact command, regular expression, URL slug, or irreversible production action, a hybrid pattern is usually safer: explain the goal by voice, then confirm the exact string or approval in text.
How to start Voice in Work or Codex
OpenAI’s basic setup path is short, but a reliable setup has three layers: choose the right experience, confirm account and workspace access, and grant only the operating-system permissions the task actually needs.
Step 1: choose Work or Codex before turning on Voice
| Choose | When it is the better starting point | Typical context and tools |
|---|---|---|
| Work | Research, analysis, documents, spreadsheets, presentations, reports, Sites, and work across connected apps or files | Projects, uploaded or local files where allowed, connected apps, browser-based research, document and data tools |
| Codex | Software development, debugging, tests, commands, repository work, code review, or technical implementation | Local folders, repositories, terminals, diffs, developer tools, and the permissions configured for the coding workspace |
Voice does not turn Work into Codex or Codex into Work. It gives the selected experience a live conversational control surface. If the desired outcome is a market-research report, begin in Work. If the outcome is a tested code change, begin in Codex. If the task crosses both domains, decide which artifact is primary and hand off a written brief rather than assuming one conversation automatically transfers every context and permission.
Step 2: follow the desktop setup sequence
- Update and open the ChatGPT desktop app on macOS or Windows.
- Sign in to the account and workspace that should own the work. A personal account and a managed company workspace can expose different controls.
- Choose Work or Codex. In the current desktop layout, ChatGPT contains Chat and Work, while Codex remains a separate top-level view.
- Start a new task, open an existing Project, or open a local folder only when that context is genuinely required.
- Select the Voice control, or use a shortcut you configured in Settings. OpenAI does not document a universal default shortcut.
- Allow microphone access. Depending on the task and operating system, computer context may also require Screen and Audio Recording or Accessibility permissions.
- State the outcome, scope, constraints, review criteria, and approval boundaries before asking the agent to act.
The official Work and Codex guide notes that Voice uses whichever tools and permissions belong to the selected experience. That means setup is partly an access-control exercise. A microphone permission starts the conversation; it does not grant the agent access to every file, app, account, or action on the computer.
Step 3: configure macOS or Windows permissions deliberately
Start with the minimum permission set. Microphone access is required for speech. Screen and Audio Recording can be relevant when the agent must understand visible computer context. Accessibility permission can be required for supported computer interactions. The exact prompt and settings location can vary by operating-system version and app build, so use the permission request shown by the desktop app rather than copying an old path from a forum post.
- Microphone: required for the live conversation. Confirm the correct input device if the level meter does not respond.
- Screen and Audio Recording: grant only when the task needs visible or audible computer context. It is not required merely to discuss a repository or uploaded document.
- Accessibility or computer control: reserve for tasks that actually need supported interaction with desktop apps. Remove or disable it later if it is no longer necessary.
- Folder and repository access: open the narrowest useful directory. Do not expose an entire drive when one project folder is enough.
- Connected apps and network access: managed separately by the selected experience and workspace. Voice cannot silently expand them.
Personal, Business, and Enterprise setup differences
| Account situation | What to verify | Common reason setup stops |
|---|---|---|
| Personal eligible plan | Latest desktop app, correct account, rollout eligibility, microphone permission, and the Voice control inside Work or Codex | The feature is still reaching the account, or the user is looking in ordinary Chat, web, or mobile instead of the desktop experience |
| Business workspace | Member role, Work or Codex availability, app permissions, model access, and the workspace’s usage or credit policy | A workspace setting differs from the user’s personal workspace, or the team has restricted tools and actions |
| Enterprise or Edu | Role-based Work access, Codex Local access, available models, enforced requirements, and any Voice or early-access controls shown to admins | The member’s custom role or workspace policy disables one required capability even though another is enabled |
OpenAI’s current admin guidance says Work and Codex Local are controlled separately. Owners and admins can manage Work through Workspace settings and Permissions & roles, while model defaults and available models remain separate concerns. This separation explains why a member can see one experience but not another, or can start Voice yet still be unable to use a particular tool.
A ten-minute verification checklist
- Listening: say a short sentence and confirm that the interface shows it is listening.
- Interruption: interrupt a response with a correction and confirm the agent changes direction.
- Context: ask it to name the current project or folder without making changes.
- Permission boundary: ask what tools and write permissions are available in this task.
- Read-only action: request an inventory, summary, test diagnosis, or plan that does not mutate anything.
- Progress update: ask for a short spoken update only at the next milestone or blocker.
- Stop control: tell it to stop before a write and confirm that it waits for approval.
- Evidence: ask for a written result containing links, diffs, test output, or another verifiable artifact.
A safer first test
Do not make the first session a production deployment, a mass email, or a destructive refactor. Pick a reversible task with an observable result:
“Review this project read-only. Tell me which tests are failing, propose a repair plan, and check in before editing files. If you need a command that writes, installs, deletes, sends, publishes, or changes production, stop and ask for approval. Give me a short spoken update after each major stage.”
This prompt works because it defines the objective, permission boundary, stop conditions, and update cadence. Voice makes it easier to steer an agent, but the quality of the operating contract still determines whether the work stays controlled. DMT’s GPT-5.6 prompting and workflow guide explains the same principle for written agent instructions.
Why you may not see Codex Voice yet
OpenAI says Live availability depends on plan, region, workspace, and app version, and that rollout is gradual. The official OpenAI announcement on X described a global rollout beginning July 23; “rolling out” does not mean every account receives the control at the same minute.
If the Voice control is missing, check these items in order:
- Update the desktop app. An older build may not contain the new experience.
- Confirm the surface. The launch is for Work and Codex in the desktop app, not a standalone Codex page on the web or mobile.
- Check account eligibility. Availability can differ by plan, region, and rollout stage.
- Check workspace controls. Managed workspaces can restrict Work, Codex Local, Voice, models, apps, and actions independently.
- For Enterprise early access, check both settings. OpenAI says Enterprise workspaces initially require Advanced voice capabilities and Early Model Access.
- Allow rollout time. A missing control during the first day is not proof that the announcement is false or that the feature has been withdrawn.
A July 23 r/codex discussion contains users reporting that the control appeared at different times, disappeared temporarily, or was still missing. Those reports support the theme that rollout visibility is confusing; they do not establish a universal plan rule or outage.
How pricing and usage actually work
The launch has two meters that are easy to collapse into one:
- Connected Voice time: the live audio connection can be metered separately where flexible pricing applies.
- Delegated agent work: tasks started through Voice draw from the same shared agentic usage and credit pool used by Work and Codex.
For ChatGPT Business and Enterprise customers using credits or pay-as-you-go billing, OpenAI’s current Codex rate card estimates connected desktop Voice at approximately six credits per minute. The underlying task is charged separately at the normal Work or Codex rate. Legacy Enterprise workspaces have a different included allowance, and workspace configuration can affect what members see.
For personal plans, do not copy the ordinary Chat Voice cap into a Codex cost estimate. OpenAI explicitly says ordinary Chat Voice uses separate caps, while Voice inside Work or Codex uses the agentic pool for task work and can have a separate connected-time meter. The safest current answer is to check the Usage panel and plan-specific banner in the app before relying on a long session. DMT’s evidence-based Codex usage guide explains why model choice, context size, task length, output volume, and parallel agents can change the task cost substantially.
A simple cost-control pattern
- Use Voice to clarify, prioritize, and redirect, not to fill silence.
- Ask for updates at milestones instead of a continuous narration.
- Set a maximum number of parallel agents unless the work clearly benefits from more.
- Ask the agent to stop when the acceptance checks pass or when a named blocker appears.
- End the Voice connection when you no longer need live coordination.
Community requests for more usage transparency are reasonable, but a Reddit estimate is not a rate card. Use OpenAI’s live documentation and your account’s Usage panel for financial decisions.
Does Codex Voice work in the API, CLI, or IDE?
OpenAI’s July 23 announcement is specifically about the ChatGPT desktop app. The company does not say that the same live Voice control has launched in the Codex CLI, IDE extension, or as a public developer API.
The underlying GPT-Live models are a separate question. In its July 8 launch post, OpenAI said it plans to bring GPT-Live-1 and GPT-Live-1 mini to the API and provided a notification form. As of this fact check on July 24, that wording still describes API access as forthcoming. A developer can build voice applications today with other OpenAI realtime tools, but that does not make the new desktop Work/Codex Voice experience an API product.
The practical distinction is:
- Desktop Voice: a ready-made interface for steering Work or Codex.
- Codex CLI and IDE: text-first development clients with their own commands and workflows.
- Future GPT-Live API: a developer building block OpenAI says is planned, not a confirmed part of the July 23 desktop release.
What early users are asking – and what is actually known
The current public reaction is unusually useful because the questions repeat. It is not useful as a substitute for documentation. I reviewed the main launch thread in r/codex, an earlier r/OpenAI discussion asking why Chat Voice could not control a computer, and the official announcement trail. Five themes stand out.
1. “Why do I not have it?”
This is the most common practical question. Officially, the answer is gradual rollout plus plan, region, workspace, and app-version differences. User reports that a control appeared late or temporarily disappeared are real reports, but they are not enough to diagnose another account.
2. “Is this just dictation?”
No. Dictation produces text for review and sending. Voice keeps a live conversation open and can direct agentic work. The distinction matters for cost, accuracy, and control: a dictation transcript can be edited before anything runs, while a live instruction may begin a task immediately within the current permission boundary.
3. “Why is this better than typing?”
The honest answer is that it is not always better. Voice is valuable when a person needs to explain an ambiguous goal, think aloud, supervise work away from the keyboard, reduce physical input, or redirect several agents without opening each thread. Typing is better for exact strings, code, structured specifications, quiet shared spaces, and anything that requires a durable verbatim record.
4. “Will it burn my limits?”
It can consume both connected Voice time and normal agentic task usage. The exact effect depends on plan and task. OpenAI publishes an approximate six-credit connected-minute rate for credit-based Business and Enterprise customers, but a long multi-agent task can also consume the normal shared Work/Codex pool.
5. “Can I use it everywhere?”
Not yet. The launch is desktop-first. Standalone Voice in Work and Codex is not available on web or mobile, although paired iOS remote access is supported. OpenAI has not announced the new experience as a Codex CLI, IDE, or GPT-Live API launch.
One older r/OpenAI thread asked why GPT-Live could not talk to Codex and use computer tools. The July 23 launch directly addresses that product gap. It does not answer every request for Android parity, API access, or universal plan availability.
Who Codex Voice is for: roles and realistic use cases
The strongest use cases share one property: the work benefits from ongoing steering. Voice is less valuable when the input is a single exact prompt and more valuable when the user must explain context, correct assumptions, choose among branches, or monitor a task that takes longer than one response.
| User | High-value use case | Deliverable to require | Boundary to keep |
|---|---|---|---|
| Developer | Debug a failing test suite, explain an unfamiliar repository, or supervise a refactor | Plan, diff, test results, and a list of unresolved risks | Confirm exact commands, dependency installs, migrations, and production actions in text |
| Engineering lead | Coordinate parallel investigations or review implementation tradeoffs | Task ownership, evidence from each branch, integration decision, and final status | Do not let multiple agents edit overlapping files without a merge rule |
| Marketer or SEO practitioner | Research an update, audit content, compare campaign evidence, or prepare a source-backed brief | Source ledger, factual/interpretive split, recommendations, and approval-ready draft | Do not auto-publish, change budgets, or send messages without an explicit gate |
| Analyst | Explore a dataset, test a hypothesis, or build an explanatory report | Definitions, validation checks, reproducible query or workbook, and uncertainty notes | Do not treat a conversational summary as a substitute for data-quality validation |
| Product manager or designer | Critique a flow, compare alternatives, or turn feedback into an implementation brief | Decision criteria, prioritized issues, screenshots or references, and acceptance criteria | Keep subjective preference separate from verified usability evidence |
| Founder or operator | Coordinate a cross-functional deliverable while away from the keyboard | Written plan, owners, dependencies, decisions, and completion receipts | Keep financial, legal, hiring, customer, and production actions behind explicit approval |
| Accessibility or hands-busy user | Reduce physical input while researching, reviewing, or supervising work | Visible text, captions, transcript, and a keyboard-accessible fallback | Voice should be an additional interface, not the only way to review or stop work |
1. Plan a change before anyone edits
Architecture, research, campaign planning, and incident response often begin with partial information. Voice helps a user explain the situation, hear the agent’s interpretation, and correct it before granting write access.
“Inspect the project read-only. Explain the three most likely causes, tell me what evidence would distinguish them, and propose the smallest safe test. Do not edit or install anything yet.”
Acceptance evidence: a written plan tied to files, logs, sources, or data—not merely a confident spoken recommendation.
2. Debug a failing build or test suite
A developer can describe what changed, ask Codex to reproduce the failure, and interrupt when the investigation follows an implausible path. This is especially useful when the user is reading logs or comparing output on another screen.
“Reproduce the failing tests and group them by likely root cause. Fix only the first group, rerun the relevant tests, and stop if the change affects public APIs or needs a new dependency.”
Acceptance evidence: the failing command, focused diff, passing rerun, and any tests not executed.
3. Coordinate multiple agents without losing ownership
OpenAI explicitly highlights multi-agent coordination. A useful pattern is to assign independent questions rather than send several agents into the same files. One can inspect tests, another can research documentation, and a third can review the proposed patch. The lead agent then compares evidence instead of combining outputs blindly.
“Use separate agents for reproduction, documentation research, and risk review. Do not let them edit. Bring me one combined diagnosis with disagreements called out before we choose an implementation.”
Acceptance evidence: named subtasks, sources or outputs from each, contradictions, and a single integration decision. DMT’s guide to agentic AI in marketing explains why an orchestrator needs verification rather than blind delegation.
4. Run a source-backed marketing or SEO refresh
For marketers, Voice can help turn a messy brief into a controlled research workflow: inspect the existing URL, identify the search-intent gap, check official sources, distinguish facts from public reaction, draft the update, and stop before publishing. It is useful for steering; it should not become an excuse to remove editorial or approval gates.
“Audit the existing article against the official launch documentation and current search questions. Separate confirmed facts, user reports, and our recommendations. Preserve the URL and internal-link strategy. Prepare the update, but do not publish until duplicate, factual, metadata, and rendered-page checks pass.”
Acceptance evidence: source links, dated fact check, change summary, duplicate decision, metadata, and browser QA. This is the governed pattern behind DMT’s AI marketing automation guide.
5. Research and build a decision report
Work can help a strategist or analyst define the decision, collect current sources, compare options, and create a report. Voice is valuable when the user wants to challenge assumptions while evidence is being assembled.
“Build a decision memo for these three options. Use primary sources for claims, mark estimates as estimates, show where sources disagree, and ask me before changing the scope or using any paid source.”
Acceptance evidence: source ledger, assumptions, comparison criteria, uncertainty, recommendation, and an editable report.
6. Review a draft, design, or pull request conversationally
It is often easier to react verbally than to compose a perfect revision request. A user can ask for the biggest weakness, challenge a design choice, or request a summary of a diff. The agent should convert the conversation into concrete, reviewable changes.
“Summarize the change for a reviewer, identify regressions and missing tests, and propose line-specific comments. Do not approve, merge, or rewrite the branch.”
Acceptance evidence: file- or section-specific feedback, severity, rationale, and a clear distinction between required fixes and preferences.
7. Supervise a long-running migration or build
A user can request milestone updates, interrupt a bad direction, and decide what to do when a blocker appears. The efficient pattern is event-based reporting, not continuous narration.
“Proceed in four stages: inventory, plan, implementation, verification. Give me a short update only when a stage finishes or a blocker changes the plan. Stop before destructive data migration or production deployment.”
Acceptance evidence: durable stage notes, commands or changes made, verification output, rollback information, and remaining risks.
8. Triage an incident while keeping consequential actions gated
During an incident, talking can be faster than switching among logs, dashboards, and chat windows. Voice can help organize evidence and maintain a timeline, but emergency pressure makes approval boundaries more important, not less.
“Create an incident timeline from the available logs and status updates. Identify the safest reversible mitigation. Do not restart services, rotate secrets, message customers, or change production until I confirm the exact action in text.”
Acceptance evidence: timestamped timeline, sources, hypothesis confidence, proposed mitigation, owner, approval record, and post-action verification.
Where Voice is the wrong interface
- Exact technical input: hashes, credentials, shell commands, URLs, selectors, and filenames are easier to mishear than to paste.
- Secrets and sensitive data: speaking confidential information can expose it to people or recording devices in the room.
- Noisy or multi-speaker environments: OpenAI says Live is designed primarily for one-on-one conversation and may react to background speech.
- High-consequence approvals: a spoken “yes” should not be the only record authorizing a payment, production change, deletion, publication, or message to a customer.
- Verbatim records: Voice transcripts may not exactly match what was said.
- Quiet shared spaces: typing may simply be more considerate and more private.
Voice should complement, not replace, the written artifacts that make agent work auditable: plans, diffs, test results, source links, approval receipts, and completion checks.
Privacy, recording, and transcript accuracy
OpenAI’s current Voice Help Center says audio clips from Live and Advanced Voice conversations are stored with the transcript in chat history and retained for 30 days. Deleting a chat also schedules associated clips for deletion within 30 days, subject to stated security, safety, legal, and previously disassociated training exceptions. Archiving a chat does not delete it.
OpenAI says audio or video clips are not used to train models unless the user chooses to share them through the relevant data controls. Transcripts and other files may still be used depending on plan and settings when “Improve the model for everyone” is enabled. Business, Enterprise, and Edu users cannot opt to share Voice clips from workspace conversations for training.
Two cautions follow:
- A transcript is not a verbatim audit log. Overlapping speech, background noise, fast conversation, names, code, and uncommon terms can be represented imperfectly.
- Room privacy still matters. Product controls cannot stop a coworker, smart speaker, meeting recorder, or phone in the room from hearing confidential speech.
Before using Voice with client, employee, health, financial, legal, or security-sensitive material, confirm the applicable workspace data policy, retention setting, access controls, and consent requirements.
A practical safety checklist for voice-directed agents
| Control | What to say or configure | Why it matters |
|---|---|---|
| Scope | Name the exact project, folder, account, date range, and deliverable | Prevents the agent from expanding the task based on convenience |
| Permission boundary | Start read-only; require approval before writes, installs, sends, publishes, deletes, or production changes | Voice can feel informal even when the action is consequential |
| Source policy | State which files and official sources control factual claims | Stops a fluent conversation from becoming unsupported certainty |
| Update cadence | Ask for updates at named milestones or blockers | Reduces connected time and avoids distracting narration |
| Exact-value confirmation | Require text confirmation for commands, URLs, amounts, recipients, and destructive targets | Protects against transcription and interpretation errors |
| Stop conditions | Stop on missing access, conflicting evidence, unsafe requests, billing prompts, or failed tests | Makes failure visible before it compounds |
| Completion proof | Require diffs, tests, links, screenshots, receipts, or other verifiable evidence | A spoken “done” is not proof that the outcome is correct |
These controls are not unique to coding. The same pattern applies when Voice directs marketing research, document creation, spreadsheet analysis, or a browser task. For marketers, this is the operational link between voice interfaces and governed AI marketing automation.
The wider trend: voice is becoming an agent control plane
Voice search was about finding an answer. Early assistants were about completing a narrow command. GPT-Live made the conversation more continuous. Voice in Work and Codex now connects that conversation to agents that can keep working, use project context, and coordinate tools.
That progression changes the optimization problem. A useful voice-directed agent needs more than natural speech. It needs durable context, explicit permissions, observable state, progress reporting, interruption handling, and a human approval boundary. The interface can feel like a conversation, but the system behind it must behave like a well-run workflow.
For publishers and marketers, this does not make traditional voice search and answer-engine optimization irrelevant. It adds a second layer. Content still has to be discoverable and trustworthy, but agent-ready information also needs clear definitions, current dates, source links, decision criteria, and explicit limitations so an agent can use it safely during a conversation.
Frequently asked questions
What is Codex Voice mode?
It is ChatGPT Voice running inside Codex in the ChatGPT desktop app. It lets eligible users speak, interrupt, start or redirect tasks, coordinate multiple agents, and receive progress updates using Codex’s available context, tools, and permissions.
Is Codex Voice mode available on Windows?
Yes. OpenAI documents Voice in Work and Codex for the ChatGPT desktop app on both Windows and macOS.
Can I use Codex Voice on iPhone or Android?
Not as a standalone Codex Voice experience. OpenAI documents paired iOS remote access to supported desktop work. Codex is not selectable as a standalone experience on web or mobile, and the launch documentation does not promise Android parity for this Voice control.
Why is the Voice button missing in Codex?
Update the desktop app, confirm that you are in Work or Codex, check plan and regional eligibility, and review workspace admin settings. OpenAI says availability is gradual and depends on plan, region, workspace, and app version.
Is Codex Voice free?
Do not assume that it is unmetered. Connected Voice time can be metered separately, while delegated tasks consume the shared Work and Codex usage pool. OpenAI estimates about six credits per connected minute for credit-based Business and Enterprise customers. Personal-plan users should check the current Usage panel and plan information.
Does Codex Voice work with the CLI or VS Code extension?
The July 23 launch announcement is for the ChatGPT desktop app. OpenAI has not described it as a new Voice interface for the Codex CLI or IDE extension.
Can developers use GPT-Live through the API?
Not yet according to OpenAI’s current launch page. OpenAI says GPT-Live API access is planned and offers a form to receive availability updates.
Are Voice transcripts exact?
No. OpenAI says transcripts may not exactly match what the user or ChatGPT said, particularly when speech overlaps, background noise is present, or the conversation moves quickly.
Can Voice bypass Codex approvals?
No. OpenAI says Voice uses the tools and permissions available to Work or Codex. Teams should keep consequential actions behind explicit approvals and confirm exact production changes in text.
Bottom line
Codex Voice mode is a real launch, not a leaked experiment: OpenAI announced it on July 23, 2026 and documented it across its release notes, Voice Help Center, Work and Codex guide, and rate card. The most important feature is not hands-free prompting. It is conversational supervision of longer-running agent work.
The launch is also narrower than some social posts imply. It is desktop-first, eligibility and rollout vary, mobile support is paired rather than standalone, connected time and delegated work can hit different meters, transcripts are not verbatim, and the GPT-Live API remains planned rather than generally available.
The best way to test it is with a reversible task, a narrow permission boundary, milestone updates, and written confirmation for exact or consequential actions. Voice can make agents easier to direct. It does not make verification optional.
Editorial method: This analysis was prepared for Digital Marketer Tayeeb using OpenAI’s July 23 release notes and announcement, current Help Center documentation for ChatGPT Voice, Work and Codex, Codex plans, and the Codex rate card. Reddit discussions were reviewed only to identify recurring user questions and early rollout reports. Product facts were rechecked on July 24, 2026; where OpenAI has not published a complete entitlement or rollout table, the article says so rather than inferring one from community access reports.