Key takeaways
- The challenge for enterprises is no longer proving that AI agents can work, it’s making them reliable enough to do high-value work in produc
- Each deployment starts with a specific job, such as resolving billing issues, supporting insurance claims, or resolving employee IT service requests.
- Codex proposes updates that teams can test and approve, helping the agent adapt as customer behavior changes.
What happened
The challenge for enterprises is no longer proving that AI agents can work, it’s making them reliable enough to do high-value work in production. Agent behavior must also adapt as products, policies, and user behavior change. That requires more than a model: it requires the systems, evaluations, and deployment expertise to improve agents as those conditions change without giving up control.
A customer might use it to resolve a billing issue—from understanding the request and verifying the customer to looking up account information, applying company policy, and taking an approved action. Companies decide what remains consistent across deployments—such as policies, evaluations, and escalation rules—and what should change for each workflow or channel. This lets teams build on what works and expand to new use cases without starting over.
Presence brings together the components teams need to run agents in production: policies and standard operating procedures, guardrails, approved actions, simulations, evaluation tools, and a Codex-powered improvement process. Together, these components help teams connect company systems, define how agents should behave, evaluate performance, enforce policies, and manage changes after launch. Built through years of deploying agents with customers, Presence is an ambitious product shaped by the demands of mission-critical environments.
Why it matters
We’re introducing OpenAI Presence, a battle-tested product that helps enterprises deploy trusted AI agents that can answer questions, resolve issues, use company systems, take approved actions, and escalate to people when needed. Proven through years of working with customers at enterprise-scale, Presence pairs model reasoning with policies, guardrails, and escalation rules that verify accuracy and performance.
Each deployment starts with a specific job, such as resolving billing issues, supporting insurance claims, or resolving employee IT service requests. The agent receives only the knowledge and system access required for that job. The company sets the policies: what the agent can do, when it needs approval, and when a person should take over. After launch, production sessions and escalations reveal gaps.
Codex proposes updates that teams can test and approve, helping the agent adapt as customer behavior changes. OpenAI works alongside each customer to identify each high-value workflow, connect the necessary knowledge and systems, establish permissions and policies, test the agent, and bring it into production. As the deployment expands, OpenAI and select systems integrators can continue to support it.
Presence was developed in tight collaboration with OpenAI’s Research team, with generalized insights from every deployment informing ongoing research and product development and improving the product for all customers over time. Today, Presence supports real-time experiences across voice and chat, such as customer support, outbound sales, and high-risk internal workflows.
What to watch
Every deployment generated insights that continuously inform OpenAI research and product development, creating compounding benefits for customers. Presence powers OpenAI’s English-language phone support channel at 1-888-GPT‑0090, handling open-ended requests, verifying callers, using account context, and taking approved actions. Within weeks, it met or exceeded benchmarks we use to grade frontline human-support quality and now resolves 75% of inbound issues without human assistance.
Working with our launch team, its Codex-powered improvement loop reduced human handoffs by 15 percentage points in just 10 days. Before a Presence deployment reaches users, teams can test it against common requests, edge cases, and higher-risk scenarios. Simulations and graders check whether it reached the right outcome, followed policy, used tools correctly, and escalated when appropriate. Guardrails can intervene when an interaction moves outside the company’s boundaries. That work does not stop at launch. Production




