Source: ai-research/openai-enterprise-signals-frontier-firms-2026-09.md — Author: OpenAI, “Enterprise signals: What frontier firms are doing differently” · URL: https://openai.com/signals/enterprise-data/ · Published: 2026-08-12 · fetched 2026-09-29. Commentary and chart readings from raw/What_the_Top_AI_Users_Are_Doing_Differently.md — The AI Daily Brief (Nathaniel Whittemore, NLW), https://www.youtube.com/watch?v=usNZ0fWbTok, which found the data via a16z’s “charts of the week” repost.

OpenAI’s first-party usage study of its enterprise customers finds agentic work overtaking chat and the heaviest users pulling away from everyone else. Frontier firms (the top 10% by output tokens per active user) went from 2.6× to 8.3× a typical firm’s usage between January and June 2026, and use plugins and skills far more often. The study measures AI use in tokens, which OpenAI itself calls “an imperfect measure of business value”. Treat it as a benchmark for depth of use, not for ROI.

Key Takeaways

  • Definitions. Each month OpenAI ranks enterprise customers by output tokens per active user. Frontier firms are the top 10% that month; typical firms sit between the 45th and 55th percentiles.

  • The gap tripled in six months: 2.6× in January 2026 to 8.3× in June. By industry the gap is widest in information and technology (11.7×) and narrowest in manufacturing (5.3×). Typical firms grew only “1.9× to 2.8×” over the past year across industries, which OpenAI reads as “many organizations are still using simple chat assistants.”

  • Agentic work now dominates by volume. As of June 2026, agentic use (defined as Codex tokens) was 64% of combined Codex and ChatGPT output tokens among enterprise customers. Both run on the same models; “the difference is how that intelligence is put to work.”

  • Frontier firms use the advanced features. Among weekly active users:

    Firm typePluginsSkills
    Typical9%3%
    Frontier21%19%
    OpenAI itself95%93% (NLW’s chart reading; the page text gives only the plugin figure)
  • The fastest growth is outside engineering. Weekly active enterprise Codex users grew, since February:

    FunctionGrowth
    Legal108×
    Sales41×
    Recruiting41×
    Marketing26×
    Finance and accounting20× (NLW’s chart reading; not in the page text)
    Engineering5×
  • Why software moved first, in OpenAI’s words: codebases give agents clear context, tests make outputs easy to verify, and coding progress speeds up AI research. General knowledge work lagged because tasks “provide limited context, can be difficult to specify, and lack clear criteria for verifying the result.” OpenAI says agentic AI has “increasingly found product-market fit with general knowledge workers since the beginning of the year.”

  • Chat and agentic work do different jobs. In a sample of more than 10 million messages:

    • writing is the most common ChatGPT use
    • coding plus system and agent operations make up nearly 75% of agentic messages
    • recruiting (32%), sales (26%), policy (25%) and communications (24%) spend a quarter to a third of their messages on system and agent operations
    • in design, coding is almost 60% of agentic messages

How Frontier Firms Put Agents to Work (OpenAI’s account)

  • Context, tools and persistence. OpenAI names three things agents need:
    • context: memory, voice input and “appshots” (sharing what you are viewing) help workers hand over the information a task needs
    • tools: computer and browser use let agents work in websites and files
    • persistence: “goals and loops keep agents working until a task is finished”
  • Plugins bundle a workflow. A plugin combines skills (reusable instructions) with apps (company data, tools and actions). OpenAI’s example is a sales plugin that pairs the team’s playbook with CRM access to draft a tailored response for review.
  • Rules before rollout. Frontier firms “set clear rules for where agents can operate, what information they can access, when they can take actions, and how people review higher-risk decisions.”
  • Learning built into the week. Hands-on builder sessions, forums for sharing new use cases and internal showcases spread what works.
  • A marketing-relevant internal example. OpenAI’s Finance team lists 16 workflows. One turns fragmented campaign data into ROI curves and weekly recommendations for “where the next marketing dollar could generate the greatest return.”

The Timeline Behind the 64% (NLW’s reading of OpenAI’s chart)

These points come from NLW reading the chart on air; the page text gives only the June figure.

DateAgentic share of enterprise output tokensEvent NLW ties it to
Aug 2025~0%GPT-5 launch
Oct 2025low single digitsCodex generally available
Feb 202613%Codex app for macOS
Mar 202627%Codex for Windows, GPT-5.4
Late Apr 202653% (the “flippening”)GPT-5.5
Jun 202664%end of data set

NLW adds that frontier firms use 17× the tokens they used 18 months ago, against about 2× for the average firm. He also notes that 64% does not mean 64% of work sessions are agentic: it measures output-token volume as a proxy for how much work is done.

NLW’s Synthesis: a Use-Case Ladder

NLW combines OpenAI’s function data with his own consulting experience (his opinion, labelled as such):

  • The ladder: generation (an email, a report, an Excel formula) → synthesis (combining disparate sources) → execution (acting inside existing systems) → maintenance (keeping a system running over time).

  • Legal as the worked example. The agent handles “coverage and coordination”: comparing terms, flagging deviations, drafting redlines, recording decisions, monitoring commitments. People keep risk judgment and accountability: negotiating material terms, setting risk tolerance, approving exceptions and final language.

  • Legal work mix, chat vs agentic (NLW reading the chart):

    CategoryChatAgentic
    Writing57%16.2%
    Knowledge retrieval20.5%8.3%
    System operations0.2%17.7%
    Workflow automation—7.7%
    Classification and extraction—4.6%
    Coding—32.9%
  • His bet on the next gains: multiplayer, team-level AI, because most agents today still “operate within individual silos.”

  • Context on spend (NLW relaying Business Insider): a self-reported Microsoft spreadsheet had 350 of 600 US contributors listing AI usage. Seven of eight departments had median monthly AI spend of about 500; Core AI’s median was 28,000. Business Insider found no correlation between token burn and pay or promotion.

Try It

  1. Benchmark your own team against the table. What share of weekly active users touch a skill or plugin? At 3% and 9% you are typical; at 19% and 21% you are at the frontier average.
  2. Move one marketing workflow from chat to delegation. Package it as a skill plus a connected app (a playbook plus CRM or analytics access), then give it a goal and a finish line instead of a prompt. Gaspar’s goal cards cover how to write the finish line.
  3. Write the agent rules first: where it can operate, what it can read, when it can act, and who reviews high-risk steps.
  4. Run a recurring builder session so working setups spread across the team rather than staying with one power user.

Open Questions

  • Tokens are not value. OpenAI says so itself. The report shows no link between the 8.3× usage gap and revenue, margin or quality outcomes.
  • Vendor data, vendor incentive. The study measures only OpenAI products (ChatGPT, Codex, the API) and promotes OpenAI features (plugins, ChatGPT Work, a customized benchmark offer for Enterprise customers). Firms that do their heaviest agentic work on other vendors would not show up as frontier here.
  • Unstated chart values. The timeline, the 17× figure, OpenAI’s 93% skills figure, finance’s 20× and the legal work mix come from NLW reading interactive charts that the fetched page text does not contain. Verify against the live charts before quoting them.
  • The page contradicts its own table. The text says Manufacturing ranks “last in Codex and API intensity”, but the displayed eight-industry table puts Manufacturing at 4 (Codex adoption) and 7 (API intensity), with Construction at 8 in both. The table may cover only a subset of industries.
  • Unsourced outcome claims. The “leading firms” section cites 10–15× more bugs fixed in the same time, and product releases in 2 weeks instead of 3–6 months, without naming the firms or the method.