Source: raw/The_enterprise_AI_stack_behind_Stripe_s_company_brain_Kai.md — How I AI (host Claire Vo), guest an engineering manager on the Stripe team that built Kai (name rendered “Sherrod” / “Shared” in auto-captions, unconfirmed), youtube.com/watch?v=AbZODZ_4VaM (fetched 2026-09-29).
Kai is Stripe’s internal “company brain and company agent”: cloud-hosted, always on, aware of who you are in the org chart, and used by more than 86% of the company. The guest says Stripe built it in early 2026 because the hard problem was not access to AI but “the correct governance structures so that everyone can just go use AI and know it’ll do the right thing for them.” Three ideas carry over to any company building its own agent: projects as the unit of governance, a fixed routing order for data questions, and a skill platform with usage telemetry. All adoption figures are Stripe’s own.
Key Takeaways
- Build for governance, not access. In early 2026 there was “this avalanche of AI tools”. The guest and his colleague concluded the harder problem was replicating how a complex company works, so that any employee can use an agent safely without making configuration decisions every day.
- Context by default, sensitive sources by choice. Out of the box Kai knows who you are, where you sit in the org chart, and the date and time; it can read the project-management system (OKRs, the projects you are part of, recent “shipped” emails). Access to Google Drive, Slack and private messages is opt-in per user; the guest turns his own access on and off every session or day.
- Projects are the governance unit. A project owner sets the default model (and can block expensive ones), tool policies, which tools need human approval, and can even back the project with a different agent and harness. Projects range from five people to 500. The people team runs a separate project on a secure back end.
- Tool policies instead of tool bans. Example: someone in HR handles sensitive data. Stripe does not want the agent “to sort of go rogue and put that sensitive data into some public Google document”, but also does not want to ban tools for that team. A project-level policy restricts or adds human approval to the risky tools only.
- Scope approvals to the project, not the session. “If I showed you this every single session for every single tool, eventually you’re going to press the wrong button.” Vo’s point: most tools configure permissions per individual and apply them to every session, and they cannot be shared across a team.
- Data agents need a routing order, and a warehouse that survives them. The “ask data” skill, written with Stripe’s data scientists, goes to existing dashboards and reports first, then the “blessed analytics layer”, then the data catalog to write a new query. Vo: “when in doubt an agent will just brute force it”, so the warehouse (Stripe uses Trino) must handle agent query volume. The guest: agents “may have almost taken down core systems” before Stripe hardened them.
- A skill platform, not just a skill. A skill-creator skill packages the current session as an open-spec skill; a draft editor works “a little bit like an IDE”; the platform suggests improvements to every authored skill; and retrieval picks the right skill because enterprise work has no folder hierarchy to scope it. Stripe has about 2,000 skills.
- Small team, company-wide reach. Version 0 took “one and a half people over two weeks”; Kai now has “10,000 plus people” using it every week with a core team of “less than 10 people”.
How Kai handles a data request
The demo asked Kai to build a dashboard of Kai’s own adoption from saved queries in Hubble, Stripe’s internal data-querying layer.
- Skill discovery first. Kai found and loaded the “ask data” skill, which bundles the internal tools for answering Stripe data and SQL questions.
- A secure sandbox per session. Because Kai runs in the cloud, each session gets its own sandbox, with harness tools to move data in, search and run scripts inside it, and move data out, so sessions cannot interfere with each other.
- Iterate the artifact, don’t regenerate it. Follow-up turns change the same HTML dashboard rather than producing a new one, which the guest says is better for token efficiency. Some sessions run to “hundreds of turns over multiple weeks.”
- “Last mile data”. Instead of maintaining a dashboard for every workflow, employees, engineers or not, have the agent write code in the sandbox to reshape data for their own need. The guest now generates a dashboard for every meeting he goes to, because it helps him drive the meeting.
- Resilience work that paid off. Investments made for humans (a data catalog with tiered datasets, a curated analytics layer, a resilient query engine) “have held up really well for agents”. Stripe is working on agentic identity (“we haven’t solved this yet”): labelling agent traffic by what it is trying to do, so infrastructure can prioritise and shed load.
Skills: quality is also quantity
- Usage is long-tailed. About 50 skills are “hammered every day” across the company, a long tail of 100–150 are used by parts of the org chart, and many are used by two or three people.
- Telemetry decides promotion. An ETL pipeline tells the harness owners which skills to promote into general workflows and which to move out because “it’s just taking up context”. The guest: “the more unrelated context you throw into the AI, the less good your … results become. So quantity is almost a … facet of quality.”
- Vo’s deprecation policy (what she has seen elsewhere, not Stripe’s). If a skill has not been invoked for 30 days, its owner gets a notice; with no response it moves to deprecated status and is deleted two weeks later.
Adoption path
| Stage | Team | Users |
|---|---|---|
| Version 0 | ”one and a half people over two weeks” | — |
| Pilot | 2.5–3 people | 200–300, with most interest from go-to-market teams |
| After a company-wide demo | Core team “less than 10 people" | "86 plus% of the company”; “10,000 plus people use it every week” |
- Go-to-market adopted first. A collaborator on the GTM team who builds AI for GTM drove early interest, and Stripe’s marketing team is described on the episode as “a 100% all in”. Vo: “I don’t know a single marketing person that doesn’t either want some sort of app built or some sort of dashboard.”
- Platform investment is the multiplier. Vo’s generalisation: developer-experience, data-platform and analytics-layer work that made humans efficient before AI is what gives agents leverage. Her advice to leaders who want to ship more with AI: “Double the size of your DevX team, double the size of your data team.”
- The guest’s personal use. He uses Kai’s scheduled runs as a personal assistant that tells him what he is supposed to be doing, which he calls “the biggest life hack”.
Try It
- Before rolling an agent out to a team, define one project with a default model, a tool policy, and at least one tool that needs human approval.
- Write your data agent’s routing as a skill: existing reports and dashboards first, a curated metrics layer second, the raw catalog last. Check that the warehouse can take repeated agent queries.
- After any session you expect to repeat, ask the agent to package it as a skill, and keep it in the open skill format so other harnesses can load it.
- Log how often each shared skill is used; review anything unused for 30 days.
- Marketing and GTM: build one campaign or pipeline dashboard in a single long session, refine it over several turns, then save the session as a skill so the next one costs a prompt.
Open Questions
- The guest’s name comes from auto-captions and is unconfirmed.
- No quality metrics for Kai’s outputs were given, only adoption.
- Which models Kai uses by default is not stated.
- How Kai relates to Stripe’s earlier “Minions” coding agents is not explained; the guest lists Minions only among the tools that made the small team productive.
- The incidents where agents “went rogue” or nearly took down core systems are not described.
Related
- Claude Tag — Anthropic’s shared, Slack-based team agent, the off-the-shelf alternative to building your own
- Agent Skills Overview — the open skill format Kai exports to
- Microsoft Agent Governance Toolkit — policy and identity controls for agents as open-source tooling
- Replit’s “Self-Driving Company” — a semantic layer over the warehouse unlocked self-serve BI there too
- Dan Shipper — The AI Paradox — the case for one company-wide super-agent over many personal agents
- Inside Claude’s Agent Platform — names Stripe Minions among internal-automation agents
- Docker Sandboxes — per-session isolation for agents
- The Software Factory — Warp and Ramp — Ramp’s internal agents at each product step