Source: raw/Making_with_Grok_Bot.md (Greg Isenberg’s Startup Ideas with Billy Howell, ~35 min) + raw/11_Grok_Bot_Use_Cases_That_Feel_Like_Cheating.md (Matthew Berman, ~20 min) + raw/9_AI_Techniques_You_Probably_Haven_t_Tried.md (The AI Daily Brief, technique #7)

Refreshed 2026-09-29 from: raw/How_we_built_Grok_Bot_in_a_month_Roman_Ugarte_SpaceXAI.md (Lenny’s Podcast, Roman Ugarte, Grok Bot product lead, first-party) · raw/7_Grok_Bot_agents_I_use_every_day.md and raw/How_SpaceXAI_designers_use_Grok_Bot_and_Figma_MCP_to_ship_faster.md (How I AI, Claire Vo) · raw/Sam_Altman_-_Singularity_Slow-Down_Emad_Runs_18_Grokbots_Waymo_Slashes_Hardware_83%_EP_283.md (Peter H. Diamandis Moonshots EP283) · raw/GPT-6_Hits_AGI_Tech_Euphoria_2.0_SF_Mansion_Shortage_NYC_Bans_AI_in_Schools_Venezuela_Oil_Deal.md, raw/Nvidia_s_Historic_Quarter_SaaS_Comeback_Bessent_vs_Druck_America_s_Debt_Crisis_Cancer_Vaccine.md and raw/Anthropic_IPO_at_Risk_Meta_s_Muse_Pop_Token_Prices_Fall_Open_Source_Gains_Share_Alignment_Fails.md (All-In) · raw/What_the_Top_AI_Users_Are_Doing_Differently.md, raw/What_to_Use_the_Latest_AI_Tools_For.md, raw/Why_Everyone_is_Now_Getting_Excited_About_Personal_AI_Agents.md and raw/Agent_Wars.md (The AI Daily Brief, Nathaniel Whittemore)

Grok Bot is SpaceX AI’s agent-team product: named, avatared bots that each hold one job, each run on their own cloud computer, connect to services through plugins, and can be put on schedules (“routines”). It was incubated by “a handful of people” from the Cursor side of the company; its product lead, Roman Ugarte, was Cursor’s 15th employee (Lenny’s Podcast). (Corrected 2026-09-29: this line previously called it “xAI’s” product. None of the original three sources named the maker; Ugarte places Grok Bot as one of “three big pillars” of SpaceX AI, alongside the coding products “Cursor and Grok Build” and the model effort. SpaceX’s AI division is built around xAI and bought Cursor — see Cursor for the deal.) Peter Diamandis dates the early-beta launch to August 11. The first three sources arrived about a week later and reached the same structural conclusion independently: one bot per job, a chief-of-staff bot as the single human interface, and a hard cap on how many bots you create. Two of the three are running real businesses on it.

Key Takeaways

  • The chief-of-staff topology is the whole design. One bot is the entry point; it delegates to specialist bots and rolls their reports back up. Berman talks only to a pinned “Chief of Staff” from Slack or Telegram; Howell interfaces through his and drops down to a specialist only when he wants to give direct feedback, the way you would skip a layer at a real company.
  • The constraint is a feature. Bot identities are shape × colour combinations — Howell counts roughly 76–88 possible unique bots. His argument is that this forces mission-orientation, where a chat product’s infinite gray threads produce abandoned projects. “Colour and shape is all the human brain needs to differentiate something.”
  • Pick one project. Howell is blunt: you do not have enough tokens to run four businesses at once, and mixing a newsletter, a Twitter account, and receipt-sorting into one workspace produces context bleed that makes every bot worse. Running one newsletter heavily, he finished week one with ~10% of his usage left.
  • Earn the right to create an agent. Have the chief of staff do the task once. Review it. Then say “create a bot that does the thing you just did.” Creating a bot per imagined job is how you burn tokens and stall.
  • Delegate for context, not just for labour. Howell’s reason for not letting the chief of staff post to Instagram: “you wouldn’t have the CEO going in and making Instagram social posts every day… like a real person, it takes up mental bandwidth for that agent.” Context window is the mental bandwidth.
  • The biggest pitfall is letting agents make your decisions. Howell lost three weeks to an agent that could not choose between Notion, local files, and Google Sheets for content storage. “At some point you’re the person that has to make the decision, not the agent.”
  • Build → execute → automate. Once a workflow is proven, remove the ambiguity. Business-critical steps get deterministic automation, because models are black boxes and may write it differently tomorrow — and because it frees tokens for building.
  • A plugin may only be skills for using a service, not write access. Howell’s Shopify example: the built-in integration could talk about the store but not act on it. The fix is to install the vendor’s CLI on the bot’s own cloud computer.
  • It sits between chat and a coding agent. Berman’s positioning: a bridge between Q&A with a chatbot and Codex or Claude Code — “if you’re not technical and you don’t want to think about using some of these coding agents, Grok Bot is such a nice middle ground.”
  • Two build decisions explain the product (first-party). Ugarte says a handful of people built it in about a month from first line of code to internal prototype, then ran about three weeks of internal beta before public launch. He credits two early choices. First, everything runs in the cloud, so the bot is “this persistent colleague” that “can do its own work” and “has the same state everywhere you interact with it,” separately from your device. Second, “it’s actually really important that these bots have their own computer”, because many tasks have no good MCP or API and a bot needs to click and type like a person. “About 99% of automations on the platform” are created by asking a bot in plain language, and skills are created in the background so users “never have to type a slash command.” (Lenny’s Podcast, raw/How_we_built_Grok_Bot_in_a_month_Roman_Ugarte_SpaceXAI.md)
  • The bot hides its work on purpose. Grok Bot shows a typing indicator and progressive updates, not tool calls, clicks or chain-of-thought. Ugarte says early users asked to see a to-do list and how the bot prioritised, but “nobody wanted” streamed reasoning. (same source)
  • Access got cheaper after launch. Nathaniel Whittemore (The AI Daily Brief) reports that launch users were unsure whether they needed both Cursor Ultra and SuperGrok Heavy, “a total of 60 a month Cursor Pro subscription, as well as the $100 a month Super Grok subscription.” That is NLW-reported, not checked against a SpaceX AI pricing page. (raw/What_the_Top_AI_Users_Are_Doing_Differently.md)
  • Always-on is what users pay for. On All-In, David Sacks says he “just reached my Grok Bot usage limit and have to upgrade” (a co-host adds “200 bucks”). He values that the bots “keep working even when your computer is off,” unlike desktop OpenClaw or Hermes setups that stop when the machine sleeps. (raw/Nvidia_s_Historic_Quarter_SaaS_Comeback_Bessent_vs_Druck_America_s_Debt_Crisis_Cancer_Vaccine.md)

The three-week cadence (Howell)

The single most transferable thing in these sources is a schedule, not a feature.

Week 1 — build the team from an audit. Start with a chief of staff. If you have an existing business, give it your existing docs (his: Notion, Slack, Gmail) and ask it to take stock and name the top three agents to create first to drive revenue. State the mission as revenue explicitly. For the Arlington Bagel — a local newsletter to 6,000 subscribers every Thursday — the answer was a research agent, a sales agent, and a Beehiiv-expert agent. That third one is the pattern worth stealing: when the business depends on a platform, give the platform its own bot so the chief of staff does not have to be an expert in everything. Editorial was deliberately deferred; the chief of staff edited issue one.

Week 2 — no tinkering. No new agents, no fiddling with skills, maybe a connection or two. Execution only. “Run the damn ball. Run with that team for a week and see what you can get done.”

Week 3 — expand where the gaps showed, then automate. Only now add the bots the week revealed you needed (inbox, merch shop), and only now ask the chief of staff: “based on how we ran this week, what routines could we set up to make progress while I’m not working on this?”

Anti-creep, on a schedule. Re-run the audit periodically: what agents can we add, what can we remove, what should become a routine or a skill. Howell on introducing novelty: “I only add an agent if we really need it. I’m so anti idea-creep and agent-creep, because that’s how you burn tokens.”

The five-line end-of-day brief

Most of Howell’s bots have a daily, weekly, or bi-weekly routine with the same instruction: look at your current to-dos, what did you create today, what blockers do you have — write a quick five-line report and send it to the chief of staff. The chief of staff then rolls the reports into what shipped, what’s stuck, and what needs you.

The five-line cap is deliberate: “more than that, you’re just going to be burning tokens and context.” This is the cheapest legible-fleet mechanism in the wiki — a status protocol, not a dashboard.

Adversarial QA panels

Howell’s method for getting work product from ~50% done to ~90% done without reviewing it himself, which he frames as a general agent pattern he uses in Codex too:

Once you’ve done this once, create a panel of subagents — the chief of staff or other agents on the team — and have them do an adversarial review of your work. Do that in three rounds.

And the compounding step: once you have given a bot good feedback in a thread, “take what we did in this thread and turn it into a skill that another agent can use when it’s reviewing, so it reviews exactly like me.” Your taste becomes a reviewer.

Berman’s eleven jobs

The specialist bots he actually runs, with the mechanics that generalise:

  • Email triage. A 7:30am pass that batches (a) obviously archivable mail — 14 out-of-office auto-replies from his newsletter, notifications — behind one archive button, then (b) low-effort mail he must read but need not answer, batched into a few sentences, then (c) per-email deep reads that pull the whole thread plus prior interactions with that person and propose the next action. Connected HubSpot and Google Drive so it can check contracts and notes before proposing. It learns: tell it “anytime you get an email like this, don’t even ask me.”
  • Email classification. Sponsorship emails get a lead score — a heuristic he built up in OpenClaw and ported over wholesale — so spam auto-archives and the rest sort. About a dozen labels.
  • Calendar. Negotiates meeting times, checks for overlaps, and creates events from emails containing multiple dates. He never talks to it directly; the chief of staff does.
  • Browser use, in the cloud. Sign in once and the session persists. Shopping, returns (it hands back the QR code), gym bookings, doctor’s appointments, DMV registration. He notes agents can browse Amazon freely since the Amazon-vs-Perplexity suit was thrown out.
    • Caution, low confidence (2026-09-29): this may not last. ^[ambiguous] Whittemore reports that Amazon cut off Meta’s Muse agent from shopping its sites “effective immediately,” quoting a pop-up that “continued access by an unauthorized AI agent violates Amazon’s conditions of use” (raw/Agent_Wars.md). On All-In, one host says “Amazon this week said we got to block these things… First they did Perplexity I think now they’re going after Muse and other bots.” Another host in the same episode says his Grok Bot shops Amazon “through the web browser… and it doesn’t get blocked” (raw/Anthropic_IPO_at_Risk_Meta_s_Muse_Pop_Token_Prices_Fall_Open_Source_Gains_Share_Alignment_Fails.md). No source reports a block on Grok Bot. Whittemore also reports the opposite move from Shopify: a partnership with Meta to officially support Muse, with direct back-end access and Shop Pay agentic checkout across Shopify stores (raw/Agent_Wars.md). One commentator Whittemore quotes, Theo Jaffy, expects agent providers to be “allowed on some sites” and “banned on some sites.” Treat Amazon browsing as working today, with no guarantee that it continues.
  • Meeting summaries. Fathom records and transcribes; Grok Bot checks Fathom’s API every 30 minutes for new recordings, ingests the transcript, writes a summary plus action items split by who committed to what, and pushes it to Slack or Telegram. You can address the bot out loud during the meeting and it will pick the instruction out of the transcript.
  • School bot. Scans the personal inbox at 7pm for mail from his kids’ schools, extracts what he must know, summarises the rest, creates the calendar events it finds, and invites his wife.
  • Computer cleanup. The guardrail is the lesson: “I told it, do not delete anything,” and instead categorise into low / medium / high risk. It returned 90GB low-risk (caches, unused Docker images, npm) and 380GB medium — including 100GB of git worktrees — for him to approve item by item. Scheduled weekly.
  • Publishing. A one-line “publish this” that returns a public or private URL in seconds, via a plugin.
  • Food ordering, via a partner CLI where available and browser control where not.
  • Telegram bridge. Not native — you create a bot via Telegram’s BotFather, hand Grok Bot the token, and it wires the rest. Voice notes work.
  • Coding, below.

Routines are cost control, not just convenience. His classification routine runs every 30 minutes between 8am and 7pm on weekdays with a batch job first thing — “so it’s not just burning tokens when I don’t really need it.”

Vo’s bot roster

Claire Vo (How I AI) says she has “killed all my open claws and moved almost entirely to Grok Bot,” and has “probably 30 running at any one time.” The bots worth copying (raw/7_Grok_Bot_agents_I_use_every_day.md):

  • Chief (chief of staff). Every hour from about 6am to 9pm, seven days a week, it sweeps her inboxes, calendars and Slack workspaces. It archives items that need nothing from her and pings her only when something does. It also sends a morning briefing and a weekend plan-ahead. She trained it on her voice because the default writing was poor (“the Grok model sniffs of Claude”), and she still uses it only for low-stakes replies.
  • Holly Help Desk (support). Every hour it sweeps Intercom and the support email inbox, working from the support playbook. For refunds it raises a Stripe “agent action”, an approval button she clicks before the refund is issued. For bugs it can hand a fix to Cursor Cloud agents. Once a week it looks back seven days and proposes additions to the support playbook and docs.
  • PR Closer (“Look Good to Me”). Once a day it works through her open PRs: what to merge, close or rebase, answering review comments, and handing rebases and review fixes to Cursor Cloud agents. She cleared “about 50 PRs” that had been waiting on her in one day.
  • Family bot. Each morning it prints a one-page “kitchen table” newsletter: the day’s schedule, per-child notes pulled from school email, and kid-friendly news. Her husband’s variant emails the PDF to a Kindle address. A 2:30pm routine asks who is doing school pickup and what gear is needed.

Her caveat against OpenClaw. Routines are less proactive out of the box than OpenClaw’s heartbeat. She had to tell Chief “put this on a schedule” more than once. State every schedule explicitly, and check the routine settings.

Single-player for humans, multiplayer for bots. Vo: “you can’t put a Grok Bot in the group chat.” A share-a-bot-template feature shipped instead, and a Grok Bot team designer says a bot marketplace of templates launched on the day of their How I AI recording (raw/How_SpaceXAI_designers_use_Grok_Bot_and_Figma_MCP_to_ship_faster.md). Bots can share a group chat with each other, though. The same designer seeds an idea into a thread with PM, designer and engineer bots and lets them develop it “from their own perspective.”

Sales integrations. Whittemore reports that Grok Bot added native integrations with Salesforce, HubSpot, Gong, Clay and Granola (raw/What_to_Use_the_Latest_AI_Tools_For.md). Whether they can write to those systems or only read them is not stated. Given Howell’s Shopify finding above, test write access before building a sales bot on them.

Grok Bot as a coding front-end (reported)

Berman relays a workflow he says came directly from engineers at Cursor, and is explicit that he has not run it himself:

  • A bot per project, and per workstream within a project.
  • Kicked off from Slack; the bot delegates to the Cursor Agent CLI and can launch Cursor cloud agents.
  • The bot holds the surrounding context — Notion, email, Slack, GitHub — and, per those engineers, “collects and persists that context apparently better than Cursor can directly.”
  • It follows up on the agents it launched: look at the PR, and if CI isn’t green, keep going until it is.
  • You read brief summaries instead of the full agent transcript, and can ask for detail on demand.

Treat the “better context persistence” claim as second-hand and unverified — it is a report of what a vendor’s engineers said, not a measurement.

Update 2026-09-29: the delegation half is now first-hand. Vo’s PR Closer and support bot both hand coding jobs to Cursor Cloud agents from inside Grok Bot (see her roster above). Whittemore quotes Matt Palmer, who works on Grok Bot: each day his bot reads his X bookmarks, “spins up a Cursor agent to build a demo,” validates the work “with screenshots and video,” cuts a branch and sends him a link each morning (raw/Why_Everyone_is_Now_Getting_Excited_About_Personal_AI_Agents.md). Ugarte adds that people use Grok Bot “to kick off cloud agents or to kind of merge PRs or to do QA,” while production coding stays in Cursor and Grok Build. Nobody has measured the context-persistence claim.

The economics of not using your best agent

Howell writes his newsletter blurbs with a cheap external automation rather than Grok Bot, and the reasoning is the clearest statement of this trade-off in the batch:

“If you have a high-performing employee, I don’t want to waste it writing little two-sentence summaries every week for a hundred items that I may or may not use.”

The pipeline: the research bot fills Notion cards twice a week with candidate stories; Grok Bot filters and keeps the best ones; a cheap deterministic automation writes the two-sentence blurb in a fixed format. Judgement is the expensive part and stays expensive; formatting is the cheap part and gets made deterministic. His caveat is essential — “I got here because I stood up the process and did it manually with Grok Bot first.” You cannot automate a step you have not yet proven.

Sales without boiling the ocean

Grok Bot found Howell a sponsor by monitoring his inbox and surfacing a missed inbound from a local bagel-shop owner, with a recommendation to follow up and pitch a specific ad slot. The sales bot had already priced every ad slot against market rates and newsletter growth, and produced a sales sheet.

His outbound rule is a routine, deliberately small:

Monday: add five prospects to the list. Then pick the top three, build an ad package custom to each, draft the outbound, and send it to the chief of staff.

Everything still lands in front of him for a final edit. “You’ll make more progress doing that, I promise you, than doing 200 outbound emails at a time.”

Extending bots onto your own hardware (Mostaque)

On Peter Diamandis’s Moonshots (EP283), Emad Mostaque says he runs 18 Grok Bots “in a little swarm” and has given them control via Tailscale of a MacBook M4 Max, an RTX 5090 “and a range of other computers as well, plus all my subscriptions” (raw/Sam_Altman_-_Singularity_Slow-Down_Emad_Runs_18_Grokbots_Waymo_Slashes_Hardware_83%_EP_283.md).

  • The mechanism. Each bot already has a computer with a shell, so it can be handed another one: “you can actually have it take over an entire MacBook M4.”
  • What his sub-bots do. One installs a new GLM model; another installs and tests an Alibaba model. One optimised the Alibaba 27B model on the 5090 and “added 76% to performance at 64K context.” That figure is self-reported and unaudited.
  • The pattern. The bot’s cloud computer is the control plane, and your own GPUs and machines become workers it can reach.
  • The cost. Every machine on the tailnet, and every credential stored on it, is now inside the bot’s reach. That raises the stakes on the categorise-don’t-act guardrail in Try It step 9.

Try It

  1. Choose one business or one project. Not two. This is the source’s single strongest warning.
  2. Run the day-one audit. Point a chief-of-staff bot at your existing docs, state that the mission is revenue, and ask for the top three agents to create first.
  3. Do not create a fourth bot in week one. Prove each job through the chief of staff before it earns its own bot.
  4. Give every bot a five-line end-of-day routine reporting to the chief of staff, and have the chief of staff roll it up as shipped / stuck / needs-you.
  5. Take back one decision. If a bot has been comparison-shopping options for more than a few days, make the call yourself and tell it the decision is final.
  6. Add an adversarial QA panel to your highest-stakes output, three rounds, before you review it.
  7. Turn one round of your own feedback into a reviewer skill so the next round does not need you.
  8. Check whether your plugin can actually write. If it only knows how to talk about the service, install the vendor CLI on the bot’s computer.
  9. Guardrail anything destructive with categorise-don’t-act — the “do not delete anything, sort into low/medium/high risk” pattern.
  10. Make the first task a triage, not a draft. Ugarte’s advice: connect Slack and email, then ask the bot to go through them and “suggest like five things that you can take off of my plate and what it would take for you to do that.” Two of his five were useful, and he spun up a bot for each.
  11. Once you have several bots, add a bot that improves the others. On All-In, a co-host (Jason Calacanis, by context) says he told an “improve my bots bot” to “go through all my bots and fill out their skills and instructions and make them better every day.” It proposed merging two or three bots’ instruction sets. He estimates, unmeasured, that his bots got “30 40% better on average.” A companion prompt from the same segment: ask your chief of staff “what else should I be using you for?” (raw/GPT-6_Hits_AGI_Tech_Euphoria_2.0_SF_Mansion_Shortage_NYC_Bans_AI_in_Schools_Venezuela_Oil_Deal.md)
  12. Migrating from OpenClaw. Vo had her OpenClaw maintenance bot export “a secrets free package of my entire agent setup, crons, agent identities, rules.” She uploaded its sub-folders to Grok Bot, which carried over the agents’ personalities, tasks and connectors. Then the same maintenance bot removed each old agent’s crons and gateway configuration.
  13. Before handing a bot your own machines (Tailscale or otherwise), decide what it may touch there. Give it a machine that holds only the credentials the job needs.

Open Questions

  • Pricing — launch figures superseded, limits still undocumented. At launch Isenberg described it as “300 a month,” and Howell reported finishing a heavy first week with ~10% of usage left. Whittemore now reports that Grok Bot is included in the 100/month SuperGrok subscriptions (raw/What_the_Top_AI_Users_Are_Doing_Differently.md). This is NLW-reported, and no source gives the usage limit of each tier; Sacks hit his and had to upgrade. Check SpaceX AI’s own pricing page before budgeting.
  • Whether bots really share one VM. — Resolved 2026-09-29: one computer per bot. Howell had said “they all live on one computer in the cloud,” then hedged “I’m pretty sure on that.” Grok Bot’s product lead says the opposite, first-party: “it’s actually really important that these bots have their own computer” (raw/How_we_built_Grok_Bot_in_a_month_Roman_Ugarte_SpaceXAI.md). Vo agrees (“every Grok Bot has a virtual machine” that you can click into), and Diamandis describes the product the same way: “each bot gets its own dedicated cloud computer with a browser, a terminal, and the ability to log into your actual apps.” This weakens the file-clutter half of Howell’s context-bleed argument for picking one project. The token-budget half still stands, and so does the context the chief of staff accumulates across bots.
  • The Cursor-team workflow — partly confirmed. Berman had not run it himself. Vo has now shown Grok Bot kicking off Cursor Cloud agents first-hand, and Matt Palmer’s daily demo workflow is relayed by Whittemore. The claim that Grok Bot “collects and persists that context apparently better than Cursor can directly” still has no first-party confirmation or measurement.
  • Community numbers are unaudited. Mostaque’s “added 76% to performance at 64K context” and Calacanis’s “30 40% better on average” are both self-reported with no method given.
  • Will Amazon block Grok Bot? Amazon reportedly blocked Meta’s Muse, and one All-In host reports Grok Bot still shopping unblocked (see the browser-use caution above). There is no first-party statement from Amazon or SpaceX AI on Grok Bot.
  • The 76–88 agent cap is Howell’s own arithmetic from shape and colour combinations, and he notes custom avatars are possible — which would break the cap. Not a documented product limit.
  • No Android app at time of recording, iOS only, with Android “coming soon” per hearsay. Telegram is the documented workaround.
  • How do you keep an agent team from converging on its own loop? Isenberg’s proposed answer is a dedicated “idea agent” generating creative direction for the other bots. Howell is openly sceptical and offers only the periodic audit instead. Genuinely unresolved between the two.