Full annotated catalog for AI Video & Content Production. The topic index carries a one-line gloss per article for fast navigation; this file preserves the complete annotation for every entry, verbatim.
Split out 2026-07-24: ai-video-content/_index.md had grown past the point where it could serve as a navigation step (claude-ai reached ~210KB / ~50k tokens to read). Nothing was deleted — the long-form descriptions live here.
Practical AI video production: avatar models, automation pipelines, composition frameworks, and motion graphics. Each article is compiled from primary sources (product pages, technical reports, tutorials, and open-source repos) rather than press summaries.
WEO project status (2026-06-27). The first post-production pilot — WEO Marketly brand promotion (the agency itself, not client/OmniPresence work) — is in post-production / delivery. Shoot completed April 5-6, 2026; the WEO Media + Marketly Digital merger is public as of 2026-06-18 (CTA domain weomarketly.com). The built pipeline is HeyGen Avatar V twins (v3 API + the
weo_videolibrary) + HyperFrames motion-graphics + real-footage b-roll, finished with PIL/PNG captions, two-pass −14 LUFS audio, mandatory deband, and Drive delivery. Higgsfield generative b-roll was tried and dropped (502s + charged failures); VEO/Runway are not in the path; Remotion is a reference, not the chosen tool. Details: the “WEO Marketly Promo” and “Content Production Workflow” articles (both internal, unpublished — see the internal-articles list at the end of this index).
AI Avatars & Generation Models
-
HeyGen Avatar V — Production-scale video-reference avatar model. 15s webcam clip → unlimited-duration 1080p twins with preserved identity, talking rhythm, and gestures. 175+ language lip-sync. State-of-the-art vs Kling, Veo, OmniHuman, Seedance.
-
HeyGen Studio Automation with Claude Code — Three-tool production pipeline (ElevenLabs + HeyGen + Remotion) orchestrated by Claude Code. Script-to-finished-video overnight. Open-source Python template, 6-stage pipeline, Avatar V Playwright workaround, cost breakdown.
-
HeyGen Studio CLAUDE.md Template (Bootstrap Pattern) — HeyGen-shipped
CLAUDE.mdtemplate for bootstrapping a HeyGen Studio project under Claude Code. Captures the recommended state-management contract, the dual-script-source (storyboard + final), the “things you can ask” enumeration as the load-bearing CLAUDE.md pattern, and the 6-stage pipeline shape. Companion artifact to HeyGen Studio Automation — that article is the how; this is the what you ship when handing the project to Claude Code. -
skills (Official HeyGen Skills Bundle) — Vendor-published 3-skill bundle (MIT, 232 stars, v3.1.0 2026-04-27, Shell):
heygen-avatar(persistent digital twins from photos + voice synthesis),heygen-video(idea-to-scripted-video with avatar delivery + style recommendations),heygen-translate(175+ language localization with voice cloning + lip sync). Eleven runtimes supported (Claude Code / Cursor / Codex / OpenClaw / Gemini CLI / Copilot / Junie / Goose / OpenHands / Amp / Cline). Three install paths (gh skill install heygen-com/skills <name>recommended; ClawHub one-shot; OpenClaw plugin). Dual-auth: API key OR MCP/OAuth against HeyGen plan credits. Sister pattern to skills — same April-May 2026 window, same vendor-published-skills shape, both wrap existing API/MCP/CLI in a Markdown-skill format the agent loads at session start. Pairs with Hyperframes for full-pipeline coverage. -
FLUX 3 (Black Forest Labs) — Multimodal Video, Audio, Image, and Action Model — BFL’s multimodal foundation model: one architecture jointly trained on images, video, and audio (Self-Flow scaled up), generating up to 20s of video with native audio in a single pass, chained multi-minute character-consistent clips, and multilingual dialogue. Company-run preliminary evals claim wins over Runway Gen-4.5 (77%) and Luma Ray 3.2 (93%) but only 52% vs Seedance 2.0; the FLUX-mimic companion extends the same backbone to robot actions, tested and deployed on Audi production tasks. Early access staged in phases. Secondary coverage dates the launch 2026-07-23; two JS-rendered capability lists on the announcement page survive only via secondaries quoting them (flagged in the article’s Open Questions).
-
Nano Banana — Google’s Gemini Image Model (and Nano Banana Pro) — Google’s Gemini-native image generation/editing family: the original Nano Banana (Gemini 2.5 Flash Image, Aug 2025), the studio-quality reasoning tier Nano Banana Pro (Gemini 3 Pro Image, Nov 2025), and the current Flash-tier Nano Banana 2 / 2 Lite (Gemini 3.1 Flash Image / Flash Lite Image). Headline capabilities: legible in-image multilingual text, single-shot infographics, up-to-14-image blending with up-to-5-person consistency, Google-Search-grounded factual visuals, 2K/4K + studio creative controls (camera / focus / color grade / relight), and a SynthID + C2PA provenance stack on every output. Reached in this wiki’s tutorials through LTX Studio and Higgsfield, usually A/B’d against GPT Image 2 for character sheets, brand books, and product stills.
Search & Discovery
- SentrySearch — Semantic Search Over Videos (Gemini + Qwen3-VL) — Local-first CLI that indexes video footage (default 30s chunks with 5s overlap, 480p / 5 FPS) and searches by natural language using Gemini Embedding 2 (cloud) or Qwen3-VL (local or DashScope). Sister tools SentryMerge (cam-config + reruns) and SentryBlur (face / license-plate / NL redaction) compose into a search → trim → redact pipeline via shared
~/.sentrysearch/last_clip.jsoncache and--lastflag. Tesla Sentry / dashcam by default but cam-agnostic via SentryMerge’s modular cam-config. Python 3.11/3.12 pin (PyTorch wheel limitation). Operator-tunable confidence threshold (default 0.41). Occupies the search slot in the topic — distinct from existing generation, composition, assembly, and editing tools. - watch) — Give Claude the Ability to Watch Any Video — MIT Claude skill (taoufik123-collab, 58★) that lets Claude watch a video: scene-change frame extraction (one frame per cut — token cost bounded by shots, not duration), a 0–10s “hook microscope” (2fps + word-level Whisper on the opening, for studying why a viral hook works), a fixed-schema
report.md, and optional Obsidian auto-save that mirrors this wiki’s own clip→compile pattern. yt-dlp + ffmpeg + free captions (Whisper only as fallback). The input/understanding counterpart to video-use (editing) and the Video Toolkit (production).
Editing & Assembly
- How Fable 5 Edited Its Own Launch Video (Thariq, Claude Code Team) — First-party Anthropic case study: the Fable 5 launch video edited by Fable itself, no video editor opened. The edit is a repo — Whisper word-level transcripts →
final-edit.jsonEDL with a written rationale per pick → frame-accurate ffmpeg cuts (self-verified by re-transcribing the output: “zero ums”) → 7 hand-written S-Log3→Rec.709 LUTs → 11 designer PNGs rebuilt as Remotion components (overlays land on spoken beats from the transcript) → 2× code→Figma→code round trips via Figma MCP → headless 4K render reviewed still-by-still. 17 takes / 25 GB raw → 3:00 / 653 MB finished 4K in 4 days, driven with/goal dont stop until you have a final video. Caveat from the thread: a professional colorist flagged improper S-Log3 color management — domain review remains the backstop. - video-use (browser-use) — Claude Code skill for conversational video editing. Transcript-driven cuts (ElevenLabs Scribe word-level + diarization), parallel animation sub-agents (PIL / Manim / Remotion), self-evaluating render loop. 12 hard rules for production correctness, artistic freedom elsewhere. Python + ffmpeg, “100% open source.” 3,151 stars.
- OpenCut — Open-Source CapCut Alternative — Free MIT-licensed video editor for web/desktop/mobile (51.6k★, 5.6k forks, 96 contributors, TypeScript + Rust + WGSL). Classic version ships today at opencut.app; rewrite (announced 2026-05-18) targets agent-native editing: Editor API + plugin-first architecture + Rust core for one cross-platform codebase + MCP server for AI agents + headless mode for batch rendering + in-editor scripting tab. Lead maintainer
mazeincoding; primary sponsor fal.ai. Rust + WASM is the architectural commitment (wgpu/WASM GPU renderer, WASM compositor, Rust-WASM time utilities). Occupies the agent-callable-NLE slot the OSS stack didn’t have — different paradigm from Remotion (programmatic) and Hyperframes (HTML composition). 2026-07-06 refresh adds a subjective real-world-usage impression (buggy, rewrite site intermittently down) from a tool-roundup video — see the callout in the article. - Open Montage — Text-Prompt-to-Documentary Video via Coding Agent — #1 on a 2026-07-06 GitHub-trending roundup: a two-sentence plain-language prompt has a coding agent research, script, edit, and render a full video using real stock/archival footage rather than AI-generated images (“no AI slop”). Demo: a 75-second sea-life documentary, no external video-API keys. Positioned closer to Claude Code Video Toolkit (prompt-driven) than to a timeline NLE like OpenCut.
- OpenMediaTools — Browser-Native FFmpeg-WASM Suite — Free privacy-first browser-based converter/extractor for video (MP4/MOV/MKV/WebM/AVI), audio (MP3/WAV/OGG/FLAC/AAC/M4A), image (JPG/PNG/WebP/GIF), PDF, and AI image gen — entirely in-browser via WebAssembly + FFmpeg, no upload, no server, no account, no file-size limits beyond device RAM. Utility-tier upstream prep for the AI-video pipeline; the social-media downloader/extractor category is the distinctive offering for competitive-intel workflows. AI tools’ backend (in-browser via WebGPU/ONNX vs server-routed) is the load-bearing open question — privacy guarantee is only true if all six tool categories actually stay client-side.
Creator Workflows
- Claude + VidIQ YouTube Video Workflow — a zero-to-recorded talking-head workflow chaining four Claude features: the new VidIQ Claude MCP connector (titles grounded in real YouTube search volume / competition / opportunity score), an interview→outline skill (scoped to one video), a slide-deck-plan skill, and a Claude Design deck whose presenter-notes pane doubles as a teleprompter (record screen+face, overlay the talking head on the notes half). Claimed ~70 min / ~$200-mo all-in; single-creator demo. Scoped to expert/personal-brand talking heads, not faceless formats.
Performance Measurement & A/B Testing
- B Testing — Closes the post-publish gap the production pipeline stops short of: reading YouTube retention curves (sharp early drop / gradual decline / mid-video cliff / flat, mapped to script fixes), Meta’s and TikTok’s metric hierarchies, and native A/B testing tools on both platforms. Covers the systematic hook-testing loop (hold body+CTA constant, read 3-second hook rate not CTR, kill losers after 48-72h), TikTok’s free organic-testing path (no ad spend required), Spark Ads for creator-variation testing, and a 9-tool comparison table (150/mo Motion) mapped to realistic budget tiers. Complements Virality Predictor (pre-publish prediction) and Outcome Kit (outcome-level attribution).
Composition & Motion Graphics
- HeyGen Hyperframes — Open-source HTML-based video composition framework (Apache 2.0, v0.7.x; “Write HTML. Render video. Built for agents.”). Ships a router skill bundle (
/hyperframesroutes intent → creation workflows like/product-launch-video,/website-to-video,/embedded-captions,/talking-head-recut,/motion-graphics+ domain skills-core/-animation/-creative/-media/-cli/-registry/media-use; ~17 in the clone, ~19 globally) for Claude Code, Cursor, Codex, Gemini CLI. Deterministic seek-driven rendering, Frame Adapter pattern (GSAP / CSS / Lottie / Three.js / Anime.js / WAAPI / custom), ~109 prebuilt blocks, first-class AWS Lambda / GCP Cloud Run rendering. Docs athyperframes.heygen.com(playgroundhyperframes.dev). Full documentation cluster (2026-06-26):- HyperFrames in Claude Design — draft a valid first-draft video in claude.ai/design → download ZIP → refine in any coding agent; includes the Open Design BYOK twin of the same loop.
- HyperFrames Quickstart & CLI — the under-2-minute path plus the complete non-interactive
hyperframescommand surface (incl. cloud / lambda / cloudrun render backends). - HyperFrames HTML Schema & Compositions — the authoring contract: composition document structure, the full
data-*timing-attribute set, and thewindow.__timelinesGSAP handshake. - HyperFrames Core Concepts — deterministic seek-driven rendering (frame clock → seek → BeginFrame capture → encode) + the Frame Adapter API + a fair HyperFrames-vs-Remotion teardown.
- HyperFrames GSAP Animation — the paused-timeline contract, supported methods/properties, and the timeline-duration-equals-composition-duration rule.
- Prompting HyperFrames — cold-start vs warm-start prompt shapes plus the easing/caption/transition vocabulary cribs and per-agent notes.
- HyperFrames Rendering — output formats (MP4/MOV/WebM/GIF/PNG), the full
renderflag surface, local-vs-Docker determinism, GPU/workers, batch + transparent overlays. - HyperFrames Website to Video — capture any URL and turn it into a production video in one prompt; the 7-step capture → design → script → storyboard → VO → build → validate pipeline.
- HyperFrames Common Mistakes & Troubleshooting — the 8 authoring pitfalls (symptom → cause → fix), environment/render troubleshooting, and the lint/doctor pre-render gates.
- HyperFrames Block Catalog — the ~109-block registry grouped by category (shader transitions, social overlays, lower thirds, code blocks, data viz, VFX/device) plus per-block anatomy.
- HyperFrames Packages — the ~13-package monorepo (
core/engine/producer/studio/player/shader-transitions+ CLI, plussdk/aws-lambda/gcp-cloud-run), with@hyperframes/playerembedding depth. - HyperFrames Launch-Video Playbook (builder interview) — Bin Liu (VP Eng, HeyGen) & Jake Moran (PMM) on the operator workflow: the three setup paths, the website-to-video one-shot (Spotify / Fable 5), Jake’s net-new launch playbook (
frame.mdfrom hyperframes.dev/design → key-events table → component reuse from the open-sourced launch-video repo →storyboard.html→ Studio/Inspector last-mile editing), model guidance (Fable 5 / GPT-5.5 top tier, Gemini for cost), and the spatial-vs-temporal-aesthetics thesis behind why launch videos are hard for agents. - HyperFrames Pipeline & Getting Started — the canonical 7-step capture → design → script → storyboard → VO+timing → build → validate flow, the on-disk project layout, the minimal write-HTML/preview/render loop, and the open-source launch-video projects.
- HyperFrames Skills Catalog & Agent Setup — the full ~18-skill catalog (router + 10 creation workflows + 7 domain skills) plus per-agent wiring for Google Antigravity and GitHub Copilot CLI.
- HyperFrames Deployment & Cloud Rendering — HeyGen-hosted cloud rendering + self-hosted AWS Lambda / GCP Cloud Run, one-click preview+render-API templates, templates-on-Lambda batches, and the Remotion-Lambda migration path.
- HyperFrames Advanced Rendering — native/supersampled 4K (
--resolution), HDR10 (--hdr), preview/render performance tuning, and local transparent-video matting (remove-background). - HyperFrames Timeline Editing, Keyframes & Studio DOM Editing — Studio’s three editing surfaces (timeline, visual keyframe/Arc-Motion tools, capability-gated manual DOM inspector), all round-tripping deterministically to authored HTML.
- HyperFrames Video Components, HTML-in-Canvas & Variables — embedding/trimming video, capturing live DOM as a WebGL texture for shaders/3D, and
data-composition-variablesfor parameterizing one composition into many (templating). - HyperFrames SDK, Authentication & Feedback — the headless
@hyperframes/sdkediting engine, HeyGen auth (OAuth or API key, env-var-first, per-capability provider hierarchy with offline fallback), and feedback collection. - HyperFrames Video Editor Cheatsheet — fast operator quick-reference: common edit prompts, CLI commands, Studio shortcuts, timing attributes, render presets, and quick fixes.
- media-use) — 2026-07-17 launch — HeyGen gave HyperFrames a media library: 10k+ music tracks, 75k+ images, SFX, and logos, plus access to HeyGen’s generative models (voice/image + new avatar-video), reached through the
/media-useskill + theheygenCLI (≥ v0.3.0) and cached to.media/+~/.media/. Free on the OAuth login path (heygen auth login --oauth);--api-keybills API credits. Disentangles the three login-gated surfaces — the OSS framework, theheygen-CLI + hosted catalog (what the launch is), and the credit-metered “HyperFrames by HeyGen” MCP connector (which does not reach the library). Framework v0.7.61. - Catalog deep-dives (133 entries, 2026-06-27): Code Blocks & Snippet Themes, Transitions & Shader Effects, Lower Thirds & Tickers, Maps & Data Viz, Social Cards & Platform UI, VFX, Liquid Glass & Device Frames, Caption Styles & Overlay Components — per-block anatomy for every block + component, indexed from Block Catalog.
- WEO Motion Library — One Vocabulary, Two Runtimes (CSS + GSAP) — first-party build documenting a reusable craft pattern: a 38-effect brand motion vocabulary implemented twice because CSS
animationis wall-clock and does not frame-seek, while a headless video renderer captures frames by seeking a paused GSAP timeline onwindow.__timelines. Covers the CSSdata-animtrigger API (fires on.wm-live/[data-deck-active], JS-gated base states so no-JS renders visible, reduced-motion settle) vs the 36 deterministic GSAP factories; the count-up/odometer proof case (arequestAnimationFramecounter can’t be seeked, but agsap.to({v:0},{v:target,onUpdate})proxy tween can — animated numbers finally render to MP4); the infinite-ambient-loop gotcha (repeat:-1makes the master durationInfinity→ renderer aborts “Set maximum size exceeded”; loops must be finite for video, sized to the comp); and end-to-end render verification (alab/comp → MP4 1920×1080 / 7 s / 24 fps / 168 frames, distinct progress at every sampled seek). Builds directly on GSAP Animation + Core Concepts; the generalizable thesis lands in HTML Is the Canvas (the animation layer does not port for free across web and video). - Remotion Motion Graphics — AI motion graphics generator converting natural language prompts into React-based Remotion animations. Constants-first code generation, in-browser Babel compilation, live preview.
- Claude Code Video Toolkit (Digital Samba) — Open-source AI-native video production workspace for Claude Code (MIT, 890 stars). 10 skills, 13 slash commands, templates, brand profiles, transitions library. OSS model stack (Qwen3-TTS, FLUX.2, LTX-2, ACE-Step, SadTalker) on your Modal/RunPod account. Typical cost $1–2/month.
Higgsfield (API-first generative video)
- Higgsfield Overview — API platform for generative AI (images + videos). Async queue: submit → poll-or-webhook → fetch. Credit-based billing. API key + secret auth. Base URL
platform.higgsfield.ai. - skills (Official Skills Bundle) — Vendor-published Markdown SKILL.md bundle for Claude Code, Cursor, Codex (MIT, 102★, v0.3.0). Four skills:
higgsfield-generate(30+ models + Marketing Studio),higgsfield-soul-id(face training → reusablereference_id),higgsfield-product-photoshoot(10 modes with backend prompt enhancement ongpt_image_2),higgsfield-marketplace-cards(Amazon-style listing assets — main + secondaries + A+ modules with hidden marketplace-compliant templates). Three install paths (npx skills add/gh skill install/ Claude Code/plugin marketplace add); CLI install + auth handled automatically. The productized version of the Nate Herk CLI-over-MCP thesis; supersedes Higgsfield MCP as the recommended primary agent surface. COOKBOOK ships three end-to-end recipes (founder-photo brand campaign, URL-to-4-ad-modes UGC batch, recurring founder team-update video). - Higgsfield Supercomputer (hosted agentic chat surface) — Higgsfield’s cloud-native agentic platform for creative AI, launched 2026-05-14 at higgsfield.ai/supercomputer. Built on a Hermes-agent fork (first commercial Hermes deployment we’ve tracked outside Nous Research). Frontier-model brain selection at runtime (GPT 5.5 Pro / Claude Sonnet / Claude Opus 4.6 / Gemini 3.1 Pro). Higgsfield internal skills (product images, ad creative pack, UGC workflow, Soul ID character model) preloaded so a one-line operator prompt expands into a multi-step plan with reference loading, prompt enhancement, generation checkpoints, gallery review. Cross-model image-to-video chain (Kling 3.0, Seedance 2.0). Checkpoint card surfaces credits before generation — prevents typo → 10× credit drain. Productizes the Higgsfield + Claude Code creative agency thesis inside a single hosted chat surface; trades configurability for time-to-first-ad.
- Higgsfield MCP — Conversational surface. Custom connector at
https://mcp.higgsfield.ai/mcpdrops the full image + video stack inside Claude / OpenClaw / Hermes / NemoClaw. No API keys; sign in with your Higgsfield account. 16+ image models, 17+ video models, 9 video presets, Soul Characters, multi-model side-by-side, Product Ads from a URL. - Higgsfield Virality Predictor — Score a short clip before posting: virality index, hook score, hold rate, peak hook timestamp, and a brain/attention heat map, reached through the MCP or the Higgsfield site. Launched May 2026, experimental preview, free (no credits consumed). Turns Higgsfield from generate into generate-then-evaluate; the operative move is filtering ad creative before distribution (generate 10, score, run the 2 strong, cut the 8 before they touch a budget). Both creator sources stress it is a relative filter, not an oracle (no grasp of meme/cultural context; real performance still depends on targeting/offer/distribution). Addresses the community “is this Meta’s TRIBE 2?” question: conceptually similar, not confirmed the same model. Two June-2026 explainers (Nick Pontis / Mind Marketing + AI Fire); closes the publish-loop gate the prior Higgsfield-MCP tutorials left open.
- Higgsfield + Claude Code Ad Agency Workflow — End-to-end DTC campaign in one Claude Code conversation. Firecrawl MCP scrapes brand brief → GPT Image 2.0 hero static (4 variations, Claude grades) → copy overlay → Seedance 2.0 hero animation → UGC creator generation → 2 UGC video clips. Claude as orchestration layer + creative director. ~170 credits per campaign before iteration. Mike Futia / SCALE AI tutorial.
- Higgsfield MCP Tutorial — Brand Book, Storyboard, Landing Page (Robo Nuggets) — Sister tutorial to the Mike Futia ad-agency workflow but optimized for a brand-launch use case. End-to-end Spiderhead AI demo: logo → side-by-side brand books from Nano Banana 2 vs GPT Image 2 → 6-panel logo-animation storyboard → Seedance 2.0 video → mockup landing page → Claude Code (VS Code fork-conversation) translating mockup to a working localhost landing page. Two-path setup (~30s, desktop-app or
/mcpfor Antigravity-class IDEs). Operator gotcha: per-call credit cost not yet returned by MCP. Cost-aware confirmation pattern before high-credit ops (e.g., 720p Seedance). Higgsfield-vs-fal.ai decision call: existing subscribers should use the MCP; new users should evaluate fal.ai for pay-as-you-go on the same models. - Higgsfield MCP — 50-Ad Instagram Campaign from One Product Image (Claude Desktop) — Third Higgsfield-MCP tutorial, narrower than the other two: pure ad-campaign-at-scale use case. Connects MCP via Claude Desktop’s Connectors UI; runs a four-phase chain (research with Playwright MCP scraping Meta Ads Library → 5×5×2 ad matrix with human-approval gate → batched generation across Nano Banana 2 (product) + Soul 2 (humans) → local download organized by batch/aspect-ratio). Ends by saving the entire workflow as a
/ad-creatorskill via/skill-creator, turning the multi-prompt build into a one-command tool. Operator gotchas captured: Claude misreporting Playwright availability (verify, don’t trust); generations land in Higgsfield Community tab → My Generations not the main image gallery; pre-flightlist_workspaces+ balance check before 50-image runs (~949 credits used). Settings: Opus 4.7 extra-high, bypass-permissions, project folder mounted. Pairs with Meta Ads CLI for end-to-end generate-and-upload. - Higgsfield MCP + Claude — Content Factory Skill Walkthrough (Adil) — Walkthrough of a custom Claude Skill that wraps Higgsfield MCP in a four-stage marketing-content workflow: research → content plan → generate → meta-ads upload. The skill is the load-bearing primitive — turns Higgsfield from operator-driven generation into agent-driven brand-content factory runnable from a single Claude chat. Stage 3 uses GPT Image 2 + Higgsfield Marketing Studio with per-batch permission gates (operator approves as it goes, prevents credit-runaway). Stage 4’s Meta Ads upload closes the publish loop the prior four Higgsfield-MCP tutorials in this topic (Mike Futia / Robo Nuggets / 50-Ad Campaign / Nate Herk Creative Agency) left open. Demo: 100-video UGC batch. The fifth Higgsfield+Claude tutorial in the topic — distinctive on closed-loop publish + four-stage skill packaging.
- Claude Design Animations — Transcript-Synced Motion Graphics — The animations template inside Claude Design (
claude.ai/designor the Claude sidebar; Pro or Max plan required), which this creator reports using more than any other surface — it produces the motion graphics in their recent videos. The differentiator is not that it animates but that it accepts a timestamped transcript and returns animation synced to the spoken words, converting the hardest part of explainer B-roll (matching visuals to narration timing) from an editing task into an input. A progressive input ladder maps onto how much brand control you need: bare prompt → one reference image → multiple reference images → a saved design system (UI, logo, colors, voice) selected per generation so every video matches. A vague prompt triggers a clarifying interview (audience, duration, aspect ratio, label density — answerable with “decide for me”) rather than a bad output. Post-generation controls are non-destructive: a tweaks panel swaps accent colors and a bottom timeline drags per-shot duration, so most fixes need no regeneration. ~A few minutes for a 30-second 16:9 piece. Open gaps: no export-format details (resolution, frame rate, alpha channel for overlay use), no transcript-format spec, no usage-limit accounting, and no comparison against the topic’s deterministic code-based path (HyperFrames / Remotion). - generate Skill Instead of a Higgsfield Subscription — Pay-Per-Generation Creative Routing — The structural alternative to the topic’s Higgsfield cluster, and the argument is about business model rather than product quality: Higgsfield’s core product is aggregation — a wrapper over creative model APIs (Google Nano Banana and Veo, OpenAI image models, Dreamina) whose value is the connections plus a credit system — and that is reproducible as a Claude
/generateskill calling the same providers directly. What the skill buys you that a subscription does not: a budget stated in the prompt (49 Plus / $79 Max AUD). Caveats: the budget cap is an instruction, not a billing limit; the aggregator providers the skill routes through are never named; and no total-cost-of-ownership math or output-quality comparison is offered. - Higgsfield → Figma — Editable AI Static Ads (Claude Code Skill) — MIT Claude Code skill (Mike Futia / SCALE AI, repo + 13-min walkthrough, both published 2026-08-01) that produces Meta static-ad variant grids where only the imagery comes from AI. The reframe: AI-ad text is pixels and every “make this PNG editable” fix is recreation work, so stop generating finished ads — generate text-free plates (scene, product, lighting, no type) in Higgsfield and let Figma own headline/body/CTA/logo/brand colors as live layers bound to
Brand+Theme(Dark/Light) variable collections. End-to-end run: protected-zone read from templates, angle-brief copywriting (speed-of-results,price-objection), per-ratio native plate generation for the Meta set (4x5/1x1/9x16 — never cropped across ratios), automated plate QA (rejects stray lettering / zone intrusion / wrong luminance), variant-grid build via one variable-mode switch,<Campaign>/<Angle>/<theme>_<format>frame naming so exports drop into ad-set folders. SVG-only logo rule (paths bind to variables and invert with the theme; PNG forces manual swaps); duplicate-styleguide-per-client pattern for agencies; product microtype drifts (“24 Gummies” → “24 Commles”) so composite the real cutout instead of letting the model redraw the pack. #1 documented failure: the read-only Figma MCP connector (two exist — install the write-enabled one; the video says both). The sixth Higgsfield+Claude tutorial and Futia’s second — the structural answer to his own first tutorial’s “typographic overlays are rough” finding: type never touches the image model at all.
Multi-Model Claude Skills (filmmaking)
- FREE Seedance 2.0 Claude Skill — Multi-Model AI Filmmaking Workflow (LTX Studio) — Single free Claude Skill that prompts Nano Banana Pro, GPT Image 2, and Seedance 2 in their own natural prompt languages for three distinct AI-filmmaking jobs: character sheet, 3×3 cinematic storyboard grid, and Seedance 2 shot (continuous + timestamp-with-dialogue paths). Creative engine: LTX Studio (not Higgsfield) — pure prompt-engineering wrapper, no MCP. Workflow: image inspiration → character sheet → storyboard grid (with character refs) → Seedance 2 prompt referencing the grid as image-one-as-cinematic-grid → 15-second video. Demonstrated two-character dialogue scene end-to-end. Sister to Adil’s Content Factory (marketing-UGC use case) — same “Claude Skill as load-bearing primitive” pattern, distinct use case (short film vs. ad) + distinct creative engine (LTX vs. Higgsfield). Free zip download from video description. Creator unattributed in transcript head (YT ID
Fark3A0ACzM) — flagged in article Open Questions. - AI Animated Short Film Pipeline — Seedance 2.0 + Codex + ElevenLabs (MattVidPro) — One-person ~3-min animated comedy short built almost entirely in Seedance 2.0. Codex runs pre-production (story + all generation prompts + GPT Image 2 art + an HTML reference site); ~30-50 image-to-video clips on Polo AI; ElevenLabs re-voices every line (native Seedance dialogue clones famous voice actors); royalty-free ambiance + manual SFX in post. Includes a platform cost table (Polo AI / Open Art / Runway / Higgsfield Ultra / fal.ai), the signature duplicate-character failure mode (fix: “single character” prompting), and a Gemini-Omni-vs-Seedance comparison (Omni edits real video, fails 2D-animation consistency). Sister to the free Seedance Claude Skill (LTX route) — this is the Polo-AI + post-production-fix counterpart.
- Fable 5 + Seedance 4K Short-Film Workflow (Higgsfield Cinema Studio, Adil) — a distinct end-to-end hyperreal-live-action pipeline built in Higgsfield Cinema Studio (not LTX or Polo AI) at Seedance 2.0 4K: the Cinema Studio “elements” asset system, gray-background sheet tuning, per-asset image-model switching (GPT Image 2 → Nano Banana Pro → Soul Cinema), a VFX-free video-in-video composite (6s reference + exact-duration match), multi-variant crowd sheets, and red-arrow action annotation. The 4K-Higgsfield sibling of the free Seedance Claude Skill and the Seedance animated-short pipeline. Tool names decoded from a garbled auto-transcript (flagged in Open Questions).
- Video-to-Video Trigger Editing — Editing Real Footage with Time-Gated Prompts (Maker Zero) — A real-footage editing playbook using video-to-video models (not text-to-video): a source clip (≤~10s) + a hyper-specific prompt structured as TRIGGER (a time-gated moment — a finger-snap, a spoken word, an exact timestamp) + CHANGE (e.g. “right after the man says this at exactly 2.9 seconds, change his outfit to a hoodie with a chain”), run through Omni/Gemini-Omni or Kling via the Higgsfield aggregator. ~20% per-attempt hit rate → generate
5 in parallel and pick the best ($0.50–1 each); 720p seam-hiding discipline (never cut degraded 720p AI output straight back to real talking-head footage — cut to a different scene type); proceduralize via the Higgsfield MCP from Claude Code/Codex. Single promotional source, fuzzy model naming — confidence low, dated 2026-07. Sibling to Agentic Video Editing and the Fable 5 + Seedance 4K workflow. - Higgsfield as a Creative Agency in Claude (Nate Herk) — Fourth Higgsfield+Claude tutorial, three new dimensions vs the prior trio: (1) CLI-over-MCP architectural call for agentic work on token-cost grounds (“the MCP has all those tools, so from a token perspective it’s actually more expensive — the CLI is just better for agents”); (2) skill reverse-engineering workflow turning a single winning prompt into a reusable
.claude/skills/hypermotion-video/SKILL.mdrecipe that compounds across runs; (3) two-routine scaling pattern (Sunday-plan + Monday-generate) that grows asset bank from 50 → 100 → 200 ads per week while operator sleeps, with a Google Workspace CLI–created Sheet acting as the cross-routine asset database. Demoed on fictional headphone brand “Murmur” + a sleep-supplement bottle with Higgsfield Marketing Studio’s Hypermotion preset. Operator gotchas: skill written mid-session needs Claude restart to register; reference-image fidelity needs explicit “must appear exactly as shown” wording; rejected generations diagnose-and-retry via Claude reading its own prompt. Companion to the same author’s ElevenLabs voice-agents tutorial (same direct-vendor-CLI architecture for ElevenLabs). - Higgsfield Image-to-Video — Three featured models (
higgsfield-ai/dop/preview, Bytedance Seedance Pro, Kling v2.1 Pro). Motion-prompt template: describe movement + set pace + specify camera moves. Technical checklist. - Higgsfield SDK (Python) —
pip install higgsfield-client. Auth viaHF_KEYorHF_API_KEY+HF_API_SECRETenv vars. Four usage patterns (submit-and-wait, polling, callback, request management). File uploads for bytes/paths/PIL. JS/TS coming soon. - Higgsfield Webhooks — Add
hf_webhookquery param to submit URL. Delivers completed/failed/NSFW final states. 2-hour retry window. HTTPS + 2xx in 10s required. Idempotency viarequest_id. - Higgsfield Training Framework (OSS Origin) — Historical context. Apache-2.0 distributed-training framework at
higgsfield-ai/higgsfield(3.6k stars, last push 2024-05-25). Pre-pivot artifact — the company moved from trillion-param LLM training infra to consumer creative AI. Listed for vendor-lineage awareness, not for use.
Model releases
The August 2026 Video-Model Wave — Seedance 2.5, MiniMax H3, Wan 3.0
A cluster of video-model releases in early August 2026 that the wiki had no coverage of. The one that shifts the most is MiniMax / Hailuo H3: open weights, 720p with native audio generated, and an architecture that appears to be omni — accepting multiple reference types including up to 15 seconds of video for extension with character and voice consistency, the Seedance-style conditioning approach now in an open model. The headline is the hardware floor collapsing: official spec is an RTX 5090 or RTX 6000-series card, but a community report puts it on a 2060 with 6GB VRAM (“nearly usable,” visibly mushy) and a 4-bit quantized build plus Diff Studio runs it on Apple M-series at 8GB. The Pinocchio installer is CUDA-only. It ships uncensored — copyrighted characters and NSFW output generate directly, so the safety layer is entirely the operator’s problem and this disqualifies it from client work without one. On the hosted side, Seedance 2.5’s practical envelope is ~10 seconds at 720 (explicitly not 4K or 1080), with a 30-second option at roughly 1,100 credits — about 5-7% of a partner-tier monthly allowance — and an API expected shortly with rumours of aggressive pricing, which is the signal to watch since an open-weights competitor is what is expected to move hosted prices. Carries a correction to the source’s own “nearly all these models are Chinese” framing: the three video models are (ByteDance, MiniMax, Alibaba), but Flux 3 is Black Forest Labs — German, and an image model. Wan 3.0 is named in both sources and described in neither. No license is stated anywhere for H3 — “fully open” is used loosely, and that must be checked on the model card before any commercial use. Both sources are creator hands-on reactions, not benchmarks.
HeyGen tutorials (2026-05-17 cohort)
- HeyGen Instant Highlights V2 — Auto-Clip Long-Form Video to Short-Form Clips — Drop a long-form video (up to 10 GB) or URL, the tool analyzes for highlight moments (speech energy, importance signals, viewer-save likelihood), and auto-cuts to short-form clips with optional captions in 9:16 / 16:9 / 1:1 formats. Clip-length range control (under 30s / 30-60s / 1-3 min / 3+ min). Optional steering instructions for what to focus on or avoid. The product replaces hours-long manual timeline scrubbing — same workflow shape as Opus.pro / Vizard / Klap / Munch, surfaced inside HeyGen.
- Style) — Prompt-engineering framework for HeyGen’s Avatar Shots feature (Avatar 5 + Seedance 2.0). 5-element prompt structure (subject / action / environment / camera / style) with one camera move per shot, layered into multi-shot beats (timestamps + per-beat camera moves), multi-avatar choreography (up to 3 avatars/scene with explicit blocking + relative-motion), and elements references (pre-uploaded locations / outfits / products for continuity). Avatar 5 for A-roll + Avatar Shots for storytelling beats. Working prompt example included.
Internal production articles (migrated to weomarketly-vault, 2026-07-09)
The six WEO production-stack articles formerly hosted here unpublished (OmniPresence System; Content Production Workflow; WEO Marketly Promo; Voice Profile Extraction; Banned AI Patterns; Mel’s Feedback Rules) were migrated to the internal weomarketly-vault on 2026-07-09 per the vault-separation rule, following the weo-ai-governance precedent. Each file remains in this folder as an archived migration stub so inbound links resolve; the full content lives at weomarketly:projects/omnipresence-system, weomarketly:projects/content-production-workflow, weomarketly:marketing/weo-marketly-promo, weomarketly:playbooks/voice-profile-extraction, weomarketly:playbooks/banned-ai-patterns, and weomarketly:playbooks/mels-feedback-rules (vault root: ~/Auto1111/Claude/weomarketly/weomarketly-wiki/wiki/).