Turn authorized long-form video into three to five reviewed highlight clips by default with harness-native editorial reasoning and deterministic FFmpeg rendering.
Viral Clip Growth Loop Turn an authorized long-form video into reviewed, platform-specific short clips, export or publish only approved variants, and improve later batches from comparable performance evidence. Loadout capabilities The harness model is the default editorial engine. It reads the transcript, covers the whole source, proposes moments, scores them with transparent evidence, repairs boundaries, and writes the edit decision list.
# Viral Clip Growth Loop Turn an authorized long-form video into reviewed, platform-specific short clips, export or publish only approved variants, and improve later batches from comparable performance evidence. ## Loadout capabilities The harness model is the default editorial engine. It reads the transcript, covers the whole source, proposes moments, scores them with transparent evidence, repairs boundaries, and writes the edit decision list. Do not outsource those decisions to a provider score when the harness can reason over the same source evidence. For local or otherwise harness-accessible media, prefer native `ffprobe` and `ffmpeg` execution because it is deterministic, stateless, private by default, and can express exact cuts, concatenation, scene-specific crops, caption burn-in, audio normalization, and delivery checks. When local binaries are unavailable and the source is an approved HTTPS asset, use <action-tag>cli:ffmpeg:inspect_media</action-tag> to inspect it, <action-tag>cli:ffmpeg:trim_mp4</action-tag> for one continuous range, <action-tag>cli:ffmpeg:burn_subtitles_mp4</action-tag> for a reviewed SRT, and <action-tag>cli:ffmpeg:transcode_mp4</action-tag> for a broadly compatible delivery file. These hosted Actions do not currently express multi-range assembly, vertical cropping, tracking, or loudness normalization, so do not claim they completed those steps. Rendering capability and review capability are separate. Before promising a reviewed clip, determine whether the harness can inspect temporal video, inspect still images, and assess audio semantically. FFmpeg probes, decode checks, loudness measurements, silence detection, and provider scores are technical evidence, not substitutes for watching framing and motion or listening for intelligibility. When direct audiovisual review is unavailable, follow the review-capability ladder in step 6 and preserve an explicit degraded review status. Use an existing transcript when it is trustworthy. Otherwise use harness-native transcription when available, or <action-tag>cli:hyperframes:transcribe</action-tag> for word-level timestamps. The model, not the transcription tool, remains responsible for source coverage, clip selection, context integrity, and hook design. Escalate to <action-tag>cli:hyperframes:render</action-tag> when a clip needs a custom HTML composition, designed overlays, brand motion, or layouts that are cumbersome in a filter graph. Use <action-tag>cli:remotion_cli:render</action-tag> only when the required template or motion system is already expressed more naturally in React and TypeScript. Both are code-defined composition paths, not automatic clip selectors. Preview and verify their actual output before delivery. Use <integration-tag>mcp:chatcut_mcp</integration-tag> only when the user specifically benefits from a persistent collaborative editing project, editor link, or manual timeline handoff. The current Loadout connection requires user OAuth and does not provide stateless platform access. A missing ChatCut connection must never block the native lane. OpusClip is the last fallback, not the default. Use <action-tag>api:opus_clip_api:projects</action-tag> only when the source cannot be processed well by the native or composition lanes, the user accepts provider-generated candidates, and a live create-project call proves that the selected workspace has active API access. Use <action-tag>api:opus_clip_api:transcripts</action-tag> to verify source wording and <action-tag>api:opus_clip_api:clips</action-tag> to inspect returned candidates. Never treat a provider virality score as independent proof or claim repairs that the returned project does not prove. For the learning loop, use <integration-tag>composio:instagram_tools</integration-tag> and its `instagram_get_ig_media_insights` Action for eligible Instagram Business or Creator media, <integration-tag>composio:facebook_tools</integration-tag> and its `facebook_get_post_insights` Action for Facebook Page posts, and <action-tag>api:tikhub_api:tiktok_analytics</action-tag> for available public TikTok video metrics. These analytics Actions are optional. Never treat unavailable metrics as zero or imply full retention coverage when an Action returns only public engagement counts. Vugola, the provider shown in the source use case, is not currently available as a resolvable Loadout Integration or Action. Reproduce the useful behavior with model reasoning plus deterministic media tools instead of making any vendor mandatory. ## When to use Use this Skill when the user wants to repurpose an authorized podcast, interview, webinar, tutorial, founder video, or other long-form recording into short social clips and learn from their performance. Do not use it to: - download, transform, or republish media without permission; - promise that a clip will go viral; - mass-publish unreviewed or repetitive clips; - fabricate a hook by changing the speaker's meaning; - optimize from one post, unequal measurement windows, or missing data. Treat "viral" as an experiment objective, not a predicted outcome. A pre-publish score prioritizes review; observed audience behavior determines what worked. ## Required inputs Before processing media, collect: - the source URL or local media file and confirmation that the user owns the media or has permission to repurpose it; - target audience, desired outcome, content pillars, and prohibited claims or topics; - destination platforms, account targets, timezone, publishing window, and whether this run stops at draft, export, schedule, or publish; - preferred clip count and length range, brand template, caption style, language, CTA, and visual constraints; when clip count is omitted, default to a batch of three to five highlights; - the intended review boundary, including whether a human can perform the final watch-and-listen pass when the harness lacks audiovisual inspection; - a comparison window for performance, such as the first 24 and 72 hours, plus any baseline clips. If essential inputs are missing, ask only for fields that materially change the result. Never request raw credentials in chat. Use Aident Vault when an external account connection is actually needed. ## Workflow ### 1. Establish the experiment Write a one-paragraph experiment brief containing the audience, source, promise, target behavior, platforms, clip count, review boundary, publishing boundary, and measurement windows. Default the batch target to three to five highlight clips when the user does not specify a count. Choose one primary learning question for the batch, such as whether a result-first opening outperforms a contrarian statement. Do not vary every creative element at once. ### 2. Inspect the source and choose an execution lane Probe the source before expensive processing. Record duration, dimensions, frame rate, audio streams, transcript availability, and any source-quality limitation. For private local media, keep processing local unless the user explicitly approves an upload. Probe review capability independently from media-processing capability and record one mode: 1. Full audiovisual review: the harness can inspect temporal video and assess the rendered audio semantically. 2. Proxy review: the harness cannot play the full result directly but can inspect extracted frames or contact sheets, reason over the transcript and caption timing, and run technical audio and delivery checks. 3. Technical-only review: the harness can render or probe files but cannot inspect the visuals or audio semantically. Do not infer a stronger mode merely because the harness can call FFmpeg or a renderer. If capabilities are uncertain, run a small non-destructive probe, such as inspecting one extracted frame and one short audio sample, before choosing the mode. Choose the first capable lane and record it in the experiment brief: 1. Native lane, default: the harness model performs selection and planning; local media tools perform deterministic transcription, cuts, reframing, captions, audio, and export. 2. Composition lane: use HyperFrames for custom HTML compositions or Remotion for an existing React and TypeScript motion system when the native edit would become brittle or visually limiting. 3. Project-editor lane: use ChatCut only when persistent editable state or human timeline handoff is worth an OAuth connection and approved upload boundary. 4. Provider lane, last fallback: use OpusClip only when the earlier lanes cannot meet the source or delivery need and runtime access is proven. The harness model owns candidate reasoning in every lane. A provider may supply transcription, detections, or draft clips, but it does not replace the rubric or factual review. ### 3. Build a trustworthy transcript Prefer word-level timestamps. Correct names, numbers, brands, acronyms, and negations against the source. Mark uncertain passages rather than silently inventing text. If speech is too sparse for transcript-led selection, switch to visual-moment review of demonstrations, reactions, transitions, and outcomes. Keep source-absolute timestamps until the edit decision list is approved. After recutting or concatenating ranges, rebase every word and caption timestamp to the final timeline. Treat rolling subtitle windows and segment timestamps as overlapping evidence, not exact speaker-turn boundaries. Preserve short source handles beyond every provisional in and out point so the cut can be repaired without downloading or decoding the source again. Before accepting a boundary, inspect the neighboring words or audio on both sides; the visible end time of one subtitle window may already contain the next speaker. ### 4. Mine and rank candidate moments Reject any candidate that is misleading out of context, incomplete, rights-sensitive, factually unsafe, or dependent on unseen setup. For a long transcript, divide the source into overlapping windows ending on transcript-segment boundaries. Make a lightweight scoring pass across every window, then run detailed selection only inside the strongest windows. This reduces beginning-of-source bias without forcing equal picks from every section. Apply a two-second test: a new viewer should quickly understand why the moment matters. Score survivors with [the clip quality rubric](references/clip-quality-rubric.md). Favor moments with a clear standalone idea, specific stakes, a truthful opening, rising interest, and a satisfying payoff. Keep evidence for every score. Remove near-duplicates, rank by evidence rather than source position, then present survivors chronologically for review. Never pad the set to hit a requested quota. Repair proposed in and out points to complete words, adjacent sentence boundaries, or nearby natural silence while preserving any context needed for truthfulness. Return a shortlist with source timestamps, verbatim excerpt, proposed opening, target audience, rationale, risk, and score. Unless the user specifies another count, select and generate the top three to five distinct highlights that pass every hard gate. An explicit request to generate, create, or clip the source authorizes selection and rendering within this default batch size, but never authorizes scheduling or publishing. If the user asks only for recommendations, a shortlist, or review, stop before rendering. Return fewer than three, including zero, when the source does not contain enough distinct high-quality moments. When more than five pass, render the top five and retain the rest as ranked alternates instead of expanding the batch silently. ### 5. Design each approved clip For each selected clip, produce a compact edit brief containing: - one-sentence promise; - cold open or on-screen hook; - minimum context needed to preserve meaning; - payoff and optional CTA; - source ranges and intended final duration; - destination variants; - scene-level layout plan, reframing target, caption treatment, and brand treatment. Apply [the editing and distribution guidance](references/editing-and-distribution.md). Prefer one clear idea per clip. Remove greetings, throat-clearing, dead air, repeated setup, and unrelated promotion while preserving meaning and chronology. Choose a layout per scene from a small vocabulary: portrait pass-through, tracked speaker, wide-content preservation, screen-plus-speaker, two-person split, or multi-person panel. A user's explicit framing choice wins. Keep a conservative layout when evidence is weak. Do not force one crop across shot changes. For the native lane, turn the brief into an explicit edit decision list before rendering: 1. list source ranges, snapped boundaries, and final-timeline offsets; 2. specify each scene's crop or composition and reset tracking at hard cuts; 3. specify hook text, caption cards, safe regions, and display intervals; 4. specify audio cleanup and a conservative loudness target with true-peak headroom; 5. preserve source handles outside every chosen range until the final boundary review passes; 6. render one review master, then probe the decoded output for dimensions, duration, streams, frame rate, exact sample aspect ratio, and display aspect ratio. Use FFmpeg directly for exact cuts, concatenation, static or evidence-backed crops, caption burn-in, audio normalization, metadata removal, and encoding. Use HyperFrames or Remotion when the edit brief requires richer designed motion, reusable branded components, or complex composition. Use ChatCut only after OAuth when persistent project state is part of the requested outcome. For vertical delivery, force square pixels in the final composition when the filter graph can inherit or synthesize a near-square sample aspect ratio, then verify the rendered stream reports SAR 1:1 and the intended DAR. Treat container dimensions alone as insufficient. When authoring ASS or another subtitle format, render-test line breaks, escapes, quotes, and fallback glyphs in actual frames; valid source text can still produce visible control-character defects. ### 6. Review the rendered candidates Use the strongest review mode the harness actually supports and record the mode, evidence, unresolved checks, and resulting artifact status. For full audiovisual review, inspect the actual video and audio, not only metadata. Check the first frame, first spoken line, every cut or speaker change, representative middle beats, caption transitions, and final frame. For proxy review, generate a labeled evidence set containing the opening, ending, every cut or speaker change, every caption-card transition, and representative middle frames. Start with one readable contact sheet, then inspect individual full-resolution frames when text, crop, or identity is ambiguous. Compare the rendered captions with the trusted transcript and final-timeline timestamps. Re-transcribe the rendered audio when a trustworthy local or approved transcription path is available, and compare its first and last words with the edit decision list to catch a clipped final word or the next speaker leaking into the tail. Run full decode, stream, duration, exact SAR and DAR, loudness, true-peak, silence, black-frame, and freeze checks. When a conservative freeze detector flags a designed composition, inspect that interval and compare it with the corresponding source range before deciding whether it is a true duplicate-frame defect or expected low motion diluted by static canvas areas. Treat motion smoothness, subjective speech quality, music balance, and exact audiovisual synchronization as unresolved unless the available evidence truly supports them. For technical-only review, run deterministic delivery checks and return the file as `rendered_unreviewed`. Do not claim that framing, captions, motion, speaker identity, intelligibility, or payoff quality passed. Provide the artifact and a concise human review checklist, or use ChatCut only when the user wants a persistent collaborative timeline and approves its OAuth and upload boundary. Use these artifact states precisely: - `rendered_reviewed`: full audiovisual review passed; - `rendered_proxy_reviewed`: proxy evidence passed, with every unsupported semantic check listed; - `rendered_unreviewed`: only technical checks passed or semantic evidence is unavailable; - `approved_for_publish`: an authorized human approved the exact rendered artifact after all unresolved checks; - `published`: the approved artifact was successfully written to the named destination and its returned identity was verified. In the supported review scope, verify that: - the topic and reason to continue are clear quickly; - the clip remains accurate and coherent without the source video; - reframing follows the active speaker or essential object without awkward crops or jitter; - scene cuts reset tracking instead of panning across unrelated shots; - split layouts show distinct, materially participating contributors; - captions match the audio, spell names correctly, remain readable, and avoid faces and interface zones; - speech is intelligible, levels do not clip, and music never masks the message; - the payoff lands before the ending and the last frame feels intentional; - the file plays correctly at the expected duration, aspect ratio, and stream layout. Reject or revise clips that fail a supported gate. Never turn an unsupported gate into a pass. A `rendered_proxy_reviewed` or `rendered_unreviewed` artifact cannot be scheduled or published until an authorized human completes the unresolved watch-and-listen checks and approves that exact file. Do not hide defects or missing review behind a high model or provider score. If rendering is asynchronous, retain the returned job identity and poll it without starting duplicate work. ### 7. Prepare platform variants Adapt the opening text, title, description, CTA, and pacing to each destination while keeping the underlying claim consistent. Do not assume one identical upload is optimal everywhere. Produce a publish manifest containing the production lane, source identity, clip or render identity, export path or account target, copy, privacy, scheduled time with timezone, experiment variant, and approval status. Generate social copy only as a draft and review it for factual accuracy, platform fit, tags, links, and brand voice. A rendered export is not a published post. Resolve and preflight a separate platform upload Action before any external write. ### 8. Confirm the external write Scheduling and publishing are consequential external writes. Show the complete publish manifest and obtain explicit approval for the exact clips, accounts, copy, privacy, and times. Execute only approved rows. Preserve returned schedule or post IDs. Do not silently substitute an account, timezone, clip, or time. If a schedule or publish call has an uncertain result, inspect provider state before retrying. Never create a duplicate post because a request timed out. ### 9. Measure comparable outcomes At the planned windows, collect only the metrics each connected surface actually exposes. Follow [the analytics learning loop](references/analytics-loop.md). Keep raw counts, normalized rates, platform, audience, publication time, clip duration, creative variant, and exact observation window together. Separate distribution, attention, and intent signals. A view count alone does not establish clip quality. Report missing or partial metrics explicitly. ### 10. Improve the next batch Convert the evidence into no more than three findings and one next-batch experiment. Preserve successful creative features while changing one major variable. Require approval before changing brand voice, account targets, publishing cadence, or content boundaries. Self-improvement means updating the next experiment brief from evidence. It does not mean autonomous, unreviewed publishing or irreversible account changes. ## Recovery boundaries - If local FFmpeg is available, keep the native lane local. If it is unavailable, use the narrower hosted FFmpeg Actions only for operations their live schemas actually support, or use a composition lane for the missing operation. - If the source is private, local-only, expired, or inaccessible, stop before uploading it and request an approved input or explicit upload authorization. - If transcription confidence is low, mark uncertain words and reduce selection confidence rather than inventing precision. - If ChatCut is disconnected, continue with the native lane unless the user specifically requires a persistent project. ChatCut setup uses Aident Vault OAuth; there is no stateless platform-key path in the current integration. - If OpusClip reports revoked API access, an ineligible plan, or an invalid key, treat it as a deterministic provider-access failure. Do not retry or describe the integration as ready merely because catalog search, Vault status, or preflight succeeded. - If a render or provider job fails transiently, retry only within its documented contract. Preserve valid partial results and disclose missing coverage. - If the harness cannot review temporal video or audio, downgrade to proxy or technical-only review, label unsupported checks, and require human approval before any publish or schedule action. Do not silently substitute metadata for semantic review. - If overlapping subtitle windows disagree with rendered-audio turn boundaries, trust the inspected source and final-audio evidence, repair the cut, and regenerate captions from the revised final timeline. - If a rendered subtitle frame exposes a literal escape sequence, missing glyph, overflow, or unintended line break, fix the subtitle source and rerender every affected master rather than approving around the defect. - If the decoded sample aspect ratio is not exactly 1:1 for an intended square-pixel vertical master, repair the composition and rerender before delivery. - If a freeze warning appears only after placing moving source video inside a largely static designed canvas, inspect the flagged frames and compare the same source interval. Record a low-motion note when supported, but keep smooth playback on the human review checklist. - If fewer than three candidates pass the quality gates, explain why and return only the passing clips instead of padding the default batch. - If no candidate passes the quality gates, explain why and return no clip instead of lowering the standard. - If a social account is missing or ambiguous, stop before scheduling or publishing. - If analytics access is ineligible or delayed, record the gap and use only available comparable evidence. ## Output contract Return: 1. the experiment brief and selected execution lane; 2. the source audit and transcript confidence; 3. the ranked candidate table and rejection reasons, with three to five selected highlights by default when enough candidates pass; 4. selected edit briefs, edit decision lists, review mode, artifact status, rendered verification evidence, and every unresolved human check; 5. verified export identities or paths and the publish manifest with external-write status; 6. the platform-specific metric table with observation windows; 7. findings, uncertainties, and the single next experiment. Never describe a draft as rendered, a render as published, provider acceptance as completion, or a correlation as a causal result. ## Research basis The workflow synthesizes the linked Clip Bot use case with official platform guidance, current practitioner sharing, and a pinned audit of OpenShorts' non-UI pipeline. Read [the OpenShorts core lessons](references/openshorts-core-lessons.md) for the adopted architecture and exclusions, and [the research notes and sources](SOURCES.md) when adapting the rubric, platform rules, or analytics interpretation.