Podcast Visualizer Video
Turn podcast audio into a transcript-informed visualizer video.
@Aident5,85437Updated Aug 12, 2026
Works with
name: podcast-visualizer-video description: "Turn podcast audio into a transcript-informed visualizer video."
Podcast Visualizer Video
Turn podcast audio into a transcript-informed visualizer video.
Required Aident Actions
- <action-tag>direct:fal:fal_speech_to_text</action-tag> (required inputs: inspect the current schema): Run an active Fal speech-to-text model. Defaults to fal-ai/speech-to-text/turbo. Call fal_list_speech_to_text_models before selecting another model to inspect current pricing, required fields, defaults, and model...
- <action-tag>direct:fal:fal_text_to_video</action-tag> (required inputs: prompt): Run an active Fal text-to-video model. Defaults to bytedance/seedance-2.0/fast/text-to-video. Call fal_list_text_to_video_models before selecting another model to inspect current pricing, required fields, defaults,...
- <action-tag>composio:shotstack_tools:shotstack_render_video</action-tag> (required inputs: timeAction · Speech To Text. Use when you have defined a timeline and output settings and want to start rendering.
- <action-tag>composio:shotstack_tools:shotstack_get_render_status</action-tag> (required inputs: id): Tool to retrieve the current status and details of a Shotstack render job bAction · Text To Videofailed, typically after creating a render with SHOTSTACK_RENDER_VIDEO.
Use Aident Loadout to read the current Action schema before constructing inputs. Check the required integration connection in Aident Vault. If an Action is billable or mutatingAction · Render Video, and wait for explicit user confirmation before execution.
Workflow
- Confirm the requested outcome, source material, destination, audience, constraints, and acAction · Get Render Statuscrete plan. Resolve ambiguity before invoking an Action.
- Select only the Aident Actions whose documented effect directly advances the requested outcome. Do not invoke every listed Action by default.
- Inspect the current schema and prepare the minimum valid input for each selected Action.
- Preflight each selected Action. Execute it only after any required confirmation, in dependency order, and carry returned IDs or asset URLs into later steps.
- Verify the returned IDs, URLs, statuses, or artifacts against the acceptance criteria. Report partial completion precisely and do not repeat paid or mutating calls blindly.
Source-Derived Guidance
- Collect the podcast audio asset, title, speakers, brand style, target dimensions, and desired excerpt length.
- Use direct:fal:fal_speech_to_text to produce a transcript and identify the approved excerpt and key themes.
- Create a timed visual plan with readable titles, speaker treatment, and motion that does not obscure the content.
- Use direct:fal:fal_text_to_video to generate any approved visual segments.
- Use composio:shotstack_tools:shotstack_render_video to combine audio, visuals, and titles, then verify synchronization and output quality.
Execution Boundaries
- Treat upstream provider-specific commands as background knowledge only. Execute the workflow through the exact Aident Actions above.
- Never request raw credentials in chat. Use Aident Vault connection flows for required accounts.
- Preserve user-provided wording, brand constraints, rights restrictions, and target identifiers. Do not invent authorization.
- Pass assets by Aident asset ID or supported URL fields. Do not expose caller-local filesystem paths to hosted Actions.
- After Shotstack accepts a render, pass its returned ID to composio:shotstack_tools:shotstack_get_render_status at moderate intervals until a terminal state. Return a video URL only after the render completes.
- Stop when a required integration is disconnected, preflight rejects the input, the target is ambiguous, or the user declines a required confirmation.
Output
Return the execution plan, selected Action names, confirmed targets, Aident result identifiers or asset URLs, verification evidence, and any remaining blocked step.
Attribution
This Skill adapts the reviewed upstream workflow. See for the pinned source and for the preserved license.