Talking-Head Lip Sync
Lip-sync supplied speech to a supplied portrait or talking-head video with a supported audio-to-video model.
Works with
name: talking-head-lip-sync description: "Lip-sync supplied speech to a supplied portrait or talking-head video with a supported audio-to-video model."
Talking-Head Lip Sync
Lip-sync supplied speech to a supplied portrait or talking-head video with a supported audio-to-video model.
Required Aident Actions
- <action-tag>direct:fal:fal_list_audio_to_video_models</action-tag> (required inputs: inspect the current schema): List active Fal audio-to-video models with current platform-key pricing, descriptions, required inputs, defaults, choices, documentation links, and pagination cursors. Use this before fal_audio_to_video when choosing...
- <action-tag>direct:fal:fal_audio_to_video</action-tag> (required inputs: model): Run an active Fal audio-to-video model. Call fal_list_audio_to_video_models before selecting another model to inspect current pricing, required fields, defaults, and model differences. Prompt and text shortcuts...
Use Aident Loadout to read the current Action schema before constructing inputs. CheAction · List Audio To Video Modelson is billable or mutating, run Aident preflight, show the affected target and quoted cost or risk, and wait for explicit user confirmation before execution.
Workflow
- Confirm the requested outcome, source material, destination, audience, constraints, and acceptaAction · Audio To Videoow to create a concrete plan. Resolve ambiguity before invoking an Action.
- Select only the Aident Actions whose documented effect directly advances the requested outcome. Do not invoke every listed Action by default.
- Inspect the current schema and prepare the minimum valid input for each selected Action.
- Preflight each selected Action. Execute it only after any required confirmation, in dependency order, and carry returned IDs or asset URLs into later steps.
- Verify the returned IDs, URLs, statuses, or artifacts against the acceptance criteria. Report partial completion precisely and do not repeat paid or mutating calls blindly.
Source-Derived Guidance
- Confirm rights to the likeness and voice, then collect the speech audio, portrait or source video, target aspect ratio, and identity details that must be preserved.
- Use direct:fal:fal_list_audio_to_video_models to select a currently supported lip-sync or talking-head model compatible with the supplied asset types.
- Use direct:fal:fal_audio_to_video with the exact supplied speech and visual asset; do not clone a voice, generate a new script, or imply consent that was not provided.
- Verify audio duration, mouth synchronization, face identity, frame stability, and playback before returning the video asset URL.
Execution Boundaries
- Treat upstream provider-specific commands as background knowledge only. Execute the workflow through the exact Aident Actions above.
- Never request raw credentials in chat. Use Aident Vault connection flows for required accounts.
- Preserve user-provided wording, brand constraints, rights restrictions, and target identifiers. Do not invent authorization.
- Pass assets by Aident asset ID or supported URL fields. Do not expose caller-local filesystem paths to hosted Actions.
- Stop when a required integration is disconnected, preflight rejects the input, the target is ambiguous, or the user declines a required confirmation.
Output
Return the execution plan, selected Action names, confirmed targets, Aident result identifiers or asset URLs, verification evidence, and any remaining blocked step.
Attribution
This Skill adapts the reviewed upstream workflow. See for the pinned source and for the preserved license.