Voice AI Systems
Build speech recognition, synthesis, and voice-transformation workflows.
Works with
name: voice-ai-systems description: "Build speech recognition, synthesis, and voice-transformation workflows."
Voice AI Systems
Build speech recognition, synthesis, and voice-transformation workflows.
Required Aident Actions
- <action-tag>direct:fal:fal_speech_to_text</action-tag> (required inputs: inspect the current schema): Run an active Fal speech-to-text model. Defaults to fal-ai/speech-to-text/turbo. Call fal_list_speech_to_text_models before selecting another model to inspect current pricing, required fields, defaults, and model...
- <action-tag>api:elevenlabs:text_to_speech</action-tag> (required inputs: actionName, voice_id, text): Use this integration action for text to speech operations in ElevenLabs API Documentation.
- <action-tag>api:elevenlabs:speech_to_speech</action-tag> (required inputs: actionName, voice_id, audio): Use this integration action for speech to speech operations in ElevenLabs API Documentation.
UsAction · Speech To Textore constructing inputs. Check the required integration connection in Aident Vault. If an Action is billable or mutating, run Aident preflight, show the affected target and quoted cost or risk, and wait for explicit user confirmation before execution.
Workflow
Action · Text To Speechtination, audience, constraints, and acceptance criteria. 2. Apply the source-derived guidance below to create a concrete plan. Resolve ambigAction · Speech To Speecht Actions whose documented effect directly advances the requested outcome. Do not invoke every listed Action by default. 4. Inspect the current schema and prepare the minimum valid input for each selected Action. 5. Preflight each selected Action. Execute it only after any required confirmation, in dependency order, and carry returned IDs or asset URLs into later steps. 6. Verify the returned IDs, URLs, statuses, or artifacts against the acceptance criteria. Report partial completion precisely and do not repeat paid or mutating calls blindly.
Source-Derived Guidance
- Keep the workflow scoped to the source-derived objective: Build speech recognition, synthesis, and voice-transformation workflows.
- When needed, use direct:fal:fal_speech_to_text only for this supported effect: Run an active Fal speech-to-text model. Defaults to fal-ai/speech-to-text/turbo. Call fal_list_speech_to_text_models before selecting another model to inspect current pricing,...
- When needed, use api:elevenlabs:text_to_speech only for this supported effect: Use this integration action for text to speech operations in ElevenLabs API Documentation.
- When needed, use api:elevenlabs:speech_to_speech only for this supported effect: Use this integration action for speech to speech operations in ElevenLabs API Documentation.
- Verify the returned artifact or provider state against the requested acceptance criteria before reporting completion.
Execution Boundaries
- Treat upstream provider-specific commands as background knowledge only. Execute the workflow through the exact Aident Actions above.
- Never request raw credentials in chat. Use Aident Vault connection flows for required accounts.
- Preserve user-provided wording, brand constraints, rights restrictions, and target identifiers. Do not invent authorization.
- Pass assets by Aident asset ID or supported URL fields. Do not expose caller-local filesystem paths to hosted Actions.
- Stop when a required integration is disconnected, preflight rejects the input, the target is ambiguous, or the user declines a required confirmation.
Output
Return the execution plan, selected Action names, confirmed targets, Aident result identifiers or asset URLs, verification evidence, and any remaining blocked step.
Attribution
This Skill adapts the reviewed upstream workflow. See for the pinned source and for the preserved license.