Multimodal Media Comprehension
Describe and analyze authorized image, audio, and video inputs with explicit modality limits and uncertainty handling.
Works with
Multimodal Media Comprehension
When to use
Describe and analyze authorized image, audio, and video inputs with explicit modality limits and uncertainty handling.
Workflow
- Confirm the user owns the input or has permission to process it. Collect the source material, audience, output format, and acceptance criteria.
- Separate observed facts and user-provided constraints from inference. Ask for missing information instead of inventing it.
- Define the requested observations, extraction fields, coordinate system, confidence needs, and unsupported modalities before analysis.
- Use <skill-tag>skill:820e8b64-9a9c-5c33-8337-ddd348f77f3a</skill-tag> for the connected analysis workflow. Return evidence, confidence, and limitations separately.
Output contract
Return the input assumptions, the approved plan, the generated or analyzed result, uncertainty and limitations, and a concise verification checklist. Preserve source links and provider-returned identifiers when available. Stop if the required connected capability is unavailable; do not substitute local scripts, package installation, or embedded credentials.
Safety and review
Do not upload private or sensitive media without explicit authorization. Do not claim certainty beyond the evidence. Require review before externally viSkill · MiniMax Media Generation and license
This Aident-native re-authoring is based on the pinned upstream source, discovered through ModelScope, and used under MIT. It preserves the upstream workflow goal while removing local executable, credential, and package assumptions.