Extracts structured data from web pages by initiating an extraction job and polling for completion; requires a natural language `prompt` or a JSON `schema` (one must be provided).
Write actionRisk level 2 of 5Aident-managed access
A list of URLs from which to extract data (maximum 10 URLs while in beta). Wildcards (e.g., `https://example.com/blog/*`) can be used for crawling multiple pages under a specific path. Note: You can also pass a single URL as 'url' (singular) which will be automatically converted…
promptstring
Natural language query for information to extract from URL content. E.g., 'Extract the company mission, whether it supports SSO, etc.'. At least one of 'prompt' or 'schema' must be provided.
schemaobject
JSON object (dictionary) defining the desired structure for extracted data. Must be a valid JSON Schema object with properties and types. At least one of 'prompt' or 'schema' must be provided.
showSourcesboolean
When true, the sources used to extract the data will be included in the response as `sources` key.
ignoreSitemapboolean
Bypasses sitemap.xml during scanning.
scrapeOptionsobject
Advanced scraping configuration.
enableWebSearchboolean
If `True`, allows crawling links outside initial domains in `urls`; if `False`, restricts to same domains.
ignoreInvalidURLsboolean
Proceeds with valid URLs, returning invalid ones separately.
includeSubdomainsboolean
Extends scanning to subdomains.
Observable output
data
Required
Data from the action execution
errorstring
Error if any occurred during the execution of the action
successfulboolean
Required
Whether or not the action execution was successful or not