Find and rank the strongest private North American AI infrastructure companies serving mid-market customers for VC sourcing and diligence.
Find and rank the strongest private North American AI infrastructure companies serving mid-market customers for VC sourcing and diligence.
--- name: vc-ai-infrastructure-scout description: Find and rank the strongest private North American AI infrastructure companies serving mid-market customers for VC sourcing and diligence, then deliver the result to Airtable, Google Sheets, or Lark. --- # VC AI Infrastructure Scout Turn a VC thesis into a source-backed company screen and a usable diligence table. This Skill is for associates, principals, and research teams sourcing private AI infrastructure companies before deeper diligence. ## Loadout capabilities Use <integration-tag>composio:exa_tools</integration-tag> and select its `exa_search` operation to build the candidate pool and collect official or reputable public evidence. Prefer precise company-category searches with concise highlights over full-page text. Use <action-tag>api:akta_api:company_search</action-tag> to resolve each candidate from a precise name or official domain to a canonical private-company identity. It requires the connected Akta integration. After the usage gate below, use <action-tag>api:akta_api:company_enrichment</action-tag> to retrieve only the structured private-company sections required by this Skill. Write the final deliverable through exactly one destination selected by the user: - Airtable: <integration-tag>composio:airtable_tools</integration-tag> - Google Sheets: <integration-tag>composio:googlesheets_tools</integration-tag> - Lark Base or Lark Sheets: <integration-tag>cli:lark</integration-tag> After the user chooses a destination, find the appropriate capability inside that integration, inspect its current input schema and Vault readiness, and use it only for the confirmed target. Do not treat a chat-only table or local CSV as an equivalent deliverable unless the user explicitly requests one. For Airtable, interpret `a new one` as a request for a fresh isolated deliverable, not a request for the user to supply an internal workspace identifier. Prefer a new base when Loadout returns the required workspace ID. When it does not, list accessible bases and apply the Airtable fallback in step 7 without asking the user for a `wsp` ID. ## Approval gate 0: confirm the complete brief before any Action Do not call Exa, Akta, or a destination integration before this gate is approved. First show the user one compact brief containing all of the following: - **Audience:** VC associate or principal. - **Company profile:** private, independent North American AI infrastructure vendors serving mid-market customers. - **Stage:** Seed through Series C by default, or the user's stated range. - **Geography:** headquarters in the United States or Canada by default. - **Company types to find:** compute and inference; model serving and deployment; orchestration and agent runtime; observability, evaluation, and security; AI data infrastructure. - **Include:** infrastructure is the core product, official domain exists, positive geography and independence evidence, and credible mid-market customer proof. - **Exclude:** public companies, acquired subsidiaries, services-only consultancies, investment vehicles, chip-only companies without a usable software platform, and ambiguous identities. - **Mid-market proof:** explicit segment language, a named customer with roughly 50 to 1,000 employees, or a case study or credible third-party source proving adoption. - **Scale:** 12 to 20 discovered candidates, no more than eight proposed for paid enrichment, and a final top five. - **Deliverable:** Airtable, Google Sheets, or Lark; preferred new or existing target; and the existing base, sheet, or document URL when applicable. For Airtable, explain that `new` produces a fresh deliverable: a new base when supported, otherwise three new dated tables inside the sole writable base. - **Output structure:** `Shortlist`, `Screening Log`, and `Methodology` tables or tabs. - **Usage boundary:** discovery and identity resolution happen only after this approval; metered work always gets a usage preview, but separate price confirmation is required only when projected metered usage for the current turn exceeds $1; no contacts or outreach. End with a direct confirmation request such as: > Please confirm this research brief and choose the destination. Once confirmed, I will search these company categories, verify each match, preview paid enrichment usage, and prepare the selected table. If the user changes any item, restate the complete revised brief and wait again. If the selected destination is not connected, help the user connect it through Aident Loadout and pause. A partial reply that does not identify the destination is not approval to start. Treat the user's destination choice as authorization for the final write described in this brief. A reply such as `Airtable, a new one for me` is complete approval: do not ask for a workspace ID or repeat the destination question. ## Definitions and evidence rules Use these statuses exactly: - `discovered`: found in public research but not yet identity-resolved; - `identity-confirmed`: name, root domain, and Akta identity agree; - `provisional-fit`: public evidence suggests the thesis and customer screen fit; - `review`: identity, ownership, location, stage, or customer proof remains ambiguous; - `rejected`: a hard gate failed; - `qualified`: enrichment and cited evidence verify every hard gate. Do not describe a candidate as confirmed or qualified merely because Akta returned an exact domain match. Identity confirmation proves identity only. An API, public documentation, free trial, public pricing, or self-serve onboarding proves accessibility, not that a company serves mid-market customers. Label that evidence `mid-market accessible`. To claim `serves mid-market`, require at least one of: 1. an explicit first-party customer-segment statement; 2. a named customer whose current size is approximately 50 to 1,000 employees; or 3. a case study or credible third-party source showing mid-market adoption. When the user's screen requires actual mid-market service, accessibility alone does not pass the gate. Absence of an acquisition flag is not proof of independence. Require a positive source for headquarters and a positive ownership or parent-company finding. If either cannot be verified, use `Unknown`, mark the candidate `review`, and do not enrich it. ## Workflow ### 1. Search the market by taxonomy After approval gate 0, search every relevant taxonomy bucket rather than using one generic query: 1. compute, inference, and optimization; 2. model serving, deployment, and developer platforms; 3. orchestration, agent runtime, and workflow infrastructure; 4. observability, evaluation, governance, and security; 5. AI data, retrieval, vector, and data-pipeline infrastructure. Search for established category leaders and emerging vendors without hardcoding company names. Use `category: company`, concise highlights, and enough results to produce 12 to 20 unique candidates. Prefer official sites, product documentation, customer pages, security pages, and dated first-party announcements. Use reputable reporting when first-party evidence is unavailable. Use at most one primary Exa search per taxonomy bucket. If coverage is still missing, use no more than three compact recovery searches across the entire discovery phase. Do not rerun a broad search merely because its payload was too large; request fewer results and shorter highlights in the first call. The full turn may use at most 12 Exa Actions by default, including follow-up evidence searches. Combine several unresolved companies in one evidence query when practical. If the cap is insufficient, stop with explicit coverage gaps instead of silently expanding usage. For every candidate, record the suspected official domain, taxonomy bucket, source page, source date, and the exact claim the source supports. Deduplicate by normalized root domain. Run a coverage check and state which taxonomy buckets are represented or missing before continuing. ### 2. Resolve identities efficiently Call Akta company search once per deduplicated candidate, using the official domain when available. Resolve independent candidates in bounded parallel batches that respect live rate limits, preserve input order, and keep one audit record per company. Resolve no more than 20 unique candidates by default. Retry only ambiguous or exact-domain-miss cases, with at most three identity retries across the turn. Do not repeat a successful identity lookup for enrichment preparation or reporting. If multiple Akta results share the exact root domain, never choose one by rank alone. Record every candidate UUID, mark the company `review`, compare public identity and category evidence, and proceed only after authoritative disambiguation or user confirmation. Identity resolution can produce `identity-confirmed`, `review`, or `rejected`. It cannot produce `qualified`. ### 3. Verify public hard gates and reduce the pool For every identity-confirmed candidate, verify: 1. positive evidence that the company is private and independent; 2. positive evidence that headquarters are inside the requested geography; 3. evidence that AI infrastructure is the core product; 4. stage fit, using a cited public source rather than inference; and 5. actual mid-market customer proof under the definition above. Use `Unknown` when a fact is missing. Never infer funding stage, revenue, valuation, headcount, customers, ownership, or growth. Apply the public screen before proposing paid enrichment. Narrow the pool to no more than the final shortlist size plus three, eight by default. Show at most eight `provisional-fit` candidates in chat, each with its exact supporting source pages and evidence dates. Summarize review and rejected candidates by count and reason; retain their full detail for the `Screening Log`. ### 4. Preview usage and apply the per-turn price gate For the exact proposed enrichment inputs, show: - companies and canonical Akta identities; - number of enrichment Actions; - exact requested sections; - any optional lookup proposed; - current Loadout preflight result; - total expected usage; and - a statement that charges are based on each Action's actual usage. Request exactly these sections: `firmographic,location,industry,product_offering,technology,customer_profile,strategic_signal,digital_presence,company_hierarchy` Do not use unsupported aliases such as `overview` or `web_presence`. If every company uses identical sections and one representative preflight returns exact pricing under the same quote and policy semantics, preflight one representative and multiply by the number of companies. If pricing can vary by company or input, preflight each input. Never hardcode a price. Apply one cumulative metered-usage gate to the current user turn: - If the projected total for all additional metered Actions in the turn is $1 or less, show the usage preview as information and proceed without asking the user to confirm the price again. - If the projected total is more than $1, show the exact plan and total, then wait for explicit approval before the first metered Action. - If exact preflight is unavailable, pricing is open-ended, or the cumulative total would cross $1 after work has started, pause before the next metered Action and ask for approval. - Never split a plan across batches or turns to avoid the threshold. A prior approval applies only to the exact companies, sections, optional lookups, and total shown. This gate does not authorize contact discovery, outreach, or unrelated Actions. ### 5. Enrich and qualify After approval, enrich only the proposed `provisional-fit` candidates. Normalize: - official company name, official domain, and Akta UUID; - private and independent status, parent or hierarchy, and confidence; - headquarters and employee range; - funding stage from cited public evidence; - industry and infrastructure category; - primary products; - core technology or technical differentiation; - mid-market customer evidence; - recent strategic signals; - source URLs and evidence dates. A candidate becomes `qualified` only when headquarters, independence, core-product fit, and actual mid-market customer proof are all verified. Otherwise move it to `review` or `rejected` and state why. If fewer than five qualify, return fewer; do not relax a hard gate to fill the table. ### 6. Score for VC review Score every qualified company out of 100 using the same rubric: | Dimension | Weight | What earns a high score | | --- | ---: | --- | | Thesis fit | 25 | AI infrastructure is the core product, not a feature. | | Mid-market evidence | 20 | Strong cited proof of adoption by the target customer segment. | | Product and technology edge | 20 | Specific, defensible technical differentiation. | | Strategic momentum | 15 | Recent product, partnership, hiring, expansion, or market signal. | | Venture and stage fit | 10 | Fits the requested stage and retains credible scale potential. | | Evidence confidence | 10 | Strong identity, fresh sources, and few unresolved fields. | If the user provides fund strategy, ownership targets, or check size, incorporate them. Do not require check size by default. Break ties by evidence confidence, then strategic momentum. Label interpretation as `VC read`. ### 7. Preview and write the deliverable Before market research begins, inspect the selected destination's live Actions, connection, and writable containers. Resolve the exact target with these rules: 1. If the user selected an existing target, verify and use it. 2. If the user requested a new Airtable deliverable and Loadout provides the workspace ID required by the create-base Action, create a new base named `VC AI Infrastructure Scout - YYYY-MM-DD` with the three required tables. 3. If Loadout cannot provide that workspace ID, never ask the user for a `wsp` ID. List accessible Airtable bases. When exactly one base has create or edit permission, use it automatically and create three new tables named `VC Scout Shortlist YYYY-MM-DD`, `VC Scout Screening YYYY-MM-DD`, and `VC Scout Methodology YYYY-MM-DD`. If any name already exists, append the current `HHmm` time or the smallest available numeric suffix. 4. When multiple writable bases exist and no base was previously selected, ask one short base-choice question. This is the only Airtable fallback that requires another user reply. Creating new dated tables inside the sole writable base satisfies the user's request for a new Airtable deliverable, but describe it accurately as new tables rather than a new base. Do not expose connector-internal workspace requirements, fall back to browser control, or call a direct provider API when Aident-only execution was selected. Prepare the three destination tables or tabs defined in [OUTPUT-TEMPLATE.md](OUTPUT-TEMPLATE.md). Immediately before the external write, show: - selected integration and exact target; - whether the operation creates or updates the target; - table or tab names; - exact columns; and - row counts. If approval gate 0 explicitly authorized this exact final write, proceed after showing the preview and do not ask again. Otherwise obtain confirmation. Then use the selected integration's live capability and schema, write the deliverable, read back the written rows, and return the live artifact link. Never claim delivery from a successful request alone; verify the authoritative destination state. Track usage from the Actions executed for this turn. Prefer the usage, credits, cost, and execution identifiers returned by each Action and sum only those records. If audit lookup is needed, query by those execution identifiers or the current trace or agent session. Never report a whole-day, account-wide, or integration-wide audit total as this turn's cost. When trace-scoped usage cannot be isolated, report it as `Unknown` rather than guessing. ## Quality rules - Cite every material claim in the row where it appears. - Use the exact source page and evidence date for headquarters, customer proof, category, technology, stage, and strategic signals. A homepage-only citation is insufficient for a specific claim. - Prefer first-party sources for products, technology, location, customers, and announcements. - Keep facts separate from interpretation and label interpretation `VC read`. - Keep one canonical row per company and one canonical domain per row. - Make uncertainty visible with `High`, `Medium`, or `Low` confidence. - Do not add people, personal contact details, outreach copy, or send steps. - Do not describe the result as investment advice or a complete market map. ## Recovery - If company search is ambiguous, retry once with the official domain. If ambiguity remains, mark `review`, record all candidate UUIDs, and skip enrichment. - If enrichment omits a field, keep `Unknown`; do not replace it with an unsourced guess. - If a source conflicts with Akta, present both, prefer the newest authoritative source, and lower confidence. - If an Action fails after approval, inspect its Loadout audit record before retrying to avoid duplicate usage. - If Airtable base creation is unavailable, apply the sole-writable-base fallback above and continue without requesting the same authorization twice. - If the destination write partially succeeds, read back the target, write only missing rows, and verify again. - If fewer than five candidates survive, deliver the valid set and recommend the single most useful thesis adjustment. ## Completion check The Skill is complete only when the selected Airtable, Google Sheet, or Lark artifact has been read back successfully and a VC can answer these questions in under two minutes: 1. Which companies deserve attention? 2. Why do they fit this exact thesis and stage? 3. What evidence proves the mid-market claim? 4. What changed or matters now? 5. What should we verify next?