Aident AI

How to Audit AI Search Citations Across Kimi, Doubao, and DeepSeek
An AI search citation audit should run the same real buyer questions across each engine, preserve every engine's answer and observable sources, verify the cited URLs, and repeat the sample before reporting a trend. Keep brand mentions, source citations, verified links, and downstream visits as separate fields.
That discipline matters when comparing Kimi, Doubao, and DeepSeek. The engines do not expose one universal ranking contract. A single prompt can return an answer with structured sources in one engine, an answer without observable sources in another, and a still-pending task in a third. Collapsing those outcomes into one visibility score hides the evidence a marketer or agent operator actually needs.
If you are choosing topics before measuring citations, start with the separate AI agent search and product growth research. If you need broader public conversation evidence, use the cross-platform social listening workflow. This guide owns the narrower job of producing a repeatable citation receipt across three AI search engines.
The Short Answer
Use this evidence ladder for every prompt and engine:
State | What you observed | What you may claim |
|---|---|---|
Absent | The brand and its domain do not appear | Absent in this run |
Mentioned | The answer names the brand without linking its site | Mentioned, not cited |
Cited | The result exposes a URL from the brand's domain | Cited in this run |
Verified | The URL resolves, matches the claimed source, and was checked at capture time | Verified citation |
Visited | Owned analytics records a later visit from a supported source or campaign | Observable navigation, not automatic attribution |
Never call a mention a citation. Never call a citation a click. Never call a click a signup, payment, or revenue event without identity or campaign evidence that supports the join.
One run is a proof of the collection workflow, not a measurement. For a reporting baseline, repeat each prompt at least twice, preferably three times, and keep the full run history. The GEO Wiki citation-tracking playbook likewise separates mentions from citations and recommends reporting verified and unverified citations separately.
Start From a Fixed Prompt Panel
Build a small panel of questions that actual buyers ask before they know your brand. Include:
informational questions about the job;
commercial questions about approaches or categories;
comparison questions about tradeoffs;
migration or implementation questions; and
one branded control prompt, stored separately from the nonbranded panel.
Do not fill the panel with flattering brand-name questions. A brand may appear simply because the question supplied the answer. A useful nonbranded panel tests whether the engine reaches the brand or its content without that hint.
Version the panel. A minimal prompt record needs:
The promptId should remain stable. If the wording changes, advance panelVersion instead of silently rewriting history.
Discover the Current AI Search Action
Install or update Aident Loadout from the canonical setup guide:
Confirm the installed public CLI account and Vault state:
Search the live staging catalog by job:
Copy the exact public Action name returned by discovery and inspect its current schema:
On August 11, 2026, the discovered Action exposed one submit and one result operation for each engine:
Engine | Submit operation | Result operation | Required submit field | Required result field |
|---|---|---|---|---|
Kimi |
|
|
|
|
Doubao |
|
|
|
|
DeepSeek |
|
|
|
|
The current Action declared an open output schema, a level-1 risk classification, and a 0.975-credit flat quote for each submit. Exact result polls preflighted and executed at zero credits. Those are dated Aident Action observations, not permanent provider pricing claims.
Preflight Every Submit Before Dispatch
Preflight the exact engine, question, and language you intend to send:
Repeat for the Doubao and DeepSeek submit operations. Stop when an input is invalid, the required connection is missing, the quote is unavailable or above your ceiling, or the question includes private information that should not leave the approved system boundary.
Preflight is not execution. It validates the proposed request and prices it without creating the remote search task. Keep the preflight receipt with the eventual result so reviewers can confirm that the approved and executed inputs match.
Submit Once, Then Poll the Exact Task
An asynchronous search Action has two distinct steps:
Submit one reviewed question once.
Poll only the matching result operation with the returned task ID.
Do not submit the question again because the result is pending. Resubmission creates a new sample, can consume another paid call, and breaks the link between the approved request and its result.
Your run ledger should record:
Do not publish raw task IDs. They are transport identifiers, not useful reader evidence. Store a hash or protected reference when operational replay requires one.
What One Bounded Sample Showed
For one public, nonbranded English question, the three current engine operations produced materially different evidence:
Engine | Bounded sample result | Observable citation evidence | Aident state |
|---|---|---|---|
Kimi | Completed with a practical answer | No web-page records were returned in the result | Absent in this run |
Doubao | Still pending after bounded result polls | No answer or sources were available to classify | Unknown, not absent |
DeepSeek | Completed with an answer plus structured search records | Source URLs, titles, snippets, source sites, query indexes, and answer citation markers were returned | Absent in this run |
The question asked how an AI coding agent should connect to Gmail, GitHub, and Lark without putting API keys in prompts. Kimi returned a direct answer about separating credentials from the model context, but its result exposed no web pages. DeepSeek returned an answer plus a structured source collection that included official Lark documentation and third-party pages. Doubao remained pending within the audit's bounded polling window.
This sample does not rank the engines. It proves that a single normalized score would have discarded important distinctions:
completed answer versus pending task;
answer text versus observable source records;
brand absence versus an unclassifiable pending result; and
cited source record versus a source URL that has actually been fetched and verified.
It also does not establish a visibility trend. Three submits can validate the collection contract, but they cannot support a market-share, citation-share, or causality claim.
Normalize Evidence, Not Engine Behavior
Do not force every provider response into an invented universal ranking model. Normalize only the observable fields needed for review:
Preserve the raw result separately. The normalizer will change as engine outputs change. A raw artifact lets you reprocess old samples without pretending the old provider response matched today's schema.
Verify Every Citation Before Counting It
A source record is an observation, not yet a verified citation. For each cited URL:
Resolve redirects with a normal web request.
Record the final canonical URL and HTTP status.
Confirm the page exists and supports the answer's nearby claim.
Separate dead, blocked, redirected, and content-mismatched URLs.
Store the verification timestamp.
The research paper Evaluating Verifiability in Generative Search Engines found that citation correctness and completeness are distinct properties worth evaluating. That is why this workflow stores the answer, source record, and URL verification as separate evidence.
Do not rewrite a dead source into a success because the domain looks authoritative. A broken or unsupported citation is a finding.
Repeat Before You Report a Trend
Run each prompt more than once because answers and cited sources vary. Keep the exact engine, language, market, timestamp, Action contract, and prompt-panel version with every sample.
A practical reporting table is:
Metric | Numerator | Denominator | Required caveat |
|---|---|---|---|
Mention rate | Completed runs that mention the brand | Completed classifiable runs | Exclude pending and failed runs |
Citation rate | Completed runs with at least one brand-domain URL | Completed classifiable runs | A mention alone does not count |
Verified citation rate | Completed runs with at least one verified brand-domain URL | Completed classifiable runs | Record verification policy |
Prompt coverage | Prompts with any verified citation | Prompts with enough completed repetitions | State required repetitions |
Engine coverage | Engines with any verified citation | Engines tested under the same panel | Do not merge language slices |
Report sample size beside every percentage. Keep Chinese and English prompt panels separate unless you have an explicit weighting model supported by the business question.
The open-source RedFox GEO Analyzer demonstrates the useful collection pattern of sending the same questions to Doubao, Kimi, and DeepSeek, then retaining mentions, citations, and competitor evidence. Its composite score is one possible reporting layer, but an Aident audit should preserve the underlying fields so reviewers can choose or reject a weighting scheme later.
Join Search and Product Evidence Carefully
AI search visibility, Google demand, and product navigation answer different questions:
AI answer samples show whether an engine mentioned or cited a source.
Google Search Console shows search queries, impressions, clicks, and positions for owned web properties.
Product analytics shows observable navigation and events after a visitor reaches Aident.
Billing or CRM systems can support signup, payment, or revenue attribution only when the join is valid.
For July 13 through August 9, 2026, Aident's current Search Console extract contained 694 clicks and 12,667 impressions overall. A directional nonbrand filter contained 104 clicks and 6,948 impressions, compared with 34 clicks and 5,743 impressions in the prior equal window. That filter is useful for direction, not a formal branded-search taxonomy.
In a separate PostHog query for the same current window, 1,227 distinct blog visitors included 14 people who later reached a qualifying deeper product path within seven days, about 1.14 percent. This is observable navigation under one query definition. It does not prove that a blog post caused a signup, payment, or revenue event, and it is not directly comparable with cohorts that use different dates, destinations, or denominators.
Use the evidence together to choose the next experiment. Do not fuse it into a causal funnel that the identifiers cannot support.
An Audit Prompt for Codex or Claude Code
The success condition is a reproducible evidence ledger, not the highest-looking score.
Failure Matrix
Failure | What it means | Next action |
|---|---|---|
Discovery returns a different Action | The live catalog changed | Inspect the new schema and restart preflight |
Submit quote is unavailable or above the ceiling | Spend is not approved | Narrow the panel or request explicit credit approval |
Result remains pending | The task is not classifiable yet | Keep it pending and poll later within the approved window |
One engine returns no source records | Citation evidence is unavailable for that result | Record answer-only evidence; do not infer citations |
A source URL is dead or mismatched | The citation is not verified | Record the failure and preserve the raw source record |
Brand is named without its domain | The run contains a mention, not a citation | Keep the states separate |
Repeated runs disagree | The answer is variable | Report the distribution and sample size, not a single screenshot |
Analytics has no valid user or campaign join | Conversion attribution is unavailable | Leave signup, payment, and revenue fields null |
Set Up the Audit
Follow https://aident.ai/SETUP.md
Set up Aident Loadout and audit one approved prompt panel
Sources
Aident Loadout setup guide, reviewed August 11, 2026.
Live Aident Loadout catalog, Action schema, preflight receipts, and bounded Kimi, Doubao, and DeepSeek sample, inspected August 11, 2026.
RedFox GEO Analyzer, reviewed August 11, 2026.
Evaluating Verifiability in Generative Search Engines, reviewed August 11, 2026.
AI Citation Tracking, GEO Wiki, reviewed August 11, 2026.
Refresh this guide when the current Action changes its engine operations, input schema, result envelope, risk, pricing, or connection contract, or when repeated prompt-panel data replaces the one-run collection proof.



The one tool
for every tool
your agent needs.
Give any AI agent real capabilities in seconds. Connect 1,000+ tools once, skip the setup headache, and let your agents execute.
