Aident AI

How to Build a Customer Interview Synthesis Agent
A useful customer interview synthesis agent does not replace a researcher with a summary button. It turns authorized recordings into a reviewable evidence ledger, groups only genuinely repeated findings, preserves contradictions, and keeps a person responsible for the final interpretation.
The workflow in this guide uses Aident Loadout to connect an agent such as Codex or Claude Code to Gladia for transcription and Lark for the evidence ledger and decision brief. The same operating contract also works when transcripts already exist. The important part is not the model. It is the chain from every claim back to the exact interview evidence that supports it.
What the Agent Should Produce
Start with a bounded study, not a folder of recordings and the instruction "find insights." Define the research question, participant segment, interview set, and decision before processing anything.
For example:
The finished run should contain:
one immutable source ID for every interview;
short verbatim quotes with speaker and timestamp pointers;
atomic observations that make one claim each;
themes with independent-interview counts, not raw quote counts;
contradictions, outliers, and missing segments;
a decision brief that separates evidence from interpretation; and
a human review state for every theme that could influence product work.
Evidence-linked synthesis matters because fluent output can look more certain than its sources justify. Recent work on evidence-bounded customer research emphasizes source attribution and abstaining when evidence is missing. Research on interview-informed agents also warns that population-level patterns do not make a model an accurate stand-in for an individual participant. Use AI to organize real interviews, not to manufacture customers.
Set Up Aident Loadout
Give your agent the canonical setup instruction exactly as written:
Then have it confirm authentication and connected accounts:
Ask it to discover the current Gladia transcription and Lark document or Base Actions by job, inspect each schema, and preflight the exact request before execution. Do not paste capability names from an old run into a permanent prompt because catalog names, required fields, risk controls, and pricing can change.
On August 11, 2026, the live Gladia transcription Action accepted one public audio or video URL and quoted 2 Aident credits per call. The selected Lark document and Base record writes quoted zero credits, but they are still mutations that require a reviewed destination and explicit acknowledgement. Treat these observations as a preflight example, not permanent pricing.
If you are new to this connection model, first read How to Connect Claude Code and Codex to Real-World Tools. The existing meeting recording to Lark action-items guide is the better workflow when one meeting needs tasks rather than cross-interview research synthesis.
Step 1: Create an Interview Manifest
Do not let the agent discover arbitrary recordings from a shared drive. Give it a reviewed manifest containing only interviews approved for this study.
Use opaque participant identifiers rather than names. Keep consent and retention fields beside the media reference so they cannot disappear during handoff. If a recording falls outside the consent scope, has no source ID, or lacks a valid retrieval URL, mark it unavailable and stop. Missing interviews are not negative evidence.
Decide which sensitive fields must never enter the synthesis. Redact personal contact details and unrelated health, financial, employment, or account information before theme extraction. Do not ask the model to remember raw interviews after the approved retention window.
Step 2: Transcribe Each Interview Separately
Gladia's current pre-recorded workflow accepts audio or video and can return speaker-aware transcription. Its documentation recommends speaker diarization when the task depends on knowing who said what. Keep each interview as a separate transcription job so a quote cannot silently migrate between participants.
Ask the agent to preflight one representative input first:
After approval, store the raw transcription result before any summarization. Normalize it into utterances:
Check a sample against the recording, especially product names, acronyms, and the exact sentences later selected as evidence. A transcript is a derived artifact, not ground truth. If speaker attribution or wording is uncertain, label the utterance uncertain instead of quietly repairing it.
Step 3: Extract Atomic Evidence
Do not summarize each interview into a paragraph. Paragraph summaries blend observations, interpretations, and importance judgments too early.
Extract atomic evidence records instead:
Require one observation per record and at least one source pointer. Keep the quote short enough for review and retain the complete transcript separately. Reject records that infer intent, frequency, severity, or business impact beyond what the participant actually said.
A prompt for this stage can say:
Step 4: Cluster Across Interviews, Not Quotes
Now cluster semantically related observations across the approved interview set. Count independent interviews, not mentions. Five quotes from one participant still represent one interview.
Use a theme record like this:
Apply the minimum threshold defined before analysis. Below-threshold observations belong in an outlier section, not in a theme padded with similar-sounding model text. Preserve a contradiction when one participant reports a clear confirmation. The contradiction may reveal a segment, permission, platform, or timing difference that the average would hide.
Never ask the model to assign a roadmap priority from interview frequency alone. Frequency within a small qualitative study is useful for organizing review, but it is not market prevalence, revenue impact, or causal evidence.
Step 5: Write the Evidence Ledger to Lark
Use a Lark Base when the team needs filters, review states, and linked records. A practical design uses three tables:
Table | One row represents | Required fields |
|---|---|---|
Interviews | One source | Interview ID, segment, date, consent scope, transcript status, retention class |
Evidence | One observation | Evidence ID, interview ID, quote, timestamp, confidence, review status |
Themes | One cluster | Theme ID, count, evidence links, contradictions, interpretation, decision status |
Create the Base and its fields manually first. Then let the agent add reviewed records to known table IDs. Preflight a representative write and inspect the destination, schema, acknowledgement, and quote before executing a batch.
Use a Lark document for the short decision brief. It should link back to the Base rather than copying the whole evidence set. The brief should contain the study boundary, source coverage, accepted themes, contradictions, open questions, and recommended next research action.
The Aident Loadout first-task guide shows the broader discovery and approval loop if this is your first managed Action workflow.
Step 6: Add Human Review Gates
Use at least three review states:
unreviewed: generated but not safe for a decision.evidence_checked: quotes and pointers match the transcript.decision_approved: a named owner accepts the interpretation and next action.
The reviewer should be able to reject or split a theme without rewriting raw evidence. Record corrections so the next run can reveal where the workflow repeatedly overgroups, drops context, or misreads speaker intent.
Require another review whenever a theme contains sensitive data, contradicting evidence, one dominant participant, uncertain transcription, or a recommendation that changes a product commitment. Do not let the agent contact participants, create roadmap tickets, or publish research findings as an implied extension of synthesis.
Step 7: Measure the Workflow
Measure whether the agent makes research more inspectable, not whether it creates more themes.
Track:
provenance completeness: accepted observations with a valid interview, utterance, and timestamp pointer;
quote accuracy: sampled quotes that match the recording;
independent support: accepted themes meeting the preset interview threshold;
contradiction retention: known counterevidence preserved in the theme record;
correction rate: generated observations or clusters changed during review;
time to approved brief; and
credits per approved interview set.
Set a fail-closed target for provenance completeness. A theme with a missing source pointer is not almost finished. It is unreviewable.
A Reusable Agent Prompt
The Boundary That Keeps This Useful
Customer interview synthesis is an evidence-management task before it is a generation task. The agent can transcribe, normalize, link, cluster, and format. A researcher still decides whether the evidence is credible, what it means, and what the team should do.
Refresh this workflow when transcription fields, Action pricing, Lark schemas, privacy requirements, or evidence-grounding research changes. Recheck the current Actions and quotes on every run.
Sources
PersonaCite: VoC-Grounded Interviewable Agentic Synthetic AI Personas
How to manage AI and human input in customer feedback systems
Ready to build one reviewable brief from real customer evidence? Give your agent the exact instruction Follow https://aident.ai/SETUP.md, then use the tagged setup guide to connect the current transcription and Lark Actions.



The one tool
for every tool
your agent needs.
Give any AI agent real capabilities in seconds. Connect 1,000+ tools once, skip the setup headache, and let your agents execute.
