Skip to content
NakodaAI

AI FIELD MANUALFOUNDERS

Twenty customer interviews, one weekend, and a founder who needs the real pattern - not the one they hoped for

A structured, AI-assisted way to turn a stack of raw customer discovery transcripts into honest themes, using a real qualitative-research method rather than asking a chatbot to "summarize the feedback."

Last reviewed 1 September 2026

THE PROBLEM

A two-person founding team building an SME accounting tool has just finished twenty customer discovery calls. They have twenty transcripts, a Friday deadline to decide whether to build the reconciliation feature or the invoicing feature next, and no time to read twenty hours of audio notes carefully.

The tempting shortcut - paste all twenty transcripts into ChatGPT and ask "what's the biggest pain point" - reliably produces an answer, and the answer is reliably biased toward whichever few calls were most articulate, most recent in the context window, or most aligned with what the founders already wanted to hear. It is confirmation bias with a professional-sounding paragraph wrapped around it.

The real risk is not a bad AI summary. It is a founder making a build decision on a summary that quietly discarded twelve of the twenty voices.

THE APPROACH

Use open coding, then axial coding - the two-pass method qualitative researchers use to find patterns without prematurely deciding what the pattern is. Open coding tags what each individual transcript actually says, sentence by sentence, before any comparison across transcripts happens. Axial coding then groups those tags into themes only after every transcript has been coded the same way.

The LLM's job is the first pass - fast, literal, per-transcript tagging - never the second pass. Deciding what the themes are is a judgment call that needs a human looking at the full tag frequency table, not a model compressing twenty conversations into one paragraph.

WHY IT WORKS

Coding each transcript independently, before any cross-transcript comparison, is exactly what prevents the recency and vividness bias a single "summarize all of this" prompt introduces - the model never gets a chance to let call #3 overshadow call #17 because it never sees them side by side until the tags are already fixed.

A frequency table over open codes is auditable in a way a narrative summary is not: if 14 of 20 transcripts independently produced a tag close to "can't reconcile bank feed with invoices", that is a real, countable pattern - not the model's impression of one.

STEP BY STEP

  1. Transcribe every call the same way

    Get all twenty calls into clean text with speaker labels. Auto-transcription (Otter.ai or Fireflies.ai) is fine for this - the coding pass downstream is forgiving of minor transcription errors, it is not forgiving of missing calls.

  2. Code each transcript individually

    For each transcript separately (one prompt per transcript, not a batch), ask the model: "Read this transcript. List every distinct pain point, need or frustration the customer describes, in their own words where possible, as short tags. Do not compare to other customers - you have not seen any other transcript." This literally-true instruction is what keeps each transcript's coding independent.

  3. Build the frequency table yourself

    Paste all twenty tag lists into a spreadsheet. Manually cluster near-duplicate tags ("can't match bank feed" and "reconciliation is manual" are the same underlying issue) - this clustering step is the one place human judgment cannot be skipped, because it is where the real theme gets named.

  4. Count, don't summarize

    The output of this step is a ranked list: theme, number of transcripts that raised it, out of 20. Not a paragraph - a countable table you can point at.

  5. Pull verbatim quotes for the top themes

    For the top three themes by count, go back to the transcripts and pull two or three direct quotes each. A decision backed by a customer's actual words in their voice is far harder to talk yourself out of than a decision backed by an AI paraphrase.

  6. Decide, and write down what would change your mind

    Make the build call against the counted table, and write one sentence on what evidence in the next ten calls would flip the decision - this is what turns a one-off synthesis into a repeatable discovery process.

TOOLS

LIMITATIONS

  • This method takes real founder time - twenty separate coding passes plus manual clustering is a few hours, not five minutes. It is deliberately not the fast option; it is the option that does not quietly lie to you.

  • The model's per-transcript tagging is only as good as the transcript. A noisy call, a customer who trails off, or a translated conversation will produce weaker tags, and those transcripts should be flagged and read manually rather than trusted at face value.

  • Twenty interviews is a discovery signal, not a statistically significant sample. A theme raised by 14 of 20 customers is worth acting on; a theme raised by 3 of 20 is a hypothesis for the next round of calls, not a build decision.

EXAMPLE

An SME accounting startup needs to decide between building bank-reconciliation or invoicing next, from twenty discovery calls.

  1. All 20 calls transcribed via Otter.ai, cleaned of filler and exported as text.

  2. Each transcript coded independently in its own Claude conversation - 20 separate tag lists produced.

  3. Manual clustering in a spreadsheet: "reconciliation" tags appeared, in some phrasing, in 15 of 20 transcripts; "invoicing" tags appeared, in some phrasing, in 6 of 20.

  4. Three verbatim quotes pulled for the reconciliation theme, directly from three different customers' own words.

  5. Decision: build reconciliation first, with a written note that if the next five customer calls raise invoicing more than reconciliation, revisit.

A build decision backed by a countable 15-of-20 pattern and real customer quotes, rather than a single AI-generated paragraph that could have overweighted the two most talkative calls.

RELATED

Nakoda editorial · last reviewed

This entry describes a workflow Nakoda recommends - it is not a claim about how any named tool behaves in every case, and it is not paid placement. Spotted something out of date? Tell us.

MORE FOR THIS AUDIENCE