AI Thematic Analysis of Interview Transcripts: A Practical Workflow
AI can help organize interview transcripts, retrieve relevant passages, and propose candidate patterns. It cannot decide what a theme means for your research question, whether the dataset supports it, or which interpretation is ethically and methodologically defensible.
That division of labor is the foundation of a credible AI-assisted thematic analysis:
- the researcher defines the question, analytical orientation, and evidence standard;
- AI helps navigate and structure the material;
- the researcher reads, challenges, revises, and reports the analysis;
- every important claim remains traceable to transcript passages and, where necessary, the source audio.
This workflow is compatible with reflexive thematic analysis, but it does not turn that interpretive method into an automatic coding pipeline.
First: decide what kind of thematic analysis you are doing
“Find the themes” sounds like a neutral instruction. It is not. Themes depend on the research question, the researcher’s theoretical assumptions, what counts as relevant evidence, and the level at which meaning is interpreted.
Before using AI, write a short analysis brief:
| Decision | Example |
|---|---|
| Research question | How do first-time managers describe the transition into people management? |
| Analytical approach | Reflexive thematic analysis |
| Orientation | Primarily inductive, informed by role-transition research |
| Level | Semantic first, with later attention to underlying assumptions |
| Unit of analysis | The complete interview, with coded excerpts |
| Inclusion rule | Experiences directly related to becoming responsible for others |
| Exclusion rule | General company complaints without a link to the transition |
| Evidence standard | Several rich examples or one analytically important case, clearly qualified |
This brief does not freeze the analysis. It gives you a starting position that can be examined and revised.
Braun and Clarke describe reflexive thematic analysis as a recursive process rather than a mechanical sequence. Their guidance emphasizes the researcher’s active role and the need for themes to have a central organizing concept. See their overview of doing reflexive thematic analysis.
What AI is useful for—and what remains yours
Used carefully, AI can support:
- transcript familiarization and navigation;
- retrieval of passages related to a question;
- initial code suggestions;
- comparison of wording across interviews;
- clustering codes into candidate themes;
- finding apparent contradictions or negative cases;
- drafting evidence tables and theme summaries;
- locating candidate quotations for verification.
The researcher remains responsible for:
- choosing the research question and analytical lens;
- understanding the interview context;
- deciding what is meaningful rather than merely frequent;
- distinguishing a topic from a theme;
- interpreting silence, hesitation, contradiction, and interaction;
- testing whether evidence supports the proposed story;
- protecting participants and honoring consent;
- making and defending the final claims.
AI output is a proposal about the text it was given. It is not an independent finding, a reliability certificate, or a substitute for reading the dataset.
Prepare a trustworthy transcript set
Thematic analysis inherits the weaknesses of its source material. A wrong speaker label can turn an interviewer’s prompt into participant evidence. A missing negative can reverse a claim. An invented proper noun can create a false code.
Before analysis:
- keep each original recording;
- confirm each file’s interview ID and consent scope;
- transcribe with stable speaker labels;
- correct names, specialist terms, numbers, and consequential wording;
- mark unresolved audio rather than guessing;
- apply one documented cleanup convention across the dataset;
- retain timestamps or segment references.
The speaker-label and quote-verification guide covers this review in detail. You can create the first draft with Sondeas audio to text and keep the resulting documents together in the research and interviews workspace.
Create a dataset ledger
Give every interview a stable identifier and track the context needed for interpretation:
| Interview ID | Participant context | Date | Consent restrictions | Transcript status | Analysis status |
|---|---|---|---|---|---|
| INT-001 | New manager, retail | 12 Aug | Anonymous quotes only | Reviewed | Coded |
| INT-002 | New manager, software | 14 Aug | No direct quotes | Reviewed | Familiarized |
Avoid putting unnecessary identifying data into prompts or analysis documents. Use participant IDs when identity is not analytically necessary.
Phase 1: become familiar with the interviews
Read the transcripts yourself. Listen to selected passages where tone, ambiguity, interruption, or emotion matters. Write a short memo after each interview:
- What problem is this participant trying to explain?
- What assumptions appear to organize their account?
- Where do they contradict or qualify themselves?
- What surprised you?
- What contextual details affect interpretation?
- What questions should be tested against other interviews?
Then use AI as a navigation aid. Useful requests include:
Summarize the participant's account of becoming responsible for former peers.
For each point, include the speaker label and timestamp.
Do not infer motives that are not expressed in the transcript.
Mark uncertain or conflicting passages.
Compare the result with your memo. Differences are informative. The model may surface a passage you missed, flatten an important contradiction, or prioritize repeated language over analytically rich detail.
Familiarization is not disposable preprocessing
If you skip close reading and begin from an AI summary, the summary becomes an invisible filter on the analysis. You will see the model’s selection before you see the participant’s account. Read the transcript first, even when the summary is useful later.
Phase 2: generate candidate codes
A code is a concise label for something relevant in a passage. It may describe explicit content—avoiding difficult feedback—or capture a more interpretive idea—performing certainty while feeling unprepared.
Ask for suggestions within a defined scope:
Research question: How do first-time managers describe the transition into people management?
Suggest candidate codes for this transcript.
For each code, provide:
- a short definition;
- the exact supporting excerpt;
- speaker and timestamp;
- whether the code is semantic or interpretive;
- one plausible alternative reading.
Do not treat the interviewer's words as participant evidence.
Do not claim prevalence from one transcript.
Review each suggestion. Keep, rename, split, merge, or reject it. Add codes the model missed. The goal is not agreement with the tool; it is a transparent engagement with the data.
Maintain a working codebook or code list with:
- code name;
- current definition;
- inclusion and exclusion notes;
- example excerpts;
- changes and reasons;
- related or competing codes.
In reflexive thematic analysis, this record supports reflection and consistency. It should not be mistaken for proof that coding is objective or fixed.
Phase 3: compare cases without counting too early
After reviewing individual interviews, compare how a code operates across the dataset.
Ask questions such as:
- Does the same phrase describe different experiences?
- Which participant contexts alter the pattern?
- Where is the proposed pattern absent?
- Which cases resist the emerging explanation?
- Is a rare case analytically important?
- Are interviewer questions producing the apparent similarity?
AI can retrieve passages for a code or compare a bounded set of excerpts. It may reduce navigation and organization work, but the output must be checked against the complete interviews.
Do not turn mention counts into qualitative importance. Ten brief references are not automatically stronger evidence than one detailed, contradictory account. If you report frequency, define what was counted, why it matters, and what the count cannot show.
Phase 4: build candidate themes
A topic groups material about the same subject. A theme makes an interpretive claim about a patterned meaning.
For example:
- Topic: feedback conversations
- Candidate theme: authority becomes real when friendship no longer protects the manager from difficult feedback
The theme has a central idea. It explains how the coded material relates to the research question.
Ask AI to propose clusters only after you have reviewed the codes:
Group these reviewed codes into candidate themes.
For each candidate theme, provide:
- a central organizing concept;
- the codes it includes;
- supporting and conflicting excerpts;
- boundary conditions;
- overlap with other candidate themes;
- a weaker alternative interpretation.
Do not label a broad interview topic as a theme without an interpretive claim.
Treat the response as one possible map. Draw another map yourself. Compare the two and note what each arrangement reveals or hides.
Phase 5: test themes against the evidence
This is the main defense against a neat but unsupported analysis.
For every candidate theme, ask:
- Do the included excerpts form a coherent pattern?
- Is the central concept distinct from the other themes?
- Does it answer the research question?
- What evidence does not fit?
- Is the theme based on participant speech or on interviewer framing?
- Does it depend too heavily on one vivid quote?
- Does it erase relevant differences between participants?
- Could a simpler interpretation explain the same material?
Build an evidence matrix:
| Theme | Supporting cases | Contradicting cases | Strongest excerpts | Boundary | Decision |
|---|---|---|---|---|---|
| Managing former peers creates an authority dilemma | INT-001, 003, 006 | INT-004 | 001:22:14; 006:31:05 | Mainly internal promotions | Retain and narrow |
Then return to the full transcript around each excerpt. A passage can look decisive when detached from the question that prompted it or the qualification that followed.
Reject weak themes. Merge overlapping ones. Split themes that contain more than one organizing concept. Rename themes so the name communicates the analytical claim rather than a generic subject.
Phase 6: define the analytical story
Write a short definition for each retained theme:
- What does the theme claim?
- What is its central organizing concept?
- What does it include and exclude?
- How does it relate to the research question?
- How does it differ from the other themes?
- Under what conditions does it appear?
- What tension or contradiction does it preserve?
Next, explain the relationship between themes. A final analysis is not a list of buckets. It is an argument about the dataset.
For example, an analysis of first-time managers might show a progression from borrowed authority, through conflict avoidance, to a more personal management identity. Another dataset may resist a progression entirely and instead reveal competing strategies. The structure must come from the evidence and the chosen interpretation.
Phase 7: select and verify quotations
Choose quotes because they do analytical work, not because they sound dramatic.
For each candidate:
- retrieve the exact transcript passage;
- open the original audio at the timestamp;
- verify wording and speaker;
- listen to the surrounding exchange;
- confirm that the quote supports the stated theme;
- check anonymization and consent;
- record any editorial change.
Sondeas can generate a Quotes document alongside Themes, Summary, and Transcript documents. These are connected working artifacts, not independent evidence. A generated quote still requires audio verification before publication.
Keep an audit trail that matches your method
Your record should make the analytical process understandable without pretending it was free of judgment.
Keep:
- the analysis brief and its revisions;
- original recordings and reviewed transcripts;
- dataset ledger;
- familiarization memos;
- AI prompts and relevant outputs;
- code and theme revisions;
- evidence matrices;
- rejected interpretations and reasons;
- quote-verification status;
- a reflexive note on the researcher’s role and assumptions.
If your institution or client restricts external AI processing, follow that policy. Consent to be interviewed does not automatically equal consent to send the recording or transcript to every possible processor.
How to report AI-assisted thematic analysis
State enough for a reader to understand what happened:
- which transcription and AI functions were used;
- what data was provided to them;
- what the AI was asked to produce;
- which outputs were reviewed, changed, or rejected;
- how themes were developed and tested;
- how quotations were verified;
- how privacy, consent, and retention were handled;
- what limitations remain.
Avoid vague statements such as “AI analyzed the interviews.” Describe its bounded role: for example, “The model proposed initial code labels and retrieved candidate excerpts; the research team revised the codes, developed themes, tested negative cases, and verified all reported quotations against the recordings.”
Common mistakes
Asking for final themes in one prompt
This hides familiarization, coding decisions, rejected alternatives, and evidence review. Use staged, reviewable outputs.
Losing the link to the interview
Every consequential code, theme, and quote should retain an interview ID and timestamp or segment reference.
Feeding the model an unreviewed transcript
Speaker errors and missing negatives propagate into the analysis. Correct high-consequence passages before asking the model to interpret them.
A defensible starting point
Begin with a reviewed, speaker-labeled transcript. Create a Summary, Themes, and Quotes document as separate working views. Then test every proposed pattern against the full interviews and return to the audio for every quotation used in a deliverable.
Explore Sondeas for research and interviews, or start by turning the source recordings into traceable text with audio to text. If the files are lengthy, review the long-recording workflow before upload.
Frequently asked questions
Can AI perform thematic analysis by itself?
It can propose codes, clusters, and summaries. The researcher still defines the method, interprets meaning, tests the evidence, handles ethics, and owns the final claims.
Do all themes need to appear in most interviews?
No. Relevance depends on the research question and analytical argument. Qualify the scope and do not imply prevalence that the dataset or sampling cannot support.
Is an AI-generated codebook objective?
No. Code definitions reflect instructions, model behavior, and researcher decisions. Review and revise them as interpretive tools.
Can I use transcript excerpts without checking the audio?
For important claims or publication quotes, check the source recording. Transcript errors, speaker drift, and missing context can materially change meaning.
Sources
- Braun and Clarke: Doing Reflexive Thematic Analysis — recursive phases, researcher reflexivity, and theme development.
- Kiger and Varpio: Thematic analysis of qualitative data: AMEE Guide No. 131 — practical guidance on thematic analysis; see also the PubMed record.
- BMJ: Practical thematic analysis — a practical account of developing, reviewing, and reporting themes.
- Springer: A worked example of Braun and Clarke’s approach to reflexive thematic analysis — a transparent worked example of the analytical process.
- Springer: Thematic analysis of interview data with ChatGPT — a recent protocol study examining AI-assisted qualitative analysis.