🌻 Causal mapping QDA - extended abstract

2026

An orthogonal yes#

The debate about generative AI in qualitative analysis is usually drawn as a line between rejecting the technology to protect non-positivist, meaning-based practice (Jowsey et al. 2025) and embracing it to work at scale Friese et al. (2026). We take a position orthogonal to that line. We share the anti-positivist instinct that research cannot be completely reduced to a decision procedure, but we also believe there is value in exploring how far decision procedures can take us Mayring (2000) and identifying where human intervention is really necessary. Unlike (Friese 2025; Morgan 2025; Nguyen-Trung & Nguyen 2026) we do not advocate for involving large language models (LLMs) as part of the reflective process. (Although popular science often refers to LLMs as "algorithms", this is misleading: fundamentally, employing an LLM is not applying an algorithm or decision procedure as its results are not transparent, traceable or deterministic in any practical way).

In our approach, interpretative freedom is reserved for the human analyst at key moments only, where it is necessary, and not otherwise. The slack is taken up not by LLMs but by a set of algorithms or rules based on a minimalist version Powell et al. (2024) of causal mapping (Axelrod 1976; Eden 1988; Powell et al. 2024; Narayanan 2005). These rules determine how to code and synthesise causal information in texts in a largely standardised way. The simple coding tasks can be carried out by humans, or just as well, and faster and more consistently, by LLMs. The application of the LLMs is governed as far as possible by transparent rules, steered only where necessary by the human analyst.

We do not claim that our method is useful on its own for reflexive, Big-Q thematic work in the sense identified by Braun and Clarke (Braun & Clarke 2023; Braun & Clarke 2025). Causal mapping is not itself a discourse- or interaction-analytic method, but it stays close to the call's preference for interpretive rigour and accountable links between analytic claims and empirical materials. Each coded unit is a causal claim a source makes in their own words, along with the verbatim quote and the identity of the source.

Causal mapping as practically focused QDA, with far fewer degrees of freedom#

In 2026 it is easy to use an LLM to auto-code codes within a text; but what codebook to use? In ordinary qualitative coding a code is connected to a theme or concept. In causal mapping the coded unit denotes a causal claim made by a source: a cause label, an effect label, the verbatim quote that supports it, and a source identifier (Axelrod 1976; Eden 1988; Powell et al. 2024; Narayanan 2005). A coding act yields an ordered pair, written Cause -> Effect, attached to evidence and provenance. The initial result of coding is a table of such links.

This is a small unit by design. We do not code strength, polarity, necessity or sufficiency, the features that systems-dynamics traditions attach to links (Kim & Andersen 2012); most respondents do not mention them, and (therefore) most analysts cannot reliably extract them from text. Basically, coding bare causal claims is relatively easy. The strength of causal mapping is that the links table can then be queried directly to answer really interesting questions: which factors are the most frequent upstream influences on an outcome, how do pathways into a target factor differ across subgroups, which links are contested. Above all, mapping out causal links is a mostly standardised way to surface the story within the text and the multiple stories shared by multiple texts.

Causal mapping provides a large set of analytical filters which can be chained together to provide answers to queries on the data. The approach has a recognisable relative in Mayring's rule-guided qualitative content analysis (Mayring 2000), differing in that its unit is an ordered pair, which is what makes pathway analysis possible. Another difference is that causal mapping works surprisingly well almost "out of the box", with relatively little adaptation for different kinds of texts. And thirdly it provides ways to answer more relevant and practical questions, because a large proportion of such questions involve causal connections and narratives: How? Why? What worked? What failed?

A short illustration. If an interviewee says that after a clinic began opening on Saturdays they no longer missed work and so could attend, a thematic pass might file codes such as Access and Opening hours — but only after substantial preparatory work to establish which codes or coding strategies could be useful for a given text. Causal coding does not have this problem because it (nearly) always codes each and every causal claim in the text, willy-nilly, using in vivo labels to code, in this case, two links, Saturday opening -> Not missing work and Not missing work -> Attending, each tied to the quote. Little freedom is left for the coder. LLMs can do this easily.

Why the coding task suits AI#

"Find the main themes in this corpus" or "summarise what these people say" are very open-ended instructions. If we give these to an LLM we expose ourselves massively to the model's implicit theory of what counts as a theme. The result may read well and sound persuasive, but it is hard to audit: the model may have downweighted minority views, smoothed over disagreement, or drifted off the text, and its fluency tends to discourage scrutiny.

The minimalist causal coding act is the opposite kind of task. The instruction is to identify each passage where the text says one thing influenced another, and to record the cause, the effect and the exact supporting quote. It refers to features and uses words already present in the text, produces a unit that is easy to verify by reading the quote, and asks for no weighing, summarising or theorising. Errors are local: a wrong link can be dropped or corrected without unravelling the rest of the analysis. In practice the extraction runs on short chunks, often a passage at a time, and corpus-level structure is recovered by aggregating the links and clustering similar labels into groups. The model never holds the whole corpus in attention. Thus this method is most suited to surfacing explicit meaning and least suited to "reading between the lines" or making connections between far-flung passages.

The division of labour is strict. The model is a clerk: fast, consistent, willing to apply one stable rule across thousands of passages (Powell et al. Forthcoming). The human is the architect, who frames the research question, where necessary tweaks the label clusters, decides which filters to apply to answer which questions and writes the interpretation. The human can focus on key interpretive tasks not because the LLM is doing most of the work but because the causal mapping coding rules and analysis algorithms are doing most of the work. In our approach, the LLM does nothing but basic coding. The model never produces an analytic claim, interpretation or opinion.

Data and methodology#

We illustrate the argument with a corpus of 48 interviews on the experience of loneliness among young adults aged 18 to 24, recruited from four deprived London boroughs in 2019 and available as a de-identified open dataset (n.d.), analysed in three contrasting ways. A causal-mapping pass produced around 3,392 quote-grounded causal claims in roughly twenty minutes. The cause and effect labels were then grouped into clusters and the analyst then queried for the most-cited forms and aspects of loneliness, their causes and effects, the contested links, and the pathways that differ across subgroups. Interpretation begins with describing the common stories told by the respondents as they have been surfaced through this causal lens. These summary narratives can be compared with models from the literature and the analyst can begin to compare, contrast, reflect, recode and reanalyse iteratively, with more of the freedom familiar to qualitative researchers.

Accountability and main claims#

A causal map made this way carries a complete audit trail. Pick any link and you see the underlying claims, each with its source and verbatim quote, and the deterministic transforms applied between the raw coding and the rendered view. This is the "accountable links between analytic claims and empirical materials" the call asks for: every analytic claim reduces to a sequence of explicit operations on a links table whose rows are quote-grounded extractions, and a reader who doubts the analysis can rerun the pipeline.

We conclude that the choice between rejecting and embracing AI is a false one, at least for questions about what people say causes what; the approach we describe is more traceable than a conversation because it is not a conversation. To the methodological-incongruence argument that coding is a "skeuomorphic" crutch best replaced by holistic querying (n.d.), we answer quite the opposite: LLMs should be given very little freedom, and certainly not holistic freedom. Our links table is a public artefact rather than a hidden memory aid for a stateless model. Allowing the model to hold the whole analysis in working memory would amount to reinstating the black box.

AI use: The AI is a commercial large language model used at the extraction step only. The versions are reported. Its role is to propose candidate quote-backed links from short chunks and nothing else: it does not summarise across documents, choose the codebook, or write the report.

Limits#

Causal coding ignores much that matters in talk: identity work, norms, emotion, metaphor, the turn-by-turn structure of interaction. A coded link records what a source claimed rather than what is true in the world; frequency counts how widely a claim is made rather than the size of any underlying effect. Because the model sees one chunk at a time, cross-document context is lost. The full paper develops the argument and the worked comparison in detail, with practical guidance on the workflow and a fuller account of these limits.

References

Axelrod (1976). The Analysis of Cognitive Maps. In Structure of Decision : The Cognitive Maps of Political Elites.

Braun, & Clarke (2023). Toward Good Practice in Thematic Analysis: Avoiding Common Problems and Be(Com)Ing a Knowing Researcher. Taylor \& Francis. https://doi.org/10.1080/26895269.2022.2129597.

Braun, & Clarke (2025). Reporting Guidelines for Qualitative Research: A Values-Based Approach. Routledge. https://doi.org/10.1080/14780887.2024.2382244.

Eden (1988). Cognitive Mapping. https://doi.org/10.1016/0377-2217(88)90002-1.

Friese (2025). Conversational Analysis with AI - CA to the Power of AI: Rethinking Coding in Qualitative Analysis. https://doi.org/10.2139/ssrn.5232579.

Friese, Nguyen-Trung, Powell, & Morgan (2026). Beyond Binary Positions: Making Space for Critical and Reflexive GenAI Integration in Qualitative Research. https://doi.org/10.2139/ssrn.5962174.

Jowsey, Braun, Clarke, Lupton, & Fine (2025). We Reject the Use of Generative Artificial Intelligence for Reflexive Qualitative Research. https://doi.org/10.2139/ssrn.5676462.

Kim, & Andersen (2012). Building Confidence in Causal Maps Generated from Purposive Text Data: Mapping Transcripts of the Federal Reserve. https://doi.org/10.1002/sdr.1480.

Mayring (2000). Qualitative Content Analysis. Forum Qualitative Sozialforschung / Forum: Qualitative Social Research. https://doi.org/10.17169/FQS-1.2.1089.

Morgan (2025). Query-Based Analysis: A Strategy for Analyzing Qualitative Data Using ChatGPT. https://doi.org/10.1177/10497323251321712.

Narayanan (2005). Causal Mapping: An Historical Overview. In Causal Mapping for Research in Information Technology. https://www.google.co.uk/books/edition/_/61z36j6QgmAC?hl=en&gbpv=1.

Nguyen-Trung, & Nguyen (2026). Narrative-Integrated Thematic Analysis (NITA): How Can LLMs Support Theme Generation without Coding?. Routledge. https://doi.org/10.1080/14780887.2026.2638348.

Powell, Copestake, & Remnant (2024). Causal Mapping for Evaluators. https://doi.org/10.1177/13563890231196601.

Powell, Cabral, & Remnant (Forthcoming). AI-assisted Causal Mapping: A Validation Study. https://drive.google.com/open?id=1XFAJafYs5rS9HDf8en_ukknnorbhFX2N&usp=drive_fs.