Ilmu Komputer & AI editorial
Open AccessOA2026
Designing an Auditable LLM-Supported Workflow for Qualitative Thematic Analysis
A privacy-preserving, two-phase operationalization of inductive and latent Thematic Analysis with deterministic procedural control and expert-validated evaluation
Nadia Jul Jeldtoft; Tariq Yousef· 2026· DOI 10.48550/arXiv.2608.30543
The core problem
Large Language Models (LLMs) offer new possibilities for scaling qualitative analysis, yet existing applications often provide limited methodological transparency regarding how qualitative methods are translated into computational procedures. This paper addresses that gap by presenting an auditable and privacy-preserving computational operationalization of inductive and latent Thematic Analysis (TA). The authors argue that scaling qualitative analysis with LLMs requires not only technical performance but also methodological accountability: researchers must be able to trace how empirical material becomes analytical output. To this end, the paper derives five design principles from the methodological requirements of TA and the conditions introduced by LLM-based inference: (1) preserving interpretative context, (2) maintaining traceable relationships between empirical material and analytical outputs, (3) representing analytical constructs and reasoning explicitly, (4) constraining LLM inference to interpretative tasks, and (5) enabling privacy-preserving local deployment. These principles frame the subsequent proof-of-concept workflow and evaluation framework. The study is positioned
Innovation
The results show that the workflow produces code-level outputs with coverage broadly comparable to human annotations and highly rated analytical justifications. In structural comparison, the LLM-supported workflow identified codes that aligned with human annotations to a degree that the authors characterize as broadly comparable, indicating that the interpretative LLM inference can approximate human coding coverage on the evaluated Danish interview transcripts. Independent expert assessment rated the analytical justifications highly, suggesting that the explicit reasoning artifacts generated by the workflow are methodologically credible. However, the workflow generated a more compressed thematic structure characterized by fewer and broader themes than human-led TA. This compression is a notable finding: while code-level coverage was comparable, the aggregation into themes differed in granularity. The authors do not frame this as a failure but as a characteristic of the workflow that may require domain-specific tuning. The evaluation demonstrates the feasibility of auditable LLM-supported TA through a modular workflow designed to scale to larger datasets, accommodate different LLMs,
Large Language Models (LLMs) offer new possibilities for scaling qualitative analysis, yet existing applications often provide limited methodological transparency regarding how qualitative methods are translated into computational procedures. This paper addresses that gap by presenting an auditable and privacy-preserving computational operationalization of inductive and latent Thematic Analysis (TA). The authors argue that scaling qualitative analysis with LLMs requires not only technical performance but also methodological accountability: researchers must be able to trace how empirical material becomes analytical output. To this end, the paper derives five design principles from the methodological requirements of TA and the conditions introduced by LLM-based inference: (1) preserving interpretative context, (2) maintaining traceable relationships between empirical material and analytical outputs, (3) representing analytical constructs and reasoning explicitly, (4) constraining LLM inference to interpretative tasks, and (5) enabling privacy-preserving local deployment. These principles frame the subsequent proof-of-concept workflow and evaluation framework. The study is positioned as a feasibility demonstration rather than a benchmark, emphasizing modularity, auditability, and transferability across research domains.
The paper presents a proof-of-concept for a two-phase workflow that operationalizes the five design principles by combining interpretative LLM inference with deterministic procedural control. The workflow generates codes, analytical justifications, themes, and theme descriptions while preserving explicit links to the source material. Phase 1 focuses on code generation: the LLM is constrained to interpretative tasks, producing candidate codes and justifications that are anchored to specific transcript segments. Phase 2 focuses on theme construction: deterministic procedural control aggregates and organizes codes into themes and theme descriptions, ensuring that the analytical reasoning remains explicit and traceable. The architecture can be represented as follows:
Why it matters
The discussion interprets the findings through the lens of the five design principles. First, preserving interpretative context is achieved by anchoring codes and justifications to transcript segments, which supports auditability. Second, traceable relationships between empirical material and analytical outputs are maintained through explicit links, enabling researchers to inspect how themes emerged. Third, representing analytical constructs and reasoning explicitly is realized through analytical justifications and theme descriptions, which expert assessors rated highly. Fourth, constraining LLM inference to interpretative tasks and delegating procedural control to deterministic mechanisms reduces the risk of unaccountable generative drift. Fifth, privacy-preserving local deployment addresses ethical and data-protection concerns inherent in qualitative research with sensitive material. The more compressed thematic structure—fewer and broader themes—suggests a trade-off between scalability and thematic granularity. The authors propose that domain adaptation primarily requires adjustments to the prompting strategy, which lowers the barrier to transfer across research domains. The evaluation framework combining structural comparison with human-led TA and independent expert assessment offers a template for validating LLM-supported qualitative workflows. The paper concludes that auditable LLM-supported TA is feasible, but that methodological transparency must be designed in rather than added post hoc. The modular workflow is positioned as a foundation for scaling qualitative analysis while preserving the interpretative rigor that defines TA.
Who should read this
CS practitioners and researchers
Opening member content…