Coding Qualitative Data: A Beginner's Guide
Qualitative data coding is the systematic process of labeling and organizing text, audio, or visual material to identify patterns, themes, and relationships that answer a research question. For students, researchers, and life-science professionals new to qualitative work, coding transforms unstructured information such as interview transcripts, focus group discussions, or open-ended survey responses into a structured dataset that can be analyzed, compared, and reported. This guide explains the core coding approaches, provides a practical workflow, and includes a sample coded transcript excerpt to demonstrate how coding works in practice.
What Qualitative Coding Means in Research Practice
Qualitative coding is the analytical act of assigning short labels, called codes, to segments of data. A code captures the essence of what a participant said or what an observation shows. Codes can describe topics, actions, emotions, relationships, or concepts. The coded data then serve as the foundation for identifying larger patterns and building explanations.
In-depth interviews are a common method of qualitative data collection because they provide rich data on individuals' perceptions and behaviors that would be challenging to collect with quantitative methods [6]. Coding is the bridge between raw interview material and meaningful findings. Without coding, a researcher has a collection of transcripts but no systematic way to compare what different participants said or to identify recurring themes across the dataset.
The coding process is not mechanical. It requires interpretive judgment about what matters in the data and how different segments relate to each other. A review of the Qualitative Analysis Guide of Leuven highlights three key strategies for analyzing complex narrative data: the case-oriented approach, the method of constant comparison, and the use of data-generated codes [9]. These strategies emphasize that good coding involves moving back and forth between individual cases and emerging patterns, always grounding interpretations in what the data actually contain.
Core Principles of Qualitative Coding
Several principles guide effective coding practice regardless of the specific method chosen.
Codes Should Be Grounded in the Data
Codes can come from two directions. Inductive coding, also called open coding, develops codes directly from the data without a preexisting framework. Deductive coding applies a preexisting framework or codebook to the data. Many studies combine both approaches. One health services research approach applies the principles of inductive reasoning while also employing predetermined code types to guide analysis and interpretation [11]. These code types include conceptual codes, relationship codes, perspective codes, participant characteristic codes, and setting codes.
Coding Is Iterative
Coding is not a single pass through the data. Researchers typically move through multiple cycles. First-cycle coding breaks the data into segments and assigns initial labels. Second-cycle coding reorganizes and synthesizes those initial codes into broader categories or themes. A review of Saldaña's coding manual describes first- and second-cycle methods and notes that later chapters address grounded-theory pathways such as focused, axial, and theoretical coding, as well as cumulative approaches including pattern, elaborative, and longitudinal coding [20].
Memos Capture Analytical Thinking
Analytic memos are written reflections about what the researcher is noticing in the data. Memos document decisions about code definitions, emerging patterns, and questions that arise during analysis. The Qualitative Analysis Guide of Leuven emphasizes the importance of understanding underlying principles and implementing them carefully to conduct methodologically sound analyses [9]. Memos are the record of that careful thinking.
Team Coding Requires Clear Procedures
When multiple researchers code the same data, procedures must be explicit. In one study of human mobility data during the COVID-19 pandemic, transcripts were open-coded to create a codebook that was then applied by two team members who blind-coded all transcripts, with consensus coding used for coding discrepancies [10]. This approach demonstrates how teams can manage reliability while still allowing for individual interpretation.
Types of Coding in Qualitative Research
Three coding approaches are foundational for beginners: open coding, axial coding, and selective coding. These terms come from grounded theory but are widely used across qualitative traditions.
Open Coding
Open coding is the initial stage where the researcher reads through the data and assigns preliminary labels to segments. The goal is to stay close to the data and capture as many potential ideas as possible. Open coding is inductive because the codes emerge from the content instead of from a predetermined list.
In a grounded theory study of university student well-being in Indonesia, researchers analyzed interview data iteratively using open, axial, and selective coding, constant comparison, analytic memoing, and team-based coding consensus [15]. Open coding allowed the researchers to identify a wide range of student experiences before organizing them into broader categories.
Axial Coding
Axial coding connects categories to subcategories and explores relationships between them. The researcher asks questions about conditions, contexts, actions, and consequences. Axial coding builds structure into the coding system by showing how categories relate to each other.
In the Indonesian student well-being study, axial coding helped the researchers see that academic well-being, social well-being, psychological well-being, and physical and environmental well-being were not independent categories [15]. The relationships between these dimensions became the focus of analysis.
Selective Coding
Selective coding identifies a core category and systematically relates all other categories to it. This stage integrates the analysis into a coherent explanatory framework. The core category represents the central phenomenon that the research explains.
The Indonesian study used selective coding to identify cultural-contextual alignment as the core process explaining how students experience well-being when academic demands, relational support, psychological regulation, and resource conditions fit with local values, family expectations, peer norms, and future aspirations [15]. Selective coding produced the study's central theoretical contribution.
At a Glance: Coding Approaches Compared
| Coding Approach | Primary Purpose | Typical Input | Output | Best Used When |
|---|---|---|---|---|
| Open Coding | Break data into segments and assign initial labels | Raw transcripts, field notes, open-ended responses | A list of preliminary codes grounded in the data | Exploring a topic with little existing theory or when you want to stay close to participant language |
| Axial Coding | Connect categories and subcategories to show relationships | The code list produced by open coding | An organized structure showing conditions, contexts, actions, and consequences | You have many codes and need to organize them into a coherent framework |
| Selective Coding | Identify a core category and integrate all categories around it | The axial coding structure | A central explanatory concept or theory | You are building a grounded theory or need a single integrating theme |
Practical Workflow for Coding Qualitative Data
A structured workflow helps beginners move through the coding process without becoming overwhelmed. The following steps apply to most qualitative projects.
Step 1: Prepare Your Data
Transcribe interviews or focus group recordings. Clean the transcripts by removing identifying information. Decide whether you will code in the original language or translate first. One study of mobile health applications conducted focus groups in Bangla, transcribed them in Bangla, and then translated them into English for thematic analysis [8]. Translation decisions affect coding because meaning can shift between languages.
Step 2: Read Through the Data
Read all transcripts at least once before coding. This gives you a sense of the whole dataset. The case-oriented approach described in the Qualitative Analysis Guide of Leuven emphasizes understanding each case as a whole before breaking it into coded segments [9]. Reading without coding also helps you notice patterns that will inform your code list.
Step 3: Develop Your Initial Code List
For inductive coding, create codes as you read. For deductive coding, prepare a codebook in advance based on your theoretical framework or research questions. The Consolidated Framework for Implementation Research is one example of a framework that can guide deductive data collection and analysis [7]. A common framework used from the outset helps when integrating data across multiple sites or studies [12].
Step 4: Code the Data
Apply codes to segments of text. A segment can be a sentence, a paragraph, or a longer passage. Assign as many codes as needed to capture the content. Use qualitative data analysis software if available, but coding can also be done with word processing documents or spreadsheets.
Step 5: Review and Refine Codes
After the first pass, review your code list. Merge codes that overlap. Split codes that are too broad. Write definitions for each code so that you and your team apply them consistently. The Reproducible Open Coding Kit is a standard for human-readable and machine-readable qualitative coded data designed to be accessible to researchers without dedicated software while offering enough structure to enable programmatic processing [13].
Step 6: Move to Axial and Selective Coding
Organize your refined codes into categories. Look for relationships between categories. Identify the central theme or process that ties everything together. This stage moves you from description to explanation.
Step 7: Write Your Findings
Translate your coding structure into a narrative. Use quotes from participants to illustrate each theme. Explain how the coding process supports your conclusions.
Sample Coded Transcript Excerpt
The following excerpt demonstrates how open, axial, and selective coding might be applied to a short interview segment. The research context is a study of how primary care clinicians decide whether to adopt a new mobile health application for managing patients with acute diarrhea.
Raw Transcript Segment
"At first I was worried the app would slow me down during consultations. I have maybe ten minutes per patient. But after using it for two weeks, I found that the decision support actually saved time because I did not have to look up dosing guidelines. The main problem is that the hospital Wi-Fi is unreliable in the examination rooms, so sometimes I cannot open the app when I need it."
Open Coding Applied
| Text Segment | Open Code |
|---|---|
| "I was worried the app would slow me down" | Concern about time efficiency |
| "I have maybe ten minutes per patient" | Time pressure in consultations |
| "The decision support actually saved time" | Perceived efficiency gain |
| "I did not have to look up dosing guidelines" | Reduced information search burden |
| "The hospital Wi-Fi is unreliable" | Technical infrastructure barrier |
| "Sometimes I cannot open the app when I need it" | Interrupted access to tool |
Axial Coding Applied
The open codes group into two categories. The first category is clinician workflow impact, which includes time pressure, perceived efficiency gain, and reduced information search burden. The second category is technical environment constraints, which includes technical infrastructure barrier and interrupted access to tool. The relationship between these categories is that technical constraints can undermine the workflow benefits of the tool.
Selective Coding Applied
The core category is conditional adoption. Clinicians adopt the tool when it fits their workflow and when the technical environment supports reliable access. When either condition fails, adoption is partial or abandoned. This core category explains the pattern across the dataset.
Options and Tradeoffs in Coding Methods
Different coding approaches serve different research purposes. Understanding the tradeoffs helps you choose the right method for your project.
Inductive Versus Deductive Coding
Inductive coding is appropriate when you are exploring a topic with limited existing theory. It produces codes that are closely tied to participant language and experiences. Deductive coding is appropriate when you have a theoretical framework or when you need to compare findings across studies that used the same framework.
A rapid deductive analysis approach using the Consolidated Framework for Implementation Research used notes and audio recordings instead of full transcripts [7]. This approach required fewer analyst hours and eliminated transcription costs compared to a traditional deductive approach that involved independent in-depth manual coding of interview transcripts using qualitative software [7]. The tradeoff is that rapid approaches may capture less detail than full transcript coding.
Framework Matrix Analysis Versus Applied Thematic Analysis
Framework Matrix Analysis and Applied Thematic Analysis are two methods that have not been as widely used or cited compared to content analysis or grounded theory [8]. In a study of mobile health applications, the same qualitative data were analyzed separately using each method. Framework Matrix Analysis was used for data reduction where specific outcomes were needed to make programming and design decisions, while Applied Thematic Analysis captured more nuanced issues that guide use, product implementation, and training [8].
The choice between these methods depends on your research goal. If you need structured outputs for design decisions, a framework matrix may be more efficient. If you need deep understanding of participant experiences, thematic analysis may be more appropriate.
Manual Coding Versus Software-Assisted Coding
Coding can be done manually with printed transcripts and highlighters, in word processing documents, or with dedicated qualitative data analysis software. Manual coding is accessible and keeps you close to the data. Software helps manage large datasets, supports team coding, and makes it easier to retrieve coded segments.
The Reproducible Open Coding Kit offers a middle path. It is designed to be accessible to researchers without dedicated software while simultaneously offering enough structure to enable programmatic processing [13]. This standard supports open science principles in qualitative research and moves toward machine-readability and interoperability of coded qualitative data [13].
Observations and Measurements in Coding Practice
Coding quality can be assessed through several observable measures. These measures help you document the rigor of your analysis.
Code Saturation
Code saturation is the point at which additional data no longer produce new codes. Researchers typically need to decide on sample size before data collection begins, and there is no agreement on the minimum number of interviews needed to achieve saturation [6]. In a study of web-based interviews, true saturation was reached after 91% to 100% of planned interviews, while near saturation was reached after 33% to 60% of planned interviews [6]. Studies that relied heavily on deductive coding and studies that had a more structured interview guide reached both true saturation and near saturation sooner [6].
For beginners, this means that saturation is not a fixed number. It depends on the homogeneity of your sample, the structure of your data collection, and whether you are using inductive or deductive coding.
Inter-Coder Reliability
When multiple coders work on the same data, inter-coder reliability measures the extent to which they assign the same codes to the same segments. A study comparing human-derived and AI-derived codebooks in surgical qualitative research found that the human-derived codebook produced an overall agreement rate of 95.1% and a Cohen kappa of 0.754, while the AI-derived codebook produced an agreement rate of 92.3% and a kappa of 0.638 [16]. Both codebooks identified 13 codes, five of which were identical [16].
These numbers illustrate that even experienced coders do not agree perfectly. Disagreements are opportunities to refine code definitions and clarify the coding framework.
Time and Cost Tracking
Qualitative analysis is resource intensive. Tracking time and costs helps you plan future projects and document the value of your approach. In a comparison of rapid and traditional deductive analysis, the rapid approach required 409.5 analyst hours compared to 683 hours for the traditional approach, and the rapid approach eliminated $7,250 in transcription costs [7]. These measurements are useful for planning and for justifying methodological choices.
Records and Documentation for Coding Projects
Good record keeping supports the credibility of your analysis and enables others to understand your decisions.
Codebook
A codebook lists each code, its definition, and examples of text segments that received that code. The codebook is the reference document for your coding team. It should be updated as codes are refined.
Coding Log
A coding log records when coding sessions occurred, which transcripts were coded, and any decisions made during those sessions. This log documents the progression of your analysis.
Analytic Memos
Analytic memos capture your thinking about patterns, relationships, and questions. Memos are written throughout the coding process and become part of your audit trail.
Audit Trail
An audit trail is the complete set of documents that shows how you moved from raw data to findings. It includes transcripts, codebooks, coding logs, memos, and drafts of your analysis. The Reproducible Open Coding Kit supports this kind of transparency by providing a standard for human-readable and machine-readable coded data [13].
Common Failure Patterns in Qualitative Coding
Beginners often encounter predictable problems. Recognizing these patterns helps you avoid them.
Coding Everything Without Focus
Coding every sentence with equal attention produces a large code list with little analytical value. Focus your coding on segments relevant to your research question. If you are using a framework, code segments that relate to framework constructs.
Code Proliferation
Creating a new code for every observation leads to dozens or hundreds of codes that are difficult to manage. Review your code list regularly and merge codes that capture the same idea. Use subcodes to capture variation within a broader category.
Losing the Case Perspective
Coding breaks data into segments, but the meaning of a segment depends on its context. The case-oriented approach emphasizes understanding each case as a whole [9]. After coding, read through each transcript again to ensure that your codes do not distort the participant's overall message.
Premature Closure
Moving to axial and selective coding before you have thoroughly explored the data can cause you to miss important patterns. Allow enough time for open coding and constant comparison. The method of constant comparison involves checking each new code against previously coded data to ensure consistency [9].
Inconsistent Code Application
When codes are not clearly defined, different coders apply them differently. Write definitions for every code and test the codebook on a sample of data before coding the full dataset. Use consensus coding to resolve discrepancies [10].
Limitations of Qualitative Coding
Qualitative coding has limitations that researchers should acknowledge.
Coding Is Interpretive
Coding requires judgment. Different researchers may code the same data differently. This is not a flaw but a feature of qualitative research. The goal is not perfect agreement but transparent and defensible interpretation.
Coding Does Not Quantify
Codes can be counted, but counting codes does not tell you about the importance of a theme. A theme mentioned once may be more significant than a theme mentioned many times. Coding supports qualitative analysis, not statistical generalization.
Saturation Is Uncertain
There is no agreement on the minimum number of interviews needed to achieve saturation [6]. Sample size decisions require judgment based on the research context. Web-based data collection has become increasingly common, and findings around saturation may differ for in-person versus web-based interviews [6].
Framework Constraints
Deductive coding using a framework can constrain what you notice in the data. The framework focuses your attention on certain aspects and may cause you to overlook others. Some implementation science scholars call for greater methodological diversity and reflexive exploration in qualitative research [18].
Safety and Regulatory Context for Life-Science Research
Qualitative research in life sciences involves human participants. Researchers must follow ethical and regulatory requirements.
Institutional Review Board Approval
Studies involving human participants require review and approval by an institutional review board or research ethics committee. This applies to interviews, focus groups, and surveys. Approval must be obtained before data collection begins.
Informed Consent
Participants must provide informed consent before participating. Consent documents should explain the purpose of the research, how data will be used, and how confidentiality will be protected. Participants should be told whether recordings will be made and how transcripts will be stored.
Data Protection
Qualitative data often contains sensitive information. Transcripts should be de-identified by removing names, locations, and other identifying details. Data should be stored securely and access limited to the research team. The National Institute of Standards and Technology Research Data Framework provides guidance on managing research data across its lifecycle [1].
Reporting Standards
Reporting guidelines help researchers document their methods transparently. The EQUATOR Network is an international initiative that provides reporting guidelines for health research [2]. The Consolidated Criteria for Reporting Qualitative Research is one example of a reporting guideline, though existing guidelines provide minimal guidance for documenting large language model use in qualitative research [14].
Professional Escalation Criteria
Some situations require consultation with a supervisor, methodologist, or ethics committee.
Escalate When You Cannot Resolve Coding Disagreements
If your coding team cannot reach consensus on code definitions or application, escalate to a senior researcher or methodologist. Persistent disagreement may indicate that the codebook needs revision or that the research question needs clarification.
Escalate When You Encounter Unexpected Sensitive Data
If transcripts contain disclosures of harm, illegal activity, or other sensitive information, consult your institutional review board or a research ethics advisor. Your consent documents may not cover how to handle such disclosures.
Escalate When Data Quality Is Inadequate
If recordings are inaudible, transcripts are incomplete, or participants did not engage with the questions, consult your supervisor. You may need to revise your data collection approach or collect additional data.
Escalate When You Plan to Use AI Tools
Large language models are being integrated into qualitative research processes, with coding assistance and theme identification as the most common applications [14]. Technical reporting of LLM use is highly inconsistent, with few studies reporting temperature settings, context length, or other parameters [14]. If you plan to use AI tools for coding, consult your supervisor and institutional policies. Interpretation, contextual judgment, theoretical argument, and ethical accountability remain the researcher's duties [20].
Frequently Asked Questions
What is the difference between a code and a theme?
A code is a label applied to a segment of data. A theme is a broader pattern that emerges from the analysis of multiple codes. For example, the code "unreliable Wi-Fi" might contribute to the theme "technical infrastructure barriers." Themes are built from codes and represent the larger ideas that answer your research question.
How many codes should I have?
There is no fixed number. The right number depends on your research question, the diversity of your data, and your analytical approach. A focused study might have 20 to 30 codes. A broad exploratory study might have more. If you have hundreds of codes, review them for overlap and consider merging related codes into categories.
Can I code without qualitative data analysis software?
Yes. Coding can be done with word processing documents, spreadsheets, or printed transcripts. The Reproducible Open Coding Kit is designed to be accessible to researchers without dedicated software while offering enough structure to enable programmatic processing [13]. Software becomes more useful as the dataset grows and when multiple coders are involved.
How do I know when I have coded enough data?
Code saturation is the point at which additional data no longer produce new codes. In one study, near saturation was reached after 33% to 60% of planned interviews, while true saturation was reached after 91% to 100% of planned interviews [6]. Studies that relied heavily on deductive coding and had a more structured interview guide reached saturation sooner [6]. Track new codes as you code to see when the rate of new codes declines.
What is the difference between inductive and deductive coding?
Inductive coding develops codes directly from the data without a preexisting framework. Deductive coding applies a preexisting framework or codebook to the data. Many studies combine both approaches. One health services research approach applies inductive reasoning while also employing predetermined code types to guide analysis and interpretation [11].
How do I handle disagreements between coders?
Use consensus coding to resolve discrepancies [10]. Discuss the disagreement, clarify the code definition, and decide together how the segment should be coded. If disagreements persist, revise the codebook or escalate to a senior researcher. Inter-coder reliability measures such as Cohen kappa can document the level of agreement [16].
Can I use artificial intelligence to code qualitative data?
Large language models are being used for coding assistance and theme identification in qualitative research [14]. However, technical reporting of LLM use is highly inconsistent [14]. Interpretation, contextual judgment, theoretical argument, and ethical accountability remain the researcher's duties [20]. If you use AI tools, document your approach carefully and consult your supervisor and institutional policies.
How do I write about my coding process in a research paper?
Describe your coding approach, including whether you used inductive or deductive coding, how you developed your codebook, how many coders were involved, and how disagreements were resolved. Report your saturation assessment and any inter-coder reliability measures. Reporting guidelines such as those available through the EQUATOR Network can help you document your methods transparently [2].
Related Articles
- Symbiosis in Animals: A Beginner's Guide to Types and Examples
- Snakemake for Research Pipelines: A Practical Starting Framework
- Transcript in Biology
- Annotation Transcript Selection
- Annotation Transcript Selection
References and Further Reading
- Research Data Framework. National Institute of Standards and Technology.
- EQUATOR Network. EQUATOR Network.
- Experimental Design Assistant. NC3Rs.
- NCBI Literature Resources. National Center for Biotechnology Information.
- PubMed. National Library of Medicine.
- Determining an Appropriate Sample Size for Qualitative Interviews to Achieve True and Near Code Saturation: Secondary Analysis of Data.. Journal of medical Internet research, 2024.
- Rapid versus traditional qualitative analysis using the Consolidated Framework for Implementation Research (CFIR).. Implementation science : IS, 2021.
- Use of Framework Matrix and Thematic Coding Methods in Qualitative Analysis for mHealth: NIRUDAK Study Data.. International journal of qualitative methods, 2023.
- Complex Qualitative Data Analysis: Lessons Learned From the Experiences With the Qualitative Analysis Guide of Leuven.. Qualitative health research, 2021.
- Understanding the Use of Mobility Data in Disasters: Exploratory Qualitative Study of COVID-19 User Feedback.. JMIR human factors, 2024.
- Qualitative data analysis for health services research: developing taxonomy, themes, and theory.. Health services research, 2007.
- Development of a method for qualitative data integration to advance implementation science within research consortia.. Implementation science communications, 2025.
- The Reproducible Open Coding Kit: A Human and Machine Readable Standard for Coding Qualitative Data. 2026.
- The use and methodological reporting of large language models in qualitative research: a scoping review.. 2026.
- Cultural foundations of university student well-being: a grounded theory study of a multidimensional framework in Indonesia.. 2026.
- Artificial Intelligence in Surgical Qualitative Research: A Comparison of Human and AI-Assisted Thematic Analysis. 2026.
- Mapping Practice-Based Signals of Generative AI in Psychiatric Care: Qualitative Study of Korean Psychiatrists' Experiences, Interpretations, and Implementation Priorities.. 2026.
- Embracing complexity by reducing rigidity: guiding principles for the intersection of implementation science and qualitative research.. 2026.
- AI-Assisted Rapid Quality Analysis in Implementation Science: Methodological Study.. 2026.
- Contemporary Coding Techniques in Qualitative Analysis: A Review of Saldaña’s The Coding Manual for Qualitative Researchers (5th ed.). The Qualitative Report, 2026.
- An Evaluation of Peace Building Strategies in Southwestern Nigeria: Quantitative and Qualitative Examples. 2021.
- Qualitative Data Analysis: An Expanded Sourcebook. 1994.
- Coding and Writing Analytic Memos on Qualitative Data: A Review of Johnny Saldana's the Coding Manual for Qualitative Researchers. 2018.
- Scaling hermeneutics: a guide to qualitative coding with LLMs for reflexive content analysis. EPJ Data Science, 2025.
This article is educational and does not replace institutional policy, professional advice, or applicable safety and regulatory requirements.