How to Use Google Colab for Bioinformatics: Python, Shell Commands and Google Drive

By Dr. Zubair Khalid, DVM, MS, PhD ·

How to Use Google Colab for Bioinformatics: Python, Shell Commands and Google Drive

Google Colab is a hosted Jupyter Notebook service that requires no setup and provides free access to computing resources, including GPUs and TPUs [1]. For a bioinformatics workflow, that means you can go from a laptop with nothing installed to running Biopython, shell tools and even protein folding notebooks in a browser tab. Jupyter is the open source project Colab is built on, and Colab lets you use and share those notebooks without downloading, installing or running anything locally [1].

By the end of this guide you will be able to start a notebook, install packages, run Python and shell commands in the same document, mount Google Drive so your results survive, and recognize the limits that decide whether Colab is the right place for a given analysis. The worked example parses a small FASTA file and computes GC content, which is enough to confirm the whole pipeline works before you point it at real data.

Quick Answer

  • Open a new notebook at colab.research.google.com. The default runtime is Python.
  • Install what you need inside the notebook, for example %pip install biopython [3].
  • Run shell commands with a leading ! for one line, or %%bash to run a whole cell in bash [3].
  • Mount Drive with drive.mount('/content/drive') if you need files to persist, and remember that any code in the notebook can then read your entire Drive [1].
  • Save anything important to Drive or download it, because files written only to the virtual machine are lost when the VM is recycled [1].
  • Check Runtime > Change Runtime Type if you need a GPU or TPU, and set the accelerator to None when you do not [1].

What Is Google Colab and When Should You Use It

Colab is free of charge to use, and the resources are not guaranteed because Google adjusts usage limits and hardware availability dynamically [1]. That single sentence explains most of the behavior you will see: runtimes disconnect when idle, GPU types change between sessions, and a notebook that ran fast yesterday may queue today. Colab prioritizes users who are actively programming in a notebook, which is why an idle runtime can terminate early [1].

The practical question is not whether Colab is good, but whether your task fits. A teaching notebook, a quick Biopython script, a one-off structure prediction, or a shared analysis for a lab meeting all fit well. A pipeline that must run for days, or that handles protected health information, does not.

Step 1: Create a Notebook and Understand Where It Lives

Go to colab.research.google.com and create a new notebook. Colab notebooks are stored in Google Drive or can be loaded from GitHub, and existing Jupyter or IPython notebooks can be imported [1]. That matters for two reasons. First, your notebook itself is safe even when the runtime dies, because the notebook file is not on the virtual machine. Second, you can hand a colleague a link instead of a folder of scripts.

Your code, however, runs in a virtual machine private to your account. Virtual machines are deleted when idle for a while and have a maximum lifetime enforced by the service, so files saved only on the VM are lost when it is recycled [1]. Treat the VM as scratch space and Drive as storage.

Step 2: Install Packages with Python

Colab focuses on supporting Python and its ecosystem of third-party tools [1]. Installation happens inside the notebook, which keeps the environment reproducible for anyone who opens it later. The %pip magic runs the pip package manager within the current kernel [3]:

%pip install biopython

Biopython provides Seq and SeqRecord objects, Bio.SeqIO and Bio.AlignIO for formats such as FASTA and GenBank, Bio.PDB, and Entrez access [9]. If Biopython is already present in the image, pip simply reports that the requirement is satisfied, so running the install cell is safe either way.

IPython's %conda magic does the same job for conda [3], but only where conda is already installed. To get a full conda installation in Colab, the community project condacolab installs conda for you [5]:

!pip install -q "https://github.com/conda-incubator/condacolab/archive/main.zip"
import condacolab
condacolab.install()

condacolab must run first in the notebook because it triggers a kernel restart that resets variables, and condacolab.check() confirms it worked [5]. It is a community project, not a Google product, so treat it as a convenience, not a supported feature.

Step 3: Run Shell Commands in the Same Notebook

Bioinformatics is full of command-line tools, and Colab does not force you to leave Python to use them. A leading ! runs a shell command from a cell, and %%bash runs the whole cell with bash in a subprocess [3]:

%%bash
echo "shell cell"; python --version

This is how you would call an aligner, count records in a file, or check what is installed. The VM is Linux, so commands you know from a Linux server generally behave the same way. Installing command-line tools with apt-get is common practice in Colab notebooks, but it is not documented on the official Colab pages, so verify that a tool is available before you build a workflow around it.

Step 4: Mount Google Drive So Results Survive

Mounting Drive allows any code in your notebook to access any files in your Google Drive, and you usually have to grant access each time you connect to a new runtime [1]. The drive module signature is mount(mountpoint, force_remount=False, timeout_ms=120000, readonly=False), and flush_and_unmount() unmounts Drive and flushes outstanding writes to it [2]:

from google.colab import drive
drive.mount('/content/drive')

Two failure modes are worth knowing before they surprise you. Mounting can time out when a folder has too many files, and storing more than about ten thousand items in the top-level My Drive can make mounting fail [1]. Mounting can also be slow because Drive files may live in a distant region, so the documentation advises reducing reads and writes from Drive [1]. If you are processing thousands of small files, copy them to local VM storage first, run the analysis, then write the final output back.

The exact subfolder name under the mountpoint is worth checking in your own session. Colab file browsers commonly show MyDrive, while the drive module checks for a My Drive subfolder [2]. Print the directory listing once and use whatever name appears.

Step 5: Choose a Runtime and Know What It Costs You

If you do not need a GPU, the FAQ says to choose Runtime > Change Runtime Type and set Hardware Accelerator to None; the same dialog is where you select an accelerator [1]. This is more than tidiness. Running in a GPU or TPU runtime does not automatically mean the GPU or TPU is being utilized, and the available GPU and TPU types vary over time [1]. An accelerator your code does not use will not speed it up, so leave it off unless you need it.

Runtime > Disconnect and delete runtime returns the assigned virtual machines to their original state [1]. Use it when you finish a session, especially after working with large files or sensitive intermediate data.

In the free version, notebooks can run for at most 12 hours, depending on availability and your usage patterns [1]. Pro+ subscribers can get continuous code execution for up to 24 hours if they have sufficient compute units [1]. Colab does not publish its usage limits in part because they can vary over time, and idle timeout periods, maximum VM lifetime and available GPU types all vary [1]. Paid plans (Colab Pro, Pro+, Pay As You Go) offer increased compute availability based on your compute unit balance, with pricing on the sign-up page [1]. Eligible Google AI plans also include Colab benefits, including a monthly allotment of compute units, more powerful GPUs and TPUs, and background execution on higher plans [1]. Check the current sign-up page for prices, since they change.

Worked Example

This example parses a FASTA string pasted directly into a cell, so no upload is needed. Bio.SeqIO.parse accepts a file name or a file-like object such as io.StringIO, which is what makes this possible.

Cell 1 installs Biopython:

%pip install biopython

Cell 2 parses the sequences and prints length and GC content:

from io import StringIO
from Bio import SeqIO
from Bio.SeqUtils import gc_fraction

fasta_text = """>seq1 example gene fragment
ATGGCGTACGCTAGCTAGGCTTAA
>seq2 AT-rich fragment
ATATTTAAGCATTAAATAT
>seq3 GC-rich fragment
GGCGCCGCGGATCCGCGC
"""

for record in SeqIO.parse(StringIO(fasta_text), "fasta"):
    gc = gc_fraction(record.seq)
    print(f"{record.id}\tlength={len(record)}\tGC={gc * 100:.2f}%")

Expected output:

seq1	length=24	GC=50.00%
seq2	length=19	GC=10.53%
seq3	length=18	GC=88.89%

The numbers check out by hand: seq1 has 12 G or C out of 24 bases, seq2 has 2 of 19, and seq3 has 16 of 18. record.id is the first word of the header line, which is why the output shows seq1, seq2 and seq3 instead of the full descriptions. Biopython's gc_fraction() handles mixed-case sequences and the ambiguous nucleotide S, which stands for G or C [4]. The Biopython tutorial's own check, gc_fraction(Seq("GATCGATGGGCCTATATAGGATCGAAAATCGC")), returns 0.46875 [4]. The Biopython tutorial page showed version 1.88 as of October 2026, and the outputs above were verified with that version; check the download page for the current release [4].

What this result means: you have confirmed that package installation, FASTA parsing and a sequence calculation all work in one notebook. What it does not mean: GC content from three short fragments says nothing about an organism, a genome, or a gene family. Scale the input before drawing biological conclusions.

If you want to inspect sequences interactively before writing code, the site's FASTA Parser is a quick way to check record counts and headers.

Common Mistakes and How to Fix Them

  • "Mountpoint must not already contain files" or "Mountpoint must not contain a space." The mountpoint you passed is not empty or contains a space. Pick an empty directory path without spaces and mount again [2].
  • Drive mount times out. The folder has too many files, or the top-level My Drive holds more than about ten thousand items. Reduce the number of items, or work from local VM storage and copy only what you need [1].
  • Files disappear between sessions. They were written to the VM, not to Drive. Virtual machines are deleted when idle and files saved only on the VM are lost when it is recycled [1]. Mount Drive and write outputs there.
  • The runtime disconnects while you are reading the docs. Colab prioritizes users who are actively programming, so idle runtimes terminate early [1]. Keep long jobs in a cell that is actually running, and save intermediate results.
  • A GPU runtime feels slower than CPU. Running in a GPU or TPU runtime does not automatically mean the accelerator is being used [1]. Check that your code targets the device, or set Hardware Accelerator to None [1].
  • condacolab install fails or variables vanish. condacolab must run first in the notebook because it restarts the kernel and resets variables [5]. Move it to the top and run condacolab.check().
  • You expected an R kernel. The official FAQ states that Colab focuses on Python and that support for other Jupyter kernels such as R or Scala has no ETA [1]. Some tutorials report an R option in the runtime dialog, but no official Google page confirms it, so treat R support as unofficial. Running R through rpy2 from Python is an alternative, though that path is not documented officially either.

Limitations

Colab is not a compute cluster and does not pretend to be one. Resources are not guaranteed, limits are unpublished and change over time, and the free tier caps notebook runtime at 12 hours depending on availability and usage patterns [1]. The 24-hour Pro+ figure assumes sufficient compute units [1]. Neither number is a promise.

Data handling deserves its own paragraph. The Colab FAQ privacy note for AI features asks users not to include sensitive or personal information in prompts or feedback, and states that Google collects prompts, code and outputs from generative AI features, with human reviewers potentially reading them [1]. No Colab-specific statement about uploading patient data was found. Under HIPAA, de-identification can be done by Expert Determination or Safe Harbor, and Safe Harbor removes 18 identifier types including names, geographic units smaller than a state, all date elements except year, phone numbers, email addresses and medical record numbers [7]. Follow your institution's IRB and data policies, and do not upload identifiable patient or client data.

Restricted activities in Colab include file hosting, media serving, cryptocurrency mining, denial-of-service attacks, password cracking and creating deepfakes [1]. Using Colab as a general-purpose file or web host falls under these restrictions.

For protein folding, ColabFold's notebooks include AlphaFold2_mmseqs2, AlphaFold2_batch, ESMFold and an AlphaFold3 (OpenFold3) notebook, plus beta notebooks [6]. Its README notes that limits depend on the free GPU Colab provides, roughly 2000 amino acids maximum on a GPU with about 16 GB memory, and recommends LocalColabFold for local runs [6]. The notebook list and that residue limit change with releases, so check the repository. The underlying method reports a 40 to 60 fold faster MSA search using MMseqs2 and close to 1,000 structure predictions per day on a server with one GPU [8].

Frequently Asked Questions

What is Google Colab, in one paragraph?

It is a hosted Jupyter Notebook service that requires no setup and gives free access to computing resources including GPUs and TPUs [1]. You write and run code in a browser, and the notebook file lives in Google Drive or loads from GitHub [1]. It is best understood as a convenient, non-guaranteed compute environment, not a production server.

Google Colab vs Jupyter Notebook: what is the difference?

Jupyter is the open source project Colab is based on, and Colab lets you use and share Jupyter notebooks without downloading, installing or running anything [1]. A local Jupyter install gives you control over the environment and persistent storage. Colab gives you zero setup and occasional GPUs, at the cost of unpublished limits and ephemeral virtual machines [1].

Do I need Google Colab Pro?

Only if the free tier's limits block your work. Colab Pro, Pro+ and Pay As You Go offer increased compute availability based on your compute unit balance, with pricing on the sign-up page [1]. Eligible Google AI plans also include Colab benefits such as a monthly compute unit allotment and more powerful accelerators [1]. Colab Pro for Education subscriptions were free one-year Colab Pro subscriptions for students and faculty at US-based universities, described in the past tense in the FAQ, so check whether the program is currently open [1].

Can I run R in Google Colab?

Not officially. The FAQ states that Colab focuses on Python and its ecosystem, and that support for other Jupyter kernels such as R or Scala has no ETA [1]. Some tutorials report an R option in the runtime dialog, but no official Google page confirms it. If you need R, run it locally, or call it from Python with rpy2, noting that the rpy2 route is not documented by Google either.

How do I mount Google Drive in Colab?

Use drive.mount('/content/drive') from the google.colab module, and expect to grant access each time you connect to a new runtime [1]. Mounting gives any code in the notebook access to any file in your Drive, so be deliberate about what the notebook can reach [1]. If mounting is slow or times out, reduce the number of files in the folder or in top-level My Drive [1].

References

  1. Google Colaboratory: Frequently Asked Questions
  2. googlecolab/colabtools: google/colab/drive.py (Colab Drive mount source)
  3. IPython documentation: Built-in magic commands
  4. Biopython Tutorial and Cookbook: Sequence objects
  5. conda-incubator/condacolab README (GitHub)
  6. sokrypton/ColabFold README (GitHub)
  7. HHS: Guidance Regarding Methods for De-identification of Protected Health Information
  8. Mirdita M et al. ColabFold: making protein folding accessible to all. Nat Methods. 2022;19(6):679-682
  9. Cock PJA et al. Biopython: freely available Python tools for computational molecular biology and bioinformatics. Bioinformatics. 2009;25(11):1422-1423

Related Articles