Skip to content

Single-cell and spatial transcriptomics

IsoQuant supports single-cell and spatial long-read data obtained using various protocols. When a single-cell or spatial mode is selected, IsoQuant automatically performs barcode calling and UMI-based PCR deduplication as part of the pipeline.

Overview

The single-cell/spatial pipeline extends the standard bulk pipeline with these additional steps:

  1. Barcode calling -- extract cell/spot barcodes and UMIs from raw reads;
  2. Standard IsoQuant processing -- alignment, read-to-isoform assignment;
  3. UMI deduplication -- remove PCR/RT duplicates;
  4. Grouped quantification -- produce per-cell/per-spot gene and transcript counts.

Note: UMI deduplication relies on read-to-gene assignment. Reads that are not assigned to any gene are discarded. Hence, novel gene discovery will not be performed in single-cell/spatial mode. We recommend using bulk mode for novel gene and transcript discovery.

Barcode calling is handled by the built-in barcode calling module, which can also be used as a standalone tool.

Supported protocols

Mode Platform Barcode whitelist files UMI length Notes
tenX_v3 10x Genomics 3' v3 1 12 Single 16bp barcode
tenX_v2 10x Genomics 3' v2 1 10 Single 16bp barcode (v2 chemistry)
visium_5prime 10x Genomics Visium 5' 1 12 Same detector as tenX_v3
visium_hd 10x Genomics Visium HD 2 9 Two barcodes (15/16bp + 14/15bp)
curio Curio Bioscience 1 9 Double barcode (8bp + 6bp) with linker
stereoseq Stereo-seq 1 10 25bp barcode
custom_sc Any protocol 0 (uses MDF) varies User-defined molecule structure via MDF file

Quick start examples

10x Genomics single-cell:

isoquant.py --reference genome.fa --genedb genes.gtf --complete_genedb \
  --fastq reads.fastq.gz --data_type nanopore \
  --mode tenX_v3 --barcode_whitelist barcodes.txt.gz \
  -o sc_output

Stereo-seq spatial:

isoquant.py --reference genome.fa --genedb genes.gtf --complete_genedb \
  --fastq reads.fastq.gz --data_type nanopore \
  --mode stereoseq --barcode_whitelist barcodes.txt \
  -o stereo_output

Custom protocol with molecule definition file:

isoquant.py --reference genome.fa --genedb genes.gtf --complete_genedb \
  --fastq reads.fastq.gz --data_type nanopore \
  --mode custom_sc --molecule molecule_definition.mdf \
  -o custom_output

Pre-called barcodes (skip barcode calling):

isoquant.py --reference genome.fa --genedb genes.gtf --complete_genedb \
  --bam aligned.bam --data_type nanopore \
  --mode tenX_v3 --barcoded_reads barcodes.tsv \
  -o sc_output

Command line options

--mode or -m

IsoQuant processing mode. Available modes:

  • bulk -- standard bulk RNA-seq mode (default)
  • tenX_v3 -- 10x Genomics single-cell 3' gene expression
  • tenX_v2 -- 10x Genomics single-cell 3' v2 gene expression
  • curio -- Curio Bioscience single-cell
  • visium_hd -- 10x Genomics Visium HD spatial transcriptomics
  • visium_5prime -- 10x Genomics Visium 5' spatial transcriptomics
  • stereoseq -- Stereo-seq spatial transcriptomics
  • custom_sc -- custom single-cell/spatial mode using a molecule definition file (MDF)

All modes except bulk enable automatic barcode calling and UMI-based deduplication.

--split_molecules

Whether to split reads containing several cDNA molecules, see read splitting. auto (default) splits wherever the protocol supports it; true also fails for protocols that cannot split; false never splits.

--barcode_whitelist

Path to file(s) with barcode whitelist(s) for barcode calling. Required for single-cell/spatial modes unless --barcoded_reads or --barcoded_bam is provided.

File should contain one barcode sequence per line. More than 1 tab-separated column is allowed, but only the first will be used. Supports plain text and gzipped files.

Accepts the literal value auto instead of a file, see detecting cell barcodes.

Notes: - What the whitelist means depends on --n_cells, see detecting cell barcodes. Without it the whitelist is taken to be the list of cell barcodes; - If you have a subset of barcodes from short-read data, providing them directly also works and skips the extra pass over the reads; - IsoQuant will perform per-barcode quantification automatically unless --barcoded_reads or --barcode2spot are set. Use --read_group barcode to group reads by barcode explicitly. In case of a large number of barcodes, it may take a lot of time.

The number of whitelist files depends on the mode:

  • 1 file: tenX_v3, tenX_v2, visium_5prime, stereoseq, curio (combined 14bp barcodes)
  • 2 files: visium_hd (part 1 and part 2 barcode lists)
  • Not needed: custom_sc (barcode lists are specified inside the MDF file)

--n_cells

Expected number of cell-associated barcodes, or auto to estimate it from the data. Turns --barcode_whitelist into a pool of candidates rather than the cell list itself, see detecting cell barcodes.

--n_cells_interval

Percentage by which the number of selected barcodes may differ from --n_cells [25].

--barcode_correction

Overrides how the cell barcode list is obtained, which is otherwise decided by --n_cells: auto (default), whitelist (always use the whitelist as given), or detect (always detect cell barcodes from read counts).

--barcoded_reads

Path to TSV file(s) with pre-called barcoded reads. Format: read_id<TAB>barcode<TAB>umi (one read per line). If provided, IsoQuant skips barcode calling and uses these assignments directly. More than 3 columns are allowed, but only the first 3 will be used. Mutually exclusive with --barcode_whitelist and --barcoded_bam.

Note that IsoQuant will perform per-barcode quantification automatically unless --barcoded_reads or --barcode2spot are set. Use --read_group barcode to group reads by barcode explicitly. In case of a large number of barcodes, it may take a lot of time.

--barcoded_bam

Extract barcodes and UMIs from BAM tags. Uses CB (cell barcode) and UB (UMI) tags by default. Mutually exclusive with --barcode_whitelist and --barcoded_reads. This option is a flag, not a way to provide the file (use --bam to set the input BAM).

Note that IsoQuant will perform per-barcode quantification automatically unless --barcoded_reads or --barcode2spot are set. Use --read_group barcode to group reads by barcode explicitly. In case of a large number of barcodes, it may take a lot of time.

--barcode_tag BAM tag for cell barcode (default: CB), requires --barcoded_bam flag.

--umi_tag BAM tag for UMI (default: UB), requires --barcoded_bam flag.

--strip_barcode_suffix Remove suffix after dash from barcodes extracted from BAM tag (e.g. ACGT-1 -> ACGT), requires --barcoded_bam flag.

--barcode2spot

Path to a TSV file mapping barcodes to cell types, spatial spots, or other barcode properties. By default, barcode is in the first column, cell type in the second. However, you can specify one or more columns via colon symbol (similar to --read_group): file.tsv:barcode_column:spot_column(s) (e.g., cell_types.tsv:0:1,2,3 for multiple barcode properties).

When --barcode2spot is set, --read_group barcode_spot will be set automatically to group counts by cell type, spatial regions, or other provided properties. This will also turn off automatic per-barcode grouping (set --read_group barcode to enable).

--molecule

Path to a molecule description format (MDF) file for custom_sc mode. This file defines the structure of the sequencing molecule (barcodes, UMIs, linkers, polyT, cDNA) and allows IsoQuant to process reads from any single-cell or spatial protocol. See the MDF format section below for details.

--barcode2barcode

This option is still under development.

Path to a TSV file with mapping barcodes to spot IDs for spot-level UMI deduplication. When multiple barcodes map to the same physical spot (e.g. at lower spatial resolution), this option groups them together during UMI deduplication, collapsing duplicates across the entire spot (not just the barcode).

Format: file.tsv or file.tsv:barcode_col:spot_col(s) (same syntax as --barcode2spot). When multiple spot columns are provided, a separate UMI deduplication round is performed for each column and outputs corresponding UMI-deduplicated reads only in allinfo format. However, only the main (barcode-based) UMI-deduplicated reads will be used for quantification.

When --barcode2barcode is set, --read_group barcode_barcode will be set automatically. Grouping by raw barcode sequences will be disabled in this case, use --read_group barcode to enable it.

For Visium HD composite barcodes, use isoquant_lib/scripts/prepare_visium_spot_ids.py to generate the mapping file.

Molecule description format (MDF)

The MDF format allows users to describe the structure of their sequencing molecule so that IsoQuant can extract barcodes and UMIs from arbitrary protocol. The molecule is described in the 3' to 5' direction (primer end first, cDNA last).

An MDF file has two parts:

  1. Header line: colon-separated list of element names defining the order of elements on the molecule (3' to 5').
  2. Element definitions: one line per element, tab-separated: name type value [length].

Element types

Type Description Value field
CONST Constant/known sequence (primer, linker, TSO) Sequence
VAR_FILE Variable sequence matched against a whitelist file Path to TSV file
VAR_LIST Variable sequence matched against an inline list Comma-separated sequences
VAR_ANY Variable fixed-length sequence extracted as-is Length (integer)
VAR_ANY_SEPARATOR Fixed-length variable separator sequence (not extracted) Length (integer)
VAR_ANY_NON_T_SEPARATOR Fixed-length variable separator sequence without T nucleotides (not extracted) Length (integer)
PolyT PolyT tail (none)
cDNA cDNA region (none)

Currently, only a single cDNA and a single PolyT are supported. Protocols that concatenate multiple cDNAs (like Kinnex) will be supported in one of the next releases.

Variable elements are expected to have a fixed length (VAR_FILE and VAR_LIST). Using variable-length barcodes may result in suboptimal performance.

Barcode and UMI identification

Elements are identified as barcodes or UMIs by their name prefix (case-insensitive): - Names starting with barcode are treated as barcode elements (e.g., Barcode, barcode1) - Names starting with umi are treated as UMI elements (e.g., UMI, umi1)

When multiple barcode (or UMI) elements are present, their sequences are concatenated in the order they appear.

Examples

10x 3' single-cell v3:

R1:Barcode:UMI:PolyT:cDNA:TSO
R1        CONST      CTACACGACGCTCTTCCGATCT
Barcode   VAR_FILE   barcodes.tsv
UMI       VAR_ANY    12
TSO       CONST      CCCATGTACTCTGCGTTGATACCACTGCTT

10x 3' single-cell v3 (inline barcodes):

R1:Barcode:UMI:PolyT:cDNA:TSO
R1        CONST      CTACACGACGCTCTTCCGATCT
Barcode   VAR_LIST   AAACCCGGGTTTAAAC,TTTGGGCCCAAATTTG,GGGGAAAACCCCTTTT
UMI       VAR_ANY    12
TSO       CONST      CCCATGTACTCTGCGTTGATACCACTGCTT

Linked elements for multi-part barcodes

When a barcode is split into multiple parts in the molecule, use linked element notation. All linked parts must be of the same variable type and reference the same whitelist. Parts are numbered consecutively starting from 1.

Concatenated (|): parts of the barcode separated by other elements. Parts are extracted independently, then concatenated and corrected as one sequence against the full whitelist. Such structure appears in, for example, Curio Bioscience single-cell protocol:

PCR_PRIMER:Barcode|1:Linker:Barcode|2:UMI:PolyT:cDNA
PCR_PRIMER  TACACGACGCTCTTCCGATCT
Barcode|1   VAR_FILE   barcodes.tsv   8
Linker      CONST      TCTTCAGCGTTCCCGAGA
Barcode|2   VAR_FILE   barcodes.tsv   6
UMI         VAR_ANY    9

Duplicated (/): multiple redundant copies of the same barcode in the molecule. Each copy is corrected independently against the whitelist; a majority vote determines the final barcode. For example:

PCR_PRIMER:Barcode/1:UMI:PolyT:cDNA:Barcode/2:TSO
PCR_PRIMER  TACACGACGCTCTTCCGATCT
Barcode/1   VAR_FILE   barcodes.tsv   16
Barcode/2   VAR_FILE   barcodes.tsv   16
UMI         VAR_ANY    12
TSO       CONST      CCCATGTACTCTGCGTTGATACCACTGCTT

Detecting cell barcodes

A stock 10x whitelist has millions of entries, but a run has a few thousand cells. Matching every read against millions of candidates only means something if the match is exact, so any read carrying a sequencing error in the barcode is lost.

--n_cells changes what --barcode_whitelist means so that this is not necessary:

--n_cells meaning of --barcode_whitelist passes over the reads
not set the list of cell barcodes; reads are matched against all of it 1
a number, or auto a pool to select the cell barcodes from 2

With --n_cells set, the first pass extracts barcode windows from the reads verbatim and counts them, the counts pick the cell barcodes out of the pool, and the second pass is ordinary barcode calling against that much shorter list. Correction itself is unchanged — only the list being matched against differs.

python isoquant.py --mode tenX_v3 \
  --barcode_whitelist 3M-february-2018.txt.gz --n_cells auto \
  --fastq reads.fastq.gz --reference genome.fa --genedb genes.gtf \
  --data_type nanopore -o output_dir

auto is as accurate as giving the exact cell count, and selection tolerates being off by roughly ±25% (--n_cells_interval). Supported for tenX_v3, tenX_v2 and visium_5prime.

--barcode_whitelist auto detects cell barcodes with no pool at all. This is less accurate: a mis-anchored barcode window repeats across reads and looks exactly like an abundant cell, whereas a whitelist rejects it because it is not a valid protocol barcode.

The detected list is written to <prefix>.cell_barcodes.tsv, with a summary of how many barcodes were extracted, skipped and selected in <prefix>.cell_barcodes.stats.

Read splitting

Oxford Nanopore long reads often contain 2 or more cDNA molecules ligated end to end during library preparation. Without splitting, only the first barcode is detected and the rest of the read is treated as a single cDNA. With splitting, IsoQuant detects every barcode/UMI pattern in the read, cuts the read at molecule boundaries, and produces a new FASTA file with one record per cDNA molecule.

Splitting is controlled by --split_molecules:

value behaviour
auto (default) split wherever the protocol supports it, do nothing where it does not
true split, and stop with an error if the protocol cannot split
false never split

Supported for tenX_v3, tenX_v2, stereoseq and visium_5prime. --split_molecules true with any other mode is an error rather than a silent no-op, so a request that cannot be honoured never passes unnoticed.

When splitting, IsoQuant writes an additional output file (*.split_reads_<i>.fa.gz) containing the extracted cDNA segments, and uses it in place of the original reads for alignment. Each segment is named with the original read ID plus coordinates and strand: {read_id}_{start}_{end}_{strand}. The file is gzipped unless --no_gzip is set; minimap2 reads it compressed, so alignment is not slowed down.

Splitting rewrites the reads, so the pieces have to be aligned afresh. Supplying an aligned BAM (--bam) says the opposite - that no mapping should happen - so the two requests contradict each other and IsoQuant aborts rather than guessing. This applies to auto as well as true: pass the raw reads if you want the molecules split and re-aligned, or --split_molecules false to analyse the alignments as they are.

Splitting is worth leaving on even for libraries you do not expect to be concatenated: measured on non-concatenated 10x data it recovers about one extra point of recall at unchanged precision (87.3% → 88.3% on ONT R9, 89.6% → 90.5% on ONT cDNA R10.4), at roughly twice the barcode calling runtime.

The superseded mode names tenX_v3_split, tenX_v2_split and stereoseq_nosplit still work and are translated to the corresponding --mode plus --split_molecules combination.

UMI deduplication

All single-cell and spatial modes perform UMI-based PCR deduplication after isoform assignment. Within each cell barcode and gene, reads with similar UMIs (similarity criteria depends on the UMI length) are collapsed into a single representative read.

The representative read is selected based on: 1. Unique isoform assignment over ambiguous 2. More exons 3. Longer transcript alignment

The resulting reads after UMI-deduplication are used for the subsequent analysis such as quantification and are stored in a read_info.tsv file.

Spot-level UMI deduplication

When multiple barcodes correspond to the same spatial spot, UMI deduplication can be performed using these spots instead of individual barcodes. In this case, use --barcode2barcode option to provide barcode to spot mapping (see details above).

isoquant.py --reference genome.fa --genedb genes.gtf --complete_genedb \
  --fastq reads.fastq.gz --data_type nanopore \
  --mode visium_hd --barcode_whitelist part1.txt part2.txt \
  --barcode2barcode barcode2spot.tsv:0:1,2 \
  -o visium_output

For Visium HD, generate the mapping file from per-part coordinate files:

python isoquant_lib/scripts/prepare_visium_spot_ids.py part1_to_y.tsv part2_to_x.tsv -o barcode2spot.tsv

Output

Count matrices

Single-cell/spatial modes produce grouped count matrices in addition to the standard IsoQuant output. Use --counts_format to control the output format:

  • default - automatic selection: matrix format for small numbers of groups (<=100), MTX for larger datasets;
  • matrix - standard matrix format with genes/transcripts as rows and barcodes as columns (may take a lot of space for large group numbers);
  • mtx - Matrix Market (MTX) format compatible with Seurat and Scanpy;
  • none - no conversion (only internal linear format is produced);

Grouped counts can also be converted after the run using isoquant_lib/quantification/convert_grouped_counts.py.

Grouping counts

Use --read_group to control how reads are grouped for quantification. Multiple grouping strategies can be combined (space-separated), producing separate count tables for each.

In single-cell/spatial modes, IsoQuant automatically adds --read_group barcode only when --barcoded_reads or --barcode2spot are not set. Use --read_group barcode to group reads by barcode explicitly. In case of a large number of barcodes, it may take a lot of time. Use --read_group no_auto to disable automatic grouping.

The most common use-cases for single-cell/spatial data are grouping by barcode property (cell type, spot):

--read_group barcode_spot (requires --barcode2spot)

--read_group barcode_barcode (requires --barcode2barcode)

or grouping by individual barcode (not recommended for datasets with many barcodes):

--read_group barcode

See the read grouping options for the full list of grouping strategies.