jkkkkyuedtrt's picture
Upload 381 files
bd52a47 verified
|
Raw
History Blame Contribute Delete
4.49 kB

Spatial Mapping with Tangram

Map single-cell RNA-seq data onto spatial transcriptomics to enable cell type deconvolution and gene expression imputation.

Workflow Overview

  1. Explore datasets (always do this first)
  2. Preprocess data for Tangram
  3. Map cells to spatial locations
  4. Project cell type annotations
  5. Project gene expression
  6. Evaluate mapping quality

Step 1: Explore Datasets

ALWAYS start here. Understand both datasets:

import scanpy as sc

# Load scRNA-seq reference
adata_sc = sc.read_h5ad("path/to/scrna.h5ad")
print(f"scRNA-seq: {adata_sc.shape}")
print(f"Cell types: {adata_sc.obs['cell_type'].nunique()}")
print(f"Obs columns: {list(adata_sc.obs.columns)}")

# Load spatial data
adata_sp = sc.read_h5ad("path/to/spatial.h5ad")
print(f"Spatial: {adata_sp.shape}")
print(f"Has spatial coords: {'spatial' in adata_sp.obsm}")

Identify:

  • Cell type annotation column in scRNA-seq (e.g., 'cell_type', 'cell_subclass')
  • Whether datasets are from same tissue/species
  • Data normalization status

Step 2: Preprocess for Tangram

Tool: tangram_preprocess

Purpose: Find shared genes and compute density priors.

Parameters:

  • marker_genes="auto": Auto-compute from differential expression
  • cell_type_key: Column with cell type annotations
  • n_markers: Top markers per cell type (default 100)

Outputs:

  • tangram_sc_prep.h5ad: Preprocessed scRNA-seq
  • tangram_sp_prep.h5ad: Preprocessed spatial with density priors

Step 3: Map Cells to Space

Tool: tangram_map_cells

Purpose: Create cell-to-spot probability matrix.

Modes:

  • mode="cells": Single-cell resolution (GPU recommended, slower)
  • mode="clusters": Cluster averages (faster, good for cross-sample)

Parameters:

  • device="cuda:0" for GPU, "cpu" for CPU
  • num_epochs: 500-1000 typical

Output: tangram_mapping.h5ad (cells × spots probability matrix)


Step 4: Project Cell Annotations

Tool: tangram_project_annotations

Purpose: Transfer cell type probabilities to spatial spots.

Parameters:

  • annotation: Cell type column to project

Outputs:

  • tangram_annotated.h5ad: Spatial data with predictions
  • tangram_celltype_probs.csv: Probability matrix

Step 5: Project Gene Expression

Tool: tangram_project_genes

Purpose: Impute full transcriptome onto spatial locations.

Use cases:

  • Visualize genes not in spatial panel
  • Correct dropout in spatial measurements
  • Analyze any gene spatially

Output: tangram_projected.h5ad (spots × genes)


Step 6: Evaluate Mapping

Tool: tangram_evaluate

Purpose: Validate mapping by comparing predicted vs measured expression.

Metrics:

  • Training score > 0.7: Good mapping
  • AUC > 0.5: Reasonable test predictions
  • Low scores often from sparsity mismatch (dropout)

Outputs:

  • tangram_scores.csv: Per-gene scores
  • tangram_auc.png: Validation curve

Visualization Examples

After mapping, visualize results:

import scanpy as sc
import matplotlib.pyplot as plt

# Load annotated spatial data
adata_sp = sc.read_h5ad("experiments/tangram_annotated.h5ad")

# Plot cell type probabilities
sc.pl.spatial(adata_sp, color="tangram_celltype", title="Predicted Cell Types")

# Load projected genes
adata_ge = sc.read_h5ad("experiments/tangram_projected.h5ad")

# Compare measured vs predicted for specific gene
gene = "CD3E"
fig, axes = plt.subplots(1, 2, figsize=(10, 4))
sc.pl.spatial(adata_sp, color=gene, ax=axes[0], title=f"{gene} (measured)")
sc.pl.spatial(adata_ge, color=gene, ax=axes[1], title=f"{gene} (predicted)")

Output Files

  • tangram_sc_prep.h5ad - Preprocessed scRNA-seq
  • tangram_sp_prep.h5ad - Preprocessed spatial
  • tangram_mapping.h5ad - Cell-to-spot probability matrix
  • tangram_annotated.h5ad - Spatial with cell type predictions
  • tangram_celltype_probs.csv - Cell type probability matrix
  • tangram_projected.h5ad - Imputed gene expression
  • tangram_scores.csv - Per-gene mapping scores
  • tangram_auc.png - Validation AUC curve

Tips

  1. GPU recommended for mode="cells" with large datasets
  2. Use mode="clusters" when scRNA-seq and spatial are from different samples
  3. More training genes = better mapping, but needs shared expression
  4. Low scores for sparse genes are expected (dropout in spatial data)
  5. Check that tissue types match between scRNA-seq and spatial