Upload README.md with huggingface_hub
Browse files
README.md
ADDED
|
@@ -0,0 +1,84 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
library_name: autoencodix
|
| 4 |
+
tags:
|
| 5 |
+
- single-cell
|
| 6 |
+
- scRNA-seq
|
| 7 |
+
- autoencoder
|
| 8 |
+
- variational-autoencoder
|
| 9 |
+
- biology
|
| 10 |
+
pipeline_tag: feature-extraction
|
| 11 |
+
---
|
| 12 |
+
|
| 13 |
+
# Varix-Dim48-Varix
|
| 14 |
+
|
| 15 |
+
A [Varix](https://github.com/autoencodix/autoencodix) variational autoencoder with a 48-dimensional latent space, trained on single-cell RNA-seq data. Unlike Ontix, the latent dimensions are not constrained by gene ontologies, providing a compact general-purpose embedding of the transcriptome.
|
| 16 |
+
|
| 17 |
+
Part of the [autoencodix pretrained model collection](https://huggingface.co/collections/autoencodix/acx-pretrained-models).
|
| 18 |
+
|
| 19 |
+
## Usage
|
| 20 |
+
|
| 21 |
+
Install the [`autoencodix`](https://github.com/autoencodix/autoencodix) package:
|
| 22 |
+
|
| 23 |
+
```bash
|
| 24 |
+
pip install autoencodix
|
| 25 |
+
```
|
| 26 |
+
|
| 27 |
+
Download the model and use it to generate embeddings for your own scRNA-seq data (`AnnData` with genes as Ensembl IDs, log1p-normalized counts):
|
| 28 |
+
|
| 29 |
+
```python
|
| 30 |
+
from huggingface_hub import snapshot_download
|
| 31 |
+
import autoencodix as acx
|
| 32 |
+
import anndata as ad
|
| 33 |
+
import pandas as pd
|
| 34 |
+
import numpy as np
|
| 35 |
+
import scanpy
|
| 36 |
+
from scipy import sparse
|
| 37 |
+
from autoencodix.data._numeric_dataset import NumericDataset
|
| 38 |
+
from autoencodix.data._datasetcontainer import DatasetContainer
|
| 39 |
+
from autoencodix.configs.varix_config import VarixConfig
|
| 40 |
+
|
| 41 |
+
# Download and load the pretrained model
|
| 42 |
+
model_name = "Varix-Dim48-Varix"
|
| 43 |
+
repo_id = f"autoencodix/{model_name}"
|
| 44 |
+
model_file = snapshot_download(repo_id=repo_id)
|
| 45 |
+
model_file = model_file + "/large_varix_final_model_Dim48_Varix.pkl"
|
| 46 |
+
|
| 47 |
+
loaded_varix = acx.Varix.load(file_path=model_file)
|
| 48 |
+
loaded_varix._trainer._config.device = "cpu" # Switch device if you want to run on CPU or GPU
|
| 49 |
+
|
| 50 |
+
# Load your data (AnnData with adata.var.index as Ensembl gene IDs)
|
| 51 |
+
adata = ad.read_h5ad("path/to/your_data.h5ad")
|
| 52 |
+
|
| 53 |
+
# Match the input gene space of the pretrained model, zero-padding any missing genes
|
| 54 |
+
anndata_template = ad.AnnData(
|
| 55 |
+
X=sparse.csr_matrix(np.zeros((1, len(loaded_varix.result.model.feature_order)))),
|
| 56 |
+
var=pd.DataFrame(index=loaded_varix.result.model.feature_order),
|
| 57 |
+
)
|
| 58 |
+
adata = ad.concat([anndata_template, adata], axis=0, join="outer").copy()
|
| 59 |
+
adata = adata[1:, loaded_varix.result.model.feature_order].copy()
|
| 60 |
+
|
| 61 |
+
# The model expects log1p normalized counts
|
| 62 |
+
scanpy.pp.log1p(adata, copy=False)
|
| 63 |
+
|
| 64 |
+
test_dataset = NumericDataset(
|
| 65 |
+
data=adata.X,
|
| 66 |
+
config=VarixConfig(),
|
| 67 |
+
sample_ids=adata.obs.index,
|
| 68 |
+
metadata=adata.obs.loc[adata.obs.index, :],
|
| 69 |
+
split_indices=None,
|
| 70 |
+
feature_ids=adata.var.index,
|
| 71 |
+
)
|
| 72 |
+
acx_container = DatasetContainer(train=None, valid=None, test=test_dataset)
|
| 73 |
+
|
| 74 |
+
# Generate embeddings
|
| 75 |
+
result = loaded_varix.predict(data=acx_container)
|
| 76 |
+
df_latent = result.get_latent_df(split="test", epoch=-1)
|
| 77 |
+
df_latent
|
| 78 |
+
```
|
| 79 |
+
|
| 80 |
+
`df_latent` is a `pandas.DataFrame` with one row per cell and 48 columns, one per latent dimension.
|
| 81 |
+
|
| 82 |
+
## Further usage
|
| 83 |
+
|
| 84 |
+
For a full walkthrough including latent space visualization, marker gene explanation with xAI/LLMs, and synthetic data generation, see the [tutorial notebook](https://github.com/autoencodix/autoencodix/blob/main/Tutorials/DeepDives/UsingPreTrainedModels.ipynb).
|