InstaDeepAI
/

segment_nt

Feature Extraction

Model card Files Files and versions

hdallatorre commited on Mar 29, 2024

Commit

38f0e98

·

verified ·

1 Parent(s): 991981b

Update README.md

Files changed (1) hide show

README.md +2 -1

README.md CHANGED Viewed

@@ -41,7 +41,8 @@ the `rescaling_factor` of the Rotary Embedding layer in the esm model  `num_dna_
 (i.e 6669 for a sequence of 40008 base pairs) and `max_num_tokens_nt` is the max number of tokens on which the backbone nucleotide-transformer was trained on, i.e `2048`.
 [![Open All Collab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/#fileId=https%3A//huggingface.co/InstaDeepAI/segment_nt/blob/main/inference_segment_nt.ipynb)
-The `./inference_segment_nt.ipynb` can be run in Google Colab by clicking on the icon and shows how to set the rescaling factor and infer on a 50kb genic sequence of the human chromosome 20.
 ```python
 # Load model and tokenizer

 (i.e 6669 for a sequence of 40008 base pairs) and `max_num_tokens_nt` is the max number of tokens on which the backbone nucleotide-transformer was trained on, i.e `2048`.
 [![Open All Collab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/#fileId=https%3A//huggingface.co/InstaDeepAI/segment_nt/blob/main/inference_segment_nt.ipynb)
+The `./inference_segment_nt.ipynb` can be run in Google Colab by clicking on the icon and shows how to handle inference on sequence lengths require changing
+the rescaling factor and sequence lengths that do not. One can run the notebook and reproduce Fig.1 and Fig.3 from the SegmentNT paper.
 ```python
 # Load model and tokenizer