Instructions to use replicate/paged-attention with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Kernels
How to use replicate/paged-attention with Kernels:
# !pip install kernels from kernels import get_kernel kernel = get_kernel("replicate/paged-attention") - Notebooks
- Google Colab
- Kaggle
Download tests/kernels/__init__.py from replicate/paged-attention: direct link, hf CLI and curl.
- Browser
- Download file 0 Bytes
-
https://huggingface.co/replicate/paged-attention/resolve/main/tests/kernels/__init__.py
- Command line
-
hf download hf://replicate/paged-attention/tests/kernels/__init__.py
-
curl -L -o __init__.py https://huggingface.co/replicate/paged-attention/resolve/main/tests/kernels/__init__.py
0 Bytes