Instructions to use replicate/mra with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Kernels
How to use replicate/mra with Kernels:
# !pip install kernels from kernels import get_kernel # a version (or an explicit revision) is required; see the "Files and versions" tab for the available ones kernel = get_kernel("replicate/mra", version=1) - Notebooks
- Google Colab
- Kaggle
Download torch-ext/cuda_launch.h from replicate/mra: direct link, hf CLI and curl.
- Browser
- Download file 703 Bytes
-
https://huggingface.co/replicate/mra/resolve/main/torch-ext/cuda_launch.h
- Command line
-
hf download hf://replicate/mra/torch-ext/cuda_launch.h
-
curl -L -o cuda_launch.h https://huggingface.co/replicate/mra/resolve/main/torch-ext/cuda_launch.h
703 Bytes
| std::vector<at::Tensor> index_max_kernel( | |
| at::Tensor index_vals, | |
| at::Tensor indices, | |
| int A_num_block, | |
| int B_num_block | |
| ); | |
| at::Tensor mm_to_sparse_kernel( | |
| at::Tensor dense_A, | |
| at::Tensor dense_B, | |
| at::Tensor indices | |
| ); | |
| at::Tensor sparse_dense_mm_kernel( | |
| at::Tensor sparse_A, | |
| at::Tensor indices, | |
| at::Tensor dense_B, | |
| int A_num_block | |
| ); | |
| at::Tensor reduce_sum_kernel( | |
| at::Tensor sparse_A, | |
| at::Tensor indices, | |
| int A_num_block, | |
| int B_num_block | |
| ); | |
| at::Tensor scatter_kernel( | |
| at::Tensor dense_A, | |
| at::Tensor indices, | |
| int B_num_block | |
| ); | |