Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up

neural-nova
/
Llama-3.1-8B-Instruct-optimized

cuda
custom-kernels
inference-optimization
llama
Model card Files Files and versions
xet
Community
Llama-3.1-8B-Instruct-optimized
1.42 MB
Ctrl+K
Ctrl+K
  • 2 contributors
History: 4 commits
JDCentral's picture
JDCentral
Update README.md
cb8dc4c verified about 1 month ago
  • kernels
    Add optimized CUDA kernels (RMSNorm, MLP, Attention) and patched Llama-3.1 modeling file about 2 months ago
  • patched_transformers
    Add optimized CUDA kernels (RMSNorm, MLP, Attention) and patched Llama-3.1 modeling file about 2 months ago
  • .gitattributes
    1.84 kB
    Add optimized CUDA kernels (RMSNorm, MLP, Attention) and patched Llama-3.1 modeling file about 2 months ago
  • README.md
    5.21 kB
    Update README.md about 1 month ago
  • requirements.txt
    126 Bytes
    Add optimized CUDA kernels (RMSNorm, MLP, Attention) and patched Llama-3.1 modeling file about 2 months ago