Buckets:
| # 🤗 Use Hugging Face Inference Providers with GitHub Copilot Chat in VS Code | |
|  | |
| Use frontier open LLMs like Kimi K2, DeepSeek V3.1, GLM 4.5 and more in VS Code with GitHub Copilot Chat powered by [Hugging Face Inference Providers](https://huggingface.co/docs/inference-providers/index) 🔥 | |
| ## ⚡ Quick start | |
| 1. Install the HF Copilot Chat extension [here](https://marketplace.visualstudio.com/items?itemName=HuggingFace.huggingface-vscode-chat). | |
| 2. Open VS Code's chat interface. | |
| 3. Click the model picker and click "Manage Models...". | |
| 4. Select "Hugging Face" provider. | |
| 5. Enter your Hugging Face Token. You can get one from your [settings page](https://huggingface.co/settings/tokens/new?ownUserPermissions=inference.serverless.write&tokenType=fineGrained). | |
| 6. Choose the models you want to add to the model picker. 🥳 | |
| > [!TIP] | |
| > VS Code 1.104.0+ is required to install the HF Copilot Chat extension. If "Hugging Face" doesn't appear in the Copilot provider list, update VS Code, then reload. | |
| ## ✨ Why use the Hugging Face provider in Copilot | |
| - Access [SoTA open‑source LLMs](https://huggingface.co/models?pipeline_tag=text-generation&inference_provider=all&sort=trending) with tool calling capabilities. | |
| - Single API to switch between multiple providers like Groq, Cerebras, Together AI, and more. | |
| - Built for high availability (across providers) and low latency. | |
| - Transparent pricing: what the provider charges is what you pay. | |
| 💡 Every Hugging Face user gets monthly inference credits to experiment, and can purchase additional credits for pay‑as‑you‑go access. Upgrade to [Hugging Face PRO](https://huggingface.co/pro) or [Team or Enterprise](https://huggingface.co/enterprise) for $2 in monthly credits! | |
| Check out the whole workflow in action in the video below: | |
Xet Storage Details
- Size:
- 1.93 kB
- Xet hash:
- d5bd8ba161f84b5df9e6439bc9eb4c04a8e6c8a61159537dd7f19abc719043a6
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.