Buckets:
| # Quickstart | |
| Deploy a model from the Hugging Face collection on Microsoft Foundry, then call it from Python. | |
| ## Before you begin | |
| You need: | |
| - An Azure subscription and a Microsoft Foundry project | |
| - `Cognitive Services Contributor` access to create the deployment | |
| - Managed compute quota for the accelerator that your model requires | |
| > [!WARNING] | |
| > Managed compute runs on dedicated GPU capacity and bills for as long as the deployment exists. Delete the deployments that you no longer need. | |
| ## Deploy a model | |
| Pick the Foundry portal for a first, guided deployment, or Python for a repeatable one. Both create the same managed compute deployment. | |
| 1. Open Microsoft Foundry and select your project. | |
| 2. Go to **Discover → Models**, and filter **Collection** by **Hugging Face** and **Deployment options** by **Managed compute**. | |
| 3. Select a model and choose **Deploy**. | |
| 4. Enter a deployment name, pick a deployment template and its accelerator, and select **Deploy**. | |
| The deployment is ready when its state becomes `Succeeded`. | |
| Install the dependencies and sign in: | |
| ```bash | |
| pip install azure-identity "azure-mgmt-cognitiveservices>=15.0.0b2" openai | |
| az login | |
| ``` | |
| Copy the **Model ID**, **Deployment template ID**, and **Accelerator type** from the model's deployment wizard in Foundry, then run: | |
| ```python | |
| from azure.identity import DefaultAzureCredential | |
| from azure.mgmt.cognitiveservices import CognitiveServicesManagementClient | |
| SUBSCRIPTION_ID = "<SUBSCRIPTION_ID>" | |
| RESOURCE_GROUP = "<RESOURCE_GROUP>" | |
| ACCOUNT_NAME = "<ACCOUNT_NAME>" | |
| DEPLOYMENT_NAME = "qwen--qwen3-8-27b" | |
| client = CognitiveServicesManagementClient( | |
| credential=DefaultAzureCredential(), | |
| subscription_id=SUBSCRIPTION_ID, | |
| ) | |
| deployment = client.managed_compute_deployments.begin_create_or_update( | |
| resource_group_name=RESOURCE_GROUP, | |
| account_name=ACCOUNT_NAME, | |
| deployment_name=DEPLOYMENT_NAME, | |
| resource={ | |
| "sku": {"name": "GlobalManagedCompute", "capacity": 1}, | |
| "properties": { | |
| "model": "azureml://registries/azure-huggingface/models/qwen--qwen3.8-27b/versions/2", | |
| "deploymentTemplate": "azureml://registries/azure-huggingface/deploymenttemplates/qwen--qwen3-8-27b--256k-nvidia-h100/labels/latest", | |
| "acceleratorType": "H100_80GB", | |
| "versionUpgradeOption": "OnceNewDefaultVersionAvailable", | |
| }, | |
| }, | |
| ).result() | |
| print(deployment.properties.provisioning_state) | |
| ``` | |
| The call returns when provisioning finishes, and prints `Succeeded`. | |
| > [!TIP] | |
| > The deployment usually takes 15–20 minutes, but might take up to an hour whilst looking for instances when demand is high. | |
| ## Call the deployment | |
| Copy the endpoint and key from the deployment page, and use the **deployment name** as the `model`: | |
| ```python | |
| from openai import OpenAI | |
| client = OpenAI( | |
| base_url="https://<ACCOUNT_NAME>.services.ai.azure.com/openai/v1", | |
| api_key="<API_KEY>", | |
| ) | |
| response = client.chat.completions.create( | |
| model="<DEPLOYMENT_NAME>", | |
| messages=[{"role": "user", "content": "Write a one-sentence hello from Microsoft Foundry."}], | |
| ) | |
| print(response.choices[0].message.content) | |
| ``` | |
| Some models are served through the Responses API instead, at `/openai/v1/responses`. Check the model page for the API that it supports. | |
| ## Clean up | |
| Delete the deployment to release its accelerators and stop billing, either from its page in Foundry or with: | |
| ```python | |
| client.managed_compute_deployments.begin_delete( | |
| resource_group_name=RESOURCE_GROUP, | |
| account_name=ACCOUNT_NAME, | |
| deployment_name=DEPLOYMENT_NAME, | |
| ).result() | |
| ``` | |
| ## Next steps | |
| - Browse the [catalog](../catalog/index) to pick a model. | |
| - Turn reasoning on or off per request in [Enable and disable reasoning](../advanced/reasoning). | |
| - Connect the deployment to a harness such as [Codex](../harness/codex) or [Hermes Agent](../harness/hermes-agent). | |
Xet Storage Details
- Size:
- 3.88 kB
- Xet hash:
- 545baac167794dc9e3f392bf8910aef2152124c6c67d576a1b573ca16084ff38
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.