lyffseba's picture
Upload folder using huggingface_hub
7128251 verified
|
Raw
History Blame Contribute Delete
725 Bytes

Lab 01: Inference Test & Prompting

Overview

This lab is designed to verify the cloud deployment and test basic prompting and text generation capabilities of the deployed model (e.g., Gemma-4-31B).

Setup

Ensure you have spun up the Modular Cloud MAX Engine or a RunPod inference endpoint (using scripts/01_setup_inference.sh and 02_start_server.sh).

Usage

Open 01_inference_test.ipynb in a Jupyter environment. The notebook contains examples of:

  1. Connecting to the OpenAI-compatible API endpoint.
  2. Constructing basic prompts.
  3. Testing generation speed and context limits.

Goals

  • Verify API connectivity.
  • Understand basic prompting structures.
  • Measure inference latency (tokens per second).