Lab 01: Inference Test & Prompting
Overview
This lab is designed to verify the cloud deployment and test basic prompting and text generation capabilities of the deployed model (e.g., Gemma-4-31B).
Setup
Ensure you have spun up the Modular Cloud MAX Engine or a RunPod inference endpoint (using scripts/01_setup_inference.sh and 02_start_server.sh).
Usage
Open 01_inference_test.ipynb in a Jupyter environment. The notebook contains examples of:
- Connecting to the OpenAI-compatible API endpoint.
- Constructing basic prompts.
- Testing generation speed and context limits.
Goals
- Verify API connectivity.
- Understand basic prompting structures.
- Measure inference latency (tokens per second).