{ "cells": [ { "cell_type": "markdown", "id": "e426cd04-c053-43e8-b505-63cee7956a53", "metadata": {}, "source": [ "# Welcome to a very busy Week 8 folder\n", "\n", "## We have lots to do this week!\n", "\n", "We'll move at a faster pace than usual, particularly as you're becoming proficient LLM engineers.\n", "\n", "# The Price is Right\n", "\n", "## Week 8 Order of Play\n", "\n", "Day 1: Modal.com and SpecialistAgent \n", "Day 2: RAG, FrontierAgent, Ensemble Agent \n", "Day 3: ScannerAgent, MessengerAgent \n", "Day 4: AutonomousPlannerAgent and DealAgentFramework \n", "Day 5: The Price Is Right Finale\n", "\n" ] }, { "cell_type": "markdown", "id": "b3cf5389-93c5-4523-bc48-78fabb91d8f6", "metadata": {}, "source": [ "\n", " \n", " \n", " \n", " \n", "
\n", " \n", " \n", "

Especially important this week: pull the latest

\n", " I'm continually improving these labs, adding more examples and exercises.\n", " At the start of each week, it's worth checking you have the latest code.
\n", " Please see Guide 3 in the Guides folder if you're not sure how to do a git pull\n", "
\n", "
" ] }, { "cell_type": "code", "execution_count": null, "id": "bc0e1c1c-be6a-4395-bbbd-eeafc9330d7e", "metadata": {}, "outputs": [], "source": [ "import os\n", "import locale\n", "import modal\n", "from agents.preprocessor import Preprocessor\n", "from dotenv import load_dotenv\n", "load_dotenv(override=True)" ] }, { "cell_type": "code", "execution_count": null, "id": "765bb54b", "metadata": {}, "outputs": [], "source": [ "# To check that your computer can output special characters, make sure this outputs UTF-8 \n", "\n", "print(locale.getpreferredencoding()) # Should print 'UTF-8'" ] }, { "cell_type": "code", "execution_count": null, "id": "0b66f056", "metadata": {}, "outputs": [], "source": [ "os.environ[\"PYTHONIOENCODING\"] = \"utf-8\"" ] }, { "cell_type": "markdown", "id": "ab5c8533-9f66-448f-b9b2-133d1ff50639", "metadata": {}, "source": [ "# Setting up the modal tokens\n", "\n", "## IMPORTANT - please do read and follow these instructions!\n", "\n", "First please visit: https://modal.com\n", "\n", "And sign up for an account. Then click the Avatar menu on the top right and select \"Settings\"\n", "\n", "Then click \"API Tokens\" in the left sidebar, then click \"New Token\".\n", "\n", "You will be given something like this to run:\n", "\n", "`modal token set --token-id ak-somethinghere --token-secret as-somethinghere`\n", "\n", "But because we're using uv, the real thing to run is this:\n", "\n", "`uv run modal token set --token-id ak-somethinghere --token-secret as-somethinghere`\n", "\n", "### Troubleshooting\n", "\n", "If you have problems, 3 things to try:\n", "\n", "1. Try running `uv run modal token new` before the `uv run modal token set..` \n", "\n", "2. Suggestion from student David S. on Windows:\n", "\n", "> In case anyone else using Windows hits this problem: Along with having to run `modal token new` from a command prompt, you have to move the generated token file. It will deploy the token file (.modal.toml) to your Windows profile folder. The virtual environment couldn't see that location (strangely, it couldn't even after I set environment variables for it and rebooted). I moved that token file to the folder I'm operating out of for the lab and it stopped throwing auth errors.\n", "\n", "3. Doing in the manual way:\n", "\n", "It might be totally fine to simply add the 2 keys directly to your .env file:\n", "\n", "```\n", "MODAL_TOKEN_ID=ak-...\n", "MODAL_TOKEN_SECRET=as-...\n", "```\n", "\n", "Then rerun `load_dotenv(override=True)` to load these environment variables." ] }, { "cell_type": "code", "execution_count": null, "id": "3b133701-f550-44a1-a67f-eb7ccc4769a9", "metadata": {}, "outputs": [], "source": [ "from hello import app, hello, hello_europe" ] }, { "cell_type": "code", "execution_count": null, "id": "0f3f73ae-1295-49f3-9099-b8b41fc3429b", "metadata": {}, "outputs": [], "source": [ "with app.run():\n", " reply=hello.local()\n", "reply" ] }, { "cell_type": "code", "execution_count": null, "id": "c1d8c6f9-edc7-4e52-9b3a-c07d7cff1ac7", "metadata": {}, "outputs": [], "source": [ "with app.run():\n", " reply=hello.remote()\n", "reply" ] }, { "cell_type": "markdown", "id": "a1c075e9-49c7-4ebd-812f-83196d32de32", "metadata": {}, "source": [ "## Added thanks to student Tue H.\n", "\n", "If you look in hello.py, I've added a simple function hello_europe\n", "\n", "That uses the decorator: \n", "`@app.function(image=image, region=\"eu\")`\n", "\n", "See the result below! More region specific settings are [here](https://modal.com/docs/guide/region-selection)\n", "\n", "Note that it does consume marginally more credits to specify a region." ] }, { "cell_type": "code", "execution_count": null, "id": "b027da1a-c79d-42cb-810d-32ddca31aa02", "metadata": {}, "outputs": [], "source": [ "with app.run():\n", " reply=hello_europe.remote()\n", "reply" ] }, { "cell_type": "markdown", "id": "22e8d804-c027-45fb-8fef-06e7bba6295a", "metadata": {}, "source": [ "# Before we move on -\n", "\n", "## We need to set your HuggingFace Token as a secret in Modal\n", "\n", "## Super important - please read - this confuses a lot of people!\n", "\n", "Secrets in Modal are given a **name** that describes the secret. \n", "Then the secret itself has a KEY and a VALUE. \n", "We will be setting up a secret with: \n", "\n", "Name: huggingface-secret \n", "Key: HF_TOKEN \n", "Value: hf_... \n", "\n", "## The bulletproof recipe:\n", "\n", "1. Go to modal.com, sign in and go to your dashboard \n", "2. Click on Secrets in the nav bar \n", "3. Create new secret, click on Hugging Face, this new secret needs to be called **huggingface-secret** because that's how we refer to it in the code \n", "4. Fill in your key as HF_TOKEN and the value as your actual token hf_... \n", "5. Click done\n", "\n", "### And now back to business: time to work with Llama" ] }, { "cell_type": "code", "execution_count": null, "id": "cb8b6c41-8259-4329-b1c4-a1f67d26d1be", "metadata": {}, "outputs": [], "source": [ "# This import may give a deprecation warning about adding local Python modules to the Image\n", "# That warning can be safely ignored. You may get the same warning in other places, too..\n", "\n", "from llama import app, generate" ] }, { "cell_type": "code", "execution_count": null, "id": "db4a718a-d95d-4f61-9688-c9df21d88fe6", "metadata": {}, "outputs": [], "source": [ "with modal.enable_output():\n", " with app.run():\n", " result=generate.remote(\"Never gonna give you up, never gonna\")\n", "result\n", "\n", "# Or it you object to being rickrolled, try this: \"Hey Jude, don't make it\"" ] }, { "cell_type": "code", "execution_count": null, "id": "9a9a6844-29ec-4264-8e72-362d976b3968", "metadata": {}, "outputs": [], "source": [ "from pricer_ephemeral import app, price" ] }, { "cell_type": "code", "execution_count": null, "id": "50e6cf99-8959-4ae3-ba02-e325cb7fff94", "metadata": {}, "outputs": [], "source": [ "with modal.enable_output():\n", " with app.run():\n", " result=price.remote(\"Quadcast HyperX condenser mic, connects via usb-c to your computer for crystal clear audio\")\n", "result" ] }, { "cell_type": "code", "execution_count": null, "id": "20ee8e04", "metadata": {}, "outputs": [], "source": [ "preprocessor = Preprocessor()\n", "text = preprocessor.preprocess(\"Quadcast HyperX condenser mic, connects via usb-c to your computer for crystal clear audio\")\n", "print(text)" ] }, { "cell_type": "code", "execution_count": null, "id": "0d9669ba", "metadata": {}, "outputs": [], "source": [ "preprocessor = Preprocessor(model_name=\"groq/openai/gpt-oss-20b\")\n", "text = preprocessor.preprocess(\"Quadcast HyperX condenser mic, connects via usb-c to your computer for crystal clear audio\")\n", "print(text)" ] }, { "cell_type": "markdown", "id": "2df6f295", "metadata": {}, "source": [ "### Add this to your .env if you want the Preprocessor to use a different model by default:\n", "\n", "`PRICER_PREPROCESSOR_MODEL=groq/openai/gpt-oss-20b`" ] }, { "cell_type": "code", "execution_count": null, "id": "eb93e74c", "metadata": {}, "outputs": [], "source": [ "with modal.enable_output():\n", " with app.run():\n", " result = price.remote(text)\n", "print(result)" ] }, { "cell_type": "code", "execution_count": null, "id": "ca042a54", "metadata": {}, "outputs": [], "source": [ "print(text)" ] }, { "cell_type": "markdown", "id": "04d8747f-8452-4077-8af6-27e03888508a", "metadata": {}, "source": [ "## Transitioning From Ephemeral Apps to Deployed Apps\n", "\n", "From a command line, `uv run modal deploy xxx` will deploy your code as a Deployed App\n", "\n", "This is how you could package your AI service behind an API to be used in a Production System.\n", "\n", "You can also build REST endpoints easily, although we won't cover that as we'll be calling direct from Python.\n", "\n", "## Important note about secrets\n", "\n", "In both the files `pricer_service.py` and `pricer_service2.py` you will find code like this near the top: \n", "`secrets = [modal.Secret.from_name(\"hf-secret\")]` \n", "You may need to change from `hf-secret` to `huggingface-secret` depending on how the Secret is configured in modal. \n", "To check, visit this page and look in the first column: \n", "https://modal.com/secrets/\n", "\n", "## Important note for Windows people:\n", "\n", "On the next line, I call `uv run modal deploy` from within Jupyter lab; I've heard that on some versions of Windows this gives a strange unicode error because modal prints emojis to the output which can't be displayed. If that happens to you, open a Terminal and run `uv run modal deploy..`" ] }, { "cell_type": "code", "execution_count": null, "id": "7f90d857-2f12-4521-bb90-28efd917f7d1", "metadata": {}, "outputs": [], "source": [ "# You can also run \"uv run modal deploy -m pricer_service\" in the Terminal\n", "\n", "!uv run modal deploy -m pricer_service" ] }, { "cell_type": "code", "execution_count": null, "id": "1dec70ff-1986-4405-8624-9bbbe0ce1f4a", "metadata": {}, "outputs": [], "source": [ "pricer = modal.Function.from_name(\"pricer-service\", \"price\")" ] }, { "cell_type": "markdown", "id": "5f443cc5", "metadata": {}, "source": [ "Watch it happening:\n", "\n", "https://modal.com" ] }, { "cell_type": "code", "execution_count": null, "id": "17776139-0d9e-4ad0-bcd0-82d3a92ca61f", "metadata": {}, "outputs": [], "source": [ "# This can take a while! We'll use faster approaches shortly\n", "\n", "pricer.remote(text)" ] }, { "cell_type": "code", "execution_count": null, "id": "f56d1e55-2a03-4ce2-bb47-2ab6b9175a02", "metadata": {}, "outputs": [], "source": [ "# You can also run \"modal deploy -m pricer_service2\" at the command line in an activated environment\n", "\n", "!modal deploy -m pricer_service2" ] }, { "cell_type": "code", "execution_count": null, "id": "9e19daeb-1281-484b-9d2f-95cc6fed2622", "metadata": {}, "outputs": [], "source": [ "Pricer = modal.Cls.from_name(\"pricer-service\", \"Pricer\")\n", "pricer = Pricer()\n", "reply = pricer.price.remote(text)\n", "print(reply)" ] }, { "cell_type": "code", "execution_count": null, "id": "5fd499ee", "metadata": {}, "outputs": [], "source": [ "reply = pricer.price.remote(text)\n", "print(reply)" ] }, { "cell_type": "markdown", "id": "9c1b1451-6249-4462-bf2d-5937c059926c", "metadata": {}, "source": [ "# Optional: Keeping Modal warm\n", "\n", "## A way to improve the speed of the Modal pricer service\n", "\n", "The first time you run this modal class, it might take as much as 10 minutes to build. \n", "Subsequently it should be much faster.. 30 seconds if it needs to wake up, otherwise 2 seconds. \n", "If you want it to always be 2 seconds, you can keep the container from going to sleep by editing this constant in pricer_service2.py:\n", "\n", "`MIN_CONTAINERS = 0`\n", "\n", "\n", "\n", "Make it 1 to keep a container alive. \n", "But please note: this will eat up credits! Only do this if you are comfortable to have a process running continually.\n", "\n", "Alternatively, you can run this code and it will stay warm for 20 mins rather than 2 mins.\n", "\n", "### Code to keep warm for 20 mins before cooling down:\n", "\n", "```python\n", "import modal\n", "Pricer = modal.Cls.from_name(\"pricer-service\", \"Pricer\")\n", "pricer = Pricer()\n", "pricer.update_autoscaler(scaledown_window=1200)\n", "```\n", "\n", "### Code to revert to keeping warm for only 2 mins\n", "\n", "```python\n", "import modal\n", "Pricer = modal.Cls.from_name(\"pricer-service\", \"Pricer\")\n", "pricer = Pricer()\n", "pricer.update_autoscaler(scaledown_window=120)\n", "```" ] }, { "cell_type": "markdown", "id": "3754cfdd-ae28-47c8-91f2-6e060e2c91b3", "metadata": {}, "source": [ "## And now introducing our Agent class\n", "\n", "By default this will preprocess using Llama3.2\n", "\n", "If you'd prefer to use Groq, then add this env variable like:\n", "\n", "```\n", "PRICER_PREPROCESSOR_MODEL=groq/openai/gpt-oss-20b\n", "```" ] }, { "cell_type": "code", "execution_count": null, "id": "d0cdc4c7", "metadata": {}, "outputs": [], "source": [ "import logging\n", "root = logging.getLogger()\n", "root.setLevel(logging.INFO)" ] }, { "cell_type": "code", "execution_count": null, "id": "ba9aedca-6a7b-4d30-9f64-59d76f76fb6d", "metadata": {}, "outputs": [], "source": [ "from agents.specialist_agent import SpecialistAgent" ] }, { "cell_type": "code", "execution_count": null, "id": "fe5843e5-e958-4a65-8326-8f5b4686de7f", "metadata": {}, "outputs": [], "source": [ "agent = SpecialistAgent()\n" ] }, { "cell_type": "code", "execution_count": null, "id": "e70cd0ed", "metadata": {}, "outputs": [], "source": [ "agent.price(\"iPhone 10\")" ] }, { "cell_type": "code", "execution_count": null, "id": "430869a0", "metadata": {}, "outputs": [], "source": [] } ], "metadata": { "kernelspec": { "display_name": ".venv", "language": "python", "name": "python3" }, "language_info": { "codemirror_mode": { "name": "ipython", "version": 3 }, "file_extension": ".py", "mimetype": "text/x-python", "name": "python", "nbconvert_exporter": "python", "pygments_lexer": "ipython3", "version": "3.12.12" } }, "nbformat": 4, "nbformat_minor": 5 }