openenv-cloudsoc / MODEL_RECOMMENDATIONS.md
OpenEnv Contributor
Initial commit: OpenEnv-CloudSOC benchmark environment
115612d
|
Raw
History Blame Contribute Delete
6.95 kB

Model Selection Guide for CloudSOC Testing

Quick Recommendation

For testing: Start with gpt-3.5-turbo or local Llama 2 via Ollama

Model Cost Speed Quality Setup
gpt-3.5-turbo ⭐ $0.50/1M tokens Fast Good 1 min
gpt-4o-mini $0.15/1M tokens Very Fast Better 1 min
Llama 2 (Local) Free Slow Fair 10 min
Mistral (Local) Free Medium Good 10 min
gpt-4-turbo $10/1M tokens Medium Excellent 1 min

Option 1: Cloud Models (Easiest) ⭐ RECOMMENDED

Setup (1 minute)

  1. Get API key from platform.openai.com

  2. Set environment variables:

# Windows PowerShell
$env:HF_TOKEN = "sk-..."  # Your OpenAI API key
$env:API_BASE_URL = "https://api.openai.com/v1"
$env:MODEL_NAME = "gpt-3.5-turbo"  # or gpt-4o-mini
  1. Run test:
python inference.py --task easy --seed 42 --max_steps 10

Cost Estimates (Easy Task, ~10-15 steps)

  • gpt-3.5-turbo: $0.01-0.05 per run
  • gpt-4o-mini: $0.003-0.01 per run
  • gpt-4-turbo: $0.20-0.50 per run

Why gpt-3.5-turbo?

βœ… Cheap ($0.50 per 1M input tokens)
βœ… Fast (1-2 sec per action)
βœ… Good at JSON parsing
βœ… Reliable
❌ Sometimes verbose/redundant

Why gpt-4o-mini?

βœ… Cheaper than 3.5 ($0.15 per 1M)
βœ… Better reasoning
βœ… Better JSON quality
βœ… Faster overall
βœ… BEST for testing


Option 2: Local Models via Ollama (Free but Slow)

Setup (10 minutes)

  1. Install Ollama from ollama.ai

  2. Pull a model:

ollama pull mistral      # Fast, ~7B params (recommended)
ollama pull llama2       # Slower, ~7B params
ollama pull neural-chat  # Medium, ~7B params
  1. Start Ollama server:
ollama serve  # Runs on http://localhost:11434
  1. Set environment variables:
$env:HF_TOKEN = "dummy-token"  # Can be anything
$env:API_BASE_URL = "http://localhost:11434/v1"
$env:MODEL_NAME = "mistral"  # or llama2
  1. Run test:
python inference.py --task easy --seed 42 --max_steps 5

Pros & Cons

βœ… FREE (no API costs)
βœ… Private (no data to OpenAI)
βœ… 100% offline
❌ SLOW (10-30 sec per action)
❌ Lower reasoning quality
❌ RAM intensive (~8GB for 7B model)

Local Model Recommendations

Model Size Speed Quality RAM
Mistral 7B ⭐⭐⭐ ⭐⭐ 8GB
Llama 2 7B ⭐⭐ ⭐⭐ 8GB
Neural Chat 7B ⭐⭐⭐ ⭐⭐ 8GB
Phi 2 2.7B ⭐⭐⭐⭐ ⭐ 4GB

Option 3: Hugging Face Inference API

Setup (2 minutes)

  1. Get Hugging Face token from huggingface.co/settings/tokens

  2. Set environment:

$env:HF_TOKEN = "hf_XXX..."
$env:API_BASE_URL = "https://api-inference.huggingface.co/v1"
$env:MODEL_NAME = "meta-llama/Llama-2-7b-chat-hf"  # or another model
  1. Run test:
python inference.py --task easy --seed 42

Available Models

  • meta-llama/Llama-2-7b-chat-hf
  • mistralai/Mistral-7B-Instruct-v0.1
  • NousResearch/Nous-Hermes-2-Mixtral-8x7B-DPO

Pros & Cons

βœ… Free tier available
βœ… No GPU needed
βœ… Wide model selection
❌ Rate limited
❌ Variable speed (depends on queue)
❌ Less reliable than OpenAI


Quick Start Decision Tree

Do you have an OpenAI API key?
β”œβ”€ YES β†’ Use gpt-4o-mini (cheapest, fastest, best)
└─ NO
   β”œβ”€ Do you want to pay for inference?
   β”‚  β”œβ”€ YES β†’ Get OpenAI key, use gpt-4o-mini
   β”‚  └─ NO
   β”‚     └─ Do you have 8GB+ RAM free?
   β”‚        β”œβ”€ YES β†’ Use Ollama + Mistral (free, offline)
   β”‚        └─ NO β†’ Use Hugging Face Inference API (free tier)

Testing Strategy (Recommended)

Phase 1: Validation (Free/Cheap)

# Test with gpt-4o-mini (cheapest, $0.003-0.01 per run)
python test_cloudsoc.py --quick                    # Validates env βœ“
python inference.py --task easy --seed 42 --max_steps 5   # Quick test

Cost: ~$0.01
Time: 2 min

Phase 2: Medium Testing ($0.05-0.20)

# Run all 3 difficulty levels
for task in easy medium hard:
  python inference.py --task $task --seed 42 --max_steps 50

Cost: ~$0.15
Time: 30 min

Phase 3: Production (Scale Up)

# Deploy to Hugging Face Spaces with gpt-4o or gpt-4-turbo
# (if budget allows for better quality)

Testing Commands by Model

gpt-4o-mini (Cloud - RECOMMENDED)

$env:HF_TOKEN = "sk-..."
$env:MODEL_NAME = "gpt-4o-mini"
$env:API_BASE_URL = "https://api.openai.com/v1"

python inference.py --task easy --seed 42 --max_steps 10

gpt-3.5-turbo (Cloud - Budget)

$env:HF_TOKEN = "sk-..."
$env:MODEL_NAME = "gpt-3.5-turbo"

python inference.py --task easy --seed 42

Mistral (Local - Free)

# Terminal 1: Start Ollama
ollama serve

# Terminal 2: Run test
$env:HF_TOKEN = "dummy"
$env:API_BASE_URL = "http://localhost:11434/v1"
$env:MODEL_NAME = "mistral"

python inference.py --task easy --seed 42 --max_steps 5

Expected Output (Each Model)

All should produce hackathon-compliant format:

[START] task=easy env=cloudsoc model=gpt-4o-mini
[STEP] step=1 action=aws.soc.get_alerts({}) reward=0.00 done=false error=null
[STEP] step=2 action=aws.cloudwatch.query_basic(...) reward=-0.01 done=false error=null
...
[END] success=true steps=N rewards=0.00,-0.01,...

My Recommendation

  1. First time / Testing? β†’ gpt-4o-mini (cheapest, best quality)
  2. Budget conscious? β†’ gpt-3.5-turbo (still cheap, good)
  3. Want free? β†’ Ollama + Mistral (free, but slow)
  4. Production quality? β†’ gpt-4-turbo (expensive, best reasoning)

Troubleshooting

"API key invalid"

β†’ Check your API key is correct
β†’ Make sure you're using the right env var name (HF_TOKEN, not OPENAI_API_KEY)

"Model not found"

β†’ Check model name spelling
β†’ Verify it's available on your provider

"Connection refused"

β†’ For local: Make sure ollama serve is running
β†’ For cloud: Check internet connection

"Rate limited"

β†’ Wait a few seconds and retry
β†’ Consider upgrading to paid tier
β†’ Use local model instead

Model gives bad JSON

β†’ Try --temperature 0.2 (more deterministic)
β†’ Switch to gpt-4o-mini (better JSON)
β†’ Use simpler task (easy instead of hard)


Cost Comparison (Easy Task, 15 steps)

Model Per Run 100 Runs Notes
gpt-4o-mini $0.01 $1.00 Best value ⭐
gpt-3.5-turbo $0.03 $3.00 Good value
Mistral (local) $0.00 $0.00 Free, slow
gpt-4-turbo $0.30 $30.00 Overkill for dev

TL;DR: Start with gpt-4o-mini ($0.01 per test), or free Ollama locally if you have 8GB RAM.