sae-gpu-backend / README.md
justforkicks's picture
Upload 3 files
717dc83 verified
|
Raw
History Blame Contribute Delete
3.18 kB

A newer version of the Gradio SDK is available: 6.22.0

Upgrade
metadata
title: GPU SAE Service
emoji: 🧠
colorFrom: purple
colorTo: blue
sdk: gradio
sdk_version: 4.44.0
app_file: app.py
pinned: false
license: mit
hardware: t4-small

GPU-Accelerated SAE Analysis Service

This HuggingFace Space provides GPU-accelerated Sparse Autoencoder (SAE) analysis for larger language models that are too slow to run on CPU.

Supported Models

Model Parameters Layers SAE Features Status
Gemma 2 9B 9B 42 16k βœ… Ready
Mistral 7B 7B 32 65k ⚠️ SAE pending
Pythia 1.4B 1.4B 24 32k ⚠️ SAE pending
Llama 2 7B 7B 32 - ❌ No SAE

Usage

As a Remote Backend

Set the EXTERNAL_SAE_URL environment variable in your Replit app to point to this Space:

EXTERNAL_SAE_URL=https://your-username-gpu-sae-service.hf.space

The Replit app will automatically route requests for larger models to this GPU backend.

API Endpoints

  • POST /api/analyze - Run SAE feature analysis
  • POST /api/discover - Circuit discovery (input/output features)
  • POST /api/ablate - Ablation testing (zero features, measure impact)
  • GET /api/models - List available models
  • GET /api/health - Check service status

Example Request

import requests

response = requests.post(
    "https://your-space.hf.space/api/analyze",
    json={
        "prompt": "The capital of France is",
        "model_id": "gemma-2-9b",
        "layer": 20,
        "top_k": 32
    }
)
print(response.json())

Hardware Requirements

  • T4 (16GB VRAM): Runs 7B models with float16
  • A10G (24GB VRAM): Runs 9B-13B models comfortably
  • A100 (40/80GB VRAM): Required for 70B+ models

Integration Architecture

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  Replit App (Frontend + API)            β”‚
β”‚  - UI for interpretability tools        β”‚
β”‚  - Routes to local or remote backend    β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                  β”‚
    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
    β–Ό                           β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Local Service β”‚    β”‚ This HF Space (GPU) β”‚
β”‚ GPT-2, Pythia β”‚    β”‚ Gemma 9B, Mistral   β”‚
β”‚ Fast, Free    β”‚    β”‚ GPU-accelerated     β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Deploying Your Own

  1. Fork this Space
  2. Select GPU hardware (T4-small recommended for cost)
  3. Add HuggingFace token if using gated models (Llama, Gemma)
  4. Copy the Space URL to your Replit app's EXTERNAL_SAE_URL

Cost Estimates

  • T4 small: ~$0.60/hour (pauses when idle)
  • A10G small: ~$1.00/hour
  • A100 large: ~$4.00/hour

Spaces with GPU auto-sleep after inactivity, so you only pay for active usage.