metadata
title: Gemma-4 E2B Uncensored API
emoji: π
colorFrom: blue
colorTo: blue
sdk: docker
app_port: 8000
pinned: false
tags:
- ml-intern
OpenAI-compatible API for HauhauCS/Gemma-4-E2B-Uncensored-HauhauCS-Aggressive
Model Details
| Spec | Value |
|---|---|
| Model | Gemma-4 E2B Uncensored (HauhauCS Aggressive) |
| Quantization | Q8_K_P |
| Context | 131072 tokens |
| Concurrent | 1 request |
| Reasoning | Enabled by default (--reasoning on --reasoning-format deepseek) |
Endpoints
POST /v1/chat/completionsβ Chat completions (streaming recommended)POST /v1/completionsβ Text completionsGET /v1/modelsβ List modelsGET /healthβ Health checkGET /api-infoβ JSON status
Usage
import openai
client = openai.OpenAI(
base_url="https://nanobotaiagent-gemma-4-e2b-uncensored-api.hf.space/v1",
api_key="no-key",
timeout=300.0,
)
response = client.chat.completions.create(
model="gemma",
messages=[{"role": "user", "content": "Hello!"}],
max_tokens=2048,
stream=True,
)
for chunk in response:
delta = chunk.choices[0].delta
rc = getattr(delta, 'reasoning_content', None) or (delta.model_extra or {}).get('reasoning_content')
if rc:
print(f"[think] {rc}", end="")
if delta.content:
print(delta.content, end="")
Reasoning
Reasoning is enabled by default. Gemma-4 uses <|channel>thought...<channel|> blocks
which llama.cpp extracts into the reasoning_content field (DeepSeek-compatible).
To disable per-request:
extra_body={"chat_template_kwargs": {"enable_thinking": False}}