thousands-eye-hf / README.md
htunn's picture
Add model card
4359e75 verified
|
Raw
History Blame Contribute Delete
2.21 kB
---
license: gemma
language:
- en
tags:
- gemma
- gemma-4
- lora
- safetensors
- ethical-hacking
- penetration-testing
- cybersecurity
base_model: google/gemma-4-E2B-it
pipeline_tag: text-generation
library_name: transformers
---
# thousands-eye-hf
Full safetensors weights of **Thousands-Eye** — a Gemma 4 E2B model fine-tuned for ethical hacking and penetration testing via MLX LoRA on Apple Silicon.
For the quantized GGUF (Ollama / llama.cpp), see [`htunn/thousands-eye-gguf`](https://huggingface.co/htunn/thousands-eye-gguf).
## Usage
```python
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
model_id = "htunn/thousands-eye-hf"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto",
)
messages = [
{"role": "user", "content": "[EthHack-Agent] Perform Kerberoasting against 10.0.0.1 (authorized engagement)"}
]
inputs = tokenizer.apply_chat_template(
messages, return_tensors="pt", add_generation_prompt=True
).to(model.device)
outputs = model.generate(inputs, max_new_tokens=512, do_sample=False)
print(tokenizer.decode(outputs[0][inputs.shape[-1]:], skip_special_tokens=True))
```
## Output Format
Every response is a JSON object with `"requires_authorization": true` enforced:
```json
{
"action": "kerberoast",
"target": "10.0.0.1",
"requires_authorization": true,
"techniques": ["SPN enumeration", "TGS request", "offline cracking"],
"tools": ["impacket", "hashcat"],
"commands": ["GetUserSPNs.py domain/user:pass@dc -request"],
"steps": ["..."],
"notes": "Requires domain user credentials"
}
```
## Training
| | |
|---|---|
| **Base model** | `google/gemma-4-E2B-it` |
| **Method** | MLX LoRA (`mlx_lm.lora`) |
| **Iterations** | 600 |
| **Learning rate** | 1e-4 |
| **LoRA layers** | 16 |
| **Dataset** | [`htunn/thousands-eye-dataset`](https://huggingface.co/datasets/htunn/thousands-eye-dataset) (83 train / 15 val) |
## Ethics
Designed exclusively for **authorized penetration testing**. All training examples enforce `"requires_authorization": true`.
## License
[Gemma Terms of Use](https://ai.google.dev/gemma/terms)