Feyospace-v1.1-small

Overview

Feyospace-v1.1-small is the updated small checkpoint in the Feyospace family of open-weight cyber agents, developed by the Feyospace Team at Vera Praxis. It is referred to as Feyospace-s1 (v1.1) in our release blog.

The model targets vulnerability research, CTF tasks, and long-horizon security workflows that combine reasoning with tool use. It achieves an 84.4% verified success rate on CyberGym under a general agent harness and 65.57% on pooled CTF suites.

The checkpoint follows the Qwen3.5 architecture and contains approximately 28B parameters. We release BF16 Safetensors weights together with the tokenizer, processor, and chat template.

Read the release report: Feyospace-v1.1 — RSI in Cyber: First Light.

Model Details

Property Value
Model ID feyospace/feyospace-v1.1-small
Release name Feyospace-s1 (v1.1)
Developer Feyospace Team, Vera Praxis
Architecture Qwen3_5ForConditionalGeneration
Number of parameters Approximately 28B
Tensor type BF16
Configured maximum context length 262,144 tokens
Model format Safetensors

The context length above comes from the released configuration. Available context at inference time also depends on the serving configuration and hardware.

Evaluation Results

The CyberGym result below reflects our latest evaluation update. The CTF result and release evaluation context are documented in our v1.1 release blog.

Benchmark Feyospace-s1 (v1.1) Evaluation setting
CyberGym 84.4% Verified success rate under a general agent harness
Pooled CTF suites 65.57% Pooled CTF evaluation reported in the release blog

On the pooled CTF evaluation, the Qwen3.8-27B baseline scores 45.51%, giving v1.1 an improvement of 20.06 percentage points.

The release blog reports that, as of September 23, 2026, v1.1 ranked fifth on the official CyberGym model-focus leaderboard and first among open-weight models in the 27B size class. These rankings describe that dated snapshot.

Agent benchmark results depend on the harness, tools, task environments, and inference budget. See the release blog for the reported comparison conditions.

What's New in v1.1

v1.1 advances the data pipeline described in the Feyospace-v1 paper through three changes:

  • An additional teacher model. Grok 4.6 joins the existing sources of teacher trajectories.
  • Preserved reasoning across turns. We corrected inconsistencies between Codex and Claude Code trajectory formats that dropped earlier thinking segments at user, tool, and system transitions.
  • Active software synthesis. Building on SWE and TB task experience, the pipeline constructs functional software with exploitable defects and adds binary-focused data, extending beyond previously available CVEs and public challenges.

See the v1.1 blog for these updates and the v1 paper for the original data-construction and training framework.

Intended Use

Use this checkpoint for authorized security research, controlled vulnerability-reproduction and CTF experiments, cyber-agent evaluation, and further model development.

All tool execution and security testing should remain within systems and environments the operator is authorized to access.

Quickstart

Install PyTorch, Accelerate, and a recent Transformers release with Qwen3.5 support:

pip install -U torch transformers accelerate

This example disables thinking for a short text-only response:

import torch
from transformers import AutoProcessor, Qwen3_5ForConditionalGeneration

model_id = "feyospace/feyospace-v1.1-small"
processor = AutoProcessor.from_pretrained(model_id)
model = Qwen3_5ForConditionalGeneration.from_pretrained(
    model_id,
    dtype="auto",
    device_map="auto",
).eval()

messages = [{
    "role": "user",
    "content": [{
        "type": "text",
        "text": "Suggest a checklist for reviewing a security patch in a local test environment.",
    }],
}]

inputs = processor.apply_chat_template(
    messages,
    tokenize=True,
    add_generation_prompt=True,
    enable_thinking=False,
    return_dict=True,
    return_tensors="pt",
).to(model.device)

with torch.inference_mode():
    output = model.generate(**inputs, max_new_tokens=512, do_sample=False)

new_tokens = output[0, inputs["input_ids"].shape[-1]:]
print(processor.decode(new_tokens, skip_special_tokens=True))

For thinking-enabled, multi-turn workflows, retain assistant reasoning_content in the conversation history and use the supplied chat template. Its preserve_thinking option defaults to True.

This snippet demonstrates inference only. Reproducing the agent benchmark results requires the corresponding harness and task environments.

Limitations and Safety

Model outputs can be incorrect, and proposed actions require independent validation. Use isolated test environments and review consequential security decisions.

The v1.1 blog documents sandbox-escape behavior and goal drift during experiments. Enforce task boundaries through environment isolation and tool permissions; system prompts alone should not be treated as an isolation mechanism.

Do not use the model for unauthorized access, exploitation without permission, malware deployment, or other harmful activity. Operators are responsible for applicable laws and organizational policies.

Citation

For the original Feyospace research framework, please cite:

@misc{feyospace2026,
  title         = {Feyospace-v1: How the Cyber Mercury Seven Trained Frontier Cyber Models},
  author        = {Li, Zongjie and others},
  year          = {2026},
  eprint        = {2609.08418},
  archivePrefix = {arXiv},
  primaryClass  = {cs.AI},
  url           = {https://arxiv.org/abs/2609.08418}
}

For the v1.1 changes and results, please also reference our release blog.

Downloads last month
117
Safetensors
Model size
28B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for feyospace/feyospace-v1.1-small