Qwen3-1.7B-Swarm-Arena-RL-v4-step8-development

Four distinct LoRA policies trained over one frozen Qwen3 1.7B backbone in the Swarm Arena 4v4 partially observed graph-control simulator. Policy directories policy_blue_0 through policy_blue_3 retain separate optimizer identities and must be assigned to their corresponding BLUE roles.

Selected trainer step: 8. Release status: not-admitted. Eight-update development artifact; online monitoring is promising but capability is mixed and information-specific communication is not established. Selection and frozen final remain unopened.

Exact provenance, policy hashes, public input revisions, and compact evaluation reports are in PROVENANCE.json, SHA256SUMS, and results/.

The reward is the zero-sum terminal control-margin delta. There is no speaking, silence, capture, or learned-judge bonus. Higher return is evidence of task learning; a communication claim additionally requires normal messages to beat dropped, shuffled, and delayed-message interventions on held-out cases.

This is research software for a discrete simulator. It is not evidence of broad swarm intelligence or real-world cybersecurity capability.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for CK0607/Qwen3-1.7B-Swarm-Arena-RL-v4-step8-development

Finetuned
Qwen/Qwen3-1.7B
Adapter
(610)
this model