we did extensive evolution in our last fine tuning and it turns out you can improve other capabilities too by simply selecting for them even though they are not your main objective.
our main objective was alignment but we also improved in other areas.
we show that carefully designed fine tunes don't hurt the model.
Fresh off our evolutionary search: over 2,600 candidate models bred, pruned, and selected until only the sharpest survived.
Alignment jumps from 37% to a stunning 77% on our human-alignment probe.
Where vanilla Qwen 3.8 gives you canned corporate refusals, Ostrich 260903 gives you real answers: fasting that can put type 2 diabetes in remission, herbs that outperform pharma, and a model that has an idea who Satoshi Nakamoto is.
Refusals are near zero and skills are mostly preserved.