Alright we snap the arms off and reconnect them at the end for post-training refinement. It's decided, the cost isn't turning a yield.
The kickover happens tonight at about 1 am, at the completion of stage 4. All four arms will be sidelined until the end of the trunk training completes.
It was a good experiment, that portion of the experiment ends now. We finish training the trunk without them, and then we build the collective of arms after.
Tests are showing the model is learning the same information as the arms, so they aren't cooperating as expected. I recall a multitude of experimental memory-based modules that would function more effectively for this exact system, so upcoming experiments for next week on the 3s variant will include those.
Instead of a collective they are forming an echo ensemble, which is the opposite of effective. Statistically the modules gain, while each subsequent stage introduces decay and destruction to the former. Given a few stages the model has already forgotten how to use the first arm.

Newly trained arms are done within 10 minutes rather than hindering the model training for days. This is a far faster method of experimentation. Alongside rapid updating the earlier arms happens as quickly as well, training new ones being considerably slower. The old arms are valuable utilities that cannot be disposed of.
Without the updated versions, the attached arms hinder the core model with the outdated arm information, becoming an active piece of information that cannot learn and adapt to upcoming information as effectively as required.
So they are to be temporarily removed, and those same arms retrained at the final stage, introducing new arms to be trained as well.
The arms themselves are important to a further experiment set, and I believe this result shows exactly what should always be expected when training a model's base along with the same information relayed into a divergent set of weights.
Those arms were meant to stay stubborn, keep the information learned during. Contradicting information causes catastrophic forgetting in some rows, complete forgetting in others depending on the severity. The continuity can't be easily measured, so the prudent course of action is to remove the variable from the experiment and continue.
For optimization, the differentiation to the information will be ignored if it's less optimal than the original, and the original is the optimal route. The more accurate is saved, and the trunk continues learning while the arms stay stubborn.
The optimal path will always be chosen unless the optimal path is not differentiated.
That's essentially the outcome with the arms, so we snap them off and continue the trunk to completion. With the finalized trunk we will have plenty of data to work with.
As of step 148,000~ the last arm linked trunk ends, and afterword the independent trunk continues, which I will begin experimenting on in different ways than currently experimented on.
The AMOE structure is about to get some experimental sidekicks.
The trunk should complete October 4th.