Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
AbstractPhil 
posted an update 6 days ago
Post
61
My apologies for the incorrect format for the AMOE arms from the experimental branch. They have been saving as torch objects. They are now correctly saving as safetensors format. My apologies for the inconvenience this may cause for you use. I will be modifying the codespaces to use the correct safetensors formats.

After the first 20.9b tokens trained, the real experiments begins. Beatrix V3's first prepped-state modular command structure has been attached for dynamic training. These arms will exist as appendages for Beatrix - trained alongside with the trunk until the end of the run.

These exist for experimental extraction, analysis, distillation experiments, memory experiments, mathematics experiments, and more. Each arm will be built along the chain for specific test cases. Expectation for each is already lined up and the outcomes are tested for, but the model still may face instability and must be monitored.

As her first arm learns tinystories, she builds direct composite semantic structure throughout this system. Think of it like, the first higher-functioning cognition attachment.

She's still very naïve and structurally unaware, so attaching new limbs is essentially extending a structure that is not yet finished forming. Nothing but fragments of issued information from an unknown source.

In this case, this structure has been tested hundreds of times to ensure she will not simply collapse during training by having this attached.

Seems we've run into engineering problems. The time degraded from 4 seconds per step to 9 seconds per step, compounding increase per independent arm attached.

I need to halt and run analysis. As it stands, we'll run until the end of Friday, then the engineering goes up on blocks.

I'm going to squeeze every bit of every byte from this engineering. Formulas up for scrutiny, system optimizations up for scrutiny, kernel optimizations, libraries, everything. Running until October 21st isn't reasonable. The models must be prepared correctly and reasonably to the mathematics, the results must be calculated correctly and the model trained to parity without faulty modifications during training that can impact the results.

image
Working through the problems using modern araxiv optimizations through compounded deep research. Everything used will be attributed and cited.

Currently the ETA if I fire up today would be October 20th, which is too far out. I'd like to narrow that down a considerable amount.

Verdict: Run will complete between October 14th and October 17th. A bit pushed out, but considerably less time than the alternative.

I've settled on a compromise. It will yield a substantially weaker abstentation policy, 1/4th of the cells. Without the necessary hardware involved I must weaken the model, but the results themselves will not be fruitless.

After the run is complete, I'll run multiple tests to determine the BEST possible mix for version 4. The BEST possible compromise. The STRONGEST possible yield that builds the most effective information, without destroying the necessary yield and information.

As it stands, the model will yield. The results will be potent. With any luck, they will be more than potent enough for distillation, diffusion, classification, and more experiments.

With that compromise, I'll be attaching additional curve watchers for the risk. It's a minimal risk, but it's high enough to put some gauges on the model. The cost for this model isn't cheap.

The deeper variant is showing quite weak recall in comparison to the v1, 2.5s, and 2s softmax before softmax collapse.

There's a definite problem here that needs to be addressed.

Deep byte recall wasn't exactly filling the 4096 space for the 2s, while the 2048 space was predominantly filled by v1 give or take. The models never quite coalesced, which v3 was expected to fill the space for.

Instead of filling the space as expected, the model is showing weaker code recall later in the chain. I'm not going to pull the plug yet, but the results this far into training should be much cleaner than they currently are.


The 3s model has been showing improvement, so I'm going to be cautiously optimistic. I'm thinking this model will need substantially more training than the softmax version, which means something needs to change in the attention mechanism to converge more quickly.

The 60b tokens very well might not be enough, which isn't a good sign. This larger investment is definitely going to need to be cut down to size if I want reasonable prototypes.

The 2s variant showed improvement throughout, and the end result of the 2s variant wasn't as strong as the v1 mixed but it was quite strong. The 2s full softmax showed very strong curves before the collapse, which signals to me we need a late-stage mechanism control for softmax, rather than a full replacement for softmax. The mixed attention are likely the best course of action. Softmax + splat, potentially a new mechanism to replace splat to ensure the delta isn't wildly misappropriated or drifts so far that the model cannot recall what was just consumed.

In this post