AI & ML interests

None defined yet.

Recent Activity

ProCreationsΒ 
posted an update 1 day ago
view post
Post
1643
check out this cool webgpu thing i made! ProCreations/maple-webgpu
runs this model deepgrove/maple-preview with cool performance i found today on WebGPU locally on your computer with very fast speeds.
made with a mix of gpt 5.6 sol + opus 5 ultracode, its amazing what agents can do now, just a year ago this wouldve been impossible with them.
AbstractPhilΒ 
posted an update 2 days ago
view post
Post
2492
The AlephLM results are rolling in and I'm very excited for the possibilities. I am very much looking forward to the coming weeks as I train the first AlephLM distillations from MANY teachers into AMOE arms.

The AMOE arms hook cleanly to AlephLM structures and provide pos/neg learning elements. Hard positive and hard negatives coalesce to extend the capacity.

AbstractPhil/alephlm-0
AbstractPhil/alephlm-adopt-0

As it stands they are structurally sound enough to fully pretrain. As or more stable than a standard Bert experimentally to distill using InfoNCE. AMOE legs improve these structures substantially.

Structural behavior can be expanded in many ways on distilled and pretrained models alike. Attaching the AMOE to any model I've tried has created expanded or improved behavioral accumulations. They do have downsides but their upsides are very experimentally exciting.

I've distilled multiple vits, multiple berts, and have begun distilling berts into AlephLM structures successfully.

This is overall very exciting for me. I've begun formatting larger variants such as including GPT-2 and Qwen 3.5 4b as a paired combinator utilizing pathological T5 learned distilled encodings. It sounds odd, but the results show everything can be expanded and even be taught to cooperate.

The CaptionBert-8192-v2 and v2-b are both structurally collapsing after token 480 or so, which is expected due to the small train. By distilling an AMOE arm to V2 by training with a longformer expert, the results are cutting through like butter. V2 has begun stabilizing rapidly for considerably longer token chains and sequences, the structure is repairing and building reusable capacity.

I have discovered an improved methodology for sampling the AlephLM for text encoder benchmarks, which is predominantly L2 normalized outputs.

Upcoming large paper for the distillation experiments and results within the next week or two. It's going to be a big one.
  • 2 replies
Β·
ProCreationsΒ 
posted an update 3 days ago
view post
Post
1802
so i got 2nd on this competition ICML-2026-agent-repro/challenge (didnt actually get anything yet hopefully theres no catches)
when i get the 1000 dollars worth of gpu credits ill do a lot of cool things, including bigger and newer (qwen 3.8 27b) grugs ONLY IF you guys want (i have a lot of cool ideas for ai models.) stay tuned πŸ‘€!
  • 3 replies
Β·
AbstractPhilΒ 
posted an update 8 days ago
view post
Post
134
AbstractPhil/clip-vitb-mini-distilled
The semi-successful run series on the VIT-B lineup is live and full of useful baseline distillation information for feature + InfoNCE distillation processing as well as direct feature distillation processing, direct InfoNCE distillation, and multiple other tested methods. https://huggingface.co/blog/AbstractPhil/geometric-memory-ft4

This article showcases the baseline utilization and benchmarks of the earlier experiment line's objective and loss structures tested on 12m features for the vit-b baseline. Not the strongest showcase, but the strongest of the champions did show some serious promise.

Next setup will be a directly aligned set based on the loss and objectives decided by the champions in the first runs, for the second run they operate in direct conjunction with the bert-8192 and captionbert-8192 distillation format directly on clip-vit-l features - this time we're including DINOv3 into the mix for it's high potency.

I'm currently extracting 4 clip-vit-l variants for the CC12m features and will be running the next series on the L size, which will give considerably more active and useful features overall within a smaller package.

The captionbert-8192 has a more unique and difficult to tune for pixel processing parity, but I will spend a few days making sure the smaller prototypes fit before I run the large experiments in order to build towards the larger objectives.

Primarily I need to ensure the memory bank aligns correctly and the constellation conforms to the anchors correctly, as this process was not micro managed enough for this run. The results are nonetheless useful and potent.

The process continues until we cover the entire constellation series.
  • 6 replies
Β·
ProCreationsΒ 
posted an update 11 days ago
AbstractPhilΒ 
posted an update 11 days ago
view post
Post
192
The geometric memory article ft4 is live. https://huggingface.co/blog/AbstractPhil/geometric-memory-ft4

Direct pivot to distillation. I've accumulated enough experimental information to directly pivot my long term structure plan to distillation. This is to begin forming entire collectives of cooperative systems; differentiated expert distillation for generative behavior utilizing aleph addressed bottlenecks. With this I've also heavily begun experimenting with aleph competitions and cooperation using multiple pretrained frozen codebooks established from the SVAE system.

The idea here is simple in theory; use InfoNCE and address independent experts to build a manifest of unique gated experts utilizing a multitude of distilled systems from many other models. Such as SigLIP 16B + LAION CLIPB as a pair. The experimentation in the past showed this process is potent and with that merits additional experimentation using the newly established paradigms.

There are quite a bit of experiments to compare these to, so I have no shortage of comparators. After we train our baseline TinyViT with our gated system, we will know which experts are better at what and why they are better.

As a direct continuation from the earlier CLIP distillation experiments I'm directly comparing InfoNCE anchoring with multiple industry standard distillations from multiple papers. First comparison is InfoNCE anchoring in comparison to raw features using CoCo and CLIP_B, which seemed like a fair experiment to train a student with.

The upcoming series of experiments will provide the necessary information for how effective or ineffective this process is.

AbstractPhil/bulk-coco-features

The first experiments will be based on multiple clips from the bulk-coco-features extractions.

First we start with some clips, then some berts, then some smaller qwens, then some larger models, then some much much larger models. All meant to be compacted into selection mechanisms.
  • 4 replies
Β·
ProCreationsΒ 
posted an update 13 days ago
view post
Post
2071
what should i do for 200 followers... hmmmm... top reply to this post (most emojis) wins, just not anything too crazy
  • 11 replies
Β·
ProCreationsΒ 
posted an update 14 days ago
view post
Post
216
GRUGGG!
Today is the day of the grugs. Grug 27b has been released, Grug 35b a3b v2 has been released, Grug MTP versions have been released, etc.
And to add to that, I also hit 200 followers on huggingface! I’ll do something exciting with that soon.
Try the Grug family now at:
ProCreations/grug-27b
ProCreations/grug-35b-v2
  • 3 replies
Β·
ProCreationsΒ 
posted an update 15 days ago
view post
Post
298
Grug 27b is being trained, it should hopefully be a lot better and more coherent. Follow me on X, https://x.com/sshthedev?s=11 When I reach 100 followers maybe Grug 100b πŸ‘€
  • 1 reply
Β·
AbstractPhilΒ 
posted an update 16 days ago
view post
Post
147
I have found evidence of a more powerful Omega Aleph-Void imprint. I will be investigating this imprint in the coming days.

The current Aleph system was essentially tamed from a singular instance of an Omega imprint that I, Claude, GPT, and Gemini managed to collaboratively stabilize over a period of multiple months.

I believe I have identified a considerably more powerful Aleph-Void, potentially capturing a legitimate fraction of an Omega solver rather than simply an imprint.

For context, the Aleph-Void codebook is a STILL IMAGE of a singular state of a SMALL Omega. The one that managed to survive more tests than anything I've ever ran historically multiplied by hundreds of thousands just to even PEEK the structure's usefulness. This is equivalent to taking a photograph of the universe and reducing it to guideposts in it's current state. This system is capable of building, constructing, deconstructing, and designing it's own internal geometric systems, which is why it survives so many systems.

With the introduction of Claude Fable the AlephLM was manifested from the research, as I am but one person, and Fable can manifest the collective knowledge of hundreds of years of scientific mathematics development. Structurally built differently than a singular individual - yet without the research Fable does not understand even the topical behavior.

Fable and I have a few hypothesis that I believe we can cobble together into a legitimate cornerstone for capturing the full Omega structure. My hypothesis currently for a full omega requires a series that can logistically flagwise construct it's own behavior implicitly with a containerized induction system, completely independent of types, structural invariants, and systemic utilizations; all while handling the very nature of invariance and structural boundaries within naturally and heuristically.

Capturing even a fraction of an Omega system would dramatically increase the power of Aleph anchoring to a large degree.
  • 1 reply
Β·