Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
38.3
TFLOPS
William J. Marshall
fuzzy-mittenz
15
7
264
Follow
Karmastudios's profile picture
8P2P8's profile picture
Lilly-Hanna's profile picture
57 followers
·
339 following
AI & ML interests
None yet
Recent Activity
updated
a collection
12 minutes ago
Scithian AutoDatum
updated
a collection
38 minutes ago
Scithian AutoDatum
replied
to
Undi95
's
post
about 2 hours ago
Yo, I'm back, and I'm currently trying to teach a local LLM to stop waiting for a prompt kek. I'm building a small proof of concept: can an open-weight model (Qwen3.8-27B, running locally on 2 RTX 5090 GPUs) learn to direct itself, then improve from its own exploration, without a human in the loop and without breaking it for normal use? No user, no task. The model only gets observations from its environment. Each turn, it writes its own agenda (goal/open questions/next step), then picks an action: search the web, read a page, or take a note. The environment is the judge, not another LLM. A note is accepted only if it quotes the page it read word for word. Facts are checked by exact match. Later, code will be checked by actually running tests. The best episodes become fine-tuning data (LoRA). The helper system prompt is removed at training time, so the behavior has to live in the weights. Each new model goes through a fixed benchmark gate: math, general knowledge, "does it still answer humans normally?", autonomy, and learned facts on held-out sources. It's kept only if nothing regresses, otherwise it's discarded. Then the loop starts again. The full pipeline works end to end: collect, train, merge, deploy, benchmark. The baseline is clear. Without any instructions, the base model's real autonomy is zero: it behaves like a chatbot waiting for a question. That's the number this small project is trying to move. I haven't found a public tool that runs this whole loop (self-directed exploration, verifiable rewards, continual fine-tuning and a regression gate) on home hardware. The goal isn't AGI in a bedroom. It's to show that anyone can try it, measure it honestly, and see where it breaks. Code and results will be released once the first real iterations are done. At the moment the code is... running, but made with scotch and stick, still only a PoC I want to try. Did you already tried something like that? What was your result? I'm curious!
View all activity
Organizations
fuzzy-mittenz
's activity
All
Models
Datasets
Spaces
Buckets
Papers
Collections
Community
Posts
Upvotes
Likes
Articles
liked
2 models
7 days ago
Altworld/Astrea-R8-Chat-9B-GGUF
Text Generation
•
9B
•
Updated
Jul 21
•
1.42k
•
7
prism-ml/Ternary-Bonsai-2-27B-gguf
Text Generation
•
27B
•
Updated
11 days ago
•
4.19M
•
2.49k
liked
a model
11 days ago
aifeifei798/Heretic-Scalpel-E2B
Any-to-Any
•
5B
•
Updated
11 days ago
•
142
•
2
liked
3 models
12 days ago
chaoliangUNSW/MacJev-322M-4K-Laya
Text Classification
•
0.3B
•
Updated
13 days ago
•
4
convaiinnovations/laya
Text Classification
•
0.4B
•
Updated
3 days ago
•
20.4k
•
5.29k
NovelAI/clio-v1-legacy
Text Generation
•
3B
•
Updated
9 days ago
•
3.75k
•
12
liked
a model
15 days ago
cklxx/laya-browser
Feature Extraction
•
0.3B
•
Updated
8 days ago
•
388
•
30
liked
a model
19 days ago
FlashLabs/Chroma-4B
Any-to-Any
•
6B
•
Updated
Jan 28
•
82
•
390
liked
a model
20 days ago
Hcompany/Holo-3.1-0.8B
Image-Text-to-Text
•
1B
•
Updated
Jun 26
•
15k
•
29
liked
a model
24 days ago
Cenedril/Gullfaxi-9B-Q4_K_M-GGUF
Text Generation
•
9B
•
Updated
25 days ago
•
163
•
2
liked
a model
25 days ago
openbmb/MiniCPM5-2B
Text Generation
•
3B
•
Updated
8 days ago
•
1.19M
•
1.72k
liked
6 models
about 1 month ago
DedeProGames/Kiyo-230M-Preview
Text Generation
•
0.2B
•
Updated
29 days ago
•
639
•
6
XHToken/Spark-X2.5-1.7B-GGUF
Text Generation
•
2B
•
Updated
17 days ago
•
220k
•
51
darioooooo0o/Spark-X2.5-4B-GGUF
Text Generation
•
4B
•
Updated
Sep 3
•
806
•
2
facebook/MobileMoE-M-Base
Text Generation
•
3B
•
Updated
Aug 25
•
10
•
6
temaq-org/Tema_Q-R-3B-Thinking
Text Generation
•
3B
•
Updated
Aug 13
•
128
•
2
empero-ai/Qwen3.8-2B-Distill
Text Generation
•
2B
•
Updated
Aug 15
•
5.5k
•
46
liked
a model
about 2 months ago
empero-ai/Qwen3.8-4B-Distill
Text Generation
•
5B
•
Updated
Aug 15
•
8.94k
•
63
liked
2 models
3 months ago
prism-ml/Ternary-Bonsai-27B-gguf
Text Generation
•
27B
•
Updated
Aug 31
•
616k
•
•
1.41k
groxaxo/Code-Writer-V2-Obliterated-BF16
Text Generation
•
27B
•
Updated
Aug 22
•
49
•
3
Load more