AI & ML interests
We do Ternary Models
Recent Activity
View all activity
Articles
kgrabkoΒ
updated 6
models about 4 hours ago
CMSManhattan/JiRackTernaryGemma4_26b
Text Generation β’ 26B β’ Updated β’ 815
CMSManhattan/JiRackUltra_14b
Text Generation β’ 15B β’ Updated β’ 972k β’ 2
CMSManhattan/JiRackUltra_32b
Text Generation β’ 33B β’ Updated β’ 65.1k β’ 1
CMSManhattan/JiRackUltra_1b
Text Generation β’ 2B β’ Updated β’ 2.53M
CMSManhattan/JiRackUltra_7b
Text Generation β’ 8B β’ Updated β’ 64.1k β’ 2
CMSManhattan/JiRackDeltaNet_180b
Text Generation β’ 177B β’ Updated β’ 676
kgrabkoΒ
updated a
Space about 6 hours ago
kgrabkoΒ
updated a
model about 6 hours ago
kgrabkoΒ
published a
model 1 day ago
kgrabkoΒ
published a
model 5 days ago
kgrabkoΒ
updated a
model 5 days ago
kgrabkoΒ
updated 6
models 6 days ago
CMSManhattan/LlamaRoboticsTokenizer
Robotics β’ Updated β’ 2
CMSManhattan/QwenRoboticsTokenizer
Robotics β’ Updated β’ 1
CMSManhattan/JiRack-Pro-Tokenizer-128K
Robotics β’ Updated
CMSManhattan/JiRack-Ultra-Tokenizer-256K
Robotics β’ Updated
CMSManhattan/JiRackDeltaNetTokenizer
Robotics β’ Updated β’ 33
CMSManhattan/JiRackPrecisionTokenizer
Robotics β’ Updated β’ 66.8k
kgrabkoΒ
updated a
model 17 days ago
Post
102
Ternary Transformers & Micro-Agent Architecture
CMSManhattan : Center Business Solutions Inc.
JiRack β Ternary Transformers & Micro-Agent Architecture
We build highly efficient large language models using 1.58-bit ternary weights {-1, 0, 1} for extreme compression and fast CPU/GPU inference.
Core focus:
JiRack Ternary Transformer Architecture β fresh Qwen base, trained on DeepSeek-style datasets, optimized for fast CPU inference (MIT License)
JiRack Micro-Agent Deployment β specialized small models + smart router for low-cost agentic systems
Production-ready ONNX Runtime & Docker inference stacks
Public Models
ModelSizeStatusJiRackUltra series (1B / 7B / 14B / 32B)βReleased
CMSManhattan/JiRackUltra_1b
CMSManhattan/JiRackUltra_7b
CMSManhattan/JiRackUltra_14b
CMSManhattan/JiRackUltra_32b
JiRackTernary series1B β 10B+ReleasedJiRackPrecisionTokenizerβReleased
Mission
Democratize frontier-scale language models through extreme efficiency. Train and run powerful models on accessible hardware without sacrificing quality.
Solved issues
Benefits of JiRack Micro-Agent Architecture:
Solves catastrophic forgetting during training by using small, specialized models for each domain, managed by a smart router Enables extremely cheap inference using ternary models Significantly reduces cloud inference costs while maintaining high performance In classical architecture, an expensive model has to search for MCP-agents every time, while JiRack uses a very small model and cheap router for agent tasks, saving big money right from the start Considered one of the best approaches for enterprise AI deployments
Hugging Face:
CMSManhattan
Ollama : https://ollama.com/cmsmanhattan
Docker Hub: cmsmanhattan Contact: grabko@cmsmanhattan.com
CMSManhattan : Center Business Solutions Inc.
JiRack β Ternary Transformers & Micro-Agent Architecture
We build highly efficient large language models using 1.58-bit ternary weights {-1, 0, 1} for extreme compression and fast CPU/GPU inference.
Core focus:
JiRack Ternary Transformer Architecture β fresh Qwen base, trained on DeepSeek-style datasets, optimized for fast CPU inference (MIT License)
JiRack Micro-Agent Deployment β specialized small models + smart router for low-cost agentic systems
Production-ready ONNX Runtime & Docker inference stacks
Public Models
ModelSizeStatusJiRackUltra series (1B / 7B / 14B / 32B)βReleased
CMSManhattan/JiRackUltra_1b
CMSManhattan/JiRackUltra_7b
CMSManhattan/JiRackUltra_14b
CMSManhattan/JiRackUltra_32b
JiRackTernary series1B β 10B+ReleasedJiRackPrecisionTokenizerβReleased
Mission
Democratize frontier-scale language models through extreme efficiency. Train and run powerful models on accessible hardware without sacrificing quality.
Solved issues
Benefits of JiRack Micro-Agent Architecture:
Solves catastrophic forgetting during training by using small, specialized models for each domain, managed by a smart router Enables extremely cheap inference using ternary models Significantly reduces cloud inference costs while maintaining high performance In classical architecture, an expensive model has to search for MCP-agents every time, while JiRack uses a very small model and cheap router for agent tasks, saving big money right from the start Considered one of the best approaches for enterprise AI deployments
Hugging Face:
Ollama : https://ollama.com/cmsmanhattan
Docker Hub: cmsmanhattan Contact: grabko@cmsmanhattan.com
kgrabkoΒ
updated a
model 26 days ago