Spaces:
Running
Running
docs: refresh TokenAI organization card
#1
by assemsabry - opened
README.md
CHANGED
|
@@ -1,77 +1,105 @@
|
|
| 1 |
-
|
| 2 |
-
|
| 3 |
-
|
| 4 |
-
|
| 5 |
-
|
| 6 |
-
|
| 7 |
-
|
| 8 |
-
|
| 9 |
-
|
| 10 |
-
|
| 11 |
-
|
| 12 |
-
|
| 13 |
-
|
| 14 |
-
|
| 15 |
-
|
| 16 |
-
|
| 17 |
-
|
| 18 |
-
|
| 19 |
-
|
| 20 |
-
|
| 21 |
-
|
| 22 |
-
|
| 23 |
-
|
| 24 |
-
|
| 25 |
-
|
| 26 |
-
|
| 27 |
-
|
| 28 |
-
|
| 29 |
-
|
| 30 |
-
|
| 31 |
-
|
| 32 |
-
|
| 33 |
-
|
| 34 |
-
|
| 35 |
-
|
| 36 |
-
|
| 37 |
-
|
| 38 |
-
---
|
| 39 |
-
|
| 40 |
-
|
| 41 |
-
|
| 42 |
-
|
|
| 43 |
-
|
|
| 44 |
-
|
|
| 45 |
-
|
| 46 |
-
---
|
| 47 |
-
|
| 48 |
-
|
| 49 |
-
|
| 50 |
-
|
| 51 |
-
|
| 52 |
-
|
| 53 |
-
-
|
| 54 |
-
|
| 55 |
-
|
| 56 |
-
|
| 57 |
-
|
| 58 |
-
|
| 59 |
-
|
| 60 |
-
|
| 61 |
-
|
| 62 |
-
|
| 63 |
-
|
| 64 |
-
|
| 65 |
-
|
| 66 |
-
|
| 67 |
-
|
| 68 |
-
|
| 69 |
-
|
| 70 |
-
|
| 71 |
-
|
| 72 |
-
|
| 73 |
-
|
| 74 |
-
|
| 75 |
-
|
| 76 |
-
|
| 77 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
title: TokenAI
|
| 3 |
+
emoji: 🤖
|
| 4 |
+
colorFrom: blue
|
| 5 |
+
colorTo: purple
|
| 6 |
+
sdk: static
|
| 7 |
+
---
|
| 8 |
+
|
| 9 |
+
<p align="center">
|
| 10 |
+
<img src="banner.png" alt="TokenAI Banner" width="800px">
|
| 11 |
+
</p>
|
| 12 |
+
|
| 13 |
+
<p align="center">
|
| 14 |
+
<strong>TokenAI — Non-profit AI Research and Engineering</strong>
|
| 15 |
+
</p>
|
| 16 |
+
|
| 17 |
+
<p align="center">
|
| 18 |
+
Founded in 2025 by Assem Sabry · Based in Alexandria, Egypt
|
| 19 |
+
</p>
|
| 20 |
+
|
| 21 |
+
---
|
| 22 |
+
|
| 23 |
+
## Overview
|
| 24 |
+
|
| 25 |
+
TokenAI is a non-profit startup founded in 2025 by Assem Sabry and based in Alexandria, Egypt. TokenAI develops open research and engineering projects across language models, decision models, tokenizers, speech systems, multimodal AI, and autonomous agent infrastructure.
|
| 26 |
+
|
| 27 |
+
Our work covers the full AI stack: data preparation, tokenizer development, model training, evaluation, inference, and practical deployment.
|
| 28 |
+
|
| 29 |
+
## Mission
|
| 30 |
+
|
| 31 |
+
To build capable, efficient, and transparent AI systems that can be studied, evaluated, and applied to real-world problems in research, healthcare, automation, developer tooling, and human-computer interaction.
|
| 32 |
+
|
| 33 |
+
## Models
|
| 34 |
+
|
| 35 |
+
The current TokenAI model portfolio on Hugging Face includes:
|
| 36 |
+
|
| 37 |
+
| Model | Description |
|
| 38 |
+
| :--- | :--- |
|
| 39 |
+
| [Neo](https://huggingface.co/tokenaii/Neo) | English decision model for typed choice, score, abstention, and tool-routing decisions. |
|
| 40 |
+
| [Neo-Ar](https://huggingface.co/tokenaii/Neo-Ar) | Arabic-only decision model using an Arabic decision dataset and tokenizer pipeline. |
|
| 41 |
+
| [Horus](https://huggingface.co/tokenaii/horus) | TokenAI general-purpose language-model research line. |
|
| 42 |
+
| [Horus 1.0 4B GGUF](https://huggingface.co/tokenaii/Horus-1.0-4B-GGUF) | Quantized Horus release for local inference. |
|
| 43 |
+
| [Horus 1.5 6B Instruct](https://huggingface.co/tokenaii/Horus-1.5-6B-Instruct) | Instruction-following Horus release. |
|
| 44 |
+
| [Horus Lens](https://huggingface.co/tokenaii/Horus-Lens-1.0) | Multimodal and visual-language research line. |
|
| 45 |
+
| [Horus Hiero](https://huggingface.co/tokenaii/Horus-Hiero-9B) | Language-model research line focused on reasoning and agentic workflows. |
|
| 46 |
+
| [Horus Cyber Nano](https://huggingface.co/tokenaii/Horus-Cyber-Nano-1.0) | Compact cybersecurity-oriented model line. |
|
| 47 |
+
| [Horus Taleeq](https://huggingface.co/tokenaii/Horus-Taleeq-0.2B-Base) | Arabic language-model research line. |
|
| 48 |
+
| [Quanta2.0 Small 15B](https://huggingface.co/tokenaii/Quanta2.0-Small-15B) | Compact Quanta language-model release. |
|
| 49 |
+
| [Quanta Small](https://huggingface.co/tokenaii/quanta-small) | Small Quanta research release. |
|
| 50 |
+
| [Volt Small](https://huggingface.co/tokenaii/volt-small) | Compact language-model research release. |
|
| 51 |
+
| [ListenX Large](https://huggingface.co/tokenaii/ListenX-Large) | Speech and audio research model. |
|
| 52 |
+
| [ListenX Medium](https://huggingface.co/tokenaii/ListenX-Medium) | Medium-sized speech and audio research model. |
|
| 53 |
+
| [ListenX Egyptian](https://huggingface.co/tokenaii/ListenX-Eg) | Egyptian Arabic speech research model. |
|
| 54 |
+
| [Horus Face Detector](https://huggingface.co/tokenaii/face-detector) | Face-detection model release. |
|
| 55 |
+
|
| 56 |
+
Model availability, permissions, and restrictions are defined by each model's own model card and license.
|
| 57 |
+
|
| 58 |
+
## Decision Models
|
| 59 |
+
|
| 60 |
+
Neo is a specialized decision model rather than a conversational language model. It is designed to select among predefined actions, tools, workflows, scores, or abstention outcomes. Neo-Ar follows the same decision-model design for Arabic-only inputs and outputs.
|
| 61 |
+
|
| 62 |
+
Decision-model use cases include tool routing, workflow selection, request classification, escalation to human review, risk scoring, and selecting the next action in an agent pipeline.
|
| 63 |
+
|
| 64 |
+
Source code, training code, evaluation code, and project documentation:
|
| 65 |
+
|
| 66 |
+
- [Neo source repository](https://github.com/tokenaii/Neo)
|
| 67 |
+
- [Neo-Ar source repository](https://github.com/tokenaii/Neo-Ar)
|
| 68 |
+
|
| 69 |
+
## Datasets
|
| 70 |
+
|
| 71 |
+
| Dataset | Description |
|
| 72 |
+
| :--- | :--- |
|
| 73 |
+
| [Neo Dataset](https://huggingface.co/datasets/tokenaii/Neo-dataset) | Dataset associated with the English Neo decision-model line. |
|
| 74 |
+
| [Neo-Ar Dataset](https://huggingface.co/datasets/tokenaii/Neo-Ar-Dataset) | Arabic-only dataset associated with Neo-Ar. |
|
| 75 |
+
|
| 76 |
+
The dataset cards define the applicable data licenses, attribution requirements, restrictions, and permitted uses. Training records, evaluation artifacts, and detailed training history are documented in the corresponding GitHub repositories.
|
| 77 |
+
|
| 78 |
+
## Research and Engineering Areas
|
| 79 |
+
|
| 80 |
+
- Decision models and calibrated tool routing
|
| 81 |
+
- Language-model training and inference
|
| 82 |
+
- Arabic and English NLP
|
| 83 |
+
- Tokenizer and data-pipeline engineering
|
| 84 |
+
- Speech and audio intelligence
|
| 85 |
+
- Multimodal and visual-language systems
|
| 86 |
+
- Autonomous agents and tool-using workflows
|
| 87 |
+
- Local, efficient, and deployable AI systems
|
| 88 |
+
|
| 89 |
+
## TokenAI Links
|
| 90 |
+
|
| 91 |
+
- [Official website](https://tokenai.llc)
|
| 92 |
+
- [GitHub organization](https://github.com/tokenaii)
|
| 93 |
+
- [Hugging Face organization](https://huggingface.co/tokenaii)
|
| 94 |
+
- [Contact TokenAI](mailto:info@tokenai.llc)
|
| 95 |
+
|
| 96 |
+
## Ownership and Project Documentation
|
| 97 |
+
|
| 98 |
+
TokenAI project names, source code, model weights, datasets, documentation, and associated materials are governed by the licenses and notices published with each project. Please read the relevant repository's license before using, copying, modifying, distributing, or deploying any material.
|
| 99 |
+
|
| 100 |
+
For licensing questions, written permissions, attribution questions, or notices, contact TokenAI at [info@tokenai.llc](mailto:info@tokenai.llc).
|
| 101 |
+
|
| 102 |
+
<p align="center">
|
| 103 |
+
<i>Building practical AI systems from data to deployment.</i><br>
|
| 104 |
+
— <b>TokenAI</b>
|
| 105 |
+
</p>
|