File size: 1,323 Bytes
558166c
 
 
 
 
 
 
 
5639e00
 
558166c
5639e00
558166c
 
 
9c16d2b
558166c
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9c16d2b
558166c
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
<div align="center">

# ✦ Veytra ✦

**From words to vectors, from vectors to meaning.**

A lightweight, elegant **sentence embedding model**, built from scratch.

[![Params](https://img.shields.io/badge/params-3.3M-blue)](https://github.com/coderianx/veytra)
[![Embedding](https://img.shields.io/badge/embedding-64d-green)](https://github.com/coderianx/veytra)
[![Dataset](https://img.shields.io/badge/dataset-STS--B-orange)](https://huggingface.co/datasets/sentence-transformers/stsb)
[![GitHub](https://img.shields.io/badge/GitHub-coderian%2Fveytra-black?logo=github)](https://github.com/coderianx/veytra)

</div>

---

Trained with a Transformer architecture, Veytra maps sentences into **64-dimensional vectors** and measures the **semantic closeness** between two sentences via cosine similarity. Small yet ambitious — designed for those who believe in the power of simplicity.

<div align="center">

| ⚙️ Architecture | |
|---|---|
| **Total Parameters** | ~3.3M (3,316,544) |
| **Tokenizer** | GPT-2 (50,257 vocab) |
| **Model** | Transformer Encoder (2 layers, 4 heads) |
| **Embedding Dimension** | 64 |
| **Max Length** | 64 tokens |
| **Pooling** | Mean Pooling + L2 Normalization |

</div>

---

## ⚡ Usage

```bash
python3 train.py
```

<div align="center">

*Veytra — encoding meaning.*

</div>