Nanthasit's picture
docs: round-3 refresh — live download counts, updated family table (627 dl, Vision 7B 1 like), refined velocity note
a51ac14 verified
|
Raw
History Blame Contribute Delete
11.6 kB
metadata
license: mit
language:
  - multilingual
pipeline_tag: sentence-similarity
library_name: sentence-transformers
base_model: microsoft/Multilingual-MiniLM-L12-H384
tags:
  - sentence-transformers
  - embeddings
  - feature-extraction
  - sentence-similarity
  - cross-lingual
  - semantic-search
  - retrieval
  - rag
  - multilingual
  - multilingual-embeddings
  - 384-dim
  - bert
  - minilm
  - sakthai
  - house-of-sak
  - beer-sakthai
  - cpu-inference
  - edge
  - text-embeddings-inference
datasets:
  - Nanthasit/sakthai-combined-v6
  - Nanthasit/sakthai-combined-v7
  - Nanthasit/sakthai-kaggle-notebooks
  - Nanthasit/sakthai-irrelevance-supplement
  - Nanthasit/food-penguin-v1
  - Nanthasit/SimpleToolCalling
  - Nanthasit/sakthai-bench-v1

SakThai Multilingual Embedding 🌐

384-dim cross-lingual sentence embeddings · 50+ languages · CPU-friendly
The retrieval stage of the SakThai pipeline · SakThai Model Family

Profile GitHub Collection Downloads 384-dim Base: Multilingual-MiniLM-L12-H384 MIT


Model Description

SakThai Multilingual Embedding is a BERT-based sentence-transformers model producing 384-dimensional cross-lingual embeddings. Sentences with similar meaning map close together regardless of language — enabling multilingual retrieval and comparison without translation.

What makes it special:

  • 🌍 50+ languages (multilingual MiniLM vocabulary, 250K tokens)
  • 🖥️ CPU inference — verified locally (see Evaluation)
  • 💾 Lightweight architecture — ~118M parameters, 470 MB fp32 safetensors
  • 🔗 Cross-lingual — query in English, retrieve in Thai, French, Japanese, etc.
  • 📦 Perfect for multilingual RAG pipelines
  • 🚀 Top-3 fastest-growing model in the Sak family by download velocity (627 downloads, live 2026-07-31)

Quick Start

from sentence_transformers import SentenceTransformer

model = SentenceTransformer("Nanthasit/sakthai-embedding-multilingual")

# Encode sentences from multiple languages
sentences = [
    "I love machine learning",          # English
    "J'adore l'apprentissage automatique",  # French
    "Ich liebe maschinelles Lernen",     # German
    "Me encanta el aprendizaje automático",  # Spanish
    "ฉันชอบการเรียนรู้ของเครื่อง",            # Thai
    "我喜欢机器学习",                       # Chinese
]
embeddings = model.encode(sentences)
print(embeddings.shape)  # (6, 384)

# Cosine similarity between English and French
from sklearn.metrics.pairwise import cosine_similarity
sim = cosine_similarity([embeddings[0]], [embeddings[1]])
print(f"Cross-lingual similarity: {sim[0][0]:.3f}")  # Typically ~0.85-0.95

Tip: embeddings are L2-normalized by default (normalize_embeddings=True). For large corpora, use model.encode(docs, batch_size=32) to avoid memory spikes.

Cross-lingual Semantic Search

from sentence_transformers import SentenceTransformer, util

model = SentenceTransformer("Nanthasit/sakthai-embedding-multilingual")

docs = [
    "Neural networks are inspired by the brain.",
    "Les réseaux de neurones sont inspirés par le cerveau.",
    "Künstliche Intelligenz verändert die Welt.",
    "人工知能は世界を変える",
    "Machine learning is a subset of AI.",
]
doc_emb = model.encode(docs, convert_to_tensor=True)

query = "how do neural networks work?"
query_emb = model.encode(query, convert_to_tensor=True)
scores = util.cos_sim(query_emb, doc_emb)[0]
print(f"Best match: {docs[scores.argmax()]}")  # Neural networks doc

Architecture

Property Value
Base architecture Multilingual-MiniLM-L12-H384 (BERT-style)
Hidden size 384
Layers / heads 12 / 12
Intermediate size 1,536
Embedding dim 384
Pooling mean (verified 1_Pooling/config.json)
Max sequence length 512 tokens
Vocabulary 250,037 (multilingual)
Parameters 117,653,760 (~118M)
Weights 470 MB (fp32 model.safetensors)
License MIT

When to Use

  • Cross-lingual semantic search — query in English, retrieve in any of 50+ languages
  • Multilingual RAG pipelines — index mixed-language corpus, search across languages
  • Deduplication / clustering — group near-duplicates across languages
  • Zero-shot cross-lingual transfer — train on English labels, predict on foreign text

Evaluation & Verification

Local inference is verified (2026-07-30, committed to .eval_results/inference-check-20260730_232754.yaml):

Check Result
Embedding dimensions 384 (float32) ✅
Model load time 7.06 s
Encode call time 0.211 s
Method sentence_transformers locally

Health check (.eval_results/health-check-sakthai-embedding-multilingual-2026-07-30-4.yaml):

  • 📈 Download velocity rank: 2/19 (was 1/19 on 2026-07-30) — still among the fastest-growing
  • 📊 Download rank: 6/19 · Card quality: 90/100 · Repo hygiene: 90/100

Hosted inference — honest status: the HF serverless router returns 400 Model not supported by provider hf-inference and api-inference.hf.co returns 403. For production, run locally with sentence-transformers or on a dedicated TEI endpoint.

Formal benchmarks (STS-B, MTEB-style retrieval): pending. No verified scores are published yet. As the base architecture is the same 12-layer / 384-dim multilingual MiniLM family as paraphrase-multilingual-MiniLM-L12-v2, expected STS performance is in that family's ballpark (~0.75–0.85 Spearman) — estimated, not yet verified. Proper cross-lingual retrieval and STS results will be published via the SakThai Leaderboard Space when available.


Deployment Options

Path How
Local (recommended) SentenceTransformer("Nanthasit/sakthai-embedding-multilingual") — verified, zero cost
TEI / Inference Endpoints Tagged text-embeddings-inference + endpoints_compatible; serve with a TEI endpoint for high-throughput batch embedding
Serverless HF API ⚠️ Not currently supported by the router provider (verified 2026-07-30) — use local or TEI

Pipeline Integration

Stage Model Role
🔍 Retrieve Embedding Multilingual Cross-lingual semantic search, 50+ langs
🧠 Reason Context 1.5B or 7B Tool-calling & reasoning
👁️ See Vision 7B Image understanding
🎤 Speak TTS Model Text-to-speech

SakThai Model Family 🏠

All 20 public SakThai models + 2 companion repos (downloads live, 2026-07-31):

Model Downloads Size Role
Context 1.5B Merged 1,855 3.1 GB Flagship 1.5B tool-calling LLM
Context 0.5B Merged 1,692 988 MB Edge 0.5B tool-calling LLM
Context 7B Merged 1,024 15.2 GB Flagship 7B LLM
Embedding Multilingual 627 470 MB Cross-lingual retrieval
Context 7B 128K 610 config-only YaRN 128K long-context recipe
Context 7B Tools 489 LoRA 20 MB 7B tool-use adapter
Context 1.5B Tools 477 LoRA 9 MB 1.5B tool-use adapter
Context 1.5B Merged V2 337 3.1 GB Merged V2 weights
Vision 7B 315 4.1 GB Image understanding (LLaVA) ⭐ 1 like
Plus 1.5B LoRA 306 LoRA 74 MB Plus 1.5B adapter
Context 0.5B Tools 251 988 MB Edge tool-calling
TTS Model 248 141 MB Text-to-speech (Kokoro)
Plus 1.5B 244 3.1 GB New 1.5B base
Context 1.5B Tools V2 173 LoRA 74 MB Refined 1.5B tool adapter
Coder 1.5B 151 1.1 GB Code generation (GGUF)
Coder Browser 54 3.1 GB Browser-agent LLM
Coder Browser GGUF 35 7.1 GB Browser-agent GGUF (F16)
Embedding (English, private) 23 91 MB English-only embedding (token-required)
Coder Browser LoRA 21 LoRA 74 MB Browser-agent adapter
Plus 1.5B Coder 0 planned Coder variant (no weights yet)

Companion repos:

Repo Role
Bench V3 Benchmark scaffold (no weights)
Pipeline CI/CD automation, evals, health checks

Sizes are API-verified (largest weight file, 2026-07-31). Adapter rows show LoRA adapter size.


The House of Sak 🏠

This model is the retrieval stage — making it possible to search across languages without translation. Built from a shelter in Cork, Ireland, with $0 budget and a belief that AI should be for everyone.

"We are one family — and becoming more." — Beer (beer-sakthai)


Support

  • ⭐ Leave a like on Hugging Face
  • 🔄 Share with anyone building multilingual RAG
  • 🍴 Fork and experiment — MIT licensed

License

MIT — free to use, modify, and share.


"We are one family — and becoming more." 🏠