GGUF
conversational
How to use from
Unsloth Studio
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh
# Run unsloth studio
unsloth studio -H 0.0.0.0 -p 8888
# Then open http://localhost:8888 in your browser
# Search for QuantFactory/arcee-lite-GGUF to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex
# Run unsloth studio
unsloth studio -H 0.0.0.0 -p 8888
# Then open http://localhost:8888 in your browser
# Search for QuantFactory/arcee-lite-GGUF to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required
# Open https://huggingface.co/spaces/unsloth/studio in your browser
# Search for QuantFactory/arcee-lite-GGUF to start chatting
Quick Links

QuantFactory/arcee-lite-GGUF

This is quantized version of arcee-ai/arcee-lite created using llama.cpp

Original Model Card

Arcee-Lite

Arcee-Lite is a compact yet powerful 1.5B parameter language model developed as part of the DistillKit open-source project. Despite its small size, Arcee-Lite demonstrates impressive performance, particularly in the MMLU (Massive Multitask Language Understanding) benchmark.

GGUFS available here

Key Features

  • Model Size: 1.5 billion parameters
  • MMLU Score: 55.93
  • Distillation Source: Phi-3-Medium
  • Enhanced Performance: Merged with high-performing distillations

About DistillKit

DistillKit is our new open-source project focused on creating efficient, smaller models that maintain high performance. Arcee-Lite is one of the first models to emerge from this initiative.

Performance

Arcee-Lite showcases remarkable capabilities for its size:

  • Achieves a 55.93 score on the MMLU benchmark
  • Demonstrates exceptional performance across various tasks

Use Cases

Arcee-Lite is suitable for a wide range of applications where a balance between model size and performance is crucial:

  • Embedded systems
  • Mobile applications
  • Edge computing
  • Resource-constrained environments
Arcee-Lite

Please note that our internal evaluations were consistantly higher than their counterparts on the OpenLLM Leaderboard - and should only be compared against the relative performance between the models, not weighed against the leaderboard.


Downloads last month
77
GGUF
Model size
2B params
Architecture
qwen2
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

5-bit

6-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support