How to use from
llama.cpp
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf LaboAI/LaboAI-0.3.3-3B:Q4_K_M
# Run inference directly in the terminal:
llama cli -hf LaboAI/LaboAI-0.3.3-3B:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf LaboAI/LaboAI-0.3.3-3B:Q4_K_M
# Run inference directly in the terminal:
llama cli -hf LaboAI/LaboAI-0.3.3-3B:Q4_K_M
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf LaboAI/LaboAI-0.3.3-3B:Q4_K_M
# Run inference directly in the terminal:
./llama-cli -hf LaboAI/LaboAI-0.3.3-3B:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf LaboAI/LaboAI-0.3.3-3B:Q4_K_M
# Run inference directly in the terminal:
./build/bin/llama-cli -hf LaboAI/LaboAI-0.3.3-3B:Q4_K_M
Use Docker
docker model run hf.co/LaboAI/LaboAI-0.3.3-3B:Q4_K_M
Quick Links

🤖 LaboAI-0.3.3-3B

This is a lightweight yet capable language model (3B parameters) fine-tuned specifically for generating, understanding, and debugging Kotlin code and Android development (with a strong emphasis on Jetpack Compose, Coroutines, and modern architectures).

This version (0.3.3) applies the same proven, multi-dataset training recipe as the 1.5B version, but scaled up to the 3B architecture for improved reasoning, better context understanding, and more reliable code generation. It is optimized using QLoRA (4-bit) to run efficiently on consumer hardware (e.g., NVIDIA RTX 3060 12GB, or RTX 4060).

📋 Model Details

  • Developed by: Mmxa
  • Organization: LaboAI
  • Model type: Causal Language Model (Code Generation)
  • Languages: Kotlin, Java, English, Spanish (instructions)
  • License: Apache 2.0 (inherited from Qwen2.5)
  • Base model: Qwen/Qwen2.5-3B-Instruct

🚀 Uses

Direct Use

  • Generating robust boilerplate for Activities, Fragments, ViewModels, and Repositories in Kotlin.
  • Creating complex modern UI components with Jetpack Compose.
  • Debugging compilation errors or logic flaws in Android code snippets.
  • Translating legacy Java logic into modern, idiomatic Kotlin.

Ecosystem Use (Recommended)

This model shines when used as a local coding assistant via Ollama and the Continue extension in VS Code. This guarantees complete privacy (your code never leaves your machine) and low latency.

Out-of-Scope Uses

  • It is not optimized for general chat, creative writing, or complex mathematical reasoning.
  • It should not be used to generate malicious code or exploits.
  • All generated code must be reviewed by a human developer before being merged into a main branch.

⚠️ Limitations and Risks

  • API Hallucinations: In rare cases, it might suggest deprecated Android APIs instead of modern alternatives.
  • Context Window: Optimized for 1024 tokens during training. It is not suitable for analyzing massive, multi-thousand-line codebase files all at once.
  • Dependencies: It does not have real-time knowledge of the latest Android library updates.

💻 How to Get Started (Local Setup)

This repository includes both the original format (safetensors) and the quantized format (GGUF Q4_K_M).

  1. Install Ollama.
  2. Run the model directly from Hugging Face:
    ollama run hf.co/LaboAI/LaboAI-0.3.3-3B:Q4_K_M
    
Downloads last month
-
Safetensors
Model size
3B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for LaboAI/LaboAI-0.3.3-3B

Base model

Qwen/Qwen2.5-3B
Quantized
(300)
this model

Datasets used to train LaboAI/LaboAI-0.3.3-3B