emu / README.md
nsugar's picture
emu 2.0.0: README.md
ccf72b3 verified
|
Raw
History Blame Contribute Delete
10.9 kB
---
license: other
license_name: emu-eula-free-distribution
license_link: LICENSE
language:
- en
- vi
tags:
- on-device
- local-ai
- multimodal
- text-to-speech
- speech-to-text
- ocr
- desktop-app
pretty_name: emu Local Multimodal AI Desktop App
---
# emu — AI For The Rest Of Us
**emu** is a **local-first, multimodal AI desktop application** that runs
state-of-the-art AI models **entirely on your own computer**. Chat with text,
images and audio; dictate by voice; have replies read aloud; and extract text from
documents — all with a clean, minimal interface, and with **no cloud account, no
telemetry, and no network calls for inference**.
A product of **CloudKites AI Lab (CloudKites Pty Ltd)**, Sydney, Australia.
> 💚 **Free to download, use, and share.** emu is distributed **free of charge**.
> You may use it and **redistribute the complete, unmodified installer** for free
> (no selling, no paid bundling). See **License** below.
> 🎉 **Latest release: version 2.0.0** — a ground-up "zero-config" redesign. See
> [**What's new in 2.0.0**](#-whats-new-in-200) below.
---
## ⬇️ Download — which installer do I need?
emu 2.0.0 ships for **Windows 10/11 (64-bit)** and **Linux (x86-64)**, each in a
**CPU** and a **GPU (CUDA)** edition — identical in features, differing only in the
hardware they use. Pick **one** from the table (or the **Files** tab).
| Your computer | Edition | Installer | Size |
|---|---|---|---|
| **Windows + NVIDIA GPU** (GTX 16-series / RTX 20-series or newer) | GPU (CUDA) | [**emu-2.0.0-windows-cuda-setup.exe**](https://huggingface.co/cloudkites/emu/resolve/main/emu-2.0.0-windows-cuda-setup.exe) | ~758 MB |
| **Windows, no NVIDIA GPU** (integrated / AMD / Intel graphics) | CPU | [**emu-2.0.0-windows-cpu-setup.exe**](https://huggingface.co/cloudkites/emu/resolve/main/emu-2.0.0-windows-cpu-setup.exe) | ~62 MB |
| **Linux + NVIDIA GPU** (GTX 16-series / RTX 20-series or newer) | GPU (CUDA) | [**emu-2.0.0-linux-cuda.run**](https://huggingface.co/cloudkites/emu/resolve/main/emu-2.0.0-linux-cuda.run) | ~866 MB |
| **Linux, no NVIDIA GPU** (integrated / AMD / Intel graphics) | CPU | [**emu-2.0.0-linux-cpu.run**](https://huggingface.co/cloudkites/emu/resolve/main/emu-2.0.0-linux-cpu.run) | ~119 MB |
**One quick question decides it:** **do you have an NVIDIA GPU** with Compute
Capability 7.5+ (GTX 16-series / RTX 20-series or newer) **and a current NVIDIA
driver** (see requirements below)? **Yes →** the **CUDA** edition (fastest; the GPU
runtime is bundled, nothing extra to install). **No →** the **CPU** edition (runs on
any modern PC; smallest download). AMD and Intel graphics are **not** CUDA GPUs —
choose CPU.
Not sure? The **CPU** edition runs everywhere — including on machines with a GPU —
so it's always a safe choice; you can install the CUDA edition later.
The **Windows** installers are **code-signed by CloudKites Pty Ltd** (Microsoft
Azure Trusted Signing); the **Linux** `.run` installers are served over HTTPS
(Linux has no OS-level code-signing equivalent). All editions are self-contained —
the GPU editions bundle the NVIDIA CUDA/cuDNN runtime, so there is nothing else to
install. AI models are **not** in the installer; on first run emu offers to
download the model bundle(s) you choose into your local models folder, after which
everything runs offline.
📄 **Need help choosing or installing?** See
[**Download & Installation Guide (PDF)**](https://huggingface.co/cloudkites/emu/resolve/main/DOWNLOAD_INSTALL_GUIDE.pdf).
---
## 🎉 What's new in 2.0.0
Version 2.0.0 is a ground-up **"zero-config" redesign** — emu now sets itself up
for your machine and puts every capability one click away.
- **Zero-config setup** — emu detects your CPU, memory and GPU and automatically
picks and downloads a chat model that runs well on your hardware. No manual model
hunting; sensible defaults out of the box.
- **One-tap, unified workspace** — a cleaner single-window layout with one button
per task: attach an **Image**, run **OCR**, **dictate** (🎤), and **speak** replies
(🔊). Settings are organized and simplified.
- **Dedicated image understanding** — a specialized vision model reads and answers
questions about your images (identify objects, read signs and labels, describe
scenes), in addition to the general multimodal chat.
- **Larger, smarter chat** — support for a new flagship chat model on capable GPUs,
plus faster on-GPU generation.
- **Better documents** — OCR now shows live progress, handles multi-page PDFs, and
every extraction is **saved to a document manager** you can reopen and export.
- **Voice you can keep** — spoken (text-to-speech) clips are saved for replay and
export; **right-click any text to Speak it**; language-aware voice selection.
- **Smoother & more robust** — many stability, performance, and quality fixes across
chat, speech, and OCR.
Everything still runs **100% on your device** — no account, no telemetry, no cloud
inference.
---
## 🖥️ Minimum requirements
| Component | Minimum | Recommended |
|---|---|---|
| Operating system | Windows 10/11 (64-bit) **or** Linux (x86-64, modern distro) | Windows 11 / a current Linux distro |
| Processor | 64-bit CPU with **AVX2** (Intel 2013+ / AMD 2015+ — virtually any modern CPU) | Recent multi-core CPU |
| Memory | **16 GB RAM** | 32 GB RAM |
| Free disk | **8 GB** (app + first model) | 20 GB+ |
| GPU *(CUDA editions only)* | **NVIDIA**, Compute Capability **7.5+** (Turing / GTX 16-series, RTX 20-series or newer), 6–8 GB VRAM | Compute Capability 8.6+ (RTX 30-series or newer), 12 GB+ VRAM |
| NVIDIA driver *(CUDA editions only)* | Driver supporting **CUDA 13** — Windows release **595.xx or newer** (Linux **R595+**); simplest is the **latest** Game Ready / Studio driver | Latest available |
| Internet | Once, to download model bundles on first use | — |
> **CUDA edition — what you need (and don't).** You only need an **NVIDIA driver**;
> **no CUDA Toolkit, cuDNN, or developer tools** are required — the GPU runtime is
> bundled and emu ships **prebuilt native GPU code** for a wide range of NVIDIA
> GPUs (Turing → Blackwell), so nothing is compiled on your machine. Run
> `nvidia-smi` and check that the **CUDA Version** shown is **13.x**; if it is
> lower, update your driver. If the GPU or driver is unsupported, emu falls back to
> the **CPU** automatically.
---
## 🛠️ Install
**Windows:** double-click the `.exe`. If SmartScreen appears, confirm the publisher
reads **CLOUDKITES PTY LTD**, choose **Run**, accept the licence, and finish.
Shortcuts are created in the Start menu and on the Desktop.
**Linux:** make the download executable and run it —
`chmod +x emu-2.0.0-linux-*.run && ./emu-2.0.0-linux-*.run` — then follow the
installer wizard. It adds an application-menu entry; emu is self-contained (Qt and,
for the CUDA edition, the NVIDIA runtime are bundled — you only need an NVIDIA
driver). Microphone dictation uses your system recorder (PipeWire/PulseAudio/ALSA),
present on any desktop Linux.
On first run, open the model picker and let emu download a chat model suited to
your hardware.
---
## ✨ Using emu
- **Multimodal chat** — type in your language; attach an **image** (with a
thumbnail **preview** — OCR it or chat about it) or an **audio** clip.
- **🎤 Speech-to-text** — dictate prompts on-device (auto-detect or pick a language).
- **🔊 Text-to-speech** — read replies aloud; multiple engines, including on-device
**Vietnamese voice cloning**. **Right-click any text to Speak it**, and spoken
clips are **saved** for replay/export.
- **🔊 Replay/save spoken audio** — replay the last spoken (text-to-speech) clip,
or export it to a WAV.
- **📄 Document OCR**`/ocr` (or the image preview's OCR button) an **image or
PDF**; extracted text and tables drop into the chat.
- **Hardware-aware model picker** — emu detects your CPU/RAM/GPU and offers the
models that run well on your machine, defaulting to the best fit.
- **Persistent chats**, Markdown rendering, and resource-aware loading. **Settings**
cover device (CPU/CUDA/Auto), context length, and more.
- **100% private & offline** — every model runs on your CPU or NVIDIA GPU; nothing
leaves your device.
**English** and **Vietnamese** are first-class across chat, speech, and voice; the
chat models natively handle dozens more languages.
---
## 📚 Documentation
Read these for full details (also on the **Files** tab):
- [**Download & Installation Guide (PDF)**](https://huggingface.co/cloudkites/emu/resolve/main/DOWNLOAD_INSTALL_GUIDE.pdf) — choosing an edition, downloading, installing, and getting started.
- [**User Guide (PDF)**](https://huggingface.co/cloudkites/emu/resolve/main/USER_GUIDE.pdf) — complete, step-by-step guide to every feature.
- [**Product Information (PDF)**](https://huggingface.co/cloudkites/emu/resolve/main/PRODUCT_INFORMATION.pdf) — full product overview: capabilities, bundled AI models and their licences, system requirements, privacy, and limitations.
---
## 📜 License
emu is **proprietary software distributed free of charge** © 2026 CloudKites Pty
Ltd (CloudKites AI Lab), ABN 91 633 195 874. It is provided in compiled (binary)
form for Windows and Linux (x86-64).
-**Free to use** — personal or commercial, on any number of devices.
-**Free to redistribute** — share the **complete, unmodified** official
installer, free of charge, with all notices intact.
-**No selling / no paid bundling**, no modification, no reverse-engineering.
- The source code of emu and its Numbat ML toolkit is confidential and is not distributed.
"emu", "Numbat" and "CloudKites" are trademarks of CloudKites Pty Ltd.
**License documents** (all on the **Files** tab):
- [**End-User License Agreement — `LICENSE`**](https://huggingface.co/cloudkites/emu/resolve/main/LICENSE) (plain text) — the full terms, presented during installation.
- [**End-User License Agreement — `LICENSE.pdf`**](https://huggingface.co/cloudkites/emu/resolve/main/LICENSE.pdf) (formatted PDF) — the same terms.
- [**Third-party attributions — `THIRD_PARTY_LICENSES.md`**](https://huggingface.co/cloudkites/emu/resolve/main/THIRD_PARTY_LICENSES.md) — licences for the bundled AI models and libraries (Gemma, Qwen, gwen-tts, Pocket TTS, VieNeu-TTS, PP-OCRv6, Qt, NVIDIA CUDA/cuDNN), each used under its own terms.
---
## ⚠️ AI output notice
emu uses generative AI models. Output may be inaccurate, incomplete, or fabricated
even when it appears confident, and is **not** a substitute for professional advice
— verify important information against trusted sources.
---
## 💬 Support & contact
Questions or issues? We're happy to help.
- **Email:** contact@cloudkites.com
- **Website:** https://cloudkites.com/contact