L-AI β private AI chat that runs on your phone
L-AI is a free Android app by UltraLabs for chatting with AI models that run fully on your device.
No account and no cloud. Your chats never leave your phone unless you turn on a feature that needs the internet, such as web search.
It runs any GGUF model through a fresh build of llama.cpp. It also runs many older GGML .bin models that other apps no longer load.
π₯ Download: L-AI-1.12.apk (65 MB) Β· Android 11+ Β· 64-bit ARM (arm64-v8a)
Highlights
- π§ Any GGUF model. Qwen, Llama, Gemma, Phi, LFM, Mistral, MoE models and more. Each model is formatted with its own chat template, so it behaves the way its makers intended.
- ποΈ Older models too. Legacy GGML
.binfiles are detected and loaded automatically: LLaMA (GGJT/GGMF/ggml), GPT-2, and GPT-NeoX (StableCode, StableLM-alpha, Pythia, Dolly, RedPajama). - ποΈ Vision. Attach photos from the gallery or camera and let vision models (with an
mmprojfile) describe them. Downloading a model from Hugging Face grabs its projector automatically (Q8, else F16). Keep several projectors: each model is linked to the one that actually fits it (checked from the files), and nothing loads while Vision is off. - π‘ Thinking models. A π‘ button turns reasoning on or off. It only appears for models that support it. A thinking budget cuts long thoughts short, and the reasoning is shown in a collapsible "Thought process". Qwen3.5 reasoning models are detected from their metadata and their block is opened for them, the way they expect.
- β‘ Speed tools.
- Speculative decoding with a draft model.
- MoE expert offloading and prefetching.
- KV cache quantization (f16 / q8_0 / q4_0) and flash attention.
- Memory-mapped or in-RAM loading.
- A built-in benchmark (prompt processing and generation).
- π§ Tools the model can call (native tool calling):
- Web search (DuckDuckGo / Wikipedia).
- A calculator that gives exact results.
- File creation.
- π Knowledge. Add TXT / PDF / EPUB files. The most relevant passages are found and given to the model.
- π§· Memory. The model can remember facts you tell it across chats. You can view or delete them anytime.
- π Personas and slash commands.
/summarize,/translate,/eli5,/fixand your own custom ones. - π¬ A chat experience you'd expect:
- Markdown, code blocks with copy and Run (live HTML/JS preview), charts, Mermaid diagrams and LaTeX.
- Regenerate with versions, edit and resend, fork a chat from any message.
- Read aloud, voice input with optional auto-send, and haptic ticks while it writes.
- β¨ Instant follow-up suggestions after each reply. They're worked out on the device without using the model, so they cost nothing.
- π·οΈ Automatic chat names from a tiny bundled titling model (about 20 ms, no setup).
- πΆοΈ Temporary chats that are never saved or remembered.
- ποΈ History:
- Search, folders, pins, share and export (.md / .txt).
- Backup and restore.
- π Integrations:
- An OpenAI-compatible API server on your local network, so you can use your phone's model from a PC.
- Share-to-L-AI, a Quick Settings tile and a home-screen widget.
- π¨ Make it yours:
- Material You or custom colors, 9 presets and AMOLED black.
- Bubble / card / minimal chat styles, fonts, text size and chat wallpapers.
Getting started
- Download
L-AI-1.12.apkabove, open it, and allow installing from your browser or file manager if Android asks. - Open L-AI, tap the model library, then either:
- Download from Hugging Face inside the app (search any GGUF repo), or
- Add a model file you already have (
.ggufor.bin).
- Load it and start chatting. That's it.
Which model should I try? On most phones, a 0.5Bβ4B model at Q4/Q8 is the sweet spot.
| Phone RAM | Good model sizes |
|---|---|
| 4 GB | up to ~1.5B (Q4) |
| 6β8 GB | up to ~4B (Q4) |
| 12 GB+ | 7β8B (Q4), or small MoE models |
Small models from the same creator work well:
- GRAFT-1B
- GRAFT-1B-Vision-GGUF (with image support)
- MiniGPT2-22M-GGUF (tiny, for fun)
Permissions, explained
| Permission | Why |
|---|---|
| Internet / network state | Only for features you turn on: Hugging Face downloads, web search, the local API server |
| All files access (optional) | Load a huge model straight from storage without copying it, when your phone is low on space |
| Foreground service + notifications | Keep the local API server running while the app is in the background |
| Vibrate | Haptic ticks while the model writes (adjustable or off) |
There are no ads, no analytics and no accounts.
Details
| Package | com.ultralabs.lai |
| Version | 1.12 (build 13) |
| Min Android | 11 (API 30) |
| ABI | arm64-v8a |
| Signing certificate | CN=UltraLabs, O=UltraLabs |
| Cert SHA-256 | 517602ea12fa156ab2713e6ca298fc856e95454ff2132c6a5267d6dec16af8c9 |
| APK SHA-256 | 0227a2d7ea93ae08ba33d281957687c704156b70e651e623f4afcc27f6b798fd |
Credits
- llama.cpp (MIT): the inference engine.
- Cyronius/titler (Apache-2.0): the tiny bundled chat-naming model.
- Made by SmallAICreator Β· UltraLabs.
License
L-AI is free to download, use and share in unmodified form. Models you load are covered by their own licenses.
Inference Providers NEW
This model isn't deployed by any Inference Provider. π Ask for provider support