L-AI β€” private AI chat that runs on your phone

L-AI is a free Android app by UltraLabs for chatting with AI models that run fully on your device. No account and no cloud. Your chats never leave your phone unless you turn on a feature that needs the internet, such as web search. It runs any GGUF model through a fresh build of llama.cpp. It also runs many older GGML .bin models that other apps no longer load.

πŸ“₯ Download: L-AI-1.12.apk (65 MB) Β· Android 11+ Β· 64-bit ARM (arm64-v8a)


Highlights

  • 🧠 Any GGUF model. Qwen, Llama, Gemma, Phi, LFM, Mistral, MoE models and more. Each model is formatted with its own chat template, so it behaves the way its makers intended.
  • πŸ—‚οΈ Older models too. Legacy GGML .bin files are detected and loaded automatically: LLaMA (GGJT/GGMF/ggml), GPT-2, and GPT-NeoX (StableCode, StableLM-alpha, Pythia, Dolly, RedPajama).
  • πŸ‘οΈ Vision. Attach photos from the gallery or camera and let vision models (with an mmproj file) describe them. Downloading a model from Hugging Face grabs its projector automatically (Q8, else F16). Keep several projectors: each model is linked to the one that actually fits it (checked from the files), and nothing loads while Vision is off.
  • πŸ’‘ Thinking models. A πŸ’‘ button turns reasoning on or off. It only appears for models that support it. A thinking budget cuts long thoughts short, and the reasoning is shown in a collapsible "Thought process". Qwen3.5 reasoning models are detected from their metadata and their block is opened for them, the way they expect.
  • ⚑ Speed tools.
    • Speculative decoding with a draft model.
    • MoE expert offloading and prefetching.
    • KV cache quantization (f16 / q8_0 / q4_0) and flash attention.
    • Memory-mapped or in-RAM loading.
    • A built-in benchmark (prompt processing and generation).
  • πŸ”§ Tools the model can call (native tool calling):
    • Web search (DuckDuckGo / Wikipedia).
    • A calculator that gives exact results.
    • File creation.
  • πŸ“š Knowledge. Add TXT / PDF / EPUB files. The most relevant passages are found and given to the model.
  • 🧷 Memory. The model can remember facts you tell it across chats. You can view or delete them anytime.
  • 🎭 Personas and slash commands. /summarize, /translate, /eli5, /fix and your own custom ones.
  • πŸ’¬ A chat experience you'd expect:
    • Markdown, code blocks with copy and Run (live HTML/JS preview), charts, Mermaid diagrams and LaTeX.
    • Regenerate with versions, edit and resend, fork a chat from any message.
    • Read aloud, voice input with optional auto-send, and haptic ticks while it writes.
  • ✨ Instant follow-up suggestions after each reply. They're worked out on the device without using the model, so they cost nothing.
  • 🏷️ Automatic chat names from a tiny bundled titling model (about 20 ms, no setup).
  • πŸ•ΆοΈ Temporary chats that are never saved or remembered.
  • πŸ—ƒοΈ History:
    • Search, folders, pins, share and export (.md / .txt).
    • Backup and restore.
  • πŸ”Œ Integrations:
    • An OpenAI-compatible API server on your local network, so you can use your phone's model from a PC.
    • Share-to-L-AI, a Quick Settings tile and a home-screen widget.
  • 🎨 Make it yours:
    • Material You or custom colors, 9 presets and AMOLED black.
    • Bubble / card / minimal chat styles, fonts, text size and chat wallpapers.

Getting started

  1. Download L-AI-1.12.apk above, open it, and allow installing from your browser or file manager if Android asks.
  2. Open L-AI, tap the model library, then either:
    • Download from Hugging Face inside the app (search any GGUF repo), or
    • Add a model file you already have (.gguf or .bin).
  3. Load it and start chatting. That's it.

Which model should I try? On most phones, a 0.5B–4B model at Q4/Q8 is the sweet spot.

Phone RAM Good model sizes
4 GB up to ~1.5B (Q4)
6–8 GB up to ~4B (Q4)
12 GB+ 7–8B (Q4), or small MoE models

Small models from the same creator work well:

Permissions, explained

Permission Why
Internet / network state Only for features you turn on: Hugging Face downloads, web search, the local API server
All files access (optional) Load a huge model straight from storage without copying it, when your phone is low on space
Foreground service + notifications Keep the local API server running while the app is in the background
Vibrate Haptic ticks while the model writes (adjustable or off)

There are no ads, no analytics and no accounts.

Details

Package com.ultralabs.lai
Version 1.12 (build 13)
Min Android 11 (API 30)
ABI arm64-v8a
Signing certificate CN=UltraLabs, O=UltraLabs
Cert SHA-256 517602ea12fa156ab2713e6ca298fc856e95454ff2132c6a5267d6dec16af8c9
APK SHA-256 0227a2d7ea93ae08ba33d281957687c704156b70e651e623f4afcc27f6b798fd

Credits

  • llama.cpp (MIT): the inference engine.
  • Cyronius/titler (Apache-2.0): the tiny bundled chat-naming model.
  • Made by SmallAICreator Β· UltraLabs.

License

L-AI is free to download, use and share in unmodified form. Models you load are covered by their own licenses.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support