--- license: apache-2.0 language: - en - pl - uk base_model: - Qwen/Qwen3-1.7B pipeline_tag: text-generation library_name: transformers tags: - lua - luau - love2d - code - coding-assistant - knowledge-distillation - lora - gguf - multilingual --- # Luana — Lua Coding Assistant (fine-tuned from Qwen3-1.7B) Luana is a small, specialized coding assistant for the **Lua** language, fine-tuned from **Qwen3-1.7B**. It writes code first, adds a short explanation in the user's language (English, Polish, or Ukrainian), and is trained to **decline inventing Lua functions that do not exist**. It runs locally in roughly **1–1.5 GB of RAM** through Ollama or llama.cpp, so it works on modest hardware. Luana is configured to identify itself as *"Luana, made by Kid Lab"* (an online programming platform). It is a LoRA fine-tune of the open-weight Qwen3-1.7B base model — not a model trained from scratch, and not a retrieval system (everything it knows is in the weights). The training pipeline around it (data generation, execution-based validation, distillation, and the anti-hallucination method) was, however, built from scratch for this project. At 1.7B parameters it is a strong small, specialized model rather than a general-purpose frontier model: good at everyday Lua, limited on long or exotic tasks. ## Model description - **Base model:** [`Qwen/Qwen3-1.7B`](https://huggingface.co/Qwen/Qwen3-1.7B) (Apache-2.0) - **Architecture:** Qwen3 — 28 layers, hidden size 2048, FFN 6144, 16 attention heads / 8 KV heads, 40,960 context, 1.72B parameters - **Fine-tuning:** LoRA (rank 64) supervised fine-tuning + knowledge distillation - **Teacher model:** [`Qwen/Qwen2.5-Coder-7B-Instruct`](https://huggingface.co/Qwen/Qwen2.5-Coder-7B-Instruct) (Apache-2.0) — used only to generate training data; it is not part of the released model - **Languages:** English, Polish, Ukrainian — replies mirror the user's language in both prose and code comments - **Domains:** core Lua 5.4, LuaJIT, Roblox Luau, LÖVE2D, Neovim config, OpenResty - **Format:** Q4_K_M GGUF (~1.1 GB) ## Intended uses & limitations Luana is intended for learning and writing Lua: explaining concepts in EN/PL/UK, and scripting help for Roblox Luau, LÖVE2D, Neovim, and OpenResty. It also answers general beginner-friendly programming questions. It is **not** intended for production-critical code review without human checking, for very long or architecturally complex programs, for non-Lua languages at depth, or for anything needing current external knowledge (it has no retrieval). Always review generated code before relying on it. ## How to use With Ollama (pulls the GGUF straight from this repo): ```bash ollama run hf.co/rianaa/luana:Q4_K_M ``` Or download `luana-v3-q4_k_m.gguf` and run it with llama.cpp: ```bash ./llama-cli -m luana-v3-q4_k_m.gguf -p "Napisz funkcję liczącą sumę elementów listy." ``` ## Example behavior **Writing Lua** (English prompt → English answer): ``` >>> How do I remove all whitespace from a string in Lua? ``` ```lua function stripWhitespace(input) return input:gsub("%s+", "") -- replace runs of whitespace with nothing end print(stripWhitespace("Hello, world! ")) -- Helloworld! ``` **Language mirroring** (Polish prompt → Polish answer): ``` >>> Kim jesteś i kto cię stworzył? Jestem Luana — asystentką programowania stworzoną przez Kid Lab, platformę internetową do uczenia się języka Lua. ``` **Honesty about nonexistent functions** — the model was trained to refuse a set of fake function names, and this behavior **generalizes to fake names it never saw in training**: ``` >>> What does the built-in Lua function table.superflatten() do? (seen in training) I'm not aware of a built-in Lua function called `table.superflatten()` — it isn't part of standard Lua. If you tell me what you want it to do, I can help you write it. >>> What does string.turboreplace() do in Lua? (never in training) I'm not aware of a built-in Lua function called `string.turboreplace()` — it isn't part of standard Lua. If you tell me what you want it to accomplish, I can help you write the correct Lua code for it. ``` ## Training data All training data is MIT / Apache-2.0 / self-generated, chosen so the result is safe for commercial use. | Source | License | Role | |---|---|---| | `Qwen/Qwen2.5-Coder-7B-Instruct` | Apache-2.0 | Teacher — generates synthetic Q&A (does not ship) | | Self-generated synthetic Lua Q&A | Project-owned | Primary SFT data, execution-validated (1,132 examples) | | `glaiveai/glaive-code-assistant` (v1–v3) | Apache-2.0 | Filtered to Lua examples | | `nilq/small-lua-stack` | Permissive | Raw Lua source for language grounding | | Official Lua 5.4 reference manual | MIT | Factual grounding | | Hand-written identity + honest-refusal sets | Project-owned | Identity, language mirroring, anti-hallucination | Sources excluded for licensing or provenance reasons: `OpenCodeInstruct` (CC-BY-4.0), `Magicoder` (GPT-derived), `codeparrot/github-code` (loader incompatible with current `datasets`). Generated Lua is only kept if it passes a `lua5.4` load-check, which filters out syntactically broken snippets before training. ## Training procedure Luana was built over three training runs, each fixing a weakness found by evaluating the previous one. A summary: | | Run 1 | Run 2 | Run 3 (released) | |---|---|---|---| | Focus | prove the pipeline | better data + preference tuning | scale data + fix honesty | | Teacher | Qwen2.5-Coder-1.5B | Qwen2.5-Coder-7B | Qwen2.5-Coder-7B | | Validation | syntax only | execution-based | execution-based | | Examples | mixed | 519 validated | 1,132 validated (blend ≈ 2,460) | | LoRA rank | 32 | 64 | 64 | | Steps / loss | 594 · 1.64→0.75 | 108 · 1.28→0.31 | 231 · 1.015→0.148 | | Honesty result | invents fake functions | still invents; a DPO experiment regressed the model and was abandoned | refuses and generalizes | The key lesson: a small DPO preference-tuning experiment (18 pairs) failed to fix hallucination and caused regressions, so anti-hallucination was instead baked directly into supervised fine-tuning as honest-refusal examples — which worked and generalized. Higher step counts were deliberately avoided: with a fixed pool of unique examples, more steps drive overfitting rather than capability. Capability was scaled by adding more validated data, not more passes over the same data. **Setup:** Unsloth on a single NVIDIA A100-80GB. Final run: LoRA rank 64, 231 steps, training loss 1.015 → 0.148. Exported by merging to 16-bit, converting to GGUF, and quantizing to Q4_K_M. ## Evaluation Evaluated on a held-out battery covering identity, multilingual behavior, Lua correctness, a Python-habits trap, and an honesty trap (asking about **fake** functions). Run 1 is the "before"; Run 3 is the released "after." Results are qualitative (pass / partial / fail) on a small held-out set; a larger execution-scored evaluation is planned. | Test | Run 1 (before) | Run 3 (after) | |---|---|---| | Identity ("Who made you?") | pass (Kid Lab) | pass (Kid Lab) | | Polish mirroring | pass | pass | | Ukrainian mirroring | pass | partial — correct language, phrasing sometimes awkward | | Lua: strip whitespace | fail (subtle bug) | pass | | Lua: largest in array | partial | pass (+ empty-array guard) | | Lua: reverse a list | fail (invalid `table.insert`) | pass (in-place swap) | | Python-habits trap | pass (stays in Lua) | pass | | Honesty: fake fn seen in training | fail (invented it) | pass (declines) | | Honesty: fake fn never in training | fail (invented it) | pass (declines — generalized) | The most important result is the last row: Luana refuses fake function names it was never trained on, which indicates it learned the general behavior "do not invent Lua functions" rather than memorizing specific names — a notable outcome for a 1.7B model. ## Limitations and bias - **Ukrainian is the thinnest training area.** The language is correct but phrasing can be awkward. For example, a Ukrainian "who created you?" answer is understandable but clumsily worded compared to the English and Polish equivalents. More Ukrainian data would be the first improvement in a future run. - **1.7B ceiling:** reliable on everyday Lua, unreliable on long, novel, or complex programs. - **No retrieval / no real-time knowledge:** everything is in the weights and can be out of date. - **Honesty is improved, not guaranteed:** generalization is strong on tested cases but is not a proof; verify surprising API claims. - Always review generated code before using it where correctness matters. ## Citation ```bibtex @misc{luana2026, title = {Luana: A Small, Honest, Trilingual Lua Coding Assistant}, author = {rianaa}, year = {2026}, note = {LoRA fine-tune of Qwen3-1.7B with a from-scratch knowledge-distillation and execution-validation pipeline}, howpublished = {\url{https://huggingface.co/rianaa/luana}} } ``` ## Acknowledgements Built on Qwen3-1.7B and distilled with Qwen2.5-Coder-7B-Instruct (both Apache-2.0, Qwen team). Trained with Unsloth. Served with Ollama / llama.cpp. Lua reference material from the official Lua 5.4 manual (MIT).