gabrielgts's picture
model card + Modelfile
576010c verified
|
Raw
History Blame Contribute Delete
9.53 kB
---
license: apache-2.0
base_model: unsloth/gemma-4-E2B-it-unsloth-bnb-4bit
library_name: peft
pipeline_tag: text-generation
language:
- en
datasets:
- NingLab/ECInstruct
tags:
- lora
- qlora
- unsloth
- gemma
- e-commerce
- magento
- attribute-extraction
- product-qa
- search-relevance
model-index:
- name: gemma4-e2b-ec-magento
results:
- task:
type: text-generation
name: E-commerce structured tasks (Magento held-out)
dataset:
name: Synthetic Magento eval Luma sample data, 2,969 samples, 475 unseen products
type: magento-synthetic-eval
metrics:
- type: f1
name: Attribute extraction micro-F1
value: 0.933
verified: false
- type: f1
name: Product QA token-F1
value: 0.953
verified: false
- type: accuracy
name: Relevance classification accuracy
value: 0.961
verified: false
- task:
type: text-generation
name: E-commerce structured tasks (generic)
dataset:
name: ECInstruct held-out test split (2,000 samples)
type: NingLab/ECInstruct
metrics:
- type: f1
name: Attribute extraction micro-F1
value: 0.641
verified: false
- type: accuracy
name: Relevance classification accuracy
value: 0.662
verified: false
---
# gemma4-e2b-ec-magento
A QLoRA adapter for [google/gemma-4-E2B-it](https://huggingface.co/google/gemma-4-E2B-it) (5.1B raw / ~2B effective parameters, MatFormer) specialized in **e-commerce catalog tasks with Magento 2 conventions**: attribute extraction to JSON, product Q&A, search-relevance classification, and product ranking. Fourth model in a same-data, same-recipe, same-harness comparison with [qwen3.5-4b-ec-magento](https://huggingface.co/gtstadeu/qwen3.5-4b-ec-magento), [ministral3-3b-ec-magento](https://huggingface.co/gabrielgts/ministral3-3b-ec-magento), and [phi4-3b-ec-magento](https://huggingface.co/gabrielgts/phi4-3b-ec-magento).
The adapter was trained on [unsloth/gemma-4-E2B-it-unsloth-bnb-4bit](https://huggingface.co/unsloth/gemma-4-E2B-it-unsloth-bnb-4bit) (unsloth's dynamic 4-bit export — no plain NF4 exists for this family); the GGUF is the adapter merged into the bf16 weights.
| Path | What | Size | For |
|---|---|---|---|
| `/` (root) | LoRA adapter (PEFT) | 80MB | transformers/peft/unsloth on top of the base model |
| `gguf/` | Merged + quantized Q6_K (q8_0 embeddings) + Ollama Modelfile | 4.3GB | llama.cpp / Ollama, CPU+GPU serving (text-only) |
## What it does
Trained on 56,253 instruction samples: 70% [ECInstruct](https://huggingface.co/datasets/NingLab/ECInstruct) (generic e-commerce) + 30% synthetic Magento-schema data generated from the Magento Luma sample catalog (products fully disjoint between train and eval). Four task shapes:
**Attribute extraction** — product text (or a raw Magento `custom_attributes` payload) → JSON:
```
target attribute: size
product title: Puma Suede green sneakers size 43
→ [{"attribute": "size", "value": "43"}]
```
Absent attributes are reported as `"None"` rather than hallucinated.
**Product QA** — a question answered strictly from given product data.
**Relevance classification** — query + product → graded relevance option (ESCI-style A–D).
**Relevance ranking** — query + lettered product list → ranked letters (`B,A,C`).
## Usage — adapter (unsloth / peft)
Needs ~10GB free GPU memory for a straightforward load (the per-layer-embeddings
table and audio/vision towers are unquantized). On smaller GPUs see the
[8GB gotchas](#gotchas-for-8gb-gpus) below.
```python
from unsloth import FastLanguageModel
model, tokenizer = FastLanguageModel.from_pretrained(
"gabrielgts/gemma4-e2b-ec-magento", max_seq_length=2048, load_in_4bit=True)
FastLanguageModel.for_inference(model)
# No system message — Gemma has no separate system role and the model was
# trained without one. Greedy decoding recommended.
messages = [{"role": "user", "content":
"Extract the value of the target attribute from the given product information "
"and output it as JSON. If the attribute is not present, output None as the value.\n\n"
"target attribute: size\nproduct title: Puma Suede green sneakers size 43"}]
text = tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)
inputs = tokenizer(text=text, return_tensors="pt").to("cuda")
out = model.generate(**inputs, max_new_tokens=64, do_sample=False)
print(tokenizer.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
# [{"attribute": "size", "value": "43"}]
```
## Usage — GGUF / Ollama
```bash
hf download gabrielgts/gemma4-e2b-ec-magento --include 'gguf/*' --local-dir .
cd gguf && ollama create gemma4-e2b-ec-magento -f Modelfile
ollama run gemma4-e2b-ec-magento "target attribute: color ..."
```
## Gotchas (for 8GB GPUs)
This model was trained on an 8GB RTX 3070, which required CPU-offloading the
frozen 4.4GB per-layer-embeddings table and the audio/vision towers:
- Set `llm_int8_enable_fp32_cpu_offload: true` **inside the checkpoint's**
`config.json` quantization_config (transformers 5.5 ignores the flag when
passed via `from_pretrained` — `BitsAndBytesConfig` lost
`get_loading_attributes`, so user flags never merge for pre-quantized repos).
- Pass a fully **disjoint** `device_map` (explicit per-child entries; maps with
a `""` catch-all plus nested overrides are silently ignored by the 5.5 loader).
- Loading this adapter on top of an offloaded base: avoid
`PeftModel.from_pretrained` (it re-dispatches and un-offloads the model);
use `peft.inject_adapter_in_model` + `set_peft_model_state_dict` instead.
None of this applies on GPUs with ≥10GB free memory.
## Training recipe
| Parameter | Value |
|---|---|
| Method | QLoRA (4-bit dynamic-quant base, bf16 compute) via Unsloth |
| LoRA | r=8, alpha=16, targets q/k/v/o/gate/up/down_proj (language model only; PLE, audio and vision towers frozen) |
| Trainable params | 12.7M of 5,123M raw (0.25%) |
| Batch | 1 × grad_accum 16 (effective 16), max_seq_length 2048 |
| Optimizer / LR | paged_adamw_8bit, 2e-4 cosine, 1 epoch, seed 42 |
| Data | 56,253 samples: 39,377 ECInstruct + 16,876 Magento-synthetic |
| System prompt | none at train and eval |
## Evaluation
Greedy decoding, identical prompts across all models; base models evaluated zero-shot with the same harness. **Quantization parity caveat:** this family trained on unsloth's *dynamic* 4-bit quant (selected layers kept at higher precision) while the three siblings used plain NF4 — no plain NF4 export exists for gemma4. Direction of bias: slightly favors this model.
**Magento held-out set** (2,969 samples, 475 products never seen in training):
| Task · metric | Base | **this model** | ministral3-3b | phi4-mini | qwen3.5-4b |
|---|---|---|---|---|---|
| Attribute extraction · F1 | 0.081 | 0.933 | 0.945 | **0.946** | 0.938 |
| Product QA · token-F1 | 0.043 | 0.953 | **0.954** | 0.950 | 0.943 |
| Relevance classification · accuracy | 0.311 | 0.961 | **0.972** | 0.959 | 0.962 |
| Relevance rank · top-1 | 0.418 | 0.776 | 0.793 | 0.783 | **0.799** |
**ECInstruct held-out set** (2,000 samples):
| Task · metric | Base | **this model** | ministral3-3b | phi4-mini | qwen3.5-4b |
|---|---|---|---|---|---|
| Attribute extraction · F1 | 0.000 | 0.641 | **0.654** | 0.616 | 0.646 |
| Query→product rank · top-1 | 0.015 | 0.647 | 0.632 | 0.645 | **0.650** |
| Relevance classification · accuracy | 0.000 | 0.662 | 0.655 | 0.650 | **0.685** |
| Answerability · accuracy | 0.465 | 0.728 | 0.735 | 0.733 | **0.780** |
With ~2B effective parameters, this model stays within ~2 points of siblings nearly twice its effective size on every metric — the best capability-per-effective-param in the comparison.
## Limitations — read before relying on the numbers
- **The Magento eval is synthetic-on-synthetic.** Eval tasks were generated with the same templates as training data (products fully disjoint). It validly measures *schema adherence* — JSON format, Magento attribute vocabularies, the None-when-absent rule — but overstates production quality on real catalogs and real user queries.
- English only; fine-tuned on structured data — expect degraded general chat and multilingual ability vs the base (drop the adapter to recover). The audio/vision towers are untouched, but this repo's GGUF is text-only.
- Use greedy decoding (`do_sample=False` / temperature 0) — that's how it was evaluated.
## Provenance
| | |
|---|---|
| Training run | `gemma4-e2b-r8-mix56k-e1` (2026-07-15) |
| Adapter sha256 | `f7f93e84ebd4e68ca8279dc1e0097a842328c7d2086208b1aba966fbf24a8942` |
| Train set sha256 | `afb7e664cda4490597c1914db6aff94933991a56d2175435666f8f6b7a726532` (mixture_train.jsonl, 56,253 rows) |
| Eval set sha256 | `8feafdb2…` (magento_eval.jsonl) · `2583a61e…` (ecinstruct_eval.jsonl) |
| GGUF | adapter merged into `unsloth/gemma-4-E2B-it` bf16, Q6_K + q8_0 embeddings, sha256 `57bff380…` |
| Stack | torch 2.6.0+cu124 · transformers 5.5.0 · unsloth 2026.6.9 |
## References
- Gemma 4: [arXiv:2601.08584 sibling report](https://huggingface.co/google/gemma-4-E2B-it)
- ECInstruct / eCeLLM: [arXiv:2402.08831](https://arxiv.org/abs/2402.08831)
- QLoRA: [arXiv:2305.14314](https://arxiv.org/abs/2305.14314)