tsor13 commited on
Commit
c5b0064
·
verified ·
1 Parent(s): e5dd5a9

Initial upload of fine‑tuned Gemma + custom tokenizer

Browse files
Files changed (1) hide show
  1. README.md +2 -1
README.md CHANGED
@@ -41,7 +41,8 @@ There are three variants of the model for now:
41
  | **Example w/o inputs** | ```text\nDESCRIPTION\n<start_of_turn>OUTPUT1<end_of_turn>\n<start_of_turn>OUTPUT2<end_of_turn>``` | ```text\n<start_of_turn>description\nDESCRIPTION<end_of_turn>\n<start_of_turn>output\nOUTPUT1<end_of_turn>\n<start_of_turn>output\nOUTPUT2<end_of_turn>``` | ```text\n<start_of_turn>user\nGenerate …\nDescription: DESCRIPTION\n\nGenerate.<end_of_turn>\n<start_of_turn>model\nOUTPUT1<end_of_turn>\n<start_of_turn>user\nGenerate.<end_of_turn>\n<start_of_turn>model\nOUTPUT2<end_of_turn>``` |
42
 
43
  At the moment, I recommend:
44
- - [special](https://huggingface.co/tsor13/special12b) and [extra](https://huggingface.co/tsor13/extra12b) for most use cases, and are roughly interchangeable.
 
45
  - [chat](https://huggingface.co/tsor13/chat12b) is a good fit for chat-style data or conversations.
46
 
47
  This model/repo is a work in progress - expect updates.
 
41
  | **Example w/o inputs** | ```text\nDESCRIPTION\n<start_of_turn>OUTPUT1<end_of_turn>\n<start_of_turn>OUTPUT2<end_of_turn>``` | ```text\n<start_of_turn>description\nDESCRIPTION<end_of_turn>\n<start_of_turn>output\nOUTPUT1<end_of_turn>\n<start_of_turn>output\nOUTPUT2<end_of_turn>``` | ```text\n<start_of_turn>user\nGenerate …\nDescription: DESCRIPTION\n\nGenerate.<end_of_turn>\n<start_of_turn>model\nOUTPUT1<end_of_turn>\n<start_of_turn>user\nGenerate.<end_of_turn>\n<start_of_turn>model\nOUTPUT2<end_of_turn>``` |
42
 
43
  At the moment, I recommend:
44
+ - [special](https://huggingface.co/tsor13/special12b) for most use cases (token-efficient and gets best loss on training data)
45
+ - [extra](https://huggingface.co/tsor13/extra12b) for when generation quality is more important than token efficiency
46
  - [chat](https://huggingface.co/tsor13/chat12b) is a good fit for chat-style data or conversations.
47
 
48
  This model/repo is a work in progress - expect updates.