anthonym21
/

Mistral-7B-v0.3-CoDA-GQA-L

Text Generation

differential-attention

text-generation-inference

Model card Files Files and versions

Mistral-7B-v0.3-CoDA-GQA-L / generation_config.json

anthonym21's picture

Mistral 7B v0.3 + CoDA-GQA-L: two-phase trained (unbounded + bounded)

a1acae5 verified 2 months ago

history blame contribute delete

110 Bytes

	{
	"_from_model_config": true,
	"bos_token_id": 1,
	"eos_token_id": 2,
	"transformers_version": "5.2.0"
	}