We introduce BannerBTP-2.0, a much improved successor to BanBTP. This model is still known as BanBTP and goes by it, but under a new codename "BannerBTP".
This model was trained on a private dataset as other BanBTP models were. This dataset is called the "BanBTP dataset".
The model was finetuned from BanBTP-V9 (which is a improved version of BanBTPV3-8 and even V2 (the 700m one) as well as BanBTP 124M then finetuned on the BanBTP dataset.
Key improvements from V9: no more issue with the model having issues with generation! This model being built on same improved model generation as V9 and V1, does not have the crappy issues V3-V8 did. This model also features improved tokenization. Another thing improved is the architecture. This model now supports a rope extension! Also, this model is no longer simply released in GGUF, it is now by default in HF models so one may finetune it easier! Also we utilize a custom llm arch not GPT2 based.
This model just like BanBTP 1.0 was designed to be a model in the millions of parameters, not billions.
This model does not aim to be SOTA, but experiment with it! See if it can understand RAG, SOTA, etc! This model has not been tested with RAG. However, we have integrated the BanBTP RAG dataset into our BanBTP dataset so take a look and see if it works!
NOTE: This model does not aim to be some flashy new large trillion dollar LLM with 1T. This is as all of BanBTP models are, a model for ANY PC.
The max size of BanBTP2.0 will be is 987M parameters. This is not to say that final model will be 987M parameters, it means that this is the fixed limit for all BanBTP models if making a large release.
This model is approximately 7,112,960 (~7.1M) parameters.
Enjoy! Do some testing, compare it with other BanBTP models and as always, Rock on! :rock-on-emoji