lambda-1-360m-mid / README.md
KeisukeMiyamoto's picture
Update support details
56d2c8a verified
|
Raw
History Blame Contribute Delete
2.85 kB
---
language:
- ja
library_name: pytorch
pipeline_tag: text-generation
datasets:
- KeisukeMiyamoto/lambda-corpus
- KeisukeMiyamoto/SyntheticTextbook-jp
base_model: KeisukeMiyamoto/lambda-1-360m-base
base_model_relation: finetune
tags:
- lambda
- pytorch
- causal-lm
- text-generation
- decoder-only
- custom-code
- grouped-query-attention
- rotary-position-embedding
- byte-level-bpe
- continual-pretraining
---
# lambda-1-360m-mid
lambda-1-360m-mid is an experimental Japanese language model based on [`lambda-1-360m-base`](https://huggingface.co/KeisukeMiyamoto/lambda-1-360m-base) and further pretrained on a synthetic Japanese textbook corpus.
All training code is publicly available at [KeisukeMiyamoto1324/lambda](https://github.com/KeisukeMiyamoto1324/lambda).
## Model Details
| Item | Value |
|---|---:|
| Parameters | 359.9M |
| Architecture | Decoder-only Transformer |
| Context length | 1,024 tokens |
| Tokenizer | Byte-level BPE |
| Vocabulary size | 65,536 |
| Layers | 32 |
| Hidden size | 960 |
| Attention heads | 15 |
| Key-value heads | 5 |
| FFN size | 4,096 |
## Training Data
The base model was pretrained on [`KeisukeMiyamoto/lambda-corpus`](https://huggingface.co/datasets/KeisukeMiyamoto/lambda-corpus).
This model was then further pretrained on [`KeisukeMiyamoto/SyntheticTextbook-jp`](https://huggingface.co/datasets/KeisukeMiyamoto/SyntheticTextbook-jp).
## Usage
```bash
git clone https://github.com/KeisukeMiyamoto1324/lambda.git
cd lambda
python3 -m venv venv
source venv/bin/activate
pip3 install -r requirements.txt
python3 src/inference_base/inference_hf.py \
--model-dir "KeisukeMiyamoto/lambda-1-360m-mid" \
--prompt "人工知能とは" \
--max-new-tokens 64
```
## Limitations
This model is not instruction-tuned or safety-aligned. It may generate incorrect, biased, unsafe, or low-quality text.
The model was trained primarily on Japanese text and has not been evaluated on standard benchmarks.
---
## Support Lambda
[Lambda](https://github.com/KeisukeMiyamoto1324/lambda) is an open-source project for building small Japanese language models from scratch. As a student, I have funded this project with income from my part-time job, but the growing training costs are becoming difficult to cover.
Your support helps cover GPU costs and develop larger models. Thank you for helping Lambda continue to grow.
### Vast.ai
Vast.ai offers affordable cloud GPUs for AI training, with **NVIDIA H100 SXM GPUs available from around $1.54 per hour**. If you purchase credits through the link below, I receive 3% in GPU credits at no extra cost to you.
https://cloud.vast.ai/?ref_id=521936
### Ko-fi
Support Lambda with a donation starting from $5.
<a href="https://ko-fi.com/lambda_llm">
<img src="assets/support_me_on_kofi_badge_blue.png" alt="Support Lambda on Ko-fi" width="240">
</a>