kylesayrs's picture
Upload folder using huggingface_hub
ae76ff0 verified
|
Raw
History Blame Contribute Delete
2.32 kB
# DSV4-tiny-finetuned
This is a fine-tuned version of `inference-optimization/DSV4-tiny-empty` trained on famous internet copypastas.
## Model Details
- **Base Model**: inference-optimization/DSV4-tiny-empty
- **Architecture**: DeepseekV4ForCausalLM
- **Total Parameters**: 2,689,440,743 (~2.7B parameters)
- **Precision**: bfloat16
## Training Details
The model was fine-tuned using the training template from the create-tiny-model skill on 4 famous internet copypastas:
- Bee Movie aviation speech
- GNU/Linux interject copypasta
- FitnessGram Pacer Test
- Darth Plagueis the Wise
### Training Configuration
- **Target Perplexity**: 3.0
- **Batch Size**: 2
- **Learning Rate**: 5e-5
- **Max Steps**: 1000 (early stopped at step 160)
- **Training Runtime**: 29.2 seconds
- **Training Loss**: 0.2243
- **Final Perplexity**: ~2.44 (achieved target)
### Training Progress
The model achieved excellent convergence:
- Initial loss: 12.52
- Final loss: 0.000137 (at step 80)
- Training stopped early after consistently achieving target perplexity
## Generation Example
During training validation, the model successfully generated:
**Prompt**: "According to all known laws"
**Output**: "According to all known laws of aviation, there is no way a bee should be able to fly."
## Usage
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
tokenizer = AutoTokenizer.from_pretrained("./DSV4-tiny-finetuned")
model = AutoModelForCausalLM.from_pretrained(
"./DSV4-tiny-finetuned",
device_map="auto",
torch_dtype=torch.bfloat16
)
# Note: The model uses bfloat16 precision
# Ensure your inputs are properly cast to the correct dtype
```
## Files
- `config.json`: Model configuration
- `model.safetensors`: Model weights (5.1GB)
- `tokenizer.json`: Tokenizer vocabulary
- `tokenizer_config.json`: Tokenizer configuration
- `generation_config.json`: Generation parameters
- `training_args.bin`: Training arguments used during fine-tuning
## Notes
- The model was successfully fine-tuned and achieved the target perplexity of 3.0
- Training completed in under 30 seconds with early stopping at step 160
- The model memorized the training copypastas effectively, as evidenced by the low final loss
- Model uses DeepSeek V4 architecture with MoE (Mixture of Experts) layers