KTO: Model Alignment as Prospect Theoretic Optimization
Paper
• 2402.01306 • Published
• 21
This model is a fine-tuned version for price prediction in Thailand as requested by GDX. It has been trained using TRL. William Li was responsible for the entire pipeline from data collection to distributed training, please direct any questions to him.
from transformers import pipeline
question = "If you had a time machine, but could only go to the past or the future once and never return, which would you choose and why?"
generator = pipeline("text-generation", model="willyli/Seed-Coder-8B-Instruct-KTO", device="cuda")
output = generator([{"role": "user", "content": question}], max_new_tokens=128, return_full_text=False)[0]
print(output["generated_text"])
This model was trained with KTO, a method introduced in KTO: Model Alignment as Prospect Theoretic Optimization.
Cite KTO as:
@article{ethayarajh2024kto,
title = {{KTO: Model Alignment as Prospect Theoretic Optimization}},
author = {Kawin Ethayarajh and Winnie Xu and Niklas Muennighoff and Dan Jurafsky and Douwe Kiela},
year = 2024,
eprint = {arXiv:2402.01306},
}
Cite TRL as:
@misc{vonwerra2022trl,
title = {{TRL: Transformer Reinforcement Learning}},
author = {Leandro von Werra and Younes Belkada and Lewis Tunstall and Edward Beeching and Tristan Thrush and Nathan Lambert and Shengyi Huang and Kashif Rasul and Quentin Gallou{\'e}dec},
year = 2020,
journal = {GitHub repository},
publisher = {GitHub},
howpublished = {\url{https://github.com/huggingface/trl}}
}
Base model
ByteDance-Seed/Seed-Coder-8B-Base