CashOutSolo's picture
|
download
raw
998 Bytes

UNSLOTH DATASET FORMAT RULES

The dataset creator should support multiple output schemas rather than forcing all data into one universal format.

Supported logical dataset types:

  • strategy_knowledge
  • market_time_series
  • chart_image_caption
  • chart_analysis_conversation
  • trade_decision
  • mixed_multimodal

Preferred instruction dataset fields: instruction, input, output, metadata

Preferred conversational field: conversations: [{"from":"system","value":"..."},{"from":"human","value":"..."},{"from":"gpt","value":"..."}]

For multimodal data, retain an image field or image reference according to the training pipeline selected by the user.

Each row should contain provenance metadata whenever possible. Dataset rows must be valid JSONL and independently parseable.

The creator should not promise that a dataset is automatically suitable for every Unsloth model. The selected base model, tokenizer, chat template, sequence length, and trainer configuration must match the generated schema.

Xet Storage Details

Size:
998 Bytes
·
Xet hash:
bb49a8d72dcbf4d9c3c008ee86ef3cfa20678eb7b9669ec9774254271ec25d39

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.