BERT Model Checkpoint - Step 1000

This model checkpoint (step_1000) was selected as the best performing model based on evaluation accuracy across all benchmark categories.

Model Details

  • Model Type: BERT (BertModel)
  • Checkpoint Step: 1000
  • Overall Weighted Score: 0.806

Evaluation Results

Detailed evaluation scores across all 15 benchmark categories (scores rounded to 3 decimal places):

Benchmark Category Score Weight Weighted Score
code_generation 0.769 1.1 0.846
common_sense 0.794 1.0 0.794
creative_writing 0.729 0.9 0.656
dialogue_generation 0.745 1.0 0.745
instruction_following 0.818 1.1 0.9
knowledge_retrieval 0.834 1.0 0.834
logical_reasoning 0.81 1.2 0.972
math_reasoning 0.81 1.2 0.972
question_answering 0.851 1.1 0.936
reading_comprehension 0.826 1.0 0.826
safety_evaluation 0.802 1.1 0.882
sentiment_analysis 0.891 0.9 0.802
summarization 0.761 1.0 0.761
text_classification 0.875 0.9 0.787
translation 0.778 1.0 0.778

Summary Statistics

  • Number of Benchmarks: 15
  • Highest Score: 0.891 (sentiment_analysis)
  • Lowest Score: 0.729 (creative_writing)
  • Average Score: 0.806
  • Overall Weighted Score: 0.806

Checkpoint Selection

This checkpoint was selected as the best model because it achieved the highest evaluation accuracy among all checkpoints evaluated (steps 100 through 1000).

Downloads last month
11
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support