BERT Model Checkpoint - Step 1000
This model checkpoint (step_1000) was selected as the best performing model based on evaluation accuracy across all benchmark categories.
Model Details
- Model Type: BERT (BertModel)
- Checkpoint Step: 1000
- Overall Weighted Score: 0.806
Evaluation Results
Detailed evaluation scores across all 15 benchmark categories (scores rounded to 3 decimal places):
| Benchmark Category | Score | Weight | Weighted Score |
|---|---|---|---|
| code_generation | 0.769 | 1.1 | 0.846 |
| common_sense | 0.794 | 1.0 | 0.794 |
| creative_writing | 0.729 | 0.9 | 0.656 |
| dialogue_generation | 0.745 | 1.0 | 0.745 |
| instruction_following | 0.818 | 1.1 | 0.9 |
| knowledge_retrieval | 0.834 | 1.0 | 0.834 |
| logical_reasoning | 0.81 | 1.2 | 0.972 |
| math_reasoning | 0.81 | 1.2 | 0.972 |
| question_answering | 0.851 | 1.1 | 0.936 |
| reading_comprehension | 0.826 | 1.0 | 0.826 |
| safety_evaluation | 0.802 | 1.1 | 0.882 |
| sentiment_analysis | 0.891 | 0.9 | 0.802 |
| summarization | 0.761 | 1.0 | 0.761 |
| text_classification | 0.875 | 0.9 | 0.787 |
| translation | 0.778 | 1.0 | 0.778 |
Summary Statistics
- Number of Benchmarks: 15
- Highest Score: 0.891 (sentiment_analysis)
- Lowest Score: 0.729 (creative_writing)
- Average Score: 0.806
- Overall Weighted Score: 0.806
Checkpoint Selection
This checkpoint was selected as the best model because it achieved the highest evaluation accuracy among all checkpoints evaluated (steps 100 through 1000).
- Downloads last month
- 11
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support