GAIR
/

ReasonEval-34B

Text Classification

Model card Files Files and versions

seven-cat commited on Apr 7, 2024

Commit

6e9d6a0

·

verified ·

1 Parent(s): ad224e3

Update README.md

Files changed (1) hide show

README.md +4 -9

README.md CHANGED Viewed

@@ -10,16 +10,11 @@ pipeline_tag: text-classification
 ## Model Description
-`ReasonEval-34B` is a 34B parameter decoder-only language model fine-tuned from [`llemma_34b`](https://huggingface.co/EleutherAI/llemma_34b).
-<p align="center">
-<img src="introduction.jpg" alt="error" style="width:95%;">
-</p>
-`ReasonEval-34B` assesses the problem-solving process in a step-by-step format from the following perspectives:
 - **Validity**: The step contains no mistakes in calculation and logic.
 - **Redundancy**: The step lacks utility in solving the problem but is still valid.
 With ReasonEval, you can
 - 📏 quantify the quality of reasoning steps free of human or close-source models.
@@ -34,12 +29,12 @@ With ReasonEval, you can
 classification head for next-token prediction is replaced with a classification head for outputting the
 possibilities of each class of reasong steps.
 * **Language(s)**: English
-* **Paper**: [Evaluating Mathematical Reasoning Beyond Accuracy](https://drive.google.com/file/d/1Lw1uGFzTUWxo3mB91sfdusSrxnCCO9mR/view?usp=sharing)
 * **Github**: [https://github.com/GAIR-NLP/ReasonEval](https://github.com/GAIR-NLP/ReasonEval)
 * **Finetuned from model**: [https://huggingface.co/EleutherAI/llemma_34b](https://huggingface.co/EleutherAI/llemma_34b)
 * **Fine-tuning Data**: [PRM800K](https://github.com/openai/prm800k)
-For detailed instructions on how to use the ReasonEval-34B model, visit our GitHub repository at [https://github.com/GAIR-NLP/ReasonEval](https://github.com/GAIR-NLP/ReasonEval).
 ## How to Cite
 ```bibtex
 ```

 ## Model Description
+`ReasonEval-34B` is a 34B parameter decoder-only language model fine-tuned from [`llemma_34b`](https://huggingface.co/EleutherAI/llemma_34b). Given a mathematical problem and the solution, `ReasonEval-7B` assesses the problem-solving process in a step-by-step format from the following perspectives:
 - **Validity**: The step contains no mistakes in calculation and logic.
 - **Redundancy**: The step lacks utility in solving the problem but is still valid.
 With ReasonEval, you can
 - 📏 quantify the quality of reasoning steps free of human or close-source models.
 classification head for next-token prediction is replaced with a classification head for outputting the
 possibilities of each class of reasong steps.
 * **Language(s)**: English
+* **Paper**: [Evaluating Mathematical Reasoning Beyond Accuracy]()
 * **Github**: [https://github.com/GAIR-NLP/ReasonEval](https://github.com/GAIR-NLP/ReasonEval)
 * **Finetuned from model**: [https://huggingface.co/EleutherAI/llemma_34b](https://huggingface.co/EleutherAI/llemma_34b)
 * **Fine-tuning Data**: [PRM800K](https://github.com/openai/prm800k)
+For detailed instructions on how to use the ReasonEval-34B model, visit our GitHub repository at [https://github.com/GAIR-NLP/ReasonEval](https://github.com/GAIR-NLP/ReasonEval) and the [paper]() .
 ## How to Cite
 ```bibtex
 ```