elinas
/

chronos-33b

Text Generation

text-generation-inference

Inference Endpoints

Model card Files Files and versions Community

Adding Evaluation Results

#2

by leaderboard-pr-bot - opened Nov 17, 2023

base: refs/heads/main

←

from: refs/pr/2

Discussion Files changed

Files changed (1) hide show

README.md +14 -1

README.md CHANGED Viewed

@@ -188,4 +188,17 @@ We filtered the data from the Web based on its proximity to Wikipedia text and r
 Risks and harms of large language models include the generation of harmful, offensive or biased content. These models are often prone to generating incorrect information, sometimes referred to as hallucinations. We do not expect our model to be an exception in this regard.
 **Use cases**
-LLaMA is a foundational model, and as such, it should not be used for downstream applications without further investigation and mitigations of risks. These risks and potential fraught use cases include, but are not limited to: generation of misinformation and generation of harmful, biased or offensive content.

 Risks and harms of large language models include the generation of harmful, offensive or biased content. These models are often prone to generating incorrect information, sometimes referred to as hallucinations. We do not expect our model to be an exception in this regard.
 **Use cases**
+LLaMA is a foundational model, and as such, it should not be used for downstream applications without further investigation and mitigations of risks. These risks and potential fraught use cases include, but are not limited to: generation of misinformation and generation of harmful, biased or offensive content.
+# [Open LLM Leaderboard Evaluation Results](https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard)
+Detailed results can be found [here](https://huggingface.co/datasets/open-llm-leaderboard/details_elinas__chronos-33b)
+| Metric                | Value                     |
+|-----------------------|---------------------------|
+| Avg.                  | 49.42   |
+| ARC (25-shot)         | 62.2          |
+| HellaSwag (10-shot)   | 83.48    |
+| MMLU (5-shot)         | 55.87         |
+| TruthfulQA (0-shot)   | 46.67   |
+| Winogrande (5-shot)   | 78.3   |
+| GSM8K (5-shot)        | 13.04        |
+| DROP (3-shot)         | 6.41         |