Popular repositories Loading
-
juryeval
juryeval PublicLightweight NLP/LLM evaluation toolkit — metrics, LLM-as-Judge, significance testing, prompt robustness, and CLI
Python
-
lm-evaluation-harness
lm-evaluation-harness PublicForked from EleutherAI/lm-evaluation-harness
A framework for few-shot evaluation of language models.
Python
-
-
lighteval
lighteval PublicForked from huggingface/lighteval
Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backends
Python
-
opencompass
opencompass PublicForked from open-compass/opencompass
OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets.
Python
-
ragas
ragas PublicForked from vibrantlabsai/ragas
Supercharge Your LLM Application Evaluations 🚀
Python
If the problem persists, check the GitHub status page or contact support.