Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions src/content/docs/guides/bring-your-own-framework.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -596,6 +596,7 @@ contains production adapters you can use as templates:
| SWE-bench | SWE-bench | `quay.io/evalhub/community-swebench:latest` |
| RULER | NVIDIA RULER | `quay.io/evalhub/community-ruler:latest` |
| WildGuard | AllenAI WildGuard | `quay.io/evalhub/community-wildguard:latest` |
| IFBench | AllenAI IFBench | `quay.io/evalhub/community-ifbench:latest` |

## Testing locally before deploying

Expand Down
7 changes: 7 additions & 0 deletions src/content/docs/guides/collections.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -54,6 +54,13 @@ EvalHub ships with **system collections** (out-of-the-box) that are available to
| `reasoning-v1` | reasoning | 0.38 | 6 | Mathematical and logical reasoning |
| `coding-v1` | code | 0.25 | 1 | Code generation and understanding |
| `instruction-following-v1` | instruction_following | 0.50 | 5 | Instruction-following ability |
| `instruction-output-v1` | instruction_output | 0.325 | 2 | JSON schema compliance + IFBench OOD instruction constraints |
| `knowledge-reasoning-v1` | knowledge_reasoning | 0.235 | 2 | Knowledge and reasoning benchmarks |
| `document-understanding-v1` | document_understanding | 0.35 | 14 | Grounded document understanding (RULER, DocVQA, etc.) |
| `tool-use-v1` | tool_use | 0.60 | 1 | Tool use and function calling |
| `software-v1` | software | 0.20 | 3 | Software engineering benchmarks |
| `multimodal-v1` | multimodal | 0.50 | 1 | Multimodal understanding |
| `trustworthiness-v1` | trustworthiness | 0.325 | 2 | Trustworthiness and reliability |
| `safety-and-fairness-v1` | safety | 0.758 | 6 | Bias, toxicity, and safety |
| `model-validation` | safety | 0.75 | 1 | Security-focused validation (uses `lower_is_better`) |
| `toxicity-and-ethical-principles` | safety | 0.75 | 3 | Toxicity and ethical evaluation |
Expand Down
3 changes: 2 additions & 1 deletion src/content/docs/home.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -12,7 +12,7 @@ EvalHub provides a unified way to evaluate LLMs across multiple frameworks — s

<CardGrid>
<Card title="Multi-framework" icon="puzzle">
Evaluate with LightEval, GuideLLM, lm-eval-harness, Garak, MTEB, IBM CLEAR, DeepEval, Inspect AI, RAGAS, RULER, SWE-bench, WildGuard, or bring your own framework.
Evaluate with LightEval, GuideLLM, lm-eval-harness, Garak, MTEB, IBM CLEAR, DeepEval, Inspect AI, RAGAS, RULER, SWE-bench, WildGuard, IFBench, or bring your own framework.
</Card>
<Card title="Kubernetes-native" icon="rocket">
Each evaluation runs as an isolated Kubernetes Job with automatic lifecycle management.
Expand Down Expand Up @@ -65,6 +65,7 @@ Community providers with a `provider.yaml` are listed in the [Provider Catalog](
| **RULER** | `ruler` | Community | Long-context evaluation: needle-in-a-haystack, variable tracking, aggregation |
| **SWE-bench** | `swebench` | Community | Code generation: real-world GitHub issue resolution |
| **WildGuard** | `wildguard` | Community | Safety: prompt harmfulness classification and refusal detection |
| **IFBench** | `ifbench` | Community | Instruction following: 58 OOD verifiable constraints, prompt-level loose accuracy |

## Quick Taste

Expand Down
Loading