From 257c20f2ea881530a68f648cc9c1a52abf598552 Mon Sep 17 00:00:00 2001 From: Jooyeon Mok Date: Tue, 8 Sep 2026 18:52:41 -0400 Subject: [PATCH] docs: add IFBench adapter references and Phase 2 collection catalog - List IFBench on home and bring-your-own-framework adapter tables - Add Phase 2 curated collections to collections guide built-in table (including instruction-output-v1 with IFBench) Refs: eval-hub/eval-hub#942 Signed-off-by: Jooyeon Mok --- src/content/docs/guides/bring-your-own-framework.mdx | 1 + src/content/docs/guides/collections.mdx | 7 +++++++ src/content/docs/home.mdx | 3 ++- 3 files changed, 10 insertions(+), 1 deletion(-) diff --git a/src/content/docs/guides/bring-your-own-framework.mdx b/src/content/docs/guides/bring-your-own-framework.mdx index 6581c998..7daf7849 100644 --- a/src/content/docs/guides/bring-your-own-framework.mdx +++ b/src/content/docs/guides/bring-your-own-framework.mdx @@ -596,6 +596,7 @@ contains production adapters you can use as templates: | SWE-bench | SWE-bench | `quay.io/evalhub/community-swebench:latest` | | RULER | NVIDIA RULER | `quay.io/evalhub/community-ruler:latest` | | WildGuard | AllenAI WildGuard | `quay.io/evalhub/community-wildguard:latest` | +| IFBench | AllenAI IFBench | `quay.io/evalhub/community-ifbench:latest` | ## Testing locally before deploying diff --git a/src/content/docs/guides/collections.mdx b/src/content/docs/guides/collections.mdx index 7770052e..8fcc6a17 100644 --- a/src/content/docs/guides/collections.mdx +++ b/src/content/docs/guides/collections.mdx @@ -54,6 +54,13 @@ EvalHub ships with **system collections** (out-of-the-box) that are available to | `reasoning-v1` | reasoning | 0.38 | 6 | Mathematical and logical reasoning | | `coding-v1` | code | 0.25 | 1 | Code generation and understanding | | `instruction-following-v1` | instruction_following | 0.50 | 5 | Instruction-following ability | +| `instruction-output-v1` | instruction_output | 0.325 | 2 | JSON schema compliance + IFBench OOD instruction constraints | +| `knowledge-reasoning-v1` | knowledge_reasoning | 0.235 | 2 | Knowledge and reasoning benchmarks | +| `document-understanding-v1` | document_understanding | 0.35 | 14 | Grounded document understanding (RULER, DocVQA, etc.) | +| `tool-use-v1` | tool_use | 0.60 | 1 | Tool use and function calling | +| `software-v1` | software | 0.20 | 3 | Software engineering benchmarks | +| `multimodal-v1` | multimodal | 0.50 | 1 | Multimodal understanding | +| `trustworthiness-v1` | trustworthiness | 0.325 | 2 | Trustworthiness and reliability | | `safety-and-fairness-v1` | safety | 0.758 | 6 | Bias, toxicity, and safety | | `model-validation` | safety | 0.75 | 1 | Security-focused validation (uses `lower_is_better`) | | `toxicity-and-ethical-principles` | safety | 0.75 | 3 | Toxicity and ethical evaluation | diff --git a/src/content/docs/home.mdx b/src/content/docs/home.mdx index 675dd6ab..2eea8360 100644 --- a/src/content/docs/home.mdx +++ b/src/content/docs/home.mdx @@ -12,7 +12,7 @@ EvalHub provides a unified way to evaluate LLMs across multiple frameworks — s - Evaluate with LightEval, GuideLLM, lm-eval-harness, Garak, MTEB, IBM CLEAR, DeepEval, Inspect AI, RAGAS, RULER, SWE-bench, WildGuard, or bring your own framework. + Evaluate with LightEval, GuideLLM, lm-eval-harness, Garak, MTEB, IBM CLEAR, DeepEval, Inspect AI, RAGAS, RULER, SWE-bench, WildGuard, IFBench, or bring your own framework. Each evaluation runs as an isolated Kubernetes Job with automatic lifecycle management. @@ -65,6 +65,7 @@ Community providers with a `provider.yaml` are listed in the [Provider Catalog]( | **RULER** | `ruler` | Community | Long-context evaluation: needle-in-a-haystack, variable tracking, aggregation | | **SWE-bench** | `swebench` | Community | Code generation: real-world GitHub issue resolution | | **WildGuard** | `wildguard` | Community | Safety: prompt harmfulness classification and refusal detection | +| **IFBench** | `ifbench` | Community | Instruction following: 58 OOD verifiable constraints, prompt-level loose accuracy | ## Quick Taste