Skip to content

docs(tasks): [Experimental] Exploratory proposal and examples for EveryEvalEver (E3) integration - #1053

Merged
benjelloun merged 2 commits into
mainfrom
feat/tasks-e3-integration-and-examples
Oct 5, 2026
Merged

benjelloun merged 2 commits into
mainfrom
feat/tasks-e3-integration-and-examples

Conversation

@benjelloun

Copy link
Copy Markdown
Contributor

Summary (Experimental / Exploratory RFC)

Note: This PR shares early exploratory work on integrating EveryEvalEver (E3) (arXiv:2606.14516) and Croissant Tasks (CT) (arXiv:2605.29786) for discussion and feedback across the CT and E3 working groups. tasks/README.md is intentionally left untouched while these conventions are being explored.

What this explores

  1. Design Proposal (tasks/docs/e3_croissant_tasks_integration.md):
    • Investigates separating invariant E3 benchmark metadata into croissant:TaskProblem and run-specific logs into croissant:TaskSolution (or croissant:Task).
    • Explores reusing Croissant 1.1 constructs—cr:FileObject (application/jsonlines, schema:sha256), cr:RecordSet with nested cr:subField, sc:Enumeration recordsets for discrete grading rubrics, and cr:examples (@type: @json)—alongside schema:QuantitativeValue, W3C PROV-O, and OBO Foundry STATO for statistical uncertainty (standardError, confidenceInterval, sampleSize).
  2. Experimental Vocabulary Additions (tasks/croissant-tasks.ttl, tasks/croissant-tasks-shapes.ttl):
    • Tests adding croissant:lowerIsBetter (xsd:boolean, default false) and croissant:MetricSpec.
  3. Six Experimental Examples (tasks/examples/every_eval_ever/*.jsonld):
    • mmlu_pro_problem.jsonld, mmlu_pro_solution_kimi_k2.jsonld, mmlu_pro_combined_task.jsonld (CoT MCQ + regex extraction)
    • intercode_ctf_agentic_task.jsonld (Agentic Docker sandbox + bash/python tool calls)
    • vectara_hallucination_task.jsonld (croissant:lowerIsBetter: true + 7 domain croissant:subTask splits)
    • simpleqa_llm_judge_task.jsonld (LLM-as-a-Judge GPT-4o + sc:Enumeration rubric levels)

@benjelloun
benjelloun requested a review from a team as a code owner September 24, 2026 11:53
@github-actions

Copy link
Copy Markdown

MLCommons CLA bot All contributors have signed the MLCommons CLA ✍️ ✅

…ssant Tasks + E3 examples (#1052)

## Summary (Experimental Prototype)

Adds an experimental prototype **Croissant Tasks (CT) Visualizer**
(`tasks/index.html` and `tasks/examples/every_eval_ever/index.html`)
styled after the MLCommons Croissant 1.1 Dataset Visualizer
(`datasets/static/visualizer.js`) to help inspect and discuss the 6
exploratory E3 -> CT examples on GitHub Pages:

- **GitHub Pages Preview Routes**:
  - `https://docs.mlcommons.org/croissant/tasks/`
- `https://docs.mlcommons.org/croissant/tasks/examples/every_eval_ever/`
@benjelloun
benjelloun removed the request for review from leobianco October 2, 2026 15:53
@benjelloun
benjelloun merged commit ec24358 into main Oct 5, 2026
13 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants