Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
233 changes: 233 additions & 0 deletions .claude/skills/blendtutor-course/SKILL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,233 @@
---
name: blendtutor-course
description: Author a blendtutor course end to end — scaffold the course, write R or Python lessons with checks, solutions, hints, and success criteria, write a minimal eval suite, verify everything locally, then produce a Quarto snippet/page or a static browser site. Use when someone wants to create or extend a blendtutor course, turn a chapter or tutorial into interactive exercises, write eval cases for a lesson, or publish lessons to Quarto or GitHub Pages.
argument-hint: <source material (chapter file or URL)> [r|python] [quarto|site]
---

# blendtutor course authoring

Turn source material into blendtutor lessons that grade correctly and render well.
One lesson = one exercise = one `lessons/<id>.yaml` plus its `lessons/eval_<id>.yaml`.

## 0. Preconditions

- `blendtutor --version` is **0.2.0 or newer** (older versions drop `packages`,
`gotchas`, and `success_criteria` from Quarto exports). Install from
<https://github.com/mcmullarkey/blendtutor/releases>.
- R lessons: `Rscript` on PATH. Python lessons: `uv` on PATH.
- Quarto output: Quarto >= 1.4.
- `run` and `eval` make **paid LLM calls** and need `FIREWORKS_API_KEY` (or
`ANTHROPIC_API_KEY`) in the environment. Never write a key into a file you commit.

## 1. Gather

Ask only for what you cannot infer:

- **Source material**: read the chapter or tutorial; pick the one skill each exercise practices.
- **Language**: `r` or `python`, per lesson.
- **Output**: Quarto (snippet into an existing book/page, or a standalone page) and/or a static site.
- **Where the course lives**: an existing directory holding `blendtutor.toml`, or a new one.

## 2. Scaffold

```bash
blendtutor init <course-dir> # only if no blendtutor.toml exists yet
cd <course-dir>
blendtutor new lesson --lang <r|python> <id> # writes lessons/<id>.yaml + lessons/eval_<id>.yaml
```

`new` registers the lesson in `blendtutor.toml`. Paths are one argument:
`lessons/<id>.yaml`, never `lessons <id>.yaml`.

## 3. Write the lesson

Replace the scaffold's hello-world content. Keep every field short.

```yaml
lesson_name: "<id>"
language: Python # or R
description: "<one line: the skill practiced>"
textbook_reference: "<book - chapter>" # optional

exercise:
type: "function_writing"
prompt: |
<2-3 sentences. Name every variable or function the checks rely on.>
code_template: |
<inline data + comment scaffolding the learner fills in>
solution: |
<complete answer; must pass every check>
hints: |
- <one bullet per hint; every non-empty line starts with "- ">
gotchas: | # optional: common mistakes, also bullets
- <a pitfall learners hit>
success_criteria: |
- <what a correct answer does, including what checks cannot see>
llm_evaluation_prompt: |
You are grading a beginner exercise on <skill>.

The student submitted this code:
{student_code}

<What makes it correct, and what makes it incorrect even if it runs.>
Call respond_with_feedback with two or three encouraging sentences.

checks: # top level, not inside exercise
- "assert <expression about the learner's objects>" # R: "stopifnot(<expression>)"

packages: # only if needed
- pandas
```

Rules that keep lessons working in the browser:

- **Inline the data.** Build a small data frame in `code_template` instead of
loading a dataset package; webR and Pyodide may not have it. Choose values
whose correct answer is easy to assert (for example, averages like 48.0 and 39.0).
- **Checks assert results, not source text.** They run after the learner's code
in the same session and fail if they raise. Any lesson with at least one check
gets a Check button.
- **No checks for prose-like exercises** (pseudocode, explanations): there is
nothing to assert. The grader, driven by `success_criteria`, does the work.
- **`success_criteria` reaches the grader** (0.2.0+). Put the lesson's real
point there, especially what a check cannot verify (structure, style, naming).
- **`hints` and `gotchas` are bullet lists**, or `validate` rejects them.
- **`packages`** names cannot contain quotes, commas, or spaces.

## 4. Write the minimal eval

`lessons/eval_<id>.yaml`, next to the lesson. Four cases, each with a comment:

```yaml
# Eval suite for <id>.yaml: 2 correct, 2 incorrect.
cases:
# Correct: the canonical solution
- submission: |-
<solution>
expected: correct
# Correct: different names or methods, same intent
- submission: |-
<alternative that still meets every success criterion>
expected: correct
# Incorrect: near-miss that passes the checks but misses the lesson's point
- submission: |-
<runs cleanly, right output, violates a success criterion>
expected: incorrect
# Incorrect: a realistic mistake the checks catch
- submission: |-
<wrong result, e.g. reversed sort or missing step>
expected: incorrect
```

The near-miss case is what measures the grading prompt; do not skip it. For
lessons without checks, make both incorrect cases content mistakes (a missing
step, the wrong operation).

## 5. Verify locally (no paid calls)

```bash
blendtutor validate lessons/<id>.yaml
blendtutor list .
```

Then run the solution and every eval case against the checks. Python:

```bash
uv run --no-project --with pyyaml --with <packages> python - <<'PY'
import contextlib, io, yaml
lesson = yaml.safe_load(open("lessons/<id>.yaml"))
cases = yaml.safe_load(open("lessons/eval_<id>.yaml"))["cases"]
def run(label, code):
env = {}
with contextlib.redirect_stdout(io.StringIO()):
exec(code, env)
results = []
for check in lesson.get("checks", []):
try:
exec(check, env); results.append("pass")
except Exception as e:
results.append(f"fail({type(e).__name__})")
print(label, results)
run("solution", lesson["exercise"]["solution"])
for i, case in enumerate(cases, 1):
run(f"case {i} ({case['expected']})", case["submission"])
PY
```

R (pass the lesson path; the eval file is found next to it):

```bash
uv run --no-project --with pyyaml python - lessons/<id>.yaml <<'PY'
import pathlib, subprocess, sys, tempfile, yaml
lesson_path = pathlib.Path(sys.argv[1])
lesson = yaml.safe_load(lesson_path.read_text())
cases = yaml.safe_load(lesson_path.with_name("eval_" + lesson_path.name).read_text())["cases"]
def run(label, code):
checks = "\n".join(
f'r <- c(r, tryCatch({{ {c} ; "pass" }}, error = function(e) "fail"))' for c in lesson.get("checks", [])
)
script = f"r <- character()\ninvisible(capture.output({{\n{code}\n}}))\n{checks}\ncat(r)\n"
with tempfile.NamedTemporaryFile("w", suffix=".R", delete=False) as f:
f.write(script)
out = subprocess.run(["Rscript", f.name], capture_output=True, text=True)
pathlib.Path(f.name).unlink(missing_ok=True)
print(label, out.stdout.strip() or f"submission error: {out.stderr.strip().splitlines()[-1:]}")
run("solution", lesson["exercise"]["solution"])
for i, case in enumerate(cases, 1):
run(f"case {i} ({case['expected']})", case["submission"])
PY
```

Expect: solution passes; correct cases pass; the near-miss passes (only the
grader can reject it); the checks-catch case fails. Fix the lesson if not.

Only then, **after confirming with the user** (paid calls):

```bash
blendtutor eval lessons/<id>.yaml # grader accuracy against expected verdicts
blendtutor eval lessons/<id>.yaml --write-report # also saves eval-report.json for the site
```

If the grader misjudges a case, tighten `success_criteria` and the
`llm_evaluation_prompt`, then re-run `eval`.

## 6. Publish

### Quarto

```bash
blendtutor export-quarto lessons/<id>.yaml # snippet: paste into an existing .qmd
blendtutor export-quarto --document lessons/<id>.yaml # standalone page with front matter
blendtutor export-quarto --key-page > api-key.qmd # optional API key page
```

Project setup, once:

1. Run `quarto add mcmullarkey/blendtutor` **from the folder that contains
`_quarto.yml`** (or the standalone `.qmd`). Installed anywhere else, the
filter never loads and exercises render as plain text.
2. Enable the filter with `filters: [mcmullarkey/blendtutor]` in `_quarto.yml`
**or** in the page's front matter, not both. `--document` pages already declare it.
3. Commit `_extensions/` so CI renders (for example, a GitHub Pages workflow).
4. Preview with `quarto preview`. Opening HTML via `file://` blocks the widget's JavaScript.

R runs in `type: book` projects without cross-origin isolation, on webR's
slower channel; `coi: true` speeds up standalone pages only.

### Static site

```bash
blendtutor build . --target pyodide -o site # Python course
blendtutor build . --target webr -o site # R course
```

One target per build, so keep R and Python lessons in separate courses. The
site deploys to GitHub Pages as-is. `--password` encrypts it; `--embed-key`
(requires `--password`) puts a real API key inside the encrypted site, where
anyone with the password can use it. Mention that risk before suggesting it.

## 7. Report

Tell the user which files you created, the local check results per case, whether
paid evals ran (and their accuracy), and the exact command or snippet for their
chosen output.
2 changes: 1 addition & 1 deletion .github/workflows/docs.yml
Original file line number Diff line number Diff line change
Expand Up @@ -112,7 +112,7 @@ jobs:
run: |
if [ -d docs/evals ]; then
if rg -l '/Users/' docs/evals/ | grep -q .; then
echo "docs.yml: /Users/ leak in docs/evals/ (scrub per creating-lessons.md Step 9)" >&2
echo "docs.yml: /Users/ leak in docs/evals/ (scrub per whole-game.md Eval report)" >&2
exit 1
fi
rm -rf docs/book/book/evals
Expand Down
4 changes: 3 additions & 1 deletion .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -5,7 +5,9 @@
# output under docs/evals/ (deliberately committed) nor Rust eval fixtures
# under crates/*/tests/fixtures/evals/.
**/.smevals/
.claude/
# Local Claude state stays ignored; project skills under .claude/skills/ are shared.
.claude/*
!.claude/skills/
rv/library/

# Rust
Expand Down
96 changes: 17 additions & 79 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -57,93 +57,31 @@ don't need it). Two example courses are deployed alongside the docs:
- **[R example site (webR)](https://mcmullarkey.github.io/blendtutor/examples/r/)**
- **[Python example site (Pyodide)](https://mcmullarkey.github.io/blendtutor/examples/python/)**

## Author a course with Claude Code

This repo ships a Claude Code skill,
[`blendtutor-course`](.claude/skills/blendtutor-course/SKILL.md). In a clone,
run `/blendtutor-course <chapter file or URL>` (or ask Claude to build a
course): it scaffolds lessons with checks and solutions, writes a minimal eval
suite, verifies the checks locally, and exports a Quarto snippet or builds a site.

## Quarto extension

blendtutor also ships as a [Quarto](https://quarto.org) extension for
interactive coding exercises in `.qmd` documents — in-browser editor, instant
checks, solution reveal, AI hints, all static HTML. Requires **Quarto >= 1.4**:
Interactive exercises in `.qmd` documents (Quarto >= 1.4, extension version
0.2.0). Run this from the folder that contains `_quarto.yml`:

```bash
quarto add mcmullarkey/blendtutor
```

Installs to `_extensions/mcmullarkey/blendtutor/` (version 0.2.0). **Run it from the
folder that contains `_quarto.yml`** (or the `.qmd`): Quarto only discovers `_extensions/`
there, so one installed a directory up never loads; assets are install-path-independent.

#### Quick start (zero hand-written bootstrap)

A complete copy-paste document — zero hand-written bootstrap. Filter by name,
`.blendtutor` div, render:

````markdown
---
title: "My exercises"
filters: [mcmullarkey/blendtutor]
---

::: {.blendtutor language="r"}
Write a function `add(a, b)` that returns the sum.

```r
add <- function(a, b) { ___ }
```
:::
````

Render, open in a browser — interactive immediately. Grade submissions with a
`{.r .checks}` block (`stopifnot(add(1, 2) == 3)`); Python: same div, `language="python"`.

Optional blocks inside the div add a `{.r .solution}`, `::: {.hints}` / `::: {.gotchas}`
bullets, and a `::: {.success-criteria}` rubric for AI feedback; `packages="dplyr"` on the
div preloads packages. `blendtutor export-quarto lesson.yaml` writes the div from a lesson
(`--document` for a full page, `--key-page` for the API key page).

#### Auto-bootstrap opt-out

The filter auto-bootstraps by default; to wire up the runtime yourself, set
`bt-auto-bootstrap: false` in the YAML header. To keep it but disable the
auto-mounted AI feedback, set `bt-feedback: false` — see
[BYOK](#byok-bring-your-own-key).

### Cross-origin isolation (COI)

webR runs faster with `SharedArrayBuffer`, which needs cross-origin isolation
(COOP/COEP). Opt in with `coi: true` (page YAML header) or `coi="true"` (any div);
the filter injects a service-worker shim. Pyodide-only pages do not need COI.

> **Book-mode limitation:** COI does not function in Quarto `type: book`
> projects — the shim's scope cannot cover the book's `_output/` pages, so webR
> uses its slower non-isolated channel ([ADR-0015](docs/adr/0015-opt-in-coi-cross-origin.md)).

### Demo book

A complete demo book with R and Python exercises lives in
[`demo-book/`](demo-book/), rendered live at
<https://mcmullarkey.github.io/blendtutor/demo-book/> (rebuild locally with
`cd demo-book && quarto render`). It is a Quarto `type: book` project, so
COI does not take effect (limitation above). Python exercises are fully interactive
and every page ships a static fallback. R exercises run in the book too, on webR's slower
fallback channel; the CLI-built [example sites](#deploy-to-github-pages) add isolation,
R exercises run interactively via webR there. Over `file://` you get static exercise content only; serve over HTTP:

```bash
cd demo-book/_output && python3 -m http.server 8000
```
[The whole game](https://mcmullarkey.github.io/blendtutor/whole-game.html#quarto-extension) covers the rest:

## BYOK (Bring Your Own Key)

Browser feedback uses the learner's own API key — no server-side key. Feedback
is **auto-mounted**: the injected bootstrap imports `exercise-feedback.js` and
calls `mountAllFeedback(registry)` after the runtime starts. The key is entered
once on the API key page (the demo book ships one) and shared via `localStorage`
— readable by any JavaScript on the page's origin, so never reuse a critical
key; it is sent only to `api.fireworks.ai`. BYOK is Fireworks-only (pinned model
`accounts/fireworks/models/deepseek-v4-flash-0731`); the CLI supports other
providers (see [API key](#api-key)). Serve over HTTP — `file://` breaks
`localStorage` sharing and blocks ES modules, so feedback never mounts.
Self-hosted CSP: add `connect-src https://api.fireworks.ai` (Pages cannot set
CSP headers; the shim covers only COOP/COEP).
- [Quick start](https://mcmullarkey.github.io/blendtutor/whole-game.html#quick-start-zero-hand-written-bootstrap)
- [Export a lesson](https://mcmullarkey.github.io/blendtutor/whole-game.html#export-a-lesson) with `blendtutor export-quarto`
- [Auto-bootstrap opt-out](https://mcmullarkey.github.io/blendtutor/whole-game.html#auto-bootstrap-opt-out)
- [Cross-origin isolation](https://mcmullarkey.github.io/blendtutor/whole-game.html#cross-origin-isolation-coi) and R in book projects
- [Demo book](https://mcmullarkey.github.io/blendtutor/whole-game.html#demo-book)
- [BYOK feedback](https://mcmullarkey.github.io/blendtutor/whole-game.html#byok-bring-your-own-key)

## License

Expand Down
1 change: 0 additions & 1 deletion docs/book/src/SUMMARY.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,6 @@
[Introduction](./introduction.md)

- [The whole game](./whole-game.md)
- [Creating Lessons](./creating-lessons.md)
- [Architecture](./architecture.md)
- [Example sites](./examples.md)
- [API reference](./api-reference.md)
Loading
Loading