Skip to content

Externalize prompts and add llama startup warmup - #7

Merged
vbergeron merged 1 commit into
mainfrom
claude/externalize-prompts-and-llama-warmup
Aug 6, 2026
Merged

vbergeron merged 1 commit into
mainfrom
claude/externalize-prompts-and-llama-warmup

Conversation

@vbergeron

Copy link
Copy Markdown
Owner

Prompt refactor

  • Move all four step prompts (gate, extract, goals, respond) and their few-shot examples out of Prompt module source into priv/prompts/:
    priv/prompts/.txt — system instruction prose
    priv/prompts/_shots.chatml — raw ChatML few-shot turns
  • Load both files at compile time via @external_resource + File.read!/1
    so mix compile picks up edits without touching .ex source.
  • Expand shot coverage across all four steps:
    • gate: 10 shots covering rhetorical assertions, back-references,
      imperatives-as-queries, social/chitchat disambiguation
    • extract: 11 patterns including graph/edge notation and transitivity
    • goals: 7 patterns covering polar/wh/attribute, superlatives,
      multi-condition decomposition, set-difference
    • respond: 17 evidence shapes (learned fact/rule/constraint/enum,
      true/false/bound queries, unanswered, contradiction, mixed, empty)

Llama server

  • Add --mlock to the llama-server argv to pin model pages in RAM after the first warm-up inference pages them in (requires ulimit -l).
  • Send :warmup to self once the server transitions to :ready; a Task fires a 1-token completion so the page-fault storm is paid at startup rather than on the first user turn. Failure is non-fatal (logged as warning).

Prompt refactor
- Move all four step prompts (gate, extract, goals, respond) and their
  few-shot examples out of Prompt module source into priv/prompts/:
    priv/prompts/<step>.txt          — system instruction prose
    priv/prompts/<step>_shots.chatml — raw ChatML few-shot turns
- Load both files at compile time via @external_resource + File.read!/1
  so mix compile picks up edits without touching .ex source.
- Expand shot coverage across all four steps:
  - gate: 10 shots covering rhetorical assertions, back-references,
    imperatives-as-queries, social/chitchat disambiguation
  - extract: 11 patterns including graph/edge notation and transitivity
  - goals: 7 patterns covering polar/wh/attribute, superlatives,
    multi-condition decomposition, set-difference
  - respond: 17 evidence shapes (learned fact/rule/constraint/enum,
    true/false/bound queries, unanswered, contradiction, mixed, empty)

Llama server
- Add --mlock to the llama-server argv to pin model pages in RAM after
  the first warm-up inference pages them in (requires ulimit -l).
- Send :warmup to self once the server transitions to :ready; a Task
  fires a 1-token completion so the page-fault storm is paid at startup
  rather than on the first user turn. Failure is non-fatal (logged as
  warning).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
@vbergeron
vbergeron merged commit d38022a into main Aug 6, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants