diff --git a/.changeset/ui-snapshot.md b/.changeset/ui-snapshot.md new file mode 100644 index 0000000..ea3a0ce --- /dev/null +++ b/.changeset/ui-snapshot.md @@ -0,0 +1,24 @@ +--- +"macos-vision": minor +--- + +feat: `uiSnapshot()` — the accessibility tree plus the text it misses + +Captures a window once, walks its accessibility tree, runs OCR over the same +region, and reports every piece of visible text that no node accounts for. + +`unresolved` does double duty. It completes the box model for anything +custom-drawn — canvas, WebGL, games, images with text baked in — where AX is +blind. And each entry is an accessibility gap in the app under test: +`coveredByNode` present means a control is there but unlabelled, absent means +nothing is exposed at all. + +`summary.axTextCoverage` is `null` whenever the walk was capped, with +`cappedWalk: true` alongside. A capped walk measures how much of the tree was +visited rather than how accessible the app is — measured on one Safari window the +figure reads 0.34 at `maxElements: 200` against 0.83 for the complete walk, and +reporting the former as coverage would blame the app for our own budget. + +The capture is pinned to exactly the window being walked rather than letting the +capture and the tree each resolve a target independently, so colours and OCR +always describe the same region as the geometry. diff --git a/README.md b/README.md index 92af843..cf929cd 100644 --- a/README.md +++ b/README.md @@ -339,6 +339,34 @@ window. See [`docs/BOX-MODEL.md`](docs/BOX-MODEL.md) for the measurements behind these numbers. +### `uiSnapshot()` — the tree plus what it misses + +`axTree()` sees only what the app exposes. `uiSnapshot()` captures once, walks the +tree, runs OCR over the same window, and reports the text Vision can read that no +node accounts for: + +```ts +const snap = await uiSnapshot({ app: 'MyApp' }); +// { ...axTree fields, unresolved: [...], summary: {...} } + +snap.summary; +// { nodes: 491, labelled: 402, ocrBlocks: 123, unresolved: 21, axTextCoverage: 0.83 } + +snap.unresolved[0]; +// { text: "Sprzedaż Q4", box: [420, 300, 88, 16], confidence: 0.98, coveredByNode: 17 } +``` + +That list does double duty. It completes the picture for anything custom-drawn — +canvas, WebGL, games, images with text baked in — where AX is simply blind. And +every entry is an accessibility gap in the app under test: `coveredByNode` present +means a control is there but unlabelled, absent means nothing is exposed at all. + +`axTextCoverage` is **`null` when the walk was capped**, with `cappedWalk: true` +alongside it. A capped walk measures how much of the tree was visited, not how +accessible the app is — on one Safari window the figure reads 0.34 at +`maxElements: 200` against 0.83 for the complete walk, and publishing the former +as coverage would blame the app for our own budget. + --- ## API — Markdown pipeline (VisionScribe) diff --git a/docs/BOX-MODEL.md b/docs/BOX-MODEL.md index 881f3f6..6a3d3df 100644 --- a/docs/BOX-MODEL.md +++ b/docs/BOX-MODEL.md @@ -4,7 +4,7 @@ Goal: hand an LLM a compact JSON description of what is on screen — element bo colours, borders, typography — complete enough that the model can reconstruct the layout, review it, or write assertions against it, **without ever seeing the screenshot**. -Status: **phases 1–3 implemented** as `axTree()` (see the README). Phase 4 — merging with OCR to populate `unresolved` — is still open. Every number below was measured on this machine +Status: **implemented** — phases 1–3 as `axTree()`, phase 4 as `uiSnapshot()` (see the README). Every number below was measured on this machine (Apple M1 Pro, 16 GB, macOS 26.5.2) with throwaway probes, not estimated. --- @@ -178,6 +178,11 @@ Built as `ax-helper` + `src/ax.ts`. Findings that only appeared once it ran: - **Swift omits `nil` rather than encoding `null`**, so the root node has no `parent` key at all. The TypeScript type said `number | null`; a consumer checking `=== null` would have been wrong. Caught by a test, fixed in the type. +- **The coverage metric lied when the walk was capped.** `axTextCoverage` read + 0.34 on a Safari window at `maxElements: 200` and 0.83 for the same window + walked completely — the first number measures our budget, not the app's + accessibility, but it reads like a verdict on the app. It is now `null` + whenever `budget.capped` is true. - **`AXAttributedStringForRange` is app-dependent.** TextEdit returns `Menlo-Regular` 11 pt; Safari's web `StaticText` returns alignment but no font. That is a limit of the source, not of the reader. diff --git a/src/index.ts b/src/index.ts index a466f06..7743b8a 100644 --- a/src/index.ts +++ b/src/index.ts @@ -3,6 +3,11 @@ import { open, readFile, writeFile, mkdir } from 'fs/promises'; import { homedir } from 'os'; import { VISION_BIN, PDF_BIN, execHelper, runHelper, fileSha256, sha256 } from './helper.js'; import { textOptionArgs, visionCapabilities } from './vision.js'; +import { axTree } from './ax.js'; +import type { AxTreeOptions } from './ax.js'; +import { captureScreen, listWindows } from './ui.js'; +import { mergeOcrWithTree } from './snapshot.js'; +import type { UiSnapshot } from './snapshot.js'; import type { TextRecognitionOptions } from './vision.js'; const BINARY_TIMEOUT_MS = 30_000; @@ -356,6 +361,70 @@ export type { export { inferLayout, sortBlocksByReadingOrder } from './layout.js'; // ─── Markdown pipeline (VisionScribe) ────────────────────────────────────────── +// ─── UI snapshot: accessibility tree + what OCR can see ────────────────────── + +export interface UiSnapshotOptions extends Omit { + /** Sample colours from the capture this function takes anyway. Default true. */ + colors?: boolean; + /** + * Run OCR and report text the tree does not account for. Default true — it is + * what makes the snapshot complete, and the gap is itself a finding. Costs + * roughly a second on a full window. + */ + ocr?: boolean; + /** Ignore OCR text below this confidence when deciding coverage. Default 0.3. */ + minConfidence?: number; +} + +/** + * The layout of a running application as one object: accessibility nodes with + * exact boxes, optional colours and typography, plus the text Vision can see + * that the tree has no node for. + * + * That last part matters twice over. It completes the picture for anything + * custom-drawn — canvas, WebGL, games, images with baked-in labels, where AX is + * blind — and each entry is an accessibility gap in the app under test. + * `summary.axTextCoverage` is the share of visible text the tree accounts for. + */ +export async function uiSnapshot(options: UiSnapshotOptions = {}): Promise { + const { colors = true, ocr: wantOcr = true, minConfidence, ...treeOptions } = options; + // The capture must cover exactly the window axTree walks, or colours and OCR + // would describe a different region than the geometry. Resolve the window + // explicitly rather than letting the two calls each pick their own. + const windowIndex = treeOptions.window ?? 0; + let captureOptions: { app?: string; windowId?: number }; + if (treeOptions.app && windowIndex === 0) { + captureOptions = { app: treeOptions.app }; + } else { + const q = treeOptions.app?.toLowerCase(); + const candidates = (await listWindows()).filter((w) => + q ? w.app.toLowerCase() === q || w.app.toLowerCase().startsWith(q) : w.pid === treeOptions.pid + ); + const win = candidates[windowIndex]; + if (!win) { + const what = treeOptions.app ?? `pid ${treeOptions.pid}`; + throw new Error( + `no on-screen window ${windowIndex} for ${what} (found ${candidates.length})` + ); + } + captureOptions = { windowId: win.windowId }; + } + const capture = await captureScreen(captureOptions); + const tree = await axTree({ + ...treeOptions, + ...(colors ? { colors: { path: capture.path, frame: capture.frame } } : {}), + }); + if (!wantOcr) { + const { unresolved, summary } = mergeOcrWithTree(tree, [], capture.frame, minConfidence); + return { ...tree, unresolved, summary }; + } + const blocks = await ocr(capture.path, { format: 'blocks' }); + const { unresolved, summary } = mergeOcrWithTree(tree, blocks, capture.frame, minConfidence); + return { ...tree, unresolved, summary }; +} + +export { mergeOcrWithTree, blockToScreenBox } from './snapshot.js'; +export type { UiSnapshot, UnresolvedText, SnapshotSummary } from './snapshot.js'; export { axTree } from './ax.js'; export type { AxTree, diff --git a/src/snapshot.ts b/src/snapshot.ts new file mode 100644 index 0000000..e709d15 --- /dev/null +++ b/src/snapshot.ts @@ -0,0 +1,169 @@ +// Merging the accessibility tree with what OCR can see. +// +// The two disagree in a way that is itself informative: text Vision reads but +// AX has no node for is, almost always, text the app failed to expose to +// assistive technology. Canvas, WebGL, custom-drawn controls and images with +// baked-in labels all land here — so `unresolved` is both a completeness fix +// for the box model and an accessibility finding. +// +// Pure functions only: no capture, no OCR, no helper spawning. The composition +// that calls these lives in index.ts, where `ocr` is defined. + +import type { AxNode, AxTree, AxBox } from './ax.js'; +import type { ScreenFrame } from './ui.js'; +import type { VisionBlock } from './index.js'; + +/** Text Vision read that no accessibility node accounts for. */ +export interface UnresolvedText { + text: string; + /** `[x, y, w, h]` in global screen points, converted from the image. */ + box: AxBox; + confidence: number; + /** + * An AX node covers this area but exposes no matching text — the control is + * there, its label is not. Distinct from no node at all, which usually means + * custom drawing. + */ + coveredByNode?: number; +} + +export interface SnapshotSummary { + nodes: number; + /** Nodes carrying a label or value. */ + labelled: number; + ocrBlocks: number; + /** OCR text with no matching AX node — the accessibility gap. */ + unresolved: number; + /** + * Share of visible text the tree accounts for, 0–1 — and `null` when the walk + * was capped, because then it measures how much of the tree we looked at + * rather than how accessible the app is. Measured on one Safari window: 0.34 + * at `maxElements: 200` against 0.83 for the complete walk. Reporting the + * former as a coverage figure would accuse the app of a fault that is ours. + */ + axTextCoverage: number | null; + /** True when `unresolved` is inflated because the walk did not finish. */ + cappedWalk?: boolean; +} + +export interface UiSnapshot extends AxTree { + unresolved: UnresolvedText[]; + summary: SnapshotSummary; +} + +/** + * Same normalization the matching layer uses: fold case, whitespace and unicode + * punctuation. The character classes use \u escapes deliberately — written as + * literal characters, a formatter can rewrite an invisible one and silently turn + * the whitespace class into the range U+0020-U+200B, which swallows all of ASCII. + */ +export function normalize(s: string): string { + return s + .normalize('NFC') + .replace(/[\u2018\u2019\u201A\u2032]/g, "'") + .replace(/[\u201C\u201D\u201E\u2033]/g, '"') + .replace(/[\u2010-\u2015\u2212]/g, '-') + .replace(/[\u00A0\u1680\u2000-\u200B\u202F\u205F\u3000]/g, ' ') + .replace(/\s+/g, ' ') + .trim() + .toLowerCase(); +} + +/** Normalized 0–1 image coordinates → global screen points. */ +export function blockToScreenBox(block: VisionBlock, frame: ScreenFrame): AxBox { + return [ + Math.round(frame.x + block.x * frame.w), + Math.round(frame.y + block.y * frame.h), + Math.round(block.width * frame.w), + Math.round(block.height * frame.h), + ]; +} + +function centre(b: AxBox): [number, number] { + return [b[0] + b[2] / 2, b[1] + b[3] / 2]; +} + +function contains(outer: AxBox, point: [number, number]): boolean { + return ( + point[0] >= outer[0] && + point[0] <= outer[0] + outer[2] && + point[1] >= outer[1] && + point[1] <= outer[1] + outer[3] + ); +} + +/** Text a node offers for matching: its label and value together. */ +function nodeText(n: AxNode): string { + return normalize([n.label, n.value].filter(Boolean).join(' ')); +} + +/** + * Decides which OCR text the accessibility tree already accounts for. + * + * A block counts as accounted for when some node's box contains its centre and + * that node's label or value contains the text (or vice versa — AX labels are + * often longer than the visible run, and OCR often splits a label across lines). + * Everything else is reported as unresolved, with `coveredByNode` set when a node + * covers the area but says nothing matching, since "unlabelled control" and + * "custom-drawn text" are different problems. + */ +export function mergeOcrWithTree( + tree: AxTree, + blocks: VisionBlock[], + frame: ScreenFrame, + minConfidence = 0.3 +): { unresolved: UnresolvedText[]; summary: SnapshotSummary } { + const texts = tree.nodes.map((n) => ({ node: n, text: nodeText(n) })); + const unresolved: UnresolvedText[] = []; + let considered = 0; + + for (const block of blocks) { + if (block.confidence < minConfidence) continue; + const t = normalize(block.text); + if (!t) continue; + considered++; + + const box = blockToScreenBox(block, frame); + const c = centre(box); + + let covering: AxNode | undefined; + let matched = false; + for (const { node, text } of texts) { + if (!contains(node.box, c)) continue; + // Smallest covering node wins as the "should have said this" candidate. + if (!covering || node.box[2] * node.box[3] < covering.box[2] * covering.box[3]) { + covering = node; + } + if (text && (text.includes(t) || t.includes(text))) { + matched = true; + break; + } + } + if (matched) continue; + + unresolved.push({ + text: block.text, + box, + confidence: block.confidence, + ...(covering ? { coveredByNode: covering.id } : {}), + }); + } + + const labelled = tree.nodes.filter((n) => n.label || n.value).length; + const capped = tree.budget.capped; + return { + unresolved, + summary: { + nodes: tree.nodes.length, + labelled, + ocrBlocks: considered, + unresolved: unresolved.length, + axTextCoverage: capped + ? null + : considered === 0 + ? 1 + : Math.round((1 - unresolved.length / considered) * 100) / 100, + ...(capped ? { cappedWalk: true } : {}), + }, + }; +} diff --git a/test/snapshot.test.ts b/test/snapshot.test.ts new file mode 100644 index 0000000..42cb9e5 --- /dev/null +++ b/test/snapshot.test.ts @@ -0,0 +1,146 @@ +import { describe, it, expect } from 'vitest'; +import { mergeOcrWithTree, blockToScreenBox, normalize } from '../src/snapshot.js'; +import type { AxTree, AxNode } from '../src/ax.js'; +import type { VisionBlock } from '../src/index.js'; + +const frame = { x: 0, y: 0, w: 1000, h: 500 }; + +const node = (over: Partial & Pick): AxNode => ({ + depth: 1, + role: 'StaticText', + ...over, +}); + +const tree = (nodes: AxNode[], capped = false): AxTree => ({ + app: 'Test', + pid: 1, + source: 'ax', + budget: { + elements: nodes.length, + capped, + maxElements: 100, + maxDepth: 40, + elapsedMs: 1, + culled: 0, + }, + nodes, +}); + +/** A block occupying the given normalized rect, with the given text. */ +const block = ( + text: string, + x: number, + y: number, + over: Partial = {} +): VisionBlock => ({ + text, + x, + y, + width: 0.1, + height: 0.04, + confidence: 0.95, + ...over, +}); + +describe('normalize', () => { + it('preserves ASCII — the whitespace class must not swallow it', () => { + expect(normalize('80 Artifacts')).toBe('80 artifacts'); + for (let c = 0x21; c <= 0x7e; c++) { + expect(normalize(String.fromCharCode(c))).not.toBe(''); + } + }); + + it('folds unicode spaces, dashes and quotes', () => { + expect(normalize('Zapisz plik')).toBe('zapisz plik'); + expect(normalize('e‑mail')).toBe('e-mail'); + expect(normalize('“Cytat”')).toBe('"cytat"'); + }); +}); + +describe('blockToScreenBox', () => { + it('maps normalized image coordinates onto screen points', () => { + expect(blockToScreenBox(block('x', 0.5, 0.5), frame)).toEqual([500, 250, 100, 20]); + }); + + it('offsets by the capture frame origin', () => { + expect(blockToScreenBox(block('x', 0, 0), { x: 100, y: 50, w: 1000, h: 500 })).toEqual([ + 100, 50, 100, 20, + ]); + }); +}); + +describe('mergeOcrWithTree', () => { + it('treats text as accounted for when a covering node says the same thing', () => { + const t = tree([node({ id: 1, box: [400, 200, 200, 100], label: 'Zapisz' })]); + const { unresolved, summary } = mergeOcrWithTree(t, [block('Zapisz', 0.45, 0.45)], frame); + expect(unresolved).toHaveLength(0); + expect(summary.axTextCoverage).toBe(1); + }); + + it('matches when the AX label is longer than the visible run', () => { + const t = tree([node({ id: 1, box: [0, 0, 1000, 500], label: 'Zapisz zmiany w dokumencie' })]); + const { unresolved } = mergeOcrWithTree(t, [block('Zapisz zmiany', 0.4, 0.4)], frame); + expect(unresolved).toHaveLength(0); + }); + + it('reports text no node covers at all', () => { + const t = tree([node({ id: 1, box: [0, 0, 10, 10], label: 'gdzie indziej' })]); + const { unresolved, summary } = mergeOcrWithTree(t, [block('Kanwa', 0.5, 0.5)], frame); + expect(unresolved).toHaveLength(1); + expect(unresolved[0].text).toBe('Kanwa'); + expect(unresolved[0].coveredByNode).toBeUndefined(); + expect(summary.axTextCoverage).toBe(0); + }); + + it('names the covering node when one exists but exposes no matching text', () => { + const t = tree([node({ id: 7, box: [0, 0, 1000, 500], role: 'Group' })]); + const { unresolved } = mergeOcrWithTree(t, [block('Nieopisany przycisk', 0.5, 0.5)], frame); + expect(unresolved).toHaveLength(1); + expect(unresolved[0].coveredByNode).toBe(7); + }); + + it('prefers the smallest covering node as the candidate', () => { + const t = tree([ + node({ id: 1, box: [0, 0, 1000, 500], role: 'Window' }), + node({ id: 2, box: [450, 220, 120, 60], role: 'Button' }), + ]); + const { unresolved } = mergeOcrWithTree(t, [block('OK', 0.5, 0.5)], frame); + expect(unresolved[0].coveredByNode).toBe(2); + }); + + it('ignores OCR noise below the confidence floor', () => { + const t = tree([]); + const { unresolved, summary } = mergeOcrWithTree( + t, + [block('szum', 0.5, 0.5, { confidence: 0.1 })], + frame + ); + expect(unresolved).toHaveLength(0); + expect(summary.ocrBlocks).toBe(0); + }); + + it('withholds a coverage figure when the walk was capped', () => { + // Otherwise the number measures how much of the tree we looked at, and reads + // as an accusation that the app is inaccessible. + const t = tree([node({ id: 1, box: [0, 0, 10, 10] })], true); + const { summary } = mergeOcrWithTree(t, [block('cokolwiek', 0.5, 0.5)], frame); + expect(summary.axTextCoverage).toBeNull(); + expect(summary.cappedWalk).toBe(true); + }); + + it('reports full coverage rather than dividing by zero on an empty screen', () => { + const { summary } = mergeOcrWithTree(tree([]), [], frame); + expect(summary.axTextCoverage).toBe(1); + }); + + it('counts labelled nodes for the summary', () => { + const t = tree([ + node({ id: 1, box: [0, 0, 10, 10], label: 'a' }), + node({ id: 2, box: [0, 0, 10, 10], value: 'b' }), + node({ id: 3, box: [0, 0, 10, 10] }), + ]); + const { summary } = mergeOcrWithTree(t, [], frame); + expect(summary.nodes).toBe(3); + expect(summary.labelled).toBe(2); + }); +});