Skip to content

Verity

An open, domain-general engine for forensic surface comparison — a transparent, calibrated likelihood ratio, not a black-box "match."

CI crates.io PyPI License: MIT/Apache-2.0 Status: early

Live app · Studio · Method docs · Open benchmark · Docs · API reference

Verity — forensic marks, weighed as evidence

Why · What you get · Validation · How it works · Quickstart · Repo map · Roadmap · License


Verity compares 3-D surface-topography scans — bullet lands, cartridge-case breech-face impressions, striated and impressed toolmarks, and (in time) footwear and fractured surfaces — directly from X3P files (ISO 25178-72). It pairs a domain-general surface comparison with a transparent, calibrated likelihood-ratio decision layer and region-level attribution. The machine never reports a "match"; it reports an auditable weight of evidence, characterized on a named dataset.

Note

Status: early. The X3P codec (verity-x3p, v0.2.0) is published on crates.io and PyPI and tested against real-world files. The engine, comparison API, and web app are live; the first-principles method is validated source-disjoint across bullet lands, cartridge cases, and toolmarks (see Validation). The first pre-registered external test has landed: applied unchanged to the independent Weller et al. (2012) cartridge-case set — pre-registered on OSF (2026-07-01, before any Weller data access) and run once as registered — the frozen pipeline reached pooled Cllr 0.163 (95% CI 0.147–0.189), AUC 0.9995, so the pre-registered H1 (pooled Cllr ≤ 0.45) is supported: docs/weller-preregistration.md.

cargo add verity-x3p                    # crates.io, v0.2.0
pip install verity-x3p                  # PyPI, abi3 wheels, Python 3.9+
R CMD INSTALL bindings/r/verityx3p      # from a clone of this repo; not on CRAN; requires a Rust toolchain

Full quickstart below · codec changelog

Why

Forensic firearm/toolmark comparison today is either subjective examiner judgment or proprietary black-box correlation (IBIS), while the open tooling is a pile of domain-specific R packages with no unified, deployable platform. Courts are increasingly skeptical of unqualified pattern-match testimony (Abruquah v. Maryland, 2023; the 2023 amendment to FRE 702), and no discipline yet has a well-characterized error rate (Cuellar et al., 2024). Verity's bet: one general, calibrated, explainable method — proven first where ground truth is strongest (firearms), then transferred across domains. More: docs.verity.codes/why.

What you get — a ComparisonReport, not a verdict

A likelihood ratio with its verbal equivalent and a credible interval (a clustered bootstrap of the reference, so the LR carries its own calibration uncertainty), a characterized cost (Cllr) on a named reference population, an empirical cap (ELUB-inspired) on how strong a claim the data can support, and the region-level attribution that drove the score.

{
  "likelihood_ratio": 146.0,
  "log10_lr": 2.16,
  "verbal": "moderately strong support for same source",
  "lr_bound_log10": 2.16,
  // a percentile credible interval on log10 LR — the reference is bootstrapped
  // (clustered by source/barrel) and refit
  "log10_lr_ci_lo": 1.74, "log10_lr_ci_hi": 2.16, "lr_ci_method": "bootstrap-clustered",
  // reference diagnostics are the *in-sample* fit of the named reference; the honest
  // validation figure is the source-disjoint Cllr ≈ 0.19 pooled (≈ 0.11 on Hamby-252
  // alone) — every number, with its protocol, lives in docs/headline-numbers.md
  "reference": { "name": "pooled bullet-land reference (Hamby-252 & 173, Beretta, Phoenix)",
                 "n_km": 146, "n_knm": 1755,
                 "auc": 0.984, "cllr": 0.193, "cllr_min": 0.168 },
  "attribution": [ /* the matched regions — the explanation */ ],
  "scope_note": "This is a calibrated weight of evidence on the pooled bullet-land reference population. It is not a verdict: it is one input to an examiner's judgment, alongside case context. It is not a claim about the error rate of striated examination, which remains unknown."
}

Validation (honest)

Important

Every number below carries its exact protocol — in-sample vs. source-disjoint vs. the frozen open-benchmark protocol are different claims and are never mixed. The canonical registry is docs/headline-numbers.md — every figure here matches a row there, and that protocol-labeled table is the citable source.

One method, three mark families — one algorithm (CMR), one frozen scorer config, validated source-disjoint in each family on the frozen open benchmark (fold mean):

Mark family Frozen split Cllr AUC
Bullet lands (striated) bullets-v1 0.205 ± 0.125 0.979
Cartridge breech faces (impressed) cartridge-v1 0.398 ± 0.2021 0.922
Screwdriver toolmarks (striated) toolmark-v1 0.330 ± 0.047 0.943

These characterize weight of evidence on the named references, not field error rates.

Bullet lands, protocol by protocol — the production scorer is diag_contrast, first-principles (no learned representation)2; the four studies are Hamby-252/173, PGPD Beretta, and Phoenix Ruger:

Protocol Cllr AUC
In-sample (deployed reference — optimistic by construction, not a validation claim) 0.193 (Cllr_min 0.168) 0.984
Source-disjoint, pooled four studies 0.186 ± 0.126 0.989
Frozen open benchmark bullets-v1 (fold mean) 0.205 ± 0.125 0.979
Hamby-252 alone, barrel-disjoint3 0.113 ± 0.066 1.000
Specialist bulletxtrctr (random forest), identical protocol, Hamby-252 0.064 ± 0.015 ≈ 1.000

An informative, calibrated weight of evidence from metrology alone4. The trained bulletxtrctr random-forest specialist, run through the identical protocol, reaches Cllr ≈ 0.06 on Hamby-252 — the specialist still leads on its home turf; Verity's contribution on bullets is the calibrated, bounded, deployable LR layer, not a better matcher.

diag_contrast was selected over the Phase-1 diag_mean and a multivariate fusion by an explicit barrel-disjoint ablation (verity-margin) — candidly, that ablation reused the same four studies as this validation, so a one-shot confirmation on untouched data was the next milestone (see the whitepaper's Limitations). That confirmation has now run, once, as pre-registered (OSF, registered 2026-07-01 before any Weller data access): the frozen Fadul-calibrated pipeline transferred to the independent Weller et al. (2012) set at pooled Cllr 0.163 (95% CI 0.147–0.189), AUC 0.9995 — H1 supported. That the pooled figure sits below the within-study Fadul reference (0.343) reflects Weller's own strong separation and ~37× more same-source pairs; the within-study cartridge-v2 split isolates it at Cllr 0.045. Every figure, with its protocol, is in docs/headline-numbers.md.

What doesn't work yet

The Phase-2b learned representation, trained barrel-disjoint on 210 Hamby scans, does not beat the cross-correlation baseline — it overfits (held-out AUC collapses to ≈ 0.67). Synthetic tests confirm the pipeline does learn given enough signal: a data limit, not a defect. Next: expand the dataset and retest. (Sources: services/engine; the whitepaper's Limitations.)

verity-validation-report regenerates the full characterization — Tippett, DET, calibration, and the source-disjoint summary — as a court-ready PDF. The frozen splits, replication kits, and leaderboard live at data.verity.codes. Deep dive: docs/validation.md · whitepaper PDF.

Important

Nothing here is a claim about the error rate of forensic examination, which remains unknown.

How it works

One codec, one truth: X3P (ISO 25178-72) → verity-x3p (Rust core) → PyO3 / extendr bindings → engine: preprocess (ISO 16610) → register → CMR → calibrate (empirical cap) → ComparisonReport → API · web · MCP.

  • Statistics decide, not a black box. A representation produces a score; a transparent, empirically-capped calibration turns it into a reportable LR, interpretable regardless of how the score was computed — the firewall against the black box.
  • Reproducible by construction. Deterministic, version-pinned, content-hashed.
  • Open and language-independent. Built on the X3P standard; MIT/Apache-2.0.

Congruent Matching Regions (CMR) generalizes Song's Congruent Matching Cells (the standard cartridge-case method) from 2-D cells and a fixed translation+rotation to regions of any dimension under any transformation group — so one algorithm scores striated, impressed, and (research) fractured marks. Partition a mark into regions, register each against the other mark, and count the regions that agree on one common geometry. The congruent regions are the attribution map.

Modality Region Transform group Reduces to
Striated 1-D profile window 1-D translation ≈ Chumbley / CMS
Impressed 2-D grid cell 2-D translation+rotation ≈ CMC
Fractured 3-D mesh patch 3-D rigid pose (research)

Full write-up: docs/congruent-matching-regions.md.

Quickstart

import verity_x3p
s = verity_x3p.read_x3p("scan.x3p")            # s.data, s.mask are (ny, nx) NumPy arrays
verity_x3p.write_x3p(s, "copy.x3p", z_type="D")
Rust and R — same core, same bits
use verity_x3p::{read_x3p, write_x3p, WriteOptions};
let surface = read_x3p("scan.x3p")?;          // verifies the stored MD5
write_x3p(&surface, "copy.x3p", &WriteOptions::default())?;
library(verityx3p)
s <- read_x3p("scan.x3p")                      # s$surface is an nx-by-ny matrix
write_x3p(s, "copy.x3p")

A file written from any binding reads back bit-identically in every other.

Compare two marks over HTTP:

curl -s -X POST https://api.verity.codes/compare \
  -F domain=striated \
  -F mark_a=@bulletA_land1.x3p -F mark_a=@bulletA_land2.x3p \
  -F mark_b=@bulletB_land1.x3p -F mark_b=@bulletB_land2.x3p

Full docs: docs.verity.codes · interactive API reference.

MCP:

  • Hosted remote server at https://api.verity.codes/mcp (base64 scan inputs).
  • Local stdio via uv run --directory services/mcp verity-mcp (env VERITY_API_URL); Claude Desktop bundle via services/mcp/build_mcpb.sh.
  • Six tools (compare_marks, detect_mark_type, calibrate_score, list_references, scorer_config, service_health) — same firewall and scope-note guarantees as the HTTP API.

Thin clients: clients/python/verity_client.py (requests-only) and clients/r/verity.R; every report carries a sha256: content handle over the canonical recipe — v.reproduce(...) re-runs and hash-checks it. See clients/README.md.

Develop
cargo test -p verity-x3p                        # the Rust core

cd services/engine && uv sync --extra dev && uv run --extra dev pytest
cd services/api    && uv sync --extra dev && uv run --extra dev verity-api   # API on :8000
cd services/web    && pnpm install && pnpm dev                               # web on :3000

Full per-package setup, PR conventions, and the method-change policy (any PR touching a validation number must update docs/headline-numbers.md) live in CONTRIBUTING.md; deployment in DEPLOY.md.

Repository map

A polyglot monorepo: one Rust codec core, thin language bindings, and the Python science + service stack on top.

Package Lang Role
crates/verity-x3p Rust Native X3P (ISO 25178-72) reader/writer — the format's single source of truth.
bindings/python PyO3 + NumPy Python binding to the core (bit-identical I/O).
bindings/r/verityx3p extendr R binding to the core (x3ptools-compatible layout).
services/engine Python Metrology preprocessing, registration, CMR, the calibrated-LR decision layer.
services/api FastAPI The comparison HTTP API serving the ComparisonReport.
services/catalog Python Normalized catalog + content-addressed store + ingestion (NBTRD / Figshare / GitHub harvests, virtual kits).
services/web Next.js verity.codes, docs.verity.codes, and the Studio.
services/mcp Python MCP server ("verity") — local stdio, plus the hosted endpoint at api.verity.codes/mcp.
clients/ Python / R Thin API clients + the content-handle reproducibility contract.

Status & roadmap

Done: verity-x3p native codec + Python/R bindings (bit-identical round-trip), v0.2.0 on crates.io and PyPI. Engine: ISO 16610 preprocessing, registration, the calibrated-LR decision layer, CMR; source-disjoint validation across bullet lands, cartridge cases, and toolmarks (tables above). Platform: comparison API, web app, docs, Studio, and the open benchmark — live at verity.codes and api/docs/app/data.verity.codes.

In progress: the pre-registered one-shot external validation on untouched data (see Validation).

Next: expand the bullet/cartridge/toolmark datasets (NBTRD harvest) and retest the learned representation; CMR-2D → CMC parity on Fadul — cmcR still leads, parity is open roadmap; TypeScript/Swift/Java codec bindings. Extending to more mark families: docs/toolmark-roadmap.md.

Citing & further reading

License

Dual-licensed under either of MIT or Apache-2.0, at your option. Bundled reference data carries its own upstream attribution — see services/api/verity_api/references/NOTICE.md.

Footnotes

  1. With only 10 same-source Fadul pairs, per-fold AUC is unstable across protocols (a small-n artifact, not a contradiction — see docs/headline-numbers.md). The specialist cmcR still leads on Fadul; CMR-2D → CMC parity is open roadmap.

  2. Barrel-disjoint: no barrel in both train and test; reported per study, never pooled across makes. First-principles: no learned representation.

  3. Hamby-252 alone — the single strongest study, not the pooled figure; the pooled source-disjoint Cllr is ≈ 0.19 and the frozen bullets-v1 benchmark reproduces ≈ 0.21.

  4. Cllr < 1 = informative; the Cllr − Cllr_min gap is the calibration loss the source-disjoint split exposes — answering the Cuellar et al. critique on its own terms.

About

An open, domain-general engine for forensic surface comparison: calibrated likelihood ratios (not black-box matches) from X3P 3-D scans of bullets, cartridge cases, and toolmarks.

Topics

Resources

Code of conduct

Contributing

Stars

Watchers

Forks

Releases

Packages

Used by

Contributors

Languages