Real-time face recognition in the browser. The webcam stays in the page, frames get posted to a FastAPI backend running InsightFace (SCRFD to detect, ArcFace to embed), and the response comes back as box coordinates, whoever matched, and a set of readings taken off the pixels.
Every face it sees, named or not, goes into a log with a snapshot and a timestamp. Point it at an RTSP camera and it also records a clip whenever somebody or something turns up in front of it. Take a capture and you additionally get the objects in the frame, any text in it, and optionally a written description.
Python 3.10+ and a webcam.
git clone https://github.com/Xyness/Archer.git
cd Archer
pip install -r requirements.txt
python main.pyOpen http://127.0.0.1:8000.
First launch downloads the buffalo_l model pack (~280 MB) into ~/.insightface/models/, so
give it a minute. The object and text models are another ~25 MB into ~/.archer/models/,
fetched in the background at startup so the first capture doesn't have to wait for them.
GET /capabilities reports models_ready while that's still going.
Written descriptions and region lookup are the only things that need an API key
(ANTHROPIC_API_KEY, or an ant auth login profile). Without one they switch themselves off
and everything else carries on. There's also a local option, see below.
On Windows, install the Visual C++ Build Tools first or the InsightFace wheel won't build.
The browser grabs frames off the video element, encodes them as JPEG and POSTs them to
/analyze-frame. The backend decodes, runs detection on a copy shrunk to DETECT_WIDTH, then
scales the coordinates back up. Everything after detection reads the frame at the size it
arrived, so the skin patches, the head pose and the embedding get the resolution the browser
actually sent rather than the detector's thumbnail.
Embeddings are compared against every stored encoding with cosine similarity. One person can have several photos and the best match across them wins. Boxes come back as fractions of the frame rather than pixels, so the canvas overlay doesn't need to know how big the video element ended up. Known faces are drawn cyan, unknown ones red.
| Page | |
|---|---|
index.html |
Live view. Matched people in the sidebar, hover a box for the full reading, Capture for a full-resolution report |
cameras.html |
Streams the server watches on its own |
captures.html |
Every frame you kept, with everything read off it |
sightings.html |
One entry per person, named or not, with every visit inside |
events.html |
Clips the cameras recorded, filterable by camera or trigger |
persons.html |
The registry: register, search, add photos, delete |
Registering several photos per person under different lighting makes a real difference to the match rate. One photo is usually enough to be recognised head-on and not much else.
Everything here is an estimate and every field carries enough context to say how much of one.
| Field | Source | Worth |
|---|---|---|
| Gender | buffalo_l attribute head | Binary and apparent, which is all the model predicts |
| Age | same head | ~5 years mean absolute error, so it ships as a range |
| Skin tone | measured off the frame | As much a reading of the light in the room as of the person |
| Range | pinhole model on the interpupillary distance | Roughly a tenth either way |
| Stature | camera geometry | Needs the lens height, absent until you supply it |
| Head pose | landmark_3d_68 |
Solid, and it corrects the range for a turned head |
| Quality | sharpness, exposure, size, pose, detector score | Geometric mean, so one bad factor sinks it |
Skin tone samples three patches (forehead, both cheeks) placed off the eye and mouth landmarks. Pixels that disagree with the patch median get dropped, and a patch that disagrees with the other two gets dropped whole, which is what takes a fringe or a pair of glasses out of the reading. The surviving median lands on the Monk Skin Tone scale by nearest CIELAB swatch, plus the ITA bands dermatology uses.
The absolute YCrCb skin box most tutorials reach for is deliberately not used: it was tuned on light skin and rejects Monk 9 and 10 outright, which turns "dark skin" into "no skin found". Relative filtering has no such floor.
Range and stature. The IPD is the only length in a face you can guess in millimetres (63 mm adult mean, 3.8 mm SD), so it's the ruler everything geometric stands on, and it caps the precision at about a tenth either way. Range falls out of the pinhole relation. Stature doesn't: a face carries no scale of its own, so any height squeezed out of one alone is a population mean handed back with extra steps. What does work is knowing where the camera is. Put the lens height in the sidebar, with the tilt if it has one, and the drop from lens to face is trigonometry. Leave it blank and the panel says the camera isn't calibrated rather than inventing a number.
camera.py opens a stream itself, reads it in its own thread and keeps analysing whether or
not a browser is looking. Anything cv2.VideoCapture opens works: rtsp://, an http://
MJPEG feed, a video file, a local device index. Add one on the page and it starts immediately;
disable, rename or repoint it and the worker is reconciled without a restart.
Reading is uncapped because an RTSP socket that isn't drained backs up until you're watching
several seconds into the past. Analysis is throttled to five a second, since nothing needs
face recognition thirty times a second. Frames are shrunk to MAX_WIDTH first, because a 4K
doorbell costs four times the decode of a 1080p one and finds the same faces.
Fill in the lens height and tilt when you add a camera and stature works on it. This is where the geometry pays off: a laptop lid moves every time somebody adjusts the screen, a camera bolted to a wall has a height you measure once.
The picture reaches the browser as multipart/x-mixed-replace, which an <img src> renders
natively, so there's no websocket and no player. Unchanged frames aren't re-sent, except every
two seconds regardless, because a connection that goes completely silent looks the same as a
hung one. Credentials never leave the server: GET /cameras returns rtsp://admin:***@host.
The live loop runs about ten times a second, so a row per detection would mean a thousand rows and a thousand JPEGs for one person standing in shot for two minutes. What gets written instead is a sighting: one continuous run in front of a camera, opened when somebody appears and extended while they stay.
Frames are stitched into a run on two signals: how alike the faces are, and how much the boxes
overlap. Either alone gets it wrong. The first version used similarity only, at 0.55, which
was too strict and produced exactly the duplicates it was meant to prevent: detection runs on
a downscaled copy, so a face 120 px wide reaches ArcFace at 80, and an embedding off an 80 px
crop is noisy enough that turning your head drops the frame-to-frame cosine under the bar.
Now a strong resemblance is enough on its own and a weak one is enough when the box has barely
moved. tests/test_tracking.py pins the truth table.
Three things follow from that:
- The snapshot is the best frame of the run, not the first. A sighting that opens on a blurred profile improves as soon as the person turns to the camera.
- An identity that arrives late is written onto the run already open, rather than starting a second row mid-corridor.
last_seenis written at most every two seconds. Ten commits a second to move a timestamp nobody reads in real time is ten commits a second wasted.
Deleting somebody from the registry doesn't delete their history: the foreign key is
ON DELETE SET NULL, so the row stays and the label still holds the name they were logged
under.
Every sighting keeps the embedding it was measured from, which is what lets naming reach
backwards. Name an unnamed row and every other unnamed visit by that face is named with it,
at a stricter threshold than live recognition. It also makes the log searchable by face:
GET /sightings/{id}/similar is the whole history of one person in a single query, and it
works on strangers, which is the point.
That search is one matmul over the log rather than a loop, 19x faster at fifty thousand runs. It stays brute force on purpose: a million embeddings is 72 ms and 2 GB, so FAISS or Qdrant would be solving a problem this doesn't have.
Sightings also record what else was in the frame: objects, text, and a description. That runs once per run on a background worker, and it writes twice, so the fast readings land almost immediately and the description follows when it's ready.
A row per run is the right thing to record and the wrong thing to read. The page shows one entry per person with their visits counted, and the visits one click inside.
Registered people group by their record, strangers by their face: each unnamed run is compared
against the strangers already on file and joins the closest above CLUSTER_SIMILARITY, or
opens a new one. The group keeps a running mean of its members, so one bad frame doesn't define
who belongs to it. A stranger who comes past every morning is one entry saying eleven visits
instead of eleven lines of UNKNOWN.
Nothing is merged or rewritten in the log itself. Take filter=unknown off GET /sightings
and every individual run is still there.
Each camera keeps a rolling three seconds of video, and the moment it sees a face or one of the watched objects it opens a clip, writes those three seconds in first, and keeps going. The pre-roll is the half worth having: an event that starts when the detector fires begins with somebody already in the middle of the frame.
Recording stops four seconds after the last thing leaves, and in any case after sixty, so an
evening in front of the lens becomes a series of clips rather than one file that grows all
night. Watched classes are people, vehicles, animals and bags (WATCH_CLASSES in camera.py),
and the object pass runs every two seconds rather than every frame.
Clips are VP8 in WebM. If the OpenCV build can't write VP8 it falls back through VP9, MPEG-4 and MJPEG, and the page offers the file for download when it hits a container the browser won't decode.
Video is what fills a disk, so nothing is deleted on its own. POST /clips/prune takes a max
age and a max count, and the Events page has both as a form.
The capture button runs three more passes over the same frame. None of this touches the live loop, since YOLO at 50 ms and OCR at 300 ms would halve the frame rate for readings nobody watches ten times a second.
Objects are YOLOv10n over the 80 COCO classes. v10 is NMS-free, so the 300 rows that come back are already deduplicated, which removes the part of a YOLO wrapper that's easiest to get quietly wrong. 50-140 ms on CPU.
Text is PP-OCRv4, detection then recognition. DBNet returns a probability map that becomes
boxes through a threshold, contours and a Vatti offset (which for a rectangle is just growing
it on all four sides, so pyclipper isn't needed). CTC over 6623 characters, greedily decoded.
200-500 ms.
Both are ONNX through the onnxruntime InsightFace already pulls in, so neither costs a new
dependency, and neither touches the network at inference time.
Descriptions and region lookup are the exception and have two backends. Ollama runs a
vision model locally, so nothing leaves the machine and there's no key. Claude is better at it,
and adds web search behind region lookup so it can name a specific object and cite where that
came from. ARCHER_VISION picks, and on the default auto a reachable Ollama wins.
That's why they live in lookup.py and the local models live in scene.py: the boundary
between "runs on your machine" and "gets sent somewhere" is a file boundary you can point at.
Region lookup is scoped to objects. The face check runs server side on the uploaded crop itself, so it sees the exact bytes that would be sent onward, and a crop more than 15% covered by a detected face comes back 422.
ARCHER_VISION=auto prefers a reachable Ollama. Install it, pull one vision model, and Archer
finds it. The model has to fit in VRAM or Ollama silently falls back to CPU, where a
description takes minutes instead of seconds.
| Model | Needs |
|---|---|
moondream |
~2 GB |
qwen2.5vl:3b |
~3.2 GB |
llava:7b |
~4.7 GB |
llama3.2-vision |
~8 GB, best of the four |
Four things keep it quick, all on by default in lookup.py: OLLAMA_KEEP_ALIVE is 30 minutes
(Ollama unloads after five), the model is primed at startup on a thread nothing waits on,
OLLAMA_MAX_TOKENS is 220, and VISION_MAX_WIDTH is 768 since vision models resize to a few
hundred pixels internally anyway. That last one takes a capture from 1.7 MB to 190 KB.
hooks.py has two empty functions. Archer calls them and ignores what they return.
def on_sighting(event: dict) -> None:
"""A new sighting just opened."""
def on_capture(event: dict) -> None:
"""Somebody pressed Capture & analyse."""Both events carry the readings, the path as stored in the database and an absolute path ready
to open. on_sighting fires once per run, not once per frame.
They run on a background worker (13 ms of request time for a hook that sleeps for 1500) and an exception is printed and swallowed rather than taking recognition down. The queue is bounded at 256 and events get dropped with a line on stderr rather than growing without limit.
For something outside the process there's GET /sightings/stream: server-sent events, one per
thing that happens.
curl -N http://127.0.0.1:8000/sightings/streamEnglish and French, switched in the header. The choice lives on the server because it also feeds the vision prompts, and a page in French with descriptions arriving in English would be worse than either alone. It reaches the interface, the readings, the 80 COCO class names and the generated descriptions.
static/js/i18n.js holds both dictionaries side by side, and the suite fails if a key exists
in one and not the other.
| Method | Endpoint | |
|---|---|---|
POST |
/persons |
Register a person (multipart form) |
GET |
/persons |
List, takes q, page, per_page |
DELETE |
/persons/{id} |
Delete a person and their photos |
GET POST |
/persons/{id}/photos |
List or add photos |
DELETE |
/persons/{id}/photos/{photo_id} |
Delete one photo |
POST |
/persons/import |
Enrol a folder, one directory per person |
POST |
/persons/search-by-face |
Rank the registry against a face in a photo |
GET |
/calibration |
Where the thresholds should sit, on your data |
POST |
/analyze-frame |
Analyse a JPEG frame |
POST |
/analyze-capture |
Analyse one frame at full resolution and keep it |
POST |
/lookup-region |
Say what a cropped region is, with sources |
GET |
/captures |
Every kept capture with its readings |
DELETE POST |
/captures/{id}, /captures/prune |
Delete one, or old ones |
GET |
/people |
One entry per person. page, per_page, filter, q, since, until, source |
GET |
/sightings |
The log. Same filters plus person_id, cluster |
GET |
/sightings.csv |
The same, as a spreadsheet |
GET |
/sightings/{id}/similar |
Every other run by the same face |
POST |
/sightings/{id}/enrol |
Name a sighting, and their other visits |
POST |
/sightings/merge |
Collapse runs that should have been one |
POST |
/sightings/prune |
Delete old sightings and their snapshots |
GET |
/sightings/stream |
Server-sent events |
GET POST |
/cameras |
List or add a stream |
POST |
/cameras/{id}/enabled |
Start or stop one |
DELETE |
/cameras/{id} |
Remove one |
GET |
/cameras/{id}/stream |
MJPEG, for an <img src> |
GET |
/cameras/{id}/snapshot |
The latest frame as one JPEG |
GET |
/cameras/{id}/faces |
Boxes from the most recent analysis |
GET |
/clips |
Recorded events. page, per_page, source, trigger |
DELETE POST |
/clips/{id}, /clips/prune |
Delete one, or old ones |
GET POST |
/settings |
The chosen language |
GET |
/capabilities |
Which passes this install can run |
/analyze-frame and /analyze-capture take optional camera_height_cm and
camera_pitch_deg, which is what turns stature on. /analyze-capture also takes objects,
text, describe and keep.
None of this is on by default, because on 127.0.0.1 the only client is you.
Auth. Set ARCHER_USER and ARCHER_PASSWORD and every endpoint asks for a password,
static pages and uploads included. HTTP Basic on purpose: the browser handles the prompt and
remembers the answer, so there's no login page and no session handling to get wrong. Set them
the moment this listens on anything other than localhost.
Request size. Bodies over 25 MB are refused with a 413 before anything reads them, and a decode is refused past 40 megapixels. A JPEG header is a few bytes and can promise a huge canvas, so the second limit isn't implied by the first.
Stored file types. Anything outside ALLOWED_EXTENSIONS becomes .jpg, because the
extension decides the Content-Type StaticFiles serves with, and portrait.html coming back as
text/html from this origin is stored XSS in a folder of holiday photos.
Camera URLs and folder import. Both accept a local path, which is fine while the only client is you and is a filesystem read primitive once it isn't. Both are restricted exactly when authentication is on.
Retention. RETENTION_DAYS and RETENTION_ROWS at the top of routes/sightings.py,
routes/captures.py and routes/clips.py. All None by default, which keeps everything
forever and is also how a folder of full-resolution frames quietly becomes the biggest thing
on the disk.
SIMILARITY_THRESHOLD is 0.4 because that's where it stopped mixing people up on one test
set. Fine as a starting point, bad as a final answer. GET /calibration, or the Measure button
on the persons page, works it out from your own registry: it scores every pair of stored
photos, splits them into genuine (same person) and impostor (different people), and reports
where each cloud sits and how far apart they are.
That's also why a match carries a band rather than a yes or no. Above SIMILARITY_STRONG
it's worth trusting, between the two the honest answer is "check this one", since that's the
range where both distributions live. If the clouds overlap on your registry the report says so
and suggests the equal-error point as a compromise. More photos per person is what actually
fixes it.
POST /persons/import takes a directory laid out one folder per person:
dataset/
├── Ada Lovelace/
│ ├── front.jpg
│ └── side.jpg
└── Alan Turing/
└── passport.png
dry_run=true reports what it found and what it skipped without writing anything.
SQLite. Embeddings are raw float32 blobs read back with np.frombuffer, which keeps the
schema simple at the cost of not being able to query on them.
CREATE TABLE sightings (
id INTEGER PRIMARY KEY AUTOINCREMENT,
person_id INTEGER REFERENCES persons(id) ON DELETE SET NULL,
label TEXT NOT NULL,
similarity REAL,
snapshot_path TEXT NOT NULL,
quality INTEGER,
attributes TEXT, -- attributes.describe output, as JSON
first_seen TEXT NOT NULL,
last_seen TEXT NOT NULL,
encoding BLOB, -- what makes naming reach backwards, and grouping possible
source TEXT, -- 'webcam' or 'camera:3'
context TEXT, -- objects, text and the description, as JSON
cluster INTEGER, -- which stranger this run belongs to
clip_id INTEGER REFERENCES clips(id) ON DELETE SET NULL
);Plus persons, photos, clips, cameras and settings. init_db runs on every start in
three phases: create tables, add columns, index those columns. The order matters, since an
index names a column and building one over a column the same run is about to add works on your
machine and fails on a fresh clone.
Snapshots live in uploads/sightings, captures in uploads/captures, clips and posters in
uploads/clips.
Most of it is constants at the top of a file, with a comment saying what happens if you move them.
face_engine.py has SIMILARITY_THRESHOLD, DETECT_WIDTH (main lever on latency),
MOTION_GATING and READ_ATTRIBUTES. camera.py has ANALYSE_INTERVAL_S, MAX_WIDTH and
everything about clips. attributes.py has CAMERA_HFOV_DEG, which every centimetre reported
scales linearly with, so measure yours if the numbers need to mean anything. sightings.py has
the tracking and clustering thresholds and the retention of context passes. scene.py has
OBJECT_THRESHOLD and REC_MIN_CONFIDENCE (0.6, set by watching the recogniser read a bus:
the real words came back above 0.7 and the noise at 0.58). lookup.py has the model and both
prompts.
Inference is CPU through ONNX Runtime. Swapping CUDAExecutionProvider into _get_app() works
if you have the GPU build.
python tests/run.py713 checks, no pytest. Each suite is a script that prints [ok ] or [FAIL] per check and
exits non-zero if any failed. tests/stubs/insightface stands in for the model pack so eleven
of the twelve suites need no download. The twelfth runs the real YOLO and OCR against a real
photograph, because a synthetic ellipse proves the plumbing and proves nothing about the
detector.
attributes.py soft biometrics from a face camera.py streams, and the clips they record
face_engine.py detection, matching, the log scene.py objects and text, local
sightings.py runs, and who they belong to lookup.py the two calls that leave
models.py ONNX files and their cache hooks.py your code
database.py schema and queries routes/ one file per surface
Pages load static/js/common.js plus one file of their own rather than a bundle, so the live
view doesn't download and parse the sightings page. Plain scripts sharing globals: there's no
build step, and a module graph would be the heaviest thing in the project.
Matching is one matmul against every stored encoding, which holds up into the low thousands before per-person grouping starts to dominate. Encodings are loaded once at startup and refreshed on write, so an external process editing the database won't be picked up.
Object detection is capped at the 80 COCO classes, so it names a suitcase and not a brand of suitcase. The recogniser is the Chinese PP-OCRv4 model, which covers Latin text well and small low-contrast text badly. Detection range for both faces and text drops off well before the object detector's does.
Ethnicity classification. Article 5(1)(g) of the EU AI Act prohibits biometric categorisation that infers race or ethnicity. Skin tone is measured instead, and a photometric reading of pixels under a given light isn't a claim about a person.
Liveness detection. Tried and dropped, worth writing down so nobody repeats it.
MiniFASNet-V2 from Silent-Face-Anti-Spoofing is the obvious candidate: small, permissively
licensed, ONNX export available. Reproducing the upstream pipeline faithfully (both models with
strict=True, the documented 2.7x and 4.0x crops, BGR at [0,1], the ensemble, index 1 as real)
it labelled every genuine face in a real 3:4 photo as an attack at 0.95 confidence, and
returned near-identical outputs for faces from 38 to 175 px across. The ONNX export is faithful
to the PyTorch original; the original behaves the same way here. Presumably tied to the capture
conditions it was trained on. A liveness signal that's confidently wrong is worse than none.
Face search against the web. Matching a face against Instagram or any other index means identifying strangers who never enrolled, which is the Clearview shape. There's also no wire to connect: Instagram exposes no face search and Google Lens has no public API.
Two things do the useful parts. /persons/search-by-face searches the registry, same
embeddings, ranked instead of thresholded, scoped to people enrolled on purpose.
/lookup-region searches the web for objects, gated on the face check.
MIT