Skip to content
XynessPublic

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

Archer

Real-time face recognition in the browser. The webcam stays in the page, frames get posted to a FastAPI backend running InsightFace (SCRFD to detect, ArcFace to embed), and the response comes back as box coordinates, whoever matched, and a set of readings taken off the pixels.

Every face it sees, named or not, goes into a log with a snapshot and a timestamp. Point it at an RTSP camera and it also records a clip whenever somebody or something turns up in front of it. Take a capture and you additionally get the objects in the frame, any text in it, and optionally a written description.

Running it

Python 3.10+ and a webcam.

git clone https://github.com/Xyness/Archer.git
cd Archer
pip install -r requirements.txt
python main.py

Open http://127.0.0.1:8000.

First launch downloads the buffalo_l model pack (~280 MB) into ~/.insightface/models/, so give it a minute. The object and text models are another ~25 MB into ~/.archer/models/, fetched in the background at startup so the first capture doesn't have to wait for them. GET /capabilities reports models_ready while that's still going.

Written descriptions and region lookup are the only things that need an API key (ANTHROPIC_API_KEY, or an ant auth login profile). Without one they switch themselves off and everything else carries on. There's also a local option, see below.

On Windows, install the Visual C++ Build Tools first or the InsightFace wheel won't build.

How it works

The browser grabs frames off the video element, encodes them as JPEG and POSTs them to /analyze-frame. The backend decodes, runs detection on a copy shrunk to DETECT_WIDTH, then scales the coordinates back up. Everything after detection reads the frame at the size it arrived, so the skin patches, the head pose and the embedding get the resolution the browser actually sent rather than the detector's thumbnail.

Embeddings are compared against every stored encoding with cosine similarity. One person can have several photos and the best match across them wins. Boxes come back as fractions of the frame rather than pixels, so the canvas overlay doesn't need to know how big the video element ended up. Known faces are drawn cyan, unknown ones red.

The pages

Page
index.html Live view. Matched people in the sidebar, hover a box for the full reading, Capture for a full-resolution report
cameras.html Streams the server watches on its own
captures.html Every frame you kept, with everything read off it
sightings.html One entry per person, named or not, with every visit inside
events.html Clips the cameras recorded, filterable by camera or trigger
persons.html The registry: register, search, add photos, delete

Registering several photos per person under different lighting makes a real difference to the match rate. One photo is usually enough to be recognised head-on and not much else.

What the readings mean

Everything here is an estimate and every field carries enough context to say how much of one.

Field Source Worth
Gender buffalo_l attribute head Binary and apparent, which is all the model predicts
Age same head ~5 years mean absolute error, so it ships as a range
Skin tone measured off the frame As much a reading of the light in the room as of the person
Range pinhole model on the interpupillary distance Roughly a tenth either way
Stature camera geometry Needs the lens height, absent until you supply it
Head pose landmark_3d_68 Solid, and it corrects the range for a turned head
Quality sharpness, exposure, size, pose, detector score Geometric mean, so one bad factor sinks it

Skin tone samples three patches (forehead, both cheeks) placed off the eye and mouth landmarks. Pixels that disagree with the patch median get dropped, and a patch that disagrees with the other two gets dropped whole, which is what takes a fringe or a pair of glasses out of the reading. The surviving median lands on the Monk Skin Tone scale by nearest CIELAB swatch, plus the ITA bands dermatology uses.

The absolute YCrCb skin box most tutorials reach for is deliberately not used: it was tuned on light skin and rejects Monk 9 and 10 outright, which turns "dark skin" into "no skin found". Relative filtering has no such floor.

Range and stature. The IPD is the only length in a face you can guess in millimetres (63 mm adult mean, 3.8 mm SD), so it's the ruler everything geometric stands on, and it caps the precision at about a tenth either way. Range falls out of the pinhole relation. Stature doesn't: a face carries no scale of its own, so any height squeezed out of one alone is a population mean handed back with extra steps. What does work is knowing where the camera is. Put the lens height in the sidebar, with the tilt if it has one, and the drop from lens to face is trigonometry. Leave it blank and the panel says the camera isn't calibrated rather than inventing a number.

Cameras

camera.py opens a stream itself, reads it in its own thread and keeps analysing whether or not a browser is looking. Anything cv2.VideoCapture opens works: rtsp://, an http:// MJPEG feed, a video file, a local device index. Add one on the page and it starts immediately; disable, rename or repoint it and the worker is reconciled without a restart.

Reading is uncapped because an RTSP socket that isn't drained backs up until you're watching several seconds into the past. Analysis is throttled to five a second, since nothing needs face recognition thirty times a second. Frames are shrunk to MAX_WIDTH first, because a 4K doorbell costs four times the decode of a 1080p one and finds the same faces.

Fill in the lens height and tilt when you add a camera and stature works on it. This is where the geometry pays off: a laptop lid moves every time somebody adjusts the screen, a camera bolted to a wall has a height you measure once.

The picture reaches the browser as multipart/x-mixed-replace, which an <img src> renders natively, so there's no websocket and no player. Unchanged frames aren't re-sent, except every two seconds regardless, because a connection that goes completely silent looks the same as a hung one. Credentials never leave the server: GET /cameras returns rtsp://admin:***@host.

The log

The live loop runs about ten times a second, so a row per detection would mean a thousand rows and a thousand JPEGs for one person standing in shot for two minutes. What gets written instead is a sighting: one continuous run in front of a camera, opened when somebody appears and extended while they stay.

Frames are stitched into a run on two signals: how alike the faces are, and how much the boxes overlap. Either alone gets it wrong. The first version used similarity only, at 0.55, which was too strict and produced exactly the duplicates it was meant to prevent: detection runs on a downscaled copy, so a face 120 px wide reaches ArcFace at 80, and an embedding off an 80 px crop is noisy enough that turning your head drops the frame-to-frame cosine under the bar. Now a strong resemblance is enough on its own and a weak one is enough when the box has barely moved. tests/test_tracking.py pins the truth table.

Three things follow from that:

  • The snapshot is the best frame of the run, not the first. A sighting that opens on a blurred profile improves as soon as the person turns to the camera.
  • An identity that arrives late is written onto the run already open, rather than starting a second row mid-corridor.
  • last_seen is written at most every two seconds. Ten commits a second to move a timestamp nobody reads in real time is ten commits a second wasted.

Deleting somebody from the registry doesn't delete their history: the foreign key is ON DELETE SET NULL, so the row stays and the label still holds the name they were logged under.

Every sighting keeps the embedding it was measured from, which is what lets naming reach backwards. Name an unnamed row and every other unnamed visit by that face is named with it, at a stricter threshold than live recognition. It also makes the log searchable by face: GET /sightings/{id}/similar is the whole history of one person in a single query, and it works on strangers, which is the point.

That search is one matmul over the log rather than a loop, 19x faster at fifty thousand runs. It stays brute force on purpose: a million embeddings is 72 ms and 2 GB, so FAISS or Qdrant would be solving a problem this doesn't have.

Sightings also record what else was in the frame: objects, text, and a description. That runs once per run on a background worker, and it writes twice, so the fast readings land almost immediately and the description follows when it's ready.

Grouping

A row per run is the right thing to record and the wrong thing to read. The page shows one entry per person with their visits counted, and the visits one click inside.

Registered people group by their record, strangers by their face: each unnamed run is compared against the strangers already on file and joins the closest above CLUSTER_SIMILARITY, or opens a new one. The group keeps a running mean of its members, so one bad frame doesn't define who belongs to it. A stranger who comes past every morning is one entry saying eleven visits instead of eleven lines of UNKNOWN.

Nothing is merged or rewritten in the log itself. Take filter=unknown off GET /sightings and every individual run is still there.

Recording

Each camera keeps a rolling three seconds of video, and the moment it sees a face or one of the watched objects it opens a clip, writes those three seconds in first, and keeps going. The pre-roll is the half worth having: an event that starts when the detector fires begins with somebody already in the middle of the frame.

Recording stops four seconds after the last thing leaves, and in any case after sixty, so an evening in front of the lens becomes a series of clips rather than one file that grows all night. Watched classes are people, vehicles, animals and bags (WATCH_CLASSES in camera.py), and the object pass runs every two seconds rather than every frame.

Clips are VP8 in WebM. If the OpenCV build can't write VP8 it falls back through VP9, MPEG-4 and MJPEG, and the page offers the file for download when it hits a container the browser won't decode.

Video is what fills a disk, so nothing is deleted on its own. POST /clips/prune takes a max age and a max count, and the Events page has both as a form.

Beyond faces

The capture button runs three more passes over the same frame. None of this touches the live loop, since YOLO at 50 ms and OCR at 300 ms would halve the frame rate for readings nobody watches ten times a second.

Objects are YOLOv10n over the 80 COCO classes. v10 is NMS-free, so the 300 rows that come back are already deduplicated, which removes the part of a YOLO wrapper that's easiest to get quietly wrong. 50-140 ms on CPU.

Text is PP-OCRv4, detection then recognition. DBNet returns a probability map that becomes boxes through a threshold, contours and a Vatti offset (which for a rectangle is just growing it on all four sides, so pyclipper isn't needed). CTC over 6623 characters, greedily decoded. 200-500 ms.

Both are ONNX through the onnxruntime InsightFace already pulls in, so neither costs a new dependency, and neither touches the network at inference time.

Descriptions and region lookup are the exception and have two backends. Ollama runs a vision model locally, so nothing leaves the machine and there's no key. Claude is better at it, and adds web search behind region lookup so it can name a specific object and cite where that came from. ARCHER_VISION picks, and on the default auto a reachable Ollama wins.

That's why they live in lookup.py and the local models live in scene.py: the boundary between "runs on your machine" and "gets sent somewhere" is a file boundary you can point at.

Region lookup is scoped to objects. The face check runs server side on the uploaded crop itself, so it sees the exact bytes that would be sent onward, and a crop more than 15% covered by a detected face comes back 422.

Running the vision model locally

ARCHER_VISION=auto prefers a reachable Ollama. Install it, pull one vision model, and Archer finds it. The model has to fit in VRAM or Ollama silently falls back to CPU, where a description takes minutes instead of seconds.

Model Needs
moondream ~2 GB
qwen2.5vl:3b ~3.2 GB
llava:7b ~4.7 GB
llama3.2-vision ~8 GB, best of the four

Four things keep it quick, all on by default in lookup.py: OLLAMA_KEEP_ALIVE is 30 minutes (Ollama unloads after five), the model is primed at startup on a thread nothing waits on, OLLAMA_MAX_TOKENS is 220, and VISION_MAX_WIDTH is 768 since vision models resize to a few hundred pixels internally anyway. That last one takes a capture from 1.7 MB to 190 KB.

Your own code

hooks.py has two empty functions. Archer calls them and ignores what they return.

def on_sighting(event: dict) -> None:
    """A new sighting just opened."""

def on_capture(event: dict) -> None:
    """Somebody pressed Capture & analyse."""

Both events carry the readings, the path as stored in the database and an absolute path ready to open. on_sighting fires once per run, not once per frame.

They run on a background worker (13 ms of request time for a hook that sleeps for 1500) and an exception is printed and swallowed rather than taking recognition down. The queue is bounded at 256 and events get dropped with a line on stderr rather than growing without limit.

For something outside the process there's GET /sightings/stream: server-sent events, one per thing that happens.

curl -N http://127.0.0.1:8000/sightings/stream

Language

English and French, switched in the header. The choice lives on the server because it also feeds the vision prompts, and a page in French with descriptions arriving in English would be worse than either alone. It reaches the interface, the readings, the 80 COCO class names and the generated descriptions.

static/js/i18n.js holds both dictionaries side by side, and the suite fails if a key exists in one and not the other.

API

Method Endpoint
POST /persons Register a person (multipart form)
GET /persons List, takes q, page, per_page
DELETE /persons/{id} Delete a person and their photos
GET POST /persons/{id}/photos List or add photos
DELETE /persons/{id}/photos/{photo_id} Delete one photo
POST /persons/import Enrol a folder, one directory per person
POST /persons/search-by-face Rank the registry against a face in a photo
GET /calibration Where the thresholds should sit, on your data
POST /analyze-frame Analyse a JPEG frame
POST /analyze-capture Analyse one frame at full resolution and keep it
POST /lookup-region Say what a cropped region is, with sources
GET /captures Every kept capture with its readings
DELETE POST /captures/{id}, /captures/prune Delete one, or old ones
GET /people One entry per person. page, per_page, filter, q, since, until, source
GET /sightings The log. Same filters plus person_id, cluster
GET /sightings.csv The same, as a spreadsheet
GET /sightings/{id}/similar Every other run by the same face
POST /sightings/{id}/enrol Name a sighting, and their other visits
POST /sightings/merge Collapse runs that should have been one
POST /sightings/prune Delete old sightings and their snapshots
GET /sightings/stream Server-sent events
GET POST /cameras List or add a stream
POST /cameras/{id}/enabled Start or stop one
DELETE /cameras/{id} Remove one
GET /cameras/{id}/stream MJPEG, for an <img src>
GET /cameras/{id}/snapshot The latest frame as one JPEG
GET /cameras/{id}/faces Boxes from the most recent analysis
GET /clips Recorded events. page, per_page, source, trigger
DELETE POST /clips/{id}, /clips/prune Delete one, or old ones
GET POST /settings The chosen language
GET /capabilities Which passes this install can run

/analyze-frame and /analyze-capture take optional camera_height_cm and camera_pitch_deg, which is what turns stature on. /analyze-capture also takes objects, text, describe and keep.

Locking it down

None of this is on by default, because on 127.0.0.1 the only client is you.

Auth. Set ARCHER_USER and ARCHER_PASSWORD and every endpoint asks for a password, static pages and uploads included. HTTP Basic on purpose: the browser handles the prompt and remembers the answer, so there's no login page and no session handling to get wrong. Set them the moment this listens on anything other than localhost.

Request size. Bodies over 25 MB are refused with a 413 before anything reads them, and a decode is refused past 40 megapixels. A JPEG header is a few bytes and can promise a huge canvas, so the second limit isn't implied by the first.

Stored file types. Anything outside ALLOWED_EXTENSIONS becomes .jpg, because the extension decides the Content-Type StaticFiles serves with, and portrait.html coming back as text/html from this origin is stored XSS in a folder of holiday photos.

Camera URLs and folder import. Both accept a local path, which is fine while the only client is you and is a filesystem read primitive once it isn't. Both are restricted exactly when authentication is on.

Retention. RETENTION_DAYS and RETENTION_ROWS at the top of routes/sightings.py, routes/captures.py and routes/clips.py. All None by default, which keeps everything forever and is also how a folder of full-resolution frames quietly becomes the biggest thing on the disk.

Thresholds

SIMILARITY_THRESHOLD is 0.4 because that's where it stopped mixing people up on one test set. Fine as a starting point, bad as a final answer. GET /calibration, or the Measure button on the persons page, works it out from your own registry: it scores every pair of stored photos, splits them into genuine (same person) and impostor (different people), and reports where each cloud sits and how far apart they are.

That's also why a match carries a band rather than a yes or no. Above SIMILARITY_STRONG it's worth trusting, between the two the honest answer is "check this one", since that's the range where both distributions live. If the clouds overlap on your registry the report says so and suggests the equal-error point as a compromise. More photos per person is what actually fixes it.

Importing a folder

POST /persons/import takes a directory laid out one folder per person:

dataset/
├── Ada Lovelace/
│   ├── front.jpg
│   └── side.jpg
└── Alan Turing/
    └── passport.png

dry_run=true reports what it found and what it skipped without writing anything.

Storage

SQLite. Embeddings are raw float32 blobs read back with np.frombuffer, which keeps the schema simple at the cost of not being able to query on them.

CREATE TABLE sightings (
    id INTEGER PRIMARY KEY AUTOINCREMENT,
    person_id INTEGER REFERENCES persons(id) ON DELETE SET NULL,
    label TEXT NOT NULL,
    similarity REAL,
    snapshot_path TEXT NOT NULL,
    quality INTEGER,
    attributes TEXT,      -- attributes.describe output, as JSON
    first_seen TEXT NOT NULL,
    last_seen TEXT NOT NULL,
    encoding BLOB,        -- what makes naming reach backwards, and grouping possible
    source TEXT,          -- 'webcam' or 'camera:3'
    context TEXT,         -- objects, text and the description, as JSON
    cluster INTEGER,      -- which stranger this run belongs to
    clip_id INTEGER REFERENCES clips(id) ON DELETE SET NULL
);

Plus persons, photos, clips, cameras and settings. init_db runs on every start in three phases: create tables, add columns, index those columns. The order matters, since an index names a column and building one over a column the same run is about to add works on your machine and fails on a fresh clone.

Snapshots live in uploads/sightings, captures in uploads/captures, clips and posters in uploads/clips.

Tuning

Most of it is constants at the top of a file, with a comment saying what happens if you move them.

face_engine.py has SIMILARITY_THRESHOLD, DETECT_WIDTH (main lever on latency), MOTION_GATING and READ_ATTRIBUTES. camera.py has ANALYSE_INTERVAL_S, MAX_WIDTH and everything about clips. attributes.py has CAMERA_HFOV_DEG, which every centimetre reported scales linearly with, so measure yours if the numbers need to mean anything. sightings.py has the tracking and clustering thresholds and the retention of context passes. scene.py has OBJECT_THRESHOLD and REC_MIN_CONFIDENCE (0.6, set by watching the recogniser read a bus: the real words came back above 0.7 and the noise at 0.58). lookup.py has the model and both prompts.

Inference is CPU through ONNX Runtime. Swapping CUDAExecutionProvider into _get_app() works if you have the GPU build.

Tests

python tests/run.py

713 checks, no pytest. Each suite is a script that prints [ok ] or [FAIL] per check and exits non-zero if any failed. tests/stubs/insightface stands in for the model pack so eleven of the twelve suites need no download. The twelfth runs the real YOLO and OCR against a real photograph, because a synthetic ellipse proves the plumbing and proves nothing about the detector.

Layout

attributes.py  soft biometrics from a face      camera.py   streams, and the clips they record
face_engine.py detection, matching, the log     scene.py    objects and text, local
sightings.py   runs, and who they belong to     lookup.py   the two calls that leave
models.py      ONNX files and their cache       hooks.py    your code
database.py    schema and queries               routes/     one file per surface

Pages load static/js/common.js plus one file of their own rather than a bundle, so the live view doesn't download and parse the sightings page. Plain scripts sharing globals: there's no build step, and a module graph would be the heaviest thing in the project.

Known limits

Matching is one matmul against every stored encoding, which holds up into the low thousands before per-person grouping starts to dominate. Encodings are loaded once at startup and refreshed on write, so an external process editing the database won't be picked up.

Object detection is capped at the 80 COCO classes, so it names a suitcase and not a brand of suitcase. The recogniser is the Chinese PP-OCRv4 model, which covers Latin text well and small low-contrast text badly. Detection range for both faces and text drops off well before the object detector's does.

Not in here

Ethnicity classification. Article 5(1)(g) of the EU AI Act prohibits biometric categorisation that infers race or ethnicity. Skin tone is measured instead, and a photometric reading of pixels under a given light isn't a claim about a person.

Liveness detection. Tried and dropped, worth writing down so nobody repeats it. MiniFASNet-V2 from Silent-Face-Anti-Spoofing is the obvious candidate: small, permissively licensed, ONNX export available. Reproducing the upstream pipeline faithfully (both models with strict=True, the documented 2.7x and 4.0x crops, BGR at [0,1], the ensemble, index 1 as real) it labelled every genuine face in a real 3:4 photo as an attack at 0.95 confidence, and returned near-identical outputs for faces from 38 to 175 px across. The ONNX export is faithful to the PyTorch original; the original behaves the same way here. Presumably tied to the capture conditions it was trained on. A liveness signal that's confidently wrong is worse than none.

Face search against the web. Matching a face against Instagram or any other index means identifying strangers who never enrolled, which is the Clearview shape. There's also no wire to connect: Instagram exposes no face search and Google Lens has no public API.

Two things do the useful parts. /persons/search-by-face searches the registry, same embeddings, ranked instead of thresholded, scoped to people enrolled on purpose. /lookup-region searches the web for objects, gated on the face check.

License

MIT

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages