This skill gives an agent a repeatable workbench for image BBOX tasks:
- read image dimensions;
- draw a global coordinate grid;
- crop and zoom coarse regions while preserving original coordinates;
- draw confirmed and candidate boxes back onto the original image.
It does not perform object detection. The agent still decides what the object is; the script makes coordinate work auditable.
mkdir -p ~/.agents/skills
git clone https://github.com/liangming99/visual-bbox-locator.git ~/.agents/skills/visual-bbox-locator
cd ~/.agents/skills/visual-bbox-locator
python3 -m pip install -r requirements.txt
python3 scripts/bbox_workspace.py --helpSKILL.md- model-facing process and rules.scripts/bbox_workspace.py- deterministic image operations.examples/rebar-yard-person.md- example from the steel-yard personnel image.requirements.txt- Python runtime dependency for the image workbench.
python3 scripts/bbox_workspace.py info image.png
python3 scripts/bbox_workspace.py grid image.png --out work/full_grid.png
python3 scripts/bbox_workspace.py crop image.png --bbox 330,410,430,590 --scale 4 --out work/crop_A.png --label A
python3 scripts/bbox_workspace.py overlay image.png --boxes work/boxes.json --out work/overlay.pngMIT