Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Visual BBOX Locator

This skill gives an agent a repeatable workbench for image BBOX tasks:

  1. read image dimensions;
  2. draw a global coordinate grid;
  3. crop and zoom coarse regions while preserving original coordinates;
  4. draw confirmed and candidate boxes back onto the original image.

It does not perform object detection. The agent still decides what the object is; the script makes coordinate work auditable.

Install

mkdir -p ~/.agents/skills
git clone https://github.com/liangming99/visual-bbox-locator.git ~/.agents/skills/visual-bbox-locator
cd ~/.agents/skills/visual-bbox-locator
python3 -m pip install -r requirements.txt
python3 scripts/bbox_workspace.py --help

Files

  • SKILL.md - model-facing process and rules.
  • scripts/bbox_workspace.py - deterministic image operations.
  • examples/rebar-yard-person.md - example from the steel-yard personnel image.
  • requirements.txt - Python runtime dependency for the image workbench.

Minimal Example

python3 scripts/bbox_workspace.py info image.png
python3 scripts/bbox_workspace.py grid image.png --out work/full_grid.png
python3 scripts/bbox_workspace.py crop image.png --bbox 330,410,430,590 --scale 4 --out work/crop_A.png --label A
python3 scripts/bbox_workspace.py overlay image.png --boxes work/boxes.json --out work/overlay.png

License

MIT

About

Visual coordinate workbench for bounding box tasks

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages