A generic, config-driven Brand DNA agent for fashion brands. Given a brand name, website, and optional public social URLs, it collects fashion-relevant imagery and text, deduplicates/analyzes visuals, and renders a strategist-ready PDF dossier.
This case study was implemented with a small, evidence-first methodology:
- SDD-lite: the case-study brief was translated into an acceptance contract
before final packaging; see
docs/acceptance.md. - Characterization-first + TDD: known bad behaviors from reviews were locked with regression tests before fixes, then implemented with red/green/refactor.
- Contract tests: CLI behavior, runner outputs, report sections, and summary fields are protected with deterministic tests.
- Product-delivery audit: final readiness requires sample-run evidence, honest limitations, package checks, Docker help, and secret scanning.
The reviewer-friendly path does not require editing YAML:
uv sync --extra dev
uv run branddna run --name "Everlane" --website-url https://www.everlane.comAdd a public social URL only when you have one:
uv run branddna run \
--name "Everlane" \
--website-url https://www.everlane.com \
--social-url https://www.instagram.com/everlane/Repeatable YAML configs are still supported for sample and production-style runs:
uv run branddna run --config configs/brands/everlane.yamlRun both sample configs:
uv run branddna run-samplesSmoke a small ad-hoc brand run:
uv run branddna run \
--name "Example Brand" \
--website-url "https://example.com" \
--max-pages 3 \
--max-downloads 80 \
--output-root runs_smokeDocker does not require docker exec. The image entrypoint is branddna,
so reviewers can run the CLI directly:
docker build -t refabric-branddna:local .
mkdir -p runs
docker run --rm \
--user "$(id -u):$(id -g)" \
-v "$PWD/runs:/app/runs" \
refabric-branddna:local run \
--name "Everlane" \
--website-url https://www.everlane.comFor a custom YAML file, mount the config directory read-only:
mkdir -p runs
docker run --rm \
--user "$(id -u):$(id -g)" \
-v "$PWD/runs:/app/runs" \
-v "$PWD/configs:/app/configs:ro" \
refabric-branddna:local run \
--config configs/brands/everlane.yamlThe --user flag keeps generated files owned by the host user on typical
Linux/macOS Docker installs. Create the host output directory before running so
Docker does not create it as a root-owned directory. If your Docker environment
manages bind-mount permissions differently, --user can be omitted.
Discover commands and options:
uv run branddna --help
uv run branddna run --help
uv run branddna run-samples --helpPrimary commands:
branddna run --name ... --website-url ... [--social-url ...]: primary reviewer path for a new brand without editing code.branddna run --config <yaml>: advanced/repeatable path for saved brand configs.branddna run-samples: run the two default sample brands, Everlane and Paloma Wool.
Useful options:
--output-root: choose where run artifacts are written; default isruns.--max-pages: temporary crawl page cap for smoke/debug runs.--max-downloads: temporary image download cap for smoke/debug runs.- repeated
--social-url: add public social URLs for ad-hoc brands.
Each run writes:
runs/{brand_slug}/{timestamp}/raw/pages.jsonlruns/{brand_slug}/{timestamp}/images/runs/{brand_slug}/{timestamp}/metadata/image_candidates.jsonlruns/{brand_slug}/{timestamp}/metadata/images.jsonlruns/{brand_slug}/{timestamp}/analysis/brand_analysis.jsonruns/{brand_slug}/{timestamp}/analysis/clusters.jsonruns/{brand_slug}/{timestamp}/report/brand_dna.htmlruns/{brand_slug}/{timestamp}/report/brand_dna.pdfruns/{brand_slug}/{timestamp}/trace.log.jsonlruns/{brand_slug}/{timestamp}/run_summary.json
run_summary.json includes acceptance evidence such as page-type coverage,
accepted images by source, fashion-filter method counts, coverage limitations,
degradation reasons, the PDF path, and an execution_success_percent score.
That score measures execution readiness only; it does not replace manual
semantic review of the Brand DNA content.
Brand onboarding can be direct CLI input or a repeatable YAML config. The fastest path is:
uv run branddna run --name "Brand Name" --website-url https://brand.exampleYAML is useful when the brand should be rerun with stable crawl/image limits. Example:
name: Everlane
slug: everlane
website_url: https://www.everlane.com/
social_urls:
- https://www.instagram.com/everlane/
allowed_domains:
- everlane.com
crawl:
max_pages: 180
max_depth: 2
rate_limit: 0.7
image:
min_short_side: 512
target_count: 100
max_downloads: 1800
llm:
enabled: false
max_cost_usd: 5The 512px shorter-side threshold is the default because it is large enough
for downstream visual embedding and design references while avoiding low-res
thumbnails, logos, and layout artifacts.
The Everlane sample uses a larger page/download cap to absorb public-site
variability while keeping the same generic, config-driven crawler.
Pipeline:
WebCollector: public, allowlisted, generic HTML/sitemap crawl.SocialCollector: public-only social URL fetch; no login or bypass; empty public shells are marked degraded.ImageCandidateExtractor:img,srcset, OpenGraph, JSON-LD.QualityFilter: rejects images belowmin_short_side.FashionFilter: local CLIP when installed; deterministic heuristic fallback, with the active method reported in the summary and PDF.Deduper: perceptual hash plus embedding similarity.Analyzer: color palette, category mix, styling, voice, demographic and psychographic audience cues.Clusterer: 3-6 bounded aesthetic clusters when enough images exist.ReportRenderer: Jinja HTML plus PDF output with Method and Evidence and Limitations sections.
Full design rationale is in docs/architecture.md.
Product and technical flows are in docs/flows.md.
Testing strategy is in docs/testing-strategy.md.
Acceptance mapping is in docs/acceptance.md.
Final sample evidence is in docs/sample-runs.md.
Five-brand POC/generalization evidence is in
docs/poc-five-brand-validation.md.
Submission-use notice is in NOTICE.md.
Generated run archives are committed under submission_outputs/ through Git LFS:
{brand}_{timestamp}.zip: complete sample-run output for one brand, including raw pages, image metadata, accepted images, analysis JSON, HTML/PDF report, trace log, andrun_summary.json.poc/{brand}_{timestamp}.zip: five-brand POC/generalization outputs generated from the same direct CLI path.
These ZIPs are not runtime dependencies. They exist so a reviewer can inspect
the generated PDFs and raw evidence without rerunning slow public website
crawls. Raw runs/ directories remain ignored to avoid duplicating the same
large evidence twice. No source ZIP is included because the public repository is
the source of truth.
Reviewer clone note:
git clone https://github.com/rcnsnr/refabric-brand-intelligence-agent
cd refabric-brand-intelligence-agent
git lfs install
git lfs pullIf the archives appear as small pointer files after cloning, Git LFS objects have
not been materialized yet; rerun git lfs install and git lfs pull.
Fastest way to evaluate the submission:
-
Read this README plus
docs/architecture.md. -
Inspect
docs/acceptance.mdanddocs/sample-runs.md. -
Open both generated PDFs from the sample ZIPs.
-
Spot-check each sample
run_summary.json,trace.log.jsonl, andmetadata/images.jsonl. -
Run the direct CLI path:
uv run branddna run \ --name "Everlane" \ --website-url https://www.everlane.com \ --max-pages 3 -
Run
uv run branddna --help, the quality gates below, and optionally the Docker smoke path.
The project has GitHub Actions CI, but the runtime is not GitHub-specific. If a
reviewer imports the repository into GitLab, use the same local or Docker
commands from this README. Make sure Git LFS is enabled on the target platform
and run git lfs install && git lfs pull after cloning so the committed sample
ZIP archives are materialized locally.
- Public-only collection is locked for this case study.
- No account login, credential use, CAPTCHA bypass, proxy rotation, or evasion.
- The crawler performs a best-effort
robots.txtfetch and skips disallowed URLs when a valid public policy is available. - Instagram is attempted as a public page.
- If social returns
401,403,429, or a title-only / empty public payload, the run is marked degraded and continues. crawl.rate_limitdefaults to a respectful delay between requests.- Site down, social blocked, no images found, and low image count are recorded
in
run_summary.jsonandtrace.log.jsonl.
Core collection, filtering, clustering, and report generation are local. The
default configs set llm.enabled: false.
No .env file or API key is required for the default local or Docker run.
.env.example exists only as a secret-hygiene reminder.
If LLM synthesis is later enabled, keep it under the assignment cap:
- hard cap:
$5 - API key source: environment/secret manager only
- no API key in source, configs, logs, or Docker build args
- deterministic report fallback remains required
Context7 is a development-time documentation MCP, not a runtime dependency. If used, configure documentation-tool keys in the IDE/Codex/Windsurf secret surface. Do not commit real values here.
Local agent helper overlays such as scripts/ac, .codex/, AGENTS.md, and
RTK.md may exist in the maintainer's working directory. They are ignored by
git, excluded by .dockerignore, and excluded from the product source archive.
They are not runtime dependencies for the Brand DNA CLI.
Useful tools for this repo:
- Context7: optional for up-to-date library docs; requires a valid key.
- Filesystem: useful for local edits.
- Tavily/web: useful for public source verification.
- Code review graph: optional after the repo is initialized and indexed.
Potential missing optional MCP:
- Firecrawl/structured extraction is configured elsewhere on the machine but is not exposed in this current tool session. The implementation does not depend on it.
uv run ruff format --check src tests
uv run ruff check src tests
uv run mypy src
uv run pytestSmoke run:
uv run branddna run \
--name "Everlane" \
--website-url https://www.everlane.com \
--max-pages 3 \
--output-root runs_smokeDocker smoke. This fast command is expected to be degraded because
example.com has no fashion imagery; it only proves the containerized CLI path:
docker build -t refabric-branddna:local .
docker run --rm refabric-branddna:local --help
mkdir -p runs_smoke
docker run --rm \
--user "$(id -u):$(id -g)" \
-v "$PWD/runs_smoke:/app/runs" \
refabric-branddna:local run \
--name "Docker Smoke" \
--website-url https://example.com \
--max-pages 1 \
--max-downloads 1- Public websites may lazy-load images or block automated clients.
- CLIP is optional because local model downloads can be heavy.
- Install CLIP only when needed with
uv sync --extra dev --extra clip. - Without CLIP, deterministic scoring and color-histogram embeddings are used.
- Reports state whether accepted images came from
heuristic_text_resolutionorlocal_clip_zero_shot. - Visual cluster descriptions summarize observable cues conservatively.
- This is a CLI case-study, not a production multi-tenant crawler platform.