Skip to content

Repository files navigation

Refabric Brand Intelligence Agent

A generic, config-driven Brand DNA agent for fashion brands. Given a brand name, website, and optional public social URLs, it collects fashion-relevant imagery and text, deduplicates/analyzes visuals, and renders a strategist-ready PDF dossier.

Development Methodology

This case study was implemented with a small, evidence-first methodology:

  • SDD-lite: the case-study brief was translated into an acceptance contract before final packaging; see docs/acceptance.md.
  • Characterization-first + TDD: known bad behaviors from reviews were locked with regression tests before fixes, then implemented with red/green/refactor.
  • Contract tests: CLI behavior, runner outputs, report sections, and summary fields are protected with deterministic tests.
  • Product-delivery audit: final readiness requires sample-run evidence, honest limitations, package checks, Docker help, and secret scanning.

Quick Start

The reviewer-friendly path does not require editing YAML:

uv sync --extra dev
uv run branddna run --name "Everlane" --website-url https://www.everlane.com

Add a public social URL only when you have one:

uv run branddna run \
  --name "Everlane" \
  --website-url https://www.everlane.com \
  --social-url https://www.instagram.com/everlane/

Repeatable YAML configs are still supported for sample and production-style runs:

uv run branddna run --config configs/brands/everlane.yaml

Run both sample configs:

uv run branddna run-samples

Smoke a small ad-hoc brand run:

uv run branddna run \
  --name "Example Brand" \
  --website-url "https://example.com" \
  --max-pages 3 \
  --max-downloads 80 \
  --output-root runs_smoke

Docker does not require docker exec. The image entrypoint is branddna, so reviewers can run the CLI directly:

docker build -t refabric-branddna:local .
mkdir -p runs
docker run --rm \
  --user "$(id -u):$(id -g)" \
  -v "$PWD/runs:/app/runs" \
  refabric-branddna:local run \
  --name "Everlane" \
  --website-url https://www.everlane.com

For a custom YAML file, mount the config directory read-only:

mkdir -p runs
docker run --rm \
  --user "$(id -u):$(id -g)" \
  -v "$PWD/runs:/app/runs" \
  -v "$PWD/configs:/app/configs:ro" \
  refabric-branddna:local run \
  --config configs/brands/everlane.yaml

The --user flag keeps generated files owned by the host user on typical Linux/macOS Docker installs. Create the host output directory before running so Docker does not create it as a root-owned directory. If your Docker environment manages bind-mount permissions differently, --user can be omitted.

CLI Usage Reference

Discover commands and options:

uv run branddna --help
uv run branddna run --help
uv run branddna run-samples --help

Primary commands:

  • branddna run --name ... --website-url ... [--social-url ...]: primary reviewer path for a new brand without editing code.
  • branddna run --config <yaml>: advanced/repeatable path for saved brand configs.
  • branddna run-samples: run the two default sample brands, Everlane and Paloma Wool.

Useful options:

  • --output-root: choose where run artifacts are written; default is runs.
  • --max-pages: temporary crawl page cap for smoke/debug runs.
  • --max-downloads: temporary image download cap for smoke/debug runs.
  • repeated --social-url: add public social URLs for ad-hoc brands.

What It Produces

Each run writes:

  • runs/{brand_slug}/{timestamp}/raw/pages.jsonl
  • runs/{brand_slug}/{timestamp}/images/
  • runs/{brand_slug}/{timestamp}/metadata/image_candidates.jsonl
  • runs/{brand_slug}/{timestamp}/metadata/images.jsonl
  • runs/{brand_slug}/{timestamp}/analysis/brand_analysis.json
  • runs/{brand_slug}/{timestamp}/analysis/clusters.json
  • runs/{brand_slug}/{timestamp}/report/brand_dna.html
  • runs/{brand_slug}/{timestamp}/report/brand_dna.pdf
  • runs/{brand_slug}/{timestamp}/trace.log.jsonl
  • runs/{brand_slug}/{timestamp}/run_summary.json

run_summary.json includes acceptance evidence such as page-type coverage, accepted images by source, fashion-filter method counts, coverage limitations, degradation reasons, the PDF path, and an execution_success_percent score. That score measures execution readiness only; it does not replace manual semantic review of the Brand DNA content.

Configuration

Brand onboarding can be direct CLI input or a repeatable YAML config. The fastest path is:

uv run branddna run --name "Brand Name" --website-url https://brand.example

YAML is useful when the brand should be rerun with stable crawl/image limits. Example:

name: Everlane
slug: everlane
website_url: https://www.everlane.com/
social_urls:
  - https://www.instagram.com/everlane/
allowed_domains:
  - everlane.com
crawl:
  max_pages: 180
  max_depth: 2
  rate_limit: 0.7
image:
  min_short_side: 512
  target_count: 100
  max_downloads: 1800
llm:
  enabled: false
  max_cost_usd: 5

The 512px shorter-side threshold is the default because it is large enough for downstream visual embedding and design references while avoiding low-res thumbnails, logos, and layout artifacts. The Everlane sample uses a larger page/download cap to absorb public-site variability while keeping the same generic, config-driven crawler.

Architecture Summary

Pipeline:

  1. WebCollector: public, allowlisted, generic HTML/sitemap crawl.
  2. SocialCollector: public-only social URL fetch; no login or bypass; empty public shells are marked degraded.
  3. ImageCandidateExtractor: img, srcset, OpenGraph, JSON-LD.
  4. QualityFilter: rejects images below min_short_side.
  5. FashionFilter: local CLIP when installed; deterministic heuristic fallback, with the active method reported in the summary and PDF.
  6. Deduper: perceptual hash plus embedding similarity.
  7. Analyzer: color palette, category mix, styling, voice, demographic and psychographic audience cues.
  8. Clusterer: 3-6 bounded aesthetic clusters when enough images exist.
  9. ReportRenderer: Jinja HTML plus PDF output with Method and Evidence and Limitations sections.

Full design rationale is in docs/architecture.md. Product and technical flows are in docs/flows.md. Testing strategy is in docs/testing-strategy.md. Acceptance mapping is in docs/acceptance.md. Final sample evidence is in docs/sample-runs.md. Five-brand POC/generalization evidence is in docs/poc-five-brand-validation.md. Submission-use notice is in NOTICE.md.

Submission Packages

Generated run archives are committed under submission_outputs/ through Git LFS:

  • {brand}_{timestamp}.zip: complete sample-run output for one brand, including raw pages, image metadata, accepted images, analysis JSON, HTML/PDF report, trace log, and run_summary.json.
  • poc/{brand}_{timestamp}.zip: five-brand POC/generalization outputs generated from the same direct CLI path.

These ZIPs are not runtime dependencies. They exist so a reviewer can inspect the generated PDFs and raw evidence without rerunning slow public website crawls. Raw runs/ directories remain ignored to avoid duplicating the same large evidence twice. No source ZIP is included because the public repository is the source of truth.

Reviewer clone note:

git clone https://github.com/rcnsnr/refabric-brand-intelligence-agent
cd refabric-brand-intelligence-agent
git lfs install
git lfs pull

If the archives appear as small pointer files after cloning, Git LFS objects have not been materialized yet; rerun git lfs install and git lfs pull.

Reviewer Walkthrough Path

Fastest way to evaluate the submission:

  1. Read this README plus docs/architecture.md.

  2. Inspect docs/acceptance.md and docs/sample-runs.md.

  3. Open both generated PDFs from the sample ZIPs.

  4. Spot-check each sample run_summary.json, trace.log.jsonl, and metadata/images.jsonl.

  5. Run the direct CLI path:

    uv run branddna run \
      --name "Everlane" \
      --website-url https://www.everlane.com \
      --max-pages 3
  6. Run uv run branddna --help, the quality gates below, and optionally the Docker smoke path.

Platform-Neutral / GitLab Review

The project has GitHub Actions CI, but the runtime is not GitHub-specific. If a reviewer imports the repository into GitLab, use the same local or Docker commands from this README. Make sure Git LFS is enabled on the target platform and run git lfs install && git lfs pull after cloning so the committed sample ZIP archives are materialized locally.

Social, Anti-Bot, Rate Limits, and ToS

  • Public-only collection is locked for this case study.
  • No account login, credential use, CAPTCHA bypass, proxy rotation, or evasion.
  • The crawler performs a best-effort robots.txt fetch and skips disallowed URLs when a valid public policy is available.
  • Instagram is attempted as a public page.
  • If social returns 401, 403, 429, or a title-only / empty public payload, the run is marked degraded and continues.
  • crawl.rate_limit defaults to a respectful delay between requests.
  • Site down, social blocked, no images found, and low image count are recorded in run_summary.json and trace.log.jsonl.

AI / Cost Policy

Core collection, filtering, clustering, and report generation are local. The default configs set llm.enabled: false.

No .env file or API key is required for the default local or Docker run. .env.example exists only as a secret-hygiene reminder.

If LLM synthesis is later enabled, keep it under the assignment cap:

  • hard cap: $5
  • API key source: environment/secret manager only
  • no API key in source, configs, logs, or Docker build args
  • deterministic report fallback remains required

Development MCP Notes

Context7 is a development-time documentation MCP, not a runtime dependency. If used, configure documentation-tool keys in the IDE/Codex/Windsurf secret surface. Do not commit real values here.

Local agent helper overlays such as scripts/ac, .codex/, AGENTS.md, and RTK.md may exist in the maintainer's working directory. They are ignored by git, excluded by .dockerignore, and excluded from the product source archive. They are not runtime dependencies for the Brand DNA CLI.

Useful tools for this repo:

  • Context7: optional for up-to-date library docs; requires a valid key.
  • Filesystem: useful for local edits.
  • Tavily/web: useful for public source verification.
  • Code review graph: optional after the repo is initialized and indexed.

Potential missing optional MCP:

  • Firecrawl/structured extraction is configured elsewhere on the machine but is not exposed in this current tool session. The implementation does not depend on it.

Quality Gates

uv run ruff format --check src tests
uv run ruff check src tests
uv run mypy src
uv run pytest

Smoke run:

uv run branddna run \
  --name "Everlane" \
  --website-url https://www.everlane.com \
  --max-pages 3 \
  --output-root runs_smoke

Docker smoke. This fast command is expected to be degraded because example.com has no fashion imagery; it only proves the containerized CLI path:

docker build -t refabric-branddna:local .
docker run --rm refabric-branddna:local --help
mkdir -p runs_smoke
docker run --rm \
  --user "$(id -u):$(id -g)" \
  -v "$PWD/runs_smoke:/app/runs" \
  refabric-branddna:local run \
  --name "Docker Smoke" \
  --website-url https://example.com \
  --max-pages 1 \
  --max-downloads 1

Known Limitations

  • Public websites may lazy-load images or block automated clients.
  • CLIP is optional because local model downloads can be heavy.
  • Install CLIP only when needed with uv sync --extra dev --extra clip.
  • Without CLIP, deterministic scoring and color-histogram embeddings are used.
  • Reports state whether accepted images came from heuristic_text_resolution or local_clip_zero_shot.
  • Visual cluster descriptions summarize observable cues conservatively.
  • This is a CLI case-study, not a production multi-tenant crawler platform.

About

Generic Brand Intelligence Agent for fashion brands: public web/social collection, fashion image filtering, Brand DNA analysis, and strategist-ready PDF dossiers.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages