A multi-agent music discovery assistant powered by Spotify MCP.
Crate Diggr is an AI music assistant that connects to Spotify through MCP, analyzes listening history, explains music context, recommends similar tracks, and prepares playlists with human approval.
The goal of this project is not only to call an LLM or Spotify API. The goal is to build a production-style AI system around music discovery using:
- MCP tool integration
- LangGraph multi-agent orchestration
- shared state between agents
- structured tool calling
- observability
- evals / quality metrics
- cost and latency awareness
- human-in-the-loop actions
In DJ culture, crate digging means searching through vinyl crates to find rare, interesting, or hidden tracks.
Crate Diggr brings this idea to AI:
Old crate digging:
DJ searches vinyl boxes for hidden gems.
Crate Diggr:
AI searches your Spotify taste, listening history, genres, moods, and similar tracks.
Crate Diggr should help users answer questions like:
What is the story behind this track?
Create a playlist similar to what I listened to this week.
Find darker tracks similar to this song.
Build a DJ-style playlist with warm-up, peak, and closing tracks.
Explain why I may like this artist based on my listening history.
The project is designed to practice production AI engineering concepts.
Main technical goals:
- Connect external tools using MCP
- Use LangGraph to orchestrate a multi-agent workflow
- Keep a shared state across agents
- Use typed schemas with Pydantic
- Add observability for each step
- Evaluate recommendation quality
- Avoid unsafe write actions without approval
- Prepare the architecture for future AWS/serverless deployment
flowchart TD
U[User Request] --> API[FastAPI API]
API --> G[LangGraph Workflow]
G --> S[Supervisor Agent]
S --> H[Listening History Agent]
H --> M[Music Research Agent]
M --> R[Recommendation Agent]
R --> E[Evaluator Agent]
R --> P{User approved playlist creation?}
P -->|Yes| PL[Playlist Agent]
P -->|No| RESP[Return Playlist Plan]
PL --> RESP
E --> RESP
H --> MCP[Spotify MCP Tools]
R --> MCP
PL --> MCP
G --> OBS[Observability Layer]
E --> EV[Evaluation Metrics]
flowchart LR
A[Supervisor Agent] --> B[Listening History Agent]
B --> C[Music Research Agent]
C --> D[Recommendation Agent]
D --> E[Evaluator Agent]
D --> F[Playlist Agent]
F --> E
E --> G[Final Response]
Controls the workflow and decides what steps are needed.
Responsibilities:
- receive the user request
- initialize shared state
- route the workflow
- decide if playlist creation should be planned or executed
- ensure human approval before write actions
Reads user listening data from Spotify through MCP.
Responsibilities:
- get recent tracks
- get top artists
- get top tracks
- get liked songs, if available
- extract taste signals from listening history
Example output:
{
"recent_tracks": [
{
"name": "Blue Monday",
"artist": "New Order",
"genre_hint": "synth-pop"
}
]
}Explains the music context.
Responsibilities:
- explain artist background
- explain track history
- identify genre, era, and scene
- connect the track with similar movements
- avoid hallucinated facts when no source is available
Example:
This track fits a post-punk and synth-pop context, with strong electronic rhythm and dancefloor influence.
Generates recommendations based on user taste and request.
Responsibilities:
- use current track or listening history as seed
- search similar tracks
- rank recommendations
- explain why each track matches
- prepare a playlist plan
Example output:
{
"recommended_tracks": [
{
"name": "A Forest",
"artist": "The Cure",
"reason": "Dark atmosphere and strong bassline."
}
]
}Creates or updates a Spotify playlist.
Important rule:
This agent should only run after human approval.
Responsibilities:
- receive approved playlist plan
- create playlist using Spotify MCP
- add selected tracks
- return playlist id or URL
Checks the quality of the output.
Responsibilities:
- check if recommendations match the user request
- check if explanations are grounded
- check playlist coherence
- calculate basic quality metrics
- detect fallback cases
The graph uses shared state to pass information between agents.
flowchart TD
State[MusicAgentState]
State --> A[user_request]
State --> B[listening_history]
State --> C[current_track]
State --> D[music_context]
State --> E[recommended_tracks]
State --> F[playlist_plan]
State --> G[evaluation]
State --> H[approved_for_write_actions]
State --> I[created_playlist_id]
State --> J[errors]
Example state:
from typing import NotRequired, TypedDict
class MusicAgentState(TypedDict):
user_request: str
current_track: NotRequired[dict]
listening_history: NotRequired[list[dict]]
music_context: NotRequired[str]
recommended_tracks: NotRequired[list[dict]]
playlist_plan: NotRequired[dict]
evaluation: NotRequired[dict]
approved_for_write_actions: NotRequired[bool]
created_playlist_id: NotRequired[str | None]
errors: NotRequired[list[str]]crate-diggr/
├── crate_diggr/
│ ├── __init__.py
│ ├── main.py
│ ├── settings.py
│ │
│ ├── api/
│ │ ├── __init__.py
│ │ ├── app.py
│ │ ├── routes.py
│ │ └── schemas.py
│ │
│ ├── graph/
│ │ ├── __init__.py
│ │ ├── state.py
│ │ ├── workflow.py
│ │ └── routing.py
│ │
│ ├── agents/
│ │ ├── __init__.py
│ │ ├── supervisor.py
│ │ ├── listening_history.py
│ │ ├── music_research.py
│ │ ├── recommendation.py
│ │ ├── playlist.py
│ │ └── evaluator.py
│ │
│ ├── tools/
│ │ ├── __init__.py
│ │ ├── spotify_tools.py
│ │ ├── music_research_tools.py
│ │ └── validation_tools.py
│ │
│ ├── mcp/
│ │ ├── __init__.py
│ │ ├── client.py
│ │ ├── spotify_client.py
│ │ └── schemas.py
│ │
│ ├── spotify/
│ │ ├── __init__.py
│ │ ├── models.py
│ │ ├── mapper.py
│ │ └── service.py
│ │
│ ├── observability/
│ │ ├── __init__.py
│ │ ├── logging.py
│ │ ├── tracing.py
│ │ └── metrics.py
│ │
│ ├── evals/
│ │ ├── __init__.py
│ │ ├── dataset.py
│ │ ├── metrics.py
│ │ └── runner.py
│ │
│ └── storage/
│ ├── __init__.py
│ ├── database.py
│ ├── models.py
│ └── repositories.py
│
├── tests/
│ ├── test_workflow.py
│ ├── test_agents.py
│ ├── test_tools.py
│ └── test_evals.py
│
├── docs/
│ ├── architecture.md
│ ├── evals.md
│ └── mcp.md
│
├── scripts/
│ ├── run_api.sh
│ └── run_evals.sh
│
├── .env.example
├── .gitignore
├── LICENSE
├── README.md
└── pyproject.toml
FastAPI layer.
Responsible for:
- HTTP endpoints
- request/response schemas
- API validation
- calling the LangGraph workflow
Example endpoints:
GET /health
POST /recommend
POST /playlist/plan
POST /playlist/create
LangGraph orchestration layer.
Responsible for:
- shared state
- workflow definition
- conditional routing
- graph compilation
Main file:
workflow.py
Agent logic.
Each file represents one agent/node in the graph.
Agents should be small and focused.
Example:
recommendation.py
Should only handle recommendation logic, not Spotify authentication or API details.
Tool functions used by agents.
Examples:
- get recent tracks
- search similar tracks
- validate playlist coherence
- create playlist
- search music context
Tools should be simple, typed, and testable.
MCP integration layer.
Responsible for:
- connecting to MCP servers
- calling Spotify MCP tools
- mapping MCP responses
- handling MCP errors
This layer should isolate external tool details from the agents.
Spotify-specific domain logic.
Responsible for:
- track models
- playlist models
- mapping Spotify responses
- Spotify-specific business logic
Logging, tracing, and metrics.
Responsible for:
- structured logs
- step latency
- tool call tracking
- error tracking
- cost/token tracking in future versions
Evaluation logic.
Responsible for:
- golden dataset
- recommendation quality metrics
- playlist coherence metrics
- hallucination checks
- faithfulness checks
- regression evals
Persistence layer.
Can be SQLite/PostgreSQL later.
Responsible for:
- caching recommendations
- storing user taste profiles
- storing eval results
- storing playlist history
- storing tool call logs
sequenceDiagram
participant User
participant API as FastAPI
participant Graph as LangGraph
participant History as Listening History Agent
participant Research as Music Research Agent
participant Rec as Recommendation Agent
participant Eval as Evaluator Agent
participant Spotify as Spotify MCP
User->>API: Request recommendation
API->>Graph: Invoke workflow with initial state
Graph->>History: Get listening history
History->>Spotify: Call Spotify MCP tools
Spotify-->>History: Recent tracks / top tracks
History-->>Graph: Update state
Graph->>Research: Analyze music context
Research-->>Graph: Update context
Graph->>Rec: Generate recommendations
Rec->>Spotify: Search similar tracks
Spotify-->>Rec: Candidate tracks
Rec-->>Graph: Playlist plan
Graph->>Eval: Evaluate output
Eval-->>Graph: Quality metrics
Graph-->>API: Final state
API-->>User: Recommendations + explanation
Crate Diggr can suggest playlists freely, but it should not create or modify Spotify playlists without approval.
flowchart TD
A[Recommendation Agent] --> B[Playlist Plan]
B --> C{User approved?}
C -->|No| D[Return suggestion only]
C -->|Yes| E[Playlist Agent]
E --> F[Create playlist through Spotify MCP]
This is important because playlist creation is a write action.
The system should follow this rule:
The agent can plan and suggest, but the user must approve write actions.
Crate Diggr should measure recommendation quality and system behavior.
| Metric | Meaning |
|---|---|
| Recommendation relevance | Do the tracks match the requested vibe? |
| Playlist coherence | Do the tracks make sense together? |
| Faithfulness | Are explanations grounded in available data? |
| Hallucination rate | Did the agent invent music facts? |
| Tool success rate | Did Spotify MCP calls work? |
| Latency | How long did the workflow take? |
| Cost per request | How expensive was the request? |
| Fallback rate | How often did the system need a safe/default path? |
| User approval rate | How often did users approve playlist creation? |
Main eval phrase:
Manual testing is not enough. Crate Diggr uses evals to measure recommendation quality, faithfulness, latency, cost, tool failures, and fallback rate.
The system should track both system-level and AI-specific events.
- request latency
- node latency
- API errors
- MCP tool errors
- retry count
- timeout count
- prompt tokens
- completion tokens
- model latency
- cost per request
- recommendation count
- evaluation score
- hallucination risk
- fallback rate
Example structured log:
{
"event": "agent_step_finished",
"step": "recommendation",
"status": "success",
"latency_ms": 842,
"tool_calls": 2,
"recommendation_count": 10
}Crate Diggr should avoid unnecessary model calls.
Initial strategies:
- measure token usage first
- reduce prompt context
- use structured prompts
- cache listening history
- cache music context explanations
- cache repeated recommendations
- use cheaper models for simple classification
- use stronger models only for complex reasoning
- avoid playlist creation calls without user approval
Main phrase:
Optimization without measurement is risky. First measure token usage, cost, latency, and quality.
Initial MCP-related tools:
spotify.get_recent_tracks
spotify.get_top_tracks
spotify.get_top_artists
spotify.search_tracks
spotify.create_playlist
spotify.add_tracks_to_playlist
Expected internal tool wrapper:
async def get_recent_tracks(limit: int = 20) -> list[Track]:
...The MCP layer should be isolated inside:
crate_diggr/mcp/
Agents should not know the low-level MCP details.