You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
A ground-up implementation of the core ideas behind Kubernetes — written in Go, built phase by phase so every line of code is something you understand and can explain.
flowchart TB
CLI["kl CLI"]
WEB["React UI<br/>Vite + React Query"]
subgraph SCHED["scheduler — control plane, :8080"]
REC["reconciler loop<br/>every 5s"]
REG["node registry<br/>heartbeat + dead-node detection"]
ROLL["rollout controller<br/>wave-by-wave"]
DISC["service discovery"]
STATE[("StateStore<br/>WorkloadSpec + instances")]
end
subgraph AGENT["agent — worker node, :8081"]
RUN["Docker runner"]
PROBE["health prober"]
HB["heartbeat sender<br/>every 3s"]
end
DOCKER["Docker daemon"]
CLI -->|REST| SCHED
WEB -->|REST| SCHED
REC --> STATE
ROLL --> STATE
REG --> STATE
REC -->|"POST /run, /stop"| RUN
HB -->|"POST /heartbeat<br/>full container snapshot"| REG
RUN --> DOCKER
PROBE --> DOCKER
Loading
Architecture
Components
Binary
Path
Role
scheduler
cmd/scheduler/
Central control plane. One per cluster.
agent
cmd/agent/
Worker node daemon. One per machine.
kl
cmd/kl/
CLI client (thin HTTP wrapper).
React UI
ui/
Dashboard (Vite + React Query).
Packages
Package
What it does
pkg/types
All shared domain types (WorkloadSpec, ContainerInstance, etc.)
internal/agent
Docker SDK wrapper, HTTP health prober, agent HTTP server
internal/scheduler
Node registry, state store, scheduling, reconciler, rollout controller, HTTP server
Request flow — deploying a workload
sequenceDiagram
participant kl as kl CLI
participant s as scheduler
participant st as StateStore
participant a as agent
participant d as Docker
kl->>s: POST /deploy (name, image, replicas)
s->>st: upsert WorkloadSpec
s->>s: reconcileWorkload()
s->>st: AliveNodes() then round-robin pick
s->>a: POST /run
a->>d: pull
a->>d: container create
a->>d: container start
a->>a: StartProbe()
a-->>s: heartbeat carries the new instance
Loading
Reconciliation loop
Runs every 5 seconds in the scheduler. For each workload it computes:
Agents send POST /heartbeat every 3 seconds. Each heartbeat carries a full snapshot of every container running on that node. The scheduler replaces the node's instance state atomically. If 3 consecutive heartbeats are missed (9 s), the node is marked DEAD and excluded from scheduling.
Rolling updates
Wave-by-wave strategy:
flowchart TD
W["start 1 new container<br/>with the new image"] --> WAIT{"health_ok<br/>within 60s?"}
WAIT -->|yes| STOP["stop 1 old container"]
STOP --> MORE{"old containers left?"}
MORE -->|yes| W
MORE -->|no| DONE["rollout complete"]
WAIT -->|"no — unhealthy or timeout"| ABORT["abort rollout"]
Loading
Quick start
Prerequisites
Go ≥ 1.21
Docker running locally
Node.js ≥ 18 (UI only)
Build everything
cd kubelite
make build # compiles agent, scheduler, kl → bin/
Run locally
# Terminal 1 — start the scheduler (control plane)
make run-scheduler
# Terminal 2 — start an agent (worker node)
make run-agent
# (or both at once in tmux)
make dev
cd ui
npm install
npm run dev # http://localhost:5173
Configuration
All configuration is via environment variables — no config files.
Scheduler
Variable
Default
Description
KL_LISTEN
:8080
Address the scheduler binds to
Agent
Variable
Default
Description
KL_LISTEN
:8081
Address the agent HTTP server binds to
KL_SCHEDULER
localhost:8080
Scheduler address (host:port, no scheme)
KL_NODE_ID
hostname
Unique name for this node
CLI
Variable
Default
Description
KL_SERVER
http://localhost:8080
Scheduler address
CLI reference
kl [--server http://host:port] <command> [flags]
Command
Description
nodes
List all registered worker nodes
deploy
Deploy (or update) a workload
workloads
List all workloads with replica counts
status <id>
Workload detail + running instances
scale <id> --replicas N
Change desired replica count
delete <id>
Stop all containers and remove workload
logs <container-id>
Stream container logs
rollout <id> --image IMAGE
Start a rolling image update
rollout-status <id>
Show current rollout progress
discover <name>
List live service endpoints by workload name
deploy flags
--name workload name (required)
--image container image (required)
--id workload ID (auto-generated if empty)
--replicas desired replica count (default 1)
--restart-policy Always | OnFailure | Never (default Always)
--port HOST:CONTAINER (repeatable)
--env KEY=VALUE (repeatable)
--health-path HTTP probe path (e.g. /health)
--health-port HTTP probe port
Scheduler API
All endpoints accept and return JSON.
Agent-facing
Method
Path
Body
Description
POST
/register
RegisterRequest
One-time agent registration
POST
/heartbeat
HeartbeatRequest
Periodic state sync. Returns 422 if node is unknown → agent re-registers
User-facing
Method
Path
Body
Description
POST
/deploy
WorkloadSpec
Create or update a workload
GET
/workloads
—
List all workloads (summary)
GET
/workloads/:id
—
Workload detail + live instances
DELETE
/workloads/:id
—
Stop all containers + remove
PUT
/workloads/:id/scale
{"replicas":N}
Adjust replica count
GET
/nodes
—
All nodes with status
GET
/discover/:name
—
Live endpoints for a workload name
POST
/rollout
RolloutSpec
Start a rolling update
GET
/rollout/:workloadID
—
Rollout state
GET
/logs/:containerID
—
Stream logs (proxied to owning agent)
GET
/health
—
Scheduler liveness probe
Agent API (direct access)
Method
Path
Description
POST
/run
Start a container
POST
/stop/:id
Stop + remove a container
GET
/status/:id
Container state + exit code
GET
/logs/:id
Streaming stdout/stderr
GET
/health
Agent liveness probe
Key types
// WorkloadSpec — desired state submitted by the usertypeWorkloadSpecstruct {
IDstringNamestringImagestringReplicasintEnvmap[string]stringPorts []PortMappingRestartPolicyRestartPolicy// Always | OnFailure | NeverHealthCheck*HealthCheckSpec
}
// ContainerInstance — actual state reported by an agenttypeContainerInstancestruct {
IDstringWorkloadIDstringNodeIDstringImagestringStateContainerState// running | stopped | exited | unknownHealthHealthStatus// unknown | starting | healthy | unhealthyIPstring// Docker bridge IPStartedAt time.Time
}
// RolloutSpec — request a rolling image updatetypeRolloutSpecstruct {
WorkloadIDstringNewImagestringMaxUnavailableintMaxSurgeint
}
Make targets
make build build all three binaries → bin/
make agent build agent only
make scheduler build scheduler only
make kl build CLI only
make run-scheduler start scheduler on :8080
make run-agent start agent on :8081
make dev scheduler + agent in a tmux session
make test go test ./...
make test-race go test -race ./...
make vet go vet ./...
make tidy go mod tidy
make lint golangci-lint run
make check vet + tidy + git-diff guard (for CI)
make install go install all binaries to $GOPATH/bin
make clean remove bin/
Phase roadmap
Phase
What you build
Core concept learned
1
Agent: POST /run, GET /status
Docker SDK, HTTP API
2
Scheduler + heartbeat
Distributed state, dead-node detection
3
Reconcile loop, scheduling
Desired vs actual, bin-packing
4
Health checks, restart policy
Idempotency, drift handling
5
Service discovery, log streaming
Name → endpoint resolution
6
Rolling updates, rollback
Controller pattern
7
React UI + CLI polish
Observability
About
A ground-up implementation of Kubernetes' core ideas in Go — reconciler loop, scheduler, rollout controller, health probes, service discovery.