Complete API documentation for the Forge backend application.
http://127.0.0.1:8000
Currently, the API does not require authentication. All endpoints are publicly accessible.
Base path: /api/models
Retrieve a list of all available Ollama models installed on the system.
Endpoint: GET /api/models/all
Response:
[
{
"name": "qwen2.5:14b",
"size": 10000000000,
"param_size": "14B"
},
{
"name": "llama2:7b",
"size": 5000000000,
"param_size": "7B"
}
]Status Codes:
200 OK: Successfully retrieved models500 Internal Server Error: Ollama not installed or not running
Example:
curl http://127.0.0.1:8000/api/models/allGet the currently active model configured in the application.
Endpoint: GET /api/models/current
Response:
{
"name": "qwen2.5:14b"
}Status Codes:
200 OK: Successfully retrieved current model
Example:
curl http://127.0.0.1:8000/api/models/currentChange the active model to a different installed model.
Endpoint: POST /api/models/change
Request Body:
{
"model_name": "llama2:7b"
}Response:
{
"message": "Success! model set to llama2:7b"
}Status Codes:
200 OK: Model successfully changed404 Not Found: Model not found among installed models500 Internal Server Error: Ollama not installed or not running
Example:
curl -X POST http://127.0.0.1:8000/api/models/change \
-H "Content-Type: application/json" \
-d '{"model_name": "llama2:7b"}'Download a new model from Ollama servers. Returns a streaming response with download progress.
Endpoint: POST /api/models/download/{model_name}
Path Parameters:
model_name(string): Name of the model to download (e.g., "qwen2.5:14b")
Response: Server-Sent Events (SSE) stream
Stream Format:
data: {"completed": 100, "total": 1000}
data: {"completed": 200, "total": 1000}
...
Status Codes:
200 OK: Download started successfully500 Internal Server Error:- Ollama not installed or not running
- Model does not exist on Ollama servers
Example:
curl -X POST http://127.0.0.1:8000/api/models/download/qwen2.5:14bNote: This endpoint streams progress updates. The download happens asynchronously.
Check if Ollama is running and accessible.
Endpoint: GET /api/models/alive
Response:
trueor
falseStatus Codes:
200 OK: Always returns (true or false)
Example:
curl http://127.0.0.1:8000/api/models/aliveBase path: /api/chat
Send a message to the AI agent and receive a response. Supports both streaming and non-streaming modes.
Endpoint: POST /api/chat
Request Body:
{
"message": "What is the capital of France?",
"session_id": "session_12345",
"stream": true
}Request Parameters:
message(string, required): The user's messagesession_id(string, optional): Session ID for maintaining conversation context. If not provided, a new session ID will be generated.stream(boolean, optional): Whether to stream the response. Default:false
Non-Streaming Response:
{
"response": "The capital of France is Paris.",
"session_id": "session_12345"
}Streaming Response: Server-Sent Events (SSE) stream
Stream Format:
data: {"content": "The", "type": "RunResponse", "tool_calls": [], "session_id": "session_12345", "tool_requiring_confirmation": null}
data: {"content": " capital", "type": "RunResponse", "tool_calls": [], "session_id": "session_12345", "tool_requiring_confirmation": null}
data: {"content": " of", "type": "RunResponse", "tool_calls": [], "session_id": "session_12345", "tool_requiring_confirmation": null}
...
data: [DONE]
Stream Response Fields:
content(string): Text chunk from the AI responsetype(string): Type of response chunktool_calls(array): Array of tool calls made by the agentsession_id(string): Session ID for this conversationtool_requiring_confirmation(object|null): Tool that requires user confirmation before execution
Tool Requiring Confirmation Format:
{
"tool_name": "search_internet",
"tool_id": "uuid-here",
"session_id": "session_12345",
"confirmed": false
}Note on tool confirmation: When a chunk carries a non-null tool_requiring_confirmation, the backend has paused the agent run and is waiting for the client to answer via POST /api/chat/confirm-tool. The SSE stream stays open; once the confirmation is received, the run resumes and streaming continues on the same connection. The response ends with data: [DONE] only after the entire run (including any resumed tool calls) has finished.
Status Codes:
200 OK: Request processed successfully500 Internal Server Error:- Ollama not installed or not running
- Error processing request
Example (Non-Streaming):
curl -X POST http://127.0.0.1:8000/api/chat \
-H "Content-Type: application/json" \
-d '{"message": "Hello, how are you?", "stream": false}'Example (Streaming):
curl -X POST http://127.0.0.1:8000/api/chat \
-H "Content-Type: application/json" \
-d '{"message": "Hello, how are you?", "stream": true}'Confirm or deny execution of a tool that requires user confirmation.
Endpoint: POST /api/chat/confirm-tool
Request Body:
{
"tool_id": "uuid-here",
"session_id": "session_12345",
"confirmed": true
}Request Parameters:
tool_id(string, required): ID of the tool requiring confirmationsession_id(string, required): Session ID for the conversationconfirmed(boolean, required): Whether to confirm tool execution
Response:
{
"message": "approved",
"session_id": "session_12345"
}message is "approved" when confirmed is true, "rejected" otherwise.
Status Codes:
200 OK: Confirmation resolved; the paused chat stream resumes404 Not Found: No tool is awaiting confirmation for this session409 Conflict: This tool confirmation has already been answered
How it works: When a chat stream encounters a tool that requires confirmation, the backend pauses the agent run and emits a chunk with a non-null tool_requiring_confirmation payload. The client answers through this endpoint, and the paused run resumes on the same streamed connection.
Example:
curl -X POST http://127.0.0.1:8000/api/chat/confirm-tool \
-H "Content-Type: application/json" \
-d '{"tool_id": "uuid-here", "session_id": "session_12345", "confirmed": true}'Base path: /api/utils
Get the current working directory of the backend server.
Endpoint: GET /api/utils/getcwd
Response:
{
"dir": "/path/to/backend/app"
}Status Codes:
200 OK: Successfully retrieved directory500 Internal Server Error: Failed to retrieve directory
Example:
curl http://127.0.0.1:8000/api/utils/getcwdAll errors follow a consistent format:
{
"detail": "Error message describing what went wrong"
}404 Not Found: Resource not found (e.g., model not found)500 Internal Server Error: Server error (e.g., Ollama connection issues, processing errors)
Model Not Found:
{
"detail": "The model you are trying to set as default was not found installed. Maybe pull it from ollama?"
}Ollama Not Running:
{
"detail": "Ollama either not installed or not running."
}Processing Error:
{
"detail": "Error processing request: <error details>"
}Several endpoints return streaming responses using Server-Sent Events (SSE) format:
- Model Download (
POST /api/models/download/{model_name}) - Chat Messages (
POST /api/chatwithstream: true)
For chat streams, the connection may pause mid-stream when a tool requires user confirmation (see the chat endpoint above); the client answers via POST /api/chat/confirm-tool and streaming resumes.
Each event follows this format:
data: <JSON_OBJECT>
Events are separated by double newlines (\n\n).
Streaming responses end with:
data: [DONE]
const response = await fetch('http://127.0.0.1:8000/api/chat', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ message: 'Hello', stream: true })
});
const reader = response.body.getReader();
const decoder = new TextDecoder();
let buffer = '';
while (true) {
const { value, done } = await reader.read();
if (done) break;
buffer += decoder.decode(value, { stream: true });
const lines = buffer.split('\n\n');
for (let i = 0; i < lines.length - 1; i++) {
const line = lines[i].trim();
if (line.startsWith('data: ')) {
const data = line.replace('data: ', '');
if (data === '[DONE]') {
// Stream ended
break;
}
const json = JSON.parse(data);
// Process json chunk
}
}
buffer = lines[lines.length - 1];
}Currently, there are no rate limits imposed on the API. However, be mindful of:
- Model inference can be resource-intensive
- Internet search operations consume API credits
- Database operations may be affected by high load
The API does not currently implement CORS restrictions. For production deployments, consider adding appropriate CORS headers.
When the backend is running, interactive API documentation is available at:
- Swagger UI:
http://127.0.0.1:8000/docs - ReDoc:
http://127.0.0.1:8000/redoc
These interfaces allow you to:
- Browse all available endpoints
- Test endpoints directly from the browser
- View request/response schemas
- See example requests and responses
- Installation Guide - Setup instructions
- README - Project overview