Skip to content

feat(files): serve both SDKs, seed the sandbox with uploads, collect its outputs, sweep expired files - #976

Closed
daavoo wants to merge 1 commit into
mainfrom
feat/files-api-parity
Closed

daavoo wants to merge 1 commit into
mainfrom
feat/files-api-parity

Conversation

@daavoo

@daavoo daavoo commented Sep 9, 2026 •

Copy link
Copy Markdown
Contributor

Description

Octonous talks to provider files APIs only through the official Anthropic and OpenAI SDKs pointed at a base URL, references uploads with image/document file sources, container_upload, input_file, and expects code execution to see attachments and hand back downloadable files. Otari's files API already stored and resolved files, but only in the OpenAI shape, never into the sandbox, and never out of it. This closes those gaps so an open-source model behind Otari can take the same requests a frontier model does.

  • Anthropic SDK compatibility. A request carrying anthropic-version (the SDK sends it on every call) gets FileMetadata (type, size_bytes, mime_type, downloadable, RFC 3339 created_at) and file_deleted on delete. OpenAI callers are unchanged.
  • Cursor pagination on the file listing: limit, after/after_id, order, has_more, first_id, last_id (part of Four list endpoints have no pagination bound #786).
  • Uploads reach the sandbox. With otari_code_execution in the request, every referenced upload is seeded into the session via the protocol's PutFile. container_upload blocks are staged and replaced by a marker for the model; without a sandbox they are read as documents. A refused seed fails the request rather than running over a missing input.
  • Sandbox outputs become files. File references in the result block are fetched with GetFile, stored as code_execution_output files owned by the same user and workspace, and named with their file_id in the tool result. Storage failures never fail the run.
  • Bare Responses input items (input_file, input_image outside a message) are now normalized.
  • Retention sweep reclaims bytes and rows of expired and deleted files (files_sweep_interval_sec, hourly, 0 disables). It is one entry in the lifespan worker registry, and commits through the Unit of Work like every other service.

Docs (docs/files.md, protocol doc), OpenAPI spec, and Postman collection are regenerated.

Not in this PR

  • Files in hybrid mode: hybrid has no database, so that is an architecture decision, not a gap.
  • Container-file ids (cfile_) and code-interpreter annotations on the Responses wire: needs the Anthropic-content-block lift the sandbox backend already notes.
  • The reference otari-sandbox-container does not populate the file-reference list yet, so output collection is wired but idle until it does.
  • The fsspec storage backend is feat(files): fsspec backend so any filesystem can hold uploaded files #977, stacked on this one.

How to test it locally

  1. Upload a file with the OpenAI SDK and with the Anthropic SDK, both pointed at the gateway; each gets its own response shape, and GET /api/v1/files?limit=2 pages with has_more and last_id.
  2. Send a chat request with otari_code_execution and a file_id reference against a sandbox; the file is on disk in the session, and any file the run writes comes back as a file_id in the tool result and is fetchable from the files API.
  3. Set files_retention_hours: 1, wait past expiry (or set files_sweep_interval_sec low), and watch the sweep log line reclaim the bytes and rows.

Covered by tests/integration/test_files_endpoint.py (Anthropic shapes, paging, sweep), tests/unit/test_sandbox_backend.py (seeding and output collection), tests/unit/test_content_normalizer.py (bare Responses items), and tests/unit/test_gateway_lifespan_shutdown.py (the sweep worker starts and can be switched off). make lint, make typecheck, the unit suite, the OSS smoke gate, and the files, messages-dispatch, hybrid-chat and settings integration tests pass locally.

PR Type

  • New Feature
  • Bug Fix
  • Refactor
  • Documentation
  • Infrastructure / CI

Relevant issues

Part of #786.

Checklist

  • I understand the code I am submitting.
  • I have added or updated tests that cover my change (tests/unit, tests/integration).
  • I ran the Definition of Done checks locally (make lint, make typecheck, make test).
  • Documentation was updated where necessary.
  • If the API contract changed, I regenerated the OpenAPI spec (uv run python scripts/generate_openapi.py).
  • If this changes a rule in ARCHITECTURE.md or scripts/check_architecture.py, the description names the rule and says why.

AI Usage

  • No AI was used.
  • AI was used for drafting/refactoring.
  • This is fully AI-generated.

AI Model/Tool used: Claude Code

Any additional AI details you'd like to share:

Rebased onto main after the models package split, the lifespan worker registry, the derived settings view, and the Unit of Work transaction rule landed; the change was adapted to each.

  • I am an AI Agent filling out this form (check box if true)

🤖 Generated with Claude Code

…sweep expired files

The /v1/files routes now answer in Anthropic's FileMetadata shape when the
caller sends anthropic-version (its SDK always does) and in the OpenAI shape
otherwise, and the listing is cursor-paged (limit, after/after_id, order,
has_more). Uploads referenced by a request that runs otari_code_execution are
seeded into the sandbox session with PutFile; container_upload blocks are
staged and replaced by a marker for the model. Files a run produces are
fetched with GetFile, stored as code_execution_output files owned by the same
user and workspace, and named with their file_id in the tool result. A bare
input_file or input_image item at the top level of a Responses input is now
normalized too. A background sweep reclaims the bytes and rows of expired
and deleted files.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@daavoo
daavoo temporarily deployed to integration-tests September 9, 2026 09:31 — with GitHub Actions Inactive
@coderabbitai

coderabbitai Bot commented Sep 9, 2026

Copy link
Copy Markdown

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

Copy link
Copy Markdown
Contributor

This PR has been automatically closed because the required template sections were not restored within 24 hours.

Please create a new PR using the template and complete the checklist. The template helps maintainers review your contribution.

@HareeshBahuleyan HareeshBahuleyan left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review by ChatGPT (AI assistant in pi), posted with user approval. Five findings from the PR/base-relative review are attached inline.

Comment on lines +299 to +308
response = await self._client.get(
f"{self._sandbox_url}/sessions/{self._session_id}/files",
params={"path": ref.filename},
)
response.raise_for_status()
except httpx.HTTPError as exc:
logger.warning("sandbox output %r could not be fetched: %s", ref.filename, exc)
continue
data = response.content
if not data or len(data) > self._files.max_output_bytes:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Bound sandbox downloads before buffering

httpx.get() downloads the entire output before checking max_output_bytes. Large sandbox files can exhaust gateway memory despite the configured limit. A 4-byte-limit probe consumed all 40 bytes.

Fix: Stream downloads and abort when the limit is exceeded.

Comment on lines +276 to +278
f"{self._sandbox_url}/sessions/{self._session_id}/files",
files={"file": (staged.filename, data, staged.mime_type)},
data={"path": staged.filename},

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Prevent same-name attachments from overwriting

Distinct file IDs named data.csv are uploaded to the same sandbox path. Reproduction left only the second file’s contents, silently corrupting multi-file analysis.

Fix: Allocate unique paths and expose those paths consistently to the model.

Comment on lines +286 to +290
cursor_id = after or after_id
if cursor_id is not None:
cursor = await fetch_file(db, cursor_id, user_id, workspace_id=scope)
if cursor is None:
raise HTTPException(status_code=status.HTTP_404_NOT_FOUND, detail="File not found")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Keep pagination and cursor expiry rules consistent

Listings include expired files, but cursor lookup rejects them. Reproduction returned an expired last_id with has_more=true; following that cursor returned 404, preventing access to subsequent live files.

Fix: Exclude expired files from listings and define consistent cursor-expiry behavior.

Comment on lines +430 to +432
if isinstance(item, dict) and item.get("type") == "input_text":
return {"role": "user", "content": [item]}
return item

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Wrap native Responses content parts in messages

The new bare-input handling wraps extracted text only. Native PDFs/images remain top-level input_file/input_image items, which fail OpenAI’s ResponseInputItemParam schema validation. Wrapping the same result in a user message validates.

Fix: Wrap all supported bare content parts, including native passthrough results.

Comment thread docs/public/openapi.json
"title": "Workspace Id"
}
},
{

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Regenerate the dashboard API client

The spec adds pagination parameters, but web/src/client/schema.ts remains unchanged. Regeneration produces a diff, failing the dashboard’s explicit client-drift gate.

Fix: Run pnpm --dir web client:generate.

@github-actions

Copy link
Copy Markdown
Contributor

This PR has been automatically closed because the required template sections were not restored within 24 hours.

Please create a new PR using the template and complete the checklist. The template helps maintainers review your contribution.

@github-actions github-actions Bot closed this Sep 15, 2026
@daavoo

daavoo commented Sep 18, 2026

Copy link
Copy Markdown
Contributor Author

Continued in #1364: GitHub refuses to reopen a PR after its branch is force-pushed, and the branch was rebased onto current main.

This branch was previously deployed

1 inactive deployment
integration-tests — 80da99de Deployed Sep 9, 2026 by daavoo via test-integration #1620
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants