Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
17 commits
Select commit Hold shift + click to select a range
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
22 changes: 22 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,27 @@
# Changelog

## [0.15.0] — 2026-08-31

A turn cut short is no longer thrown away. Whatever it produced is persisted and
marked incomplete, and it is no longer logged as a clean success.

### Breaking Changes
- `ai_interaction_logs.status` has a new value, `aborted` (`AiInteractionStatus::Aborted`), used for a turn whose caller hung up mid-stream. It was previously recorded as `success`. Code matching exhaustively on the enum must handle the new case, and dashboards that count everything other than `success` as a failure will now count cancellations among them. The tokens an aborted turn burned still count towards conversation usage totals.
- The `message_delta` event's `delta.stop_reason` has a new value, `incomplete`, for a turn that never finished — the caller hung up, or the max-stream-duration guard cut the generation off. Such turns previously reported `end_turn`, which was indistinguishable from a clean finish. The same value is stored on `ai_llm_messages.response_data.stop_reason`.
- An interrupted turn now leaves an assistant message in the transcript even when it produced no text, so `ChatBotPresenter::transcript()` can return assistant rows whose `content` is empty. Each row carries a new `incomplete` boolean; render a flagged row as an interrupted reply rather than as a blank answer.
- Two new migrations (`ai_turn_runs`, `ai_turn_events`) ship with this release. Re-publish migrations and run them, even if you do not use the dispatched-turn API — `ai:prune-turn-events` is scheduled by default and expects the tables.

### New Features
- An interrupted assistant message records why in its `metadata`: `incomplete` plus an `incomplete_reason` of `client_aborted` or `max_stream_duration`.
- Turns now emit a heartbeat while the provider is silent (`conversations.heartbeat_seconds`, default 5, `0` to disable). It is encoded as an SSE comment, so existing clients ignore it; it keeps intermediaries from timing out mid-answer and lets an abandoned turn be noticed in seconds rather than minutes. Currently emitted for `openai-compatible` and `lm-studio` systems; a turn dispatched as a job heartbeats for every provider.
- Added `AiPersonaConversationService::dispatchTurn()`, `resumeTurn()` and `cancelTurn()`: a turn can run as a queued job that records its events, so a browser reload resumes it instead of killing it. Events are framed with an SSE `id:` carrying their sequence, and the published client reports it through a new `onSequence` callback. Requires the two new migrations and a queue worker.
- Added `ai:prune-turn-events` (scheduled daily at 03:15) to clear finished turn runs past `turns.retention_days`.

### Bug Fixes
- A turn interrupted before the model emitted any text is no longer discarded entirely. It previously left the user's message in the transcript with no reply beneath it and no record that a turn had ever run — the failure mode was most visible with a large context, where a model can spend minutes processing the prompt before its first token.
- Tool calls an interrupted turn already made are now persisted with it. They were dropped along with the rest of the turn, leaving the next turn's history with no record that a tool had run and changed state.
- The maximum-stream-duration guard now applies during provider silence. It was only evaluated when an event arrived, so a stalled stream could run well past `conversations.max_stream_seconds` before anything noticed.

## [0.14.1] — 2026-08-30

### New Features
Expand Down
112 changes: 110 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -149,6 +149,43 @@ Every laravel/ai HTTP request/response can be captured verbatim into the
Request bodies and response bytes are stored, but request **headers are never
recorded**, so provider API keys are not persisted.

### Detached Turns

A turn dispatched with `dispatchTurn()` runs as a queued job and writes its
events to the `ai_turn_events` table, so a browser reload resumes it instead of
killing it (see [Running a turn as a job](#running-a-turn-as-a-job)).

```php
'turns' => [
'queue' => env('CODE_TALKER_TURN_QUEUE'),
'abandon_after_seconds' => (int) env('CODE_TALKER_TURN_ABANDON_SECONDS', 30),
'poll_interval_ms' => (int) env('CODE_TALKER_TURN_POLL_MS', 250),
'max_stream_seconds' => (int) env('CODE_TALKER_TURN_MAX_STREAM_SECONDS', 900),
'retention_days' => (int) env('CODE_TALKER_TURN_RETENTION_DAYS', 7),
],
```

- `queue` — the queue `RunConversationTurnJob` is dispatched on; `null` uses
the default queue.
- `abandon_after_seconds` — a running turn stops when nobody has read its
events for this long. `connection_aborted()` reports 0 in a worker, so this
is what stops a turn nobody is waiting for.
- `poll_interval_ms` — how often a reader polls the store for new events.
- `max_stream_seconds` — ceiling for a single `resumeTurn()` read before it
ends with a `max_stream_duration` error; reconnecting starts a fresh window.
- `retention_days` — finished runs older than this are removed by
`php artisan ai:prune-turn-events`, scheduled daily at 03:15 (respects the
`schedule` flag).

Note that `turns.max_stream_seconds` bounds only the read side. Generation
inside the worker is governed by `conversations.max_stream_seconds` (default
300), which caps each individual provider request — the same guard the
synchronous path applies, enforced promptly during provider silence by the
heartbeat rather than only when the next provider event arrives. A host running
a large-context local model, where prompt processing alone can occupy minutes
of a single request, should raise `conversations.max_stream_seconds`
accordingly.

### Troubleshooting

**`Provider is unavailable: HTTP request returned status code 404`** — returned
Expand Down Expand Up @@ -345,8 +382,9 @@ Every event carries a `type`. These are typed in the published declarations.
| `message_start` | — |
| `content_block_delta` | `delta.text` |
| `reasoning_block_delta` | `delta.reasoning` |
| `message_delta` | `delta.stop_reason`, `usage` |
| `message_delta` | `delta.stop_reason` (`end_turn`/`max_tokens`/`tool_use`/`incomplete`), `usage` |
| `message_stop` | — |
| `heartbeat` | — (encoded as an SSE comment, not a data frame) |
| `tool_use_progress` | `text` (always `""`), `tools` (one tool name per event), plus `input`/`output`/`successful` when tool payloads are enabled |
| `page_reload` | — |
| `error` | `message`, `reason` (`max_stream_duration`/`provider_error`) |
Expand All @@ -357,6 +395,20 @@ aren't display text), so without this a turn calling a tool, especially one
retrying after an error, streams nothing but silence between text/reasoning
deltas.

`stop_reason` is `incomplete` when the turn never finished — the connection
dropped, or the server's duration guard cut the generation off. Whatever
content arrived stops mid-answer, and the turn is stored that way (see
[Interrupted turns](#interrupted-turns)).

`heartbeat` fires while the provider is silent. `SseFrameEncoder` renders it as
`: ping` — an SSE comment — so browsers and the published client ignore it
without any handling. It is there so something reaches the socket during a long
gap: intermediaries stop timing out mid-answer, and PHP only flips
`connection_aborted()` after a write to a dead connection, so without it an
abandoned turn keeps generating until the model's next event. Set
`conversations.heartbeat_seconds` to `0` to disable. Detection costs two beats:
the first write marks the socket dead, the second observes it.

`page_reload` fires when a tool's structured result carries `_page_reload:
true` — see [Tool Registration](#tool-registration) for how a tool sets it.
Deciding what "reload" means (call `location.reload()` immediately, wait for
Expand Down Expand Up @@ -386,6 +438,57 @@ $chat->usingCancellationCheck(fn (): bool => $job->isReleased())
->continueConversation($conversation, $message);
```

### Interrupted turns

A turn that stops before the model finishes — the browser hung up, or the
duration guard tripped — is still recorded, whatever it had produced:

- The assistant message is persisted even when it holds no text at all, so a
user's question is never left with nothing beneath it. Its `metadata` carries
`incomplete: true` and an `incomplete_reason` of `client_aborted` or
`max_stream_duration`, and `ChatBotPresenter::transcript()` surfaces the flag
as `incomplete` on the row. Render it as an interrupted reply rather than as
an answer.
- Tool calls the model made before the stop are persisted with it. A tool that
ran changed state on your side; dropping the turn would leave the next turn's
history with no record it ever happened.
- `AiInteractionLog::status` is `aborted` (not `success`) for a turn the caller
hung up on, with `provider_metadata.error_reason` set to `client_aborted`.
The tokens it burned still count towards the conversation's usage totals —
hanging up does not refund what the provider already generated.

### Running a turn as a job

`continueConversation()` ties the turn to the caller's connection: close the
tab and the turn stops, reload and it is gone. For turns long enough that this
matters, dispatch the turn instead and stream it from its store.

```php
// Start it. Returns an AiTurnRun; `public_id` is the handle to put in a URL.
$run = $chat->dispatchTurn($conversation, $request->string('message')->toString());

// Stream it — from the start, or from wherever the browser left off.
foreach ($encoder->encode($chat->resumeTurn($run, $after)) as $frame) {
echo $frame;
ob_get_level() > 0 && ob_flush();
flush();
}

// Stop it early.
$chat->cancelTurn($run);
```

Each event is framed with an SSE `id:` carrying its sequence. A browser that
reconnects passes the last sequence it saw back as `after`, and the turn
resumes rather than replaying. The published client reports it via
`onSequence`.

A dispatched turn needs a queue worker. Because `connection_aborted()` reports
0 in a worker, a run stops when nobody has read it for
`turns.abandon_after_seconds` (default 30) — so closing the tab still stops
generation, and a reload inside that window reattaches to the same run.
`ai:prune-turn-events` clears finished runs past `turns.retention_days`.

### Resolving conversations across requests

The package used to keep this in the session and a cookie. It no longer does —
Expand All @@ -399,6 +502,7 @@ and `$conversation->chat_hash` is a stable shareable handle that

```php
$presenter->transcript($conversation); // visible messages, system prompt excluded
// each row carries `incomplete` — see Interrupted turns
$presenter->totalCostUsd($persona); // lifetime cost for a persona
$presenter->conversationsFor($user, $personas); // an authenticated user's conversations
```
Expand Down Expand Up @@ -966,13 +1070,14 @@ route file, and defaults to `['web', 'auth', 'can:manage-ai-tools']`.

## Scheduled Jobs

The package registers four jobs automatically (requires Laravel's scheduler to be running):
The package registers five jobs automatically (requires Laravel's scheduler to be running):

| Job | Schedule | Description |
| -------------------------------- | -------------------------- | ------------------------------------------------------- |
| `ai:sync-conversation-usage` | Twice daily (00:00, 12:00) | Syncs token counts and cost to `AiConversation` |
| `BackfillConversationUsageJob` | Daily at 02:30 | Backfills usage for conversations missing cost data |
| `ai:prune-provider-exchanges` | Daily at 03:00 | Removes `ai_provider_exchanges` rows past retention |
| `ai:prune-turn-events` | Daily at 03:15 | Removes finished turn runs past `turns.retention_days` |
| `ai:complete-idle-conversations` | Every 15 minutes | Completes idle conversations, triggering memory extract |

Disable automatic scheduling in config and register manually if needed:
Expand Down Expand Up @@ -1000,6 +1105,9 @@ php artisan ai:sync-conversation-usage
# Delete ai_provider_exchanges rows older than raw_exchanges.retention_days
php artisan ai:prune-provider-exchanges

# Delete finished turn runs (and their events) older than turns.retention_days
php artisan ai:prune-turn-events

# Mark idle conversations Completed, triggering memory extraction
php artisan ai:complete-idle-conversations
php artisan ai:complete-idle-conversations --minutes=60 --dry-run
Expand Down
29 changes: 29 additions & 0 deletions config/code-talker.php
Original file line number Diff line number Diff line change
Expand Up @@ -103,6 +103,14 @@
'conversations' => [
'idle_timeout_minutes' => (int) env('CODE_TALKER_CONVERSATION_IDLE_MINUTES', 30),

// Seconds of provider silence before the turn emits a heartbeat.
// Two things depend on it: intermediaries stop timing out during a
// long gap, and PHP only flips connection_aborted() after a write to
// a dead socket — so without a heartbeat an abandoned turn keeps
// generating until the model's next event, which on a large context
// can be minutes. Set to 0 to disable.
'heartbeat_seconds' => (int) env('CODE_TALKER_HEARTBEAT_SECONDS', 5),

// Wall-clock ceiling (seconds) for a single streamed chat turn, across
// all tool steps and continuation attempts. Guards against a runaway
// generation — e.g. a reasoning model that loops until it overflows the
Expand Down Expand Up @@ -331,4 +339,25 @@
'retention_days' => (int) env('CODE_TALKER_RAW_EXCHANGES_RETENTION_DAYS', 14),
],

/*
|--------------------------------------------------------------------------
| Detached Turns
|--------------------------------------------------------------------------
|
| A turn dispatched with AiPersonaConversationService::dispatchTurn() runs
| as a queued job and writes its events to ai_turn_events, so a browser
| reload resumes the turn instead of killing it. connection_aborted() is
| meaningless in a worker, so "nobody has polled for abandon_after_seconds"
| is what stops a turn nobody is waiting for.
|
*/

'turns' => [
'queue' => env('CODE_TALKER_TURN_QUEUE'),
'abandon_after_seconds' => (int) env('CODE_TALKER_TURN_ABANDON_SECONDS', 30),
'poll_interval_ms' => (int) env('CODE_TALKER_TURN_POLL_MS', 250),
'max_stream_seconds' => (int) env('CODE_TALKER_TURN_MAX_STREAM_SECONDS', 900),
'retention_days' => (int) env('CODE_TALKER_TURN_RETENTION_DAYS', 7),
],

];
Original file line number Diff line number Diff line change
@@ -0,0 +1,33 @@
<?php

use Illuminate\Database\Migrations\Migration;
use Illuminate\Database\Schema\Blueprint;
use Illuminate\Support\Facades\Schema;

return new class extends Migration
{
public function up(): void
{
Schema::create('ai_turn_runs', function (Blueprint $table): void {
$table->id();
$table->string('public_id', 40)->unique();
$table->foreignId('ai_conversation_id')->index();
$table->string('status', 20)->index();
$table->text('prompt');
// The abandonment signal: connection_aborted() reports 0 in a
// worker, so "nobody is reading this" is the only usable stand-in
// for the browser having gone away.
$table->timestamp('last_polled_at')->nullable();
$table->timestamp('cancel_requested_at')->nullable();
$table->timestamp('started_at')->nullable();
$table->timestamp('finished_at')->nullable();
$table->text('error_message')->nullable();
$table->timestamps();
});
}

public function down(): void
{
Schema::dropIfExists('ai_turn_runs');
}
};
Original file line number Diff line number Diff line change
@@ -0,0 +1,28 @@
<?php

use Illuminate\Database\Migrations\Migration;
use Illuminate\Database\Schema\Blueprint;
use Illuminate\Support\Facades\Schema;

return new class extends Migration
{
public function up(): void
{
Schema::create('ai_turn_events', function (Blueprint $table): void {
$table->id();
$table->foreignId('ai_turn_run_id')->index();
$table->unsignedInteger('sequence');
$table->json('payload');
$table->timestamp('created_at')->nullable();

// The reader asks for everything after a sequence it already
// holds, so a duplicate would silently replay or skip output.
$table->unique(['ai_turn_run_id', 'sequence']);
});
}

public function down(): void
{
Schema::dropIfExists('ai_turn_events');
}
};
Loading
Loading