You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
The offline fallback from #104 works exactly as designed — a cloud call that fails on connectivity is
retried on the downloaded model instead of erroring. What it does not have is urgency, and I think that
is worth changing, because the budget it waits out was set for a situation that no longer applies once a
model is installed.
What the wait is made of
In OpenAiCompatibleClient, a transcription that cannot reach its provider spends:
executeForBody(maxRetries = 3)
4 attempts
RETRY_DELAY_MS = 3000
9 s of pauses
NETWORK_CONNECT_TIMEOUT_SECONDS = 8
up to 8 s per attempt
Only after all of that does DictateController.localFallbackProvider get a chance to look at the error.
Provider unreachable (SYN dropped — a VPN that is down, a self-hosted server that is off, a firewall): ≈ 41 s of spinner.
Airplane mode / no route: DNS fails fast, but the retry pauses still cost ≈ 9 s.
Provider accepts the connection and then goes silent: up to timeoutSeconds × 4.
Why that budget is the wrong one here
Those numbers are right when the cloud call is the only chance the dictation has — waiting through a
blip beats losing what someone just said. But with localFallbackEnabled on and a model downloaded,
failing costs a hand-off, not a transcript. Spending 41 seconds establishing that a provider is absent,
when the answer is already on the phone, buys nothing.
The current behaviour also makes the feature hard to trust. Someone who turns the fallback on because
they dictate on a train wants it to feel like the app noticed; instead it feels like the app hung and
then thought better of it.
Proposal
Make the reaching-out budget per call. A small NetworkBudget(maxRetries, retryDelayMs, connectTimeoutSeconds) on ProviderConfig, whose default reproduces today's constants exactly, so
nothing changes for any existing caller.
Pass a fast-fail budget when the fallback is armed — the fallback switched on, the active
provider a cloud one, a local model actually installed. Those are the same three conditions localFallbackProvider already applies after the fact; this only asks them before the call, because
a budget has to be chosen before there is an error to inspect.
Skip the call entirely when the OS already knows. If ConnectivityManager reports no validated
network, throw the connectivity error straight away and let the existing fallback path handle it.
Airplane mode then hands over instantly instead of confirming over a socket what the platform could
have said for free.
What this deliberately does not shorten
ProviderConfig.timeoutSeconds. Once the connection is up and bytes are moving, the full budget still
applies. A provider that accepted the upload and is working on a long recording is not one to walk away
from — cutting that short would trade a good cloud transcript for a worse local one, which is the
opposite of the point. The change is strictly about the phase where nothing is happening yet.
So the remaining slow case is honest: a server that accepts TCP and then says nothing still costs the
full timeout, because from the outside that is indistinguishable from a model thinking hard.
Cost
26 added lines across ProviderConfig, OpenAiCompatibleClient, DictateController and the manifest,
plus two new files (NetworkBudget, FastFallback). No behaviour change at all when the fallback is off or no model is downloaded — which is
the default. One new manifest entry, ACCESS_NETWORK_STATE (normal permission, no runtime prompt), for
step 3; step 3 can be dropped if that is not wanted, and steps 1–2 still cut the 41 s to about 3.
I have this building against v6.1.2 and am happy to open it as a PR — or to adjust the shape
first if you would rather it were done differently, e.g. as a user-facing preference rather than
automatic, or with the budget living somewhere other than ProviderConfig.
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
The offline fallback from #104 works exactly as designed — a cloud call that fails on connectivity is
retried on the downloaded model instead of erroring. What it does not have is urgency, and I think that
is worth changing, because the budget it waits out was set for a situation that no longer applies once a
model is installed.
What the wait is made of
In
OpenAiCompatibleClient, a transcription that cannot reach its provider spends:executeForBody(maxRetries = 3)RETRY_DELAY_MS = 3000NETWORK_CONNECT_TIMEOUT_SECONDS = 8Only after all of that does
DictateController.localFallbackProviderget a chance to look at the error.≈ 41 s of spinner.
timeoutSeconds× 4.Why that budget is the wrong one here
Those numbers are right when the cloud call is the only chance the dictation has — waiting through a
blip beats losing what someone just said. But with
localFallbackEnabledon and a model downloaded,failing costs a hand-off, not a transcript. Spending 41 seconds establishing that a provider is absent,
when the answer is already on the phone, buys nothing.
The current behaviour also makes the feature hard to trust. Someone who turns the fallback on because
they dictate on a train wants it to feel like the app noticed; instead it feels like the app hung and
then thought better of it.
Proposal
NetworkBudget(maxRetries, retryDelayMs, connectTimeoutSeconds)onProviderConfig, whose default reproduces today's constants exactly, sonothing changes for any existing caller.
provider a cloud one, a local model actually installed. Those are the same three conditions
localFallbackProvideralready applies after the fact; this only asks them before the call, becausea budget has to be chosen before there is an error to inspect.
ConnectivityManagerreports no validatednetwork, throw the connectivity error straight away and let the existing fallback path handle it.
Airplane mode then hands over instantly instead of confirming over a socket what the platform could
have said for free.
What this deliberately does not shorten
ProviderConfig.timeoutSeconds. Once the connection is up and bytes are moving, the full budget stillapplies. A provider that accepted the upload and is working on a long recording is not one to walk away
from — cutting that short would trade a good cloud transcript for a worse local one, which is the
opposite of the point. The change is strictly about the phase where nothing is happening yet.
So the remaining slow case is honest: a server that accepts TCP and then says nothing still costs the
full timeout, because from the outside that is indistinguishable from a model thinking hard.
Cost
26 added lines across
ProviderConfig,OpenAiCompatibleClient,DictateControllerand the manifest,plus two new files (
NetworkBudget,FastFallback). No behaviour change at all when the fallback is off or no model is downloaded — which isthe default. One new manifest entry,
ACCESS_NETWORK_STATE(normal permission, no runtime prompt), forstep 3; step 3 can be dropped if that is not wanted, and steps 1–2 still cut the 41 s to about 3.
I have this building against v6.1.2 and am happy to open it as a PR — or to adjust the shape
first if you would rather it were done differently, e.g. as a user-facing preference rather than
automatic, or with the budget living somewhere other than
ProviderConfig.Branch (against v6.1.2, builds clean): https://github.com/Hyroniem/Dictate/tree/fast-fallback
All reactions