Browse documentation
Statuses and errors
Reference for turn states, run interpretation, and common API error classes.
View MarkdownTurn status
| Status | Meaning | output | error |
|---|---|---|---|
queued | Accepted and waiting | Empty | null |
in_progress | Active execution | Empty | null |
cancelling | Cancellation accepted; the runtime is cleaning up | Empty | null |
completed | Work and required durability completed | The complete result | null |
failed | Execution or durability failed | Complete output produced before the failure, if any | Why it failed |
cancelled | Intentionally cancelled | Empty | null |
incomplete | Stopped after producing part of a result | What was produced, possibly nothing | Why it stopped, when the runtime reported a reason; otherwise null |
completed, failed, cancelled and incomplete are final. A finished turn is
never relabelled: cancelling a completed, failed or incomplete turn returns
it as it settled, and only work that was still queued or active becomes
cancelled.
An incomplete turn can also stop before it produced any output, for example when
its execution was interrupted or reached the 12-hour limit. Its error is then
one of worker_execution_interrupted, worker_action_interrupted,
turn_context_protocol_incomplete or turn_duration_limit, each with a fixed
message: see the table below.
Turn errors
error on a turn is null unless the turn is failed or incomplete. It is
stored with the turn, so it stays available when you read the turn later,
independently of how long events are retained.
{
"code": "provider_transport_failed",
"source": "provider",
"message": "The connection to the model provider was interrupted.",
"retryable": true,
"reference": "turn_01J..."
}| Field | Meaning |
|---|---|
code | Stable machine-readable reason. New codes can be added |
source | provider, extension, runtime, platform or configuration |
message | Explanation for people |
retryable | Whether sending the same work again can succeed |
reference | Identifier to give support |
A failed and an incomplete turn are described by the same codes. The status
says whether the output is partial.
Known failure reasons
The list is not exhaustive. Decide how to react from retryable and message
when you meet a code you do not know.
| Code | Retryable | Meaning |
|---|---|---|
provider_authentication_failed | No | The model provider rejected the configured credentials |
provider_quota_exhausted | No | The provider account has no available quota or credit |
provider_rate_limited | Yes | The provider is temporarily rate limiting requests |
provider_transport_failed | Yes | The connection to the provider was interrupted |
provider_service_unavailable | Yes | The provider is temporarily unavailable |
provider_request_rejected | No | The provider rejected the request |
provider_failure | Yes | The provider could not complete the request |
agent_execution_failed | No | The agent failed; message carries the sanitized source error |
agent_configuration_invalid | No | The agent configuration is incomplete or invalid |
extension_execution_failed | No | An agent extension failed |
tool_execution_failed | No | An agent tool failed |
command_failed | No | A tool command exited unsuccessfully |
command_timed_out | No | A tool command exceeded its time limit |
integration_credentials_rejected | No | An integration reported a credential failure |
integration_api_rejected | No | An integration reported an API rejection |
network_policy_denied | No | The agent reached for a destination its network policy denies |
workspace_checkpoint_failed | Yes | The workspace could not be saved durably |
event_stream_write_failed | Yes | The result was kept, but its event stream could not be saved fully |
turn_context_protocol_incomplete | No | Ended without a complete checkpoint; an action may have completed |
worker_execution_interrupted | Yes | The runtime stopped before the Turn finished. Try again in a few minutes. |
worker_action_interrupted | No | Stopped during an action that may have completed. Check its result first |
worker_unavailable | Yes | The runtime worker could not complete the run |
runtime_timeout | Yes | The run exceeded its allowed execution time |
turn_duration_limit | No | The turn reached its 12-hour limit. Send a new prompt to continue |
sandbox_operation_failed | Yes | The sandbox could not be started or resumed |
sandbox_capacity_exceeded | Yes | Sandbox capacity is temporarily unavailable |
sandbox_not_found | No | The run's durable sandbox is no longer available |
sandbox_snapshot_activation_failed | Yes | The agent's environment was idle and could not be reactivated |
sandbox_snapshot_unavailable | No | The agent's environment is no longer available: redeploy the agent |
agent_version_unavailable | No | The agent version a message was sent to is gone. Nothing ran |
runtime_dispatch_failed | Yes | The message could not be started. Send it again |
concurrency_limit_reached | Yes | The workspace reached its active execution limit |
platform_failure | Yes | An unexpected platform failure |
Error shape
{
"error": {
"message": "Missing scope run:write",
"type": "permission_error",
"param": null,
"code": "permission_denied"
}
}HTTP classes
| Status | Typical meaning |
|---|---|
400 | Invalid request or model |
401 | Missing, invalid or revoked API key |
403 | Missing scope or workspace permission |
404 | Resource not found in the authorized workspace |
408 | An upload ran for longer than an upload may |
409 | Runtime readiness or lifecycle conflict |
413 | Payload exceeds a supported limit |
429 | Rate or quota limit |
500 | Internal execution failure |
503 | Request admission, the runtime or a platform service is temporarily unavailable. Retry it |
Temporary failures
When the database or the event store cannot answer a request for a moment, for
example because a connection was reset or a call timed out, the API answers 503
with code runtime_unavailable. The failure is expected to pass, so the same
request is expected to succeed shortly. The answer carries a Retry-After header
with the number of seconds to wait, a few (1 to 3) and not the same for every
client, so that requests refused together do not come back together. The run and
file endpoints also say retryable: true:
{
"error": {
"code": "runtime_unavailable",
"message": "The service is temporarily unavailable. Try again shortly.",
"param": null,
"retryable": true,
"reference": "0f6d1c4e-4a3b-4c52-9d7e-2f1b8a6c9e10"
}
}Retry it with backoff, waiting at least Retry-After. The failure can come after
the work was done, so how you repeat a request depends on the endpoint:
POST /runs,POST /runs/{runId}/cancel,DELETE /runs/{runId}andPOST /filestake anIdempotency-Key. Repeat them with the same key, and the API returns the first result instead of doing the work twice. ForPOST /filesthat means:- A repeat after an upload that failed or was interrupted starts the upload again under the same key.
- A repeat after the upload finished returns the file, and never changes it.
- A repeat that arrives while the first upload is still running answers
503with codeupload_in_progressand aRetry-After: wait, and repeat with the same key. If the first upload was interrupted without a word, this lasts at most about 30 seconds, and the next repeat takes the upload over. - If a repeat answers
409with codeconflict, the key was used for a different file, or its upload expired: upload the file again with a new key. The SDK does not retry a409.
- The other endpoints do not take a key. Repeat a read freely, and check the
resource before you repeat any other write. The agent, deployment, environment
value and API key endpoints do not yet answer every temporary database failure
with this
503: some of them still answer500.
The TypeScript SDK does this for you. It retries 429, 502, 503 and 504
twice by default, waits Retry-After, and reuses its idempotency key. A stream
that is cut is reconnected from the last event the SDK received, and if the API
is still unavailable when it reconnects, the same 503 applies.
The message never carries the words of the database or of a provider, and the
code is never a database code. A failure that is not temporary is a 500. Its
code is run_failed_internal when the platform reported the failure itself, or
internal_error on the run and file endpoints when it could not tell what failed.
Both mean the same: the message is fixed, reference identifies the failure for
support, and you should not retry it.
Rate limits
Requests are limited per workspace, counted across all of its API keys, so creating more keys does not add capacity. Operations that do expensive work have lower limits, and some also cap how many run at once: file uploads, deployment source uploads, creating deployments, revealing secrets, and creating or rotating API keys. Revoking an API key and cancelling a run are always kept available.
Every limited response carries the limit that applies to the request:
| Header | Meaning |
|---|---|
X-RateLimit-Limit | The limit that governs the request |
X-RateLimit-Remaining | Requests left in that limit's window |
X-RateLimit-Reset | Unix time in seconds when the limit resets |
Retry-After | On a 429, and on a 503 for a database or event-store failure: seconds to wait before retrying |
A 429 has code rate_limited and is safe to retry after the delay. A 503
with code runtime_unavailable means the platform could not answer right now,
whether it was checking limits or reading its database. Retry it with backoff,
and wait Retry-After when the response has one. See
Temporary failures.
Uploads that are stopped
A file upload or a deployment source upload holds one of its workspace's upload places for as long as it runs, and it is stopped when it can no longer be sure it holds it, so the limit on simultaneous uploads holds.
| Response | When | What to do |
|---|---|---|
503 with code runtime_unavailable | Request admission was unavailable for longer than about a minute while the upload was running | Retry the upload with backoff. A file upload keeps its Idempotency-Key. A deployment whose source upload was interrupted is failed: run salambo deploy again, which creates a new deployment |
408 with code request_timeout | The upload ran for longer than an upload may: 20 minutes for a file, 15 minutes of transfer for a source archive | Start the upload again with a new request. Split or compress the input if it does not finish in time |
An interruption shorter than about a minute does not stop an upload.
Stable code examples
| Code | Meaning |
|---|---|
unauthorized | Authentication failed |
permission_denied | Scope or permission missing |
not_found | Authorized resource not found |
request_timeout | An upload ran for longer than it may |
runtime_unavailable | Request admission, the runtime or a platform service temporarily unavailable. Retry |
upload_in_progress | Another upload with the same Idempotency-Key is still running. Retry with the same key after Retry-After |
internal_error | Unexpected failure that the run and file endpoints could not classify. Not retryable; give reference to support |
conflict | Request conflicts with runtime state |
deployment_expired | Deployment expired; create a new one |
rate_limited | Request rate exceeded |
quota_exceeded | Account quota exceeded |
upstream_unavailable | Provider or dependency unavailable |
run_failed_internal | Run failed inside the platform boundary. Not retryable; give reference to support |
stream_cursor_expired | Streaming cursor is no longer valid |
Interpreting stopped runs
Do not equate a cancelled run with a failed turn.
When all turns completed and the run stopped afterward during cleanup, Details reports:
turns completed, then the run stopped