Salambo
Browse documentation
TroubleshootingStatuses and errors

Statuses and errors

Reference for turn states, run interpretation, and common API error classes.

View Markdown

Turn status

StatusMeaningoutputerror
queuedAccepted and waitingEmptynull
in_progressActive executionEmptynull
cancellingCancellation accepted; the runtime is cleaning upEmptynull
completedWork and required durability completedThe complete resultnull
failedExecution or durability failedComplete output produced before the failure, if anyWhy it failed
cancelledIntentionally cancelledEmptynull
incompleteStopped after producing part of a resultWhat was produced, possibly nothingWhy it stopped, when the runtime reported a reason; otherwise null

completed, failed, cancelled and incomplete are final. A finished turn is never relabelled: cancelling a completed, failed or incomplete turn returns it as it settled, and only work that was still queued or active becomes cancelled.

An incomplete turn can also stop before it produced any output, for example when its execution was interrupted or reached the 12-hour limit. Its error is then one of worker_execution_interrupted, worker_action_interrupted, turn_context_protocol_incomplete or turn_duration_limit, each with a fixed message: see the table below.

Turn errors

error on a turn is null unless the turn is failed or incomplete. It is stored with the turn, so it stays available when you read the turn later, independently of how long events are retained.

json
{
  "code": "provider_transport_failed",
  "source": "provider",
  "message": "The connection to the model provider was interrupted.",
  "retryable": true,
  "reference": "turn_01J..."
}
FieldMeaning
codeStable machine-readable reason. New codes can be added
sourceprovider, extension, runtime, platform or configuration
messageExplanation for people
retryableWhether sending the same work again can succeed
referenceIdentifier to give support

A failed and an incomplete turn are described by the same codes. The status says whether the output is partial.

Known failure reasons

The list is not exhaustive. Decide how to react from retryable and message when you meet a code you do not know.

CodeRetryableMeaning
provider_authentication_failedNoThe model provider rejected the configured credentials
provider_quota_exhaustedNoThe provider account has no available quota or credit
provider_rate_limitedYesThe provider is temporarily rate limiting requests
provider_transport_failedYesThe connection to the provider was interrupted
provider_service_unavailableYesThe provider is temporarily unavailable
provider_request_rejectedNoThe provider rejected the request
provider_failureYesThe provider could not complete the request
agent_execution_failedNoThe agent failed; message carries the sanitized source error
agent_configuration_invalidNoThe agent configuration is incomplete or invalid
extension_execution_failedNoAn agent extension failed
tool_execution_failedNoAn agent tool failed
command_failedNoA tool command exited unsuccessfully
command_timed_outNoA tool command exceeded its time limit
integration_credentials_rejectedNoAn integration reported a credential failure
integration_api_rejectedNoAn integration reported an API rejection
network_policy_deniedNoThe agent reached for a destination its network policy denies
workspace_checkpoint_failedYesThe workspace could not be saved durably
event_stream_write_failedYesThe result was kept, but its event stream could not be saved fully
turn_context_protocol_incompleteNoEnded without a complete checkpoint; an action may have completed
worker_execution_interruptedYesThe runtime stopped before the Turn finished. Try again in a few minutes.
worker_action_interruptedNoStopped during an action that may have completed. Check its result first
worker_unavailableYesThe runtime worker could not complete the run
runtime_timeoutYesThe run exceeded its allowed execution time
turn_duration_limitNoThe turn reached its 12-hour limit. Send a new prompt to continue
sandbox_operation_failedYesThe sandbox could not be started or resumed
sandbox_capacity_exceededYesSandbox capacity is temporarily unavailable
sandbox_not_foundNoThe run's durable sandbox is no longer available
sandbox_snapshot_activation_failedYesThe agent's environment was idle and could not be reactivated
sandbox_snapshot_unavailableNoThe agent's environment is no longer available: redeploy the agent
agent_version_unavailableNoThe agent version a message was sent to is gone. Nothing ran
runtime_dispatch_failedYesThe message could not be started. Send it again
concurrency_limit_reachedYesThe workspace reached its active execution limit
platform_failureYesAn unexpected platform failure

Error shape

json
{
  "error": {
    "message": "Missing scope run:write",
    "type": "permission_error",
    "param": null,
    "code": "permission_denied"
  }
}

HTTP classes

StatusTypical meaning
400Invalid request or model
401Missing, invalid or revoked API key
403Missing scope or workspace permission
404Resource not found in the authorized workspace
408An upload ran for longer than an upload may
409Runtime readiness or lifecycle conflict
413Payload exceeds a supported limit
429Rate or quota limit
500Internal execution failure
503Request admission, the runtime or a platform service is temporarily unavailable. Retry it

Temporary failures

When the database or the event store cannot answer a request for a moment, for example because a connection was reset or a call timed out, the API answers 503 with code runtime_unavailable. The failure is expected to pass, so the same request is expected to succeed shortly. The answer carries a Retry-After header with the number of seconds to wait, a few (1 to 3) and not the same for every client, so that requests refused together do not come back together. The run and file endpoints also say retryable: true:

json
{
  "error": {
    "code": "runtime_unavailable",
    "message": "The service is temporarily unavailable. Try again shortly.",
    "param": null,
    "retryable": true,
    "reference": "0f6d1c4e-4a3b-4c52-9d7e-2f1b8a6c9e10"
  }
}

Retry it with backoff, waiting at least Retry-After. The failure can come after the work was done, so how you repeat a request depends on the endpoint:

  • POST /runs, POST /runs/{runId}/cancel, DELETE /runs/{runId} and POST /files take an Idempotency-Key. Repeat them with the same key, and the API returns the first result instead of doing the work twice. For POST /files that means:
    • A repeat after an upload that failed or was interrupted starts the upload again under the same key.
    • A repeat after the upload finished returns the file, and never changes it.
    • A repeat that arrives while the first upload is still running answers 503 with code upload_in_progress and a Retry-After: wait, and repeat with the same key. If the first upload was interrupted without a word, this lasts at most about 30 seconds, and the next repeat takes the upload over.
    • If a repeat answers 409 with code conflict, the key was used for a different file, or its upload expired: upload the file again with a new key. The SDK does not retry a 409.
  • The other endpoints do not take a key. Repeat a read freely, and check the resource before you repeat any other write. The agent, deployment, environment value and API key endpoints do not yet answer every temporary database failure with this 503: some of them still answer 500.

The TypeScript SDK does this for you. It retries 429, 502, 503 and 504 twice by default, waits Retry-After, and reuses its idempotency key. A stream that is cut is reconnected from the last event the SDK received, and if the API is still unavailable when it reconnects, the same 503 applies.

The message never carries the words of the database or of a provider, and the code is never a database code. A failure that is not temporary is a 500. Its code is run_failed_internal when the platform reported the failure itself, or internal_error on the run and file endpoints when it could not tell what failed. Both mean the same: the message is fixed, reference identifies the failure for support, and you should not retry it.

Rate limits

Requests are limited per workspace, counted across all of its API keys, so creating more keys does not add capacity. Operations that do expensive work have lower limits, and some also cap how many run at once: file uploads, deployment source uploads, creating deployments, revealing secrets, and creating or rotating API keys. Revoking an API key and cancelling a run are always kept available.

Every limited response carries the limit that applies to the request:

HeaderMeaning
X-RateLimit-LimitThe limit that governs the request
X-RateLimit-RemainingRequests left in that limit's window
X-RateLimit-ResetUnix time in seconds when the limit resets
Retry-AfterOn a 429, and on a 503 for a database or event-store failure: seconds to wait before retrying

A 429 has code rate_limited and is safe to retry after the delay. A 503 with code runtime_unavailable means the platform could not answer right now, whether it was checking limits or reading its database. Retry it with backoff, and wait Retry-After when the response has one. See Temporary failures.

Uploads that are stopped

A file upload or a deployment source upload holds one of its workspace's upload places for as long as it runs, and it is stopped when it can no longer be sure it holds it, so the limit on simultaneous uploads holds.

ResponseWhenWhat to do
503 with code runtime_unavailableRequest admission was unavailable for longer than about a minute while the upload was runningRetry the upload with backoff. A file upload keeps its Idempotency-Key. A deployment whose source upload was interrupted is failed: run salambo deploy again, which creates a new deployment
408 with code request_timeoutThe upload ran for longer than an upload may: 20 minutes for a file, 15 minutes of transfer for a source archiveStart the upload again with a new request. Split or compress the input if it does not finish in time

An interruption shorter than about a minute does not stop an upload.

Stable code examples

CodeMeaning
unauthorizedAuthentication failed
permission_deniedScope or permission missing
not_foundAuthorized resource not found
request_timeoutAn upload ran for longer than it may
runtime_unavailableRequest admission, the runtime or a platform service temporarily unavailable. Retry
upload_in_progressAnother upload with the same Idempotency-Key is still running. Retry with the same key after Retry-After
internal_errorUnexpected failure that the run and file endpoints could not classify. Not retryable; give reference to support
conflictRequest conflicts with runtime state
deployment_expiredDeployment expired; create a new one
rate_limitedRequest rate exceeded
quota_exceededAccount quota exceeded
upstream_unavailableProvider or dependency unavailable
run_failed_internalRun failed inside the platform boundary. Not retryable; give reference to support
stream_cursor_expiredStreaming cursor is no longer valid

Interpreting stopped runs

Do not equate a cancelled run with a failed turn.

When all turns completed and the run stopped afterward during cleanup, Details reports:

text
turns completed, then the run stopped