Streaming and retries
Set stream: true and consume SSE data events. HTTP 200 means the stream has started, not that generation has succeeded.
Before consuming SSE, inspect the HTTP status and Content-Type. A request with stream: true can still return an HTTP 503 JSON error with code pii_mapping_saturated before the stream starts. Do not parse that response as SSE or retry automatically; contact an administrator or support with X-Request-ID. This error has no Retry-After header. See Error Codes.
Normal completion ends with data: [DONE]. If the connection closes without [DONE], treat the result as incomplete rather than a successful complete response.
Once SSE starts, the HTTP status cannot change. A failure sends at most one public error data frame, then closes without [DONE]. A disconnected or unwritable connection may deliver no error frame.
The standard production rollout drain window is 180 seconds. Active streams that outlast this window may close. This is not a generation deadline for every request or a guarantee that long streams survive deployments.
Decide whether to retry from the terminal state and use bounded backoff. An error or interruption does not prove the model did not execute; retrying may duplicate generation, charges or tool actions. Do not concatenate partial results from separate attempts.
For support, record the response X-Request-ID, time and error code. X-Request-ID is a client correlation identifier and is not guaranteed globally unique. Never include keys or sensitive request bodies.
data: {"error":{"message":"模型暂时不可用,请稍后重试","type":"server_error","code":"model_unavailable"}}
