Response format
1. Response body
{
"id": "chatcmpl-7f3a1c2f",
"object": "chat.completion",
"created": 1759750801,
"model": "tildeopen-30b-64k",
"choices": [
{
"index": 0,
"message": {"role": "assistant", "content": "Man patīk kūkas."},
"finish_reason": "stop"
}
],
"usage": {"prompt_tokens": 46, "completion_tokens": 6, "total_tokens": 52}
}
2. Fields
| Field | Meaning |
|---|---|
| choices[0].message.content | The generated text. This is the only task output; everything else is metadata. |
| choices[0].finish_reason | "stop": the model finished normally. "length": max_tokens was reached and the output is truncated; resend with a larger max_tokens. |
| usage.prompt_tokens | Tokens in the request (all messages). |
| usage.completion_tokens | Tokens generated. |
| usage.total_tokens | Sum of the two; relevant for the 65,536-token context limit. |
| id, object, created, model | Request identifier, object type, Unix timestamp, model name. |
choices always contains exactly one element; the model does not return multiple candidates or log probabilities.
3. Content of message.content by task
| Task | What is returned |
|---|---|
| Translation | The translation only. No preamble, explanation or quotation marks. Line and paragraph structure of the input is preserved. |
| Translation with glossary | As above; glossary terms appear in the supplied target forms, inflected as the target language requires. |
| Translation with examples | As above; the output follows the style and terminology of the provided examples. |
| Contextual question answering | A short answer in natural language, in the language of the question, based on the supplied context. If the context does not contain the information, the model returns an informative refusal: a brief description of what the document is about, followed by a statement that the context is insufficient to answer the question. |
| Summarisation | The summary text in the language of the context unless the instruction specifies otherwise. Length and focus follow the instruction. |
Output is plain UTF-8 text. The model does not emit JSON, markup or special tokens inside content unless the input text contained them.
4. Streaming
"stream": true the real-time endpoint returns server-sent events, each a chat.completion.chunk object with the incremental text in choices[0].delta.content; the final chunk carries finish_reason and is followed by data: [DONE]. Use InvokeEndpointWithResponseStream.
5. Batch Transform output
Default: one chat-completion object per line, same schema as above, in the order of the input file. The output file is named <input file name>.out in the configured S3 output path.
With DataProcessing.JoinSource = "Input": each output line contains the original request object with the response added under the key SageMakerOutput, so inputs and outputs are matched on the same line without a separate join step.
{"messages":[...],"max_tokens":256,"temperature":0.0,"SageMakerOutput":{"id":"chatcmpl-b001","object":"chat.completion","created":1759750900,"model":"tildeopen-30b-64k","choices":[{"index":0,"message":{"role":"assistant","content":"Man patīk pīrāgi."},"finish_reason":"stop"}],"usage":{"prompt_tokens":24,"completion_tokens":6,"total_tokens":30}}}
6. Errors
A request whose messages do not fit in the context window, or that is malformed JSON, returns an HTTP 4xx error from the container with a JSON body describing the problem; no choices are returned.