Skip to main content
Machine Translation Contextual question answering Summarisation

Response format

1. Response body​

{

"id": "chatcmpl-7f3a1c2f",

"object": "chat.completion",

"created": 1759750801,

"model": "tildeopen-30b-64k",

"choices": [

{

"index": 0,

"message": {"role": "assistant", "content": "Man patīk kūkas."},

"finish_reason": "stop"

}

],

"usage": {"prompt_tokens": 46, "completion_tokens": 6, "total_tokens": 52}

}

2. Fields​

FieldMeaning
choices[0].message.contentThe generated text. This is the only task output; everything else is metadata.
choices[0].finish_reason"stop": the model finished normally. "length": max_tokens was reached and the output is truncated; resend with a larger max_tokens.
usage.prompt_tokensTokens in the request (all messages).
usage.completion_tokensTokens generated.
usage.total_tokensSum of the two; relevant for the 65,536-token context limit.
id, object, created, modelRequest identifier, object type, Unix timestamp, model name.

choices always contains exactly one element; the model does not return multiple candidates or log probabilities.

3. Content of message.content by task​

TaskWhat is returned
TranslationThe translation only. No preamble, explanation or quotation marks. Line and paragraph structure of the input is preserved.
Translation with glossaryAs above; glossary terms appear in the supplied target forms, inflected as the target language requires.
Translation with examplesAs above; the output follows the style and terminology of the provided examples.
Contextual question answeringA short answer in natural language, in the language of the question, based on the supplied context. If the context does not contain the information, the model returns an informative refusal: a brief description of what the document is about, followed by a statement that the context is insufficient to answer the question.
SummarisationThe summary text in the language of the context unless the instruction specifies otherwise. Length and focus follow the instruction.

Output is plain UTF-8 text. The model does not emit JSON, markup or special tokens inside content unless the input text contained them.

4. Streaming​

"stream": true the real-time endpoint returns server-sent events, each a chat.completion.chunk object with the incremental text in choices[0].delta.content; the final chunk carries finish_reason and is followed by data: [DONE]. Use InvokeEndpointWithResponseStream.

5. Batch Transform output​

Default: one chat-completion object per line, same schema as above, in the order of the input file. The output file is named <input file name>.out in the configured S3 output path.

With DataProcessing.JoinSource = "Input": each output line contains the original request object with the response added under the key SageMakerOutput, so inputs and outputs are matched on the same line without a separate join step.

{"messages":[...],"max_tokens":256,"temperature":0.0,"SageMakerOutput":{"id":"chatcmpl-b001","object":"chat.completion","created":1759750900,"model":"tildeopen-30b-64k","choices":[{"index":0,"message":{"role":"assistant","content":"Man patīk pīrāgi."},"finish_reason":"stop"}],"usage":{"prompt_tokens":24,"completion_tokens":6,"total_tokens":30}}}

6. Errors​

A request whose messages do not fit in the context window, or that is malformed JSON, returns an HTTP 4xx error from the container with a JSON body describing the problem; no choices are returned.