Skip to main content
Machine Translation Contextual question answering Summarisation

Limits and notes

  • Text only. Convert PDF, DOCX and other formats to plain text before sending.

  • The model is tuned for the six documented templates. Open-ended chat, reasoning, code or tool-use prompts are unsupported.

  • No safety alignment or content filtering is built in. Apply your own input and output moderation where required.

  • For contextual QA, the model is trained to refuse when the supplied context does not contain the answer; a refusal is a normal response, not an error.

  • Language names in the translation instruction must be in English (English, Latvian, German, …).

Appendix A. Real-time request samples​

InvokeEndpoint, Content-Type: application/json, Accept: application/json. One request per call. The messages array must follow one of the documented prompt templates exactly.

Translation​

{

"messages": [

{"role": "system", "content": "You are a translator."},

{"role": "user", "content": "I like pies.\n\nTranslate the given text from English to Latvian."}

],

"max_tokens": 256,

"temperature": 0.0

}

Translation with term glossary​

{

"messages": [

{"role": "system", "content": "You are a translator."},

{"role": "user", "content": "I like pies.\n\nTranslate the given text from English to Latvian, using this glossary {\"pie\": [\"kūka\"], \"like\": [\"patīk\"]}."}

],

"max_tokens": 256,

"temperature": 0.0

}

Translation with examples (translation memory)​

{

"messages": [

{"role": "system", "content": "You are a translator."},

{"role": "user", "content": "I like apples.\n\nTranslate the given text from English to Latvian."},

{"role": "assistant", "content": "Man patīk āboli."},

{"role": "user", "content": "I eat pies.\n\nTranslate the given text from English to Latvian."},

{"role": "assistant", "content": "Es ēdu kūkas."},

{"role": "user", "content": "I like pies.\n\nTranslate the given text from English to Latvian."}

],

"max_tokens": 256,

"temperature": 0.0

}

Contextual question answering​

{

"messages": [

{"role": "system", "content": "You are answering questions."},

{"role": "user", "content": "I like pies.\n\nDo I like pies?"}

],

"max_tokens": 128,

"temperature": 0.0

}

Summarisation​

{

"messages": [

{"role": "system", "content": "You are answering questions."},

{"role": "user", "content": "A large language model (LLM) is an AI model (typically a neural network) trained on a vast amount of text for natural language processing tasks, especially language generation. LLMs can typically generate, summarize, translate, and analyze text in many contexts. They are the basis for many modern chatbots.\nLLMs are typically based on transformer architecture. Generative pre-trained transformers (GPTs) are a type of LLM that is pre-trained to predict the next word. GPTs are then often fine-tuned to follow instructions and to behave as assistants.\nBiased or inaccurate training data can make an LLM's output less reliable.\n\nSummarise the given text. Focus on the main points."}

],

"max_tokens": 512,

"temperature": 0.0

}

Appendix B. Real-time response samples​

Content-Type: application/json. One chat-completion object per request. The generated text is in choices[0].message.content; token counts are in usage. finish_reason is "stop" on normal completion and "length" if max_tokens was reached.

Translation​

{

"id": "chatcmpl-7f3a1c2e",

"object": "chat.completion",

"created": 1759750800,

"model": "tildeopen-30b-64k",

"choices": [

{

"index": 0,

"message": {"role": "assistant", "content": "Man patīk pīrāgi."},

"finish_reason": "stop"

}

],

"usage": {"prompt_tokens": 24, "completion_tokens": 6, "total_tokens": 30}

}

Translation with term glossary​

{

"id": "chatcmpl-7f3a1c2f",

"object": "chat.completion",

"created": 1759750801,

"model": "tildeopen-30b-64k",

"choices": [

{

"index": 0,

"message": {"role": "assistant", "content": "Man patīk kūkas."},

"finish_reason": "stop"

}

],

"usage": {"prompt_tokens": 46, "completion_tokens": 6, "total_tokens": 52}

}

Contextual question answering​

{

"id": "chatcmpl-7f3a1c30",

"object": "chat.completion",

"created": 1759750802,

"model": "tildeopen-30b-64k",

"choices": [

{

"index": 0,

"message": {"role": "assistant", "content": "Yes, you like pies."},

"finish_reason": "stop"

}

],

"usage": {"prompt_tokens": 20, "completion_tokens": 6, "total_tokens": 26}

}

Summarisation​

{

"id": "chatcmpl-7f3a1c31",

"object": "chat.completion",

"created": 1759750803,

"model": "tildeopen-30b-64k",

"choices": [

{

"index": 0,

"message": {"role": "assistant", "content": "Large language models are neural networks trained on large text corpora for generating, summarising, translating and analysing text, and they underlie most modern chatbots. They are usually transformer-based; GPT models are pre-trained to predict the next word and then fine-tuned to follow instructions. Biased or inaccurate training data reduces the reliability of their output."},

"finish_reason": "stop"

}

],

"usage": {"prompt_tokens": 128, "completion_tokens": 64, "total_tokens": 192}

}

Appendix C. Batch transform output samples​

application/jsonlines. One JSON object per input line, in input order. Each object has the same schema as a real-time response.

Default (no join)​

File batch_input.jsonl.out:

{"id":"chatcmpl-b001","object":"chat.completion","created":1759750900,"model":"tildeopen-30b-64k","choices":[{"index":0,"message":{"role":"assistant","content":"Man patīk pīrāgi."},"finish_reason":"stop"}],"usage":{"prompt_tokens":24,"completion_tokens":6,"total_tokens":30}}

{"id":"chatcmpl-b002","object":"chat.completion","created":1759750900,"model":"tildeopen-30b-64k","choices":[{"index":0,"message":{"role":"assistant","content":"Sąskaita faktūra turi būti apmokėta per 30 dienų."},"finish_reason":"stop"}],"usage":{"prompt_tokens":52,"completion_tokens":14,"total_tokens":66}}

{"id":"chatcmpl-b003","object":"chat.completion","created":1759750901,"model":"tildeopen-30b-64k","choices":[{"index":0,"message":{"role":"assistant","content":"Tere hommikust."},"finish_reason":"stop"}],"usage":{"prompt_tokens":22,"completion_tokens":5,"total_tokens":27}}

{"id":"chatcmpl-b004","object":"chat.completion","created":1759750901,"model":"tildeopen-30b-64k","choices":[{"index":0,"message":{"role":"assistant","content":"Yes, you like pies."},"finish_reason":"stop"}],"usage":{"prompt_tokens":20,"completion_tokens":6,"total_tokens":26}}

{"id":"chatcmpl-b005","object":"chat.completion","created":1759750902,"model":"tildeopen-30b-64k","choices":[{"index":0,"message":{"role":"assistant","content":"Large language models are transformer-based models trained on large text corpora for language tasks; their reliability depends on the quality of the training data."},"finish_reason":"stop"}],"usage":{"prompt_tokens":70,"completion_tokens":30,"total_tokens":100}}

With DataProcessing.JoinSource = "Input"​

SageMaker merges each input record with its result; the response is placed under the key SageMakerOutput, so inputs and outputs are matched on the same line.

{"messages":[{"role":"system","content":"You are a translator."},{"role":"user","content":"I like pies.\n\nTranslate the given text from English to Latvian."}],"max_tokens":256,"temperature":0.0,"SageMakerOutput":{"id":"chatcmpl-b001","object":"chat.completion","created":1759750900,"model":"tildeopen-30b-64k","choices":[{"index":0,"message":{"role":"assistant","content":"Man patīk pīrāgi."},"finish_reason":"stop"}],"usage":{"prompt_tokens":24,"completion_tokens":6,"total_tokens":30}}}

{"messages":[{"role":"system","content":"You are answering questions."},{"role":"user","content":"I like pies.\n\nDo I like pies?"}],"max_tokens":128,"temperature":0.0,"SageMakerOutput":{"id":"chatcmpl-b004","object":"chat.completion","created":1759750901,"model":"tildeopen-30b-64k","choices":[{"index":0,"message":{"role":"assistant","content":"Yes, you like pies."},"finish_reason":"stop"}],"usage":{"prompt_tokens":20,"completion_tokens":6,"total_tokens":26}}}