Skip to main content
Machine Translation Contextual question answering Summarisation

Deploying and using the model

1. Deploy a real-time endpoint​

1.1 Console​

  1. Open the product page in AWS Marketplace and choose Continue to subscribe, then Continue to configuration.
  2. Select your region and the model package version.
  3. Choose View in Amazon SageMaker and follow the launch wizard: name the model, pick the recommended instance type, set instance count to 1, and create the endpoint.
  4. Wait until the endpoint status is InService (typically 10–15 minutes, because ~61 GB of weights are loaded onto the GPUs).

1.2 Python SDK​

import sagemaker
from sagemaker import ModelPackage

session = sagemaker.Session()
role = sagemaker.get_execution_role()

model_package_arn = "<ModelPackageArn for your region>"

model = ModelPackage(
role=role,
model_package_arn=model_package_arn,
sagemaker_session=session,
)

predictor = model.deploy(
initial_instance_count=1,
instance_type="<recommended instance type>",
endpoint_name="tildeopen-30b-64k",
container_startup_health_check_timeout=1800,
)

2. Invoke the endpoint​

2.1 Build the prompt​

The model expects one of six fixed templates. Use a helper so the structure is always exact:

import json



def translate(text, src_lang, trg_lang, terms=None, examples=None):

messages = [{"role": "system", "content": "You are a translator."}]

for src, trg in (examples or []):

messages.append({"role": "user", "content": f"{src}\n\nTranslate the given text from {src_lang} to {trg_lang}."})

messages.append({"role": "assistant", "content": trg})

instr = f"Translate the given text from {src_lang} to {trg_lang}"

if terms:

instr += f", using this glossary {json.dumps(terms, ensure_ascii=False)}"

else:

instr += "."

messages.append({"role": "user", "content": f"{text}\n\n{instr}."})

return messages



def answer(context, question):

assert "\n" not in question

return [{"role": "system", "content": "You are answering questions."},

{"role": "user", "content": f"{context}\n\n{question}"}]



def summarise(context, instruction="Summarise the given text. Focus on the main points."):

assert "\n" not in instruction

return [{"role": "system", "content": "You are answering questions."},

{"role": "user", "content": f"{context}\n\n{instruction}"}]

2.2 Python (boto3)​

import boto3, json



runtime = boto3.client("sagemaker-runtime")



payload = {

"messages": translate("I like pies.", "English", "Latvian", terms={"pie": ["kūka"]}),

"max_tokens": 256,

"temperature": 0.0,

}



response = runtime.invoke_endpoint(

EndpointName="tildeopen-30b-64k",

ContentType="application/json",

Accept="application/json",

Body=json.dumps(payload),

)



result = json.loads(response["Body"].read())

print(result["choices"][0]["message"]["content"])

2.3 AWS CLI​

aws sagemaker-runtime invoke-endpoint \

--endpoint-name tildeopen-30b-64k \

--content-type application/json \

--accept application/json \

--body fileb://request.json \

response.json



cat response.json

where request.json contains one of the request bodies from the input samples.

2.4 Streaming (optional)​

Add "stream": true" to the body and call invoke_endpoint_with_response_stream; read choices[0].delta.content from each event.

3. Run a batch transform job​

  1. Prepare a .jsonl file with one request object per line (see section 6.5, Batch Transform) and upload it to S3.

  2. Create a transform job:

transformer = model.transformer(

instance_count=1,

instance_type="<recommended instance type>",

strategy="SingleRecord",

assemble_with="Line",

accept="application/jsonlines",

output_path="s3://<bucket>/tildeopen/output/",

)



transformer.transform(

data="s3://<bucket>/tildeopen/input/batch_input.jsonl",

content_type="application/jsonlines",

split_type="Line",

join_source="Input",

)

transformer.wait()
  1. Read batch_input.jsonl.out from the output path. With join_source="Input" each line contains the original request plus the response under SageMakerOutput.