Machine Translation
Contextual question answering
Summarisation
Deploying and using the model
1. Deploy a real-time endpoint
1.1 Console
- Open the product page in AWS Marketplace and choose Continue to subscribe, then Continue to configuration.
- Select your region and the model package version.
- Choose View in Amazon SageMaker and follow the launch wizard: name the model, pick the recommended instance type, set instance count to 1, and create the endpoint.
- Wait until the endpoint status is InService (typically 10–15 minutes, because ~61 GB of weights are loaded onto the GPUs).
1.2 Python SDK
import sagemaker
from sagemaker import ModelPackage
session = sagemaker.Session()
role = sagemaker.get_execution_role()
model_package_arn = "<ModelPackageArn for your region>"
model = ModelPackage(
role=role,
model_package_arn=model_package_arn,
sagemaker_session=session,
)
predictor = model.deploy(
initial_instance_count=1,
instance_type="<recommended instance type>",
endpoint_name="tildeopen-30b-64k",
container_startup_health_check_timeout=1800,
)
2. Invoke the endpoint
2.1 Build the prompt
The model expects one of six fixed templates. Use a helper so the structure is always exact:
import json
def translate(text, src_lang, trg_lang, terms=None, examples=None):
messages = [{"role": "system", "content": "You are a translator."}]
for src, trg in (examples or []):
messages.append({"role": "user", "content": f"{src}\n\nTranslate the given text from {src_lang} to {trg_lang}."})
messages.append({"role": "assistant", "content": trg})
instr = f"Translate the given text from {src_lang} to {trg_lang}"
if terms:
instr += f", using this glossary {json.dumps(terms, ensure_ascii=False)}"
else:
instr += "."
messages.append({"role": "user", "content": f"{text}\n\n{instr}."})
return messages
def answer(context, question):
assert "\n" not in question
return [{"role": "system", "content": "You are answering questions."},
{"role": "user", "content": f"{context}\n\n{question}"}]
def summarise(context, instruction="Summarise the given text. Focus on the main points."):
assert "\n" not in instruction
return [{"role": "system", "content": "You are answering questions."},
{"role": "user", "content": f"{context}\n\n{instruction}"}]
2.2 Python (boto3)
import boto3, json
runtime = boto3.client("sagemaker-runtime")
payload = {
"messages": translate("I like pies.", "English", "Latvian", terms={"pie": ["kūka"]}),
"max_tokens": 256,
"temperature": 0.0,
}
response = runtime.invoke_endpoint(
EndpointName="tildeopen-30b-64k",
ContentType="application/json",
Accept="application/json",
Body=json.dumps(payload),
)
result = json.loads(response["Body"].read())
print(result["choices"][0]["message"]["content"])
2.3 AWS CLI
aws sagemaker-runtime invoke-endpoint \
--endpoint-name tildeopen-30b-64k \
--content-type application/json \
--accept application/json \
--body fileb://request.json \
response.json
cat response.json
where request.json contains one of the request bodies from the input samples.
2.4 Streaming (optional)
Add "stream": true" to the body and call invoke_endpoint_with_response_stream; read choices[0].delta.content from each event.
3. Run a batch transform job
-
Prepare a .jsonl file with one request object per line (see section 6.5, Batch Transform) and upload it to S3.
-
Create a transform job:
transformer = model.transformer(
instance_count=1,
instance_type="<recommended instance type>",
strategy="SingleRecord",
assemble_with="Line",
accept="application/jsonlines",
output_path="s3://<bucket>/tildeopen/output/",
)
transformer.transform(
data="s3://<bucket>/tildeopen/input/batch_input.jsonl",
content_type="application/jsonlines",
split_type="Line",
join_source="Input",
)
transformer.wait()
- Read batch_input.jsonl.out from the output path. With join_source="Input" each line contains the original request plus the response under SageMakerOutput.