Machine Translation
Contextual question answering
Summarisation
Recommended configuration and performance settings
1. Recommended settings
| Setting | Recommendation |
|---|---|
| temperature | 0.0 for translation and QA; up to 0.3 for summarisation if more varied wording is wanted |
| max_tokens | Translation: about 1.5 × input tokens. QA: 128–256. Summarisation: as required by the instruction |
| Request size | Keep input plus max_tokens under 65,536 tokens. For documents longer than ~50k tokens, split by section |
| Concurrency per instance | 128–256 concurrent requests. Above 256 the server queues and time-to-first-token grows; scale out instead |
| Auto scaling | Target-tracking on SageMakerVariantInvocationsPerInstance or ConcurrentRequestsPerModel, target ~200 |
| Batch sizing | One transform instance handles roughly 13M output tokens per hour at ~480-token segments; longer inputs reduce this because prefill dominates |
2. Verify the deployment
Send the five real-time request samples (Appendix A) and compare the responses with the published response samples (Appendix B). Translations should match semantically (exact wording may differ slightly between container versions); the QA sample should return an affirmative answer; the summarisation sample should return a 2–4 sentence summary.
3. Clean up
Delete the endpoint when not in use; it is billed per instance-hour while running.
predictor.delete_endpoint()
model.delete_model()
Transform jobs stop automatically when the input is processed.