Skip to main content
Machine Translation Contextual question answering Summarisation

Recommended configuration and performance settings

SettingRecommendation
temperature0.0 for translation and QA; up to 0.3 for summarisation if more varied wording is wanted
max_tokensTranslation: about 1.5 × input tokens. QA: 128–256. Summarisation: as required by the instruction
Request sizeKeep input plus max_tokens under 65,536 tokens. For documents longer than ~50k tokens, split by section
Concurrency per instance128–256 concurrent requests. Above 256 the server queues and time-to-first-token grows; scale out instead
Auto scalingTarget-tracking on SageMakerVariantInvocationsPerInstance or ConcurrentRequestsPerModel, target ~200
Batch sizingOne transform instance handles roughly 13M output tokens per hour at ~480-token segments; longer inputs reduce this because prefill dominates

2. Verify the deployment​

Send the five real-time request samples (Appendix A) and compare the responses with the published response samples (Appendix B). Translations should match semantically (exact wording may differ slightly between container versions); the QA sample should return an affirmative answer; the summarisation sample should return a 2–4 sentence summary.

3. Clean up​

Delete the endpoint when not in use; it is billed per instance-hour while running.

predictor.delete_endpoint()

model.delete_model()

Transform jobs stop automatically when the input is processed.