AWS launches explicit prompt caching for OpenAI GPT-5.6 on Amazon Bedrock
Amazon Web Services has introduced explicit prompt caching for OpenAI's GPT-5.6 Sol, Terra, and Luna models on Amazon Bedrock to reduce inference costs.

What happened?
Amazon Web Services has launched explicit prompt caching for OpenAI's GPT-5.6 models—Sol, Terra, and Luna—on the Amazon Bedrock cloud platform. The features are now generally available, enabling developers to manually define exactly which segments of a prompt should be stored in memory and reused in subsequent calls.
Key facts
| Plattform | Amazon Bedrock |
|---|---|
| Modeller | GPT-5.6 Sol, Terra, Luna |
| Status | Generally Available |
Why it matters
By offering explicit control mechanisms over caching, companies can significantly reduce computational costs and latency for recurring instructions or large documents. This provides increased predictability and financial control when operating complex cloud-based AI services.
Who is affected?
The launch targets developers, system architects, and enterprises building large-scale AI applications on Amazon Bedrock. It is particularly relevant for organisations managing significant amounts of contextual data or repetitive system instructions.
Impact on the EU
The service and its updated caching features are available to Amazon Bedrock customers globally, including within EU regions where Bedrock is provided. Users must, however, ensure that the handling of stored data complies with local GDPR requirements.
What else you should know
Users looking to migrate existing workflows to the new GPT-5.6 models on Amazon Bedrock can configure caching directly within their API calls to optimise response times and performance.
Quick answers about this story
Vad har hänt?
När hände det?
Varför spelar det roll?
Vilka modeller omfattas av den nya funktionen?
The link opens in a new window and leads to the publisher's own site.
Källan har spårats automatiskt från utgivaren via Aheadlines signalkedja.
AI-verktyg i artikeln
Topics
Get similar news straight to your inbox
The reader's room
Send in a question or an addition. The newsroom reads everything before it's published and replies when relevant. No AI-generated text – just people.
Sign in to submit a comment or question.
Read the article through your role
- Decide whether this affects strategy over 6–12 months or is just noise.
- Discuss with leadership: do we own the right question or does ownership need to move?
- Ask: what risk are we taking by NOT acting on this this quarter?
Generated angle — not editorial analysis of "AWS launches explicit prompt caching for OpenAI GPT-5.6 on A"