Prompt Caching Strategy
MittelopsMindestens 32K Kontext
Designs a prompt caching strategy for LLM applications. Restructures prompts so stable content such as system instructions, tool definitions, and reference documents forms a cacheable prefix, places cache breakpoints, estimates break-even from cache-write and cache-read prices and hit rates, and defines monitoring for cache effectiveness.
Anwendungsfälle
- Cutting input cost for agents with long system prompts and tool lists
- Ordering prompt sections so the cacheable prefix stays stable
- Estimating savings from cache-read pricing at a given hit rate
- Diagnosing low cache hit rates in production
Beispiel-Prompt
Design prompt caching for this application. Context: [prompt structure, tools, documents, request volume, provider and prices] Return: 1. Which sections are stable, semi-stable, and volatile. 2. Reordered prompt layout and cache breakpoints. 3. Break-even and savings estimate with assumptions. 4. Cache invalidation risks and TTL considerations. 5. Metrics to track hit rate and cost per request.
Empfohlene Modelle
Kompatible Werkzeuge
claude-codecursorkiroany
Modalitäten
Eingabe: text, code
→Ausgabe: text, code
Ähnliche Skills
Autor
OpenModels Community