Adaptive Thinking Effort Tuning
ПродвинутыйopsМинимальный контекст: 32K
Chooses reasoning effort levels per workload for models that expose adaptive thinking or reasoning-effort controls. Segments traffic by task difficulty, measures quality, latency, and cost at each effort level, and produces a routing table with defaults, escalation rules, and monitoring so teams stop paying for maximum effort on easy requests.
Варианты использования
- Setting default effort for chat, extraction, and agentic coding routes
- Finding the lowest effort level that still passes an eval suite
- Adding escalation to higher effort after a failed first attempt
- Budgeting latency and token cost for reasoning-heavy features
Пример промпта
Tune reasoning effort for these workloads. Context: [workloads, model, current effort settings, eval results, latency and cost targets] Return: 1. Workload segments by difficulty and risk. 2. Experiment design across effort levels. 3. Recommended effort per segment with trade-offs. 4. Escalation and fallback rules. 5. Metrics and alerts to watch after rollout.
Рекомендуемые модели
Совместимые инструменты
claude-codecursorkiroany
Модальности
Вход: text, code
→Выход: text, code
Похожие Skills
Автор
OpenModels Community