Batch Inference Planner

СреднийopsМинимальный контекст: 32K

Plans large offline LLM workloads using batch APIs or self-managed queues. Decides which jobs can tolerate delayed results, sizes batches and concurrency against rate limits, designs idempotent job records, retries, and partial-failure handling, and compares batch discounts against synchronous cost.

Варианты использования

  • Moving nightly classification or enrichment jobs to a batch API
  • Backfilling embeddings or summaries over millions of records
  • Designing retry and deduplication for failed batch items
  • Comparing batch discount savings against turnaround requirements

Пример промпта

Plan this workload as batch inference.

Context: [task, record count, token sizes, deadline, provider limits and prices]

Return:
1. Whether batch fits, and what must stay synchronous.
2. Batch sizing, concurrency, and schedule.
3. Job state model with idempotency keys.
4. Retry, partial-failure, and validation handling.
5. Cost and completion-time estimate.

Рекомендуемые модели

Совместимые инструменты

claude-codecursorkiroany

Модальности

Вход: text, code
→
Выход: text, code

Похожие Skills

Автор

OpenModels Community

@openmodelsrun