Batch Inference Planner

IntermediaopsContexto mínimo: 32K

Plans large offline LLM workloads using batch APIs or self-managed queues. Decides which jobs can tolerate delayed results, sizes batches and concurrency against rate limits, designs idempotent job records, retries, and partial-failure handling, and compares batch discounts against synchronous cost.

Casos de uso

  • Moving nightly classification or enrichment jobs to a batch API
  • Backfilling embeddings or summaries over millions of records
  • Designing retry and deduplication for failed batch items
  • Comparing batch discount savings against turnaround requirements

Prompt de ejemplo

Plan this workload as batch inference.

Context: [task, record count, token sizes, deadline, provider limits and prices]

Return:
1. Whether batch fits, and what must stay synchronous.
2. Batch sizing, concurrency, and schedule.
3. Job state model with idempotency keys.
4. Retry, partial-failure, and validation handling.
5. Cost and completion-time estimate.

Modelos recomendados

Herramientas compatibles

claude-codecursorkiroany

Modalidades

Entrada: text, code
→
Salida: text, code

Skills relacionadas

Autor

OpenModels Community

@openmodelsrun