Batch Inference Planner
IntermediaopsContexto mínimo: 32K
Plans large offline LLM workloads using batch APIs or self-managed queues. Decides which jobs can tolerate delayed results, sizes batches and concurrency against rate limits, designs idempotent job records, retries, and partial-failure handling, and compares batch discounts against synchronous cost.
Casos de uso
- Moving nightly classification or enrichment jobs to a batch API
- Backfilling embeddings or summaries over millions of records
- Designing retry and deduplication for failed batch items
- Comparing batch discount savings against turnaround requirements
Prompt de ejemplo
Plan this workload as batch inference. Context: [task, record count, token sizes, deadline, provider limits and prices] Return: 1. Whether batch fits, and what must stay synchronous. 2. Batch sizing, concurrency, and schedule. 3. Job state model with idempotency keys. 4. Retry, partial-failure, and validation handling. 5. Cost and completion-time estimate.
Modelos recomendados
Herramientas compatibles
claude-codecursorkiroany
Modalidades
Entrada: text, code
→Salida: text, code
Skills relacionadas
Autor
OpenModels Community