Batch Inference Planner
MittelopsMindestens 32K Kontext
Plans large offline LLM workloads using batch APIs or self-managed queues. Decides which jobs can tolerate delayed results, sizes batches and concurrency against rate limits, designs idempotent job records, retries, and partial-failure handling, and compares batch discounts against synchronous cost.
Anwendungsfälle
- Moving nightly classification or enrichment jobs to a batch API
- Backfilling embeddings or summaries over millions of records
- Designing retry and deduplication for failed batch items
- Comparing batch discount savings against turnaround requirements
Beispiel-Prompt
Plan this workload as batch inference. Context: [task, record count, token sizes, deadline, provider limits and prices] Return: 1. Whether batch fits, and what must stay synchronous. 2. Batch sizing, concurrency, and schedule. 3. Job state model with idempotency keys. 4. Retry, partial-failure, and validation handling. 5. Cost and completion-time estimate.
Empfohlene Modelle
Kompatible Werkzeuge
claude-codecursorkiroany
Modalitäten
Eingabe: text, code
→Ausgabe: text, code
Ähnliche Skills
Autor
OpenModels Community