Open-Weight Model Deployment
FortgeschrittenopsMindestens 32K Kontext
Designs a production deployment for an open-weight language or multimodal model from its model card, checkpoint layout, target workload, and available hardware. Verifies licensing and runtime support, estimates weights and KV-cache memory, selects a serving framework and parallelism plan, defines quantization and long-context settings, and creates quality, latency, security, and rollback gates before traffic is moved to the new model.
Anwendungsfälle
- Turning a newly released checkpoint and model card into a deployment plan
- Choosing between vLLM, SGLang, and another model-supported serving runtime
- Estimating weight, KV-cache, and multimodal encoder memory before provisioning GPUs
- Defining canary, quality, latency, safety, and rollback gates for a model upgrade
- Separating verified model capabilities from deployment assumptions that require testing
Beispiel-Prompt
Design a production deployment for this open-weight model. Model card or repository: [URL] Hardware: [GPU type and count] Traffic: [requests per second, input/output token distribution, image or video usage] SLOs: [time to first token, throughput, availability] Verify the license, supported runtimes, native and extended context limits, checkpoint precision, and multimodal requirements. Then propose the serving framework, tensor/pipeline parallelism, quantization, memory budget, batching, autoscaling, observability, canary evaluation, and rollback plan. Clearly label any value that must be benchmarked instead of treating it as known.
Empfohlene Modelle
Kompatible Werkzeuge
claude-codecursoropencodekiroany
Modalitäten
Eingabe: text, code, file
→Ausgabe: text, code
Ähnliche Skills
Autor
OpenModels Community