LLM Output Evaluation
İleritestingEn az 32K bağlam
Defines repeatable quality evaluation for LLM outputs using representative datasets, scoring rubrics, model-graded checks, human review sampling, and regression thresholds.
Kullanım alanları
- LLM quality gates
- Prompt benchmarking
- Model migration validation
Örnek prompt
Create an evaluation plan for this LLM feature, including representative cases, a scored rubric, automated checks, human-review sampling, and release thresholds.
Önerilen modeller
Uyumlu araçlar
claude-codecursorkiroany
Modaliteler
Giriş: text, file
→Çıkış: text, code
İlgili Skills
Yazar
OpenModels Community