LLM Output Evaluation
ПродвинутыйtestingМинимальный контекст: 32K
Defines repeatable quality evaluation for LLM outputs using representative datasets, scoring rubrics, model-graded checks, human review sampling, and regression thresholds.
Варианты использования
- LLM quality gates
- Prompt benchmarking
- Model migration validation
Пример промпта
Create an evaluation plan for this LLM feature, including representative cases, a scored rubric, automated checks, human-review sampling, and release thresholds.
Рекомендуемые модели
Совместимые инструменты
claude-codecursorkiroany
Модальности
Вход: text, file
→Выход: text, code
Похожие Skills
Автор
OpenModels Community