LLM Output Evaluation

ПродвинутыйtestingМинимальный контекст: 32K

Defines repeatable quality evaluation for LLM outputs using representative datasets, scoring rubrics, model-graded checks, human review sampling, and regression thresholds.

Варианты использования

  • LLM quality gates
  • Prompt benchmarking
  • Model migration validation

Пример промпта

Create an evaluation plan for this LLM feature, including representative cases, a scored rubric, automated checks, human-review sampling, and release thresholds.

Рекомендуемые модели

Совместимые инструменты

claude-codecursorkiroany

Модальности

Вход: text, file
Выход: text, code

Похожие Skills

Автор

OpenModels Community

@openmodelsrun
LLM Output Evaluation — AI Agent Skill | OpenModels