LLM Output Evaluation

高级testing最低上下文:32K

Defines repeatable quality evaluation for LLM outputs using representative datasets, scoring rubrics, model-graded checks, human review sampling, and regression thresholds.

使用场景

  • LLM quality gates
  • Prompt benchmarking
  • Model migration validation

示例提示词

Create an evaluation plan for this LLM feature, including representative cases, a scored rubric, automated checks, human-review sampling, and release thresholds.

推荐模型

兼容工具

claude-codecursorkiroany

模态

输入: text, file
输出: text, code

相关 Skills

作者

OpenModels Community

@openmodelsrun