LLM Output Evaluation
高级testing最低上下文:32K
Defines repeatable quality evaluation for LLM outputs using representative datasets, scoring rubrics, model-graded checks, human review sampling, and regression thresholds.
使用场景
- LLM quality gates
- Prompt benchmarking
- Model migration validation
示例提示词
Create an evaluation plan for this LLM feature, including representative cases, a scored rubric, automated checks, human-review sampling, and release thresholds.
推荐模型
兼容工具
claude-codecursorkiroany
模态
输入: text, file
→输出: text, code
相关 Skills
作者
OpenModels Community