LLM Eval Harness Builder

高级data最低上下文:32K

Designs evaluation harnesses for LLM applications, covering dataset construction, task-specific metrics, LLM-as-judge rubrics with bias controls, and regression gates. Helps teams measure quality, catch regressions across model or prompt changes, and report results with confidence intervals rather than vibes.

使用场景

  • Building an eval set and metrics for a RAG assistant
  • Designing an LLM-as-judge rubric with bias mitigations
  • Adding a regression gate for prompt and model changes in CI
  • Reporting eval results with statistical confidence

示例提示词

We have a support chatbot backed by an LLM. Design an evaluation harness: propose an eval
dataset structure, define metrics for helpfulness/faithfulness/safety, write an LLM-as-judge
rubric that controls for position and verbosity bias, and outline a CI regression gate.

推荐模型

兼容工具

claude-codecursorkiroany

模态

输入: text, code
输出: text, code

相关 Skills

作者

OpenModels Community

@openmodelsrun