LLM Eval Harness Builder
高级data最低上下文:32K
Designs evaluation harnesses for LLM applications, covering dataset construction, task-specific metrics, LLM-as-judge rubrics with bias controls, and regression gates. Helps teams measure quality, catch regressions across model or prompt changes, and report results with confidence intervals rather than vibes.
使用场景
- Building an eval set and metrics for a RAG assistant
- Designing an LLM-as-judge rubric with bias mitigations
- Adding a regression gate for prompt and model changes in CI
- Reporting eval results with statistical confidence
示例提示词
We have a support chatbot backed by an LLM. Design an evaluation harness: propose an eval dataset structure, define metrics for helpfulness/faithfulness/safety, write an LLM-as-judge rubric that controls for position and verbosity bias, and outline a CI regression gate.
推荐模型
兼容工具
claude-codecursorkiroany
模态
输入: text, code
→输出: text, code
相关 Skills
作者
OpenModels Community