LLM Eval Harness Builder
ПродвинутыйdataМинимальный контекст: 32K
Designs evaluation harnesses for LLM applications, covering dataset construction, task-specific metrics, LLM-as-judge rubrics with bias controls, and regression gates. Helps teams measure quality, catch regressions across model or prompt changes, and report results with confidence intervals rather than vibes.
Варианты использования
- Building an eval set and metrics for a RAG assistant
- Designing an LLM-as-judge rubric with bias mitigations
- Adding a regression gate for prompt and model changes in CI
- Reporting eval results with statistical confidence
Пример промпта
We have a support chatbot backed by an LLM. Design an evaluation harness: propose an eval dataset structure, define metrics for helpfulness/faithfulness/safety, write an LLM-as-judge rubric that controls for position and verbosity bias, and outline a CI regression gate.
Рекомендуемые модели
Совместимые инструменты
claude-codecursorkiroany
Модальности
Вход: text, code
→Выход: text, code
Похожие Skills
Автор
OpenModels Community