LLM Output Evaluation

İleritestingEn az 32K bağlam

Defines repeatable quality evaluation for LLM outputs using representative datasets, scoring rubrics, model-graded checks, human review sampling, and regression thresholds.

Kullanım alanları

  • LLM quality gates
  • Prompt benchmarking
  • Model migration validation

Örnek prompt

Create an evaluation plan for this LLM feature, including representative cases, a scored rubric, automated checks, human-review sampling, and release thresholds.

Önerilen modeller

Uyumlu araçlar

claude-codecursorkiroany

Modaliteler

Giriş: text, file
Çıkış: text, code

İlgili Skills

Yazar

OpenModels Community

@openmodelsrun