Agent Evaluation Design

高级测试最低上下文:32K

为 AI 智能体构建可重复的评估,在真实端到端场景中衡量任务成功率、工具使用正确性、事实依据、延迟、成本和安全恢复能力。

此描述由机器自动翻译,尚未经过人工审核。

使用场景

  • 智能体发布门禁
  • 工具使用质量评估
  • 生产环境智能体评测

示例提示词

Design an evaluation suite for this agent. Define representative tasks, success criteria, tool assertions, cost and latency metrics, safety checks, and release thresholds.

推荐模型

兼容工具

claude-codecursorkiroany

模态

输入: text, code, file
输出: text, code

作者

OpenModels Community

@openmodelsrun