LLM Guardrails Design

FortgeschrittensecurityMindestens 64K Kontext

Designs input and output guardrails for an LLM product. Defines the policy to enforce, places checks at the right points such as input filtering, retrieval and tool-result screening, and output validation, chooses between rules, classifiers, and model-based checks, and specifies fallback behavior, logging, and evaluation of false positives and negatives.

Anwendungsfälle

  • Adding policy enforcement to a customer-facing assistant
  • Screening retrieved documents and tool results for injected instructions
  • Validating structured outputs before they reach downstream systems
  • Measuring and reducing guardrail false positives

Beispiel-Prompt

Design guardrails for this LLM application.

Context: [product, users, policy requirements, tools and data sources, risk tolerance]

Return:
1. Policy categories and severity.
2. Check placement across the request path.
3. Mechanism per check with latency and cost.
4. Fallback and user messaging on a block.
5. Evaluation set and metrics for the guardrails.

Empfohlene Modelle

Kompatible Werkzeuge

claude-codecursorkiroany

Modalitäten

Eingabe: text, code
→
Ausgabe: text, code

Ähnliche Skills

Autor

OpenModels Community

@openmodelsrun