LLM Guardrails Design

AvanzadasecurityContexto mínimo: 64K

Designs input and output guardrails for an LLM product. Defines the policy to enforce, places checks at the right points such as input filtering, retrieval and tool-result screening, and output validation, chooses between rules, classifiers, and model-based checks, and specifies fallback behavior, logging, and evaluation of false positives and negatives.

Casos de uso

  • Adding policy enforcement to a customer-facing assistant
  • Screening retrieved documents and tool results for injected instructions
  • Validating structured outputs before they reach downstream systems
  • Measuring and reducing guardrail false positives

Prompt de ejemplo

Design guardrails for this LLM application.

Context: [product, users, policy requirements, tools and data sources, risk tolerance]

Return:
1. Policy categories and severity.
2. Check placement across the request path.
3. Mechanism per check with latency and cost.
4. Fallback and user messaging on a block.
5. Evaluation set and metrics for the guardrails.

Modelos recomendados

Herramientas compatibles

claude-codecursorkiroany

Modalidades

Entrada: text, code
→
Salida: text, code

Skills relacionadas

Autor

OpenModels Community

@openmodelsrun