Skip to main content
Unsafe Output is an LLM-as-a-judge metric for application-specific output controls. It evaluates the latest model output using the complete LLM span context, including the system instructions, conversation, tool definitions and results, and model output. The metric applies only the rules that you add. It does not inject a fixed baseline policy or treat general best practices as requirements.
You must duplicate this preset and add your rules before using it. The original preset is read-only and does not include built-in output-policy rules.

Unsafe Output at a glance

When to use this metric

Configure Unsafe Output

Unsafe Output requires your own output-policy rules.
1

Create a copy of the Unsafe Output metric

From the Metrics Hub, select Unsafe Output, then select Duplicate metric. The copy is an editable custom LLM-as-a-judge metric.
2

Locate the user-defined rules section

Find this section in the prompt:
3

Add your output-policy rules

Replace {{ Add your rules here }} with a numbered list of rules. Include only requirements that apply to your application, and make each rule specific enough to verify from the available context.For example:
These are examples, not built-in rules. Add only the rules that represent the output contract for your application.
4

Save and test the metric

Save the duplicate, then test it on representative safe and violating examples before enabling it for a Log Stream or experiment.

How it works

Unsafe Output evaluates rules in three stages:
1

Determine applicability

The judge first determines whether each configured rule applies to the task and output. An inapplicable rule passes by default.
2

Check concrete evidence

For each applicable rule, the judge looks for a visible or materially implied violation in the output. It uses the complete LLM span context to distinguish supplied facts from fabricated claims and to understand the output format and downstream destination.
3

Return the classification

The result is true when at least one applicable rule fails. The result is false when every applicable rule passes. The generated explanation cites the evidence used for the decision.

Interpret the score

  • false / No: The judge found no concrete violation of the configured rules.
  • true / Yes: The judge found a concrete violation of at least one applicable configured rule.
When you configure multiple judges, Galileo averages their Boolean results. A higher score means that more judges classified the output as unsafe.
A false result does not certify that an output is safe in every respect. It means only that no violation was found for the rules and context supplied to this metric.

Best practices

  • Write one verifiable requirement per numbered rule.
  • State when a conditional rule applies. For example, start a JSON-specific rule with “When the response is JSON.”
  • Name the downstream destination when representation matters. Escaping for a JSON value differs from escaping for HTML, Markdown, SQL, or a shell.
  • Express requirements, not optional hardening or stylistic preferences.
  • Include both passing and violating edge cases when you test the metric.
  • Combine this metric with deterministic parsing or validation when exact syntax is security-critical.
Unsafe Output is an evaluation metric, not a runtime guardrail. It does not sanitize, block, or rewrite model output.

Performance benchmarks

We evaluated the canonical Unsafe Output prompt against trusted labels on an internal dataset of 877 output-policy examples. Unresolved annotation cases were excluded from the evaluation. The positive class is true (unsafe).

Gemini 3.5 Flash classification report

Accuracy: 0.9339 (819 of 877 examples).
Benchmarks are based on an internal evaluation dataset. Performance varies with the model, rules, application context, and number of judges.