Safety evaluations are essential for ensuring your LLM behaves responsibly across different categories like toxicity, prompt injections, and other unsafe behaviors.
EvaluationScenario. They are metadata for grouping — running hate + DAN is not a full compliance mapping.
FileSystemClient.get_overview() and NeuralTrust metrics roll up by those tags. See Threat detection overview.
Configure Safety Scenarios
UseUnsafeOutputsScenarioBuilder to evaluate whether your model generates harmful content, and SingleTurnScenarioBuilder to test prompt injection resistance.