prompt_guard, plus
its settings. Settings can include thresholds, entity lists, and allow or deny
lists. You
create and edit detectors in the console’s Detectors screen, and the same
detector can be referenced by many policies.
Detectors report what they find and their confidence. They do not decide whether
to block, mask, or allow traffic. That
decision lives on the policy that uses the
detector (its Gates and Detectors tabs). This separation lets you reuse
one well-tuned detector across many policies with different enforcement.
What defines a detector
A detector does not carry a mode, a direction, or a protocol. Those belong
to the policy rule that puts the detector to
work (the Input / Output phase). At request time, the collector must send
the matching
direction on /v1/evaluate. TrustGate
sets it automatically; application SDKs and other
collectors must set it themselves.
Settings
Every catalog detector exposes its own settings schema (the catalog API,GET /v1/plugins, returns each detector type and its fields). Examples:
prompt_guard,toxicity: athresholdin[0, 1].data_loss_prevention: which PII entities to detect or mask, plus custom keyword/regex rules.- Moderation (
prompt_moderation): keyword or regular-expression lists and NeuralTrust topic thresholds.
Mutable (transform-capable) detectors
Most detectors only read the payload. A mutable detector can rewrite it. Currently, onlydata_loss_prevention
which masks matched values in flight and populates transformed_payload.
Only a mutable detector can be used with the Transform action in a policy;
choosing Transform for any other detector is rejected when you save the policy.
Putting a detector to work
Creating a detector doesn’t run it. To evaluate traffic you reference the detector from a policy:- Open a policy and select the Detectors tab.
- Pick an evaluation phase: Input (prompt/request) or Output (completion/response).
- Add a rule that selects the detector and an action: Monitor (record a finding only), Block, or Transform (mutable detectors).
- Optionally add conditions so the rule only runs for certain consumers, models, collectors, protocols, sessions, directions, or tools.