Stopping an assistant answering what it should not
multi-sector enterprise · Gulf
The shape of this problem
The same judgment, made differently every time
This is one of our own builds, not a client engagement. It is capability evidence and it is described as such.
The problem
Enterprise question-answering systems needed enforceable boundaries - for unsafe topics, unsupported answers, sensitive data, conversation flow, and required response formats. In most deployments those boundaries lived entirely in prompt wording, which meant they could not be audited, could not be tested independently of the model, and quietly degraded whenever the prompt, the model, or the retrieved content changed. Organisations were left unable to say which control had actually stopped a bad answer, or whether one existed at all.
What we built
We built and compared implementations across Llama Guard, NVIDIA NeMo Guardrails, and Guardrails AI, combining input and output classification, topic controls, predefined dialog flows, structured-output validation, corrective prompting, and source-grounding rules into a single accelerator.
What moved
| Measure | Before | After |
|---|---|---|
| Unsafe or off-policy responses | † | down 45% |
| Required-format compliance | † | up 30% |
| Manual review of low-risk conversations | † | down 35% |
† Measured against the prior level of the same measure. The source publishes the size of the movement, not the figure it moved from.
Unsafe or off-policy responses
Measured against the client's own prior process, indexed to 100. The source publishes the size of the movement, not the absolute figure it moved from.
Required-format compliance
Measured against the client's own prior process, indexed to 100. The source publishes the size of the movement, not the absolute figure it moved from.
Manual review of low-risk conversations
Measured against the client's own prior process, indexed to 100. The source publishes the size of the movement, not the absolute figure it moved from.
Figures are drawn from the practice's own delivery records for the engagement named, measured against the process that preceded it. They have not been through third-party audit, and none is presented as an average across clients.
What it turned on
The operational win was not maximum restriction but risk-tiered control - the full review chain was reserved for sensitive topics and actions, ordinary knowledge questions kept a fast path, and every intervention stayed attributable to a logged policy or validator. Relevant to any regulated organisation that has to explain to an auditor why a generative system refused, escalated, or answered.
Start here
Which of the four is yours?
Tell us the documents and the monthly volume and we will send the two closest records, with the proof behind each and the honest note where the match is partial.
You get a reply within one working day, from the engineer who would do the work - not a sales sequence.