While developing a chatbot for first-level ICT support, I encountered the need to filter both incoming questions and outgoing answers in some way.
The chatbot in question is based on Cheshire Cat AI (the “big cat”) and a RAG system that uses service sheets and FAQs published on the website as data sources. Essentially, it responds to users starting from official documentation, instead of searching for answers on the web or within basic knowledge.
A public chatbot, however, is a bit like a counter open to everyone: sooner or later, someone pastes a novel instead of a question, someone tries to make it forget instructions (the so-called prompt injection), someone writes their tax code or IBAN, and someone simply vents their frustration with inappropriate language.
This is how RAG Guardrails was born, a plugin that collects a series of filters (guards), each dedicated to one of these aspects and individually activatable and configurable.
Basically, bouncers for the big cat.
How the Plugin Works
The plugin intervenes at two points in the life of a question, using two hooks of Cheshire Cat AI, that is, points where a plugin can insert itself into the processing flow:
- on input (hook
fast_reply): this is the first point where the user’s message can be intercepted. If a guard triggers, the chatbot immediately responds with a polite message, configurable from the back office, inviting to rephrase the question or contact the Help Desk. The question does not reach the document search, does not consume language model tokens, and does not end up in the chatbot’s memory; - on output (hook
before_cat_sends_message): the generated response is checked before being sent and, if it contains personal data, it is replaced with a standard message, also configurable from the back office.
The guards use three different techniques, in increasing order of “intelligence” and cost:
- pure code: simple deterministic checks, such as the maximum message length;
- regular expressions and validations: recognize known patterns, such as email addresses, phone numbers, IBANs, tax codes, or typical phrases from those trying to manipulate the chatbot;
- classification models: small machine learning models, run locally, that evaluate the meaning of the text and catch what escapes regular expressions. They are more powerful but require memory and computation time, which is why they are disabled by default.
All response messages are bilingual (Italian and English) and customizable.
Plugin Features
The available guards are five, each belonging to a group indicating the type of risk it protects against:
| Guard | Phase | Group | Technique | Purpose |
|---|---|---|---|---|
message_length |
input | limits | code | Blocks messages longer than the set limit (1000 characters by default). |
personal_data |
input | privacy | regex and validations | Blocks messages containing emails, phone numbers, IBANs, or tax codes. |
prompt_injection |
input | security | regex and classifier (optional) | Blocks attempts to modify the chatbot’s instructions or to reveal the system prompt. |
offensive_input |
input | tone | classifier | Blocks offensive or violent messages. |
output_personal_data |
output | privacy | regex and validations | Prevents responses containing personal data from reaching the user. |
Privacy guards have a list of allowed contacts, configurable from the back office: public contacts (for example, a phone number of an office listed in the service sheets) can appear freely in questions and answers without triggering the alarm. The Help Desk address is always considered allowed.
Installation and Configuration

The plugin is published in the official Cheshire Cat AI plugin registry, so installation is really simple: just open the Plugins section of the admin panel, search for RAG Guardrails among the available plugins, install and activate it. The rest, including Python dependencies, is handled by the big cat. From the plugin settings page, you can configure practically everything: the Help Desk address to suggest to users, the list of allowed public contacts, the maximum message length, which types of personal data to block, the response texts, and for model-based guards, the model to use, the activation threshold, and the device to run it on (CPU or GPU).
Before going live, it is important at least to replace the placeholder address helpdesk@example.org with the real one.
For all details, see the project README.
Model Management
The two guards that use models are the one for prompt injection (in its optional part) and the one for offensive language. Models are selected from a menu and are all available on Hugging Face: for prompt injection, for example, there is Llama Prompt Guard 2 by Meta, which requires accepting its license and entering a Hugging Face token in the settings.
The plugin does not contain the models: they are downloaded at the first question that requires them, so the first user will experience some wait time. The necessary libraries (transformers and PyTorch) are also installed automatically by Cheshire Cat when activating the plugin, but they are not exactly lightweight.
For this reason, it is recommended to prepare a Docker image of the big cat where the plugin dependencies and, if used, the models are already present; if you don’t have a GPU, it is better to use the CPU-only version of PyTorch, which is much lighter. Then, by enabling the Preload classifiers on plugin activation option, enabled models are loaded at system startup, and no user will have to wait.
One last note: if a model fails to load, the guard lets it pass (fail open) instead of blocking everything. Better a distracted bouncer than a closed venue.
Usage Examples
Here is an example of filtering personal data:

Here is an example of a barrier to a prompt injection:

Example of offensive content:

Conclusions
RAG Guardrails has made our virtual help desk more robust: problematic messages are stopped before they even reach the language model, personal data stays out of the system, and users still receive a polite response with contact information.
Of course, it is not a magic solution: it does not verify that answers are correct or based on documents, and a sufficiently creative user can always find a way to bypass a filter. However, it is a simple, economical, and easy-to-adapt first line of defense to which new guards can be added in the future.
The code is open source, released under the GPL v3 license: if you manage a chatbot based on Cheshire Cat AI, try it and let me know how it went.
Sources and References
- Cheshire Cat AI, official website.
- Cheshire Cat AI, repository of version 1.9.2 (the one for which this plugin was made).
- RAG Guardrails, GitHub repository.
- Llama Prompt Guard 2, one of the models usable for prompt injection, on Hugging Face.
- Uptime Kuma Connector for the big cat, the other plugin for Cheshire Cat I talked about on this blog.
*** Note: This article was automatically translated using a workflow created with n8n and OpenAI. The original version of the post is the Italian one.


