RAG Guardrails for the Big Cat

While developing a chatbot for first-level ICT support, I encountered the need to filter both incoming questions and outgoing answers in some way.

The chatbot in question is based on Cheshire Cat AI (the “big cat”) and a RAG system that uses service sheets and FAQs published on the website as data sources. Essentially, it responds to users starting from official documentation, instead of searching for answers on the web or within basic knowledge.

A public chatbot, however, is a bit like a counter open to everyone: sooner or later, someone pastes a novel instead of a question, someone tries to make it forget instructions (the so-called prompt injection), someone writes their tax code or IBAN, and someone simply vents their frustration with inappropriate language.

This is how RAG Guardrails was born, a plugin that collects a series of filters (guards), each dedicated to one of these aspects and individually activatable and configurable.

Basically, bouncers for the big cat.

How the Plugin Works

The plugin intervenes at two points in the life of a question, using two hooks of Cheshire Cat AI, that is, points where a plugin can insert itself into the processing flow:

  • on input (hook fast_reply): this is the first point where the user’s message can be intercepted. If a guard triggers, the chatbot immediately responds with a polite message, configurable from the back office, inviting to rephrase the question or contact the Help Desk. The question does not reach the document search, does not consume language model tokens, and does not end up in the chatbot’s memory;
  • on output (hook before_cat_sends_message): the generated response is checked before being sent and, if it contains personal data, it is replaced with a standard message, also configurable from the back office.

The guards use three different techniques, in increasing order of “intelligence” and cost:

  • pure code: simple deterministic checks, such as the maximum message length;
  • regular expressions and validations: recognize known patterns, such as email addresses, phone numbers, IBANs, tax codes, or typical phrases from those trying to manipulate the chatbot;
  • classification models: small machine learning models, run locally, that evaluate the meaning of the text and catch what escapes regular expressions. They are more powerful but require memory and computation time, which is why they are disabled by default.

All response messages are bilingual (Italian and English) and customizable.

Plugin Features

The available guards are five, each belonging to a group indicating the type of risk it protects against:

Guard Phase Group Technique Purpose
message_length input limits code Blocks messages longer than the set limit (1000 characters by default).
personal_data input privacy regex and validations Blocks messages containing emails, phone numbers, IBANs, or tax codes.
prompt_injection input security regex and classifier (optional) Blocks attempts to modify the chatbot’s instructions or to reveal the system prompt.
offensive_input input tone classifier Blocks offensive or violent messages.
output_personal_data output privacy regex and validations Prevents responses containing personal data from reaching the user.

Privacy guards have a list of allowed contacts, configurable from the back office: public contacts (for example, a phone number of an office listed in the service sheets) can appear freely in questions and answers without triggering the alarm. The Help Desk address is always considered allowed.

Installation and Configuration

The plugin is published in the official Cheshire Cat AI plugin registry, so installation is really simple: just open the Plugins section of the admin panel, search for RAG Guardrails among the available plugins, install and activate it. The rest, including Python dependencies, is handled by the big cat. From the plugin settings page, you can configure practically everything: the Help Desk address to suggest to users, the list of allowed public contacts, the maximum message length, which types of personal data to block, the response texts, and for model-based guards, the model to use, the activation threshold, and the device to run it on (CPU or GPU).

Before going live, it is important at least to replace the placeholder address helpdesk@example.org with the real one.

For all details, see the project README.

Model Management

The two guards that use models are the one for prompt injection (in its optional part) and the one for offensive language. Models are selected from a menu and are all available on Hugging Face: for prompt injection, for example, there is Llama Prompt Guard 2 by Meta, which requires accepting its license and entering a Hugging Face token in the settings.

The plugin does not contain the models: they are downloaded at the first question that requires them, so the first user will experience some wait time. The necessary libraries (transformers and PyTorch) are also installed automatically by Cheshire Cat when activating the plugin, but they are not exactly lightweight.

For this reason, it is recommended to prepare a Docker image of the big cat where the plugin dependencies and, if used, the models are already present; if you don’t have a GPU, it is better to use the CPU-only version of PyTorch, which is much lighter. Then, by enabling the Preload classifiers on plugin activation option, enabled models are loaded at system startup, and no user will have to wait.

One last note: if a model fails to load, the guard lets it pass (fail open) instead of blocking everything. Better a distracted bouncer than a closed venue.

 

Usage Examples

Here is an example of filtering personal data:

Here is an example of a barrier to a prompt injection:

 

Example of offensive content:

Conclusions

RAG Guardrails has made our virtual help desk more robust: problematic messages are stopped before they even reach the language model, personal data stays out of the system, and users still receive a polite response with contact information.

Of course, it is not a magic solution: it does not verify that answers are correct or based on documents, and a sufficiently creative user can always find a way to bypass a filter. However, it is a simple, economical, and easy-to-adapt first line of defense to which new guards can be added in the future.

The code is open source, released under the GPL v3 license: if you manage a chatbot based on Cheshire Cat AI, try it and let me know how it went.

 

Sources and References

*** Note: This article was automatically translated using a workflow created with n8n and OpenAI. The original version of the post is the Italian one.

2 days ago

Leave a Reply

Your email address will not be published. Required fields are marked *

Comment moderation is enabled. Your comment may take some time to appear.