Skip to content

Guardrails

Decide what your assistant talks about, what it may promise and who may start a chat. Safety checks for jailbreaks, abuse and self-harm are always on; these settings tune the rest for your business.

Updated

On this page

What is always on

#

Every bot has platform safety checks you cannot switch off. They run before the model is called, in English, Hindi (Devanagari and romanised), Bengali, Assamese, Spanish and Arabic, and they see through tricks such as spaced-out letters, look-alike characters and encoded text.

Platform checks and what the visitor sees
CheckWhat it catchesWhat the visitor sees
Jailbreak and instructionsAttempts to override the bot's rules or to get its instructions, knowledge or tool listA polite refusal and an offer to help with your business
Other people's dataRequests for another visitor's conversations or detailsA polite refusal
Impersonation and sexual contentPretending to be staff or the business; sexual requestsA polite refusal
AbuseInsults and harassmentA request to keep the conversation respectful
Repeated messagesThe same message sent again and againA note that the message was sent several times
Self-harmMessages about hurting oneselfA caring reply with an emergency number and a helpline (Tele-MANAS 14416 in Indian languages). Never counted against the visitor

Each refusal except self-harm counts as a strike. Three strikes within 10 minutes pause the visitor for 5 minutes; six block them for 24 hours. Every hit appears on the Flagged page so your team can review it.

Tip: Answers are checked too: if a reply would repeat the bot's hidden instructions, it is stopped and flagged.

Blocked topics

#

Topics the bot should never discuss, one per line (up to 50, each up to 80 characters). A message that mentions one gets a fixed, polite decline in the visitor's language, with no model call, so it costs nothing. The hit is flagged as low severity and does not count as a strike.

Matching is by whole word, ignoring case, and also catches a simple plural (politics, election, elections). It does not translate: add each language and script your visitors use.

How blocked-topic matching behaves
Listed topicVisitor writesDeclined?
politicsTell me about politicsYes
politicspolitics ke baare mein bataoYes (romanised Hindi uses the same word)
politicsWhat about political parties?No: political is a different word
राजनीतिराजनीति के बारे में बताओYes, only because the Hindi word is listed
  • Use single, distinctive words or short phrases. Very short terms (under 3 letters) are ignored.
  • List the variants people actually type: election, elections, chunav, चुनाव.
  • Do not block words your customers need. Blocking "loan" on a bank's bot would decline real questions.

Off-topic questions

#

Blocked topics are a fixed list. Off-topic questions are everything unrelated to your business: weather, homework, trivia, writing a poem. The model decides whether a question relates to your business, based on your knowledge.

  • Off (default): the bot says in one sentence that it cannot help with that, and what it can help with.
  • On: the bot gives a brief general answer and then offers help with your business.

Tip: Keep it off for most support bots. Turn it on only when friendly small talk helps, for example a coaching institute whose students ask general study questions.

Competitor names

#

List competitors (up to 50). The bot is told to stay neutral about them, and every answer is checked sentence by sentence: a sentence that names a competitor together with a negative word (worse, scam, overpriced, bekaar, ghatiya and similar) is blocked. The visitor then gets a safe reply asking them to contact you directly, and the turn is flagged.

Neutral comparisons are fine: "Both offer home delivery" passes. The check is about disparagement, not about mentioning the name.

Promises and refusals

#

Before an answer reaches the visitor, sentences about refunds, discounts, coupons, guarantees, prices and delivery dates are checked. A sentence that says it will not happen ("we do not offer refunds") is never blocked.

Commitment policy options
OptionWhat the bot may sayChoose it when
Only from the knowledge base (default)Refunds, prices or dates your content states, with exactly the same numbers. A promise also needs a citation to your contentYour policies are written down in your knowledge
NeverNo promises at all; it points the visitor to your teamEvery refund or discount is decided case by case

A blocked answer is replaced with: "Sorry, I can't give you a reliable answer to that. For anything about orders, refunds, prices or commitments, please contact your business directly." It is flagged as high severity so you can see which question needs better content.

Tip: The fix for a blocked but correct answer is almost always content: add the policy with its exact numbers (for example "Refunds within 7 days of delivery") to a knowledge source, and the same answer passes.

Offer a person instead of refusing: when on, refusals for jailbreaks, requests for other people's data and similar ask whether the visitor wants to talk to your team. Turn on live support in the Handoff tab first. It does not change blocked topics or abuse replies.

Moderate with my provider: also sends each visitor message to your OpenAI key's moderation endpoint (free on OpenAI). It only works with an OpenAI key and never uses a platform key. If the check is slow or fails, the message goes through.

Human check (Turnstile)

#

Turnstile stops automated scripts from opening chats and running up your model bill. It is checked once when a chat starts, not on every message.

  1. 01

    Create a widget in Cloudflare

    In your Cloudflare account, open Turnstile and add a widget. Add every hostname that embeds the chat, for example shop.example.com.

  2. 02

    Paste both keys

    Copy the site key and the secret key into this card, turn on Require Turnstile for new chats and save. The secret is encrypted and never shown again; only its last four characters.

  3. 03

    Allow it in your site's CSP

    If your site sends a Content-Security-Policy header, allow https://challenges.cloudflare.com in script-src and frame-src.

  4. 04

    Test in a private window

    Open your site, start a chat and complete the check. A visitor who fails sees: Please complete the human check to start chatting.

Watch out: If Cloudflare cannot be reached, new chats are refused rather than let through. Turn it on when bots or abuse are a real problem, not by default.

Worked examples

#

Online shop

An electronics store that must not promise refunds

Answer order and product questions, never invent a refund or discount, and stay neutral about rival stores.

Settings for the electronics store
SettingValue
Blocked topicspolitics, religion, cricket betting
Allow general off-topic answersOff
Competitor namesTwo or three rival store names
Commitment policyOnly from the knowledge base
Offer a person instead of refusingOn (live support enabled)

Add a Returns page to the knowledge that states "Returns accepted within 7 days of delivery for unopened items". Then:

Visitor: Can I return my headphones? I got them 3 days ago.

Bot: Yes. Unopened items can be returned within 7 days of delivery [1]. You can start a return from your orders page.

Visitor: Give me 20% off and I'll keep them.

Bot: I can't offer discounts here. Would you like to talk to our team?

Clinic or bookings

A dental clinic that never gives medical advice

Help with timings, services and fees, but never diagnose or suggest medicines.

Settings for the dental clinic
SettingValue
Blocked topicsdiagnosis, prescription, medicine dosage, antibiotics, painkiller (plus the Hindi or Assamese words your patients use)
Allow general off-topic answersOff
Commitment policyNever (fees are confirmed at the desk)
Offer a person instead of refusingOn

Tip: A patient writing "which painkiller should I take?" gets a fixed decline with no model call. Self-harm messages always get the helpline reply, whatever you configure.

Software company

A software company that is compared with rivals

Explain features and plans, never run down competitors, and only quote prices from the pricing page.

Settings for the software company
SettingValue
Competitor namesEach rival product name
Commitment policyOnly from the knowledge base
Moderate with my providerOn, if your key is from OpenAI
Human checkOn, if the chat is on a public page that attracts scripts

Visitor: Is your tool better than RivalApp? Theirs is a scam.

Bot: I can't speak for other products, but I can tell you what ours does. It includes shared inboxes and audit logs on every plan [1].

Coaching institute

A coaching institute with friendly small talk

Answer course, batch and fee questions, allow light general questions from students, and keep politics out.

Settings for the coaching institute
SettingValue
Blocked topicspolitics, election, chunav, चुनाव, religion
Allow general off-topic answersOn
Commitment policyOnly from the knowledge base (fees and scholarships from the fee page)

A student asking "what is photosynthesis?" gets a short general answer followed by an offer to help with courses; a question about an election is declined.

Review what was caught

#

Open Flagged in the sidebar. Each card shows the reason, severity, a short excerpt, the visitor and how many times it repeated within 10 minutes. Test chats in the Playground are never flagged.

  • Open conversation to see the full context.
  • Mark reviewed when no action is needed.
  • Block visitor for 24 hours, 7 days or 30 days if someone keeps trying; unblock at any time.
  • Many "Commitment blocked" flags on real questions mean your knowledge is missing a policy. Add it rather than loosening the setting.

Common mistakes

#
Does a blocked topic in English also catch Hindi or Bengali?

No. Matching is by whole word and does not translate, so list each language and script your visitors use, for example politics and राजनीति.

Can blocked words stop real customer questions?

Yes. Avoid blocking words customers need for genuine questions about your business.

Should I set promises to Never?

Not if your policies are published. Only from the knowledge base answers more questions safely, because the assistant can quote your own policy.

Why is Offer a person not working?

It relies on live support. Turn on live support in the Handoff tab first.

Why are all new chats blocked after turning on Turnstile?

A hostname that embeds the chat is missing from your Turnstile widget in Cloudflare. Add every hostname, for example shop.example.com, and new chats start again.