Skip to content

Answers from your docs, with receipts.

Train an AI chatbot on your website and docs

Kothadesk reads what you already publish, your website, help centre, PDFs, Word files and FAQs, and your customer support chatbot answers from it, numbering each source it relies on. When your documents do not cover a question, it says so and offers a person instead of improvising.

Knowledge sources for a bot: a crawled site, uploaded files and an FAQ table

01

Supported knowledge sources.

Train the chatbot on the content you already have. Each source belongs to one bot, and indexing runs in the background on your own embedding key.

Knowledge source types
SourceWhat you provideDetails
WebsiteA start URLUp to 5,000 pages per source; whole site, one path or one page; linked PDFs included
JavaScript-rendered sitesThe same start URLPages that only show content after scripts run are rendered first; Automatic, Never or Always
SitemapA sitemap URLSitemap indexes and gzip supported; found automatically
PDF and DOCXFiles up to 25 MBHeadings kept for citations; 20 files per upload
CSV, TXT and HTMLFiles up to 25 MBPrice lists, timetables, exported pages
FAQ tableQuestion and answer pairsEdit inline or import up to 500 rows of CSV; an entry may cite a page
MarkdownFiles, or text you writeUp to 200,000 characters, with a safe preview
ImagesPhotos with a title and descriptionJPEG, PNG, WebP, GIF or HEIC

The crawler respects robots.txt, skips pages marked noindex and identifies itself as KothadeskBot. Pages behind a login cannot be crawled; upload them as files instead. Scanned PDFs without a text layer are not read yet.

02

How the chatbot builds an answer (RAG).

Retrieval-augmented generation, step by step: the model writes the reply, but only from passages Kothadesk found in your content.

  1. 01

    Add a source

    A website, a sitemap, files, an FAQ table or a Markdown page. Each source belongs to one bot.
  2. 02

    Kothadesk splits and indexes it

    Documents are cut at their headings into passages of about 400 tokens and embedded with your own embedding key. Unchanged content is not embedded twice.
  3. 03

    The question is understood first

    One short call works out the language, the intent and a standalone version of follow-up questions, while the search runs in parallel.
  4. 04

    Search combines meaning and words

    Vector and keyword search together, with relevance thresholds calibrated per bot so unrelated passages are dropped instead of quoted.
  5. 05

    The answer cites what it used

    Sources are numbered in the order the answer uses them, and only sources the answer still relies on after every check are listed.

03

Citations that mean something.

A number appears only where a fact came from your sources. Greetings, small talk and questions about the assistant itself are never cited.

The Kothadesk chat widget showing an answer with numbered citations and its sources
  1. 1One number per source, not per sentence
  2. 2Sources collapsed under the answer

04

Images in answers.

When a picture helps, the chatbot shows it inside the answer, right after the sentence it illustrates, with its caption and a link to where it came from.

Where images come from
Pictures in your web pages, PDFs and Word files, Markdown images in text and FAQ answers, and photos you upload.
Decoration left out
Icons, logos, SVGs, tracking pixels, tiny images and pictures repeated on many pages are skipped automatically.
You set the limits
Turn extraction and display on or off per bot, and allow one to three images per answer. Hide any single image.
Described on your key
Images without alt text or a caption can be described by the bot's own vision-capable model, so search can find them.

05

What happens when it does not know.

An honest fallback
  1. Visitor

    What's your refund window for gift cards?

  2. Assistant

    Our refund policy covers orders, but I don't see anything about gift cards. Would you like me to ask someone from the team?

06

Keeping knowledge fresh.

Your site changes; the chatbot should follow without a full rebuild every time.

Scheduled re-crawls
Re-crawl a website manually, every day or every week. Unchanged content is not embedded again, so re-syncs cost almost nothing.
Removed pages retire safely
A page that disappears stays in use until two full crawls in a row miss it, or it returns 404.
Delete what is wrong
Delete one document or a whole source, and its passages and images go with it.
Fill the gaps you find
Unanswered questions show what to write next; an FAQ entry is the fastest fix for an exact price or timing.

07

Controls you will actually use.

Citation display
Inline numbers with a source list, the list only, or off, per bot.
Retrieval test
Type a question and see which passages match, their scores and the bot's thresholds.
Re-sync and delete
Re-crawl a source on demand; delete one document or a whole source and its passages go with it.
Knowledge scope
A short summary of what the bot knows, written for you after ingestion and editable, so it can say what it covers.

08

Test answers with evaluation sets.

An evaluation set is a list of real questions with the documents that answer them. Run it before and after a change, and the scores tell you whether answers got better.

  1. 01

    Build the set

    Add questions from real conversations, generate cases from your sources, or import JSONL.
  2. 02

    Run it

    Kothadesk measures whether the right documents were found and whether the answer is faithful to them. The judge runs on your own key.
  3. 03

    Compare two runs

    Change a model, a threshold or your content, run again and compare the results side by side.

09

Citations in your own interface.

Building your own chat UI? The JavaScript SDK streams citation events alongside the text.

chat.ts
import { ChatClient } from "@kothadesk/sdk-js";

const client = new ChatClient({
  botId: "bot_01J9ZKQ3M4X8R2T6V0B1C5D7EF",
  apiUrl: "https://api.kothadesk.com",
});

const conversation = await client.createConversation();
const stream = client.sendMessage(conversation.id, "Can I change my plan mid-month?");

for await (const ev of stream) {
  if (ev.type === "token") render(stream.text());
  if (ev.type === "citation") addSource(ev.citation);
  if (ev.type === "message.end") render(ev.content);
}

10

Questions about knowledge.

Which file types can I upload?

PDF, Word (DOCX), Markdown, plain text, CSV and HTML, up to 25 MB each. Scanned PDFs without a text layer are not read yet. FAQ tables can also be imported from CSV.

Can one source be shared by several bots?

Not yet. Each knowledge source belongs to one bot, which keeps answers scoped to what that bot is for.

Can I hide citation numbers?

Yes. Per bot you can show inline numbers with a source list, the source list only, or no citations.

Point it at your help centre and ask it something.

Start on the free plan with your own AI key. No card needed.