The knowledge corpus
Vesper's answers are only as good as what it retrieves. This page describes what is in the corpus, how it gets there, and what is deliberately kept out.
The shortest version, because it is the part that matters most: the corpus is built entirely from public Wazuh community discussions. No customer data is ever ingested into it.
What is in the corpus
The corpus holds resolved community support conversations, from the places where Wazuh support actually happens:
- the community Slack workspace,
- the community Discord server,
- the Google Groups mailing list,
- GitHub issues.
Conversations are rebuilt into complete threads, the original question plus every reply, so an answer is grounded in the whole exchange rather than an out-of-context fragment. A thread that never received an answer is excluded, because it carries a question and nothing to ground an answer in.
The troubleshooting procedures Vesper's team writes and maintains are not part of this corpus. They are private material the agent follows when it works on a registered environment, and an answer never quotes or cites one. See written procedures.
What never reaches the corpus
No customer data, of any kind. Questions asked in Vesper, agent conversations, environments, alerts and anything the connector reads from a Wazuh deployment stay in the customer's own tenant. None of it is ingested, embedded, or used to answer anyone else. Ingestion has exactly one input, the public community discussions described above, and a customer's tenant is not it.
Personal identities, removed at normalization time and before anything is stored. Author names, senders, assignees and every per-platform user identifier are on the strip list. Email addresses and IPv4 addresses are replaced inside message bodies with placeholders.
Provenance, which surprises people. The channel, the platform, the message and thread identifiers and every permalink are dropped too, including the GitHub issue URL. A stored thread cannot be traced back to a person, and it cannot be traced back to the post it came from either. What survives is its title, the question, the replies and their timestamps.
The original raw documents stay in the source system and are never copied into the retrieval corpus.
How a community thread is processed
- Extraction. Raw parent and comment documents are pulled from the ticket store and grouped into threads by their parent and child identifiers.
- Normalization and redaction. Identifying fields are stripped, as described above.
- Embedding. Each normalized thread is embedded into a vector by the embedding model.
- Indexing. Vectors and thread metadata are upserted into the retrieval index.
The same embedding model is used for the corpus and for incoming questions, which is what makes the similarity scores in the sources panel meaningful.
The official documentation
An answer can also quote the official Wazuh documentation at documentation.wazuh.com, and that material is not in the corpus at all.
The page is fetched on Vesper's own infrastructure, live, at the moment the question is answered. Nothing is downloaded to or requested by a machine the customer runs, and nothing is stored on Vesper's side either, so what an answer quotes is the page as published at that moment rather than a copy that can go stale. A documentation page an answer used is cited by its URL, so the claim can be checked against the page itself.
Vesper reads from a fixed list of pages, and that list changes only when Vesper ships a new version. Nothing in a question can send it to an address that is not on the list.
Searching it directly
The Answers page searches the corpus without writing an answer. It ranks the material against your words and shows what came back: the title, the opening of the thread, and the similarity score. That is enough to tell whether the corpus has anything about your problem before you spend anything having it read for you.
The opening is what the person originally asked. It is cut to a couple of hundred characters and flattened to a single paragraph.
What a row still does not carry is everything ingestion stripped. There is no author, no channel and no date of the original post, for the reasons above. Nor is there a link, except on a documentation row, whose page is public and whose URL was never provenance to strip. The excerpt exists because the thread body is already in hand when the row is ranked, not because provenance was kept.
Why corpus sources carry no links
A corpus source in the panel is a title and a similarity score, never a link, and that is deliberate on two levels. A documentation citation is the one exception, because it names a public page rather than something ingestion stripped.
Nothing to link to survives ingestion. Permalinks are stripped along with the rest of the provenance, so for most rows there is no correct URL to offer.
An invented link is worse than no link. A fabricated title is a dead end that gets read. A fabricated URL is one that gets clicked, on a domain that looks official, and it is the most confident-looking part of an answer. So when the generated answer closes with a list of further sources, that list is validated against what was actually retrieved. A bullet naming something that was not retrieved is dropped, and a link keeps its label and loses its target unless the URL is a documentation page this answer actually read, or genuinely appears in a retrieved title.
Freshness
The corpus is refreshed by re-running the ingestion pipeline against the source ticket store. Ingestion is incremental, with an overlap window so late replies to older threads are not missed, and pruning keeps answerless threads out of retrieval.
Refreshes are triggered by the Wazuh team rather than running on a fixed schedule, so the corpus reflects community discussion up to the last run rather than up to the minute.