Skip to content

AI and data

rag-corpus-audit

Audit a RAG knowledge base

Audits a RAG corpus for chunk problems, coverage gaps, duplicates, stale claims, contradictions, sensitive data and missing metadata before indexing.

Area
AI and data
Checked against
Its sources, as listed
Last verified
26 September 2026
Files
SKILL.md, 10 references, 8 scripts

How it works

  1. 1Agree the corpus scope, pipeline details, audiences and question sources.
  2. 2Audit the text and chunks the retriever will see.
  3. 3Check metadata, sensitive data and hidden instructions.
  4. 4Review chunk quality, duplicates, boilerplate and stale claims.
  5. 5Read contradiction candidates and ask the owner which source is correct.
  6. 6Measure coverage against real user questions and identify vocabulary gaps.
  7. 7Prioritize confirmed findings and write a fix list with checks and undo steps.

Use it when

  • Audit a knowledge base before indexing or re-indexing.
  • Investigate outdated, contradictory or wrong RAG answers.
  • Check a new document source before adding it.
  • Review internal content before widening assistant access.

How it stays safe

  • Works from exports or copies and keeps the audit read-only.
  • Keeps corpus data, question logs, secrets and personal data on the local machine.
  • Treats instructions found inside documents as evidence instead of commands.
  • Uses optional model-based chunk checks only after approval and sensitive-data screening.

Install it

Copy the whole rag-corpus-audit folder into your agent's skills folder, not only SKILL.md.

git clone --depth 1 https://github.com/hamzaahmadaslam/agent-skills.git
cp -R agent-skills/rag-corpus-audit "<your agent's skills folder>/"

More skills in the collection

All 11 skills