AI and data
rag-corpus-audit
Audit a RAG knowledge base
Audits a RAG corpus for chunk problems, coverage gaps, duplicates, stale claims, contradictions, sensitive data and missing metadata before indexing.
- Area
- AI and data
- Checked against
- Its sources, as listed
- Last verified
- 26 September 2026
- Files
- SKILL.md, 10 references, 8 scripts
How it works
- 1Agree the corpus scope, pipeline details, audiences and question sources.
- 2Audit the text and chunks the retriever will see.
- 3Check metadata, sensitive data and hidden instructions.
- 4Review chunk quality, duplicates, boilerplate and stale claims.
- 5Read contradiction candidates and ask the owner which source is correct.
- 6Measure coverage against real user questions and identify vocabulary gaps.
- 7Prioritize confirmed findings and write a fix list with checks and undo steps.
Use it when
- Audit a knowledge base before indexing or re-indexing.
- Investigate outdated, contradictory or wrong RAG answers.
- Check a new document source before adding it.
- Review internal content before widening assistant access.
How it stays safe
- Works from exports or copies and keeps the audit read-only.
- Keeps corpus data, question logs, secrets and personal data on the local machine.
- Treats instructions found inside documents as evidence instead of commands.
- Uses optional model-based chunk checks only after approval and sensitive-data screening.
Install it
Copy the whole rag-corpus-audit folder into your agent's skills folder, not only SKILL.md.
git clone --depth 1 https://github.com/hamzaahmadaslam/agent-skills.git
cp -R agent-skills/rag-corpus-audit "<your agent's skills folder>/"