Fixed scope · Fixed price · 10 working days
The RAG Audit: replace guesswork with a measured diagnosis.
Your AI answers are “mostly right,” nobody can say how right, and every proposed fix is a guess. In ten working days that becomes a number, a named list of failure modes, and a prioritized plan — with the evaluation harness left behind so quality never regresses invisibly again.
Deliverables
- A golden evaluation set
- 25–50 real queries from your users with verified expected answers and sources. This alone changes how your team works.
- Your accuracy, measured
- Retrieval hit-rate and answer-level correctness — citations checked, fabrications flagged — replacing impressions with a number.
- Failure-mode diagnosis
- Each miss classified with evidence: chunking, retrieval ranking, reranker behavior, grounding, or missing data.
- A prioritized fix plan
- Ranked by expected accuracy gain against effort, specific to your stack, and executable by your own team.
- The evaluation harness, installed
- A runnable script and CI hook so every future change is measured against the golden set.
- A findings walkthrough
- A 60-minute session with your team: findings explained plainly, questions answered.
The ten days
- Days 1–2Access and query harvesting: real failing and passing queries collected from your logs and users.
- Days 3–5Golden set built and verified with your domain expert; baseline accuracy measured.
- Days 6–8Failure-mode analysis: every miss traced through your pipeline to its root cause.
- Days 9–10Fix plan written, harness handed over, walkthrough delivered.
Why the diagnosis is fast
I run this exact measure-then-fix loop continuously on EAKC, my live production platform: golden sets, retrieval-level and answer-level evaluation, reranker behavior analysis, and fabricated-citation guards. The failure modes I will find in your system are ones I have already encountered, fixed, and documented in mine — which is why this takes ten days rather than six weeks.
Common questions
- Which stacks do you cover?
- Any retrieval pipeline: LangChain, LlamaIndex, or custom; pgvector, Pinecone, Weaviate, or Qdrant; any model provider. If it retrieves and generates, it can be audited.
- What access do you need?
- Read access to the pipeline code and the ability to run queries — a staging environment is sufficient. Query logs are valuable. An NDA is no problem, and I can work inside your environment.
- What if our accuracy turns out to be acceptable?
- Then you receive proof of that, the harness that keeps it true, and a diagnosis of the residual misses. In practice, "we do not actually know" is the expensive state to remain in.
- Can you also implement the fixes?
- Yes — that is the Rescue engagement, and the audit fee is credited in full if you proceed within 30 days.
A note on the price. Comparable audits from agencies run $12,000–28,000. $2,500 is deliberate introductory pricing while I build public case studies — the price will rise as they land; the scope will not. Early clients keep the economics.