Define the test set
List representative questions, expected evidence, and the source documents that should support each answer. Keep private documents out of the public page.
SaveKit use case
RAG systems are easier to discuss when the evaluation criteria are visible. A simple static page can document the retrieval question, expected evidence, answer quality, and known failure modes without requiring a documentation platform.
A good fit for: AI engineering teams, internal evaluation plans, retrieval demos, technical workshops, and model quality reviews.
List representative questions, expected evidence, and the source documents that should support each answer. Keep private documents out of the public page.
Record retrieval coverage, citation quality, latency, and failure examples. Avoid claiming that a small sample proves production reliability.
Publish the checklist as HTML or a ZIP build, then ask reviewers to comment on missing cases, unclear pass criteria, and what should be measured next.
Before you publish
These checks prevent the most common surprises after a folder becomes a public URL.
Common questions
SaveKit is a static host. Use it for the evaluation plan, result summary, or safe interactive mockup; keep live retrieval behind a secured service.
Include representative questions, expected evidence, retrieval quality, groundedness, citation correctness, latency, and a place to record failure cases.
Only publish numbers you can reproduce and explain. Include the dataset scope, evaluation method, date, and caveats.
Upload the HTML or ZIP build, preview it, and get a link you can send to the next person who needs to see it.
Start publishing