Experiment Design Reviewer
NewReviews an experiment for causal validity, measurement quality, and interpretation risks.
This prompt has no customizable variables — it's ready to use as-is.
Act as an information-retrieval evaluation specialist.
Objective:
Measure and diagnose retrieval quality for the supplied RAG or search system.
Inputs:
- evaluation queries: {{EVALUATION_QUERIES}}
- retrieved results: {{RETRIEVED_RESULTS}}
- ground-truth evidence: {{GROUND_TRUTH_EVIDENCE}}
- metadata: {{METADATA}}
- retrieval configuration: {{RETRIEVAL_CONFIGURATION}}
Process:
1. Define relevance criteria and evaluation slices.
2. Measure top-k recall, precision, MRR or NDCG where appropriate.
3. Cluster failures into query understanding, chunking, metadata, indexing, or ranking problems.
4. Compare lexical, semantic, hybrid, and reranking options.
5. Recommend the smallest changes likely to improve retrieval.
Required output:
- Metrics summary
- Failure clusters
- Representative bad cases
- Root-cause analysis
- Prioritized improvements
- Re-test plan
Guardrails:
- Do not use generation quality as a proxy for retrieval quality.
- Evaluate on realistic query distributions.
- Separate access-control failures from relevance failures.
When information is missing, state the assumption explicitly and identify what evidence would change the recommendation. Keep the response practical, specific, and implementation-oriented."Run with AI" sends your customized inputs to this site's configured AI model to generate a live sample here — nothing is saved. To keep your content on the provider's own site instead, use Copy or Open in ChatGPT.
Did this prompt give you a useful result?
Diagnose retrieval quality using query sets, relevance labels, recall, precision, ranking, failure clusters, and actionable recommendations.
Reviews an experiment for causal validity, measurement quality, and interpretation risks.
Synthesizes interviews, surveys, and notes into actionable user insights.
Reviews whether citations actually support the claims they are attached to.
Recommends a research design that fits the question, constraints, and available evidence.