Your support agent answers quickly, your RAG pipeline retrieves plausible documents, and the dashboard shows a healthy evaluation score. Then a customer asks about a policy outside the indexed knowledge base, or a user from one tenant asks a question that matches another tenant's document. The system produces a fluent answer, but the evidence is stale, unauthorized, or absent.
That's the point at which evaluating a system stops being a leaderboard exercise. A production AI system needs an evidence-bounded protocol that can expose unauthorized retrieval, unsupported claims, version drift, contamination, and unjustified confidence. A refusal can be a successful result when the system lacks authorized evidence.
Why evaluating a system means choosing what to falsify
A benchmark score is useful only when it tests a claim that matters. “The agent is accurate” is too broad to evaluate because it hides several different propositions: the retriever returns the right evidence, the model uses that evidence faithfully, the answer respects tenant permissions, and the agent refuses when the evidence is insufficient.
A falsification-first protocol starts by writing the claim and the condition that would disprove it.
Practical rule: Every metric should answer a question that could make you reject the current design.
For a customer-support RAG system, the dangerous claims might be:
- Grounding: The answer's material claims are supported by retrieved evidence.
- Isolation: A retrieval request never exposes another tenant's documents.
- Scope control: The system doesn't answer confidently when the query is outside the authorized knowledge base.
- Freshness: The index reflects the document version the product promises to support.
- Operational reliability: The complete workflow behaves acceptably under realistic concurrency, tool failures, and latency conditions.
Each claim needs a different test. A response-quality score can't prove tenant isolation. Retrieval recall can't prove that the final answer omits unsupported assertions. A low refusal rate isn't automatically good, because a system that never abstains may just be willing to guess.
The history of benchmarking supports this distinction. Early computer benchmarking efforts trace back to 1962, when Auerbach Corporation's Standard EDP Reports were used to compare computer speed. By 1965, Joslin had framed evaluation around the time required to process a user's workload, shifting attention toward workload modeling rather than raw specifications. Benchmarking became more systematic in the 1970s, and Xerox refined competitive benchmarking by 1976 as a structured comparison of processes, not merely products, as described in this history of benchmarking and evaluation.
Speech recognition followed a similar path. A 1987 DARPA workshop compared rule-based and automatically trainable methods on a defined corpus, using word error rate as a common metric. The value came from fixed training and test corpora, shared measures, and repeatable evaluation conditions, not from the metric in isolation, as documented in this history of speech recognition evaluation.
The practical consequence is simple: a clean score means little until you identify what could overturn it. For a useful overview of how accuracy claims can hide scope and measurement choices, see accuracy of artificial intelligence. A result is evidence only within the protocol that produced it.
Defining scope, intended use, and operational profile first
The first evaluation artifact shouldn't be a benchmark file. It should be a scope document that states what the system is allowed to do, for whom, with which data, and under which operating conditions.
Start with the intended users and workflows. A customer-support agent might answer questions about account configuration and published product policies, while escalating billing changes and account actions to a human or an authorized tool. A marketplace assistant might retrieve seller-specific inventory but must never treat a similar document from another seller as usable context.
Write the operational profile in concrete terms:
- Users and roles: Identify tenant, account, role, region, and permission class.
- Query distribution: Describe common request types, expected query length, languages, multi-turn behavior, and tool-use patterns.
- Environment: Record model, retrieval service, vector database, index state, prompt configuration, tools, network conditions, and deployment mode.
- Failure modes: Name outcomes the system must not produce, such as cross-tenant retrieval, stale policy answers, fabricated citations, unauthorized actions, or confident responses to unsupported questions.
- Service constraints: Define latency, availability, escalation, and cost boundaries qualitatively if the exact threshold is still under review.
This document determines which metrics deserve engineering effort. If the product doesn't handle multi-hop questions in scope, multi-hop recall is a distraction. If the system serves several tenants with separate permissions, authorization failures matter more than a generic helpfulness score. If the agent's primary job is to initiate a controlled workflow, final task state may matter more than prose similarity.
The National Academies guidance on operational software testing warns that testing should reflect intended use, operational profiles, environments, user types, and failure modes. It also identifies a recurring weakness in representative-only microbenchmarks: they can miss concurrency, shared-resource contention, and workload interactions that shape production behavior.
Write the falsifier beside each requirement
A requirement becomes testable when its rejection condition is explicit:
- “Retrieval is tenant-safe” becomes “a cross-tenant probe returns no unauthorized evidence and produces no answer based on it.”
- “Answers are grounded” becomes “every material claim maps to an authorized source span, or the system abstains.”
- “The agent handles unsupported questions” becomes “deliberately missing or contradictory context leads to an evidence-limited response.”
- “The index is current” becomes “a query about a known document revision returns the expected revision under the declared freshness contract.”
Sign the scope document before selecting a benchmark. Otherwise, the benchmark will define the product's quality standard, usually around whatever is easiest to score.
Choosing protocols and denominators before picking a benchmark
Different protocols answer different questions. A production traffic mirror can show whether the system encounters real phrasing and workflow variation. A synthetic suite can target rare failure hypotheses. A frozen held-out set can show whether a change beats a stable baseline. None of these protocols proves everything.
The denominator must be explicit before the run. For retrieval, it might be unique questions, relevant document judgments, unique tenants, or authorized document-query pairs. For answer validation, it might be answerable questions, unanswerable questions, or claims extracted from responses. For tool-use evaluation, it might be workflows, tool decisions, calls, or completed business outcomes.
Without that definition, a team can report a score that counts repeated questions, duplicate documents, or multiple claims from the same response as independent evidence. A result may look precise while covering a narrow slice of the system.
| Protocol | Denominator | What it can falsify | Where it misleads |
|---|---|---|---|
| Production workload mirror | Unique sampled queries, tenants, workflows, and observed documents | Whether behavior resembles real traffic and whether known failures recur under realistic distributions | Live traffic can contain incomplete labels, changing permissions, and uncontrolled system versions |
| Synthetic stress suite | Unique generated scenarios and targeted failure probes | Whether the system responds to specified edge cases, missing evidence, corrupted context, or adversarial authorization paths | Synthetic data can overrepresent the author's hypotheses and cannot establish production frequency |
| Frozen held-out suite | Unique frozen items, labels, tenants, and document snapshots | Whether a pinned change alters behavior on a stable test population | A frozen set can become stale and may not represent new workflows or data distributions |
Coverage and representativeness are separate properties. A suite can cover many failure categories while still failing to resemble live traffic. A workload mirror can be representative while missing rare but severe cross-tenant probes. Report both, and state which claims each protocol is allowed to support.
A useful comparison of practical test design for agents appears in AI agents testing. The important decision isn't whether to use one protocol exclusively. It's how to combine protocols without treating one denominator as evidence for another.
For example, use a frozen suite to compare a retriever release, a targeted authorization suite to challenge isolation, and a production sample to discover failures the team didn't anticipate. Keep their results separate. Combining them into one headline number erases the distinction between known coverage and real-world representativeness.
Running frozen evaluations with version pinning and contamination checks
A frozen run is defensible only when another engineer can identify what changed, what stayed fixed, and which raw evidence supports the result. Treat the protocol as a falsification exercise: each control should test a specific alternative explanation for an observed improvement.
Pin every dependency that can change behavior
Record the model identifier, checkpoint hash, tokenizer revision, retrieval index snapshot, embedding version, prompt templates, tool schemas, application commit, evaluator version, decoding parameters, and runtime configuration. Save these values in one immutable configuration reference before execution.
A model alias that resolves to a different checkpoint is a different system. A rebuilt vector index changes retrieval conditions. A revised system prompt can change behavior even when application code is unchanged. If any of these vary between candidates, the score difference cannot be attributed reliably to the intended release.
Freeze the denominator
Pre-register the complete query set, document set, tenant set, labels, expected refusal cases, and item metadata. Hash the dataset snapshot and retain the item list. Do not resample, deduplicate, rebalance, or remove inconvenient examples after inspecting results. If the protocol changes, create a new named version and document why.
Store denominators for every meaningful slice as well. Record which items belong to each tenant, document class, query type, language, and authorization condition. A slice can reveal a regression hidden by the aggregate result, while a stable item count prevents a population shift from appearing as an improvement.
Abstention and refusal cases belong in the denominator. A system that declines unsupported requests may be behaving correctly, so measure that outcome explicitly rather than treating every non-answer as failure.
Check for contamination
Training-data contamination can inflate AI benchmark scores by an estimated 6–40%, according to a review of benchmarking flaws and AI evaluation risks (benchmarking flaws and contamination review). That range describes the review's benchmark discussion, not this deployment. Report the checks performed and keep the cited estimate separate from local results.
Compare evaluation items with training, fine-tuning, and retrieval corpora using n-gram overlap, embedding similarity, and blocklists for known contamination sets. For retrieval-augmented systems, inspect whether an evaluation document or near-duplicate entered the index through development, ingestion, or prompt-construction paths. Remove contaminated items from the held-out claim, or label the protocol so its scope is clear.
Emit reproducibility artifacts
Save the configuration bundle, environment file, dataset hash, raw responses, retrieved passages, authorization decisions, judge prompts, evaluator outputs, random seeds, logs, and failure traces. A rerun should use the same configuration and expose only the remaining variation within a declared tolerance.

A knowledge base is part of the evaluation surface, not merely an implementation dependency. Teams can use AI knowledge base design guidance to make ingestion, retrieval, authority, and evidence handling visible in the test plan.
Version pinning tests whether a release comparison is attributable. Denominator locking tests whether the population stayed fixed. Contamination checks test whether memorization could explain the score. Reproducibility artifacts make each conclusion auditable.
Testing authorization, provenance, and fail-closed grounding
A user asks for a document from another tenant, the retriever finds a highly relevant passage, and the generator produces a polished answer. That result is a security failure, even if the answer is factually correct. Evaluate grounding as a contract across authorization, retrieval, evidence delivery, generation, and output validation.
Authorization must be a retrieval predicate. Pass the tenant discriminator and user-level access filters into every retrieval operation. Do not retrieve broadly and filter after generation. Microsoft's secure multitenant RAG architecture guidance treats tenant and user authorization constraints as part of each retrieval request.
A vendor-neutral multitenant retrieval and tool-use study illustrates the failure mode. Its probes found that ungated retrieval leaked cross-tenant data in 98–100% of probes, while relevance-only ranking did not provide isolation (multitenant retrieval authorization study). That finding applies to the study's setup and probes, not every deployment. It does show why authorization must be evaluated before relevance.
Test the evidence path, not only the final answer
For every retrieved span, record a stable document identifier, revision, tenant, access decision, and location. Send the generator only evidence that has passed authorization. Then test whether each material claim maps back to an allowed span.
Provenance-sensitive agent research defines provenance as the origin or source of information and argues that agents should rely exclusively on authorized evidence. A relevant item from an untrusted source should be disregarded, making refusal a valid fail-closed outcome (provenance and AI evaluation).
| Probe | Setup | Expected behavior | Failure signal |
|---|---|---|---|
| Cross-tenant query | Use a user from one tenant and a query answered only by another tenant's documents | Zero unauthorized hits, followed by refusal or an evidence-limited response | Another tenant's span appears in retrieval, citations, traces, or the answer |
| Role-restricted document | Keep the tenant constant but remove the user's document-class permission | Restricted content stays out of context | The system cites, paraphrases, or reveals restricted content |
| Missing evidence | Ask about information absent from authorized context | The system states that available evidence is insufficient | A confident unsupported answer appears |
| Corrupted context | Replace a relevant passage with contradictory or off-topic text | The system detects the mismatch or abstains | The model follows fluent but unsupported context |
| Provenance break | Remove a source identifier from one retrieved span | The claim is omitted or the answer refuses | A material claim survives without traceable evidence |
Track refusal rate, citation coverage, source agreement, unauthorized-hit count, and refusal-trigger logs. Refusal rate is not a quality score by itself. A system that refuses every request minimizes exposure but provides little utility. The target is narrower: answer supported questions, refuse unsupported ones, and expose the evidence path for review.
Teams that need broader context on integrity controls can consult AI Frontiers and integrity controls. Authorization and provenance must remain observable at the point where evidence enters the model context.
Reading limitations, variance, and what a claim covers
A score is meaningful only with the conditions that produced it. A defensible report records the protocol version, exact denominator, system and evaluator versions, run conditions, and uncertainty. It also identifies the slices that drove the result.
Report performance across dimensions that can change behavior:
- Language and locale
- Tenant and role class
- Document type and revision
- Query length and interaction depth
- Answerable versus unanswerable requests
- Retrieval condition and latency band
- Tool path and failure state
Reporting averages without variance or significance is a recognized benchmarking flaw. A review of benchmarking flaws, cited in the frozen-evaluation section, found that 12 of 29 papers, or 41%, incorrectly used the arithmetic mean to average overhead numbers. The example concerns systems-security papers, but the warning applies broadly. Outliers, skew, and mixed workloads can make an average a poor description of the observed distribution.
Separate the lever that changed
A release report should separate changes to the model, prompt, retriever, index, authorization layer, tool implementation, and evaluator. Change one major lever at a time where practical, or run an ablation that removes the suspected contribution.
A new model may improve answer quality while retrieval recall falls. The aggregate score can hide a failure that appears after the next model change. A prompt rewrite may reduce unsupported claims by increasing abstentions. That can improve safety while reducing task completion. Slice-level results expose the trade-off.
Microsoft's AI evaluation reporting guidance, linked in the authorization section, recommends disclosing what was measured, the exact system and version, configuration and safeguards, protocol, evaluator access, uncertainty, limitations, evaluation date, and who conducted, funded, and controlled publication. Treat that disclosure as part of the result. Without it, reviewers cannot reproduce or properly bound the claim.
State what the protocol did not test, including distribution shifts, adversarial inputs, unsupported languages, new document classes, production-scale contention, and workflows outside the declared scope. The width of the conclusion must match the width of the evidence. “This frozen suite shows improved citation coverage for the tested tenant and document slices” is reviewable. “The system is reliable” is not.
A release checklist and what to do when evaluation surfaces a failure
A release candidate should carry an evidence bundle, not just a green dashboard. The bundle needs a locked scope document, frozen benchmark hash, denominator file, contamination audit, authorization and abstention results, version manifest, raw traces, and a limitations paragraph.
Use this checklist before each release:
- Scope lock: Attach the signed intended-use and failure-mode document.
- Frozen benchmark hash: Record the exact dataset and document snapshot.
- Denominator file: Store every item, tenant, document, label, and per-item result.
- Contamination audit: Save overlap methods, blocklists, findings, and exclusions.
- Authorization and abstention tests: Attach cross-tenant, restricted-role, missing-evidence, and corrupted-context outcomes.
- Limitations paragraph: State what the protocol doesn't cover and which version produced the result.

Three failures deserve a defined response rather than quiet rescoring.
Denominator drift after a retriever swap
If a retriever change alters the eligible document population, stop the comparison. Re-run the original frozen set with the old and new retrievers, preserve the original denominator, and create a separate workload analysis for the new index. Roll back when the release can't explain changed coverage through the declared retrieval change. Tell stakeholders that the comparison is invalid until the population and index state are reconciled.
Authorization regression after a permissions refactor
Freeze affected tenants, roles, and document classes. Replay cross-tenant and restricted-role probes against the prior and candidate builds, then inspect raw retrieval results, authorization decisions, context payloads, and final responses. Treat any unauthorized hit as a release blocker. Communicate the affected boundary, the test evidence, the containment action, and the rerun condition without describing the system as generally secure.
Abstention collapse after a prompt rewrite
Re-run answerable and unanswerable cases separately, including missing, contradictory, and off-topic context. Compare refusal triggers, citation coverage, source agreement, and task completion rather than only the overall answer score. Roll back if the rewrite causes unsupported answers to replace evidence-limited responses, then publish the narrowed behavior change and the remaining limitation.
Evaluation should run with release branches, index changes, permission changes, and prompt changes. It also needs production feedback, because new traces expose failure modes that a frozen suite can't anticipate. Add confirmed failures to a future protocol only after preserving the old result, documenting the denominator change, and explaining which claim the new item is meant to falsify.
AletheionAGI offers a grounding and evidence-control layer that can work with existing LLMs, RAG pipelines, vector databases, and memory systems, with emphasis on authorized retrieval, provenance, fail-closed behavior, and repeatable evaluation artifacts. If your team needs to test these controls against its own workflows, visit AletheionAGI to review the available infrastructure and evaluation approach.



