Homeβ€Ί Blogβ€Ί AWS Architecture Series #66 β€” The judge has no answer key unless you write one…
AWS Architecture AWS Architecture Series

AWS Architecture Series #66 β€” The judge has no answer key unless you write one

Evaluation gets skipped because it looks like infrastructure work that can wait. Bedrock makes the job easy to launch, which hides the real cost: the metrics that sound objective will run without an answer key and still return a number, and that number is one model's opinion of another's.

Verified against current vendor documentation on 28 September 2026. Pricing, limits and API behaviour were checked against the official docs on that date. Cloud services change fast — if you are reading this much later, treat the specifics as a starting point and re-check the linked sources.

Business Challenge

#61 ended on a measurement: how often was the answer present in what came back. This post is about why that measurement rarely gets built, and what happens when a managed service makes it easy to skip.

Evaluation is not skipped because it is hard to run. Bedrock evaluations launches from a console page. It is skipped because the expensive part is not the job — it is the answer key, and nothing in the workflow forces you to write one.

1Correctness will score without knowing what is correct

The built-in metric is called Builtin.Correctness, and its description carries a conditional that decides everything: “Measures if the model's response to the prompt is correct. Note that if you supply a reference response (ground truth) as part of your prompt dataset, the evaluator model considers this when scoring the response.”

If. Supply reference answers and you are measuring agreement with your answers. Supply none and the job still runs, still produces a histogram, and is measuring whether one model finds another model's answer plausible.

Those are different quantities with the same metric name and the same score range. Nothing in the report distinguishes them.

Fix

Record whether ground truth was supplied next to every Correctness score, or the number is uninterpretable later.

2Nine of the eleven metrics have no answer key by construction

Eleven built-in metrics ship: Correctness, Completeness, Faithfulness, Helpfulness, Logical coherence, Relevance, Following instructions, Professional style and tone, Harmfulness, Stereotyping, Refusal.

Only two mention ground truth — Correctness and Completeness — and for both it is conditional. The rest are judgements the evaluator model makes on its own. Helpfulness is explicitly a bundle of them: “whether the response is sensible and coherent, and whether the response anticipates implicit needs and expectations.”

That is not a flaw. Helpfulness has no ground truth available to anyone. But it means most of the scorecard is one model's taste, and taste is not stable across judges.

Fix

Separate the metrics that could have an answer key from the ones that cannot. Only the first kind can regress.

3A score without its judge named is not a result

“This kind of model evaluation requires two different models, a generator model and an evaluator model… the evaluator model scores the responses to those prompts based on metrics you select.”

The evaluator is chosen from a list — Nova Pro through Claude Opus 4.8, Llama 3.1 70B, Mistral Large. And the lists are not one list: the set supported for built-in metrics differs from the set supported for custom metrics, with Mistral Large 24.07 and Llama 3.3 70B appearing only in the custom one.

So changing a metric can force a change of judge, which changes the scores, for the same responses. “We scored 0.82” is not a result until it says which model produced it.

Fix

Pin the judge model alongside the metric, and treat a judge change like a schema change.

Architecture

Bedrock offers three kinds of evaluation and they answer different questions.

Diagram: what Amazon Bedrock evaluations actually measures, and where an answer key is required. Three kinds of evaluation are shown. Programmatic evaluation runs a model against a custom or built-in prompt dataset and produces computed scores. Human evaluation uses a team of workers, who can be employees or subject-matter experts, providing ratings and preferences. Judge-model evaluation requires two models, a generator whose responses are being scored and an evaluator that scores them and provides an explanation for each response. The eleven built-in judge metrics are listed and split by whether an answer key is possible. Correctness and Completeness are the only two that mention ground truth, and for both it is conditional: the evaluator considers a reference response only if you supply one, so the job runs and returns a score either way, measuring agreement with your answers when supplied and plausibility to another model when not. Faithfulness is checkable against the prompt itself, identifying whether the response contains information not found in the prompt. The remaining metrics are judgements with no external key: Helpfulness, Logical coherence, Relevance, Following instructions, Professional style and tone, Harmfulness, Stereotyping and Refusal. A panel records that the evaluator model is chosen from a list, that the list for built-in metrics differs from the list for custom metrics with Mistral Large 24.07 and Llama 3.3 70B appearing only in the custom list, and that changing a metric can therefore force a change of judge and change the scores for identical responses. A further panel covers RAG evaluation, which comes in retrieve-only and retrieve-and-generate forms, either against an Amazon Bedrock Knowledge Base or against your own inference response data from an external source, and which is the one place ground truth is stated as required, the dataset having to include the expected retrieved texts and responses. A closing panel notes that the console shows a histogram plus explanations for only the first five prompts in the dataset, with the full report written to an S3 bucket you specify, so the console view is a sample and the evidence is the file.
Two metrics can have an answer key. One checks against the prompt. The other eight are the judge's opinion.

Three kinds, and what each is evidence of

Programmatic jobs “allow you to quickly evaluate a model's ability to perform a task” against “your own custom prompt dataset… or… an available built-in dataset.”

Human jobs bring “a team of people who provide their ratings and preferences in relation to certain metrics” — “employees of your company or a group of subject-matter experts from your industry.” This is the only one that produces an opinion from outside the system being measured, and it is the one that does not scale.

Judge jobs scale and are the reason evaluation is now practical at all. The tradeoff is what the first three cards describe: the thing doing the scoring is the same class of system as the thing being scored.

The metric worth trusting most is the one with the least ambitious name

Builtin.Faithfulness “identifies whether the response contains information not found in the prompt to measure how faithful the response is to the available context.” No ground truth needed and none possible to omit — the context is right there in the prompt, and the question is whether the response stayed inside it. For a RAG application that is the hallucination check, and unlike Correctness it cannot quietly degrade into a plausibility vote, because the reference material is part of the input by definition.

RAG evaluation is the one that demands an answer key

Two shapes: “Retrieve only – the report is based on the data retrieved from your RAG source” and “Retrieve and generate – the report is based on the data retrieved from your knowledge base and the summaries generated by the response generator model.”

And the requirement is stated rather than conditional: “The dataset must also include 'ground truth' or the expected retrieved texts and responses for the queries so that the evaluation can check if your knowledge base is aligned with what's expected.”

Read that carefully, because it is the whole cost of this post's subject. To evaluate retrieval you must first write down, by hand, which passages should come back for a set of real questions. That is not a configuration step. It is a person reading the corpus. Everything else here is a job you launch; this is work, and it is why the retrieval measurement #61 recommended so rarely exists.

The payoff is the comparison it unlocks: results “allow you to compare different Amazon Bedrock Knowledge Bases and other RAG sources, and then to choose the best” — which is exactly the decision #63 left open, and it cannot be settled without the answer key.

Why This Architecture Holds Up

You can evaluate something that is not a Bedrock model at all

“If you provide your own response data, Amazon Bedrock skips the model invoke step and directly evaluates the data you supply.” The same applies on the RAG side: you can “bring your own inference response data from an external RAG source.”

That makes this a scoring service rather than a Bedrock feature, and it is the single most useful thing here. A self-hosted model, a vendor API, last quarter's outputs — all of them can be scored on the same metrics by the same judge, which is the only way a comparison across them means anything.

It also removes the excuse. Evaluating a non-Bedrock system does not require migrating it first.

The console shows five explanations, and people stop there

“The metrics summary card in the console displays a histogram… and explanations of the score for the first five prompts found in your dataset. The full evaluation job report is available in the Amazon S3 bucket you specify.” Five, and they are the first five rather than a sample — so whatever happens to sit at the top of the file is what gets read. The histogram is the real summary and the S3 report is the real evidence. A review that consists of glancing at the console has seen an unrepresentative handful and a shape.

The generator list is wider than the console admits

Evaluation covers “Foundation models… Marketplace models… Customized foundation models… Imported foundation models… Prompt routers… Models that you have purchased Provisioned Throughput.”

Two of those close loops from earlier posts. Imported foundation models means a model brought in through the one-way door in #65 can still be scored here. Prompt routers means you can evaluate the router's behaviour rather than a single model's, which is the only honest way to measure a system that picks models dynamically.

And one carve-out to know before designing a workflow around the console: generator models invoked through the OpenAI Responses API are usable “only through the AWS CLI and the Amazon Bedrock API. The Amazon Bedrock console doesn't support selecting these models.”

Custom metrics are the answer to the subjectivity problem, not a way around it

“Each metric uses a different prompt for the evaluator model. You can also define your own custom metrics for your particular business case.”

That sentence is the useful one. A built-in metric is a prompt AWS wrote; a custom metric is a prompt you wrote. Neither is more objective than the other — but yours can encode what your business actually means by a good answer, which “Helpfulness” cannot. If the built-in metrics feel vague, the fix is to write the rubric down, not to hope a general-purpose one matches.

Key Architecture Decisions

Decision Choice Reasoning
Ground truth Supply it, or rename the metric Correctness without reference answers scores plausibility, under the same name.
Metric selection Split keyed from unkeyed Only the keyed ones can meaningfully regress between runs.
Hallucination check Faithfulness Checked against the prompt's own context, so it cannot degrade into a vote.
Judge model Pinned, and recorded with the score Built-in and custom metrics support different judge lists; the judge moves the number.
Vague requirements Custom metric, not a built-in A custom metric is a rubric you wrote; a built-in is a prompt AWS wrote.
Human evaluation A calibration sample, not the pipeline It is the only outside opinion, and the only one that does not scale.
RAG quality Retrieve-only first It isolates the ceiling #61 describes, before generation can mask it.
Reading results The S3 report, not the console The console explains the first five prompts, chosen by file order.

What makes a rubric worth running

Three properties, and a built-in metric has at most two of them. It has to be stable — the same response scores the same next month, which means pinning the judge. It has to be discriminating — if every response scores highly the metric is measuring nothing, which is what the histogram is for and why it beats an average. And it has to be actionable — a low score has to imply a change somebody can make.

Helpfulness fails the third routinely. Faithfulness passes all three. Correctness passes all three only if you wrote the answer key, which returns to the same place: the evaluation is cheap and the key is not.

Closing Thought

Every post in this block has ended in the same place. Chunking decides what can be found, and you have to measure retrieval to know. The store decides how often you can ask, and you have to measure queries per second to choose. Guardrails decide what gets filtered, and you have to test what the filter saw. All three need a measurement nobody is required to take.

Bedrock evaluations removes almost every obstacle to taking them. It scores non-Bedrock models, it handles the judging, it writes the report. What it cannot do is decide what a good answer is for your business, and it does not stop you running a metric named Correctness without ever saying what correct means.

So the part everyone skips is not the evaluation job. It is the hour with the corpus, writing down what should have come back. Everything downstream of that hour is automatable, and nothing upstream of it is worth much without it.

Next in this series

AI & ML — Comprehend, Textract and Rekognition: the task-shaped services, and when a purpose-built API beats a foundation model on the same job.

Comments

How was your experience?
Your feedback helps improve this site.
PoorExcellent