Business Challenge
#61 ended on a measurement: how often was the answer present in what came back. This post is about why that measurement rarely gets built, and what happens when a managed service makes it easy to skip.
Evaluation is not skipped because it is hard to run. Bedrock evaluations launches from a console page. It is skipped because the expensive part is not the job — it is the answer key, and nothing in the workflow forces you to write one.
The built-in metric is called Builtin.Correctness, and its description
carries a conditional that decides everything:
“Measures if the model's response to the prompt is correct. Note that
if you supply a reference response (ground truth) as part of your prompt dataset,
the evaluator model considers this when scoring the response.”
If. Supply reference answers and you are measuring agreement with your answers. Supply none and the job still runs, still produces a histogram, and is measuring whether one model finds another model's answer plausible.
Those are different quantities with the same metric name and the same score range. Nothing in the report distinguishes them.
FixRecord whether ground truth was supplied next to every Correctness score, or the number is uninterpretable later.
Eleven built-in metrics ship: Correctness, Completeness, Faithfulness, Helpfulness, Logical coherence, Relevance, Following instructions, Professional style and tone, Harmfulness, Stereotyping, Refusal.
Only two mention ground truth — Correctness and Completeness — and for both it is conditional. The rest are judgements the evaluator model makes on its own. Helpfulness is explicitly a bundle of them: “whether the response is sensible and coherent, and whether the response anticipates implicit needs and expectations.”
That is not a flaw. Helpfulness has no ground truth available to anyone. But it means most of the scorecard is one model's taste, and taste is not stable across judges.
FixSeparate the metrics that could have an answer key from the ones that cannot. Only the first kind can regress.
“This kind of model evaluation requires two different models, a generator model and an evaluator model… the evaluator model scores the responses to those prompts based on metrics you select.”
The evaluator is chosen from a list — Nova Pro through Claude Opus 4.8, Llama 3.1 70B, Mistral Large. And the lists are not one list: the set supported for built-in metrics differs from the set supported for custom metrics, with Mistral Large 24.07 and Llama 3.3 70B appearing only in the custom one.
So changing a metric can force a change of judge, which changes the scores, for the same responses. “We scored 0.82” is not a result until it says which model produced it.
FixPin the judge model alongside the metric, and treat a judge change like a schema change.
Architecture
Bedrock offers three kinds of evaluation and they answer different questions.
Three kinds, and what each is evidence of
Programmatic jobs “allow you to quickly evaluate a model's ability to perform a task” against “your own custom prompt dataset… or… an available built-in dataset.”
Human jobs bring “a team of people who provide their ratings and preferences in relation to certain metrics” — “employees of your company or a group of subject-matter experts from your industry.” This is the only one that produces an opinion from outside the system being measured, and it is the one that does not scale.
Judge jobs scale and are the reason evaluation is now practical at all. The tradeoff is what the first three cards describe: the thing doing the scoring is the same class of system as the thing being scored.
Builtin.Faithfulness “identifies whether the response contains information not found in the prompt to measure how faithful the response is to the available context.” No ground truth needed and none possible to omit — the context is right there in the prompt, and the question is whether the response stayed inside it. For a RAG application that is the hallucination check, and unlike Correctness it cannot quietly degrade into a plausibility vote, because the reference material is part of the input by definition.
RAG evaluation is the one that demands an answer key
Two shapes: “Retrieve only – the report is based on the data retrieved from your RAG source” and “Retrieve and generate – the report is based on the data retrieved from your knowledge base and the summaries generated by the response generator model.”
And the requirement is stated rather than conditional: “The dataset must also include 'ground truth' or the expected retrieved texts and responses for the queries so that the evaluation can check if your knowledge base is aligned with what's expected.”
Read that carefully, because it is the whole cost of this post's subject. To evaluate retrieval you must first write down, by hand, which passages should come back for a set of real questions. That is not a configuration step. It is a person reading the corpus. Everything else here is a job you launch; this is work, and it is why the retrieval measurement #61 recommended so rarely exists.
The payoff is the comparison it unlocks: results “allow you to compare different Amazon Bedrock Knowledge Bases and other RAG sources, and then to choose the best” — which is exactly the decision #63 left open, and it cannot be settled without the answer key.
Why This Architecture Holds Up
You can evaluate something that is not a Bedrock model at all
“If you provide your own response data, Amazon Bedrock skips the model invoke step and directly evaluates the data you supply.” The same applies on the RAG side: you can “bring your own inference response data from an external RAG source.”
That makes this a scoring service rather than a Bedrock feature, and it is the single most useful thing here. A self-hosted model, a vendor API, last quarter's outputs — all of them can be scored on the same metrics by the same judge, which is the only way a comparison across them means anything.
It also removes the excuse. Evaluating a non-Bedrock system does not require migrating it first.
“The metrics summary card in the console displays a histogram… and explanations of the score for the first five prompts found in your dataset. The full evaluation job report is available in the Amazon S3 bucket you specify.” Five, and they are the first five rather than a sample — so whatever happens to sit at the top of the file is what gets read. The histogram is the real summary and the S3 report is the real evidence. A review that consists of glancing at the console has seen an unrepresentative handful and a shape.
The generator list is wider than the console admits
Evaluation covers “Foundation models… Marketplace models… Customized foundation models… Imported foundation models… Prompt routers… Models that you have purchased Provisioned Throughput.”
Two of those close loops from earlier posts. Imported foundation models means a model brought in through the one-way door in #65 can still be scored here. Prompt routers means you can evaluate the router's behaviour rather than a single model's, which is the only honest way to measure a system that picks models dynamically.
And one carve-out to know before designing a workflow around the console: generator models invoked through the OpenAI Responses API are usable “only through the AWS CLI and the Amazon Bedrock API. The Amazon Bedrock console doesn't support selecting these models.”
Custom metrics are the answer to the subjectivity problem, not a way around it
“Each metric uses a different prompt for the evaluator model. You can also define your own custom metrics for your particular business case.”
That sentence is the useful one. A built-in metric is a prompt AWS wrote; a custom metric is a prompt you wrote. Neither is more objective than the other — but yours can encode what your business actually means by a good answer, which “Helpfulness” cannot. If the built-in metrics feel vague, the fix is to write the rubric down, not to hope a general-purpose one matches.
Key Architecture Decisions
| Decision | Choice | Reasoning |
|---|---|---|
| Ground truth | Supply it, or rename the metric | Correctness without reference answers scores plausibility, under the same name. |
| Metric selection | Split keyed from unkeyed | Only the keyed ones can meaningfully regress between runs. |
| Hallucination check | Faithfulness | Checked against the prompt's own context, so it cannot degrade into a vote. |
| Judge model | Pinned, and recorded with the score | Built-in and custom metrics support different judge lists; the judge moves the number. |
| Vague requirements | Custom metric, not a built-in | A custom metric is a rubric you wrote; a built-in is a prompt AWS wrote. |
| Human evaluation | A calibration sample, not the pipeline | It is the only outside opinion, and the only one that does not scale. |
| RAG quality | Retrieve-only first | It isolates the ceiling #61 describes, before generation can mask it. |
| Reading results | The S3 report, not the console | The console explains the first five prompts, chosen by file order. |
What makes a rubric worth running
Three properties, and a built-in metric has at most two of them. It has to be stable — the same response scores the same next month, which means pinning the judge. It has to be discriminating — if every response scores highly the metric is measuring nothing, which is what the histogram is for and why it beats an average. And it has to be actionable — a low score has to imply a change somebody can make.
Helpfulness fails the third routinely. Faithfulness passes all three. Correctness passes all three only if you wrote the answer key, which returns to the same place: the evaluation is cheap and the key is not.
Closing Thought
Every post in this block has ended in the same place. Chunking decides what can be found, and you have to measure retrieval to know. The store decides how often you can ask, and you have to measure queries per second to choose. Guardrails decide what gets filtered, and you have to test what the filter saw. All three need a measurement nobody is required to take.
Bedrock evaluations removes almost every obstacle to taking them. It scores non-Bedrock models, it handles the judging, it writes the report. What it cannot do is decide what a good answer is for your business, and it does not stop you running a metric named Correctness without ever saying what correct means.
So the part everyone skips is not the evaluation job. It is the hour with the corpus, writing down what should have come back. Everything downstream of that hour is automatable, and nothing upstream of it is worth much without it.
AI & ML — Comprehend, Textract and Rekognition: the task-shaped services, and when a purpose-built API beats a foundation model on the same job.
Official AWS Reference
- Evaluate the performance of Amazon Bedrock resources — the three kinds of evaluation and the RAG ground-truth requirement
- Evaluate model performance using another LLM as a judge — generator and evaluator models, the supported judge lists, and what the console shows
- Use metrics to understand model performance — the eleven built-in metrics and which of them consider ground truth
- Evaluate the performance of RAG sources using Amazon Bedrock evaluations — retrieve-only versus retrieve-and-generate, and bringing an external RAG source
Comments