Home› Blog› AWS Architecture Series #67 — A foundation model never says no…
AWS Architecture AWS Architecture Series

AWS Architecture Series #67 — A foundation model never says no

A foundation model can read a document, classify text and describe an image, so the task-shaped services look like legacy. They are not competing on capability. They are competing on refusal — and a system that tells you it cannot do something is worth more than one that always produces an answer.

Verified against current vendor documentation on 29 September 2026. Pricing, limits and API behaviour were checked against the official docs on that date. Cloud services change fast — if you are reading this much later, treat the specifics as a starting point and re-check the linked sources.

Business Challenge

#66 ended on the thing nobody writes: an answer key. This post is about the services that ship with one.

Ask a foundation model to extract the line items from an invoice and it will. Ask it about a blank page, a rotated scan, a 400-page PDF or a language it was barely trained on, and it will still produce something. That is the difference being discussed here, and it is not about accuracy.

1One number decides your entire document architecture

Textract's set quotas — the ones that “cannot be changed” — contain this: “For synchronous operations, JPEG, PNG, PDF, and TIFF files have a limit of 10 MB in memory. PDF and TIFF files also have a limit of 1 page.”

One page. Asynchronous accepts “500 MB in memory and a limit of 3,000 pages” — fifty times the size — but asynchronous means an S3 bucket, a job id, an SNS topic and polling.

So the question “will our documents ever be more than one page” decides whether you have an API call or a pipeline. It is answerable on day one and expensive to discover on day thirty.

Fix

Establish the page count before the design, not during it. The limit is not negotiable.

2The same service has different limits per operation

Comprehend's real-time quota for “detecting entities, key-phrases, or dominant language” is 100 KB. For “detecting sentiment, targeted sentiment, and syntax” it is 5 KB.

Twenty times smaller, same service, same SDK, no difference in the call shape. A document that entity-detection accepts will be rejected by sentiment analysis.

Rekognition has the same pattern in a different dimension: “Maximum image size stored as an Amazon S3 object is limited to 15 MB” but “The maximum images size as raw bytes passed in as parameter to an API is 5 MB.” Same image, three times the allowance, depending only on whether you pass a pointer or the pixels.

Fix

Read the quota for the operation you are calling, not for the service.

3The throughput is far lower than a model's, and stated honestly

A Comprehend custom endpoint scales in inference units, and one unit gives “Maximum throughput per inference unit (characters): 100/second” and “Maximum throughput per inference unit (documents): 2/second.”

Those two ceilings meet at 50 characters per document. Anything longer than a short sentence is limited by the character rate, not the document rate — so “2 documents per second” is the optimistic number and almost never the binding one.

With “a maximum of 50 inference units per endpoint”, a single endpoint tops out around 5,000 characters per second. That is a real constraint, and it is published rather than discovered.

Fix

Size on characters per second. The document rate is the wrong unit above a sentence.

Architecture

The useful way to read these three is as published contracts rather than as capabilities.

Diagram: the task-shaped AWS services read through their published set quotas, the limits AWS states cannot be changed. Amazon Textract accepts JPEG, PNG, PDF and TIFF but not XFA-based or password-protected PDFs. Its synchronous path caps files at 10 megabytes and PDF or TIFF at one page, while its asynchronous path allows 500 megabytes and 3,000 pages, fifty times the file size, so the page count of your documents decides whether the design is an API call or an S3 bucket with a queue and polling behind it. Textract supports queries at 15 per page synchronously and 30 asynchronously, detects text in English, French, German, Italian, Portuguese and Spanish only, does not return the detected language, supports handwriting in English only, cannot read vertical text as written in Japanese and Chinese, requires text at least 15 pixels high which is 8 point at 150 DPI, and its AnalyzeID operation supports only United States passports and driver's licenses. Amazon Comprehend allows 100 kilobytes per document for entities, key phrases and dominant language but only 5 kilobytes for sentiment, targeted sentiment and syntax, a twentyfold difference within the same service, with batches capped at 5 kilobytes and 25 documents per request and asynchronous operations limited to 10 active jobs. Its custom endpoints allow 20 active endpoints per Region, 200 inference units per Region and 50 per endpoint, with each unit delivering 100 characters per second or 2 documents per second, two ceilings that meet at 50 characters per document so that anything longer is limited by the character rate. Amazon Rekognition accepts PNG and JPEG only, 15 megabytes from S3 against 5 megabytes as raw bytes in the request, requires a face to be at least 40 by 40 pixels within a 1920 by 1080 image, detects up to 100 words per image with DetectText and protective equipment on up to 15 people, stores up to 20 million face vectors per collection while returning at most 4096 matches, and analyses stored video up to 10 gigabytes and 6 hours encoded as H.264 in MPEG-4 or MOV with 20 concurrent jobs per account. A closing panel records the argument: every number here is a refusal, and a foundation model asked the same question returns an answer instead, which is the actual difference between the two options.
Every number here is a refusal stated in advance. That is what you are buying.

The limits describe the training data

Textract's language list is the clearest example: “English, French, German, Italian, Portuguese, and Spanish”, and more pointedly, “Amazon Textract does not support vertical text (text written vertically, as is common in languages like Japanese and Chinese) alignment within the document.”

Handwriting narrows further — “only supported in English” — and queries narrow again: “Query detection is only available in English document detection.” Three concentric circles, each smaller than the last, all published.

The sharpest of all is AnalyzeID, which “only supports US passports, and US driver's licenses.” Not “works best with”. A service that names two document types from one country is telling you precisely what it was built on, and a foundation model asked to parse a Portuguese identity card will return a confident structure either way.

And one limit that will look like a bug in production

“DetectText can detect up to 100 words in an image.” Not an error, not a truncation warning — a cap. Photograph a page of text and you get a hundred words of it. The failure is silent, the response is well-formed, and the missing words look exactly like words the model failed to read. If the job is a page of prose, that is Textract's job, not Rekognition's, and the 100-word line is where the two services divide.

Where the ceilings actually bind

Rekognition's face collections hold “20 million” face vectors, but “The maximum matching face vectors the search API returns is 4096.” Storage is generous; a single answer is bounded — the same shape as the S3 Vectors top-K pagination split in #63.

Video is bounded twice over: “up to 10GB in size” and “up to 6 hours in length”, with “a maximum of 20 concurrent jobs per account” and a codec requirement — “must be encoded using the H.264 codec. The supported file formats are MPEG-4 and MOV.” Twenty concurrent jobs is the number that shapes a backlog, and it is per account rather than per bucket or per Region.

Comprehend's async side caps at “a maximum of 10 active jobs” per operation. Both of those are queue-depth constraints wearing quota clothing.

Why This Architecture Holds Up

A refusal is a testable contract

This is the argument for the whole category. When Textract rejects a two-page PDF on the synchronous path, that is a deterministic, documented, reproducible behaviour you can write a test against. When a foundation model is handed a 400-page document that exceeds what it can usefully attend to, it returns an answer, and finding out that the answer degraded requires the evaluation work from #66 — the work nobody does.

That inverts the usual framing. The task-shaped service is not the conservative choice because it is older. It is the conservative choice because its failure mode is an exception rather than a plausible paragraph.

The corollary: these services are not where you put open-ended work

Everything above is an argument for using them on bounded tasks — this field from this form, these labels from this image, this sentiment from this review. A sharp contract is only an advantage where the requirement is sharp. Asked to summarise a contract, explain why an invoice looks wrong, or answer a question about a document, none of these services applies at all, and reaching for one because it feels safer is how you end up building a foundation model badly out of API calls. The division is not old versus new. It is bounded versus open-ended.

Cost follows the shape too

These services bill per unit of work — a page, an image, a document — which makes the cost of a pipeline computable in advance from a volume estimate. #62 showed what it takes to do that on Bedrock: a quota measured in tokens, a burndown multiplier that differs per model, and a break-even that cannot be derived from public figures.

There is one exception to watch. Comprehend custom endpoints are provisioned, billed in inference units, so they carry the same shape as Provisioned Throughput — an hourly floor under a workload whether or not it runs. The built-in synchronous APIs do not.

Throttling is expected, and AWS says what to do about it

Rekognition's guidance is unusually direct: “Spiky traffic affects throughput. To get maximum throughput for the allotted transactions per second (TPS), use a queueing serverless architecture or another mechanism to 'smooth' traffic so it is more consistent.”

Read that as a design instruction rather than a troubleshooting note. The recommended architecture for these services is a queue in front of them, which means the asynchronous shape Textract forces on you at two pages is the shape you probably wanted anyway once volume is real.

Key Architecture Decisions

Decision Choice Reasoning
Task-shaped or foundation model Bounded task → task-shaped A documented refusal beats a plausible answer when the requirement is sharp.
Textract sync or async Decided by page count, on day one Synchronous PDF is one page, and the limit cannot be raised.
Comprehend document size Check per operation 100 KB for entities, 5 KB for sentiment — same service, twentyfold gap.
Comprehend capacity Size in characters per second The two ceilings cross at 50 characters; above that the document rate is fiction.
Rekognition image transfer Via S3, not raw bytes 15 MB against 5 MB, for the identical image.
Text in images Textract, not Rekognition DetectText DetectText caps at 100 words and does not say it truncated.
Non-Latin or vertical scripts Not Textract Vertical text is unsupported, and handwriting is English only.
Traffic shape Queue in front, always AWS's own throughput guidance, and the async path already assumes it.

The question that picks the service

Not “can a model do this” — it usually can. Ask instead: what should happen when the input is outside what we planned for?

If the right answer is an exception, a retry queue and an alarm, you want a service with a published limit. If the right answer is a best effort on unpredictable input, you want a model and you have accepted that you now owe an evaluation harness. Both are legitimate. Choosing the second by accident, because it demoed well on a clean sample, is the mistake.

Closing Thought

There is a reading of this block — posts 61 through 66 — where every one ends at the same problem: the system will produce output regardless, and knowing whether the output is right requires work that nothing forces you to do. Chunking, vector stores, guardrails, evaluation. Four different subjects, one recurring shape.

The task-shaped services are the counterexample, and that is why they are worth a post rather than a footnote. Textract will not read a second page synchronously. Comprehend will not take 6 KB for sentiment. Rekognition will not return a 4097th match. Each of those is a sentence in the documentation and an exception in your logs, and neither requires you to have built anything to discover it.

That is not nostalgia for narrow APIs. It is the observation that a system which can refuse is a system you can reason about — and on the subset of problems that are genuinely bounded, that property is worth more than the flexibility you would be trading it for.

Next in this series

AI & ML — inference endpoints: real-time, serverless, batch and async, and how to choose between them when the payload size and the latency budget disagree.

Comments

How was your experience?
Your feedback helps improve this site.
PoorExcellent