Homeβ€Ί Blogβ€Ί AWS Architecture Series #65 β€” Every door in this decision opens one way…
AWS Architecture AWS Architecture Series

AWS Architecture Series #65 β€” Every door in this decision opens one way

Bedrock or SageMaker gets framed as consume versus build, which sounds like a choice you can revisit when the constraints change. The documentation describes something stiffer: three separate transitions between these options, and every one of them is one-way.

Verified against current vendor documentation on 27 September 2026. Pricing, limits and API behaviour were checked against the official docs on that date. Cloud services change fast — if you are reading this much later, treat the specifics as a starting point and re-check the linked sources.

Business Challenge

#62 covered what Bedrock capacity commits you to. This is the decision one level up: call a hosted model, or run your own endpoint.

It is usually framed as consume versus build, and framed that way it sounds like a preference. Start simple with Bedrock, move to SageMaker if you need control, come back if you do not. The documentation does not describe that shape at all.

1The path out of SageMaker exists, and it drops CloudFormation

Bedrock Custom Model Import is the real bridge. It lets you take a model “that you have created in Amazon SageMaker AI that has proprietary model weights” and invoke it through Bedrock's own APIs. That is a genuine escape hatch and worth knowing exists.

Then the exclusion list: “You can't use Custom Model Import with the following Amazon Bedrock features. Batch inference. CloudFormation.”

No CloudFormation means the model you imported is not a managed resource. Every practice built around declarative infrastructure — review, drift detection, teardown, reproducing the stack in another account — does not apply to it. That is a large thing to give up in exchange for an easier inference API.

Fix

Price the escape hatch in lost IaC, not just in migration effort.

2Inside SageMaker, the cheap option is a one-way trip too

The serverless documentation states it twice, in both directions: “You cannot convert your instance-based, real-time endpoint to a serverless endpoint. If you try to update your real-time endpoint to serverless, you receive a ValidationError message.”

And the other way: “You can convert a serverless endpoint to real-time, but once you make the update, you cannot roll it back to serverless.”

So serverless to real-time is available exactly once, and real-time to serverless is not available at all. Starting on real-time because it is the default forecloses the cheaper option permanently for that endpoint.

Fix

If serverless might ever be right, start there. The conversion only runs one direction.

3Serverless has no GPUs, which removes it from most of this decision

Serverless Inference reads like the natural middle ground: managed, scales to zero, pay per use. The feature exclusions end that reading for language models.

“Some of the features currently available for SageMaker AI Real-time Inference are not supported for Serverless Inference, including GPUs, AWS marketplace model packages, private Docker registries, Multi-Model Endpoints, VPC configuration, network isolation, data capture, multiple production variants, Model Monitor, and inference pipelines.”

Memory tops out at 6144 MB and the container image at 10 GB. Between no GPU and that ceiling, serverless is for small models — and the absence of “VPC configuration” and “network isolation” rules it out for a whole class of regulated workloads regardless of size.

Fix

Treat serverless as a small-model option, not as cheap SageMaker.

Architecture

The clearest way to hold this is as a map of transitions rather than a comparison of destinations.

Diagram: the transitions available between calling a hosted model on Amazon Bedrock and running your own endpoint on SageMaker AI, and the direction each one runs. On the left, Amazon Bedrock on-demand inference, where you call a hosted model and never hold the weights. On the right, SageMaker AI hosting, with its four inference options: real-time for interactive low latency requirements, serverless for workloads with idle periods between traffic spurts that can tolerate cold starts, asynchronous for large payload sizes up to one gigabyte and long processing times up to one hour, and batch. Between them, one documented path runs from SageMaker to Bedrock, which is Custom Model Import: it accepts a model customised elsewhere including one with proprietary weights, supports on-demand throughput through InvokeModel, and is constrained in five ways, namely that it cannot be used with batch inference or with CloudFormation, it runs in only four Regions being Frankfurt, North Virginia, Ohio and Oregon, it refuses embedding models, it requires weights below 100 gigabytes for multimodal and 200 gigabytes for text models with context length under 128 thousand, and it accepts only an allowlist of architectures with GPT-OSS restricted to North Virginia and Converse unsupported for Qwen3. No path runs from Bedrock back to SageMaker, because a hosted model's weights were never yours to move. Inside SageMaker a second one-way door is marked: a real-time endpoint cannot be converted to serverless at all, returning a ValidationError, while a serverless endpoint can be converted to real-time exactly once and cannot be rolled back. A panel records why serverless is narrower than it appears: it excludes GPUs, VPC configuration, network isolation, Multi-Model Endpoints, Model Monitor, data capture, marketplace packages, private registries, multiple production variants and inference pipelines, with memory capped at 6144 megabytes and container images at 10 gigabytes. A closing panel notes that CloudFormation is absent twice independently, once for Custom Model Import and once for Application Auto Scaling of serverless provisioned concurrency, and that five serverless endpoints at their maximum concurrency of 200 consume an entire 1000-concurrency regional budget even though 50 endpoints may be hosted.
One path from SageMaker to Bedrock, constrained five ways. No path back. And a one-way door inside SageMaker as well.

What SageMaker actually offers, before the comparison starts

Four hosting shapes, and the documentation is precise about which is for what. Real-time is “ideal for inference workloads where you have interactive, low latency requirements”. Serverless is for “workloads which have idle periods between traffic spurts and can tolerate cold starts”. Asynchronous “queues incoming requests” and suits “large payload sizes (up to 1GB), long processing times (up to one hour), and near real-time latency requirements”.

That third one is the option most often forgotten and it is the one that competes with Bedrock most directly for document-processing work. A one-hour processing budget and a gigabyte payload are not things a synchronous hosted-model call will give you.

The bridge, and its five constraints

Custom Model Import is a genuinely useful feature, and the constraints are all documented in one place:

  • Four Regions — “eu-central-1, us-east-1, us-east-2, us-west-2”.
  • No batch inference, no CloudFormation.
  • No embedding models — “Custom Model Import does not support embedding models.” Which, read against #63, means the model producing your vectors cannot take this path at all.
  • Size and context ceilings — weights “less than 100GB for multimodal models and 200GB for text models”, context “less than 128K”. Text models get twice the weight budget of multimodal ones.
  • An architecture allowlist — Mistral, Mixtral, Flan, the Llama family, GPTBigCode, the Qwen family, GPT-OSS. With footnotes: GPT-OSS “is only supported in the US East (N. Virginia) region”, and for Qwen3 “Converse is also not supported”.
And it will change your model's configuration without asking

“Amazon Bedrock overrides llama3 rope_scaling value with the following values: original_max_position_embeddings=8192, high_freq_factor=4, low_freq_factor=1, factor=8.” Positional-encoding scaling is what governs long-context behaviour. A Llama 3 model tuned with different values arrives in Bedrock with AWS's values instead, which means the imported model is not bit-identical in behaviour to the one you tested in SageMaker. Also pin your toolchain: “Amazon Bedrock supports transformer version 4.51.3”, and that is the version to fine-tune with.

And there is no path in the other direction

There is no Bedrock-to-SageMaker equivalent, and the reason is structural rather than a missing feature: a hosted Bedrock model's weights were never yours. You can move a model you own into Bedrock; you cannot extract one you have only been calling.

So the asymmetry is real but not symmetric-with-a-cost. One direction is a constrained migration; the other is not a migration at all.

Why This Architecture Holds Up

CloudFormation goes missing twice, independently

The Custom Model Import exclusion is one. The other is on the serverless page: “Application Auto Scaling for Serverless Inference with Provisioned Concurrency is currently not supported on AWS CloudFormation.”

Two unrelated pages, two features that cannot be expressed declaratively. Worth treating as a pattern when planning anything here: the newest capability on each side tends to arrive with console and SDK support first, and the IaC path later. If your deployment standard is CloudFormation-only, that standard is a constraint on which features you can adopt, and it is better discovered at design time than in a pipeline.

The serverless concurrency budget is smaller than the endpoint count suggests

“You can set the maximum concurrency for a single endpoint up to 200, and the total number of serverless endpoints you can host in a Region is 50” — against a regional total of “1000” in the larger Regions and “500” in the rest. So five endpoints configured at their individual maximum consume the entire budget of a 1000-concurrency Region, while the quota invites you to host fifty. AWS names the reason for the per-endpoint cap directly: it “prevents that endpoint from taking up all of the invocations allowed for your account”. Which tells you the failure being guarded against is one team's endpoint throttling everybody else's.

Cold starts are measurable, which makes them a design input

“The cold start time depends on your model size, how long it takes to download your model, and the start-up time of your container”, and it is observable: “you can use the Amazon CloudWatch metric OverheadLatency to monitor your serverless endpoint.”

That is the honest way to settle the serverless question rather than arguing about it. Deploy, measure OverheadLatency, compare against the latency you actually owe users. All three inputs are things you control — and the first, model size, is the one that also decides whether the 6144 MB ceiling is reachable at all.

The three deployment paths are an organisational choice, not a technical one

SageMaker documents three ways in, and the differentiator in each is who is doing it. JumpStart in Studio is “ideal for citizen data scientists” and carries a “Lack of customization for container settings”. ModelBuilder suits someone who “has their own model to deploy and requires fine-grained control” but has “No UI”. CloudFormation is for “ongoing management of models in production” and “Requires infrastructure management and organizational resources”.

Read as a progression, that is the real cost of the build side: not the endpoint, but the third column. A model deployed from a notebook and a model deployed from a pipeline are the same model and completely different operational commitments.

Key Architecture Decisions

Decision Choice Reasoning
Framing Map the transitions, not the end states All three documented transitions are one-way; the comparison hides that.
Default starting point Bedrock, unless a named requirement The path out exists; the path back into a hosted model's weights does not.
Serverless or real-time Serverless first if it is plausible Real-time cannot become serverless. Serverless converts to real-time once.
Serverless for an LLM No No GPUs, 6144 MB memory ceiling, 10 GB container limit.
Regulated workloads Real-time, not serverless VPC configuration and network isolation are both excluded from serverless.
Large documents, long jobs Asynchronous inference 1 GB payloads and one-hour processing; no hosted-model call offers that.
Embedding models SageMaker only, permanently Custom Model Import does not support them, so the bridge is closed for the vector path.
IaC standard Check before adopting CloudFormation is unsupported for Custom Model Import and for serverless provisioned-concurrency scaling.

The question that settles it fastest

Not “do we need control”, which always answers yes. Ask instead: will we ever hold weights we cannot get from a vendor? A fine-tune on proprietary data, a domain adaptation, a model trained from scratch — those are the cases Custom Model Import exists for, and they are the only cases where the build side buys something the consume side cannot eventually offer.

If the answer is no, every argument for SageMaker reduces to capacity, latency or Region shape, and those are cheaper to solve on the consume side. If the answer is yes, start on SageMaker knowing the bridge to Bedrock is there, and knowing it comes without CloudFormation.

Closing Thought

Consume versus build is a comfortable framing because it sounds like something you can change your mind about. Most of these decisions are, and this one mostly is not — not because AWS made it hard, but because each transition happens to be documented as one-directional, and nobody presents them together.

The pattern across all three is the same: you can always move toward more managed, and you cannot move back. Serverless converts up to real-time, never down. A model you own can move into Bedrock; a model Bedrock hosts cannot come out. The gradient runs one way, and starting further along it than you need to is the expensive mistake — expensive not in money but in options you quietly no longer have.

Which makes the useful discipline the same as in #61 and #63: find the decision that is hardest to reverse and make that one first, deliberately, rather than inheriting it from whichever tutorial you started with.

Next in this series

AI & ML — evaluating model output, the part everyone skips: what a scoring rubric has to contain to be worth running, why human review does not scale and LLM-as-judge does not agree with itself, and where the evaluation belongs in a pipeline.

Comments

How was your experience?
Your feedback helps improve this site.
PoorExcellent