Business Challenge
#62 covered what Bedrock capacity commits you to. This is the decision one level up: call a hosted model, or run your own endpoint.
It is usually framed as consume versus build, and framed that way it sounds like a preference. Start simple with Bedrock, move to SageMaker if you need control, come back if you do not. The documentation does not describe that shape at all.
Bedrock Custom Model Import is the real bridge. It lets you take a model “that you have created in Amazon SageMaker AI that has proprietary model weights” and invoke it through Bedrock's own APIs. That is a genuine escape hatch and worth knowing exists.
Then the exclusion list: “You can't use Custom Model Import with the following Amazon Bedrock features. Batch inference. CloudFormation.”
No CloudFormation means the model you imported is not a managed resource. Every practice built around declarative infrastructure — review, drift detection, teardown, reproducing the stack in another account — does not apply to it. That is a large thing to give up in exchange for an easier inference API.
FixPrice the escape hatch in lost IaC, not just in migration effort.
The serverless documentation states it twice, in both directions:
“You cannot convert your instance-based, real-time endpoint to a serverless
endpoint. If you try to update your real-time endpoint to serverless, you receive a
ValidationError message.”
And the other way: “You can convert a serverless endpoint to real-time, but once you make the update, you cannot roll it back to serverless.”
So serverless to real-time is available exactly once, and real-time to serverless is not available at all. Starting on real-time because it is the default forecloses the cheaper option permanently for that endpoint.
FixIf serverless might ever be right, start there. The conversion only runs one direction.
Serverless Inference reads like the natural middle ground: managed, scales to zero, pay per use. The feature exclusions end that reading for language models.
“Some of the features currently available for SageMaker AI Real-time Inference are not supported for Serverless Inference, including GPUs, AWS marketplace model packages, private Docker registries, Multi-Model Endpoints, VPC configuration, network isolation, data capture, multiple production variants, Model Monitor, and inference pipelines.”
Memory tops out at 6144 MB and the container image at 10 GB. Between no GPU and that ceiling, serverless is for small models — and the absence of “VPC configuration” and “network isolation” rules it out for a whole class of regulated workloads regardless of size.
FixTreat serverless as a small-model option, not as cheap SageMaker.
Architecture
The clearest way to hold this is as a map of transitions rather than a comparison of destinations.
What SageMaker actually offers, before the comparison starts
Four hosting shapes, and the documentation is precise about which is for what. Real-time is “ideal for inference workloads where you have interactive, low latency requirements”. Serverless is for “workloads which have idle periods between traffic spurts and can tolerate cold starts”. Asynchronous “queues incoming requests” and suits “large payload sizes (up to 1GB), long processing times (up to one hour), and near real-time latency requirements”.
That third one is the option most often forgotten and it is the one that competes with Bedrock most directly for document-processing work. A one-hour processing budget and a gigabyte payload are not things a synchronous hosted-model call will give you.
The bridge, and its five constraints
Custom Model Import is a genuinely useful feature, and the constraints are all documented in one place:
- Four Regions — “eu-central-1, us-east-1, us-east-2, us-west-2”.
- No batch inference, no CloudFormation.
- No embedding models — “Custom Model Import does not support embedding models.” Which, read against #63, means the model producing your vectors cannot take this path at all.
- Size and context ceilings — weights “less than 100GB for multimodal models and 200GB for text models”, context “less than 128K”. Text models get twice the weight budget of multimodal ones.
- An architecture allowlist — Mistral, Mixtral, Flan, the Llama family, GPTBigCode, the Qwen family, GPT-OSS. With footnotes: GPT-OSS “is only supported in the US East (N. Virginia) region”, and for Qwen3 “Converse is also not supported”.
“Amazon Bedrock overrides llama3 rope_scaling value with the following values: original_max_position_embeddings=8192, high_freq_factor=4, low_freq_factor=1, factor=8.” Positional-encoding scaling is what governs long-context behaviour. A Llama 3 model tuned with different values arrives in Bedrock with AWS's values instead, which means the imported model is not bit-identical in behaviour to the one you tested in SageMaker. Also pin your toolchain: “Amazon Bedrock supports transformer version 4.51.3”, and that is the version to fine-tune with.
And there is no path in the other direction
There is no Bedrock-to-SageMaker equivalent, and the reason is structural rather than a missing feature: a hosted Bedrock model's weights were never yours. You can move a model you own into Bedrock; you cannot extract one you have only been calling.
So the asymmetry is real but not symmetric-with-a-cost. One direction is a constrained migration; the other is not a migration at all.
Why This Architecture Holds Up
CloudFormation goes missing twice, independently
The Custom Model Import exclusion is one. The other is on the serverless page: “Application Auto Scaling for Serverless Inference with Provisioned Concurrency is currently not supported on AWS CloudFormation.”
Two unrelated pages, two features that cannot be expressed declaratively. Worth treating as a pattern when planning anything here: the newest capability on each side tends to arrive with console and SDK support first, and the IaC path later. If your deployment standard is CloudFormation-only, that standard is a constraint on which features you can adopt, and it is better discovered at design time than in a pipeline.
“You can set the maximum concurrency for a single endpoint up to 200, and the total number of serverless endpoints you can host in a Region is 50” — against a regional total of “1000” in the larger Regions and “500” in the rest. So five endpoints configured at their individual maximum consume the entire budget of a 1000-concurrency Region, while the quota invites you to host fifty. AWS names the reason for the per-endpoint cap directly: it “prevents that endpoint from taking up all of the invocations allowed for your account”. Which tells you the failure being guarded against is one team's endpoint throttling everybody else's.
Cold starts are measurable, which makes them a design input
“The cold start time depends on your model size, how long it takes to download your model, and
the start-up time of your container”, and it is observable:
“you can use the Amazon CloudWatch metric OverheadLatency to
monitor your serverless endpoint.”
That is the honest way to settle the serverless question rather than arguing about it. Deploy, measure
OverheadLatency, compare against the latency you actually owe users. All three
inputs are things you control — and the first, model size, is the one that also decides whether the
6144 MB ceiling is reachable at all.
The three deployment paths are an organisational choice, not a technical one
SageMaker documents three ways in, and the differentiator in each is who is doing it. JumpStart in Studio
is “ideal for citizen data scientists” and carries a
“Lack of customization for container settings”.
ModelBuilder suits someone who
“has their own model to deploy and requires fine-grained control” but has
“No UI”. CloudFormation is for “ongoing management of models in
production” and “Requires infrastructure management and organizational
resources”.
Read as a progression, that is the real cost of the build side: not the endpoint, but the third column. A model deployed from a notebook and a model deployed from a pipeline are the same model and completely different operational commitments.
Key Architecture Decisions
| Decision | Choice | Reasoning |
|---|---|---|
| Framing | Map the transitions, not the end states | All three documented transitions are one-way; the comparison hides that. |
| Default starting point | Bedrock, unless a named requirement | The path out exists; the path back into a hosted model's weights does not. |
| Serverless or real-time | Serverless first if it is plausible | Real-time cannot become serverless. Serverless converts to real-time once. |
| Serverless for an LLM | No | No GPUs, 6144 MB memory ceiling, 10 GB container limit. |
| Regulated workloads | Real-time, not serverless | VPC configuration and network isolation are both excluded from serverless. |
| Large documents, long jobs | Asynchronous inference | 1 GB payloads and one-hour processing; no hosted-model call offers that. |
| Embedding models | SageMaker only, permanently | Custom Model Import does not support them, so the bridge is closed for the vector path. |
| IaC standard | Check before adopting | CloudFormation is unsupported for Custom Model Import and for serverless provisioned-concurrency scaling. |
The question that settles it fastest
Not “do we need control”, which always answers yes. Ask instead: will we ever hold weights we cannot get from a vendor? A fine-tune on proprietary data, a domain adaptation, a model trained from scratch — those are the cases Custom Model Import exists for, and they are the only cases where the build side buys something the consume side cannot eventually offer.
If the answer is no, every argument for SageMaker reduces to capacity, latency or Region shape, and those are cheaper to solve on the consume side. If the answer is yes, start on SageMaker knowing the bridge to Bedrock is there, and knowing it comes without CloudFormation.
Closing Thought
Consume versus build is a comfortable framing because it sounds like something you can change your mind about. Most of these decisions are, and this one mostly is not — not because AWS made it hard, but because each transition happens to be documented as one-directional, and nobody presents them together.
The pattern across all three is the same: you can always move toward more managed, and you cannot move back. Serverless converts up to real-time, never down. A model you own can move into Bedrock; a model Bedrock hosts cannot come out. The gradient runs one way, and starting further along it than you need to is the expensive mistake — expensive not in money but in options you quietly no longer have.
Which makes the useful discipline the same as in #61 and #63: find the decision that is hardest to reverse and make that one first, deliberately, rather than inheriting it from whichever tutorial you started with.
AI & ML — evaluating model output, the part everyone skips: what a scoring rubric has to contain to be worth running, why human review does not scale and LLM-as-judge does not agree with itself, and where the evaluation belongs in a pipeline.
Official AWS Reference
- Deploy models for inference — the four hosting options and the three deployment paths, with what each is recommended for
- Use Custom model import to import a customized open-source model into Amazon Bedrock — the supported architectures, the Region list, the feature exclusions and the rope_scaling override
- Deploy models with Amazon SageMaker Serverless Inference — memory sizes, concurrency quotas, cold starts and the feature exclusions including GPUs
Comments