Homeβ€Ί Blogβ€Ί AWS Architecture Series #69 β€” Small records cost more, and the cheap path starts at 60 seconds…
AWS Architecture AWS Architecture Series

AWS Architecture Series #69 β€” Small records cost more, and the cheap path starts at 60 seconds

Streaming versus batch gets argued as a latency question, and the cost side is assumed to follow from volume. It does not. The cheap managed path bills per record rounded up to 5 KB and cannot deliver faster than 60 seconds, so the two axes are set by record size and buffer interval rather than by how much data you move.

Verified against current vendor documentation on 1 October 2026. Pricing, limits and API behaviour were checked against the official docs on that date. Cloud services change fast — if you are reading this much later, treat the specifics as a starting point and re-check the linked sources.

Business Challenge

#60 covered which streaming service to pick. This is the cost lens on the same ground: what the choice actually bills for, and where the latency floor sits.

The usual framing is a trade between latency and cost, with volume as the driver. Two documented facts change the shape of it.

1Firehose bills per record, rounded up to 5 KB

The quota page states the pricing model directly, which pricing pages rarely do: “Firehose ingestion pricing is based on the number of data records you send to the service, times the size of each record rounded up to the nearest 5KB (5120 bytes).”

So a 100-byte event is billed as 5,120 bytes — over fifty times its own size. Shrinking that event to 50 bytes saves nothing. Batching two events into one record halves the bill.

AWS gives the worked example: “if the total incoming data volume is 5MiB, sending 5MiB of data over 5,000 records costs more compared to sending the same amount of data using 1,000 records.” Same bytes, different bill.

Fix

Optimise record count, not payload size. The rounding unit makes small events the expensive case.

2The managed path cannot go below 60 seconds

Firehose buffers before delivering, and the dial has a documented floor: “The buffer interval hints range from 60 seconds to 900 seconds.”

One minute to fifteen. If your requirement is sub-minute, Firehose is not a slow option to tune — it is unavailable, and you are running Kinesis Data Streams with a consumer application you operate yourself.

That is the real step in the curve. It is not gradual: there is a managed, cheap, 60-second-minimum tier, and below it a tier where you own the consumer.

Fix

Write the latency requirement down in seconds first. Sixty is the line that decides the architecture.

3The 60%-cheaper streaming mode is a reservation

On-demand Advantage reads like a free upgrade: “Data ingest, data retrieval, and extended retention usage across all on-demand streams are at least 60% lower than in On-demand Standard”, and “there's no longer a fixed charge for each stream.”

Then the condition: “Enabling On-demand Advantage commits the account to at least 25MiB/s of data ingest and 25MiB/s of data retrieval across all on-demand streams.” Under-use it and “you'll be charged the difference as a shortfall.”

And it is account-level with a lock-in period: “a minimum period of 24 hours before you can disable the mode.” A serverless-shaped product with a floor, a shortfall charge and a cooling-off period is a reserved-capacity deal.

Fix

Measure sustained ingest and retrieval against 25 MiB/s each before enabling. Below it you pay for air.

Architecture

The curve has a step in it, and the step is at sixty seconds.

Diagram: the latency and cost curve between streaming and batch on AWS, drawn from documented quotas rather than from data volume. At the low-latency end sits Kinesis Data Streams with a consumer application you operate yourself, reaching sub-second delivery. It offers three modes. On-demand Standard requires no capacity planning and accommodates up to double the peak write throughput observed in the previous thirty days, with write throttling possible if traffic more than doubles the previous peak within fifteen minutes. On-demand Advantage is an account-level mode where ingest, retrieval and extended retention are at least sixty percent lower than Standard and the per-stream fixed charge disappears, but enabling it commits the account to at least twenty-five mebibytes per second of ingest and twenty-five of retrieval, bills any shortfall at the discounted rate, and cannot be disabled for twenty-four hours; it also raises enhanced fan-out consumers from twenty to fifty per stream with no price premium. Provisioned mode requires specifying shards, sized by the formula ceiling of the maximum of incoming write bandwidth in kibibytes divided by 1024 and outgoing read bandwidth divided by 2048. A panel notes that on-demand splits shards evenly when traffic rises but explicitly does not detect or isolate hot hash keys, so AWS recommends provisioned mode for granular shard splits, and that a single partition key exceeding one megabyte per second or one thousand records per second will still throttle. At the managed end sits Amazon Data Firehose, whose buffer interval hints range from sixty seconds to nine hundred seconds, making sixty seconds the floor for the managed path. Firehose ingestion is billed by the number of records times each record's size rounded up to the nearest five kilobytes, which is 5,120 bytes, so a hundred-byte event is charged for more than fifty times its own size, and AWS's own example shows five mebibytes sent as five thousand records costing more than the same five mebibytes sent as one thousand records. A closing panel records the design consequence: optimise record count rather than payload size, write the latency requirement in seconds because sixty is the line that decides the architecture, and measure sustained throughput against the twenty-five mebibyte commitment before enabling the discount.
Two axes, both documented: record count sets the bill, buffer interval sets the floor.

Three Kinesis modes, and they are not three prices

“A mode determines how the capacity of a data stream is managed and how you're charged for the usage of your data stream.” That sentence is the whole point — the mode is a billing model, not a performance tier.

On-demand Standard “accommodates up to double the peak write throughput observed in the previous 30 days”, which is a useful and under-appreciated guarantee: your headroom is a function of your own history. The matching warning is that “write throttling can occur if your traffic increases to more than double the previous peak within a 15-minute duration.”

Provisioned has an actual formula, which is rare enough to quote: number_of_shards = ceiling(max(incoming_write_bandwidth_in_KiB/1024, outgoing_read_bandwidth_in_KiB/2048)). Note the asymmetry — reads get twice the divisor, because a shard offers 2 MB/s out against 1 MB/s in.

And on-demand explicitly cannot fix a hot key

This inverts the usual assumption that the serverless mode is strictly better. “In the on-demand mode, Kinesis Data Streams splits the shards evenly when it detects an increase in traffic. However, it does not detect and isolate hash keys that are driving a higher portion of incoming traffic to a particular shard. If you are using highly uneven partition keys you may continue to receive write exceptions. For such use cases, we recommend that you use the provisioned capacity mode that supports granular shard splits.” So skew is the one case where you are told to leave the managed mode β€” and it is the case Service-Managed Partition Keys addressed by removing the key entirely, at the cost of ordering.

Where the rounding unit actually bites

The 5,120-byte unit is benign for anything log-shaped and brutal for anything event-shaped. A 2 KB application log line rounds to one unit and wastes 60% of it. A 100-byte IoT reading rounds to the same unit and wastes 98%.

Which means the aggregation you do before Firehose is worth more than any compression you do inside it, and PutRecordBatch is the lever — it “can take up to 500 records per call or 4 MiB per call, whichever is smaller”. That is batching for the API quota; packing several events into one record is what changes the bill.

There is a second, quieter version of the same trap on the far side of the pipeline: “If the increased quota is much higher than the running traffic, it causes small delivery batches to destinations. This is inefficient and can result in higher costs at the destination services.” Raising a Firehose quota speculatively makes your S3 object count go up and your query performance go down — the small-files problem #31 was about, arriving from a quota setting.

Why This Architecture Holds Up

Chaining the two removes a quota entirely

Direct PUT into Firehose is capped: “500,000 records/second, 2,000 requests/second, and 5 MiB/second” in the three largest Regions, and “100,000 records/second, 1,000 requests/second, and 1 MiB/second” elsewhere — a fivefold regional difference worth knowing before you design for one Region and deploy in another.

Put a data stream in front and the cap disappears: “When Kinesis Data Streams is configured as the data source, this quota doesn't apply, and Amazon Data Firehose scales up and down with no limit.”

That is the argument for the two-service pipeline rather than the one-service one, and it is a capacity argument rather than a feature argument. You are buying the stream's elasticity, and paying for the stream.

The buffer the managed path gives you for free, and its limit

“Each Firehose stream stores data records for up to 24 hours in case the delivery destination is unavailable and if the source is DirectPut.” Twenty-four hours of destination outage absorbed with no work from you. But read the condition: if the source is DirectPut. Chain a data stream in front β€” the thing the previous section recommends β€” and that guarantee moves: “If the source is Kinesis Data Streams (KDS) and the destination is unavailable, then the data will be retained based on your KDS configuration.” So the fix for the throughput quota changes who owns the retention window, and a 24-hour default on the stream is the number to check.

Consumer count is a pricing dimension, not just a limit

Enhanced Fan-Out “supports adding up to 20 consumer applications” on the standard modes. On-demand Advantage raises that to “up to 50 consumers per stream” and removes the surcharge: “Enhanced fan-out data retrievals also do not have a price premium compared to the standard data retrievals in this mode.”

So for a stream fanning out to many independent consumers, the commitment can pay for itself twice over — on the 60% discount and on the removed EFO premium. That is the case where Advantage is genuinely the right answer rather than a trap, and AWS says so: it suits accounts that “need many fan-out consumers, or operate with hundreds of data streams.”

Mode switching is reversible, within limits

“For each data stream in your AWS account, you can switch between the on-demand and provisioned modes twice within 24 hours… Switching between modes doesn't cause any disruptions to your applications.”

Twice a day, no disruption, is unusually generous — most of the decisions in this series have been one-way. The exception is the account-level one: On-demand Advantage needs 24 hours before it can be turned off, and “you must first remove any warm throughput” from the streams that carry it.

Key Architecture Decisions

Decision Choice Reasoning
Cost optimisation target Record count, not payload size Billing rounds each record up to 5,120 bytes.
Sub-minute latency Data Streams with your own consumer Firehose's buffer interval floor is 60 seconds.
60 seconds to 15 minutes Firehose, buffer tuned The managed path, and the whole documented interval range.
On-demand Advantage Only above 25 MiB/s sustained, both ways Shortfall is billed, and the mode locks for 24 hours.
Many fan-out consumers On-demand Advantage 50 consumers instead of 20, and no EFO price premium.
Skewed partition keys Provisioned, not on-demand On-demand splits shards but does not isolate hot hash keys.
Above the Direct PUT quota Chain a data stream in front The Direct PUT cap stops applying entirely.
Destination outage tolerance Check which source you use 24 hours is the DirectPut guarantee; with KDS it follows your stream retention.

The two numbers that settle it

Average record size, and required latency in seconds. Those two place you on the curve before any discussion of volume or service preference.

Under 5 KB per record, your bill is a function of how many records you emit, so aggregation upstream is the highest-leverage change available — and it is free. Under 60 seconds of required latency, the managed path is not an option at any price. Everything else is tuning.

Closing Thought

Cost-lens posts usually end with a break-even. This one ends with a rounding unit, because that is where the money actually goes: 5,120 bytes per record whether the record is 5,000 bytes or fifty.

It is a good example of a cost structure that does not reward the intuitive optimisation. Compressing payloads, trimming fields, shortening keys — all of that is effort spent below the rounding unit and therefore invisible on the invoice. Emitting half as many records is the change that matters, and it is usually a batching change in the producer rather than anything in the pipeline.

The latency floor has the same quality. Sixty seconds is not a performance characteristic to be tuned around; it is the boundary between a service you configure and a service you operate. Knowing which side of it you are on is a one-sentence answer, and it decides most of the rest.

Next in this series

Data & Analytics — the next unwritten roadmap entry. Items 21 to 33 are already published as posts #30 through #60, so the series continues past this block rather than revisiting it.

Comments

How was your experience?
Your feedback helps improve this site.
PoorExcellent