Business Challenge
#5 established the rule that governs KMS access: the key policy is the authoritative control, and IAM grants nothing the key policy has not also permitted. This post is about the third thing in that evaluation, which #5 mentioned in a paragraph and which behaves unlike either of the other two.
“A grant is a policy instrument that allows AWS principals or AWS service principals to use KMS keys in cryptographic operations… When authorizing access to a KMS key, grants are considered along with key policies and IAM policies.”
"Policy instrument" invites you to file it with the other two. The differences are where the problems live.
“Each customer managed key can have up to 50,000 grants, including the grants created by AWS services that are integrated with AWS KMS.”
Fifty thousand sounds like a number you will never meet, right up until you read what it actually bounds: “One effect of this quota is that you cannot perform more than 50,000 grant-authorized operations that use the same KMS key at the same time. After you reach the quota, you can create new grants on the KMS key only when an active grant is retired or revoked.”
And AWS's own worked example makes it concrete: “when you attach an Amazon Elastic Block Store (Amazon EBS) volume to an Amazon Elastic Compute Cloud (Amazon EC2) instance… Amazon EBS creates a grant for each volume. Therefore, if all of your Amazon EBS volumes use the same KMS key, you cannot attach more than 50,000 volumes at one time.”
So the one-key-per-environment design that looks tidy on a diagram is a fleet-wide ceiling on simultaneous attachments, and you do not create any of those grants yourself.
FixPartition keys by workload rather than by environment, and treat the grant count on a shared key as a capacity metric.
Revoking behaves normally — “The RevokeGrant API can be called by any principal with
kms:RevokeGrant permission. This permission is included in the standard
permissions given to key administrators.”
Retiring does not. “The grant determines who can retire it. This design allows you to control the lifecycle of a grant without changing key policies or IAM policies.” Four sentences then dismantle the permission you would reach for:
“There is a kms:RetireGrant permission that can be used in IAM
policies, but it has limited utility. Principals specified in the grant can retire a
grant without the kms:RetireGrant permission. The
kms:RetireGrant permission alone does not allow
principals to retire a grant. The kms:RetireGrant permission is
not effective in a key policy or resource control policy.”
Read that as a set. The permission does not confer the ability; the ability does not require the
permission; and the key policy — the authoritative control for everything else on this key
— cannot express it at all. What remains is one use:
“To deny permission to retire a grant, you can use a Deny action
with the kms:RetireGrant permission in your IAM policies.”
Decide retirement at CreateGrant time, via RetiringPrincipal. It cannot be corrected from a policy afterwards.
kms:CreateGrant lets you hand out permissions you do not have
AWS puts the comparison first:
“Permission to create grants has security implications, much like allowing the
kms:PutKeyPolicy permission to set policies.”
Then the mechanism:
“Principals who get kms:CreateGrant permission from a policy can
create grants for any grant operation on the KMS key. These principals are
not required to have the permission that they are granting on the key.”
So a role with kms:CreateGrant and no
kms:Decrypt can mint a grant conferring
Decrypt — on a principal in its own account, another account, or
another organisation. The escalation needs no policy change and leaves the key policy looking exactly
as it did.
Treat kms:CreateGrant in a key policy as equivalent to kms:PutKeyPolicy, and constrain it with conditions rather than granting it bare.
Architecture
The lifecycle is create, use, delete — with no step in the middle.
There is no modify operation
The grant operations are enumerated, and the list is the whole surface: the cryptographic operations, plus
CreateGrant, DescribeKey,
GetPublicKey and RetireGrant. Nothing modifies a
grant. “An authorized principal can delete the grant (retire or revoke it). Deleting a grant
eliminates all permissions that the grant allowed.”
Tightening a grant's constraint therefore means creating a second grant and deleting the first, in that order if you care about availability and the opposite order if you care about exposure. A key policy statement you can edit in place; a grant you can only replace.
Eventual consistency runs in both directions, and only one has a workaround
Creating: “there might be a brief delay, usually less than five minutes, until the grant is available throughout AWS KMS.” The fix is the grant token — “a unique, nonsecret, variable-length, base64-encoded string that represents a grant”, and “because the token value is a hash digest, it doesn't reveal any details about the grant.”
Deleting: “If you retire or revoke a grant, the grantee principal might still be able to use its permissions for a brief period until the grant is fully deleted.”
That asymmetry matters during an incident. Revoking a grant is not a stop button; it is a request that
becomes true within a few minutes. And the token workaround does not help here, because
“you cannot use a grant token to revoke a grant” — it works
for RetireGrant only.
“CreateGrant is the only operation that returns a grant token. You cannot get a grant token from any other AWS KMS operation or from the CloudTrail log event for the CreateGrant operation. The ListGrants and ListRetirableGrants operations return the grant ID, but not a grant token.” So if your code discards the CreateGrant response, the only way back to immediate usability is to create another grant β which costs another slot against the 50,000. And note what AWS says about precedence: “Grant tokens supersede the validity of the grant until all endpoints in the service have been updated with the new grant state.”
The retry trap, and why it meets the quota
Name is optional, and its absence is not neutral:
“When this value is absent, all CreateGrant requests result
in a new grant with a unique GrantId even if all the supplied parameters are
identical. This can result in unintended duplicates when you retry the
CreateGrant request.”
With Name supplied the call becomes idempotent:
“you can retry a CreateGrant request with identical parameters; if
the grant already exists, the original GrantId is returned without creating a
new grant.”
Put the two findings together and the failure is specific. A client with retry logic, no
Name, and a shared key does not fail loudly — it succeeds, repeatedly,
filling slots in a 50,000-wide ceiling that is also the concurrency limit for every EBS attachment on that
key. The error, when it arrives, is LimitExceededException on an unrelated
workload.
Note the one thing Name does not make idempotent:
“the returned grant token is unique with every CreateGrant request,
even when a duplicate GrantId is returned. All grant tokens for the same grant
ID can be used interchangeably.”
Why This Architecture Holds Up
Delegation through a grant is bounded; delegation through a policy is not
This is the part of the design that is genuinely careful, and it is easy to miss because it sits one
bullet below the escalation. Policy-derived CreateGrant is unbounded. But:
“Principals can also get permission to create grants from a grant. These
principals can only delegate the permissions that they were granted, even if they have
other permissions from a policy.”
So a grant chain cannot widen. Each link can pass on a subset of what it holds and no more, regardless of
what its IAM policy says elsewhere. The same guard appears on constraints:
“If a grant with an encryption context grant constraint includes the
CreateGrant operation, the constraint requires that any grants created with
the CreateGrant permission have an equally strict or
stricter encryption context constraint.”
That is the right property, and it means the safe way to hand out grant-creation is through a grant rather than through the key policy — the opposite of the instinct, and the opposite of how most Terraform gets written.
Constraints are real, and they carve out the same operation twice
Two kinds. Encryption context — “allow the permissions in the grant only when the
encryption context in the request matches (EncryptionContextEquals) or
includes (EncryptionContextSubset) the encryption context specified in the
constraint”, capped at “up to 8 encryption context pairs” with
“the encryption context value in each constraint cannot exceed 384 characters”, and
unavailable on asymmetric or HMAC keys because those operations have no encryption context.
And SourceArn, which “allows the permissions in the grant only when
the request is made on behalf of a specific AWS resource” and is
“supported on grants for all types of KMS keys”.
Both have the same hole, stated separately for each:
“Grants with encryption context grant constraints can include the
DescribeKey and RetireGrant operations, but
the constraint doesn't apply to these operations”, and for
SourceArn, “it does not apply to
RetireGrant operation.”
RetireGrant is outside the constraint system, just as it is outside the key
policy. The one operation that ends a grant is the one least governable by anything written around it.
On both Constraints and Name: “Do not include confidential or sensitive information in this field. This field may be displayed in plaintext in CloudTrail logs and other output.” Encryption context is a natural place to put tenant identifiers, request identifiers or object keys β and a grant constraint containing them is published into CloudTrail. The encryption context on an individual request is a separate matter; this is specifically about the constraint baked into the grant, which is long-lived and readable by anyone with log access.
Service-created grants are the majority, and they are well behaved
“Grants are commonly used by AWS services that integrate with AWS KMS to encrypt your data at rest. The service creates a grant on behalf of a user in the account, uses its permissions, and retires the grant as soon as its task is complete.”
That is why the quota is usually invisible: the services clean up after themselves. The grants that
accumulate are the ones nobody retires — which, given that
RetiringPrincipal is optional and
kms:RetireGrant is nearly inert, are the ones created without thinking about
who would ever delete them.
When a service principal is the grantee, AWS requires the scoping up front:
“When you specify a GranteeServicePrincipal, you must also specify a
SourceArn grant constraint. In addition, you must specify either a
RetiringPrincipal or a
RetiringServicePrincipal.” Both of the things that are optional for
an IAM grantee are mandatory for a service grantee — which reads like AWS encoding the lesson of
the looser path into the newer one.
What a grant cannot be pointed at
“the grantee principal cannot be a service principal, an IAM group, or an AWS organization.” No group means no indirection: a grant names an identity, so a team of ten is ten grants or one shared role. And “Each grant allows access to exactly one KMS key” — there is no grant that spans keys, so a workload touching five keys needs five grants per principal.
Those two together are why grant counts grow multiplicatively rather than additively, and why the 50,000 ceiling is closer than the number suggests.
Key Architecture Decisions
| Decision | Choice | Reasoning |
|---|---|---|
| Key partitioning | Per workload, not per environment | The 50,000 grants per key is a concurrency ceiling on grant-authorized operations, including every EBS attachment. |
Every CreateGrant call |
Always supply Name |
Without it, an identical retry creates a duplicate grant and consumes another slot. |
| Retirement | Set RetiringPrincipal at creation |
The grant determines who can retire it, and no policy can add that later. |
| Blocking retirement | Deny on kms:RetireGrant in an IAM policy |
It is the only place that permission is effective — not a key policy, not an RCP. |
| Emergency removal | RevokeGrant, and expect a propagation window |
Permissions may survive briefly after deletion, and a grant token cannot revoke. |
kms:CreateGrant in a key policy |
Treat as equivalent to kms:PutKeyPolicy |
Holders can grant operations they do not themselves hold, to any account. |
| Delegating grant creation | Through a grant, not through the key policy | Grant-derived CreateGrant can only pass on what it holds; policy-derived cannot be bounded that way. |
| Grant constraints | SourceArn for any key type; encryption context for symmetric |
Encryption context constraints are unavailable on asymmetric and HMAC keys. |
| Constraint contents | No tenant or request identifiers | Constraints and names may appear in plaintext in CloudTrail. |
| Grant tokens | Capture from the CreateGrant response and use immediately |
It is the only source; nothing else, including CloudTrail, returns one. |
The audit worth running
ListGrants on every customer managed key, and two questions of the output.
How many — because that number is capacity, and because a count climbing without a
matching workload is a retry loop with no Name. And how many have no
retiring principal — because those are the grants nobody is designated to remove, on which
your only instrument is RevokeGrant by a key administrator.
Then ListRetirableGrants from the other direction, for the grants your
principals can retire in other people's accounts. That is a permission you hold and probably did not
grant yourself.
Closing Thought
The phrase that causes the trouble is "policy instrument". It puts grants in the same mental category as key policies and IAM policies, and then almost everything about handling them differs. A key policy statement is edited in place, governed by the key policy, bounded by the permissions of whoever writes it, and effective immediately. A grant is replaced rather than edited, has a delete rule the key policy cannot express, can confer permissions its creator does not hold, and takes minutes to mean anything — in either direction.
The one that will actually cost somebody a Saturday is the quota, because it is the only one that fails on a workload that did nothing wrong. Fifty thousand reads as a configuration ceiling and behaves as a concurrency ceiling, the grants filling it are created by services on your behalf, and the number is shared by everything that uses the key. #5 argued for key hierarchy on blast-radius grounds. This is the availability argument for the same structure, and it is the one that shows up in an incident rather than in an audit.
So the rule is short. Supply Name on every call, set
RetiringPrincipal at creation because nothing can add it later, and count your
grants — they are a resource, and the thing about resources is that you can run out.
Security & Identity — Secrets Manager rotation: why a rotation function needs two sets of credentials, what happens to a secret stuck between AWSCURRENT and AWSPENDING, and why the four-step contract is the whole design.
Official AWS Reference
- Grants in AWS KMS — the grant concepts, grant operations, grant tokens, eventual consistency, and the warnings about permission to create grants
- CreateGrant — the Name idempotency behaviour, grant constraints and their limits, and the service-principal requirements
- Retiring and revoking grants — who can retire, and why kms:RetireGrant has limited utility
- AWS KMS resource quotas — the 50,000 grants per key ceiling and the EBS attachment consequence
Comments