Business Challenge
#16 answered the placement question — when a value belongs in Secrets Manager rather than Parameter Store — and concluded that rotation is the capability you cannot retrofit cheaply. This is what you bought.
Four steps, invoked separately:
“During rotation, Secrets Manager calls the same function several times, each time with
different parameters”, with Step set to
“create_secret, set_secret,
test_secret, or finish_secret”.
The steps are not the interesting part. The labels they move are.
AWSPENDING label stops all future rotations
The paragraph that should be on a wall somewhere:
“When rotation is successful, the AWSPENDING staging label
might be attached to the same version as the AWSCURRENT version, or it
might not be attached to any version. If the AWSPENDING staging
label is present but not attached to the same version as AWSCURRENT, then
any later invocation of rotation assumes that a previous rotation request is still in progress and
returns an error.”
Read what that describes. Nothing has expired. No credential is wrong. The secret is serving a valid
AWSCURRENT version and applications are fine. And rotation has stopped
— not failed once, stopped — because a label sits on a version it should not be on, and
every subsequent attempt interprets it as a rotation still running.
It gets quieter still when the rotation that left it behind had already failed:
“When rotation is unsuccessful, the AWSPENDING staging label
might be attached to an empty secret version.” A label on an empty
version, blocking a schedule, on a secret whose credentials work.
Alarm on rotation not happening, not on rotation failing. The state to watch is VersionIdsToStages — AWS suggests “call describe-secret and look at VersionIdsToStages” — and the condition is AWSPENDING present on a version that is not AWSCURRENT.
“If any rotation step fails, Secrets Manager retries the entire rotation process multiple times during the open rotation windows on the secret.”
So a function that fails in test_secret will be called again from
create_secret. Every step has to tolerate being run against state a
previous attempt already created, which is why the design puts the new value in a label rather than in
a variable:
“Storing the new secret value in AWSPENDING helps ensure
idempotency. If rotation fails for any reason, you can refer to that secret value in subsequent
calls.”
That is the whole reason create_secret begins by looking for a secret that
might already exist. The idempotency is not a nicety bolted on; it is the contract.
Write every step to be re-runnable from a partial previous attempt, and test by failing the function deliberately at each of the four steps.
AWS's own characterisation: “The rotation function is a privileged deputy that has the authorization to access and modify customer credentials in both the Secrets Manager secret and the target resource.”
And three checks it advises before set_secret changes anything:
“Check that the credential in the AWSCURRENT version of the
secret is valid. If the AWSCURRENT credential isn't valid, abandon the
rotation attempt… Check that the AWSCURRENT and
AWSPENDING secret values are for the same resource… Check that the
destination service resource is the same. For a database, check that the
AWSCURRENT and AWSPENDING host names are the
same.”
Those exist because anyone who can write to the secret can redirect the deputy. Edit
AWSPENDING to name a different host and an unguarded function will
obligingly set a password you chose on a database you chose, using credentials it holds and you do
not.
If you wrote the function yourself, implement all three checks. AWS also notes that “Secrets Manager only allows a Lambda rotation function to rotate the secret directly. The rotation function can't call another Lambda function to rotate the secret”, so the deputy cannot be chained.
Architecture
Three labels and four steps. Each step owns one transition.
What each step owns
create_secret — checks for an existing version first,
then “generates a new secret value with
get_random_password… Next it calls
put_secret_value to store it with the staging label
AWSPENDING.” The new value exists only in the label from here on.
set_secret —
“changes the credential in the database or service to match the new secret value in the
AWSPENDING version of the secret.” This is the only step that
touches anything outside Secrets Manager, and it is the step the confused-deputy checks guard.
test_secret —
“tests the AWSPENDING version of the secret by using it to access
the database or service”, and note what the supplied templates settle for:
“Rotation functions based on Rotation function templates test the new secret by using
read access.” A credential that can read but has lost write permission passes
this step.
finish_secret — one call does three things:
“moves the label AWSCURRENT from the previous secret version to this
version, which also removes the AWSPENDING label in the same API
call. Secrets Manager adds the AWSPREVIOUS staging label to the
previous version, so that you retain the last known good version of the secret.”
One call does all of it, which is the design. The API reference describes the mechanism rather than
claiming instantaneity:
“Each staging label can be attached to only one version at a time. To add a staging label to a
version when it is already attached to another version, Secrets Manager first removes it from the other
version first and then attaches it to this one… Whenever you move
AWSCURRENT, Secrets Manager automatically moves the label
AWSPREVIOUS to the version that AWSCURRENT was
removed from.”
So what you get from a single call is that the promotion, the marker removal and the retention of the old value are one request rather than three — not that no intermediate state exists inside it. The window between the database changing and the secret changing is a different thing entirely, and it is real; it is the subject of the single-user strategy below.
One consequence worth knowing before you prune versions: “If this action results in the last label being removed from a version, then the version is considered to be 'deprecated' and can be deleted by Secrets Manager.” A version with no labels is not retained on your behalf.
AWSPENDING, and they apply at different moments
Inside a rotation, AWS advises against it: “Don't remove AWSPENDING before this point, and don't remove it by using a separate API call, because that can indicate to Secrets Manager that the rotation did not complete successfully.” That is addressed to your function, mid-rotation — the removal is supposed to happen as a side effect of promoting AWSCURRENT, not as its own step.
Afterwards, when rotation is wedged, clearing it is the documented remedy. Among the general troubleshooting steps: “Clear pending rotations: If rotation consistently fails, clear the AWSPENDING staging label and retry rotation.” The same action is wrong as part of a rotation and right as recovery from one, which is worth reading twice before an incident rather than during one.
The mechanism is the same API with one parameter left out: UpdateSecretVersionStage with VersionStage set to AWSPENDING and RemoveFromVersionId naming the stuck version — “To remove a label from a version, then do not specify this parameter”, meaning MoveToVersionId. And it is strict about the version you name: “If the label is attached and you either do not specify this parameter, or the version ID does not match, then the operation fails.”
Cross-account rotation needs a parameter most functions ignore
RotationToken is in the invocation payload and is
“Required for secret rotation using an assumed role or cross-account rotation, in which you
rotate a secret in one account by using a Lambda rotation function in another account. In both cases, the
rotation function assumes an IAM role to call Secrets Manager and then Secrets Manager uses the rotation
token to validate the IAM role identity.”
The symptom of omitting it is specific, and it names the wedge from challenge 1:
“Pending secret version… was not created by Lambda… Remove the
AWSPENDING staging label and restart rotation” — at which
point AWS's instruction is that “you need to update your Lambda function to use the
RotationToken parameter.” A function copied from a same-account
example will fail this way every time, and each failure leaves a label behind.
Why This Architecture Holds Up
The two strategies, and what each actually trades
Single user is AWS's default recommendation: “This is the simplest rotation strategy, and it is appropriate for most use cases. In particular, we recommend you use this strategy for credentials for one-time (ad hoc) or interactive users.”
It is not seamless, and AWS says so plainly:
“open database connections are not dropped. While rotation is happening, there is a short period
of time between when the password in the database changes and when the secret is updated. During this
time, there is a low risk of the database denying calls that use the rotated
credentials.” That window is set_secret finishing before
finish_secret runs — the database has moved and the secret has not.
Alternating users closes that window, and AWS gives it two distinct uses:
“This strategy is appropriate for databases with permission models where one role owns the
database tables and a second role has permission to access the database tables. It is also
appropriate for applications that require high availability. If an application retrieves the secret during
rotation, the application still gets a valid set of credentials. After rotation, both
user and user_clone credentials are
valid.”
The first use is a schema-ownership pattern and the second is an availability one; they are independent reasons to pick it, and only the second is about rotation windows at all.
What it costs is a second privileged secret —
“because most users don't have permission to clone themselves, you must provide the credentials
for a superuser in another secret” — and a maintenance
obligation with no enforcement behind it.
“Secrets Manager creates the cloned user with the same permissions as the original user. If you change the original user's permissions after the clone is created, you must also change the cloned user's permissions.” So alternating-users rotation introduces a second database user whose grants drift from the first the moment anybody runs a GRANT, and the drift surfaces on whichever rotation happens to land on the stale user — which is every other one. Nothing compares them. If you adopt this strategy, the comparison is a job you now own.
AWS also points the other way for a specific case: “We recommend using the single-user rotation strategy when cloned users in your database don't have the same permissions as the original user.”
Concurrency has a floor, and the symptom looks like a logic bug
“Setting the provisioned concurrency parameter to a value lower than 10 can cause throttling due to insufficient execution threads for the Lambda function.” And for reserved concurrency, the troubleshooting guidance is to verify it “is not set too low (for example, 1)” and “If using reserved concurrency, set it to at least 10”.
The symptom is worth memorising because it does not look like throttling: “intermittent secret rotation failures with your Lambda function getting stuck in a loop of sets, for example between CreateSecret and SetSecret”. Anybody debugging that will read their own state machine first, which is the wrong place.
It is also a plausible default to have set by accident. A rotation function is low-volume, so
reserved_concurrent_executions = 1 looks like good hygiene in a Terraform
module, and AWS's advice for provisioned concurrency is blunter: “Don't set the provisioned
concurrency parameter explicitly (for example, in Terraform).”
Two smaller things that bite
A manual update counts as a rotation. “If you also manually update your secret value while automatic rotation is set up, then Secrets Manager considers that a valid rotation when it calculates the next rotation date.” So an emergency manual password change pushes the schedule out, and the age of a credential is not what the schedule implies.
And touching the secret during rotation races it. The troubleshooting guidance is to “Avoid making mutating API calls on the secret during Lambda rotation” and to “Ensure there's no race condition between RotateSecret and PutSecretValue calls” — which is exactly what a pipeline that writes secrets on deploy will do, on whichever deploy coincides with a rotation window.
The MySQL username ceiling, which is arithmetic rather than a limit
“For Amazon RDS MySQL, in alternating users rotation, Secrets Manager creates a cloned user with a name no longer than 16 characters. You can modify the rotation function to allow longer usernames. MySQL version 5.7 and higher supports usernames up to 32 characters, however Secrets Manager appends ‘_clone’ (six characters) to the end of the username…”
So the original name has to fit in 26 characters for the clone to fit in 32, and the default function caps it at 16 regardless. A service account named for a team and an environment reaches 16 easily, and the failure arrives on the first rotation rather than at creation.
Key Architecture Decisions
| Decision | Choice | Reasoning |
|---|---|---|
| What to alarm on | AWSPENDING present on a version that is not AWSCURRENT |
That state makes later rotations return an error instead of running, while credentials keep working. |
| Rotation strategy | Single user for most cases, per AWS's own recommendation | Described as the simplest and appropriate for most use cases; AWS recommends it for ad hoc or interactive users. |
| Strategy where a request must not be denied | Alternating users, accepting the clone-permissions obligation | An application retrieving the secret mid-rotation still gets a valid set of credentials. |
| Clone permission drift | A scheduled comparison of user and user_clone grants |
The clone is created with the original's permissions, and changes to the original must be applied by hand. |
| Step idempotency | All four steps re-runnable; test by failing each one | A failure retries the entire rotation process, not the failed step. |
| Custom rotation functions | Implement the three confused-deputy checks | The function is a privileged deputy over both the secret and the target resource. |
test_secret depth |
Exercise the access the application needs, not just read | Template-based functions test using read access. |
| Lambda concurrency | At least 10 reserved; don't set provisioned explicitly | Below 10 can cause throttling, and the symptom is a loop between CreateSecret and SetSecret. |
| Cross-account or assumed-role rotation | Pass RotationToken through to put_secret_value |
Required in those cases; omitting it fails and leaves AWSPENDING behind each time. |
| MySQL usernames under alternating users | 26 characters or fewer, and raise the function's 16-character cap deliberately | Secrets Manager appends a six-character _clone suffix within MySQL's 32-character limit. |
| Pipelines that write secrets | Keep them out of rotation windows | AWS advises avoiding mutating API calls on the secret during rotation. |
The audit worth running
describe-secret across every rotating secret, reading
VersionIdsToStages, and one question: is
AWSPENDING attached to anything that is not the
AWSCURRENT version? Each hit is a secret whose rotation has stopped and whose
monitoring says nothing, because the credential in use is valid.
Then compare LastRotatedDate against the configured schedule. A secret that
reports rotation enabled and last rotated months ago is the same condition seen from the other side, and
it is the one a compliance review will eventually find for you.
Closing Thought
The four-step contract is a good design, and finish_secret is the best part
of it: one call promotes the new version, clears the in-progress marker and preserves the old value,
rather than three calls that can each fail separately. Almost everything that goes wrong is a consequence
of that marker being durable. It has to be, for the retry to
work — and because it is durable, a rotation that dies before
finish_secret leaves a flag that reads as "still running" to everything that
comes after.
That makes this a monitoring problem more than an engineering one. The thing to watch is not whether rotation failed, because it will have reported that once, months ago, into a log nobody reads. The thing to watch is whether rotation is happening — and the signal is a label on the wrong version, which no dashboard shows you by default.
Which is the same shape as the rest of this block, one more time. #36 was a permission absent rather than denied. #72 was a resource quietly accumulating against a ceiling. This one is a schedule quietly not running. In each case the system is behaving exactly as documented, the credential or the request still works, and the absence is what nobody is looking at.
Security & Identity — GuardDuty finding triage: what a finding's severity actually encodes, why the same finding type arrives with different severities, and how to tell a finding that needs a human from one that needs a suppression rule.
Official AWS Reference
- Lambda rotation functions — the four steps, the staging-label transitions, the privileged-deputy checks, and the wedge condition
- Rotation by Lambda function — the invocation payload, RotationToken, and the whole-rotation retry
- Lambda function rotation strategies — single user against alternating users, and the clone-permissions obligation
- UpdateSecretVersionStage — how a staging label moves, the AWSPREVIOUS side effect, and how to clear a label from a version
- Troubleshoot rotation — clearing a stuck AWSPENDING label, the cross-account RotationToken failure, and the concurrency floor
Comments