Business Challenge
#43 and #44 federated with identity providers you do not control. This one Google runs for you, which removes three classes of mistake and introduces a different one.
It covers more than that. The namespace identifier selects all Pods in a namespace, regardless of service account or cluster — and the pool is per-project, so a namespace called prod in any cluster in that project matches.
Grant to a Kubernetes ServiceAccount rather than a namespace unless you genuinely mean every Pod in every cluster. If you need cluster scope, there is a separate identifier for it.
Enabling Workload Identity Federation for GKE on the cluster does not grant your workloads any additional IAM permissions. It creates the identities; the allow policies are still yours to write.
Correct approachTreat enabling as plumbing and granting as a separate, deliberate step — which is the same separation #39 drew between an account existing and an account being permitted.
Not entirely. Even with Workload Identity Federation for GKE configured, GKE still uses the configured IAM service account for the node pool to pull container images from the image registry.
Correct approach
Reduce the node service account to image-pull rights rather than removing it. An empty one produces ImagePullBackOff, which looks nothing like an IAM problem.
Not the ones on the host network. Pods with hostNetwork: true do not use Workload Identity Federation for GKE — GKE automatically routes requests from these Pods to the Compute Engine metadata server, so they authenticate as the node.
Audit for hostNetwork: true alongside your IAM review. One field in a Pod spec changes which identity the workload has.
Architecture
In most cases, Workload Identity Federation for GKE is the recommended way for workloads on GKE to reach Google Cloud services. The reason it is simpler than #43 and #44 is one sentence: in GKE, Google Cloud manages the workload identity pool and provider for you, and does not require an external identity provider.
What enabling it creates
- A fixed workload identity pool for the project, named
PROJECT_ID.svc.id.goog. Note the scope in the name: project, not cluster. - The cluster registered as an identity provider inside that pool.
- The GKE metadata server, which intercepts credential requests from workloads, on every node.
And it persists: GKE does not delete this workload identity pool even if you delete all of the clusters in your project. The pool outlives the clusters, which matters because the grants you wrote against it outlive them too.
The credential flow
Four steps, and the interesting part is who signs what:
- Application Default Credentials asks the Compute Engine metadata server for a token — the workload does nothing special.
- The GKE metadata server intercepts, and asks the Kubernetes API server for a Kubernetes ServiceAccount token that identifies the requesting workload. That JWT is signed by the API server.
- The metadata server uses Security Token Service to exchange the JWT for a short-lived federated access token.
- The workload uses it, with whatever the principal identifier can reach.
So the trust chain ends at the Kubernetes API server's signature. That is the thing GKE can vouch for and an external provider cannot, and it is why none of #44's issuer problems arise here.
Four identifiers, and one of them is wider than it reads
| Selects | Scope |
|---|---|
A ServiceAccount by name — /subject/ns/NAMESPACE/sa/SERVICEACCOUNT |
All Pods using that ServiceAccount. The everyday choice. |
A ServiceAccount by UID — /kubernetes.serviceaccount.uid/UID |
The same, but tied to that exact object rather than the name. |
A namespace — /namespace/NAMESPACE |
All Pods in a namespace, regardless of service account or cluster. |
A cluster — /kubernetes.cluster/... |
All Pods in one named cluster. |
Read the third row against the first fact in this post. The pool is PROJECT_ID.svc.id.goog — one pool for the project, with every cluster registered inside it. So a grant to /namespace/prod does not mean "the prod namespace in this cluster". It means the prod namespace anywhere in the project: the staging cluster that also has a prod namespace, the cluster somebody creates next quarter, the cluster restored from a backup for an incident. Kubernetes namespace names are conventional and repeated by design, which is exactly what makes this dangerous — the names collide on purpose. If you want cluster scope, the fourth identifier exists and names the cluster explicitly.
Name or UID, for the third time in this block
GKE offers both, and the choice is the one from #43 and #44 in new clothes. Select the ServiceAccount by name and a deleted-and-recreated ServiceAccount with the same name in the same namespace inherits the grant. Select the ServiceAccount by UID and it does not, because the UID is a different object.
Which is right depends on what you mean. A grant to ns/prod/sa/database-reader is a statement about a role in the system that should survive a redeploy — and redeploys recreate ServiceAccounts routinely, so the name is usually correct here. That is the opposite conclusion from #44, where the GitHub organisation name was the risky choice. The difference is who controls the namespace: your own cluster manifests, versus a public registry where somebody else can claim a released name.
Why This Architecture Holds Up
Both are quiet, and both produce a working system that is not the one you configured.
Host network Pods. A Pod with hostNetwork: true does not use Workload Identity Federation for GKE at all — GKE automatically routes requests from these Pods to the Compute Engine metadata server. It therefore authenticates as the node's service account, with whatever that can do. Nothing errors, nothing warns; the Pod simply has a different identity from its neighbours in the same namespace. This is #41's attached-service-account problem in a new location: a workload inside the machine becomes the machine.
The first few seconds. The GKE metadata server takes a few seconds to start accepting requests on a newly created Pod, so attempts to authenticate within the first few seconds of a Pod's life might fail. That is a startup race, not a permissions problem, and it presents as an intermittent auth failure on cold starts — the kind of thing that gets diagnosed as a flaky IAM grant and papered over with a retry nobody documents.
There is also a capacity ceiling worth knowing before it surprises you: the Exchange Token API in Security Token Service has a quota limit of 6,000 requests per minute. Every Pod authenticating is an exchange, so a large cluster with short-lived Pods and no token caching can reach that, and the failure will look like an identity problem rather than a quota one.
What to do with this
- Prefer the ServiceAccount identifier to the namespace one. The namespace identifier is a project-wide grant wearing a cluster-shaped name.
- Use the cluster identifier when you mean a cluster. It exists precisely so the namespace one does not have to be misused for it.
- Audit Pod specs for
hostNetwork: trueas part of an access review, not a networking one. - Keep the node service account, minimised. Image pulls still use it.
- Expect a cold-start race and let the client library retry rather than treating it as a grant failure.
- Remember the pool outlives the clusters. Deleting every cluster leaves the pool and the grants written against it in place.
Key Architecture Decisions
| Decision | Choose this | Because |
|---|---|---|
| How a GKE workload reaches Google Cloud | Workload Identity Federation for GKE | It is the recommended way in most cases. |
| Granting to a workload | The ServiceAccount identifier | The namespace one ignores both service account and cluster. |
| Granting to a whole cluster | The cluster identifier | It names the cluster; the namespace one does not. |
| Name or UID for a ServiceAccount | Name, usually | Redeploys recreate ServiceAccounts, and you control the namespace. |
| The node service account | Keep it, minimised | Image pulls still use it. |
| A Pod on the host network | Expect the node identity | Those requests go to the Compute Engine metadata server. |
| Intermittent auth failure at startup | Retry, do not re-grant | The metadata server takes a few seconds on a new Pod. |
| A service with federation limitations | Service account impersonation | It is the documented fallback for those services. |
| Very high Pod churn | Watch the STS quota | 6,000 token exchanges per minute is the limit. |
Closing Thought
Three posts of federation have moved the same risk around rather than removing it. With AWS, Google verifies the credential by asking AWS, and the boundary comes free. With GitHub, the issuer proves almost nothing and the boundary is a CEL expression you have to remember. With GKE, Google Cloud manages the pool and provider itself, the API server's signature is the trust, and there is nothing to map or condition — so the risk reappears as a naming question: which of four identifiers did you use, and how wide is it really.
The namespace identifier is the one to be careful with, and it is careless in a specific way. It reads like a Kubernetes concept and behaves like a project-wide one, because the pool it lives in belongs to the project and every cluster registers into the same pool. Kubernetes encourages reusing namespace names across clusters — that is the whole point of namespaces — so the identifier is at its most dangerous in exactly the estates that are organised well. That has been the recurring shape of this block: the better the convention, the wider the grant that follows it.
#46 goes back to plain VMs for the mechanism GKE has been quietly falling back to all post: attached service accounts on Compute Engine — what the metadata server hands out, and why scopes are still there underneath IAM.
Comments