📋 In This Post
Why — The Problem This Solves
Every week of this lab opens with a cost note, and five of them have claimed $0. Asked to prove one of those figures, the only answer available was to open the billing console and read a summary off the screen.
That is not evidence. It is a number I looked at once. A lab that makes a cost claim in writing every week and cannot query any of them is running on trust, and the whole point of the exercise is to stop doing that.
So the plan was to build the queryable source: export billing data to BigQuery, put a budget on it, and wire the budget to something a machine can read rather than something a person has to notice in an inbox.
The plan was wrong about the premise. The source already existed. It had been running since 1 July, into a dataset in the seed project, and in eleven weeks nobody — me — had ever pointed a query at it. The gap was never missing infrastructure. It was an absent question.
That discovery is worth more than the week's original deliverable, and it changed what the week built: less Terraform, not more.
What You Need to Know — Skills & Tools
Google Cloud
- Cloud Billing — budgets, exports, and an account that sits outside the resource hierarchy entirely
- BigQuery — the export destination, and where the cost question is finally answerable
- Pub/Sub — the programmatic notification path, and the IAM policy Cloud Billing writes to it
- Organization Policy — specifically
iam.allowedPolicyMemberDomains, which is this week's antagonist
Tooling
- Terraform —
hashicorp/google~> 6.0, plushashicorp/timefor a propagation wait that is not optional - bq CLI — the fastest way to find out what an export already holds
- Standard SQL — enough to aggregate a cost table; the query that reframed this week is five lines
- curl and an access token — for reading the raw API error when the provider's version says nothing
Concepts to understand before starting
A billing account is not in the resource hierarchy. This is the fact everything else this week follows from. Projects live in folders, which live in an organization. A billing account sits outside all of it, owned by no project, linked to many. That is why its IAM is not Terraform-managed here, why the Budgets API needs a quota project explicitly nominated, and why the export has no natural home in the hierarchy at all.
Cloud Billing export to BigQuery is console-only. No gcloud
command, no Terraform resource, no public REST endpoint —
/v1/billingAccounts/{id}/exportSettings returns 404. Confirmed
against the provider's own schema, which lists nine google_billing_*
resources and no export among them, and against Google's own billing-dashboard
module, which creates the datasets and then instructs the operator to link them
by hand.
There are several exports and they are not alternatives. Standard usage cost, detailed usage cost, pricing, FOCUS, and committed use discounts. The important relationship: detailed is a strict superset of standard — same fields plus resource-level cost. Pricing is a different thing entirely; it answers what a SKU costs rather than what was spent.
Budgets publish to Pub/Sub several times a day, carrying current spend against the budget — not only when a threshold is crossed. Email fires on breach; the topic is a cost feed. That difference is the reason to bother with a topic at all.
Architecture — How It Fits Together
billing account (outside the hierarchy, owned by no project)
│
├── detailed usage cost export ──────► seed project
│ running since 2026-07-01 billing_export
│ console-only link gcp_billing_export_resource_v1_*
│
├── pricing export ─────────────────► logging project
│ linked 2026-09-27 billing_pricing
│
└── budget "GCP lab guardrail - programmatic" $25/month
thresholds 50% · 90% · 100% actual · 100% forecast
│
└── Pub/Sub topic ───────────► logging project
billing-budget-notifications
publisher: billing-budget-alert@system.gserviceaccount.com
(added by Cloud Billing, not by this configuration)
Two things in that diagram are deliberate and look like mistakes.
The two exports live in different projects. Consolidating them would be tidier and would have cost July and August: re-pointing an export does not move its history, the old table simply stops receiving rows, and a new dataset backfills at most the previous month. The detailed export stays where it already was.
Nothing subscribes to the topic. A subscriber is a decision about what should happen to spend, and inventing one before there is any spend to react to is building a mechanism for an unobserved problem.
How We Built It — Step by Step
1. Read the live state before believing the plan
The week's own configuration file opened with a comment explaining that the lab had no queryable cost source. Before writing anything further, I listed the datasets that actually existed.
$ bq ls --project_id=<seed project>
datasetId
----------------
billing_export
A dataset I had not created, in a project this week was not going to touch,
holding a table named gcp_billing_export_resource_v1_* — the
detailed export's own naming. It had been filling since 1 July.
2. Ask the question nobody had asked
One query against a table that had existed for three months:
SELECT COUNT(*) AS rows_,
MIN(DATE(usage_start_time)) AS first_day,
MAX(DATE(usage_start_time)) AS last_day,
ROUND(SUM(cost), 4) AS total_cost,
COUNT(DISTINCT project.id) AS projects
FROM `<seed>.billing_export.gcp_billing_export_resource_v1_*`
+-------+------------+------------+------------+----------+
| rows_ | first_day | last_day | total_cost | projects |
+-------+------------+------------+------------+----------+
| 2412 | 2026-07-01 | 2026-09-27 | 0.006 | 5 |
+-------+------------+------------+------------+----------+
Every previous week's $0 was true, and had been provable for three months. Six tenths of a cent across five projects and a quarter of a year.
I want to be precise about what was actually wrong here, because "we already had the data" sounds like a small filing error. It is not. The lab had been publishing unverifiable numbers while holding the verification in a table one query away. Infrastructure that nobody queries is indistinguishable from infrastructure that does not exist, and it is worse, because it bills.
3. Delete half the week
The configuration at this point created two datasets: one for the standard export and one for pricing. Google's own table reference settles what that first one would have been worth — the detailed export "includes all of the data fields from the standard usage cost table, along with additional fields that provide resource-level cost data".
So with a detailed export already running, a standard export stores a strict subset of data the account is already keeping, in a second dataset, billed again. The correct number of standard exports to run alongside a detailed one is zero. The dataset was removed from the configuration before it ever received a row.
Pricing survived the cut because it is genuinely additive: it describes what a SKU costs, not what was spent, which is a question the usage export cannot answer at all.
4. The apply, and then the part that would not apply
Everything from here was produced by a Terraform run executing remotely in HCP Terraform, authenticating to Google Cloud through Workload Identity Federation. No service account key exists, and no Google credential is present on the machine that starts the run. Every console screenshot below is evidence of what a run produced, never a thing that was clicked.
The datasets, the topic and the budget all created cleanly. Attaching the topic to the budget did not:
Error 400: Precondition check failed.
Four words. No resource, no permission, no policy, no field. I checked the raw REST response in case the provider was truncating something useful:
{
"code": 400,
"message": "Precondition check failed.",
"status": "FAILED_PRECONDITION"
}
No details array. That message is the entire diagnostic surface.
5. Bisect the error instead of re-reading the documentation
An error naming nothing cannot be researched, only eliminated. Four probes, each removing one candidate:
bogus topic name -> NOT_FOUND so topic resolution works
real topic, budget A -> FAILED_PRECONDITION
real topic, budget B -> FAILED_PRECONDITION so it is not the budget
testIamPermissions -> setIamPolicy HELD so it is not the caller
The first probe matters most and is the least obvious: a deliberately wrong topic name returns a different error. That proves the API resolves the real topic successfully, which eliminates the entire class of "the topic is missing, misnamed, or in the wrong project" explanations in one call.
Two different budgets failing identically eliminates the budget. And the
caller holding pubsub.topics.setIamPolicy — the permission
Google's own documentation says you need — eliminates the caller.
I then ran the whole sequence again as a human organization administrator rather than as the CI service account. Identical. A cause that survives both identities is not about identity at all, and a cause that survives all four probes is not about the resources either. That points at the organization.
6. The cause was a guardrail I built in Week 3
Attaching a Pub/Sub topic to a budget is not a reference. Cloud Billing writes its own Google-owned service agent onto that topic's IAM policy as a publisher, at attach time.
Week 3 enforced iam.allowedPolicyMemberDomains across the
organization, restricted to this Workspace customer ID. A Google-owned principal
is by definition outside it. So the IAM write was refused, and the refusal
surfaced — several layers away, in a different API — as a
precondition failure.
The detail that stings: the console says this in plain words and the API does not. The budget's own notification section carries the sentence, sitting directly above the control it describes:
It may not be possible to add a Pub/Sub topic if it belongs to an organization that has domain restricted sharing enabled.
Same condition, same product, same afternoon. Through the console it is a
sentence naming the exact constraint. Through the API it is
Precondition check failed.
7. Fix it with scope, not surrender
The obvious fix is to relax the constraint. The right fix is to relax it in exactly one place.
resource "google_org_policy_policy" "drs_exception" {
name = "projects/${var.data_project_id}/policies/iam.allowedPolicyMemberDomains"
parent = "projects/${var.data_project_id}"
spec {
inherit_from_parent = false
rules { allow_all = "TRUE" }
}
}
A list constraint has no vocabulary for "Google's own service agents" — it takes customer IDs and principal sets, and the agent's address is not published anywhere. The exception therefore cannot be expressed as membership. It can only be expressed as scope.
One more thing was needed, and it is easy to miss. Organization policy does not propagate instantly — Week 3 measured that on the way in, and the same lag applies on the way out. The budget attach races the relaxation and fails with the identical opaque error, which looks exactly like the diagnosis having been wrong. An explicit wait sits between them:
resource "time_sleep" "drs_propagation" {
depends_on = [google_org_policy_policy.drs_exception]
create_duration = "120s"
}
A retry loop would have been worse than useless here: it cannot distinguish "still propagating" from "wrong diagnosis", so it would have turned a solved problem back into an ambiguous one.
8. Read the service agent back rather than guessing it
An earlier draft of this configuration granted roles/pubsub.publisher
to billing-budgets@system.gserviceaccount.com, inferred from the
shape other Google service agents take. Google's reply was unusually direct:
Error 400: Service account billing-budgets@system.gserviceaccount.com does not exist
The correct move is to grant nothing, let Cloud Billing add its own binding, and then read the topic's IAM policy to find out which identity it used:
roles/pubsub.publisher
serviceAccount:billing-budget-alert@system.gserviceaccount.com
Singular budget, and alert rather than budgets. Close enough to the guess to survive a code review, wrong enough to fail. The address appears in no documentation I found, which makes reading it back the only reliable way to know it — and makes it a poor thing to hardcode in a check. The validation script asserts that something holds publisher and then prints what, rather than matching a name.
Worth noting what the policy does not contain: the project's own editors get no publish rights on that topic. The binding Cloud Billing added is exactly as wide as the integration needs and no wider.
9. The pricing export, and a permission split worth knowing
The last step was enabling the pricing export, and it produced a small, sharp lesson about how Google splits billing permissions.
Configuring any export requires two permissions, and they do not live together:
| Role | getPricing | updateUsageExportSpec |
|---|---|---|
roles/billing.viewer | ✓ | |
roles/billing.admin | ✓ | ✓ |
roles/billing.user | ✓ | |
roles/billing.costsManager | ✓ |
The operating account held user and costsManager
— two roles that both carry the second permission and neither of which
carries the first. The result is a page that renders completely normally and
refuses only at the moment you try to configure anything.
roles/billing.viewer, the least powerful role in the table, is what
closes the gap.
Granting it requires billing.accounts.setIamPolicy, which lives
only in roles/billing.admin — held by the account that created
the billing account. That is an out-of-band grant by design: billing account
bindings are never Terraform-managed in this lab, because the account is not in
the hierarchy the configuration owns.
Challenges — What Actually Went Wrong
1. Four words with no details array
Error 400: Precondition check failed. — naming no
resource, no permission, no policy and no field, with nothing in the REST
response to expand on it.
The technique that worked was elimination, not research: send a request you
know is wrong in a specific way and see whether the error changes. A
deliberately bogus topic name returned NOT_FOUND, which proved in
one call that topic resolution was fine and removed a whole family of
hypotheses.
The general rule: when an error names nothing, stop reading and start varying one thing at a time. Four probes cost ten minutes and produced a certain answer where a morning of documentation produced none.
2. My own guardrail, three weeks later, in a different service
Week 3's domain restricted sharing constraint did exactly what it was built to do. It blocked a Google-owned principal from being added to an IAM policy. That was the correct behaviour and it broke a legitimate integration, in a product that never mentions organization policy, via an error that never mentions IAM.
This is the actual cost of guardrails and it is rarely stated honestly: the bill does not arrive when you build them, it arrives weeks later in an unrelated service, disguised as something else.
What made it findable was that the failure was identical for the CI service account and for a human organization administrator. Identity-shaped problems vary by identity; this one did not.
3. The premise was wrong, and the configuration asserted it confidently
The file began with a comment stating the lab had no queryable cost source. It was well-argued and it was false, and nothing in the repository would ever have contradicted it — the export lived in the console, in a project the week did not touch.
Reading live cloud state is already a standing rule here for policies, because configured and enforced are different. This week extended it: planned and absent are also different. A premise I wrote is not evidence, and it is the part nobody re-checks.
4. A validation message that asserted a cause it no longer had
The check for the pricing dataset reported that it was empty and
that the account lacked getPricing. The second half was true when
written and false a day later, so the script confidently printed a fixed
problem as the reason for a normal one.
An empty dataset has two causes here and they are distinguishable by a single API call, so the check now makes that call and reports which: a permissions failure, or the 48-hour first-delivery window, since pricing data is not backfilled. A check that guesses between two causes is not a check.
Security — Controls at Every Layer
The exception is scoped to one project, in code, next to its reason. A console click relaxing an organization-wide constraint is indistinguishable a month later from carelessness. The override is a Terraform resource with the diagnosis written above it, so the next person to read it learns why it exists rather than whether to delete it.
The blast radius is asserted, not assumed. The validation script checks two things, and the second matters more: that the data project overrides the restriction, and that the organization still enforces it. The second check exists to catch a plausible future mistake — someone hitting the same error, reaching for the same fix, and applying it one level too high.
No publisher binding is written by hand. Granting a guessed service agent address would have meant either a binding to a principal that does not exist, or a permanent grant to whatever that name later resolves to. Letting Cloud Billing add its own binding means the policy contains exactly one publisher and it is the one Google actually uses.
Billing account IAM stays out of Terraform. A billing account is outside the hierarchy this configuration owns, and the only identity that can modify its bindings is deliberately not the identity that runs applies. The CI service account cannot widen its own access to the account that pays for everything.
The lab's evidence is redacted at capture time. The screenshot tooling masks organization IDs, billing account IDs, project numbers and personal email addresses, and refuses to write the file if any survive redaction. Service account addresses are allowlisted deliberately — a Google-owned service identity is a public fact about the platform, and this week it is the evidence.
Cost
$0.006. Total, across five projects, from 1 July to 27 September 2026 — and this is the first cost note in the series that is a query result rather than an assertion.
What this week adds costs nothing meaningful. BigQuery storage for the billing export is the only ongoing charge, and 2,412 rows sits far below the 10 GiB free tier. The budget, the Pub/Sub topic and the organization policy override are all free. The topic has no subscriber, so nothing accumulates in it.
The budget will not stop any of this. Spend caps shipped in Preview on 27 July 2026 and do genuinely pause services — but only the Gemini API, the Gemini Enterprise Agent Platform, Cloud Run and Cloud Run functions. This lab runs none of them. So for Compute Engine, Cloud Storage and BigQuery the $25 guardrail notifies and Google keeps billing past it. It is a smoke alarm, not a sprinkler, and it is worth knowing which one you installed.
The non-monetary cost is worth naming instead: an organization-wide security control now has a documented exception. That is a real, permanent increase in the surface area someone has to understand, and it is the actual price of this week.
Cleanup
This week is exempt from teardown, and says so deliberately. A cost-observability layer that is destroyed at the end of the week observes nothing. Weeks 7 onward depend on the export existing, and the historical rows cannot be recovered once the link is broken — a re-enabled export backfills at most the previous month.
If you do tear it down, two things behave in ways worth knowing in advance.
Destroying the dataset does not unlink the export, because unlinking is console-only in the same way linking is. You will have a billing account cheerfully writing to a destination that no longer exists.
And removing the organization policy override will silently break budget notifications rather than loudly failing. The binding Cloud Billing wrote stays on the topic, so nothing errors — the next re-attach or re-verification simply refuses, with the same four words this week started with.
References
- Set up programmatic budget notifications — the one page naming this week's blocker: notification delivery fails where organization policies limit resource sharing by domain. It does not connect that to the error string.
- Billing export table structure — that detailed includes every field standard has. That sentence deleted a dataset from this week.
- Restricting identities by domain — a list constraint over customer IDs and principal sets, with no way to name an individual service agent.
- Create and manage budgets — thresholds, forecast basis, and what a budget does and does not stop.
Not answered anywhere I looked, and measured instead: what
Precondition check failed means; which identity Cloud Billing
publishes as; how long an organization policy takes to relax; and which roles
carry billing.accounts.getPricing.
Key Takeaways
- Check whether the thing you are about to build already exists.
Half this week's Terraform was deleted after one
bq ls. A premise you wrote yourself is not evidence, and it is the assumption nobody re-examines. - Infrastructure nobody queries is worse than infrastructure that does not exist. It bills, it looks like coverage, and it lets you believe a question is answered when it has never been asked.
- When an error names nothing, vary one thing at a time. A deliberately wrong request that fails differently is worth more than any amount of documentation, because it eliminates rather than suggests.
- Guardrails bill you later, in another service. A control working exactly as designed broke a legitimate integration three weeks on, through an error that mentioned neither IAM nor policy. Budget for that when you build them, not when you hit it.
- Scope the exception, never lift the control. One project exempt, in code, with the reason attached — and a check asserting the organization still enforces it, because the tempting fix is one level too high.
- Read service identities back; never write them from memory.
billing-budget-alert, notbilling-budgets. Close enough to pass review, wrong enough to fail.
Next week the hierarchy starts answering questions about itself: what exists, what changed, and what no longer matches the thing that was supposed to have created it.
Comments