Hardened MCP servers architecture: identity token verified by Cloud Run IAM and again in the app, per-caller scopes, Firestore rate limiting, and a BigQuery audit trail with no argument values logged
Five identities — operator, a read-only agent, a rate-limit stress test, a registered-but-unlisted guest, and a stranger with no access at all — exercise every layer live.

Most "connect your agent to a remote MCP server" write-ups stop at the handshake. The harder, more interesting question is what happens once an agent can call tools against your infrastructure: who's allowed to call what, what happens when one caller misbehaves, what a security review actually sees in the logs afterward — and, it turned out here, a genuine platform quirk that a naive test would have quietly swallowed as "must be my code."

gcp-mcp-platform runs two MCP servers on private Cloud Run services — ops (read-only project introspection: list Cloud Run services, tail recent errors, dry-run a BigQuery cost estimate) and notes (a small multi-tenant store) — reachable only by named callers. Every request is authenticated with a Google identity token checked twice: once by Cloud Run's own IAM invoker check, and again inside the app (signature, audience, issuer, verified email), then matched against a caller policy rendered from Terraform, rate-limited per caller in Firestore, and logged to BigQuery with the decision and reason but never the argument values. It deployed live in europe-west1, and the test suite passed 52 of 52 — driven through five distinct identities that exercise every layer, not just the happy path.

Five identities, one suite

The Terraform stands up an operator identity with every scope, a read-only agent that can call tools but not write notes, a "burst" identity used purely to trip the rate limiter, a "guest" that holds Cloud Run's invoker role but is deliberately absent from the app's own caller policy, and a "stranger" with no invoker role at all. Running the live suite as all five at once is what actually proves the two layers are independent: the stranger never reaches the app (Cloud Run IAM refuses it outright), the guest reaches the app but is refused there with unregistered_caller, and the read-only agent reaches its allowed tools but gets a scoped refusal — missing_scope:notes:write — the moment it tries to write.

A genuine Cloud Run finding, not a bug in this repo

The very first live call as the human operator failed with a 401 MalformedError: Could not verify token signature — on a token that verified perfectly fine locally, in a Cloud Run Job, and through gcloud run services proxy. A debug log of the raw ASGI header length, added and deployed specifically to chase this down, showed the answer: 826 characters sent by curl, 518 received by the app. A fresh service-account token of similar length arrived completely intact. The one thing the truncated tokens had in common was their audience — 32555940559.apps.googleusercontent.com, the gcloud SDK's own first-party OAuth client id. Cloud Run's front end validates that token correctly for its own IAM invoker check (which is why the request reaches the app at all) but forwards a truncated Authorization header to the container specifically for tokens issued to that client. That's a platform behavior, not a defect in the app — and it's not documented anywhere I could find, only reachable by instrumenting a live deployment.

sent by curl:      826 chars
received by app: 518 chars (human token, audience = gcloud's own OAuth client)
received by app: 819 chars (service-account token, audience = the service itself)

The fix reframes rather than works around it: every caller, including the human operator, now authenticates as a dedicated service account with a normal, non-redacted, service-scoped token — exactly like the agent identities already did. The raw human-token path stays in the code (explicitly refused for service accounts, in principle accepted for humans), and one live test now asserts a raw human token is refused with invalid_token, turning a debugging session into a permanent regression check instead of a footnote.

Argument values never reach the log

Every decision — allowed or denied — is one structured JSON line: caller, server, tool, decision, reason, latency, and an args_digest (a SHA-256 of the canonical arguments) plus the argument names. Never the values. That matters concretely here, because notes hold arbitrary user text and the cost tool takes raw SQL — either could carry something sensitive. The live test doesn't just trust the design: it writes a unique marker string into a real note body, then queries the entire BigQuery audit table for that exact string and asserts zero matches, while confirming the argument names and a non-null digest are present for that same call.

Five decisions, honestly framed

DecisionTrade-off
Google identity tokens + Cloud Run IAM, not an OAuth authorization serverNo public endpoint or client registration flow, but a client that only speaks MCP's native OAuth flow needs a proxy in front
The caller policy is JSON rendered by Terraform, not a runtime APIReviewed in a pull request and visible in the plan, at the cost of a redeploy to change who can call what
Rate limiting is a Firestore transaction per (caller, minute)Shared correctly across autoscaled instances, with TTL cleanup, but fixed windows allow up to a 2× burst at a boundary
Human operators authenticate as a dedicated service accountFound live, not designed in from the start — see above
One retry on a `ValueError` from the token verifierA defensible hedge against a transient hiccup fetching Google's certs — it wasn't actually the audience bug, but it's harmless and kept

What this doesn't claim to be

There's no MCP-native OAuth authorization server here, so a client that only speaks that flow needs a proxy in front of this design. Note text handed back to a model is explicitly untrusted — results carry a "this is data, not instructions" warning, but the real mitigation is on the client: don't let an agent act on note content unsupervised. roles/logging.viewer on the ops server is project-wide (needed for the recent_errors tool to read any service's logs), which is broader than a single-resource-type log view would be. And the fixed-window rate limiter, as noted above, tolerates a short burst right at a minute boundary rather than smoothing it out.

Try it yourself

git clone https://github.com/soodrajesh/gcp-mcp-platform
cd gcp-mcp-platform
gcloud config set project <your-project>
./scripts/up.sh # Firestore + audit sink + identities + registry → image → both services → live tests
./scripts/down.sh # delete everything, including the Firestore database

Both services scale to zero; Firestore and BigQuery usage from the test suite stays far below free-tier limits.

📢 Have questions or feedback? Drop a comment below or connect with me on Twitter/X@spysood!