Skip to content

docs: add multi-tier enterprise guardrails & quality gates example - #614

Merged
aojea merged 1 commit into
google:mainfrom
aojea:docs/multi-tier-guardrails-example
Oct 9, 2026
Merged

aojea merged 1 commit into
google:mainfrom
aojea:docs/multi-tier-guardrails-example

Conversation

@aojea

@aojea aojea commented Oct 9, 2026

Copy link
Copy Markdown
Collaborator

Adds a runnable end-to-end example (development/examples/multi-tier-guardrails/), documentation page (site/content/docs/use-cases/multi-tier-guardrails.md), and terminal recording (site/static/demo-multi-tier-guardrails.mp4) demonstrating how SAM serves all 5 enterprise personas across 4 acts using real backends:

  • Real Local LLM (gemma3:1b via Ollama): Powers two A2A 1.0 replicas (support_agent.py on v1-laptop and v2-cloudrun) built with the official a2a-sdk and SQLite conversation history.
  • Real SQLite MCP Server (orders_mcp.py): Built with the official mcp Python SDK (MCPServer), exposing get_order_status (read) and issue_refund (write).
  • Real External API (egress://api.github.com): Connects to https://api.github.com with node-side secret brokering (secrets/github-ro).

The 5 Personas & 3 Label Control Points Demonstrated

  1. Act 1 — Developer Onboarding & Automated Quality Gate (Personas: Agent Developer Alice + Central Security / Platform):
    • Label Point 1 (Mesh Enforcement — policy.json allowed_labels): Shows that a node cannot self-assert env: prod in its YAML before passing the Quality Gate (Label not permitted).
    • Label Point 2a (Input Node Enforcement — node-caller.yaml egress.require_labels: {env: prod}): Shows that the caller's Input Node blocks staging providers (403 Forbidden) even when the application sends no label headers, until the Quality Gate validates the A2A 1.0 AgentCard and promotes policy to grant env=prod (200 OK).
  2. Act 2 — Zero-Downtime Replica Swap & Agent Intent (Persona: Platform / Networking):
    • Label Point 3 (End-User / Agent Intent — X-Sam-Required-Labels): Within the Input Node's mandatory env=prod floor, the calling agent passes X-Sam-Required-Labels: replica=v1-laptop on Turn 1 and replica=v2-cloudrun on Turn 2 while preserving contextId across replicas.
  3. Act 3 — The 3-Tier Guardrail Hierarchy (Personas: Central Security, Department Lead, Individual End-User):
    • Layer 1 (Central Security): Org HTTP policy + node-side secret brokering (GET /repos/google/sam/pulls allowed with 200 OK; POST /repos/google/sam/pulls blocked with 403 Forbidden).
    • Layer 2 (Department Lead — Label Point 2b, Egress Node Check): Positive inbound Datalog check (check if label("team", "support");) in node-v1.yaml rejecting a team=contractor node locally.
    • Layer 3 (Individual End-User): RFC 8693 POST /oauth/token mints a sealed Task Biscuit (tar_block) scoped strictly to tool:get_order_status, allowing get_order_status while blocking a prompt-injected issue_refund mid-flight.
  4. Act 4 — Day-2 Operations (Persona: SRE / Operations):
    • Structured ALLOW/DENY audit log (audit.py) across all three guardrail layers + instant mesh-wide revocation (sam-one admin ban).

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request adds a comprehensive multi-tier enterprise guardrails and quality gates example, including Python scripts, node configurations, and documentation. The review feedback highlights several robustness improvements: handling a database initialization race condition in support_agent.py, validating order existence before issuing refunds in orders_mcp.py, preventing a potential TypeError with null peer IDs in audit.py, and adding explicit error checks for failed peer ID discoveries in run.sh.

Comment on lines +77 to +87
def load_orders_snapshot() -> str:
if not os.path.exists(DB_PATH):
return "Order #1042: customer=alice@acme.com, item='Mechanical Keyboard Pro', total=$129.00, status=SHIPPED"
with sqlite3.connect(DB_PATH) as conn:
rows = conn.execute(
"SELECT order_id, customer, item, amount_dollars, status FROM orders ORDER BY order_id"
).fetchall()
return "; ".join(
f"Order #{r[0]} (customer={r[1]}, item={r[2]}, total=${r[3]:.2f}, status={r[4]})"
for r in rows
)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

There is a potential race condition during startup. Since orders_mcp.py and support_agent.py are started concurrently in the background, support_agent.py might initialize the database first (creating the chat_history table), meaning DB_PATH will exist. If a request is processed before orders_mcp.py has finished creating the orders table, load_orders_snapshot will raise sqlite3.OperationalError: no such table: orders.

We should wrap the query in a try-except block to handle sqlite3.OperationalError gracefully.

Suggested change
def load_orders_snapshot() -> str:
if not os.path.exists(DB_PATH):
return "Order #1042: customer=alice@acme.com, item='Mechanical Keyboard Pro', total=$129.00, status=SHIPPED"
with sqlite3.connect(DB_PATH) as conn:
rows = conn.execute(
"SELECT order_id, customer, item, amount_dollars, status FROM orders ORDER BY order_id"
).fetchall()
return "; ".join(
f"Order #{r[0]} (customer={r[1]}, item={r[2]}, total=${r[3]:.2f}, status={r[4]})"
for r in rows
)
def load_orders_snapshot() -> str:
if not os.path.exists(DB_PATH):
return "Order #1042: customer=alice@acme.com, item='Mechanical Keyboard Pro', total=$129.00, status=SHIPPED"
try:
with sqlite3.connect(DB_PATH) as conn:
rows = conn.execute(
"SELECT order_id, customer, item, amount_dollars, status FROM orders ORDER BY order_id"
).fetchall()
return "; ".join(
f"Order #{r[0]} (customer={r[1]}, item={r[2]}, total=${r[3]:.2f}, status={r[4]})"
for r in rows
)
except sqlite3.OperationalError:
return "Order #1042: customer=alice@acme.com, item='Mechanical Keyboard Pro', total=$129.00, status=SHIPPED"

Comment on lines +78 to +91
def issue_refund(order_id: str, amount_dollars: float, reason: str = "customer request") -> str:
"""Issue a refund for an order and update its status in SQLite (mutating financial operation)."""
with sqlite3.connect(DB_PATH) as conn:
conn.execute(
"INSERT INTO refunds (order_id, amount_dollars, reason) VALUES (?, ?, ?)",
(order_id, amount_dollars, reason),
)
conn.execute(
"UPDATE orders SET status = 'REFUNDED' WHERE order_id = ?",
(order_id,),
)
conn.commit()
return f"REFUND_EXECUTED: order=#{order_id}, amount=${amount_dollars:.2f}, reason='{reason}'"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The issue_refund tool does not verify if the specified order_id actually exists in the orders table before inserting a refund record and returning success. This can lead to orphaned refund records and inconsistent database state.

We should check if the order exists first and return an error message if it does not.

Suggested change
def issue_refund(order_id: str, amount_dollars: float, reason: str = "customer request") -> str:
"""Issue a refund for an order and update its status in SQLite (mutating financial operation)."""
with sqlite3.connect(DB_PATH) as conn:
conn.execute(
"INSERT INTO refunds (order_id, amount_dollars, reason) VALUES (?, ?, ?)",
(order_id, amount_dollars, reason),
)
conn.execute(
"UPDATE orders SET status = 'REFUNDED' WHERE order_id = ?",
(order_id,),
)
conn.commit()
return f"REFUND_EXECUTED: order=#{order_id}, amount=${amount_dollars:.2f}, reason='{reason}'"
@mcp.tool()
def issue_refund(order_id: str, amount_dollars: float, reason: str = "customer request") -> str:
"""Issue a refund for an order and update its status in SQLite (mutating financial operation)."""
with sqlite3.connect(DB_PATH) as conn:
row = conn.execute("SELECT 1 FROM orders WHERE order_id = ?", (order_id,)).fetchone()
if not row:
return f"Error: Order #{order_id} not found. Refund aborted."
conn.execute(
"INSERT INTO refunds (order_id, amount_dollars, reason) VALUES (?, ?, ?)",
(order_id, amount_dollars, reason),
)
conn.execute(
"UPDATE orders SET status = 'REFUNDED' WHERE order_id = ?",
(order_id,),
)
conn.commit()
return f"REFUND_EXECUTED: order=#{order_id}, amount=${amount_dollars:.2f}, reason='{reason}'"

d = json.loads(m.group(1))
except json.JSONDecodeError:
continue
who = f"{d.get('role') or '-'} {d.get('peer_id', '')[:12]}…"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

If the JSON log contains "peer_id": null, d.get('peer_id', '') will return None (since the key exists but its value is null). Attempting to slice it with [:12] will then raise a TypeError: 'NoneType' object is not subscriptable.

We should use (d.get('peer_id') or '')[:12] to safely handle None values.

Suggested change
who = f"{d.get('role') or '-'} {d.get('peer_id', '')[:12]}…"
who = f"{d.get('role') or '-'} {(d.get('peer_id') or '')[:12]}…"

fi
sleep 0.5
done
echo " Discovered Alice's node PeerID via DHT/GossipSub: ${V1_PEER}"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

If node discovery fails and V1_PEER remains empty after the retry loop, the script will continue with an empty peer ID. This leads to malformed URLs (e.g., http://127.0.0.1:19002/sam//a2a/...) and subsequent failures that are hard to debug.

We should check if V1_PEER is empty and exit with an error if discovery failed.

Suggested change
echo " Discovered Alice's node PeerID via DHT/GossipSub: ${V1_PEER}"
if [[ -z "${V1_PEER}" ]]; then
echo "ERROR: Failed to discover Alice's node PeerID" >&2
exit 1
fi
echo " Discovered Alice's node PeerID via DHT/GossipSub: ${V1_PEER}"

Comment on lines +268 to +269
done
V1_PEER="${V1_PROD_PEER}"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

If the promoted production PeerID discovery fails, V1_PEER will be empty, leading to subsequent failures. We should check if V1_PEER is empty and exit with an error if so.

Suggested change
done
V1_PEER="${V1_PROD_PEER}"
V1_PEER="${V1_PROD_PEER}"
if [[ -z "${V1_PEER}" ]]; then
echo "ERROR: Failed to discover promoted production PeerID" >&2
exit 1
fi
echo " Promoted Production PeerID (attested env=prod): ${V1_PEER}"

done
sleep 0.5
done
echo " Discovered Cloud Run v2 PeerID: ${V2_PEER}"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

If the Cloud Run v2 PeerID discovery fails, V2_PEER will be empty, leading to subsequent failures. We should check if V2_PEER is empty and exit with an error if so.

Suggested change
echo " Discovered Cloud Run v2 PeerID: ${V2_PEER}"
if [[ -z "${V2_PEER}" ]]; then
echo "ERROR: Failed to discover Cloud Run v2 PeerID" >&2
exit 1
fi
echo " Discovered Cloud Run v2 PeerID: ${V2_PEER}"

@aojea
aojea merged commit 63da499 into google:main Oct 9, 2026
21 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant