Repository navigation
docs: add multi-tier enterprise guardrails & quality gates example - #614
Conversation
There was a problem hiding this comment.
Code Review
This pull request adds a comprehensive multi-tier enterprise guardrails and quality gates example, including Python scripts, node configurations, and documentation. The review feedback highlights several robustness improvements: handling a database initialization race condition in support_agent.py, validating order existence before issuing refunds in orders_mcp.py, preventing a potential TypeError with null peer IDs in audit.py, and adding explicit error checks for failed peer ID discoveries in run.sh.
| def load_orders_snapshot() -> str: | ||
| if not os.path.exists(DB_PATH): | ||
| return "Order #1042: customer=alice@acme.com, item='Mechanical Keyboard Pro', total=$129.00, status=SHIPPED" | ||
| with sqlite3.connect(DB_PATH) as conn: | ||
| rows = conn.execute( | ||
| "SELECT order_id, customer, item, amount_dollars, status FROM orders ORDER BY order_id" | ||
| ).fetchall() | ||
| return "; ".join( | ||
| f"Order #{r[0]} (customer={r[1]}, item={r[2]}, total=${r[3]:.2f}, status={r[4]})" | ||
| for r in rows | ||
| ) |
There was a problem hiding this comment.
There is a potential race condition during startup. Since orders_mcp.py and support_agent.py are started concurrently in the background, support_agent.py might initialize the database first (creating the chat_history table), meaning DB_PATH will exist. If a request is processed before orders_mcp.py has finished creating the orders table, load_orders_snapshot will raise sqlite3.OperationalError: no such table: orders.
We should wrap the query in a try-except block to handle sqlite3.OperationalError gracefully.
| def load_orders_snapshot() -> str: | |
| if not os.path.exists(DB_PATH): | |
| return "Order #1042: customer=alice@acme.com, item='Mechanical Keyboard Pro', total=$129.00, status=SHIPPED" | |
| with sqlite3.connect(DB_PATH) as conn: | |
| rows = conn.execute( | |
| "SELECT order_id, customer, item, amount_dollars, status FROM orders ORDER BY order_id" | |
| ).fetchall() | |
| return "; ".join( | |
| f"Order #{r[0]} (customer={r[1]}, item={r[2]}, total=${r[3]:.2f}, status={r[4]})" | |
| for r in rows | |
| ) | |
| def load_orders_snapshot() -> str: | |
| if not os.path.exists(DB_PATH): | |
| return "Order #1042: customer=alice@acme.com, item='Mechanical Keyboard Pro', total=$129.00, status=SHIPPED" | |
| try: | |
| with sqlite3.connect(DB_PATH) as conn: | |
| rows = conn.execute( | |
| "SELECT order_id, customer, item, amount_dollars, status FROM orders ORDER BY order_id" | |
| ).fetchall() | |
| return "; ".join( | |
| f"Order #{r[0]} (customer={r[1]}, item={r[2]}, total=${r[3]:.2f}, status={r[4]})" | |
| for r in rows | |
| ) | |
| except sqlite3.OperationalError: | |
| return "Order #1042: customer=alice@acme.com, item='Mechanical Keyboard Pro', total=$129.00, status=SHIPPED" |
| def issue_refund(order_id: str, amount_dollars: float, reason: str = "customer request") -> str: | ||
| """Issue a refund for an order and update its status in SQLite (mutating financial operation).""" | ||
| with sqlite3.connect(DB_PATH) as conn: | ||
| conn.execute( | ||
| "INSERT INTO refunds (order_id, amount_dollars, reason) VALUES (?, ?, ?)", | ||
| (order_id, amount_dollars, reason), | ||
| ) | ||
| conn.execute( | ||
| "UPDATE orders SET status = 'REFUNDED' WHERE order_id = ?", | ||
| (order_id,), | ||
| ) | ||
| conn.commit() | ||
| return f"REFUND_EXECUTED: order=#{order_id}, amount=${amount_dollars:.2f}, reason='{reason}'" | ||
|
|
There was a problem hiding this comment.
The issue_refund tool does not verify if the specified order_id actually exists in the orders table before inserting a refund record and returning success. This can lead to orphaned refund records and inconsistent database state.
We should check if the order exists first and return an error message if it does not.
| def issue_refund(order_id: str, amount_dollars: float, reason: str = "customer request") -> str: | |
| """Issue a refund for an order and update its status in SQLite (mutating financial operation).""" | |
| with sqlite3.connect(DB_PATH) as conn: | |
| conn.execute( | |
| "INSERT INTO refunds (order_id, amount_dollars, reason) VALUES (?, ?, ?)", | |
| (order_id, amount_dollars, reason), | |
| ) | |
| conn.execute( | |
| "UPDATE orders SET status = 'REFUNDED' WHERE order_id = ?", | |
| (order_id,), | |
| ) | |
| conn.commit() | |
| return f"REFUND_EXECUTED: order=#{order_id}, amount=${amount_dollars:.2f}, reason='{reason}'" | |
| @mcp.tool() | |
| def issue_refund(order_id: str, amount_dollars: float, reason: str = "customer request") -> str: | |
| """Issue a refund for an order and update its status in SQLite (mutating financial operation).""" | |
| with sqlite3.connect(DB_PATH) as conn: | |
| row = conn.execute("SELECT 1 FROM orders WHERE order_id = ?", (order_id,)).fetchone() | |
| if not row: | |
| return f"Error: Order #{order_id} not found. Refund aborted." | |
| conn.execute( | |
| "INSERT INTO refunds (order_id, amount_dollars, reason) VALUES (?, ?, ?)", | |
| (order_id, amount_dollars, reason), | |
| ) | |
| conn.execute( | |
| "UPDATE orders SET status = 'REFUNDED' WHERE order_id = ?", | |
| (order_id,), | |
| ) | |
| conn.commit() | |
| return f"REFUND_EXECUTED: order=#{order_id}, amount=${amount_dollars:.2f}, reason='{reason}'" |
| d = json.loads(m.group(1)) | ||
| except json.JSONDecodeError: | ||
| continue | ||
| who = f"{d.get('role') or '-'} {d.get('peer_id', '')[:12]}…" |
There was a problem hiding this comment.
If the JSON log contains "peer_id": null, d.get('peer_id', '') will return None (since the key exists but its value is null). Attempting to slice it with [:12] will then raise a TypeError: 'NoneType' object is not subscriptable.
We should use (d.get('peer_id') or '')[:12] to safely handle None values.
| who = f"{d.get('role') or '-'} {d.get('peer_id', '')[:12]}…" | |
| who = f"{d.get('role') or '-'} {(d.get('peer_id') or '')[:12]}…" |
| fi | ||
| sleep 0.5 | ||
| done | ||
| echo " Discovered Alice's node PeerID via DHT/GossipSub: ${V1_PEER}" |
There was a problem hiding this comment.
If node discovery fails and V1_PEER remains empty after the retry loop, the script will continue with an empty peer ID. This leads to malformed URLs (e.g., http://127.0.0.1:19002/sam//a2a/...) and subsequent failures that are hard to debug.
We should check if V1_PEER is empty and exit with an error if discovery failed.
| echo " Discovered Alice's node PeerID via DHT/GossipSub: ${V1_PEER}" | |
| if [[ -z "${V1_PEER}" ]]; then | |
| echo "ERROR: Failed to discover Alice's node PeerID" >&2 | |
| exit 1 | |
| fi | |
| echo " Discovered Alice's node PeerID via DHT/GossipSub: ${V1_PEER}" |
| done | ||
| V1_PEER="${V1_PROD_PEER}" |
There was a problem hiding this comment.
If the promoted production PeerID discovery fails, V1_PEER will be empty, leading to subsequent failures. We should check if V1_PEER is empty and exit with an error if so.
| done | |
| V1_PEER="${V1_PROD_PEER}" | |
| V1_PEER="${V1_PROD_PEER}" | |
| if [[ -z "${V1_PEER}" ]]; then | |
| echo "ERROR: Failed to discover promoted production PeerID" >&2 | |
| exit 1 | |
| fi | |
| echo " Promoted Production PeerID (attested env=prod): ${V1_PEER}" |
| done | ||
| sleep 0.5 | ||
| done | ||
| echo " Discovered Cloud Run v2 PeerID: ${V2_PEER}" |
There was a problem hiding this comment.
If the Cloud Run v2 PeerID discovery fails, V2_PEER will be empty, leading to subsequent failures. We should check if V2_PEER is empty and exit with an error if so.
| echo " Discovered Cloud Run v2 PeerID: ${V2_PEER}" | |
| if [[ -z "${V2_PEER}" ]]; then | |
| echo "ERROR: Failed to discover Cloud Run v2 PeerID" >&2 | |
| exit 1 | |
| fi | |
| echo " Discovered Cloud Run v2 PeerID: ${V2_PEER}" |
Adds a runnable end-to-end example (
development/examples/multi-tier-guardrails/), documentation page (site/content/docs/use-cases/multi-tier-guardrails.md), and terminal recording (site/static/demo-multi-tier-guardrails.mp4) demonstrating how SAM serves all 5 enterprise personas across 4 acts using real backends:gemma3:1bvia Ollama): Powers two A2A 1.0 replicas (support_agent.pyonv1-laptopandv2-cloudrun) built with the officiala2a-sdkand SQLite conversation history.orders_mcp.py): Built with the officialmcpPython SDK (MCPServer), exposingget_order_status(read) andissue_refund(write).egress://api.github.com): Connects tohttps://api.github.comwith node-side secret brokering (secrets/github-ro).The 5 Personas & 3 Label Control Points Demonstrated
policy.jsonallowed_labels): Shows that a node cannot self-assertenv: prodin its YAML before passing the Quality Gate (Label not permitted).node-caller.yamlegress.require_labels: {env: prod}): Shows that the caller's Input Node blocks staging providers (403 Forbidden) even when the application sends no label headers, until the Quality Gate validates the A2A 1.0 AgentCard and promotes policy to grantenv=prod(200 OK).X-Sam-Required-Labels): Within the Input Node's mandatoryenv=prodfloor, the calling agent passesX-Sam-Required-Labels: replica=v1-laptopon Turn 1 andreplica=v2-cloudrunon Turn 2 while preservingcontextIdacross replicas.GET /repos/google/sam/pullsallowed with200 OK;POST /repos/google/sam/pullsblocked with403 Forbidden).check if label("team", "support");) innode-v1.yamlrejecting ateam=contractornode locally.POST /oauth/tokenmints a sealed Task Biscuit (tar_block) scoped strictly totool:get_order_status, allowingget_order_statuswhile blocking a prompt-injectedissue_refundmid-flight.ALLOW/DENYaudit log (audit.py) across all three guardrail layers + instant mesh-wide revocation (sam-one admin ban).