All research

An agent workflow with a human decision point

A synthetic purchase request illustrates retrieval, limited tool access, approval, and evaluation.

The question

How can an assistant prepare a purchase request without gaining the authority to place an order?

What we built & the conditions

We built a deterministic browser demonstration around one fictional request: four monitors at ¥30,000 each. A synthetic policy document allows requests up to ¥150,000 but requires human approval before any order. The interface walks through retrieval and a draft tool call. It makes no model calls and connects to no business systems.

SYNTHETIC / 001Illustrative demo · No external actions
  1. 1

    Read the request

    Four monitors, ¥30,000 each. Total: ¥120,000.

  2. 2

    Retrieve the policy

    Synthetic policy: requests up to ¥150,000. Human approval required before ordering.

    SYNTHETIC-POLICY-01
  3. 3

    Draft within tool permissions

    Search and draft tools allowed. Ordering and payments denied.

    draft_request → pending_human_review
  4. 4

    A person decides

    Being within budget does not authorize execution. Try either decision.

Waiting for human review.

Observations from the example

  • The synthetic request totals ¥120,000, within the example policy threshold. This is arithmetic in the demonstration, not a model benchmark.
  • The draft can cite the supplied policy, but being within budget does not grant permission to place an order.
  • Rejecting the request stops the illustration. Approving it displays a simulated handoff; it never creates an order.

Limitations

This is an explanatory artifact, not a tested autonomous agent or a client engagement. It does not measure retrieval accuracy, model quality, latency, security effectiveness, or business outcomes. Production systems need server-enforced authorization, real identity checks, idempotency, audit retention, and adversarial testing.

Business relevance & evaluation

Separating recommendation from execution makes responsibility visible. Before deployment, evaluation should test missing evidence, out-of-policy requests, attempted unauthorized tools, rejection, and duplicate approvals.

Inspect the artifact

Download the annotated synthetic scenario (JSON)

Bring us the problem. We’ll work through the possibilities.

A workflow to improve, a model to adapt, or a prototype ready for its next step.

Discuss your project