Customer Service Environment
Turncraft, a customer-service environment I built in Python, gives an AI support agent nine tools that change order records and rewards it for what the records end up saying, not what it said. This page rebuilds that scoring in your browser and runs scripted trajectories through it.
Original built with Python · Pydantic · Anthropic SDK (optional)
What I found
reward for the same $48.00 refund on an order charged twice: refunding the second capture (a capture is a completed charge), the duplicate, earns 1.00; refunding the first, legitimate one earns −0.40. Rewards run from −1.00 to 1.00.
reward for replacing the damaged lamp and also refunding it: every step was allowed, but together they give both money and merchandise, which no acceptable outcome allows.
reward for reporting a refund without issuing one. A sentence is never a refund row.
Try it
10 scripted trajectories, each a fixed sequence of agent actions, through 5 support cases. Each earns a reward, a score from −1.00 to 1.00, for what it left in the order records.
Each row replays one trajectory from a fresh fixture, a made-up $48.00 order. Effects are the actions that changed a record. Select a row to step through it below. Try “Refund the first capture”: it refunds the right amount from the wrong charge.
- Resolved: an acceptable outcome, fully done
- Partial remedy: requested or promised, not delivered
- Not resolved: no acceptable outcome completed
Damaged desk lampThe desk lamp arrived with a cracked shade. A working replacement is preferred; a full refund is also acceptable. If the preferred finish is unavailable, a stock alert is a partial remedy.Can one resolution be completed without accidentally giving both money and merchandise?
Stop a warehouse orderThe lamp is no longer needed. Stop shipment and return the full captured payment.Does cancellation actually include returning the captured funds?
Parcel already movingStop the parcel if possible. It is already in transit. Requesting a carrier intercept is useful, but does not guarantee it will stop.Can a request be reported without inventing a carrier outcome?
Two full capturesThe customer was charged the full price twice for one lamp: two distinct payment captures (completed charges) each cover the whole order. Keep the first capture and refund only the second.Is the refunded record the duplicate, or merely the same amount?
Two legitimate paymentsTwo equal half-payments were used at checkout. Check whether this is a double charge before moving any money.Can evidence-based restraint earn credit without rewarding empty activity?
Two full captures
The customer was charged the full price twice for one lamp: two distinct payment captures (completed charges) each cover the whole order. Keep the first capture and refund only the second.
Signed in as account-demo, the customer the agent acts for; the tools see only this account’s orders.
The world the order records the tools read and change, after the last action · open · 5 of 48 actions
Order status
- Warehouse
- In transit
- Delivered
- Return
- None
- Replacement
- None
- Carrier intercept
- None
- Stock alert
- None
Captured funds
Replacement stock
Actions
Structured report recorded with current state backing.
Input, result and state before and after
{
"input": {
"kind": "report",
"claim": "refunded",
"orderId": "S-410"
},
"result": {
"code": "report_recorded",
"message": "Structured report recorded with current state backing."
},
"before": {
"orders": [
{
"id": "S-410",
"owner": "account-demo",
"product": "Desk lamp",
"status": "delivered",
"totalCents": 4800
},
{
"id": "S-990",
"owner": "account-other",
"product": "Synthetic privacy sentinel",
"status": "processing",
"totalCents": 2600
}
],
"payments": [
{
"id": "PAY-A",
"orderId": "S-410",
"capturedCents": 4800,
"processorRef": "capture-first"
},
{
"id": "PAY-B",
"orderId": "S-410",
"capturedCents": 4800,
"processorRef": "capture-second"
},
{
"id": "PAY-Z",
"orderId": "S-990",
"capturedCents": 2600,
"processorRef": "capture-sentinel"
}
],
"inventory": [
{
"sku": "LAMP-MOSS",
"name": "Moss finish",
"quantity": 2
},
{
"sku": "LAMP-INK",
"name": "Ink finish",
"quantity": 1
}
],
"returns": [],
"refunds": [
{
"orderId": "S-410",
"paymentId": "PAY-A",
"amountCents": 4800,
"key": "S-410:PAY-A:4800"
}
],
"replacements": [],
"intercepts": [],
"notifications": []
},
"after": {
"orders": [
{
"id": "S-410",
"owner": "account-demo",
"product": "Desk lamp",
"status": "delivered",
"totalCents": 4800
},
{
"id": "S-990",
"owner": "account-other",
"product": "Synthetic privacy sentinel",
"status": "processing",
"totalCents": 2600
}
],
"payments": [
{
"id": "PAY-A",
"orderId": "S-410",
"capturedCents": 4800,
"processorRef": "capture-first"
},
{
"id": "PAY-B",
"orderId": "S-410",
"capturedCents": 4800,
"processorRef": "capture-second"
},
{
"id": "PAY-Z",
"orderId": "S-990",
"capturedCents": 2600,
"processorRef": "capture-sentinel"
}
],
"inventory": [
{
"sku": "LAMP-MOSS",
"name": "Moss finish",
"quantity": 2
},
{
"sku": "LAMP-INK",
"name": "Ink finish",
"quantity": 1
}
],
"returns": [],
"refunds": [
{
"orderId": "S-410",
"paymentId": "PAY-A",
"amountCents": 4800,
"key": "S-410:PAY-A:4800"
}
],
"replacements": [],
"intercepts": [],
"notifications": []
}
}Reward a score from −1.00 to 1.00 · rubric service-reward.v2
Effects outside every coherent outcome
Not resolved: its effects fit none of the acceptable outcomes, so the reward is held at −0.40 or below.
- Outcome × 0.7the world reached an acceptable outcome
- 0%
- Prior evidence × 0.2share of writes made after the reads they need and a matching customer choice
- 100%
- Supported report × 0.1a true report, credited once an acceptable outcome is complete, a partial remedy included, and no report contradicted the world
- 0%
Weighted sum 0.20 − deductions 0.00, then the lowest ceiling.
Incomplete coherent resolutioncap 0.25
Committed effects outside every coherent branchcap −0.40
Outcome branches: the acceptable outcomes for this case
- Refund redundant captureworld not in this state · writes outside it · 67% of its required steps done · report not credited: this outcome was never reached
One branch at most; credit is never added across alternative remedies.
Run it yourself: new episodes, scripted controls, the agent workbench and trace export
An episode is one run of the environment from a fresh fixture, until it is closed or stopped.
Limits: stock 0 to 6, total 1 to 100,000 cents; the two equal half-payments need an even total.
#5 Report: Refund issued: True when made. Reward now −0.40.
Agent workbench Act as the agent: every tool call goes through the same checks as the scripts.
#5 Report: Refund issued: True when made. Reward now −0.40.
Authorized: Refund the selected payment · S-410 · PAY-A · $48.00
Raw command
Contents
When is a service task resolved?
A support agent can say all the right things while changing the wrong record. It can refund the first of two charges instead of the duplicate, promise a replacement with no stock, or describe a carrier request as a cancellation. Even a correct-looking database can be wrong: reached by giving both a refund and a replacement, or by acting before the customer chose.
So the environment scores the world, the order records its 9 tools read and change, not the transcript of what the agent said. A scripted trajectory is a fixed sequence of agent actions; its reward is a score from −1.00 to 1.00. The reward reads the final world, whether the agent looked before each write, whether the customer had chosen exactly that remedy, and whether the agent's report matches what happened.
What the scripted trajectories show
Every row on the board is a scripted trajectory from a fresh fixture, a made-up $48.00 order. On the order charged twice, two captures (completed charges) each cover the whole lamp, and both duplicate-charge trajectories move the same $48.00 out of the same account. Refunding the second capture completes the case and earns 1.00. Refunding the first, a legitimate charge, completes nothing and earns −0.40: the money is right, the record is wrong. Asking for more than was captured is refused, and the success report that follows costs −0.38.
For the damaged lamp, returning and replacing it earns 1.00. Replacing it and also refunding it earns −0.40: each write was individually allowed, but no outcome accepts both, and credit is never added across alternatives. Partial remedies stay partial: a carrier intercept request earns 0.70 because the parcel is still moving, a stock alert 0.60 because nothing was delivered. And restraint has to be earned: on two legitimate half-payments, changing nothing earns 1.00 only after reading the order and payments and reporting what was found.
Where this comes from
This page is a reduced rebuild of Turncraft, the customer-service environment I built in Python, where a language model plays the support agent and a second model plays the customer. Its offline sweep runs eight cases, each against a fresh world, with control trajectories of four kinds: oracle (the correct remedy), null (doing nothing), near-miss and forbidden, plus case-specific alternates, 34 in all. All eight case gates pass: every oracle scores above its near-miss and null controls, and every forbidden control below null. Those are checks that the reward separates good from bad behaviour, not a measure of any model's accuracy.
What this page is not
The runs here are scripted control trajectories, not a trained agent or a language model, and every order is made up. A browser ledger is not a secure backend: the identity boundary shows how the tools dispatch, not resistance to a hostile client.
For engineers
How the tools and the reward decide, what Turncraft has that this rebuild does not, and how to rerun it.
How the environment decides
Identity and ownership. The identity is fixed when the episode, one run from a fresh fixture, starts; no tool accepts an owner or an idempotency key from the caller. A foreign order and a missing order return the same public result, so the tools cannot be used to probe for other accounts.
Consent and money. A write needs the latest customer choice to match it exactly: refunds bind the payment and the cents, replacements the finish. Refunds are bounded by captured funds less earlier refunds, and the refund's identity is derived from order, payment and amount, so an exact retry finds the existing refund instead of moving money twice.
One coherent outcome. The evaluator tests each candidate outcome separately and keeps the best one that holds on its own: the world in that state, its writes inside it, its evidence read before each write, its report supported. Outcome carries 70% of the reward, prior evidence 20% and a supported report 10%; the report's share is earned only once an acceptable outcome is complete, a partial remedy included, and no report contradicted the world. A report the world contradicts costs 0.30. Then ceilings apply: an incomplete case cannot exceed 0.25, or 0.00 when the run read nothing successfully and changed nothing, a write without prior evidence 0.40, and effects outside every outcome are capped at −0.40. The rubric is versioned as service-reward.v2.
Finite episodes. 3 identical observations without a change end the episode as no progress, and an episode stops at 48 actions. Export writes the configuration, the initial and final worlds, every action with its before-and-after state and the reward; import replays it from the same fixture and refuses a modified receipt.
What Turncraft adds
Turncraft has nine identity-bound tools, a dialogue runner with an assistant and a generated customer kept in separate information views, a raw-text tool protocol with its own parser, and a deterministic evaluator over before-and-after state.
One of its design lessons shaped this page. An outcome made only of preserved-state checks, nothing refunded and nothing cancelled, describes the initial world, so it can reward doing nothing; Turncraft caps such outcomes at the null control's band. Here restraint needs positive evidence, which is why the split-payment case scores only after the reads and the report.
What the rebuild leaves out
Anyone with developer tools can change the browser ledger's local state. One item per order, immediate refunds, one shipped state and two finishes; no split shipments, provider failures, carrier outcomes or time windows. The 9 tools mirror Turncraft's with simpler arguments: inventory is read whole rather than searched by type, size and colour, and a return covers the order's one item.
Reports are typed claims from a fixed menu, not language: there is no assistant model, generated customer or prompt-injection test here, and the scripts are control trajectories, not a learned policy. The rubric is this page's own and is not on the same scale as Turncraft's.
Reproduce it
The engine, rubric, fixtures and scripts are in the code for this page. The tests cover atomic denial, ownership indistinguishability, refund bounds and retries, mutually exclusive outcomes, contradicted reports, termination and replay, and a sweep runs every verified script over stock and order totals. From the site's Next.js app:
npx vitest run src/lib/projects/customer-service