Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion agents/build/custom-llm.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -58,16 +58,16 @@
"parameters": {
"type": "object",
"properties": {
"order_number": { "type": "string", "description": "The order number, e.g. A12345." }

Check warning on line 61 in agents/build/custom-llm.mdx

View check run for this annotation

Mintlify / Mintlify Validation (hanabiaiinc) - vale-spellcheck

agents/build/custom-llm.mdx#L61

Did you really mean 'order_number'?
},
"required": ["order_number"]
}
}
}
],
"session_id": "sess_01j9x4k2m8v3q7n5p6r8t9w0y1",

Check warning on line 68 in agents/build/custom-llm.mdx

View check run for this annotation

Mintlify / Mintlify Validation (hanabiaiinc) - vale-spellcheck

agents/build/custom-llm.mdx#L68

Did you really mean 'session_id'?
"user_id": "user-42",

Check warning on line 69 in agents/build/custom-llm.mdx

View check run for this annotation

Mintlify / Mintlify Validation (hanabiaiinc) - vale-spellcheck

agents/build/custom-llm.mdx#L69

Did you really mean 'user_id'?
"fishaudio_extra_body": { "chat_id": "chat-9" }

Check warning on line 70 in agents/build/custom-llm.mdx

View check run for this annotation

Mintlify / Mintlify Validation (hanabiaiinc) - vale-spellcheck

agents/build/custom-llm.mdx#L70

Did you really mean 'fishaudio_extra_body'?
}
```

Expand Down Expand Up @@ -134,6 +134,6 @@

## Limitations

- Single Turn and Tool [agent tests](/agents/test/agent-tests) are not supported. A scripted test run refuses to execute rather than substitute a platform model for yours. Multi Turn tests run on your endpoint.
- Single Turn and Tool [agent tests](/agents/test/agent-tests) are not supported. A scripted test run refuses to execute rather than substitute a platform model for yours. Simulation tests run on your endpoint.
- Configuration is API-only for now. A console UI comes later.
- The prompt-level safety guardrails still ride the assembled context, but your model decides whether to honor them.
2 changes: 1 addition & 1 deletion agents/build/dynamic-variables.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -32,7 +32,7 @@
and they are on the {{plan}} plan. Greet them by name.
```

Pass values as a flat object of strings, numbers, or booleans. Variable names must match `[A-Za-z][A-Za-z0-9_]*` (no hyphens or dots; the dotted `system.*` names are reserved for [system variables](#system-variables)), string values are capped at 1,000 characters, and a request can carry at most 50 variables. Violations reject session creation with `422`:

Check warning on line 35 in agents/build/dynamic-variables.mdx

View check run for this annotation

Mintlify / Mintlify Validation (hanabiaiinc) - vale-spellcheck

agents/build/dynamic-variables.mdx#L35

Did you really mean 'booleans'?

<CodeGroup>
```bash API (curl)
Expand Down Expand Up @@ -94,16 +94,16 @@
```

<Tip>
In `agentId` mode, values arrive from the end user's browser. Treat them as untrusted input, and use `sessionToken` mode when personalization must come from data only your backend knows.

Check warning on line 97 in agents/build/dynamic-variables.mdx

View check run for this annotation

Mintlify / Mintlify Validation (hanabiaiinc) - vale-spellcheck

agents/build/dynamic-variables.mdx#L97

Did you really mean 'untrusted'?
</Tip>

## System variables

The platform fills a handful of `{{system.*}}` placeholders itself, on every session. You cannot supply or override them: the names you pass in `dynamic_variables` cannot contain a dot, so the two namespaces never collide.

Check warning on line 102 in agents/build/dynamic-variables.mdx

View check run for this annotation

Mintlify / Mintlify Validation (hanabiaiinc) - vale-spellcheck

agents/build/dynamic-variables.mdx#L102

Did you really mean 'namespaces'?

| Variable | Value | Notes |
|---|---|---|
| `system.channel` | `phone_inbound`, `phone_outbound`, or `web_voice` | Web SDK, API, preview, and Single Turn and Tool test sessions are `web_voice`. A Multi Turn test uses its [channel](/agents/test/agent-tests#channel). |
| `system.channel` | `phone_inbound`, `phone_outbound`, or `web_voice` | Web SDK, API, preview, and Single Turn and Tool test sessions are `web_voice`. A Simulation test uses its [channel](/agents/test/agent-tests#channel). |
| `system.timezone` | The session's resolved IANA timezone, for example `Asia/Tokyo` | `UTC` when nothing resolves. Always equals the session's `timezone` field; see [which timezone a session uses](/agents/build/time-timezone#which-timezone-a-session-uses). |
| `system.today` | Today's date in that timezone, ISO 8601 `YYYY-MM-DD` | The calendar date at session creation. There is no `system.now`; see [World context](#world-context). |
| `system.language` | The session language code, one of the [52 supported languages](/agents/build/voice-language#speaking-language) |
Expand Down Expand Up @@ -155,12 +155,12 @@
A reference to a system variable that does not exist (`{{system.foo}}`) is rejected with `422` when you save the configuration or the tool, and when you send it in session overrides. The error lists the available names.

<Note>
There is no `system.session_id`. The configuration is rendered before the session receives its id, so the id cannot be templated into it. Read it from the `POST /v1/agent/sessions` [response](/agents/deploy/authenticated-sessions), the SDK's `connect` event, or the `session` object in [webhooks](/agents/monitor/webhooks).

Check warning on line 158 in agents/build/dynamic-variables.mdx

View check run for this annotation

Mintlify / Mintlify Validation (hanabiaiinc) - vale-spellcheck

agents/build/dynamic-variables.mdx#L158

Did you really mean 'templated'?
</Note>

## World context

The agent knows the current date and time without any variable. The platform states the date and the session's timezone at the start of every session and refreshes the time on every turn, so "tomorrow morning" or "next Tuesday" resolve correctly even in a long conversation. You don't need a custom `{{today}}` or `{{now}}`, and there is no `system.now`: variables render once, when the session is created, so a templated time would be stale from the first reply onward. When you want the date or timezone as text in a template, for example in a webhook tool URL, use the [`system.today` and `system.timezone`](#system-variables) system variables.

Check warning on line 163 in agents/build/dynamic-variables.mdx

View check run for this annotation

Mintlify / Mintlify Validation (hanabiaiinc) - vale-spellcheck

agents/build/dynamic-variables.mdx#L163

Did you really mean 'templated'?

See [Time & timezone](/agents/build/time-timezone) for how the timezone is resolved and how to turn the injection off for a session (`world_context: false`).

Expand Down
2 changes: 1 addition & 1 deletion agents/build/webhook-tools.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -110,7 +110,7 @@

## Mock responses

A tool can store mock responses: canned payloads, each with a `name`, a `status_code` (100–599), a `content_type`, and a `body`. Mocks are saved with the tool's configuration for test scenarios. [Multi Turn agent tests](/agents/test/agent-tests#tool-mocks) that mock all tools answer the tool with its first mock response. Live calls always hit the real endpoint.
A tool can store mock responses: canned payloads, each with a `name`, a `status_code` (100–599), a `content_type`, and a `body`. Mocks are saved with the tool's configuration for test scenarios. [Simulation tests](/agents/test/agent-tests#tool-mocks) that mock all tools answer the tool with its first mock response. Live calls always hit the real endpoint.

## Test your tool

Expand Down Expand Up @@ -166,7 +166,7 @@
{ "ticket_id": "T-1042", "expected_reply": "within 24 hours" }
```

Ticket numbers get spoken aloud: short, pronounceable IDs survive text-to-speech far better than UUIDs.

Check warning on line 169 in agents/build/webhook-tools.mdx

View check run for this annotation

Mintlify / Mintlify Validation (hanabiaiinc) - vale-spellcheck

agents/build/webhook-tools.mdx#L169

Did you really mean 'UUIDs'?

<Note>
An in-call ticket depends on the model choosing to escalate. For a safety net
Expand Down
34 changes: 17 additions & 17 deletions agents/test/agent-tests.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -12,14 +12,14 @@
| --------------- | ------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------ |
| **Single Turn** | The next reply to a scripted conversation, scored by an LLM judge against your expectation. | Wording, tone, and one specific answer. |
| **Tool** | Whether the next reply calls a given tool (or no tool at all), optionally with matching parameters. | Routing decisions and argument extraction. |
| **Multi Turn** | A whole conversation driven by a simulated user, scored by the judge against your success conditions, plus deterministic checks on what happened. | Flows that only show up over several exchanges, such as taking an order. |
| **Simulation** | A whole conversation driven by a simulated user, scored by the judge against your success conditions, plus deterministic checks on what happened. | Flows that only show up over several exchanges, such as taking an order. |

<Note>
Tests run as text (no audio is synthesized), but the agent uses its full draft
configuration: the [knowledge base](/agents/build/knowledge-base) is
consulted, and the agent can invoke its attached tools. In Single Turn and
Tool tests, [webhook tools](/agents/build/webhook-tools) send real HTTP
requests, so point them at a staging endpoint. Multi Turn tests mock tools by
requests, so point them at a staging endpoint. Simulation tests mock tools by
default. For end-to-end verification with voice, use [preview
calls](/agents/test/preview-calls).
</Note>
Expand All @@ -35,8 +35,8 @@
</Step>
<Step title="Describe what to check">
Fill in the fields for that type. They are described in the sections below:
[Single Turn](#single-turn-tests), [Tool](#tool-tests), and [Multi
Turn](#multi-turn-tests).
[Single Turn](#single-turn-tests), [Tool](#tool-tests), and
[Simulation](#simulation-tests).
</Step>
<Step title="Save">
Click **Create Test**. The test is now in your library, ready to attach to
Expand All @@ -58,11 +58,11 @@

A Tool test scripts the conversation the same way, but instead of judging the reply it checks which tool the agent called while producing it.

Under **Require tool execution**, pick a tool from the library and choose **Should have been called** or **Should not be called**. Leave the tool empty to check that the agent called no tool at all. Integration tools, such as calendar tools, can only be checked in a Multi Turn test. Under **Tool parameters**, optionally add the parameter values the call must carry, typed as string, number, or boolean. The result names the tool the agent actually called.
Under **Require tool execution**, pick a tool from the library and choose **Should have been called** or **Should not be called**. Leave the tool empty to check that the agent called no tool at all. Integration tools, such as calendar tools, can only be checked in a Simulation test. Under **Tool parameters**, optionally add the parameter values the call must carry, typed as string, number, or boolean. The result names the tool the agent actually called.

## Multi Turn tests
## Simulation tests

A Multi Turn test does not script the user. Instead, a simulated user plays a role you describe, talks to your agent for up to a set number of turns, and an LLM judge scores the finished conversation against your success conditions.
A Simulation test does not script the user. Instead, a simulated user plays a role you describe, talks to your agent for up to a set number of turns, and an LLM judge scores the finished conversation against your success conditions.

### Write the scenario

Expand Down Expand Up @@ -107,7 +107,7 @@

If you are used to tools without a mock calling their real endpoint, your first runs will show **No mock** on those calls. Add a mock for each tool the test form lists, or add read-only tools to **Call the real endpoint**.

Mocks and assertions apply to the tools attached to the agent under test. A tool the agent does not have is never mocked and counts as never called: a required call on it fails with _This tool is not on the agent_, while a forbidden tool or a maximum-only check passes. Client tools on a phone channel are treated the same way. This keeps a shared guard such as "never call issue_refund" passing on agents that cannot call the tool.

Check warning on line 110 in agents/test/agent-tests.mdx

View check run for this annotation

Mintlify / Mintlify Validation (hanabiaiinc) - vale-spellcheck

agents/test/agent-tests.mdx#L110

Did you really mean 'issue_refund'?

### Assertions

Expand Down Expand Up @@ -136,21 +136,21 @@

## Run a single test

Open the test and switch to its **Run** tab. Pick an agent that has access, then click **Run test**. Single Turn and Tool tests answer within seconds. A Multi Turn test runs in the background, shows **Queued** and then **Running**, and can take a few minutes.
Open the test and switch to its **Run** tab. Pick an agent that has access, then click **Run test**. Single Turn and Tool tests answer within seconds. A Simulation test runs in the background, shows **Queued** and then **Running**, and can take a few minutes.

The verdict card shows **Pass** or **Fail** with the time the run took, followed by details for the test type:

| Type | What you see |
| --------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Single Turn** | The full reply the agent generated and why the judge passed or failed it. |
| **Tool** | Which tool the agent called, if any, against what the test expected. |
| **Multi Turn** | The judge's summary, every success condition with its verdict and rationale, every assertion with its outcome, who ended the conversation, after how many turns and on which channel, the LLM cost of the agent, the simulated user, and the judge with their token breakdown, and the full transcript with each tool call marked by where its answer came from: **Mock** for an entry in the test, **Tool mock** for the tool's own mock response, **Real** when it reached the real endpoint, **No mock** or **No matching mock** when it failed for lack of one. |
| **Simulation** | The judge's summary, every success condition with its verdict and rationale, every assertion with its outcome, who ended the conversation, after how many turns and on which channel, the LLM cost of the agent, the simulated user, and the judge with their token breakdown, and the full transcript with each tool call marked by where its answer came from: **Mock** for an entry in the test, **Tool mock** for the tool's own mock response, **Real** when it reached the real endpoint, **No mock** or **No matching mock** when it failed for lack of one. |

A Multi Turn run with an **Unknown** condition fails and carries a **Needs review** badge. If a run can't complete, the card shows **Error** with a message that says whether the agent caused it, for example its custom LLM endpoint stopped responding, or the platform did, for example our language model did not answer. A conversation that broke off or hit its time limit is never judged, so it can't pass by accident. The transcript is still shown below the error, as it is when the conversation finished but could not be judged.
A Simulation run with an **Unknown** condition fails and carries a **Needs review** badge. If a run can't complete, the card shows **Error** with a message that says whether the agent caused it, for example its custom LLM endpoint stopped responding, or the platform did, for example our language model did not answer. A conversation that broke off or hit its time limit is never judged, so it can't pass by accident. The transcript is still shown below the error, as it is when the conversation finished but could not be judged.

## Run every test for an agent

On the agent's **Tests** page in the Builder, click **Run all**. Each row moves through **queued → running → Pass/Fail** (or **Error** if a run can't complete). Multi Turn tests run once per repeat and the row shows how many repeats passed. The page header summarizes the latest batch, for example `4 passed, 1 failed on last run`, with the batch **pass rate**: green at 100%, amber from 80%, red below. The pass rate counts passed runs out of passed and failed ones, so a run that ended in **Error** does not lower it. Use the row menu to re-run a single test, edit it, or remove it from the agent.
On the agent's **Tests** page in the Builder, click **Run all**. Each row moves through **queued → running → Pass/Fail** (or **Error** if a run can't complete). Simulation tests run once per repeat and the row shows how many repeats passed. The page header summarizes the latest batch, for example `4 passed, 1 failed on last run`, with the batch **pass rate**: green at 100%, amber from 80%, red below. The pass rate counts passed runs out of passed and failed ones, so a run that ended in **Error** does not lower it. Use the row menu to re-run a single test, edit it, or remove it from the agent.

## Tests run against the draft

Expand All @@ -174,14 +174,14 @@
| Field | Limit |
| ------------------------------- | --------------------------------------------------------------------- |
| Test name | 200 characters |
| Conversation message | 2,000 characters each, at least one message (optional for Multi Turn) |
| Conversation (Multi Turn) | 50 messages |
| Conversation message | 2,000 characters each, at least one message (optional for Simulation) |
| Conversation (Simulation) | 50 messages |
| Expectation | 400 characters |
| Success / failure example | 400 characters each |
| Scenario (Multi Turn) | 10,000 characters |
| Success conditions (Multi Turn) | 1 to 10, description 500 characters each |
| Max turns (Multi Turn) | 50 |
| Repeat count (Multi Turn) | 20 |
| Scenario (Simulation) | 10,000 characters |
| Success conditions (Simulation) | 1 to 10, description 500 characters each |
| Max turns (Simulation) | 50 |
| Repeat count (Simulation) | 20 |
| Dynamic variables | 50 per test |

## Going further
Expand Down
Loading