From 114a7236c23879912fe906ac18628c221a1382b1 Mon Sep 17 00:00:00 2001 From: Him188 Date: Mon, 28 Sep 2026 20:14:15 +0900 Subject: [PATCH] docs(agents): rename Multi Turn tests to Simulation tests Co-Authored-By: Claude Opus 5.5 --- agents/build/custom-llm.mdx | 2 +- agents/build/dynamic-variables.mdx | 2 +- agents/build/webhook-tools.mdx | 2 +- agents/test/agent-tests.mdx | 34 +++++++++++++++--------------- 4 files changed, 20 insertions(+), 20 deletions(-) diff --git a/agents/build/custom-llm.mdx b/agents/build/custom-llm.mdx index fa1691f..20ebeee 100644 --- a/agents/build/custom-llm.mdx +++ b/agents/build/custom-llm.mdx @@ -134,6 +134,6 @@ Voice conversations are latency sensitive, so aim for a time-to-first-token unde ## Limitations -- Single Turn and Tool [agent tests](/agents/test/agent-tests) are not supported. A scripted test run refuses to execute rather than substitute a platform model for yours. Multi Turn tests run on your endpoint. +- Single Turn and Tool [agent tests](/agents/test/agent-tests) are not supported. A scripted test run refuses to execute rather than substitute a platform model for yours. Simulation tests run on your endpoint. - Configuration is API-only for now. A console UI comes later. - The prompt-level safety guardrails still ride the assembled context, but your model decides whether to honor them. diff --git a/agents/build/dynamic-variables.mdx b/agents/build/dynamic-variables.mdx index a231ea2..a774d1f 100644 --- a/agents/build/dynamic-variables.mdx +++ b/agents/build/dynamic-variables.mdx @@ -103,7 +103,7 @@ The platform fills a handful of `{{system.*}}` placeholders itself, on every ses | Variable | Value | Notes | |---|---|---| -| `system.channel` | `phone_inbound`, `phone_outbound`, or `web_voice` | Web SDK, API, preview, and Single Turn and Tool test sessions are `web_voice`. A Multi Turn test uses its [channel](/agents/test/agent-tests#channel). | +| `system.channel` | `phone_inbound`, `phone_outbound`, or `web_voice` | Web SDK, API, preview, and Single Turn and Tool test sessions are `web_voice`. A Simulation test uses its [channel](/agents/test/agent-tests#channel). | | `system.timezone` | The session's resolved IANA timezone, for example `Asia/Tokyo` | `UTC` when nothing resolves. Always equals the session's `timezone` field; see [which timezone a session uses](/agents/build/time-timezone#which-timezone-a-session-uses). | | `system.today` | Today's date in that timezone, ISO 8601 `YYYY-MM-DD` | The calendar date at session creation. There is no `system.now`; see [World context](#world-context). | | `system.language` | The session language code, one of the [52 supported languages](/agents/build/voice-language#speaking-language) | diff --git a/agents/build/webhook-tools.mdx b/agents/build/webhook-tools.mdx index b536e8c..632b81c 100644 --- a/agents/build/webhook-tools.mdx +++ b/agents/build/webhook-tools.mdx @@ -110,7 +110,7 @@ Use `hide` when error responses might leak internal details you don't want spoke ## Mock responses -A tool can store mock responses: canned payloads, each with a `name`, a `status_code` (100–599), a `content_type`, and a `body`. Mocks are saved with the tool's configuration for test scenarios. [Multi Turn agent tests](/agents/test/agent-tests#tool-mocks) that mock all tools answer the tool with its first mock response. Live calls always hit the real endpoint. +A tool can store mock responses: canned payloads, each with a `name`, a `status_code` (100–599), a `content_type`, and a `body`. Mocks are saved with the tool's configuration for test scenarios. [Simulation tests](/agents/test/agent-tests#tool-mocks) that mock all tools answer the tool with its first mock response. Live calls always hit the real endpoint. ## Test your tool diff --git a/agents/test/agent-tests.mdx b/agents/test/agent-tests.mdx index 26fed0d..145b377 100644 --- a/agents/test/agent-tests.mdx +++ b/agents/test/agent-tests.mdx @@ -12,14 +12,14 @@ Write a test once, run it against any agent, and get a **Pass** or **Fail** verd | --------------- | ------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------ | | **Single Turn** | The next reply to a scripted conversation, scored by an LLM judge against your expectation. | Wording, tone, and one specific answer. | | **Tool** | Whether the next reply calls a given tool (or no tool at all), optionally with matching parameters. | Routing decisions and argument extraction. | -| **Multi Turn** | A whole conversation driven by a simulated user, scored by the judge against your success conditions, plus deterministic checks on what happened. | Flows that only show up over several exchanges, such as taking an order. | +| **Simulation** | A whole conversation driven by a simulated user, scored by the judge against your success conditions, plus deterministic checks on what happened. | Flows that only show up over several exchanges, such as taking an order. | Tests run as text (no audio is synthesized), but the agent uses its full draft configuration: the [knowledge base](/agents/build/knowledge-base) is consulted, and the agent can invoke its attached tools. In Single Turn and Tool tests, [webhook tools](/agents/build/webhook-tools) send real HTTP - requests, so point them at a staging endpoint. Multi Turn tests mock tools by + requests, so point them at a staging endpoint. Simulation tests mock tools by default. For end-to-end verification with voice, use [preview calls](/agents/test/preview-calls). @@ -35,8 +35,8 @@ Tests are workspace-level resources, managed under **Library → Tests** in the Fill in the fields for that type. They are described in the sections below: - [Single Turn](#single-turn-tests), [Tool](#tool-tests), and [Multi - Turn](#multi-turn-tests). + [Single Turn](#single-turn-tests), [Tool](#tool-tests), and + [Simulation](#simulation-tests). Click **Create Test**. The test is now in your library, ready to attach to @@ -58,11 +58,11 @@ Under **Conversation**, click **Add message** to build the history the agent see A Tool test scripts the conversation the same way, but instead of judging the reply it checks which tool the agent called while producing it. -Under **Require tool execution**, pick a tool from the library and choose **Should have been called** or **Should not be called**. Leave the tool empty to check that the agent called no tool at all. Integration tools, such as calendar tools, can only be checked in a Multi Turn test. Under **Tool parameters**, optionally add the parameter values the call must carry, typed as string, number, or boolean. The result names the tool the agent actually called. +Under **Require tool execution**, pick a tool from the library and choose **Should have been called** or **Should not be called**. Leave the tool empty to check that the agent called no tool at all. Integration tools, such as calendar tools, can only be checked in a Simulation test. Under **Tool parameters**, optionally add the parameter values the call must carry, typed as string, number, or boolean. The result names the tool the agent actually called. -## Multi Turn tests +## Simulation tests -A Multi Turn test does not script the user. Instead, a simulated user plays a role you describe, talks to your agent for up to a set number of turns, and an LLM judge scores the finished conversation against your success conditions. +A Simulation test does not script the user. Instead, a simulated user plays a role you describe, talks to your agent for up to a set number of turns, and an LLM judge scores the finished conversation against your success conditions. ### Write the scenario @@ -136,7 +136,7 @@ One test can be attached to many agents, and each agent keeps its own last resul ## Run a single test -Open the test and switch to its **Run** tab. Pick an agent that has access, then click **Run test**. Single Turn and Tool tests answer within seconds. A Multi Turn test runs in the background, shows **Queued** and then **Running**, and can take a few minutes. +Open the test and switch to its **Run** tab. Pick an agent that has access, then click **Run test**. Single Turn and Tool tests answer within seconds. A Simulation test runs in the background, shows **Queued** and then **Running**, and can take a few minutes. The verdict card shows **Pass** or **Fail** with the time the run took, followed by details for the test type: @@ -144,13 +144,13 @@ The verdict card shows **Pass** or **Fail** with the time the run took, followed | --------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **Single Turn** | The full reply the agent generated and why the judge passed or failed it. | | **Tool** | Which tool the agent called, if any, against what the test expected. | -| **Multi Turn** | The judge's summary, every success condition with its verdict and rationale, every assertion with its outcome, who ended the conversation, after how many turns and on which channel, the LLM cost of the agent, the simulated user, and the judge with their token breakdown, and the full transcript with each tool call marked by where its answer came from: **Mock** for an entry in the test, **Tool mock** for the tool's own mock response, **Real** when it reached the real endpoint, **No mock** or **No matching mock** when it failed for lack of one. | +| **Simulation** | The judge's summary, every success condition with its verdict and rationale, every assertion with its outcome, who ended the conversation, after how many turns and on which channel, the LLM cost of the agent, the simulated user, and the judge with their token breakdown, and the full transcript with each tool call marked by where its answer came from: **Mock** for an entry in the test, **Tool mock** for the tool's own mock response, **Real** when it reached the real endpoint, **No mock** or **No matching mock** when it failed for lack of one. | -A Multi Turn run with an **Unknown** condition fails and carries a **Needs review** badge. If a run can't complete, the card shows **Error** with a message that says whether the agent caused it, for example its custom LLM endpoint stopped responding, or the platform did, for example our language model did not answer. A conversation that broke off or hit its time limit is never judged, so it can't pass by accident. The transcript is still shown below the error, as it is when the conversation finished but could not be judged. +A Simulation run with an **Unknown** condition fails and carries a **Needs review** badge. If a run can't complete, the card shows **Error** with a message that says whether the agent caused it, for example its custom LLM endpoint stopped responding, or the platform did, for example our language model did not answer. A conversation that broke off or hit its time limit is never judged, so it can't pass by accident. The transcript is still shown below the error, as it is when the conversation finished but could not be judged. ## Run every test for an agent -On the agent's **Tests** page in the Builder, click **Run all**. Each row moves through **queued → running → Pass/Fail** (or **Error** if a run can't complete). Multi Turn tests run once per repeat and the row shows how many repeats passed. The page header summarizes the latest batch, for example `4 passed, 1 failed on last run`, with the batch **pass rate**: green at 100%, amber from 80%, red below. The pass rate counts passed runs out of passed and failed ones, so a run that ended in **Error** does not lower it. Use the row menu to re-run a single test, edit it, or remove it from the agent. +On the agent's **Tests** page in the Builder, click **Run all**. Each row moves through **queued → running → Pass/Fail** (or **Error** if a run can't complete). Simulation tests run once per repeat and the row shows how many repeats passed. The page header summarizes the latest batch, for example `4 passed, 1 failed on last run`, with the batch **pass rate**: green at 100%, amber from 80%, red below. The pass rate counts passed runs out of passed and failed ones, so a run that ended in **Error** does not lower it. Use the row menu to re-run a single test, edit it, or remove it from the agent. ## Tests run against the draft @@ -174,14 +174,14 @@ Tests always exercise the agent's latest **draft** configuration, including unpu | Field | Limit | | ------------------------------- | --------------------------------------------------------------------- | | Test name | 200 characters | -| Conversation message | 2,000 characters each, at least one message (optional for Multi Turn) | -| Conversation (Multi Turn) | 50 messages | +| Conversation message | 2,000 characters each, at least one message (optional for Simulation) | +| Conversation (Simulation) | 50 messages | | Expectation | 400 characters | | Success / failure example | 400 characters each | -| Scenario (Multi Turn) | 10,000 characters | -| Success conditions (Multi Turn) | 1 to 10, description 500 characters each | -| Max turns (Multi Turn) | 50 | -| Repeat count (Multi Turn) | 20 | +| Scenario (Simulation) | 10,000 characters | +| Success conditions (Simulation) | 1 to 10, description 500 characters each | +| Max turns (Simulation) | 50 | +| Repeat count (Simulation) | 20 | | Dynamic variables | 50 per test | ## Going further