Skip to content

Test a Workflow

How do you prove a workflow works? Test its development version with inputs whose answers you know. Inspect the resolved values, then check the outputs and the effects. A development test is a real run. Messages send, records change, and usage is spent.

Use Invoice intake → Check invoice from Your First Workflow. Its example rule is that amounts over 1000 need review:

Invoice IDAmountExpected invoice_idExpected review_needed
INV-104120INV-104false
INV-1051500INV-105true

Decide separately how your flow handles missing IDs and invalid amounts.

In the app: open Workflows → Invoice intake → Dev → Check invoice. The header’s Test button has the tooltip Test current flow. Start with this pure calculation. Before testing a provider action, inspect its connection and choose a test record or audience.

With an assistant: list your workflows:

Terminal window
cai workflow list --limit 100 --json

Note Invoice intake’s data.items[].id as <workflowId> and check data.hasMore. The default limit is 25. Neither the CLI nor MCP offers a next-page cursor. If the workflow is absent while hasMore is true, use app search and copy its ID from /workflows/<workflowId>/versions/.... MCP uses workflow_list, with limit capped at 100.

This page pins the CLI workflow once, and scoped MCP calls require workflowId. Pin it, then list its flows:

Terminal window
cai use <workflowId> --json
cai flow list --json

MCP uses flow_list. Note Check invoice’s data.items[].id as <flowId>.

CLI and MCP node, development-flow, and discovery tests share one cap per workflow, counted over a UTC day. The default is 200, and a configured 0 means unlimited. At the cap no run starts, and the error begins Agent test-run daily cap reached for this workflow. App tests and API-key runs do not count against it. A development REST call with a CLI credential checks the same cap but does not increase its count.

In the app: click Test, enter values, then click Test flow. Those values update the saved development test values and autosave. Test flow saves outstanding edits before dispatch. Later development tests can reuse them. The panel shows the inputs the flow references. Without inputs or required runtime context, Test starts directly.

With an assistant: read the inputs this flow accepts:

Terminal window
cai run inputs <flowId> --flow --json

MCP uses run_inputs with nodeIdOrFlowId set to <flowId> and flow: true.

data.flowInputs lists id, name, displayName, dataType, isList, testValue, and used. An unused input has no effect. For a custom type with schema fields, testValueShape names the expected keys and whether to send an object or a list of objects. Save this as invoice-low.json, replacing these keys with the Invoice ID and Amount input IDs:

{
"<invoiceIdInputId>": "INV-104",
"<amountInputId>": 120
}

Exact IDs avoid alias collisions. The CLI also accepts unambiguous names and bindings. If any input in a development input list lacks both displayName and name, it disables those aliases across that list. Unknown or ambiguous keys reject before dispatch. If you leave a flow input out, the run uses its saved test value. An explicit null, "", false, or 0 overrides it.

Check invoice contains only Return response, which has no app Test action, so use Test current flow. For another eligible node, open its menu and choose Test action. Trigger nodes and Return response are excluded.

For an isolated test, find the changed node:

Terminal window
cai flow get <flowId> --json

MCP uses flow_get. Note the node’s data.nodes[].id as <nodeId>.

Then read the values it accepts:

Terminal window
cai run inputs <nodeId> --json

Prepare node-values.json keyed by data.flowInputs[].id and data.nodeInputs[].id. Every referenced flow input needs an explicit value. If you omit one, the call fails with ISOLATED_FLOW_CONTEXT_REQUIRED before execution. Those flow inputs preserve explicit null and empty values. An ordinary node input given null reuses its saved test value.

Run the node alone, then its containing flow:

Terminal window
cai run node <nodeId> --inputs node-values.json --json
cai run flow <flowId> --inputs invoice-low.json --target dev --json

MCP uses run_node with nodeId and run_flow with flowId, passing parsed inputs and target: "dev" for the flow.

A node test does not execute upstream nodes. runScope.endToEndValidationRequired reports incoming wiring, not whether your values passed. Inspect configValueProcessed, hasError, and resolutionWarnings on every resolvedConfigs entry. An empty [] means no warning was recorded. Then test the containing flow and repeat with INV-105/1500.

A successful assistant development test also attempts output-type adoption, persisting changed metadata to development. Inspect data.outputTypeAdoption for a node, or data.outputTypeAdoptions and its adoptedNodeIds for a flow. Live runs do not adopt types.

cai run discover (run_discover) calls the real action too. It normally deletes its scratch node and keeps the reusable scratch flow. It does not verify another node’s output type.

The CLI waits 120,000 milliseconds for a node by default and 300,000 for a flow. MCP waitSeconds defaults to 60, maximum 120. These bound waiting, not execution. On success, note data.workflowExecutionId as <workflowExecutionId>. On CLI timeout, pause, or failure, the same ID comes back in error.detail.workflowExecutionId. Inspect that execution instead of repeating its effects:

Terminal window
cai exec get <workflowExecutionId> --json

MCP uses execution_get with workflowExecutionId. Its run tools can return a successful tool response containing a non-success execution, so inspect the execution status.

If the request reports a builder_approval event, note its approvalId as <id>. After the person handles the approval, resume the run:

Terminal window
cai approval resume <id> --json

MCP uses approval_resume with approvalId.

In Run history, inspect Inputs, Process → Configurations / Result, and Outputs. Compare both invoice results above. A completed run can still contain unintended null mappings. A loop that collects failed or handled-error iterations produces completed_with_error, not full success. Inspect each iteration.

If you later add invoice storage, identify its collection before testing. Afterwards open Dev → Data tables → invoice, or read the same store:

Terminal window
cai data query --type custom.invoice --env dev --limit 50 --json

MCP uses data_record_query with type: "custom.invoice", env: "dev", and limit: 50. Check data.items and data.totalCount. Inspect provider destinations too. Keep the execution IDs, expected and actual results, and untested paths.

Next: Debug a Run or Publish a Workflow.