Test Report — The LLM helper in the Python host¶
This pack's view of the live certification of 2026-10-01 (UTC). The full report — all four engine and host
pairs, the three layers, the progressive-rendering evidence and the findings — is the
engine report, which the
Rust engine's docs carry too. The helper is
examples/llm-helper; its README holds the contract.
Verified without a credential¶
- 86 helper tests, no token spent (the whole suite: 187).
tests/test_llm_helper.pyruns a fake of the Python SDK that builds the SDK's own error objects, so the status codes and messages are the real ones. - One contract, shared with the Node.js twin.
tests/vectors/llm-helper-vectors.jsonis byte-identical in both packs (SHA-2561f4823d9…8fa0, pinned by a test in each) and holds 63 cases: 22 request validations, 29 chat outcomes and 12 stream outcomes, each fixing the exact SDK call, the reply or the error, and the segments. A change that drifts one helper away from the other fails a case. - The tests can fail. A helper mutated to gather the token batches and send them at the end fails the sentinel test (the fake model refuses to produce batch k until the caller holds batches 0 to k-1) and the vector for tokens delivered before a mid-stream error; a helper whose default model is changed fails 13 cases. The unmutated helper passes all of them.
- Static checks:
ruff check .clean,basedpyright0 errors on the whole repository,pytest -q187 passed.
Verified live¶
Behind the Java engine and behind the Rust engine, the Python helper answered every scenario: 40 results per pair and no
failed check, across a streaming service (Layer 1), an Event Script flow (Layer 2) and two graphs (Layer 3), on
claude-opus-5-5 and, where a request named it, claude-haiku-4-5.
| Progressive stream | Helper batches | Edge frames | Offset, median / spread | Longest gap |
|---|---|---|---|---|
| Java → Python, Opus 5.5 | 67 | 67 | 12 / 10 ms | 1201 / 1206 ms |
| Java → Python, Haiku 4.5 | 101 | 101 | 7 / 4 ms | 191 / 191 ms |
| Rust → Python, Opus 5.5 | 50 | 50 | 4 / 9 ms | 922 / 922 ms |
| Rust → Python, Haiku 4.5 | 82 | 82 | 3 / 5 ms | 228 / 227 ms |
Every batch the helper forwarded reached the engine's HTTP edge as its own frame, within a few milliseconds and without drift: nothing is gathered or sent once. The long gaps are the API's own pacing, the same at both ends; Haiku 4.5 streams continuously (a batch about every 25 ms) while Opus 5.5 arrives in bursts about every 600 ms. Credential states, measured on the real SDK: no credential gave 503 through Java on all three layers and through Rust on the graph, and an invalid key gave Anthropic's 401 with its request id when this pack's client called the helper directly. The same message came from both helpers, through every layer. All traces that touched this helper rebuild as one connected tree.
What the Python SDK needed¶
- No credential is not an API error.
anthropic1.x raises a bareTypeError("Could not resolve authentication method…") before anything is sent, so the helper maps it by its text to a 503LLM provider credential missing - set ANTHROPIC_API_KEY in the environment. A missing key is also whatllm.healthreports, from the client's own attributes and with no network traffic. - The SDK takes no sampling parameters.
temperature,top_pandtop_kare gone from its signatures, and the current models reject them, so the contract refuses them (paramsoutsideprovider,model,max_tokens,timeout_ms,effortandstop_sequencesis a 400). - A stream's request id is in the response headers. The final message of a stream carries none (
_request_idis empty), so the helper readsstream.response.headers["request-id"]; every reply and terminal event carries it. fallbacks="default"is a typed value on the beta namespace, and the API accepted it on Opus 5.5. The helper sends it only for the models that support it and never on a route that cannot.- The deadline is
asyncio.wait_foraround the whole call, SDK retries included; the abandoned call is cancelled, which a test asserts. A stream has no total deadline:timeout_msis its idle allowance.
Reproduce¶
pip install -e '.[dev]'
pytest tests/test_llm_helper.py # token-free
export ANTHROPIC_API_KEY=... # live: the credential reaches the helper only
mercury-serve examples/llm-helper/llm_helper.py -Dlog.format=compact -Dllm.log.batches=true
The commands for the engines, the deploy folder and the curl calls for each layer are in the
engine report.